Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every JEV judgment now uses one closed Claim with choice, score or noul, a natural-language context and ordered options. Claims own reusable semantics and provenance; evidence, executable Reflexes, parameter schemas, native effects and qualification retain separate lifetimes. Removed redundant Question/Request/criteria application layers, consumed markers, compilation indexes and unused generation paths. Updated protobuf, Guardrail, Web/replay and fixtures. Historical libraries with no format or claim/1 are archived byte-for-byte once, then replaced atomically with an empty claim/2 library so auto and off modes can start. Previous executable proofs are not reused; new tasks learn and qualify current definitions. Unknown versions and malformed current libraries still fail without rewriting.
Removed internal/jevwire and its public Question/Request/Answer/Response DTOs. Vendor JSON conversion now lives entirely inside the JEV Provider; extensions use only Claim/Evaluation. External API fixtures remain local to tests, and the runtime dependency rule is preserved. Regenerated Claim/JEV bindings with the pinned toolchain and included core/decision in the generated-output CI check.
From an empty library, host-recorded interactions can generate Claims and compile ordinary Reflexes. The native jev command supports Claim publication and compilation/replacement through the same executor, contracts, effect journal, recorded replay and independent semantic review. No special bootstrap Reflex or qualification exemption exists; failed replacement preserves the active source and retirement preserves Claims.
Real browser runs exposed compilation without artifacts and failed warm takeover. Compilation now preserves the complete native trajectory independently of the bounded judgment projection, inherits transient Provider retries, discards truncated output before tools execute and increases output space up to 65,536 tokens. Validation sends oversized semantic reviews and unusable entry parameters back for repair. Generated applicability and runtime schema survive decoding. Parameter extraction uses verbatim constraints and the schema, retries truncated output, supplies a fresh allocation name for newly created resources and accounts for every attempt. Runtime review reads original constraint strings and only the proposed operation's protocol. A parameter failure with no effect releases ordinary execution and avoids selecting the same failed function until input changes; reserved or unknown effects remain protected.
Total compilation defaults to unlimited, with cancellation and a 30-minute request fallback. OpenAI/Anthropic JSON and streaming paths honor the host-owned timeout. Timeouts protect stalled work; they no longer cap normal compilation at 75 seconds.
Recorded workflows now render from actual typed Claim/Evaluation events. Each Reflex takeover contains its successive judgments and tool invocations in one loop, with distinct LLM/JEV/TOOL roles and inline evidence. Removed the fixed four-stage view, swimlane toggle and duplicated detail inspector. Tool arguments, receipts, media, Markdown/code and scan results reuse the timeline components. Recorded highlights loop automatically, card selection seeks and holds its own breakpoint, focus exit resumes replay, and manual scrolling disables live following until explicitly restored.
Reasoning now stays visible inside the graph with its own bounded scroll region. Goal/retry/approval replay regressions target the unified graph and preserved content rather than removed disclosures. Batched remote REPL direction keys are translated together, preserving the draft and cursor; terminal node selection uses the exact node name so a similarly named shell session cannot steal the click.
Validation:
Latest CRM and repeated-shopping reload acceptance each passed three reuse tasks, including both formal warm pairs per scenario (six full takeovers, four formal warm tasks). CRM tasks dispatched six native calls; shopping tasks dispatched eight, with no main-model tools, Claim generation or compilation and unchanged source. The two formal warm tasks reduced foreground LLM tokens from 105,639 to 32,222 (69.5%); adding 130,856 JEV tokens makes the total 163,078, above the 105,639 baseline. Shopping formal warm tasks used 19,320 foreground plus 177,467 JEV tokens versus 91,858 baseline tokens. Net token savings and monetary savings are not established.
Browser arms equally use single native commands and structured snapshots; quoted inputs explicitly request one JSON decode. This narrows the earlier unrestricted tool-use scope. Historical failure rows and interrupted obsolete runs remain retained, with unfinished/failed request usage unknown. The expense run was interrupted after repeated upstream 500/503/auth_unavailable and reasoning exhaustion, with no qualified artifact; its incomplete usage remains unknown. Successful generated-library reload is separate from a complete latest-runtime cold/warm matrix or production SaaS performance.