Skip to content

refactor(jev): unify typed Claims and ordinary Reflex lifecycle - #170

Open
M09Ic wants to merge 12 commits into
masterfrom
refactor/typed-claims-20261006
Open

M09Ic wants to merge 12 commits into
masterfrom
refactor/typed-claims-20261006

Conversation

@M09Ic

@M09Ic M09Ic commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Every JEV judgment now uses one closed Claim with choice, score or noul, a natural-language context and ordered options. Claims own reusable semantics and provenance; evidence, executable Reflexes, parameter schemas, native effects and qualification retain separate lifetimes. Removed redundant Question/Request/criteria application layers, consumed markers, compilation indexes and unused generation paths. Updated protobuf, Guardrail, Web/replay and fixtures. Historical libraries with no format or claim/1 are archived byte-for-byte once, then replaced atomically with an empty claim/2 library so auto and off modes can start. Previous executable proofs are not reused; new tasks learn and qualify current definitions. Unknown versions and malformed current libraries still fail without rewriting.

Removed internal/jevwire and its public Question/Request/Answer/Response DTOs. Vendor JSON conversion now lives entirely inside the JEV Provider; extensions use only Claim/Evaluation. External API fixtures remain local to tests, and the runtime dependency rule is preserved. Regenerated Claim/JEV bindings with the pinned toolchain and included core/decision in the generated-output CI check.

From an empty library, host-recorded interactions can generate Claims and compile ordinary Reflexes. The native jev command supports Claim publication and compilation/replacement through the same executor, contracts, effect journal, recorded replay and independent semantic review. No special bootstrap Reflex or qualification exemption exists; failed replacement preserves the active source and retirement preserves Claims.

Real browser runs exposed compilation without artifacts and failed warm takeover. Compilation now preserves the complete native trajectory independently of the bounded judgment projection, inherits transient Provider retries, discards truncated output before tools execute and increases output space up to 65,536 tokens. Validation sends oversized semantic reviews and unusable entry parameters back for repair. Generated applicability and runtime schema survive decoding. Parameter extraction uses verbatim constraints and the schema, retries truncated output, supplies a fresh allocation name for newly created resources and accounts for every attempt. Runtime review reads original constraint strings and only the proposed operation's protocol. A parameter failure with no effect releases ordinary execution and avoids selecting the same failed function until input changes; reserved or unknown effects remain protected.

Total compilation defaults to unlimited, with cancellation and a 30-minute request fallback. OpenAI/Anthropic JSON and streaming paths honor the host-owned timeout. Timeouts protect stalled work; they no longer cap normal compilation at 75 seconds.

Recorded workflows now render from actual typed Claim/Evaluation events. Each Reflex takeover contains its successive judgments and tool invocations in one loop, with distinct LLM/JEV/TOOL roles and inline evidence. Removed the fixed four-stage view, swimlane toggle and duplicated detail inspector. Tool arguments, receipts, media, Markdown/code and scan results reuse the timeline components. Recorded highlights loop automatically, card selection seeks and holds its own breakpoint, focus exit resumes replay, and manual scrolling disables live following until explicitly restored.

Reasoning now stays visible inside the graph with its own bounded scroll region. Goal/retry/approval replay regressions target the unified graph and preserved content rather than removed disclosures. Batched remote REPL direction keys are translated together, preserving the draft and cursor; terminal node selection uses the exact node name so a similarly named shell session cannot steal the click.

Validation:

  • UI integration: all 84 JEV tests passed across the full run and corrected fixture/assertion rerun; desktop/mobile layout, continuous Claims, parallel calls, exact breakpoints, focus resume and manual-scroll regressions are covered. The retained Playwright trace renders all eight judgments and five native invocations.
  • Merge-blocker fixes: the CI architecture manifest and full changed-package build-tag manifest passed again. Focused Provider retry/batch-limit and library upgrade/publication race tests passed; full-repository golangci-lint reports zero issues, tagged go vet passes, and go mod tidy produces no drift. Pinned protobuf generation completed; generated differences were limited to correcting two tool-version headers.
  • CI follow-up regressions: all 18 targeted streaming/Goal, recap/approval, terminal and operator-journey browser tests passed across the full run and focused final rerun; the frontend and embedded CLI fixture built. The new actual-editor regression fails on the prior handler and passes with batched key handling; console race tests and lint pass. The actual-editor batch regression also passes on Linux using the production readline path after readiness.
  • Shared cyber-ui validation: 26 viewer tests and three traffic tests passed; the viewer package, integrated frontend and cyber-web executable built successfully.
  • Full manifest passed: go test -p 2 -tags 'emptytemplates forceposix full netgo noembed osusergo sqlite' ./agent/provider ./agent/provider/jev ./exts/jev ./exts/guardrail ./pkg/web/service ./cmd/aiscan -count=1 -timeout 180s.
  • Focused compiler/parameter/effect/cancellation race checks passed; the actual aiscan CLI builds successfully.
  • Real JEV choice/score/noul and gateway model protocol checks passed. The earlier fully real short pipeline generated, qualified, published, reloaded and reused an ordinary Reflex on three new tasks.
  • Real Chromium/production browser tasks generated both CRM and shopping Reflexes from empty libraries. Latest runtime reloads the actual generated libraries through production loading, freezes learning and checks new URLs, values and control addresses against private server oracles. Final measured reload outcomes and retained failures are documented in docs/jev.md.
  • Source/library SHA-256, raw requests, per-task CSV, usage classes and independent business oracles are retained outside the repository. Credentials and generated artifacts are excluded; scans cover tracked files, branch history and reports.

Latest CRM and repeated-shopping reload acceptance each passed three reuse tasks, including both formal warm pairs per scenario (six full takeovers, four formal warm tasks). CRM tasks dispatched six native calls; shopping tasks dispatched eight, with no main-model tools, Claim generation or compilation and unchanged source. The two formal warm tasks reduced foreground LLM tokens from 105,639 to 32,222 (69.5%); adding 130,856 JEV tokens makes the total 163,078, above the 105,639 baseline. Shopping formal warm tasks used 19,320 foreground plus 177,467 JEV tokens versus 91,858 baseline tokens. Net token savings and monetary savings are not established.

Browser arms equally use single native commands and structured snapshots; quoted inputs explicitly request one JSON decode. This narrows the earlier unrestricted tool-use scope. Historical failure rows and interrupted obsolete runs remain retained, with unfinished/failed request usage unknown. The expense run was interrupted after repeated upstream 500/503/auth_unavailable and reasoning exhaustion, with no qualified artifact; its incomplete usage remains unknown. Successful generated-library reload is separate from a complete latest-runtime cold/warm matrix or production SaaS performance.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant