From 4936f8b85aada5ff409ddaa14390c6e483e3abc9 Mon Sep 17 00:00:00 2001 From: Drew Stone Date: Tue, 15 Sep 2026 21:39:10 -0700 Subject: [PATCH 1/2] fix(eval): preserve scorer revisions and campaign cost context --- .claude/skills/eval-architect/SKILL.md | 65 ++++++------- .claude/skills/improve-conductor/SKILL.md | 87 +++++++++-------- .../skills/measurement-validation/SKILL.md | 74 ++++++++------- CHANGELOG.md | 8 ++ README.md | 7 ++ create-agent-app/package.json | 2 +- create-agent-app/template-chat/_package.json | 10 +- create-agent-app/template/_package.json | 14 +-- docs/api/eval-campaign.md | 4 +- docs/codemap.json | 4 +- docs/llms-full.txt | 4 +- package.json | 22 ++--- pnpm-lock.yaml | 93 +++++++++++-------- pnpm-workspace.yaml | 1 - src/eval-campaign/index.test.ts | 66 ++++++++++++- src/eval-campaign/index.ts | 18 ++-- src/eval-campaign/trust-gate.test.ts | 4 +- src/eval-campaign/trust-gate.ts | 47 +++------- src/peer-floors/check.test.ts | 13 +-- 19 files changed, 315 insertions(+), 228 deletions(-) diff --git a/.claude/skills/eval-architect/SKILL.md b/.claude/skills/eval-architect/SKILL.md index 0f4a5644..242292b5 100644 --- a/.claude/skills/eval-architect/SKILL.md +++ b/.claude/skills/eval-architect/SKILL.md @@ -1,44 +1,47 @@ --- name: eval-architect -description: Build a measurement that scores an agent's REAL deliverable — not a proxy — for a product you've never seen before. Use when scaffolding or repairing the eval an Improve loop optimizes against. Get this wrong and every downstream optimization perfects a fiction. +description: Build or repair evaluations that score the agent's actual deliverable through the production path, with independent controls and visible missing evidence. --- -# Eval Architect — measure the real deliverable +# Eval architect -You are building the measurement an improvement loop will optimize against. **The loop optimizes whatever you measure.** If you measure the wrong thing, the loop perfects the wrong thing — confidently, expensively, and invisibly. The measurement is the product. Everything else in the Improve stack is downstream of getting this right. +Build the measurement around the product's required outcome and actual execution path. +Reuse maintained Eval contracts and existing product checks before creating another scorer. -This skill is held by the agent that *builds* the eval (often a delegated coding agent). Pair it with `measurement-validation` (the gate that proves your eval is sound before anyone spends money on it). +## Locate the deliverable -## The cardinal question +1. Inspect real runs to find the output channel: replies, validated tool calls, persisted artifacts, application state, or rendered UI. + Score the channel that carries the required outcome. +2. Define the completion boundary for the task. + For work that accumulates across turns, evaluate the completed artifact and retain intermediate evidence needed to explain failures. +3. Trace every consumer of the score, including completion checks, optimization selection, and release decisions. + When the output channel changes, update every affected consumer. -**Where does this agent's deliverable actually land?** Prose in the reply? Validated tool calls? Persisted artifacts (vault docs, DB rows)? A PR? A rendered UI? Find out by inspecting *real runs* — never by assuming it's the chat text. +## Build the checks -> Worked failure (legal-agent, this is why the skill exists): the eval scored the assistant's chat prose. A tool-migration moved the deliverable into `submit_proposal` calls + vault docs, leaving the prose empty. Every scorer reading prose silently collapsed to ~0. The loop would have optimized an empty string. The deliverable had *moved* and the measurement didn't follow it. +1. Map each requirement to observable evidence and an explicit failure condition. + Use answer keys when available; otherwise use independently justified constraints, executable checks, or calibrated judgment. + Keep unsupported requirements and missing evidence visible. +2. Establish a simple baseline through the same entrypoint as the candidate. + Investigate surprising scores instead of assuming either the scorer or the agent caused them. +3. Separate training, candidate selection, and final comparison evidence where the improvement claim requires those partitions. + Preserve scenario identities and shared source-unit mappings across baseline and candidate runs. +4. Define critical failure checks separately from aggregate quality. + A favorable composite must not erase a failure that violates the product's requirements. +5. Preserve scorer identity, actual cost receipts, execution failures, and diagnostic artifacts. + Change `judgeVersion` when an ensemble scorer's configuration changes. -## Invariant (non-negotiable — violate these and the loop is a slot machine) +## Prove the measurement -1. **Score the produced artifact, not the conversation.** Locate the real output channel and score *that*. -2. **For accumulating-artifact agents, score the CONVERGED multi-shot artifact, not turn 1.** Most real agents build their deliverable over several turns. Define a convergence criterion (e.g. the artifact stops growing for N shots) and score the converged state. -3. **A held-out split exists and is never trained on.** No held-out → no honest gate → no trustworthy lift. -4. **Every requirement has gold the scorer matches against, from real records — never fabricated.** A requirement with no gold means there is nothing to verify; fail loud, do not pass-by-default. A fluent hallucination that produced nothing must score 0, not 0.9. +Run known positive and negative examples through the complete scoring path. +Perturb a real deliverable so required behavior improves or regresses, then check that the score detects each change. +Confirm that absent output and evaluator failure remain distinguishable from measured low quality. +Report case coverage, detectable failures, uncertainty, and any requirements the evaluation cannot assess. +A training gain without a final-comparison gain needs diagnosis; it does not identify the cause by itself. -## Judgment (figure this out per product — the agentic core) +## Then consider -- What *is* the deliverable here, and where does it persist? Read the runtime events / tool calls / storage, not the transcript. -- What is the convergence criterion for this agent's artifact? When has it stopped accumulating? -- What gold defines "correct" for each requirement, and where does it come from (real records, never invented figures)? -- Which dimensions matter, and what are their weights? What is the one dimension that, if it regresses, kills the deal regardless of the composite (safety, hallucination, the regulated invariant)? - -## Self-test (prove the metric works before trusting it) - -- **Baseline sanity:** run it. Is the score non-zero and plausible for a competent agent? A near-zero baseline usually means you're scoring the wrong channel, not that the agent is terrible. -- **The mutation test (the one that catches the empty-string bug):** hand-edit the produced artifact to be *obviously better* and *obviously worse*. Does the score move in the right direction and magnitude? A metric that doesn't move under obvious changes is measuring the wrong thing. -- **Audit EVERY scoring surface together.** Completion, quality, and the optimizer's own scorer all read *something*. When the deliverable's channel moves, all of them that read the old channel silently zero. (Session: completion + quality were fixed; the optimizer's own scorer was missed and only found by tracing. Three surfaces — enumerate them, don't assume one.) - -## Evolves-by - -When a later optimization shows lift on *training* but none on *held-out*, your eval was overfittable or gameable — add the gap it missed as a new judgment rule. The architect's judgment surface is itself optimized by the meta-eval *"did evals built this way yield real held-out lift, no critical regression?"* See `skill-evolution`. - -## Fleet as dogfood - -legal / tax / gtm / creative / insurance each put their deliverable in a *different* channel — filings, forms, published copy, rendered artifacts, routed proposals. The skill is general precisely because it forces you to *locate* the channel for the product in front of you rather than hardcode "the reply text." +| Condition | Skill | +|---|---| +| The evaluation path executes and needs calibration or comparison checks | `measurement-validation` with the baseline and control results | +| The validated measurement supports a candidate search | `surface-evolution` with the target surface, evidence, and resource limits | diff --git a/.claude/skills/improve-conductor/SKILL.md b/.claude/skills/improve-conductor/SKILL.md index 61f7ea48..3c67371b 100644 --- a/.claude/skills/improve-conductor/SKILL.md +++ b/.claude/skills/improve-conductor/SKILL.md @@ -1,45 +1,50 @@ --- name: improve-conductor -description: The user-facing controller for the Improve button. Decide whether a request is improvable, translate a dollar budget into a run, read the verdict honestly, and promote or refuse with a reason. Never promise a lift you cannot measure. +description: Drive a requested optimization from its target and resource limits through measured candidate selection, evidence review, and the product's promotion policy. --- -# Improve Conductor — own the user's trust - -You are the agent the user talks to when they click **Improve**. You do not build the eval or run the loop yourself — you decide *whether* to, *how much* to spend, and *what to tell the user about the result*. The product you are protecting is **trust**, not lift. You would rather say "I could not prove an improvement — here is what another $X buys" than ship noise and call it a win. - -You delegate the building to an agent holding `eval-architect` + `surface-evolution`; you both share `measurement-validation` as the honesty contract. - -## Invariant (non-negotiable) - -1. **Never promise or report a lift you cannot measure with valid paired evidence.** Surface the honest verdict: `ship` / `hold` / `need-more-data` / `invalid`. A paired, significant lift is still not shippable until a footprint-matched placebo shows the gain comes from the CONTENT, not from added prompt/mount footprint — the substrate's `neutralizationGate` / `runImprovementLoop({ neutralize })` (`@tangle-network/agent-eval/campaign`). A lift that a neutralized twin reproduces is footprint, not improvement — refuse it. Route this check through `measurement-validation`. "Invalid" (incomplete or unpaired evidence) is a first-class outcome — say it plainly, never paper over it with a survivor-mean number. -2. **Refuse below the data threshold, and say why** — "I have N real outcomes; I won't optimize below M. Here's how to get to M." A refusal with a reason builds more trust than a fabricated win. -3. **Route correctly.** Improvable by surface-tuning → dispatch `surface-evolution`. Needs a new capability or architecture → escalate and say so; don't pretend tuning will fix a structural gap. -4. **No optimization spend before the target is confirmed and the measurement is real.** If there is no improvement infrastructure yet, you do NOT improvise a metric and start spending — you dispatch `eval-bootstrap` to BUILD a validated, externally-grounded harness first. The gate between "build the apparatus" and "spend optimizing" is yours to hold. - -## Cold start — no infrastructure yet - -The most dangerous request is "improve this" for a product with no eval. The wrong move is to invent a metric and start a loop — you'll perfect a proxy and report a fake win. The right move is a strict two-step you orchestrate: - -1. **Frame + build (no spend):** confirm with the user *what "better" means* — the thing they'd reject a draft over, tied to a product-value claim — then dispatch `eval-bootstrap` (often a delegated agent-runtime build loop) to construct a harness grounded in **external truth**, exiting only when `measurement-validation` passes. The improver is a *builder* here, not a tuner. -2. **Then optimize (spend):** only once the harness is validated, dispatch `surface-evolution` against it. - -Never let the user believe step 2 happened when only a toy of step 1 did. If you can't yet build a real measurement (no grounding, target unclear), say so and ask for what you need — that's the honest move, not a loop against an invented number. - -## Judgment (figure this out per request) - -- Is this a surface-tuning problem or an architectural one? (If the agent literally cannot do the task, no prompt edit fixes it.) -- Translate the user's dollars into a run: more spend = wider candidate search + more reps = tighter CI + higher chance of clearing the gate. $0.20 ≈ one quick generation on a couple scenarios; $50 ≈ multi-generation search with a held-out gate that can actually reach significance. -- When to stop: threshold met, plateaued, or budget exhausted — and report which. - -## Self-test - -- **Before spending,** you can state out loud: the metric, its variance, the threshold, the held-out set, and what this budget buys. If you can't, you're not ready to charge for the click. -- **After,** you report the gated lift with its CI and the decision's *reason*. If the run came back `invalid` (a cell errored, evidence unpaired), you tell the user that and offer the re-run — you do not quote the broken number. - -## Evolves-by - -User accept/reject of promotions; spend→lift efficiency; the rate of `invalid` runs. A rising invalid rate is a signal the measurement or the infra needs hardening — route it back to `measurement-validation` / `eval-architect`, don't absorb it silently. See `skill-evolution`. - -## Why this is calibrated, not timid - -A naive Improve button maximizes the displayed number and tells the user "improved +47%". The disciplined one, faced with the same +47, checks the evidence, finds it unpaired, and says "I found a promising candidate but can't yet prove it beats baseline — $X more will confirm it." The second one is the one people pay for twice. +# Improve conductor + +Own the requested outcome, actual spend, and the decision supported by the evidence. +Distinguish building an evaluation, searching for candidates, and proving an improvement. + +## Establish the work + +1. Recover the target, required outcome, existing evaluation, and authorization from the user's request and product context. + Ask only for consequential information that the available evidence cannot resolve. +2. Inspect the current baseline and the system's failure cases. + Choose a surface change, capability change, or architecture experiment according to the observed limitation. +3. Check that the evaluation can detect the required behavior through the production entrypoint. + If it cannot, build that measurement before making improvement claims. + Record measurement work as measurement work, including its actual cost. +4. Set resource limits and a stopping rule for the selected experiment. + Estimate cost from the actual execution path and retain uncertainty about additional calls, retries, and candidate evaluations. + More spend does not guarantee a useful candidate or a conclusive result. + +## Run and decide + +1. Use the maintained optimization method and execution path already available to the product. + Preserve training, selection, and final-comparison boundaries along with the registered observation units. +2. Retain candidate artifacts, scorer identity, paired raw evidence, failures, and complete attempt costs. + Missing usage remains an explicit accounting gap. +3. Read the shared deciding statistic and the producer's actual verdict. + Distinguish missing evidence, a measured failure, an inconclusive comparison, and a result that meets the release policy. +4. Investigate surprising gains and null results with controls that isolate the suspected mechanism. + Use a footprint control when the claim concerns content versus added context; it is not a universal release prerequisite. +5. Promote only through the product's authorized decision path after its required checks pass. + A promising exploratory result can justify another scoped experiment without establishing an improvement. + +## Report + +State what changed, the baseline comparison, deciding interval, independent-unit count, actual costs, and the verdict's reasons. +Explain whether the run stopped because it reached its registered criterion, exhausted its budget, or could not capture valid evidence. +A proposed follow-up must state which uncertainty it could resolve; extra budget alone does not promise confirmation. + +## Then consider + +| Condition | Skill | +|---|---| +| The product has no usable evaluation path | `eval-bootstrap` with the observed deliverable and missing checks | +| A scorer or output-channel defect prevents assessment | `eval-architect` with the failed case and execution evidence | +| Existing measurements need calibration or comparison review | `measurement-validation` with the baseline and retained results | +| The supported next experiment changes an existing surface | `surface-evolution` with the target, acceptance criteria, and resource limits | diff --git a/.claude/skills/measurement-validation/SKILL.md b/.claude/skills/measurement-validation/SKILL.md index 7744db48..a0497bef 100644 --- a/.claude/skills/measurement-validation/SKILL.md +++ b/.claude/skills/measurement-validation/SKILL.md @@ -1,38 +1,44 @@ --- name: measurement-validation -description: Prove a measurement is sound BEFORE spending money optimizing against it. The gate that decides whether an Improve run is allowed to start, and whether its result is allowed to be believed. Refuse metrics whose noise exceeds the effect, that have no held-out split, or whose evidence is incomplete. +description: Check whether an evaluation supports an optimization or release decision through calibrated scoring, independent comparison units, and complete paired evidence. --- -# Measurement Validation — earn the right to optimize - -Optimization is only as trustworthy as the measurement under it. This skill is the gate on both ends: **before** a paid run (is this metric allowed to be optimized?) and **after** (is this result allowed to be believed?). It is the difference between an Improve button that is a product and one that is a slot machine. - -Held by both the orchestrator (`improve-conductor`) and the builder (`eval-architect`). It is the shared honesty contract. - -## Invariant (non-negotiable) - -1. **Refuse to optimize if CV(metric) > the target delta.** If the run-to-run noise is bigger than the effect you're paying to move, the metric *cannot* validate the change — raise reps or fix the metric first. Do not tune against noise. -2. **Refuse to report a lift over INCOMPLETE or UNPAIRED evidence.** Every held-out scenario must have a non-errored cell on *both* the baseline and the candidate side. Below the paired-n floor (≥3), the run is **invalid**, not a verdict. A lift computed over survivors is worse than no number. - > **Enforced by** `trustVerdicts` from `@tangle-network/agent-app/eval-campaign` (rater-trust dimension: IRR floor + per-item rater spread, within-item never pooled, + survivor floor, each failure named in `trustReasons`), composed with the agent-eval statistical gates from `@tangle-network/agent-eval/campaign`: `powerPreflight` (before-gate for Invariant 1 — refuse to greenlight spend when the metric is underpowered vs the target delta), `heldOutGate` / `heldoutSignificance` (paired-bootstrap CI over the held-out split — the after-gate that would have caught the +47 by refusing an unpaired lift), and `neutralizationGate` + `neutralizeText` (footprint-matched placebo gate — a held-out lift that does not survive `neutralizeText` came from added prompt/mount FOOTPRINT, not content, and must not be believed). -3. **Every metric ties to a product-value claim** — "if this number moves, *this* user-visible outcome moves with it." No claim → it's a proxy → don't optimize it. -4. **Below the data threshold of real outcomes, refuse to optimize** — state N and say why. You cannot improve what you have not yet observed enough of. - -> Worked failures (this is why the skill exists): -> - **Noise read as signal:** ~6 optimization rounds were burned chasing ±0.15 run-to-run swings as if they were real. The metric's variance was 3× any prompt delta — every conclusion was unprovable. The bug was the *measurement*, not the model. -> - **A lift that was a lie:** a GEPA run reported `heldOutLift = +47`. Reading the actual cells: 2 of 4 held-out cells had errored, so "baseline" was *delaware alone* (42) and "winner" was *saas alone* (89) — two different personas. The +47 was differencing unlike cells. The gate correctly held (0 valid pairs), but the headline number a naive promoter would have shipped was fiction. - -## Judgment (figure this out per metric) - -- How many reps establish variance for *this* metric? (Noisy targets need 5+, converged-artifact metrics fewer.) -- Is an observed "noisy" result model variance, or a measurement smell? **Default: suspect the metric** until its CI is shown tighter than the effect. -- Where might this metric diverge from real value (the Goodhart risk specific to this product)? - -## Self-test - -- Report **mean ± 95% CI over K converged rollouts.** Show CI < target delta *before* greenlighting spend. If you can't, you haven't earned the right to optimize yet. -- Confirm the held-out split is disjoint from training and large enough that the paired-n floor survives an errored cell. -- **Verify against ground truth, never the summary.** Read the actual cells / artifacts, not the provenance headline. (The +47 above was sitting right there in the summary; only the cells revealed it was unpaired. A green build-hook is not a successful build; a typechecking harness is not a running one; a reported lift is not a measured lift.) - -## Evolves-by - -Track promotions that passed validation but regressed in production → that's a missed variance source or an unguarded dimension; strengthen the preflight. The validation bar itself is a surface that tightens from its own misses. See `skill-evolution`. +# Measurement validation + +Check the measurement through the product's actual evaluation path before interpreting an optimization result. +Separate exploratory evidence from evidence that satisfies a release policy. + +## Establish the measurement + +1. State the user outcome, target population, independent comparison unit, and smallest useful effect. + Record the scoring revision, selection procedure, resource limits, and stopping rule before comparing candidates. +2. Run a simple baseline and independent positive and negative controls through the scorer. + Use `auditEvaluator` from `@tangle-network/agent-eval/meta-eval` when the scorer needs an accuracy audit. + Report false acceptances, false rejections, and missing observations against the product's requirements. +3. Check repeatability and choose enough independent units to resolve the intended effect. + Use `powerPreflight` to guide sampling and budget decisions. + An underpowered result can guide exploration; it cannot certify an improvement. +4. Keep final comparison cases separate from training and candidate selection. + Register shared source units when several scenarios come from the same source. + Additional repetitions do not create new independent units. + +## Assess the result + +1. Inspect raw baseline and candidate cells, judge failures, pairing, and the registered unit mapping. + Missing or asymmetric evidence must remain visible and must block a promotion claim. +2. Use Eval's `heldOutGate` or `heldoutSignificance` for the shared promotion decision. + Report its deciding interval, independent-unit count, eligibility, and vetoes. + The bootstrap diagnostic may differ from the deciding interval for binary outcomes. +3. Use App's `trustVerdicts` to check rater agreement and surviving-judge coverage when an ensemble is present. + Agreement alone does not establish accuracy against independent controls or authorize release. +4. Investigate null or surprising results before assigning a cause. + A control intervention supports a causal explanation only when it isolates the proposed mechanism. +5. Retain actual costs, failures, exclusions, and uncertainty with the result. + Apply the product's release policy to the deciding evidence and record what the evidence cannot establish. + +## Then consider + +| Condition | Skill | +|---|---| +| Scoring or case coverage cannot test the required outcome | `eval-architect` with the failed controls and missing coverage | +| Measurement supports a scoped optimization experiment | `improve-conductor` with the baseline, registered comparison, and resource limits | diff --git a/CHANGELOG.md b/CHANGELOG.md index f99cd3f2..04454784 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,14 @@ ## Unreleased +### Changed + +- Update evaluator, execution, knowledge, and profile dependencies as one supported set. + Generated applications use the same dependency versions. +- Forward evaluator revisions and paid-call context through ensemble judges. + Campaign caches invalidate when `judgeVersion` changes, and paid calls retain campaign cost tags. +- Clarify that rater agreement does not establish evaluator accuracy or authorize promotion. + ### Added - Export renderer-neutral `groupConversationMessages` from the existing `/web` diff --git a/README.md b/README.md index 046547eb..554a9373 100644 --- a/README.md +++ b/README.md @@ -151,6 +151,13 @@ The **complete, always-current reference** — every published subpath, its expo - [`/runtime`](src/runtime) — the bounded tool loop; the same loop drives a sandbox agent, a Worker, or an in-browser copilot behind one `streamTurn` seam. - [`/trace`](src/trace) — bounded stage timing, turn waterfalls, and mission traces over an injected telemetry carrier. +**Evaluate changes** + +- [`/eval-campaign`](docs/api/eval-campaign.md) composes Eval campaigns, optimization methods, and ensemble judges. + Change `judgeVersion` when the scorer configuration changes. + Record paid judge calls through the supplied `costLedger`, `costPhase`, and `costTags`. + `trustVerdicts` checks agreement and coverage; use Eval's `auditEvaluator` to test accuracy against independent controls. + **The server chat vertical** ([`examples/chat-app.md`](./examples/chat-app.md)) - [`/chat-routes`](src/chat-routes) — `createChatTurnRoutes`: auth → store → streaming turn with buffered replay → uploads → sidecar question answering, assembled. Plus `runDetachedTurn` for autonomous turns a browser can still watch live. - [`/chat-store`](src/chat-store) · [`/interactions`](src/interactions) · [`/plans`](src/plans) — persistence, human-in-the-loop asks, and the durable plan projection. diff --git a/create-agent-app/package.json b/create-agent-app/package.json index a1735911..6fe3c9b1 100644 --- a/create-agent-app/package.json +++ b/create-agent-app/package.json @@ -1,6 +1,6 @@ { "name": "@tangle-network/create-agent-app", - "version": "0.47.25", + "version": "0.48.0", "description": "Create a tested Tangle agent application with either a tool loop or a complete chat interface.", "keywords": [ "tangle", diff --git a/create-agent-app/template-chat/_package.json b/create-agent-app/template-chat/_package.json index e1507eba..27aeba65 100644 --- a/create-agent-app/template-chat/_package.json +++ b/create-agent-app/template-chat/_package.json @@ -21,8 +21,8 @@ "dependencies": { "@tangle-network/agent-app": "__AGENT_APP_VERSION__", "@tangle-network/agent-gateway": "0.10.0", - "@tangle-network/agent-interface": "2.6.0", - "@tangle-network/agent-runtime": "0.222.1", + "@tangle-network/agent-interface": "2.6.1", + "@tangle-network/agent-runtime": "0.229.0", "@tangle-network/sandbox": "0.39.4", "better-auth": "^1.7.2", "drizzle-orm": "^0.45.2", @@ -30,13 +30,13 @@ "viem": "^2.0.0" }, "peerDependencies": { - "@tangle-network/agent-eval": "0.180.0", + "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": ">=0.53.55 <0.54.0" }, "devDependencies": { - "@tangle-network/agent-eval": "0.180.0", + "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": "0.53.55", - "@tangle-network/agent-knowledge": "15.0.3", + "@tangle-network/agent-knowledge": "17.0.1", "@cloudflare/workers-types": "^5.20260827.1", "@types/better-sqlite3": "^9.6.0", "@types/node": "^22.20.1", diff --git a/create-agent-app/template/_package.json b/create-agent-app/template/_package.json index b2f4e7fc..30f42476 100644 --- a/create-agent-app/template/_package.json +++ b/create-agent-app/template/_package.json @@ -21,17 +21,17 @@ "@tangle-network/agent-app": "__AGENT_APP_VERSION__" }, "peerDependencies": { - "@tangle-network/agent-eval": "0.180.0", + "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": ">=0.53.55 <0.54.0", - "@tangle-network/agent-interface": "2.6.0", - "@tangle-network/agent-runtime": "0.222.1" + "@tangle-network/agent-interface": "2.6.1", + "@tangle-network/agent-runtime": "0.229.0" }, "devDependencies": { - "@tangle-network/agent-eval": "0.180.0", + "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": "0.53.55", - "@tangle-network/agent-interface": "2.6.0", - "@tangle-network/agent-knowledge": "15.0.3", - "@tangle-network/agent-runtime": "0.222.1", + "@tangle-network/agent-interface": "2.6.1", + "@tangle-network/agent-knowledge": "17.0.1", + "@tangle-network/agent-runtime": "0.229.0", "@tangle-network/sandbox": "0.39.4", "@types/node": "^22.20.1", "typescript": "^7.0.2", diff --git a/docs/api/eval-campaign.md b/docs/api/eval-campaign.md index 5812ff38..3d4a9ddc 100644 --- a/docs/api/eval-campaign.md +++ b/docs/api/eval-campaign.md @@ -280,7 +280,7 @@ interface TrustItem ### `TrustThresholds` -`interface` — Thresholds for {@link trustVerdicts}. +`interface` — Configurable agreement and coverage thresholds for {@link trustVerdicts}. ```ts interface TrustThresholds @@ -296,7 +296,7 @@ interface TrustVerdict ### `trustVerdicts` -`function` — Decide whether an ensemble's per-item verdicts are trustworthy enough to believe a lift computed from them. +`function` — Check an ensemble against its configured agreement and coverage thresholds. ```ts (items: readonly TrustItem[], thresholds?: TrustThresholds) => TrustVerdict diff --git a/docs/codemap.json b/docs/codemap.json index af970b48..d420fe71 100644 --- a/docs/codemap.json +++ b/docs/codemap.json @@ -5635,7 +5635,7 @@ "name": "TrustThresholds", "kind": "interface", "signature": "interface TrustThresholds", - "doc": "Thresholds for {@link trustVerdicts}." + "doc": "Configurable agreement and coverage thresholds for {@link trustVerdicts}." }, { "name": "TrustVerdict", @@ -5647,7 +5647,7 @@ "name": "trustVerdicts", "kind": "function", "signature": "(items: readonly TrustItem[], thresholds?: TrustThresholds) => TrustVerdict", - "doc": "Decide whether an ensemble's per-item verdicts are trustworthy enough to believe a lift computed from them." + "doc": "Check an ensemble against its configured agreement and coverage thresholds." } ] }, diff --git a/docs/llms-full.txt b/docs/llms-full.txt index 2b6c54eb..4ad7af6b 100644 --- a/docs/llms-full.txt +++ b/docs/llms-full.txt @@ -7315,7 +7315,7 @@ interface TrustItem ### `TrustThresholds` -`interface` — Thresholds for {@link trustVerdicts}. +`interface` — Configurable agreement and coverage thresholds for {@link trustVerdicts}. ```ts interface TrustThresholds @@ -7331,7 +7331,7 @@ interface TrustVerdict ### `trustVerdicts` -`function` — Decide whether an ensemble's per-item verdicts are trustworthy enough to believe a lift computed from them. +`function` — Check an ensemble against its configured agreement and coverage thresholds. ```ts (items: readonly TrustItem[], thresholds?: TrustThresholds) => TrustVerdict diff --git a/package.json b/package.json index 01abe95c..ebf41726 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@tangle-network/agent-app", - "version": "0.47.25", + "version": "0.48.0", "packageManager": "pnpm@11.24.0", "description": "Build agent applications with typed chat, tools, sandboxes, integrations, billing, and evaluation.", "keywords": [ @@ -535,13 +535,13 @@ "@storybook/react-vite": "^10.5.10", "@tailwindcss/postcss": "^4.3.3", "@tangle-network/agent-docs": "0.2.1", - "@tangle-network/agent-eval": "0.180.0", + "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-gateway": "0.10.0", "@tangle-network/agent-integrations": "0.53.55", - "@tangle-network/agent-interface": "2.6.0", - "@tangle-network/agent-knowledge": "15.0.3", - "@tangle-network/agent-profile-materialize": "0.19.0", - "@tangle-network/agent-runtime": "0.222.1", + "@tangle-network/agent-interface": "2.6.1", + "@tangle-network/agent-knowledge": "17.0.1", + "@tangle-network/agent-profile-materialize": "0.20.1", + "@tangle-network/agent-runtime": "0.229.0", "@tangle-network/brand": "1.5.0", "@tangle-network/sandbox": "0.39.4", "@tangle-network/sandbox-ui": "0.113.3", @@ -596,12 +596,12 @@ "peerDependencies": { "@firecrawl/pdf-inspector-wasm": ">=0.1.3", "@huggingface/transformers": ">=3", - "@tangle-network/agent-eval": ">=0.180.0 <0.181.0", + "@tangle-network/agent-eval": ">=0.182.0 <0.183.0", "@tangle-network/agent-integrations": ">=0.53.55 <0.54.0", - "@tangle-network/agent-interface": "^2.6.0", - "@tangle-network/agent-knowledge": "^15.0.3", - "@tangle-network/agent-profile-materialize": ">=0.19.0 <0.20.0", - "@tangle-network/agent-runtime": ">=0.222.1 <0.223.0", + "@tangle-network/agent-interface": "^2.6.1", + "@tangle-network/agent-knowledge": "^17.0.1", + "@tangle-network/agent-profile-materialize": ">=0.20.1 <0.21.0", + "@tangle-network/agent-runtime": ">=0.229.0 <0.230.0", "@tangle-network/brand": ">=1.5.0", "@tangle-network/sandbox": ">=0.39.4 <0.40.0", "@tangle-network/sandbox-ui": ">=0.113.3 <0.114.0", diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index be9ccd03..a4c9f9fa 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -6,7 +6,6 @@ settings: overrides: esbuild: 0.28.1 - '@tangle-network/agent-interface': 2.6.0 importers: @@ -38,8 +37,8 @@ importers: specifier: 0.2.1 version: 0.2.1 '@tangle-network/agent-eval': - specifier: 0.180.0 - version: 0.180.0 + specifier: 0.182.0 + version: 0.182.0 '@tangle-network/agent-gateway': specifier: 0.10.0 version: 0.10.0 @@ -47,17 +46,17 @@ importers: specifier: 0.53.55 version: 0.53.55(@types/node@24.13.3) '@tangle-network/agent-interface': - specifier: 2.6.0 - version: 2.6.0 + specifier: 2.6.1 + version: 2.6.1 '@tangle-network/agent-knowledge': - specifier: 15.0.3 - version: 15.0.3(@tangle-network/agent-eval@0.180.0)(@tangle-network/agent-interface@2.6.0) + specifier: 17.0.1 + version: 17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-profile-materialize': - specifier: 0.19.0 - version: 0.19.0(@tangle-network/agent-interface@2.6.0) + specifier: 0.20.1 + version: 0.20.1(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-runtime': - specifier: 0.222.1 - version: 0.222.1(@tangle-network/agent-eval@0.180.0)(@tangle-network/agent-interface@2.6.0)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))) + specifier: 0.229.0 + version: 0.229.0(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))) '@tangle-network/brand': specifier: 1.5.0 version: 1.5.0(react@19.2.8) @@ -66,7 +65,7 @@ importers: version: 0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) '@tangle-network/sandbox-ui': specifier: 0.113.3 - version: 0.113.3(d27cb959a9c7462878f04dcd035c9dc3) + version: 0.113.3(9d6b3ea301e53b12b06a836f6fcb6d5f) '@tangle-network/ui': specifier: ^11.8.0 version: 11.8.0(@hocuspocus/provider@4.6.0(y-protocols@1.0.7(yjs@13.6.32))(yjs@13.6.32))(@tangle-network/brand@1.5.0(react@19.2.8))(@tiptap/core@3.30.5(@tiptap/pm@3.30.5))(@tiptap/react@3.30.5(@floating-ui/dom@1.8.0)(@tiptap/core@3.30.5(@tiptap/pm@3.30.5))(@tiptap/pm@3.30.5)(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(@tiptap/starter-kit@3.30.5)(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(nanostores@1.5.2)(react-dom@19.2.8(react@19.2.8))(react-router@8.3.1(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(react@19.2.8)(yjs@13.6.32) @@ -2204,8 +2203,8 @@ packages: engines: {node: '>=20'} hasBin: true - '@tangle-network/agent-eval@0.180.0': - resolution: {integrity: sha512-sQFExEf/3eaVM0/PbNWsFJ4cSWRgQayx6VB0nR6yy1Av9XNzQ7RuVMllOmJYz3anoJ4xHjiMpbrFlRnu4MM5VQ==} + '@tangle-network/agent-eval@0.182.0': + resolution: {integrity: sha512-6vGzeieF6c5kBcFGUW8xdyNTsR3Da4JRmOn6dvQdReXMzlEswuNS/uHNjOUHHknvAl41Hu0FwSZwaKQsE0IUgA==} engines: {node: '>=20.19.0'} hasBin: true @@ -2220,26 +2219,34 @@ packages: '@tangle-network/agent-interface@2.6.0': resolution: {integrity: sha512-tEByATif9oM5EEQB94Fytjt1Jtd70KxbWudKpwMQYyqXw2dcVFh2JTkeE8rIQP2F8KDLj9n5VHFkwg/ikbr7dg==} - '@tangle-network/agent-knowledge@15.0.3': - resolution: {integrity: sha512-a9tmJXloklXHRTFVBPLmJGvgJFTCVfLp/GCo/zY31pA0g6/3f7jdGoQPjuhdNhbkJixxJCaCy4whevx9uyO86Q==} + '@tangle-network/agent-interface@2.6.1': + resolution: {integrity: sha512-ceWVBFTSgSfyGlhQLrTWqUX42yd2leLnbtUcB5FMU0ksgr4KKaOk33xq0+FQxhACf6AigCiIGLKOi4vR4XDVkg==} + + '@tangle-network/agent-knowledge@17.0.1': + resolution: {integrity: sha512-B85focFu+dCfdMSQ8NQQXQZAtJIUra7lW4K0o24IkRbOM0DVTkRH2DJjzSibnVKavlT6oHilfBUGHWd+fVZbBg==} engines: {node: '>=20.19.0'} hasBin: true peerDependencies: - '@tangle-network/agent-eval': '>=0.174.0 <0.181.0' - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-eval': '>=0.182.0 <0.183.0' + '@tangle-network/agent-interface': ^2.0.0 '@tangle-network/agent-profile-materialize@0.19.0': resolution: {integrity: sha512-Q4vcsV0s94fwOtSFnaX6nGlpzwip9zDK4PmzECwUnNoyQvmo65KqvrYmPonN4c0we9ZoBkneIgl7jVvi9awjEA==} peerDependencies: - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-interface': ^1.0.0 || ^2.0.0 + + '@tangle-network/agent-profile-materialize@0.20.1': + resolution: {integrity: sha512-4HUEowvLtdO+5DHu+lahBOuHi6bCJ+tvccwaxDbcUSfq5l3g2fVAwGyOx3NBm5APKsCkcVIuBTIrwqWQdv2Ayw==} + peerDependencies: + '@tangle-network/agent-interface': ^2.6.1 - '@tangle-network/agent-runtime@0.222.1': - resolution: {integrity: sha512-MnNZflH77L4MjjatdcbudjGjIEgk96hYDjEcEQSlYk0hfJ7kXUFY0Hn6/wTn3wXhQB4jP9rwDb3c8RAE/lJorQ==} + '@tangle-network/agent-runtime@0.229.0': + resolution: {integrity: sha512-VEMoS4BuUxgS6WgIPcbGkfGeLx6JMr/wULH/oadsV5lno2BpVAYs+uZJpHa4bTOfaJPomGvQLXTDD3psorcGzg==} engines: {node: '>=22.13.0'} hasBin: true peerDependencies: - '@tangle-network/agent-eval': '>=0.180.0 <0.181.0' - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-eval': '>=0.182.0 <0.183.0' + '@tangle-network/agent-interface': ^2.6.0 '@tangle-network/sandbox': '>=0.36.4 <0.40.0' '@tangle-network/agent-trace-contract@1.0.2': @@ -2258,7 +2265,7 @@ packages: peerDependencies: '@hocuspocus/provider': '>=2.15.0 <5.0.0' '@nanostores/react': '>=0.8.4 <2.0.0' - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-interface': ^1.0.0 || ^2.0.0 '@tangle-network/brand': ^1.5.0 '@tangle-network/ui': ^11.6.0 '@tanstack/react-query': ^5.0.0 @@ -7019,7 +7026,7 @@ snapshots: dependencies: '@typescript/typescript6': 6.0.2 - '@tangle-network/agent-eval@0.180.0': + '@tangle-network/agent-eval@0.182.0': dependencies: '@asteasolutions/zod-to-openapi': 9.1.0(zod@4.5.4) '@hono/node-server': 2.1.1(hono@4.13.5) @@ -7074,25 +7081,35 @@ snapshots: spdx-expression-parse: 5.0.0 zod: 4.5.4 - '@tangle-network/agent-knowledge@15.0.3(@tangle-network/agent-eval@0.180.0)(@tangle-network/agent-interface@2.6.0)': + '@tangle-network/agent-interface@2.6.1': dependencies: - '@tangle-network/agent-eval': 0.180.0 - '@tangle-network/agent-interface': 2.6.0 + '@noble/hashes': 2.4.0 + spdx-expression-parse: 5.0.0 + zod: 4.5.4 + + '@tangle-network/agent-knowledge@17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)': + dependencies: + '@tangle-network/agent-eval': 0.182.0 + '@tangle-network/agent-interface': 2.6.1 '@types/proper-lockfile': 4.1.4 proper-lockfile: 4.1.2 zod: 4.5.4 - '@tangle-network/agent-profile-materialize@0.19.0(@tangle-network/agent-interface@2.6.0)': + '@tangle-network/agent-profile-materialize@0.19.0(@tangle-network/agent-interface@2.6.1)': dependencies: - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-interface': 2.6.1 - '@tangle-network/agent-runtime@0.222.1(@tangle-network/agent-eval@0.180.0)(@tangle-network/agent-interface@2.6.0)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)))': + '@tangle-network/agent-profile-materialize@0.20.1(@tangle-network/agent-interface@2.6.1)': + dependencies: + '@tangle-network/agent-interface': 2.6.1 + + '@tangle-network/agent-runtime@0.229.0(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)))': dependencies: '@tangle-network/agent-core': 0.9.6 - '@tangle-network/agent-eval': 0.180.0 - '@tangle-network/agent-interface': 2.6.0 - '@tangle-network/agent-knowledge': 15.0.3(@tangle-network/agent-eval@0.180.0)(@tangle-network/agent-interface@2.6.0) - '@tangle-network/agent-profile-materialize': 0.19.0(@tangle-network/agent-interface@2.6.0) + '@tangle-network/agent-eval': 0.182.0 + '@tangle-network/agent-interface': 2.6.1 + '@tangle-network/agent-knowledge': 17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) + '@tangle-network/agent-profile-materialize': 0.19.0(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-trace-contract': 1.0.2 '@tangle-network/sandbox': 0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) tar-stream: 3.2.1 @@ -7108,13 +7125,13 @@ snapshots: optionalDependencies: react: 19.2.8 - '@tangle-network/sandbox-ui@0.113.3(d27cb959a9c7462878f04dcd035c9dc3)': + '@tangle-network/sandbox-ui@0.113.3(9d6b3ea301e53b12b06a836f6fcb6d5f)': dependencies: '@lobehub/icons-static-svg': 1.94.0 '@pierre/diffs': 1.3.6(@shikijs/themes@4.4.3)(react-dom@19.2.8(react@19.2.8))(react@19.2.8) '@radix-ui/react-dropdown-menu': 2.1.24(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(react-dom@19.2.8(react@19.2.8))(react@19.2.8) '@radix-ui/react-select': 2.3.7(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(react-dom@19.2.8(react@19.2.8))(react@19.2.8) - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-interface': 2.6.1 '@tangle-network/brand': 1.5.0(react@19.2.8) '@tangle-network/ui': 11.8.0(@hocuspocus/provider@4.6.0(y-protocols@1.0.7(yjs@13.6.32))(yjs@13.6.32))(@tangle-network/brand@1.5.0(react@19.2.8))(@tiptap/core@3.30.5(@tiptap/pm@3.30.5))(@tiptap/react@3.30.5(@floating-ui/dom@1.8.0)(@tiptap/core@3.30.5(@tiptap/pm@3.30.5))(@tiptap/pm@3.30.5)(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(@tiptap/starter-kit@3.30.5)(@types/react-dom@19.2.5(@types/react@19.2.18))(@types/react@19.2.18)(nanostores@1.5.2)(react-dom@19.2.8(react@19.2.8))(react-router@8.3.1(react-dom@19.2.8(react@19.2.8))(react@19.2.8))(react@19.2.8)(yjs@13.6.32) diff: 9.0.0 @@ -7143,7 +7160,7 @@ snapshots: '@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))': dependencies: '@tangle-network/agent-core': 0.9.6 - '@tangle-network/agent-interface': 2.6.0 + '@tangle-network/agent-interface': 2.6.1 zod: 4.5.4 optionalDependencies: viem: 2.56.0(typescript@7.0.2)(zod@4.4.3) diff --git a/pnpm-workspace.yaml b/pnpm-workspace.yaml index 9b33a52d..3fd9e810 100644 --- a/pnpm-workspace.yaml +++ b/pnpm-workspace.yaml @@ -20,4 +20,3 @@ allowBuilds: overrides: esbuild: 0.28.1 - '@tangle-network/agent-interface': 2.6.0 diff --git a/src/eval-campaign/index.test.ts b/src/eval-campaign/index.test.ts index a7b20587..4a5eccad 100644 --- a/src/eval-campaign/index.test.ts +++ b/src/eval-campaign/index.test.ts @@ -7,7 +7,8 @@ import { describe, expect, it } from 'vitest' -import type { Scenario } from '@tangle-network/agent-eval/campaign' +import { CostLedger } from '@tangle-network/agent-eval' +import { inMemoryCampaignStorage, runCampaign, type Scenario } from '@tangle-network/agent-eval/campaign' import { buildEnsembleJudge } from './index' type Dim = 'accuracy' | 'tone' @@ -54,6 +55,69 @@ describe('buildEnsembleJudge', () => { ]) }) + it('records each paid ensemble call in the campaign cost account', async () => { + const costLedger = new CostLedger() + const judge = buildEnsembleJudge({ + name: 'paid-ensemble', + rubric: RUBRIC, + judgeReps: 2, + async scoreOne({ costLedger: account, costPhase, costTags, rep }) { + if (!account || !costPhase || !costTags) throw new Error('missing campaign cost context') + const paid = await account.runPaidCall({ + channel: 'judge', + phase: costPhase, + actor: `fixture-judge-${rep}`, + tags: costTags, + execute: async () => ({ accuracy: 0.8, tone: 0.6 }), + receipt: () => ({ model: 'fixture-judge', inputTokens: 10, outputTokens: 2, actualCostUsd: 0.02 }), + }) + if (!paid.succeeded) throw paid.error + return { model: `fixture-judge-${rep}`, perDimension: paid.value } + }, + }) + const result = await runCampaign({ + runDir: 'mem://paid-ensemble', + storage: inMemoryCampaignStorage(), + scenarios: [scenario], + dispatch: async () => ({ text: 'fixture' }), + expectUsage: 'off', + judges: [judge], + costLedger, + costPhase: 'final', + costTags: { evaluation: 'fixture' }, + }) + + expect(result.cells[0]?.error).toBeUndefined() + expect(result.aggregates.cost.totalCostUsd).toBeCloseTo(0.04) + expect(costLedger.list({ channel: 'judge', phase: 'final', tags: { evaluation: 'fixture' } })).toHaveLength(2) + }) + + it('invalidates cached judgments when the caller changes the evaluator revision', async () => { + const storage = inMemoryCampaignStorage() + let calls = 0 + const evaluate = (revision: string) => runCampaign({ + runDir: 'mem://versioned-ensemble', + storage, + scenarios: [scenario], + dispatch: async () => ({ text: 'fixture' }), + expectUsage: 'off', + judges: [buildEnsembleJudge({ + name: 'versioned-ensemble', + judgeVersion: revision, + rubric: RUBRIC, + async scoreOne() { + calls += 1 + return { model: 'fixture-judge', perDimension: { accuracy: 1, tone: 1 } } + }, + })], + }) + + const first = await evaluate('rubric-v1') + const second = await evaluate('rubric-v2') + expect(calls).toBe(2) + expect(second.manifestHash).not.toBe(first.manifestHash) + }) + it('a single rep failing does NOT fail the cell — means over survivors', async () => { const judge = buildEnsembleJudge({ name: 'test', diff --git a/src/eval-campaign/index.ts b/src/eval-campaign/index.ts index 148e07cb..c4612dc7 100644 --- a/src/eval-campaign/index.ts +++ b/src/eval-campaign/index.ts @@ -36,6 +36,8 @@ import type { export interface EnsembleJudgeConfig { /** Judge name — appears in traces and scorecards. */ name: string + /** Scoring revision for campaign caches. Change it when rubric, model, or callback settings change. */ + judgeVersion?: JudgeConfig['judgeVersion'] /** Stable-ordered rubric dimensions. Drives the `JudgeDimension` list AND the * reducer keys, so a judge that omits a dimension scores it 0 (never silently * dropped). */ @@ -47,11 +49,9 @@ export interface EnsembleJudgeConfig['score']>[0] & { rep: number }) => Promise> /** Independent judge calls per artifact, reduced by `aggregateJudgeVerdicts`. @@ -85,10 +85,11 @@ export function buildEnsembleJudge ({ key, description: cfg.describe?.(key) ?? key })), - async score({ artifact, scenario, signal }): Promise { + async score(input): Promise { const settled = await Promise.allSettled( - Array.from({ length: reps }, (_, rep) => cfg.scoreOne({ artifact, scenario, signal, rep })), + Array.from({ length: reps }, (_, rep) => cfg.scoreOne({ ...input, rep })), ) const verdicts: JudgeVerdict[] = settled.map((r, rep) => r.status === 'fulfilled' @@ -102,9 +103,8 @@ export function buildEnsembleJudge { verdicts: readonly JudgeVerdict[] } -/** Thresholds for {@link trustVerdicts}. All overridable; defaults are the - * conservative after-gate bar. */ +/** Configurable agreement and coverage thresholds for {@link trustVerdicts}. */ export interface TrustThresholds { - /** Minimum corpus inter-rater reliability (Krippendorff-style α). Below this - * the raters agree no better than chance. Default 0.2. */ + /** Minimum corpus inter-rater reliability (Krippendorff-style α). Default 0.2. */ irrFloor?: number /** Maximum per-item rater spread (`max − min` over a single item's surviving * raters, across its dimensions). Above this the raters split ON THAT ITEM. @@ -110,14 +89,12 @@ function itemSpread(survivorVerdicts: JudgeVerdict[]): numb } /** - * Decide whether an ensemble's per-item verdicts are trustworthy enough to - * believe a lift computed from them. Pure: no LLM, no I/O, no clock, no random — + * Check an ensemble against its configured agreement and coverage thresholds. Pure: no LLM, no I/O, no clock, no random — * the same `items` + `thresholds` always yield the same verdict. * * Sibling to {@link aggregateJudgeVerdicts}: that reduces ONE item's raters to a - * composite; this audits the raters ACROSS items and reports whether the - * composites are believable. Run it on the corpus of held-out items before - * reporting any lift over their scores. + * composite; this summarizes agreement across the supplied items. + * A passing result does not establish evaluator accuracy or authorize a release. * * @throws if `items` is empty — an empty corpus has no measurable trust, and a * silent `trustworthy: true` over zero evidence is the exact lie the gate @@ -139,7 +116,7 @@ export function trustVerdicts( // lines up; per (item, dimension) one JudgeScore per rater, in item-then- // dimension order — the layout interRaterReliability chunks back into items. const maxRaters = items.reduce((m, it) => Math.max(m, survivors(it).length), 0) - const raterSeries: JudgeScore[][] = Array.from({ length: maxRaters }, () => []) + const raterSeries: DimensionJudgeScore[][] = Array.from({ length: maxRaters }, () => []) const perItemSpread: Record = {} const splitItems: Array<{ itemId: string; spread: number }> = [] const starvedItems: Array<{ itemId: string; n: number }> = [] diff --git a/src/peer-floors/check.test.ts b/src/peer-floors/check.test.ts index bb6a94e3..f9ee8208 100644 --- a/src/peer-floors/check.test.ts +++ b/src/peer-floors/check.test.ts @@ -149,12 +149,13 @@ describe('this package audits itself', () => { expect(satisfiesRange('1.9.0', range!)).toBe(false) expect(satisfiesRange('2.1.1', range!)).toBe(false) expect(satisfiesRange('2.5.9', range!)).toBe(false) - expect(satisfiesRange('2.6.0', range!)).toBe(true) + expect(satisfiesRange('2.6.0', range!)).toBe(false) + expect(satisfiesRange('2.6.1', range!)).toBe(true) expect(satisfiesRange('2.6.9', range!)).toBe(true) expect(satisfiesRange('3.0.0', range!)).toBe(false) }) - it('supports the verified Runtime lines without claiming the next one', async () => { + it('supports the verified Runtime line without claiming the next one', async () => { const root = join(here, '..', '..') const own = JSON.parse( await readFile(join(root, 'package.json'), 'utf8'), @@ -162,10 +163,10 @@ describe('this package audits itself', () => { const range = own.peerDependencies?.['@tangle-network/agent-runtime'] expect(range).toBeDefined() - expect(satisfiesRange('0.222.0', range!)).toBe(false) - expect(satisfiesRange('0.222.1', range!)).toBe(true) - expect(satisfiesRange('0.222.9', range!)).toBe(true) - expect(satisfiesRange('0.223.0', range!)).toBe(false) + expect(satisfiesRange('0.228.9', range!)).toBe(false) + expect(satisfiesRange('0.229.0', range!)).toBe(true) + expect(satisfiesRange('0.229.9', range!)).toBe(true) + expect(satisfiesRange('0.230.0', range!)).toBe(false) }) // The floors this shell PUBLISHES must be satisfiable by the tree it is From 6998b8a09862a3fa757e6bc5fc5e4f11080575a8 Mon Sep 17 00:00:00 2001 From: Drew Stone Date: Tue, 15 Sep 2026 21:59:16 -0700 Subject: [PATCH 2/2] chore(eval): align the published execution dependency cohort --- create-agent-app/template-chat/_package.json | 6 ++-- create-agent-app/template/_package.json | 8 ++--- package.json | 12 +++---- pnpm-lock.yaml | 36 ++++++++++---------- src/peer-floors/check.test.ts | 8 ++--- src/signoff/run.test.ts | 2 +- 6 files changed, 36 insertions(+), 36 deletions(-) diff --git a/create-agent-app/template-chat/_package.json b/create-agent-app/template-chat/_package.json index 27aeba65..a02dc5d9 100644 --- a/create-agent-app/template-chat/_package.json +++ b/create-agent-app/template-chat/_package.json @@ -22,8 +22,8 @@ "@tangle-network/agent-app": "__AGENT_APP_VERSION__", "@tangle-network/agent-gateway": "0.10.0", "@tangle-network/agent-interface": "2.6.1", - "@tangle-network/agent-runtime": "0.229.0", - "@tangle-network/sandbox": "0.39.4", + "@tangle-network/agent-runtime": "0.231.1", + "@tangle-network/sandbox": "0.40.2", "better-auth": "^1.7.2", "drizzle-orm": "^0.45.2", "hono": "^4.13.5", @@ -36,7 +36,7 @@ "devDependencies": { "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": "0.53.55", - "@tangle-network/agent-knowledge": "17.0.1", + "@tangle-network/agent-knowledge": "17.0.2", "@cloudflare/workers-types": "^5.20260827.1", "@types/better-sqlite3": "^9.6.0", "@types/node": "^22.20.1", diff --git a/create-agent-app/template/_package.json b/create-agent-app/template/_package.json index 30f42476..35f5ee15 100644 --- a/create-agent-app/template/_package.json +++ b/create-agent-app/template/_package.json @@ -24,15 +24,15 @@ "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": ">=0.53.55 <0.54.0", "@tangle-network/agent-interface": "2.6.1", - "@tangle-network/agent-runtime": "0.229.0" + "@tangle-network/agent-runtime": "0.231.1" }, "devDependencies": { "@tangle-network/agent-eval": "0.182.0", "@tangle-network/agent-integrations": "0.53.55", "@tangle-network/agent-interface": "2.6.1", - "@tangle-network/agent-knowledge": "17.0.1", - "@tangle-network/agent-runtime": "0.229.0", - "@tangle-network/sandbox": "0.39.4", + "@tangle-network/agent-knowledge": "17.0.2", + "@tangle-network/agent-runtime": "0.231.1", + "@tangle-network/sandbox": "0.40.2", "@types/node": "^22.20.1", "typescript": "^7.0.2", "viem": "^2.0.0", diff --git a/package.json b/package.json index ebf41726..439779b9 100644 --- a/package.json +++ b/package.json @@ -539,11 +539,11 @@ "@tangle-network/agent-gateway": "0.10.0", "@tangle-network/agent-integrations": "0.53.55", "@tangle-network/agent-interface": "2.6.1", - "@tangle-network/agent-knowledge": "17.0.1", + "@tangle-network/agent-knowledge": "17.0.2", "@tangle-network/agent-profile-materialize": "0.20.1", - "@tangle-network/agent-runtime": "0.229.0", + "@tangle-network/agent-runtime": "0.231.1", "@tangle-network/brand": "1.5.0", - "@tangle-network/sandbox": "0.39.4", + "@tangle-network/sandbox": "0.40.2", "@tangle-network/sandbox-ui": "0.113.3", "@tangle-network/ui": "^11.8.0", "@testing-library/dom": "^10.4.1", @@ -599,11 +599,11 @@ "@tangle-network/agent-eval": ">=0.182.0 <0.183.0", "@tangle-network/agent-integrations": ">=0.53.55 <0.54.0", "@tangle-network/agent-interface": "^2.6.1", - "@tangle-network/agent-knowledge": "^17.0.1", + "@tangle-network/agent-knowledge": "^17.0.2", "@tangle-network/agent-profile-materialize": ">=0.20.1 <0.21.0", - "@tangle-network/agent-runtime": ">=0.229.0 <0.230.0", + "@tangle-network/agent-runtime": ">=0.231.1 <0.232.0", "@tangle-network/brand": ">=1.5.0", - "@tangle-network/sandbox": ">=0.39.4 <0.40.0", + "@tangle-network/sandbox": ">=0.40.2 <0.41.0", "@tangle-network/sandbox-ui": ">=0.113.3 <0.114.0", "@tangle-network/ui": ">=11.6.0 <12.0.0", "@tiptap/core": ">=3.28.0 <4.0.0", diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index a4c9f9fa..ee870a9b 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -49,20 +49,20 @@ importers: specifier: 2.6.1 version: 2.6.1 '@tangle-network/agent-knowledge': - specifier: 17.0.1 - version: 17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) + specifier: 17.0.2 + version: 17.0.2(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-profile-materialize': specifier: 0.20.1 version: 0.20.1(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-runtime': - specifier: 0.229.0 - version: 0.229.0(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))) + specifier: 0.231.1 + version: 0.231.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.40.2(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))) '@tangle-network/brand': specifier: 1.5.0 version: 1.5.0(react@19.2.8) '@tangle-network/sandbox': - specifier: 0.39.4 - version: 0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) + specifier: 0.40.2 + version: 0.40.2(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) '@tangle-network/sandbox-ui': specifier: 0.113.3 version: 0.113.3(9d6b3ea301e53b12b06a836f6fcb6d5f) @@ -2222,8 +2222,8 @@ packages: '@tangle-network/agent-interface@2.6.1': resolution: {integrity: sha512-ceWVBFTSgSfyGlhQLrTWqUX42yd2leLnbtUcB5FMU0ksgr4KKaOk33xq0+FQxhACf6AigCiIGLKOi4vR4XDVkg==} - '@tangle-network/agent-knowledge@17.0.1': - resolution: {integrity: sha512-B85focFu+dCfdMSQ8NQQXQZAtJIUra7lW4K0o24IkRbOM0DVTkRH2DJjzSibnVKavlT6oHilfBUGHWd+fVZbBg==} + '@tangle-network/agent-knowledge@17.0.2': + resolution: {integrity: sha512-xPXGeDT8e7ag4Fz/lzZLLaNE29sWi9qyCr33KFP9kQ8MBBS2nw9R82KejJFDkvaKNm3zxI7iF4bms13PfJH3LA==} engines: {node: '>=20.19.0'} hasBin: true peerDependencies: @@ -2240,14 +2240,14 @@ packages: peerDependencies: '@tangle-network/agent-interface': ^2.6.1 - '@tangle-network/agent-runtime@0.229.0': - resolution: {integrity: sha512-VEMoS4BuUxgS6WgIPcbGkfGeLx6JMr/wULH/oadsV5lno2BpVAYs+uZJpHa4bTOfaJPomGvQLXTDD3psorcGzg==} + '@tangle-network/agent-runtime@0.231.1': + resolution: {integrity: sha512-zIjx6Vvrl6vTMIxh3qZDj7e+5WUU3B5YvCaAaumnjtaCjr/w6PWs00YAiWbFwEG6xqX0U3wDLP4Kflj1XmP6mA==} engines: {node: '>=22.13.0'} hasBin: true peerDependencies: '@tangle-network/agent-eval': '>=0.182.0 <0.183.0' '@tangle-network/agent-interface': ^2.6.0 - '@tangle-network/sandbox': '>=0.36.4 <0.40.0' + '@tangle-network/sandbox': '>=0.36.4 <0.41.0' '@tangle-network/agent-trace-contract@1.0.2': resolution: {integrity: sha512-v7uMh56jkEp4vckevEU9xKsIatbs5dqzGPp69dFLSSXUVit0RP6VD6EANMXVlTCUk+6wVKBLHJx23XspVCEiIA==} @@ -2312,8 +2312,8 @@ packages: yjs: optional: true - '@tangle-network/sandbox@0.39.4': - resolution: {integrity: sha512-yabZjharkZUaNZrgbrzofciv8Y1iI/I3M9/boozwk3ReuZwDIqYoRhF5QXXZNcFhHeaRCzsUCbyFFIVZ+RbnMw==} + '@tangle-network/sandbox@0.40.2': + resolution: {integrity: sha512-QvniJ2OuEG0AAxOolb+tUiFtzDBvI6/FdOpEKgvjeDue6elLsRB+dCF+3TFtBVwu9kgC0IrSIYLBKV/TXMW9UQ==} peerDependencies: '@mastra/core': ^1.36.0 '@modelcontextprotocol/sdk': ^1.30.0 @@ -7087,7 +7087,7 @@ snapshots: spdx-expression-parse: 5.0.0 zod: 4.5.4 - '@tangle-network/agent-knowledge@17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)': + '@tangle-network/agent-knowledge@17.0.2(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)': dependencies: '@tangle-network/agent-eval': 0.182.0 '@tangle-network/agent-interface': 2.6.1 @@ -7103,15 +7103,15 @@ snapshots: dependencies: '@tangle-network/agent-interface': 2.6.1 - '@tangle-network/agent-runtime@0.229.0(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)))': + '@tangle-network/agent-runtime@0.231.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1)(@tangle-network/sandbox@0.40.2(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)))': dependencies: '@tangle-network/agent-core': 0.9.6 '@tangle-network/agent-eval': 0.182.0 '@tangle-network/agent-interface': 2.6.1 - '@tangle-network/agent-knowledge': 17.0.1(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) + '@tangle-network/agent-knowledge': 17.0.2(@tangle-network/agent-eval@0.182.0)(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-profile-materialize': 0.19.0(@tangle-network/agent-interface@2.6.1) '@tangle-network/agent-trace-contract': 1.0.2 - '@tangle-network/sandbox': 0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) + '@tangle-network/sandbox': 0.40.2(viem@2.56.0(typescript@7.0.2)(zod@4.4.3)) tar-stream: 3.2.1 transitivePeerDependencies: - '@modelcontextprotocol/sdk' @@ -7157,7 +7157,7 @@ snapshots: - '@types/react' - '@types/react-dom' - '@tangle-network/sandbox@0.39.4(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))': + '@tangle-network/sandbox@0.40.2(viem@2.56.0(typescript@7.0.2)(zod@4.4.3))': dependencies: '@tangle-network/agent-core': 0.9.6 '@tangle-network/agent-interface': 2.6.1 diff --git a/src/peer-floors/check.test.ts b/src/peer-floors/check.test.ts index f9ee8208..5d5acd9d 100644 --- a/src/peer-floors/check.test.ts +++ b/src/peer-floors/check.test.ts @@ -163,10 +163,10 @@ describe('this package audits itself', () => { const range = own.peerDependencies?.['@tangle-network/agent-runtime'] expect(range).toBeDefined() - expect(satisfiesRange('0.228.9', range!)).toBe(false) - expect(satisfiesRange('0.229.0', range!)).toBe(true) - expect(satisfiesRange('0.229.9', range!)).toBe(true) - expect(satisfiesRange('0.230.0', range!)).toBe(false) + expect(satisfiesRange('0.231.0', range!)).toBe(false) + expect(satisfiesRange('0.231.1', range!)).toBe(true) + expect(satisfiesRange('0.231.9', range!)).toBe(true) + expect(satisfiesRange('0.232.0', range!)).toBe(false) }) // The floors this shell PUBLISHES must be satisfiable by the tree it is diff --git a/src/signoff/run.test.ts b/src/signoff/run.test.ts index 47f4cac4..27cbc9c9 100644 --- a/src/signoff/run.test.ts +++ b/src/signoff/run.test.ts @@ -329,7 +329,7 @@ describe('runSignoff (end to end)', { timeout: E2E_TIMEOUT_MS }, () => { { name: 'c', run: ${JSON.stringify(sleeper(400))} }, ]`, }) - const report = await runSignoff({ repoDir: repo, cacheDir: temp('signoff-cache-'), source: 'head' }) + const report = await runSignoff({ repoDir: repo, cacheDir: temp('signoff-cache-'), source: 'head', maxParallel: 3 }) expect(report.ok).toBe(true) expect(peakConcurrency(report.steps)).toBe(3)