diff --git a/CHANGELOG.md b/CHANGELOG.md index ad5de8324..3cee117ef 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,10 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval- ### Changed +- README and example guides include verified execution commands, public imports, and explicit limits on fixture results and release evidence. + The existing-agent quickstart is offline; its guide shows how to meter paid calls with the maintained transport and receipt helpers. +- The single-optimizer example accepts the same worker `PRICE_*` settings as the method-comparison example. + Both use one parser to validate endpoint rates before execution. - **Breaking:** Root `Scenario`, `JudgeScore`, and `GateDecision` now match `/contract`. Product workflows use `ProductScenario`, `DimensionJudgeScore`, and `HeldOutGateDecision`. - **Breaking:** Current canonical envelopes and algorithm identifiers are required for seals, attestations, and profile identities. @@ -42,6 +46,8 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval- ### Fixed +- Failure-cluster shares count all affected failed runs independently of the five displayed examples. + Multiple findings in the same cluster count once per run. - Outcome queries select the latest finite requested metric instead of an unrelated latest observation. Outcome-store corruption and unavailable evidence remain visible failures. - Calibration preserves clipped observations, measures constant predictors, and honors the requested bin count. diff --git a/README.md b/README.md index 52b08e831..6f17e97d0 100644 --- a/README.md +++ b/README.md @@ -13,6 +13,8 @@ It records outputs, failures, costs, and evidence for each comparison. ## Install +Use Node.js 20.19 or newer. + ```sh pnpm add @tangle-network/agent-eval ``` @@ -20,7 +22,7 @@ pnpm add @tangle-network/agent-eval ## Quickstart This complete example runs offline. -Replace the agent and judge with your product functions when it works. +Save it as `eval.mts`. ```ts import { defineAgentEval } from '@tangle-network/agent-eval/contract' @@ -50,11 +52,25 @@ const evalKit = defineAgentEval({ expectUsage: 'off', }) -console.log((await evalKit.evaluate()).aggregates.byJudge) -console.log( - (await evalKit.evaluate({ surface: 'Answer politely and cite the ticket id.' })).aggregates - .byJudge, -) +const baseline = await evalKit.evaluate() +const candidate = await evalKit.evaluate({ + surface: 'Answer politely and cite the ticket id.', +}) + +console.log('baseline:', baseline.aggregates.byJudge['ticket-id']?.mean) +console.log('candidate:', candidate.aggregates.byJudge['ticket-id']?.mean) +``` + +Run it with a TypeScript runner: + +```sh +pnpm add --save-dev tsx +pnpm exec tsx eval.mts +``` + +```text +baseline: 0 +candidate: 1 ``` The baseline scores `0`; the candidate scores `1` on all three cases. @@ -66,8 +82,9 @@ A **surface** is the prompt, skill, or configuration being changed. A **judge** scores the agent's result. `expectUsage: 'off'` applies because this example makes no paid calls. -Keep the default, `'assert'`, for model calls so missing cost receipts fail visibly. +Set `expectUsage: 'assert'` for paid agents so missing dispatch receipts become execution failures. The [runnable example](./examples/evaluate-a-change/) uses the same evaluation. +The [existing-agent example](./examples/foreign-agent-quickstart/) shows how to connect your agent and record model usage. ## Choose a workflow @@ -164,29 +181,17 @@ The [benchmark-book review](./docs/design/mlbenchmarks-book-review.md) records t ```sh pnpm install +pnpm build pnpm typecheck pnpm typecheck:examples pnpm typecheck:scripts +pnpm lint pnpm test -pnpm build pnpm verify:package ``` -Python compatibility tests use the locked dependencies: - -```sh -cd clients/python -uv sync --frozen --extra dev --group gepa-release -AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \ - uv run --frozen --extra dev --group gepa-release \ - pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py - -uv sync --frozen --extra dev --group skillopt-source --group gepa-source -uv run --frozen pytest - -uv sync --frozen --extra dev --extra dspy -uv run --frozen pytest tests/test_dspy_metric.py -``` +Build before checking examples because they resolve the package's generated declarations. +The [Python development guide](./clients/python/README.md#development) gives the locked commands for each optimizer environment. ## License diff --git a/clients/python/README.md b/clients/python/README.md index 1c3b023e4..016db7dd1 100644 --- a/clients/python/README.md +++ b/clients/python/README.md @@ -6,7 +6,7 @@ The Node package owns rubric execution, model calls, and scoring. ## Install -Python 3.10 or newer and Node.js 20 or newer are required. +Python 3.10 or newer and Node.js 20.19 or newer are required. Install matching package versions: ```sh @@ -19,15 +19,17 @@ Configure an OpenAI-compatible model endpoint for judge calls: ```sh export AGENT_EVAL_LLM_BASE_URL=https://api.openai.com/v1 export AGENT_EVAL_LLM_API_KEY="$YOUR_API_KEY" -export AGENT_EVAL_LLM_MODEL=gpt-4.1-mini +export AGENT_EVAL_LLM_MODEL="$YOUR_MODEL_ID" ``` `OPENAI_BASE_URL`, `OPENAI_API_KEY`, and `OPENAI_MODEL` are also accepted. The endpoint receives the content, rubric, and context passed to `client.judge()`. -The `agent-eval` binary is the only part of the package that reads a provider credential. -It is a server process, so it configures its own endpoint the way every server does; the TypeScript library holds no key and executes no paid model. -Without both a base URL and a key, `judge()` fails with `llm_not_configured` instead of calling an unintended endpoint. +The CLI resolves provider settings from environment variables. +An `OPENAI_API_KEY` or `TANGLE_API_KEY` can select that provider's default endpoint. +Set `AGENT_EVAL_LLM_BASE_URL` and `AGENT_EVAL_LLM_API_KEY` to make the route explicit. +TypeScript callers supply their transport, endpoint, and credentials directly. +Without a resolved endpoint and credential, `judge()` fails with `llm_not_configured`. ## Judge Content @@ -102,11 +104,11 @@ for rubric in Client().list_rubrics().rubrics: ## Client Options ```python -Client( - base_url: str | None = None, - cli_path: str | None = None, - transport: "auto" | "http" | "subprocess" = "auto", - timeout_s: float = 200.0, +client = Client( + base_url="http://127.0.0.1:5005", + cli_path="agent-eval", + transport="auto", + timeout_s=200.0, ) ``` @@ -178,10 +180,11 @@ python -m pip install \ "wandb" ``` -From an Agent Eval source checkout: +From `clients/python` in an Agent Eval source checkout, choose the required environment: ```sh uv sync --frozen --group gepa-release +# For source-only engines or compositions, use this instead: uv sync --frozen --group gepa-source ``` @@ -214,7 +217,7 @@ python -m pip install \ "skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90" ``` -From an Agent Eval source checkout, install the locked package with: +From `clients/python` in an Agent Eval source checkout, install the locked package with: ```sh uv sync --frozen --group skillopt-source @@ -248,6 +251,8 @@ python -m pip install "agent-eval-rpc[dspy]" ``` ```python +import os + import dspy from agent_eval_rpc import DspyJudgeMetric @@ -257,7 +262,7 @@ metric = DspyJudgeMetric(rubric_name="answer-quality") gepa = dspy.GEPA( metric=metric.feedback, - reflection_lm=dspy.LM("openai/gpt-4.1-mini"), + reflection_lm=dspy.LM(os.environ["DSPY_REFLECTION_MODEL"]), max_metric_calls=100, ) optimized = gepa.compile(program, trainset=train, valset=selection) @@ -265,6 +270,7 @@ optimized = gepa.compile(program, trainset=train, valset=selection) mipro = dspy.MIPROv2(metric=metric, auto="light") ``` +Set `DSPY_REFLECTION_MODEL` to your configured DSPy model identifier, including its provider prefix. Use `metric.feedback` for `dspy.GEPA`. It returns `dspy.Prediction(score=..., feedback=...)` with dimension scores, failure modes, wins, and rationale. Use the metric object directly for MIPROv2, SIMBA, bootstrap, and evaluation APIs that expect a number. @@ -301,17 +307,12 @@ import { analyzeTraces } from '@tangle-network/agent-eval/traces' type ModelOwner = Pick< DspyRlmTraceEngineOptions, - 'call' | 'callRef' | 'recordExecution' + 'call' | 'callRef' | 'recordExecution' | 'model' | 'pricing' > export async function analyzeRun(modelOwner: ModelOwner) { const engine = createDspyRlmTraceEngine({ ...modelOwner, - model: 'deepseek-v4-flash', - pricing: { - inputUsdPerMillion: 3, - outputUsdPerMillion: 15, - }, runner: { command: '.venv/bin/python' }, }) @@ -324,21 +325,9 @@ export async function analyzeRun(modelOwner: ModelOwner) { See [Trace Analysis](../../docs/trace-analysis.md) for custom definitions, limits, result fields, and the public quality benchmark. -DSPy 3.2.1 pins GEPA 0.0.27. -The general Optimize Anything bridge uses GEPA 0.1.4, so repository checks install them in separate environments: - -```sh -uv sync --frozen --extra dev --group gepa-release -AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \ - uv run --frozen --extra dev --group gepa-release \ - pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py - -uv sync --frozen --extra dev --group skillopt-source --group gepa-source -uv run --frozen pytest - -uv sync --frozen --extra dev --extra dspy -uv run --frozen pytest tests/test_dspy_metric.py -``` +The caller supplies the model identifier and endpoint rates with its execution callbacks. +DSPy and the Optimize Anything bridge require different GEPA versions. +The [development commands](#development) select each locked environment separately. The bridge records the installed upstream package version and source revision with each run. SkillOpt and a direct GEPA engine can restore official state only when the package revision, settings, starting candidate, described data, evaluation ID, and seed match. @@ -372,19 +361,29 @@ print(version.version, version.wire_version) ## Development +From the repository root, build Node before running cross-language tests: + ```sh +pnpm install --frozen-lockfile +pnpm build cd clients/python -pip install -e ".[dev]" -pytest ``` -Run the cross-language tests after building the Node package: +Run each compatibility suite with its locked dependencies: ```sh -cd ../.. -pnpm build -cd clients/python -pytest +uv sync --frozen --extra dev --group gepa-release +AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \ + uv run --frozen --extra dev --group gepa-release \ + pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py + +uv sync --frozen --extra dev --group skillopt-source --group gepa-source +uv run --frozen --extra dev --group skillopt-source --group gepa-source pytest + +uv sync --frozen --extra dev --extra dspy +uv run --frozen --extra dev --extra dspy pytest tests/test_dspy_metric.py ``` +Keep the same extras and groups on `uv sync` and `uv run`. +Each `uv sync` switches the local environment to that optimizer's required dependency set. The runnable Python example is [`examples/judge_anti_slop.py`](./examples/judge_anti_slop.py). diff --git a/docs/campaign-proposers.md b/docs/campaign-proposers.md index 26c7d3a64..5f2c3f7d3 100644 --- a/docs/campaign-proposers.md +++ b/docs/campaign-proposers.md @@ -15,9 +15,9 @@ That would split its search state from its own selection behavior and make budge An `OptimizationMethod` plugs into exactly two entry points. `selfImprove()` from `/contract` is the improvement entry: one call gives the method disjoint train and selection partitions, re-scores the selected surface on a held-out split, and returns a `gateDecision`. -Use it when one surface must get better. +Use it to search for a better surface and inspect the final decision. `compareOptimizationMethods()` from `/campaign` is the measurement entry: it gives every method equal inputs and scores the selected surfaces on final cases no method received. -Use it when two or more methods must be compared at equal budget. +Use it to compare selected surfaces under declared resource limits. Runnable versions: [`examples/self-improve-optimizer`](../examples/self-improve-optimizer/) and [`examples/compare-optimization-methods`](../examples/compare-optimization-methods/). ## Compose searches over a candidate @@ -75,6 +75,8 @@ Any parent receipt supplied by a custom method is also verified; it cannot repla `selfImprove({ method })` executes the complete method once and measures its selected surface on final cases. The method may select the unchanged baseline; that result returns `gateDecision: 'hold'` and an empty diff. +`winner` means the optimizer's selection, which can score worse on final cases. +Inspect `gateDecision` and its contributions before treating the selected surface as an improvement. Agent Eval does not score train and selection cases again or choose a different surface after the method finishes. The result type has two modes: @@ -216,8 +218,10 @@ Agent Eval adds the run ID, evaluation count, artifact directory, source identit - seed, - campaign defaults. -An optimization method never receives the final test cases. +The method input contains train and selection cases, without final test cases. After every method finishes, Agent Eval scores the selected surfaces on the same final cases and reports paired lift estimates. +The host must also exclude final cases from callback closures, shared files, and prior optimizer state. +The API partition does not provide process or filesystem isolation. ```ts import { @@ -227,7 +231,7 @@ import { } from '@tangle-network/agent-eval/campaign' const optimizer = { - model: 'gpt-4.1-mini', + model: optimizerModelId, // Supplied by the package that owns execution. Discovery derives these // from Runtime and one exact AgentProfile. call: optimizerExecution.call, @@ -238,10 +242,7 @@ const optimizer = { maxRequestBytes: 2_000_000, maxResponseBytes: 2_000_000, maxOutputTokensPerRequest: 32_768, - pricing: { - inputUsdPerMillion: Number(process.env.OPTIMIZER_INPUT_USD_PER_MILLION), - outputUsdPerMillion: Number(process.env.OPTIMIZER_OUTPUT_USD_PER_MILLION), - }, + pricing: optimizerTokenPricing, }, } @@ -253,7 +254,7 @@ const gepa = gepaOptimizationMethod({ kind: 'engine', run: { engine: 'gepa', - maxEvaluations: 40, + maxEvaluations: 80, maxProposerCostUsd: 5, }, }, @@ -293,54 +294,48 @@ const comparison = await compareOptimizationMethods({ }) ``` +These snippets assume caller-defined cases, dispatch, judges, `optimizerExecution`, model ID, and current token pricing. +They illustrate configuration; the linked examples provide complete scripts. +Both methods above declare the same evaluation ceiling. +Their actual evaluations, model calls, and spend can differ. + `costCeiling` is one limit shared by optimizer-model calls, train and selection evaluations, and final test scoring. +It applies to calls admitted through the cost ledger. +Leave enough capacity for every method and the final measurements. `comparison.scores` contains the final-case baseline score, selected score, lift, simultaneous interval, cost status, duration, and selected surface for each method. Official method scores contain optimizer and bridge package versions, source revisions and source-tree hashes, Python runtime, custom engine module hashes, compatible run ID, exact attempt ID, resume status, evaluation count, artifact directory, and available optimizer token usage. `comparison.pairwise` compares the highest-ranked method with every other method. -Ranking follows estimated lift, so inspect intervals before claiming a difference. +Ranking follows estimated lift. +`best` can therefore name a method whose improvement is unresolved. +Read each score's `decision` and the pairwise `favored` value before claiming a difference. +Intervals account for the method-versus-baseline contrasts and every possible method pair using a Bonferroni confidence adjustment. +The reported pairwise list contains only the observed best versus the alternatives. + +By default, final replicates are averaged within scenarios before inference. +A `claim` can group scenarios into declared independent units and set `minimumEffect`. +Read `unitScores`, `scenarioScores`, `units`, and `pairedCellN` together to retain both units and raw observation counts. +Optional `finalEvidence` records fresh final-case exposure across calls sharing its ledger. +See [evaluation integrity](./evaluation-integrity.md) for reusable claims and their boundaries. The runnable version is in [`examples/compare-optimization-methods`](../examples/compare-optimization-methods/). ## Install Official GEPA -Install the bridge and the published GEPA package for the standard engine: - -```sh -python -m pip install agent-eval-rpc -python -m pip install \ - "gepa==0.1.4" \ - "litellm>=1.83.0,<1.92" \ - "tqdm>=4.66.1" \ - "cloudpickle>=3.0.0" \ - "datasets>=2.14.6" \ - "wandb" -``` - -Do not install `gepa[full]`; its MLflow server dependency is unpatched. - -Use the tested official source revision for composed recipes and the source-only engines: - -```sh -python -m pip install \ - "gepa @ git+https://github.com/gepa-ai/gepa.git@f919db0a622e2e9f9204779b81fe00cc1b2d808f" \ - "litellm>=1.83.0,<1.92" \ - "tqdm>=4.66.1" \ - "cloudpickle>=3.0.0" \ - "datasets>=2.14.6" \ - "wandb" -``` - -From this repository: +From the repository root, install the bridge and locked standard-engine dependencies: ```sh cd clients/python uv sync --frozen --group gepa-release -uv sync --frozen --group gepa-source +cd ../.. +export OPTIMIZER_PYTHON="$PWD/clients/python/.venv/bin/python" ``` -The published package supports the standard `gepa` engine. -The composed recipes below — `sequential`, `adaptive-sequential`, `best-of`, `vote`, and `omni` — need the tested official source revision. -Move that revision only after both the release and the source compatibility tests pass. +Use `--group gepa-source` instead of `--group gepa-release` for composed recipes and source-only engines. +Those groups select different GEPA implementations and cannot coexist in one environment. +Pass the Python executable as `runner.command` when configuring a method directly. +The runnable examples read `OPTIMIZER_PYTHON` for that setting. +See the [Python GEPA guide](../clients/python/README.md#gepa) for other installation paths and compatibility checks. +Keep dependency revisions in the Python manifest and lock rather than copying them into integration code. ## Configure GEPA @@ -393,7 +388,7 @@ const method = gepaOptimizationMethod({ }, }, optimizer: { - model: 'gpt-4.1-mini', + model: optimizerModelId, call: optimizerExecution.call, callRef: optimizerExecution.callRef, budget: { @@ -402,10 +397,7 @@ const method = gepaOptimizationMethod({ maxRequestBytes: 2_000_000, maxResponseBytes: 2_000_000, maxOutputTokensPerRequest: 32_768, - pricing: { - inputUsdPerMillion: 0.4, - outputUsdPerMillion: 1.6, - }, + pricing: optimizerTokenPricing, }, }, describeScenario: (scenario) => ({ input: scenario.input }), @@ -413,43 +405,28 @@ const method = gepaOptimizationMethod({ }) ``` -Replace the rates with the exact rates charged by your endpoint. +`optimizerTokenPricing` must contain the current input and output USD rates per million tokens for the selected endpoint. If billed USD is unknown, omit `maxCostUsd`, `pricing`, and `maxProposerCostUsd`; the recorded cost remains unknown rather than becoming a guessed zero. With `optimizer`, every recipe stage must use the standard `gepa` engine or a metered agent CLI engine (below). -Agent Eval receives no provider key, enforces the declared request and token budget, and records the execution owner's exact usage and opaque finite JSON evidence. +The optimizer proxy receives no provider key. +It enforces the declared request and token limits and records the owner's usage and opaque finite JSON evidence. `maxProposerCostUsd` also limits each individual GEPA engine stage. -`optimizer.call` is always caller code. -Agent Eval owns no model transport and never receives a provider credential. - -On agent-runtime, use `profileOptimizerModelCall`, which executes one exact `AgentProfile` and reports profile-digest evidence: - -```ts -import { profileOptimizerModelCall } from '@tangle-network/agent-runtime/kernel' - -const call = profileOptimizerModelCall({ - profile: optimizerProfile, - context: 'prompt optimizer', - executor: { - backend: 'router', - routerBaseUrl: process.env.LLM_BASE_URL!, - routerKey: process.env.LLM_API_KEY!, - }, - pricing: { inputUsdPerMillion: 0.4, outputUsdPerMillion: 1.6 }, -}) -``` +`optimizer.call` supplies the model transport for this bridge. +Its execution owner holds the provider credentials. -Without agent-runtime, implement `ExternalOptimizerModelCall` over the OpenAI-compatible client you already have. -`examples/_shared/openai-compatible-owner.ts` is a complete minimal implementation to copy. -The callback resolves with one success or failure result and never rejects, because a rejection loses the execution record and fails the optimizer attempt. -The credential stays in your process; the proxy still enforces every budget and identity check. +For agent-runtime, use its maintained `profileOptimizerModelCall` adapter for the selected `AgentProfile`. +Keep runtime configuration in the execution-owning package. +For a direct endpoint, adapt [the example execution owner](../examples/_shared/openai-compatible-owner.ts) to your transport. +It implements `ExternalOptimizerModelCall` and returns a typed success or failure with a receipt and execution evidence. +The callback must resolve with that outcome; rejection loses the execution record and fails the optimizer attempt. +The optimizer proxy enforces its declared model limits around the supplied callback. ### Metered agent CLI engines The `autoresearch` and `meta_harness` engines drive a `claude` CLI subprocess. They ship only in the tested official source revision, not in the published `gepa` package (see [Install Official GEPA](#install-official-gepa)). Set `optimizer.anthropicEndpoint: true` to admit them in proxied mode. -This path is measured live: a real `claude` CLI session completes with every tool call translated and every call metered. The loopback proxy then also serves `POST /v1/messages` (Anthropic Messages API) and the bridge child receives `ANTHROPIC_BASE_URL`, an ephemeral `ANTHROPIC_AUTH_TOKEN`, and `ANTHROPIC_MODEL` in its environment. Every CLI call becomes one canonical execution-owner call with the same reservation, receipt, and budget pipeline as reflection traffic; the run fails if the receipt count differs from the admitted call count. Each agent engine run must set `engineConfig.model` to `optimizer.model`, because the engines pass `--model` and that flag beats the injected environment. @@ -486,14 +463,14 @@ recipe: { Its external model spend remains incomplete unless that engine reports it. Supply provider API keys only through `runner.env`. -An exported shell variable never reaches the bridge child. +The child does not inherit exported provider credentials automatically. The spawn builds the child environment from a fixed allowlist of benign variables (PATH, HOME, locale, `PYTHONPATH`) plus `runner.env`, so the parent environment is stripped by construction. When `optimizer` is set, `removeCredentialEnvironment` also deletes credential-shaped keys from `runner.env`; the child then receives only the loopback proxy URL and an ephemeral key inside the input JSON. Do not place credentials in `engineConfig` because run settings are persisted. `describeScenario()` controls the train and selection data sent to GEPA. `describeArtifact()` controls the execution evidence returned after a candidate is scored. -Neither callback can receive a final test case. +Agent Eval calls these callbacks only for train and selection cases; caller-owned context must respect the same boundary. A direct standard GEPA run records `provenance.gepaCandidatePopulation`. Pass that summary to `readGepaCandidatePopulationArtifact()` to verify and read every accepted candidate, its parent indices, and its selection scores. @@ -501,42 +478,42 @@ Use `readExternalOptimizerObservationArtifact()` for every distinct callback sub `provenance.evaluationCount` is the callback-metered evaluation total. `provenance.upstreamReportedEvaluations` is GEPA's self-reported total; a difference means upstream skipped, cached, or double-counted work. -## Runtime Knobs - -Each knob below has a default that works for small text campaigns and fails for agentic or slow-settling runs. -The table names the failure so you can set the knob before the run dies mid-spend. +## Runtime controls -| Knob | Default | What breaks when wrong | Where to set it | -|---|---|---|---| -| `timeoutMs` | 30 minutes | The whole bridge process tree is killed with `GEPA bridge exceeded 1800000ms`. An agentic run (40 evaluations over 45 s each) exceeds the default mid-spend. The same value bounds each callback POST and the runtime inspect pass. | `gepaOptimizationMethod({ timeoutMs })`, `skillOptOptimizationMethod({ timeoutMs })` | -| `dispatchShutdownTimeoutMs` | 5 seconds | A dispatch that cancels or settles paid calls slowly fails the cell with `CostAccountingIncompleteError` after the evaluations completed. | `runCampaign({ dispatchShutdownTimeoutMs })`; for comparisons, `compareOptimizationMethods` `optimizationRunOptions` | -| `servedModelPolicy` | `'exact'` | A router that substitutes a same-family model fails every proxied reflection call with a 502 `model substitution` error. `'allow-within-family'` accepts the substitute, keeps family-level claims, and forfeits per-model claims. | `optimizer.servedModelPolicy` | -| `reflection_lm_kwargs.num_retries` | litellm default (3) | Each failed reflection request retries 3 times inside litellm, so the proxy meters 4 request attempts per logical call and `budget.maxRequests` exhausts 4x early. Set `num_retries: 0`; the proxy already accounts each attempt. | `recipe.run.engineConfig.reflection.reflection_lm_kwargs` | -| `reflection_lm_kwargs.max_tokens` | `budget.maxOutputTokensPerRequest` | Every reflection request ships the full budget cap as `max_tokens`. A provider family with a lower completion cap rejects every call. A reasoning model also needs headroom for hidden reasoning tokens. Set a value at or below the smallest family cap; it must not exceed `budget.maxOutputTokensPerRequest`. | `recipe.run.engineConfig.reflection.reflection_lm_kwargs` | -| `maxProposerCostUsd` | unset | Without it, one engine stage can spend up to `optimizer.budget.maxCostUsd` or the campaign `costCeiling` before any limit fires. Supply it only when the execution owner can enforce billed USD. | `recipe.run.maxProposerCostUsd` | -| `maxEvaluations` (agent engines) | required, no default | An agent engine registers one aggregate evaluation that costs the full train set of callback evaluations. A value below the train-set size rejects mid-aggregate and GEPA records the candidate as `-inf`. | `recipe.run.maxEvaluations` | -| `budget.maxRequests` (agent engines) | required, no default | An agent CLI session makes tens of calls per engine run. A text-campaign-sized limit exhausts mid-run, and the CLI sees a terminal 402. | `optimizer.budget.maxRequests` | -| `expectUsage` | `'assert'` | A deterministic evaluator that makes no LLM calls records zero usage, so `'assert'` fails the run as a stub. Set `'off'` only for an evaluator with no paid calls. | `selfImprove({ expectUsage })` | +Set limits for the execution path you actually use. +Long-running dispatches and agent CLI engines can need different limits from short text evaluations. -## Install Official SkillOpt +| Setting | Default | When to change it | +|---|---|---| +| Method `timeoutMs` | 30 minutes | Bound the entire bridge run, including slow evaluations and checkpointing. | +| Campaign `dispatchShutdownTimeoutMs` | 5 seconds | Allow pending paid calls to settle after dispatch cancellation. | +| `optimizer.servedModelPolicy` | `exact` | Use `allow-within-family` only when substitutions within a model family are acceptable for the claim. | +| `reflection_lm_kwargs.num_retries` | Upstream setting | Set explicitly when bounding retry attempts in the GEPA reflection configuration. | +| `reflection_lm_kwargs.max_tokens` | Optimizer output cap | Use a limit supported by the endpoint and sufficient for the model's reasoning and output. | +| `recipe.run.maxProposerCostUsd` | Unset | Bound a GEPA stage separately when dollar accounting is available. | +| `recipe.run.maxEvaluations` | Required | Allow enough candidate-case calls for each intended aggregate evaluation. | +| `optimizer.budget.maxRequests` | Required | Bound optimizer calls; agent CLI sessions can use many calls per stage. | +| `selfImprove({ expectUsage })` | `assert` | Set `off` only for deterministic evaluation with no paid calls. | -Install the SkillOpt source revision tested by this release: +Put reflection settings under `recipe.run.engineConfig.reflection.reflection_lm_kwargs` for a direct engine recipe. +Model substitutions remain recorded; accepting one does not establish performance of the originally requested model. +Request limits count calls admitted to the execution owner. +The owner must report its internal retries and enforce their declared bounds. -```sh -python -m pip install agent-eval-rpc -python -m pip install \ - "skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90" -``` +## Install Official SkillOpt -From this repository: +From the repository root: ```sh cd clients/python uv sync --frozen --group skillopt-source +cd ../.. +export OPTIMIZER_PYTHON="$PWD/clients/python/.venv/bin/python" ``` -The published `skillopt==0.2.0` wheel omits the prompt files required by `ReflACTTrainer`. -The tested source revision contains all 21 files. +Add `--group gepa-source` to the same sync command when comparing both methods. +Use the source group selected by the lock; the bridge's compatibility checks cover that implementation. +See the [Python SkillOpt guide](../clients/python/README.md#skillopt) for package requirements and validation. `skillOptOptimizationMethod()` runs SkillOpt's official `ReflACTTrainer`. Agent Eval supplies an environment adapter that sends each candidate and case back to the TypeScript execution and judging path. @@ -554,28 +531,10 @@ Missing token usage, an oversized request or response, a wrong model, streaming, ## Use Official DSPy Optimizers -Do not convert a DSPy program into an `OptimizationMethod`. -Install `agent-eval-rpc[dspy]`, create `DspyJudgeMetric`, and pass it to official DSPy: - -```python -import dspy - -from agent_eval_rpc import DspyJudgeMetric - -dspy.configure_cache(restrict_pickle=True) -metric = DspyJudgeMetric(rubric_name="answer-quality") -gepa = dspy.GEPA( - metric=metric.feedback, - reflection_lm=dspy.LM("openai/gpt-4.1-mini"), - max_metric_calls=100, -) -mipro = dspy.MIPROv2(metric=metric, auto="light") -``` - -This keeps program compilation, traces, demos, and optimizer state inside DSPy. -Agent Eval supplies the shared rubric and returns rich feedback for `dspy.GEPA`. -DSPy 3.2.1 requires GEPA 0.0.27. -Run it in a separate Python environment from the general GEPA bridge, which uses GEPA 0.1.4. +Keep DSPy programs and their optimizer state inside DSPy. +`DspyJudgeMetric` supplies Agent Eval rubric scores and feedback to official DSPy optimizers. +Configure the judging client and use the [Python DSPy guide](../clients/python/README.md#dspy) for installation and examples. +Use a separate environment when its GEPA dependency conflicts with the general optimizer bridge. ## Resume A Compatible Run @@ -637,9 +596,10 @@ const proposer: SurfaceProposer = { Return a label and rationale when they will help later analysis. Candidate creation must not read final test results. -A proposer may also attach `attribution`: one opaque, JSON-safe record the loop carries byte-for-byte onto `GenerationCandidate.attribution` and the loop provenance record. +A proposer may attach `attribution`: an opaque JSON-safe record retained on `GenerationCandidate.attribution` and in loop provenance. The loop never interprets it. -Tag it with your own schema field and validate it on readback — `PolicyEdit` candidate records (`makePolicyEditCandidateRecord` from the analyst surface) are the first producer, which is what lets a forecast in an edit be scored against the measured delta later. +Tag it with a schema field and validate it on readback. +`makePolicyEditCandidateRecord` from `/analyst` records an edit forecast that can later be compared with the measured change. `runOptimization()` rejects a candidate whose `surfaceHash` was already admitted. This includes the baseline, an earlier generation, and another candidate in the same proposal. @@ -675,9 +635,7 @@ const result = await runOptimization({ - Final test cases may only compare surfaces after every method finishes. - The same dispatch and judges score every method. - Missing cost remains unknown. -- A method must declare bounded work before it starts. -- Credentials reach a bridge child only through `runner.env`; the child never inherits the parent process environment. +- Bound method work before it starts, including any caller-owned operations outside the ledger. +- Bridge children inherit a small environment allowlist; pass unproxied provider credentials only through `runner.env`. - The metered proxy path replaces provider credentials with a loopback URL and an ephemeral key. - Resumed state must match every input that can change the result. - -These rules make method comparisons inspectable without pretending different optimizers have identical internals. diff --git a/docs/concepts.md b/docs/concepts.md index d0bae3e05..fb2d48a8e 100644 --- a/docs/concepts.md +++ b/docs/concepts.md @@ -34,8 +34,7 @@ Keep its five values distinct because they require different actions. | `model_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the model. | Handle it; no gate in this package emits it. | | `arch_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the architecture. | Handle it; no gate in this package emits it. | -The last two are part of the taxonomy and of the composition order, but no built-in gate returns them today. -Handle all five anyway: a caller's own gate may return either, and the type will not let you ignore them. +Handle all five values when accepting caller-supplied gates. Read the contributing checks before interpreting a refused release. An unresolved comparison does not establish that the candidate is worse. @@ -85,7 +84,7 @@ Traces, datasets, optimization, statistics, and reports build on these objects. Every entry in `GateResult.contributingGates` has a `status` of `pass`, `fail`, or `not_evaluated`. `pass` and `fail` mean the check ran with sufficient input. -`not_evaluated` means the check lacked enough evidence to run. +`not_evaluated` means the check was unconfigured or lacked required input or evidence. `defaultProductionGate` always requires held-out significance. Its other checks are optional until their input is configured or their name is included in `requiredChecks`. A required check with missing or insufficient evidence remains `not_evaluated` and holds the release decision. @@ -93,7 +92,10 @@ An absent optional check records `not_evaluated` and never appears as a successf Run history is shared input only. Enable reward-hacking and canary monitoring independently with `rewardHacking` and `canary`. -`paretoSignificanceGate()` applies each objective's regression floor to its deciding confidence interval. +`selfImprove()` uses `defaultProductionGate()` unless you pass a custom `gate`. +The optional `paretoSignificanceGate()` applies each objective's regression floor to its deciding confidence interval. +Its confidence level applies per objective; the gate does not adjust for multiple objectives or repeated comparisons. +It pairs execution cells directly; custom gates must apply any grouping required by a claim. For detected binary outcomes, that interval accounts for uncertainty even when every observed pair agrees. At 95% confidence, 20 matching all-positive binary pairs leave approximately 16 percentage points of uncertainty in either direction. A declared five-point regression tolerance therefore holds that candidate; 100 matching pairs narrow the interval enough to clear that floor. @@ -136,7 +138,7 @@ that can seed memory, replay scenarios, and optimization. | **Finding** | A specific issue a judge found: file, line, severity, message. | | **Trace store** | The append-only log of every span/event during a run. Replay = read this back. | | **Composite score** | An aggregate on the judge's declared scale. Gates must use thresholds on that scale. | -| **Rubric version** | A stable hash of the rubric. Scores from different rubric versions are not comparable. | +| **Rubric version** | A stable hash of the rubric. Comparing revisions requires calibration against shared independent labels. | ### Running an evaluation @@ -146,8 +148,8 @@ that can seed memory, replay scenarios, and optimization. | **Surface** | The value being changed: a prompt, a skill, or a serialized configuration. | | **Dispatch** | The function that runs your agent on one case and returns the artifact. | | **Campaign** | One complete pass of every case, executed, scored, and cached under a run directory. | -| **Cell** | One (case × replicate) of a campaign. Cells are cached, so a rerun skips the ones that finished. | -| **Receipt** | The record of what one paid call actually cost, in dollars and tokens. Absent when nothing measured it. | +| **Cell** | One (case × replicate) of a campaign. With caching enabled, matching completed cells can be reused. | +| **Receipt** | A settled call record with cost, token usage, and flags for unknown measurements. | | **Cost ledger** | The spend account receipts are written to. A capped ledger refuses a call that would exceed the cap. | | **Provenance** | Where a number came from: the package version, the source revision, the run identity, the exact attempt. | | **`RunRecord`** | The analysis-time projection of one run: who ran, on what, with which seed, at what cost, and what it scored. | @@ -194,7 +196,7 @@ A declared sampling frame does not itself establish representative sampling. Count repetitions separately from independent units. For example, 100 retries of one incident produce 100 observations and one independent incident. Campaign aggregate `n` describes its observed scores. -Registered gates report their independent-unit count and paired-cell count separately. +The default improvement gate and method comparisons retain their independent-unit and paired-cell counts. Set `minimumEffect` when the decision concerns a practically useful change. Development and absolute-rate claims can omit it. @@ -243,22 +245,18 @@ L1 app-build Does the artifact build / typecheck / test? │ ▼ L2 app-runtime Does the artifact actually run end-to-end? - (Dynamic signal: only worth checking if L1 passed.) + (Dynamic signal: requires a runnable application.) ``` `BuilderSession` coordinates these checks. It opens at `startChat`, runs the build at `ship`, and runs the application check at `runAppScenario`. Each layer emits a trace span. -`scoreProject` combines their measured scores. - -These layers detect different failures: +`scoreProject` reports each layer's score and whether the required measurements are complete. - L0: The agent crashed during generation and left an incomplete artifact. - L1: Files exist but do not typecheck or build. - L2: Code compiles but behaves incorrectly when executed. -If you only check one layer, you ship the bugs that the other two layers would have caught. - ## How rubrics work A rubric describes: @@ -290,7 +288,7 @@ const verifier = new MultiLayerVerifier([ ]) const report = await verifier.run({ env }) -report.allPass // boolean: every layer passed +report.allPass // every layer passed and the task measurement is complete report.taskScore // complete task score, or undefined report.blendedScore // diagnostic weighted aggregate, possibly partial report.layers // per-layer status, findings, duration @@ -299,7 +297,11 @@ report.layers // per-layer status, findings, duration `env` carries the sandbox driver, the working directory, and the harness commands each layer runs. Use `taskScore` when creating task labels or training data. -An errored, timed-out, skipped, or incomplete scoring panel leaves `taskScore` undefined. +A complete scoring panel needs at least one valid score from a positive-weight layer. +Positive-weight layers that error, time out, or skip leave `taskScore` undefined. +Passing layers can omit a numeric score. +A failed layer contributes only when `failContributesToScore` is enabled and it supplies a valid score. +Zero-weight layers can still fail `allPass` without removing `taskScore`. Use `blendedScore` only to inspect the measurements that did complete. Two rules that will save you bugs: @@ -311,9 +313,9 @@ Two rules that will save you bugs: ## Judge calibration -Two questions to answer before trusting any LLM judge: +Compare a judge with independent labels and other judges before using its scores for decisions: -1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and worst-N miscalibrations vs a human golden set. +1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and the five largest errors on matched item IDs. 2. **Does it agree with other judges?** `continuousAgreement()` and `calibrateJudgeContinuous()` report agreement and bootstrap intervals on continuous scores. @@ -325,7 +327,7 @@ Each statistic answers a different question: | Spearman | Do they rank the same way? | The size of any gap | | MAE (mean absolute error) | How far apart are they, on average? | Whether the gap is systematic | | κ (Cohen's kappa) | Do they agree more than chance? | Everything below the rounding step | -| ICC(2,1) | Do they agree in absolute value, not just in shape? | — | +| ICC(2,1) | Do raters agree in absolute score under its variance model? | Shared errors against the intended outcome | Use two flavours of κ for one reason. `calibrateJudge` rounds each score to an integer first. @@ -336,8 +338,8 @@ ICC(2,1) catches a bias Pearson cannot see. If judge B always scores twice judge A, the two move together perfectly and Pearson stays near 1, while ICC drops. That drop is the signal. -These agreement intervals use bootstrap resampling. -The middle 95% of the recomputed statistics forms each reported interval. +ICC and continuous weighted κ have percentile bootstrap intervals, with `ciLevel: 0.95` by default. +Pearson, Spearman, and MAE are point estimates in these reports. Import calibration and bias functions from `/meta-eval`. diff --git a/docs/design.md b/docs/design.md index e8669f95a..97fc38657 100644 --- a/docs/design.md +++ b/docs/design.md @@ -31,7 +31,7 @@ The stack context matters only if you adopt more of it later. ## The dependency rule This section is the public rationale. -The enforceable maintainer rule lives in [`CLAUDE.md`](../CLAUDE.md#repo-layering--this-package-is-the-substrate). +The enforceable maintainer rule lives in [`CLAUDE.md`](../CLAUDE.md#dependency-and-evidence-boundaries). **`agent-eval` has zero upward dependencies on a consumer.** This is what keeps the package reusable outside our own stack: nothing in here imports from `agent-runtime`, `agent-knowledge`, or `sandbox`, whether at runtime, in development dependencies, or as type-only imports. diff --git a/docs/eval-surface-map.md b/docs/eval-surface-map.md index 76d9344c7..006cde789 100644 --- a/docs/eval-surface-map.md +++ b/docs/eval-surface-map.md @@ -8,14 +8,14 @@ This reference covers direct execution controls and specialist modules. | Primitive | Import | Use | Returns | |---|---|---|---| -| `runCampaign()` | `/campaign` | Execute and judge a scenarios × repetitions grid through caller-owned dispatch. | `CampaignResult` | -| `runEval()` | `/contract` or `/campaign` | Score one surface with campaign defaults. | `CampaignResult` | -| `runProfileMatrix()` | `/campaign` | Run the same cases across named agent profiles with provenance and backend checks. | `RunProfileMatrixResult`, including `.records`. | -| `runOptimization()` | `/campaign` | Generate, measure, and select candidates on development cases. | Generations and a winner surface. | -| `runImprovementLoop()` | `/contract` or `/campaign` | Search, compare on final cases, and apply a release gate. | Final comparison, winner, and gate decision. | -| `compareOptimizationMethods()` | `/campaign` | Compare selected surfaces from several methods under declared budgets. | Final paired contrasts, uncertainty, and costs. | -| `runEvalCampaign()` | Root | Run with a caller-supplied trace sink and emitter. | Campaign result and records. | -| `runAgentMatrix()` | `/matrix` | Schedule a general Cartesian grid without campaign scoring semantics. | Cell results. | +| [`runCampaign()`](../examples/plan-before-you-spend/) | `/campaign` | Execute and judge a scenarios × repetitions grid through caller-owned dispatch. | `CampaignResult` | +| [`runEval()`](../src/campaign/presets/run-eval.ts) | `/contract` or `/campaign` | Score one surface with campaign defaults. | `CampaignResult` | +| [`runProfileMatrix()`](../examples/profile-matrix/) | `/campaign` | Run the same cases across named agent profiles with provenance and backend checks. | `RunProfileMatrixResult`, including `.records`. | +| [`runOptimization()`](./campaign-proposers.md#write-a-custom-candidate-generator) | `/campaign` | Generate, measure, and select candidates on development cases. | Generations and a winner surface. | +| [`runImprovementLoop()`](./multi-shot-optimization.md) | `/contract` or `/campaign` | Search, compare on final cases, and apply a release gate. | Final comparison, winner, and gate decision. | +| [`compareOptimizationMethods()`](../examples/compare-optimization-methods/) | `/campaign` | Compare selected surfaces from several methods under declared budgets. | Final paired contrasts, uncertainty, and costs. | +| [`runEvalCampaign()`](../src/eval-campaign.ts) | Root | Run with a caller-supplied trace sink and emitter. | Campaign result and records. | +| [`runAgentMatrix()`](../src/matrix/runner.ts) | `/matrix` | Schedule a general Cartesian grid without campaign scoring semantics. | Cell results. | When variants of the same task run inside one `runCampaign`, give those scenarios the same `seedGroup` so each repetition uses common randomness. Use `runProfileMatrix` instead when profiles are separate campaign axes. @@ -47,6 +47,7 @@ Coverage reports missing, failed, and zero-score rows separately. | Group reusable comparisons by source unit | `/contract` or `/campaign` | The top-level `claim` option. | | Reserve fresh evidence for confirmation | `/contract` or `/campaign` | Optional `finalEvidence: { ledger, requestId, evaluatorDigest }`. | | Admit an evaluator against both error limits | `/meta-eval` | `auditEvaluator()` | +| Test a grader with known incorrect items | `/meta-eval` | [`definePlant()`, `seedPlants()`, `catchRate()`](./plants.md) | | Measure agreement and known bias patterns | `/meta-eval` | `calibrateJudgeContinuous()`, `continuousAgreement()`, `positionalBias()`, `verbosityBias()`, `selfPreference()` | | Relate scores to declared deployment outcomes | `/meta-eval` | `rubricPredictiveValidity()`, `correlationStudy()`, `calibrationFromPairs()`, `calibrationCurve()` | @@ -73,20 +74,22 @@ Use the decision diagnostics and intervals to distinguish insufficient evidence | Subpath | Use | |---|---| -| `/traces`, `/trace-attributes` | Store and inspect trace evidence; use canonical measurement attribute names. | +| `/traces`, `/trace-attributes` | Store trace evidence, [connect observability exporters](./adapters-observability.md), and use canonical measurement attribute names. | | `/analyst` | Execute declared analysts against recorded evidence. | -| `/reporting`, `/pipelines` | Compare runs, render reports, and extract recorded failure patterns. | +| `/reporting`, `/pipelines` | Compare runs, render [research reports](./research-report-methodology.md), and extract recorded failure patterns. | | `/supervisor-run` | Read recursive run directories and their evidence coverage. | | `/trace-repair`, `/trajectory-replay` | Execute proposed repairs or replay recorded shell trajectories. | | `/benchmarks`, `/fuzz` | Adapt benchmark data and explore a declared behavior space. | -| `/builder-eval`, `/multishot`, `/multishot/golden` | Evaluate generated applications and multi-turn conversations. | +| `/builder-eval`, `/multishot`, [`/multishot/golden`](./multishot-golden-records.md) | Evaluate generated applications and multi-turn conversations. | | `/matrix` | Schedule Cartesian experiment grids. | | `/rl` | Build reward, preference, and supervised datasets from eligible evidence. | | `/profile-cell` | Create and validate portable agent-profile identities. | | `/authenticity`, `/ledger-core` | Check evidence authenticity and maintain canonical hash-chained journals. | -| `/rollout`, `/storyboard` | Represent rollout trees and render recorded work. | +| [`/rollout`](./rollout.md), `/storyboard` | Serialize training rows and render recorded work. | | `/hosted`, `/wire`, `/adapters/http` | Connect hosted storage or expose evaluation through HTTP and RPC. | +For existing coding-agent transcripts, start with [session intake](./code-agent-intake.md). + Root `Scenario`, `JudgeScore`, and `GateDecision` match `/contract`. Use root `ProductScenario` and `DimensionJudgeScore` for the product-judging functions. `HeldOutGate.evaluate()` returns the separate root type `HeldOutGateDecision`. @@ -111,7 +114,7 @@ A judge that produced no score has no entry at all. An absent aggregate is the honest record of an unmeasured judge, and a zero-filled distribution would read as a measured all-zero series. `SeriesDistribution` is the one distribution summary in this package. -It is not the `ScalarDistribution` the insight report uses; see `insight-report.md` for why those two shapes stay separate. +The [insight report](./insight-report.md) explains why its `ScalarDistribution` has a separate shape. ## Planning the cell grid without a run directory @@ -123,6 +126,7 @@ Scenarios that share a `seedGroup` receive the same per-replicate seeds, which i Use `planCampaignRun()` to classify cached, pending, and blocked cells. That call reads the durable cache in a real run directory. `cellDirectory` and `cellCachePath` name a cell's location once a run directory is chosen. +Use [eval fixtures](./eval-fixtures.md) to load and fingerprint cases from folders before planning their campaign. ## Evidence receipts: `attest` @@ -160,6 +164,7 @@ runtime events -> extractProducedState(events) -> ProducedState The host supplies the correctness checker and the events. This composition shares campaign execution, capture, and reporting without another runner. +See [product patterns](./product-eval-adoption.md#product-patterns) for host adapters and [knowledge readiness](./knowledge-readiness.md) for required context checks. ### The in-band body contract diff --git a/docs/evaluation-integrity.md b/docs/evaluation-integrity.md index 4be3c255f..af9970f6a 100644 --- a/docs/evaluation-integrity.md +++ b/docs/evaluation-integrity.md @@ -9,7 +9,7 @@ The package separates three questions: | Question | Evidence | Public entry | | --- | --- | --- | | Did this change help on these cases? | Paired scores, failures, cost, and case coverage. | `defineAgentEval()` or `selfImprove()` from `/contract`. | -| Does the improvement extend to new tasks? | Representative independent tasks, a declared effect, and an appropriate comparison. | Optional `claim` on campaign comparisons; registered rules from `/experiment`. | +| Does the improvement extend to new tasks? | Representative independent tasks and an appropriate comparison; a useful effect when making an improvement decision. | Optional `claim` on campaign comparisons; registered rules from `/experiment`. | | Can fresh final evidence support this adaptive decision? | A frozen comparison, retained access boundaries, and a durable exposure record. | Optional `finalEvidence` on the same comparison. | These controls reuse the existing execution path, paired estimators, sealed experiments, and locked journal. @@ -36,7 +36,8 @@ const claim = defineEvaluationClaim({ Pass `claim` to `selfImprove()`, `runImprovementLoop()`, or `compareOptimizationMethods()`. It remains independent of final-evidence storage. `minimumEffect` is optional because development reports and absolute-rate measurements need not test an improvement threshold. -When omitted, each existing comparison keeps its documented decision threshold. +When omitted, the default `selfImprove()` gate uses a 0.05 gain threshold; method comparison uses 0. +Custom gates choose their own thresholds. `independentUnit` names a field path in each scenario or evidence row. Variants from one incident must carry the same source identity. @@ -48,7 +49,8 @@ The default self-improvement gate averages paired cells within registered units. Method comparison first averages repetitions within each scenario, then averages scenarios within each source unit. It weights source units equally. Results retain `scenarioScores`, `unitScores`, `units`, and `pairedCellN` so callers can inspect each denominator. -Custom gates receive the same measured evidence and remain responsible for their own decision rules. +Custom improvement gates receive raw cell scores and scenarios. +Pass the claim into your gate's configuration and apply its grouping and decision rules there. `fixed-roster` describes the specified cases. Repeated executions can measure execution variability on that roster. @@ -60,7 +62,9 @@ Claim metadata records the intended scope; it does not authenticate sampling or There is no universal task count that proves an improvement. The effect, outcome type, dependence, confidence level, and decision procedure determine what the evidence supports. -Paired binary decisions use the shared score interval and exact discordance check. +Paired binary decisions use the shared score interval. +At nonnegative gain thresholds, the exact discordance check can also veto promotion. +Negative thresholds ask whether a regression stays within a tolerance and do not use that veto. A sufficiently large binary gain can pass with fewer than 20 independent pairs. Continuous mean decisions require the existing bootstrap path's 20-pair eligibility threshold. That implementation threshold does not establish adequate power or guarantee interval coverage for every distribution. @@ -71,8 +75,9 @@ Method rankings describe observed lift. Each score and pairwise contrast retains its full `decision`, including the estimator, threshold, minimum, and sufficiency. An inconclusive gate leaves the selected candidate available for further development or a narrower evaluation. -Use `clusteredPower()` or a registered `power-floor` gate before an expensive population-level comparison. -Both assess power at the declared `minimumEffect`. +For a design that matches its outcome model, `clusteredPower()` simulates power at the declared `minimumEffect`. +A registered `power-floor` gate checks a supplied power curve against its minimum effect and target power. +These optional design checks do not run automatically when you pass a `claim`. High power at a much larger effect cannot substitute for power at the improvement that matters. ## Opt into fresh final evidence @@ -143,7 +148,7 @@ const interval = { // Then execute registered.interval('lift', { kind: 'rows', rows }). ``` -New-unit claims reject intervals that resample a different field. +For new-unit claims, each cluster interval's `clusterBy` must equal `claim.independentUnit`. Registered binomial intervals require one unique `unitId` per trial for new-unit claims. Opened seals capture validated rules before asynchronous execution. Caller mutation cannot change the opened experiment's rules. @@ -157,6 +162,7 @@ See [registered experiments](./experiment.md) for the complete rule language. Use `auditEvaluator()` from `/meta-eval` when admitting a new checker or model judge. Provide actual judgments of independently verified good and bad controls. Each observation names its source unit, evidence reference, expected decision, observed decision, and development exposure. +This audit is separate from `selfImprove()`; the caller decides whether to require admission before search or release. The audit measures false acceptance and false rejection separately. A source unit fails a class when any variant in that class is misjudged. diff --git a/docs/experiment.md b/docs/experiment.md index d3e029262..ca2ca6871 100644 --- a/docs/experiment.md +++ b/docs/experiment.md @@ -1,31 +1,31 @@ # The experiment subpath -`@tangle-network/agent-eval/experiment` turns an experiment's registration into the object that runs it. +`@tangle-network/agent-eval/experiment` registers decision rules as data and provides interpreters bound to a verified seal. -The registry of measured claims those experiments produce lives in [`evidence/`](../evidence/README.md); a sealed experiment's digest is the `experimentDigest` its registry record carries. +The [`evidence/` registry](../evidence/README.md) stores published measurements and their experiment identities. +Sealing does not execute an agent or publish an evidence record. -## The covenant +## What a seal enforces -1. **The registered rule is the executed rule.** - Every rule — row admission, subset selection, estimand, interval, decision table, validity gate, halt, budget, matched budget, reissue — is a typed data node, never a closure or prose. - `sealExperiment` canonicalizes and hashes the whole tree into one digest. - Every interpreter takes only a sealed node plus evidence records. - No execution surface has a parameter for alpha, threshold, metric, or stopping rule, so registered-vs-ran drift is unrepresentable rather than checked. -2. **Refusals live inside artifacts.** - An inadequate cluster count, a mismatched arm budget, a non-monotone funnel stage, a non-total decision table — each produces a typed verdict object (or a typed error), never a warning sentence beside a number. -3. **A change is a re-seal.** - `amendExperiment` verifies the current seal, validates the new spec, and appends a `{at, reason, blind[], digest}` entry. - The digest history is the audit trail; changing what is decided without a new digest is impossible. +1. `sealExperiment()` validates and hashes the specification, including its registered rules and optional claim. +2. `openSealedExperiment()` verifies that digest and captures the rules before returning the execution handle. + Its decision, interval, gate, and budget methods read their rules from that captured specification. +3. `amendExperiment()` verifies the current seal, validates the replacement specification, and records its digest, reason, time, and declared blindness. + +A seal verifies the current specification's identity. +It does not authenticate registration time, amendment history, or the origin of supplied measurements. +The host must retain evidence, invoke the required checks, and honor their refusal results. +Calling `registered.decide()` does not automatically run admission, power, budget, or halt checks. ## The objects -| object | entry point | what it closes | +| Object | Entry point | Purpose | | --- | --- | --- | -| Registered-rule AST | `src/experiment/ast.ts` (14 node families) | prose rules; every registered condition compiles to data the seal covers | -| Define / seal / execute | `defineExperiment`, `sealExperiment`, `amendExperiment`, `openSealedExperiment` | hand-written PREREG.md files; the runner executes the sealed rule itself | -| Cluster-aware power | `clusteredPower`, `assertDesignAdequate` | "4 clusters cannot certify any effect size, including 1.0" — learned by running the experiment, now refused before a dollar is spent | -| Denominator chain | `buildFunnel`, `executeAdmissionRule`, `composeFunnels`, `renderFunnelTable` | hand-assembled `20 → 15 → 14 → admitted` chains, formatted differently each run | -| Matched budgets | `verifyMatchedBudgets`, `assertMatchedBudgets` | "realized tokens must agree within 5%" verified by hand | +| Registered rules | [Rule types](../src/experiment/ast.ts) and [`ExperimentSpec`](../src/experiment/define.ts) | Describe admission, estimation, intervals, and decisions as data. | +| Define / seal / execute | `defineExperiment`, `sealExperiment`, `amendExperiment`, `openSealedExperiment` | Validate, identify, and execute the registered rules. | +| Cluster-aware power | `clusteredPower`, `assertDesignAdequate` | Assess a declared effect under a simulated outcome model and cluster-count policy. | +| Denominator chain | `buildFunnel`, `executeAdmissionRule`, `composeFunnels`, `renderFunnelTable` | Reconcile retained and excluded evidence. | +| Matched budgets | `verifyMatchedBudgets`, `assertMatchedBudgets` | Check realized tokens against a declared tolerance. | ### Sealing and execution @@ -35,15 +35,19 @@ import { sealExperiment, } from '@tangle-network/agent-eval/experiment' -const sealed = await sealExperiment(spec) // RFC 8785 + sha256 over the whole tree -const registered = await openSealedExperiment(sealed) // verifies the digest first +const sealed = await sealExperiment(spec) +const registered = await openSealedExperiment(sealed) -const admission = registered.admit(rows) // funnel + survivors, from the sealed rule +const admission = registered.admit(rows) const gate = registered.gate('power-floor', { kind: 'power-floor', curve }) -const halt = registered.halt([gate]) // refuse-spend fires before any contrast -const outcome = registered.decide(quantities) // the sealed table; non-total tables throw +const halt = registered.halt([gate]) +if (halt.fired) throw new Error(`Experiment halted: ${halt.failedGates.join(', ')}`) +// Compute quantities from the admitted evidence, then call registered.decide(quantities). ``` +This fragment assumes `spec` registers admission, the named gate, and a halt rule. +Malformed specifications and unusable evidence throw typed errors; decision and validity refusals remain in returned artifacts. + Cluster intervals register both `clusterBy` and `value` inside the sealed `IntervalSpec`. Call `registered.interval('gain95', { kind: 'rows', rows })` to apply those fields. The row evidence cannot override the registered value field. @@ -76,25 +80,32 @@ Use `buildAgentProfileCell()` and `attest()` to produce current identities from Never relabel an existing digest or reconstruct a provenance envelope from unverified metadata. A new digest cannot establish that a registration existed before its evidence was observed. -`openSealedExperiment` is the only execution surface. -A rule that is not in the sealed spec cannot run; a rule that is cannot run differently. +Use the opened handle when the result must follow a particular registration. +Direct helpers such as `computeInterval()` and `executeDecisionRule()` also accept unsealed rules for development. +They do not establish a link to a registered experiment. ### Cluster-aware power refusal -Two floors, one simulation: +`clusteredPower()` combines a cluster-count policy with a simulated power curve: -- **Closed form, zero spend.** With `C` independent clusters, the exact whole-cluster sign-flip test can never produce a two-sided p below `2^(1-C)`. - Four clusters give 0.125 and three give 0.25 — both above alpha 0.05, so those designs are refused at any effect size. - Six clusters is the smallest certifiable count at 0.05. -- **Seeded simulation.** Per-row paired contrasts are drawn under a registered effect model (base win/loss rates, optional noisy clusters), each trial takes a whole-cluster percentile bootstrap, and power is the fraction of trials whose interval excludes zero. +- The exact whole-cluster sign-flip test has a minimum two-sided p-value of `2^(1-C)` for `C` independent clusters. + At alpha 0.05, this policy requires at least six clusters; four give 0.125 and three give 0.25. +- Seeded simulations draw paired contrasts under the configured win/loss model and apply a whole-cluster percentile bootstrap. + Power is the fraction of simulated intervals that exclude zero. + +The six-cluster floor is a policy for this helper, not a universal requirement for every estimator or fixed-roster evaluation. +The helper computes both results; it does not skip simulation when the cluster-count policy fails. The refusal is a verdict inside the returned artifact (`result.refusal`), with `assertDesignAdequate` as the throwing form. Both `clusteredPower` and the registered `power-floor` gate require `minimumEffect`. +The gate evaluates a supplied curve; it does not run the simulation itself. The effect must appear exactly in the supplied grid; the API does not interpolate. Adequacy requires target power at that effect. `maxPower` describes the grid and cannot establish adequacy at a smaller effect. ```ts +import { assertDesignAdequate, clusteredPower } from '@tangle-network/agent-eval/experiment' + const power = clusteredPower({ clusterSizes: Array.from({ length: 24 }, () => 3), effects: [0.05, 0.1, 0.2], @@ -115,23 +126,32 @@ The [statistical evidence guide](./statistical-evidence.md) explains unit counts ### The funnel `buildFunnel` refuses a stage that gains rows, named exclusions that do not sum, and partitions that overdraw their source stage. -`executeAdmissionRule` runs a sealed admission rule over rows and returns the funnel, the survivors, and the partition rows in one object — the chain and the rows can never disagree. -Partitions carry `pooling: 'never'`: a secondary set is reported beside the primary chain and cannot be pooled into it. +`registered.admit(rows)` applies the sealed admission rule and returns the funnel, survivors, and partition rows together. +The standalone `executeAdmissionRule(rule, rows)` also accepts an unsealed rule. +Partitions carry `pooling: 'never'`: report each secondary set separately from the primary chain. The object is its own JSON render; `renderFunnelTable` prints the text table with the reconciliation line (`input = surviving + excluded`). ### Matched budgets `verifyMatchedBudgets` compares realized per-arm tokens under the registered tolerance and returns a verdict whose `refusal` field carries `onFail: 'refuse-contrast'` when arms diverge. -A contrast between arms that spent differently is not a contrast; the refusal is the artifact that says so. +Use this check when the claim requires matched token use. +An unequal-budget comparison answers a different question and must retain the resource difference in its interpretation. ## Acceptance: the three preregistrations -The module's acceptance suite (`tests/experiment/preregistration-acceptance.test.ts`) re-derives the week's three hand-written preregistrations as sealed specs and reproduces each recorded decision by executing the sealed rules against the recorded evidence: +The [acceptance suite](../tests/experiment/preregistration-acceptance.test.ts) encodes three historical preregistrations as sealed specifications. +It checks their recorded decisions against fixed evidence: -- **killtest-20260810** — all four validity gates fail on the recorded evidence (the rep-4 oracle flip, the 2-row population drift, the zero-call control, the 0.692 power ceiling) and the halt rule refuses the spend, matching the recorded `$0.00, contrast never run`. +- **killtest-20260810**: all four validity gates fail on the recorded evidence. + The failures are the rep-4 oracle flip, 2-row population drift, zero-call control, and 0.692 power ceiling. + The halt rule refuses spend, matching the recorded `$0.00, contrast never run`. The obligation node routes a positive interval without the registered control to `blocked-pending-registered-control`, never to `thesis-survives`. -- **freelunch-20260810** — the admission funnel reproduces the recorded `48 > 43 > 35 > 35 > 32` chain with the 3-row secondary partition; the uniform-pass budget reproduces the recorded uniform n=2; the amendment-6 ledger under the same sealed rule refuses pass 2 — the registered-vs-ran divergence the seal makes unrepresentable; the report-only decision reproduces `3/64` and `2/32`. -- **tbench-20260808 milestone 2** — the round-robin selection reproduces the recorded 20-row subset in pick order; the m3 subset filters the SEALED m2 draw (16 rows); the decision table on the recorded interval reproduces `not-certified-at-this-n`. +- **freelunch-20260810**: the admission funnel reproduces `48 > 43 > 35 > 35 > 32` with the 3-row secondary partition. + The uniform-pass budget reproduces uniform n=2; the amendment-6 ledger under the same sealed rule refuses pass 2. + The report-only decision reproduces `3/64` and `2/32`. +- **tbench-20260808 milestone 2**: round-robin selection reproduces the recorded 20-row subset in pick order. + The m3 subset filters the sealed m2 draw to 16 rows. + The decision table on the recorded interval reproduces `not-certified-at-this-n`. ## What is composed, not duplicated @@ -139,7 +159,7 @@ The statistical machinery underneath is re-exported from its existing homes; thi | family | home | | --- | --- | -| `pairedBootstrap`, `mcnemar`/`mcnemarPower`/`mcnemarRequiredN`, `pairedRiskDifference*`, `holm`, `benjaminiHochberg`, `eProcess`, `wilson`, `mulberry32`, sample-size helpers | `src/statistics.ts` | +| `pairedBootstrap`, `mcnemar`/`mcnemarPower`/`mcnemarRequiredN`, `pairedRiskDifference*`, `holm`, `benjaminiHochberg`, `eProcess`, `wilson`, `mulberry32`, sample-size helpers | [`src/statistics/index.ts`](../src/statistics/index.ts) | | `pairedEvalueSequence` (anytime-valid) | `src/sequential.ts` | | `powerPreflight` (variance-based MDE refusal) | `src/campaign/gates/power-preflight.ts` | | `sequentialPairedGate`, `sequentialDecide` (manifest-bound) | `src/campaign/gates/sequential.ts` | @@ -154,5 +174,5 @@ The trace-repair admission machinery (`buildDenominatorChain`, oracle determinis ## Where this sits -This is Wave 2 of the [charter](./charter.md): the experiment subpath, built after the kill test that re-derived the three preregistrations as decision-rule objects (verdict: extended — ten node families beyond the seed AST, no opaque node, no rule dropped). -Wave 3 wires these objects to the live-sandbox seam; the improvement receipt (Wave 4) serializes a sealed experiment's digest, gates, and refusal outcomes into one attested file. +The [charter](./charter.md) describes current package ownership and host responsibilities. +Use [evaluation claims and final evidence](./evaluation-integrity.md) when connecting a registration to an automated improvement workflow. diff --git a/docs/feature-guide.md b/docs/feature-guide.md index 245ea0d77..b357f76fb 100644 --- a/docs/feature-guide.md +++ b/docs/feature-guide.md @@ -55,10 +55,10 @@ user intent -> datasets and optimizers replay the same adapter ``` -The important part is that production and eval do not use different loops. The -adapter can swap dependencies (real user session, replay fixture, sandbox), but -the state shape, validators, actions, budgets, and stop policies should stay -the same. That is what makes benchmark gains transfer to real usage. +Keep the production state, validators, actions, budgets, and stop policies in the evaluation path. +The adapter can supply a real user session, replay fixture, or sandbox. +This tests the behavior that production executes. +Transfer to future tasks still requires representative evaluation data and a measured comparison. ### Agent Runtime Integration @@ -82,8 +82,7 @@ Implementation ownership: assignment, and optimizer row conversion in `agent-eval`. - Put product state readers, action executors, approval policy, credentials, workspace paths, and UI-specific storage in the downstream repo. -- Promote a product adapter into `agent-eval` only after at least two products - need the same adapter shape. +- Keep product execution adapters in the consuming repository. ### Code Generator diff --git a/docs/hosted-ingest-spec.md b/docs/hosted-ingest-spec.md index 53a42fbb1..d822702f4 100644 --- a/docs/hosted-ingest-spec.md +++ b/docs/hosted-ingest-spec.md @@ -178,17 +178,10 @@ It is not production storage because process restart clears its in-memory data. TENANT_KEY=dev-token TENANT_ID=acme pnpm tsx examples/hosted-ingest-server/server.ts ``` -In another terminal: - -```sh -HOSTED_ENDPOINT=http://localhost:8080 \ -HOSTED_TENANT_KEY=dev-token \ -HOSTED_TENANT_ID=acme \ -pnpm tsx examples/foreign-agent-quickstart/index.ts -``` - -The quickstart's eval-run gets POSTed to the reference receiver; the -receiver's `GET /v1/runs` lists it back. +Send events with [`createHostedClient`](../src/hosted/client.ts) from `@tangle-network/agent-eval/hosted`. +Call `client.ingestEvalRun(event)` with a valid `EvalRunEvent` after configuring the client's endpoint, tenant ID, and API key. +For `selfImprove`, configure `hostedTenant` as shown in the [receiver example](../examples/hosted-ingest-server/). +The receiver's authenticated `GET /v1/runs` endpoint lists ingested runs. --- diff --git a/docs/insight-report.md b/docs/insight-report.md index d0250a9e7..0e9493536 100644 --- a/docs/insight-report.md +++ b/docs/insight-report.md @@ -1,476 +1,208 @@ -# `InsightReport`: the report +# Read an InsightReport -The single shape every analysis call returns. `selfImprove()` embeds it in `SelfImproveResult.insight`; `analyzeRuns()` returns it directly. The hosted-tier wire format carries it on `EvalRunEvent.insightReport?`. - -Use `summarizeExecution({ runs })` when observed traces have no task-quality labels. -It returns only `execution` and `costProvenance`, so callers do not need to fabricate a quality score to report runtime facts. - -Every section is **opt-in based on what your data supports**: the function never invents signal. If your runs don't carry judge scores, `judges` is empty. If there's no baseline/candidate split, `lift` is undefined. The shape is consistent; population is honest. - -This page walks every section with a real (synthetic) example and explains how to act on it. - ---- - -## At a glance +`analyzeRuns()` summarizes captured `RunRecord` evidence and returns an `InsightReport`. +`selfImprove()` includes the same report in `result.insight`. +Analysis makes no model calls unless you supply an analyst that uses one. ```ts -interface InsightReport { - n: number // runs analyzed - execution: ExecutionInsight // duration, tokens, errors, terminal outcomes - composite: ScalarDistribution // always - perDimension: Record // when judgeScores carry dimensions - costQuality: { cost: ScalarDistribution; pareto: ParetoFigureSpec } // always - judges: Record // when runs carry judge scores - interRater?: InterRaterInsight // when raterScores supplied - lift?: LiftInsight // when baseline + candidate present - failureClasses?: FailureClassTally[] // canonical task-failure counts - failureClusters?: FailureClusterInsight // when AnalystRegistry wired - contamination?: ContaminationInsight // when canaryScenarios supplied - outcomeCorrelation?: OutcomeCorrelationInsight // when outcomeSignal supplied - release: ReleaseSummary // always - recommendations: Recommendation[] // always: read this FIRST -} -``` - ---- - -## `execution`: runtime facts, separate from quality - -Always present. -It reports duration, optional queue time, direct input, output, reasoning, cache-read, and cache-write tokens, model-call coverage, model cohorts, execution errors, terminal outcomes, and separately reported orchestration aggregates. -These fields describe what ran; they do not claim whether the task succeeded. -`executionErrors` counts child or internal errors reported by the producer. -`terminalOutcomes` reads only `RunRecord.terminalOutcome`, which must come from root-run or process evidence. -A child tool error can therefore appear in a run whose terminal outcome is `succeeded`. -Current OTel and code-agent adapters also preserve process, guardrail, judge, propagated-parent, and unknown error counts in `RunRecord.outcome.raw`. -These counters are diagnostic and never become task-quality scores. - -```jsonc -{ - "execution": { - "durationMs": { "n": 30, "p50": 5400, "p95": 82000, "min": 900, "max": 190000 }, - "queueMs": { - "n": 0, - "mean": null, - "p50": null, - "p95": null, - "stddev": null, - "min": null, - "max": null, - "histogram": [] - }, - "tokenUsage": { - "totals": { "input": 50132, "output": 471783, "reasoning": 12000, "cached": 60489565, "cacheWrite": 3032227 }, - "input": { "n": 30, "p50": 25, "p95": 56 }, - "output": { "n": 30, "p50": 230, "p95": 2651 }, - "reasoning": { "n": 12, "p50": 800, "p95": 2400 }, - "cached": { "n": 20, "p50": 94193, "p95": 310178 }, - "cacheWrite": { "n": 20, "p50": 3070, "p95": 11148 } - }, - "aggregateUsage": { - "runs": 2, - "tokenUsage": { - "totals": { "input": 5000, "output": 176829, "reasoning": 0, "cached": 0, "cacheWrite": 0 } - }, - "costUsd": { "n": 0 }, - "totalCostUsd": 0 - }, - "modelCalls": { "runs": 20, "events": 42, "reportingRuns": 30 }, - "models": [{ "model": "claude-opus@2026-07-01", "runs": 20 }], - "executionErrors": { - "runs": 2, - "fraction": 0.067, - "events": 3, - "reportingRuns": 30, - "errorSpanEvents": 3, - "errorSpanReportingRuns": 30, - "byTerminalOutcome": { - "succeeded": { "withErrors": 1, "withoutErrors": 26, "unreported": 0 }, - "failed": { "withErrors": 0, "withoutErrors": 1, "unreported": 0 }, - "cancelled": { "withErrors": 0, "withoutErrors": 0, "unreported": 0 }, - "incomplete": { "withErrors": 0, "withoutErrors": 0, "unreported": 0 }, - "unknown": { "withErrors": 1, "withoutErrors": 1, "unreported": 0 } - } - }, - "terminalOutcomes": { - "succeeded": 27, - "failed": 1, - "cancelled": 0, - "incomplete": 0, - "unknown": 2 - } - } -} -``` - -Use `distribution.n` for optional fields to distinguish an uncaptured category from a recorded zero. -When `distribution.n` is zero, `mean`, percentiles, standard deviation, minimum, and maximum are `null`. -Use `executionErrors.reportingRuns` to assess error-telemetry coverage. -`errorSpanEvents` preserves the exact child-span error count separately from other reported execution errors. -The error fraction uses `reportingRuns` as its denominator and is `null` when no run reported error telemetry, so missing telemetry is not treated as a clean run. -`byTerminalOutcome` is a cross-tab, not a causal recovery claim. -It keeps reported errors, reported zeroes, and missing error telemetry separate for every terminal result. -Missing terminal evidence counts as `unknown`, not `failed`. -Never add `aggregateUsage` to direct `tokenUsage`: orchestration spans may repeat model-call usage from other traces. -Cost remains in `costQuality`, where observed, estimated, and uncaptured USD stay separate. - ---- - -## `n` + `composite` + `perDimension`: distributional summary - -Always present. The basic "where are my numbers" view. - -```jsonc -{ - "n": 30, - "composite": { - "n": 30, - "mean": 0.683, "p50": 0.667, "p95": 1.000, "stddev": 0.231, - "min": 0.0, "max": 1.0, - "histogram": [ - { "lo": 0.0, "hi": 0.083, "count": 5 }, - { "lo": 0.083, "hi": 0.167, "count": 0 }, - // ...12 bins by default - ] - }, - "perDimension": { - "clarity": { "mean": 0.72, "p50": 0.75, "p95": 0.95, "stddev": 0.18, /* ... */ }, - "concision": { "mean": 0.65, "p50": 0.68, "p95": 0.88, "stddev": 0.21, /* ... */ } - } -} -``` - -Read `composite.mean` only when `composite.n > 0`. -A `null` mean means task quality was not measured, not that quality was zero. -When a measured mean is below 0.5, inspect the lowest-scoring runs before tuning. - -**Read next:** `perDimension`. If `clarity` is high but `concision` is low, your prompts get the right ideas in too many words: different fix than "wrong ideas." - -**Use the histogram for:** finding bimodal failure modes. A bin with `count > 0` near zero and another > 0 near 1 means your agent has two distinct behaviors, not one noisy one. - -### Why this is not the same shape as a campaign aggregate - -This report uses `ScalarDistribution`. -A campaign aggregate (`CampaignResult.aggregates`) uses `SeriesDistribution`, the value `summarizeNumberSeries` returns. -The two shapes stay separate for three reasons, and none of them is an accident of history. - -1. `ScalarDistribution` is a wire contract. - `ScalarDistributionSchema` in `src/hosted/schemas.ts:45` is a strict Zod object. - It is embedded in `InsightReportSchema`, which is embedded in `EvalRunEventSchema`, which the hosted client validates every event against before it ships (`src/hosted/client.ts:182`). - A strict object rejects an unknown key, so adding or renaming a field breaks every event a consumer already sends. -2. `ScalarDistribution` reports what a report needs and a series summary does not have: a histogram, the worst-N `tailRuns` by score, `p95` for a latency question, and `mean` plus `stddev` beside the order statistics. - `SeriesDistribution` is the in-memory summary of a plain number series with no run identity attached. -3. The two answer at different `n = 0` boundaries. - `ScalarDistribution` represents an empty series as `n: 0` with every field `null`, because a report always has a slot for a metric it did not measure. - `summarizeNumberSeries` returns `null` for an empty series, because there is no distribution to report and a zero-filled summary would read as a measured all-zero series. - -Both refuse to encode a missing measurement as a zero. -That is the shared rule; the shapes differ because the surfaces differ. - ---- - -## `costQuality`: cost-vs-quality Pareto - -Always present. `cost.histogram` is the per-run cost distribution; `pareto` is the substrate's `ParetoFigureSpec`. - -```jsonc -{ - "costQuality": { - "cost": { - "mean": 0.024, "p95": 0.041, - "histogram": [/* */] - }, - "pareto": { - "kind": "pareto-cost-quality", - "split": "holdout", - "axes": { "x": "costUsd", "y": "score" }, - "points": [ - { "candidateId": "baseline", "cost": 0.018, "quality": 0.58, "n": 20, "onFrontier": true }, - { "candidateId": "winner", "cost": 0.027, "quality": 0.65, "n": 20, "onFrontier": true } - ] - } - } -} -``` - -**Use this when:** comparing prompts, models, or candidate surfaces. The Pareto frontier is your menu of "best you can do at each cost level." - -**Render with:** any chart library: `points` is plain JSON. Hosted-tier dashboards render this as a scatter with the frontier highlighted. - ---- - -## `judges`: per-judge mean - -Populated when run records carry `outcome.judgeScores`. - -```jsonc -{ - "judges": { - "domain-expert": { "n": 30, "meanScore": 0.71 }, - "helpfulness-llm": { "n": 30, "meanScore": 0.62 } - } -} -``` - -The substrate's full judge-calibration suite (positional bias, self-preference, verbosity bias) lives in `/reporting` and operates on **paired-by-condition** inputs that `analyzeRuns` doesn't synthesize from raw `RunRecord[]`. Wire them yourself when you have the paired data; the report's `judges` map is the corpus-level slice. - -**Use this when:** comparing multiple judges over the same corpus. A big gap between two judges' means is the first signal that one of them is mis-calibrated. - ---- - -## `interRater`: multi-rater agreement and disagreement review - -Populated when `analyzeRuns({ raterScores })` is supplied: typically via `fromFeedbackTable()`. - -```jsonc -{ - "interRater": { - "raters": 3, - "jointlyRated": 30, - "kappa": 0.40, - "icc": 0.42, - "pearson": 0.43, - "spearman": 0.41, - "perPair": { - "alice::bob": 0.53, - "alice::carol": 0.47, - "bob::carol": 0.19 - }, - "disagreementCases": [ - { "runId": "claim-7", "range": 1.00, - "ratings": [{"rater":"alice","score":1},{"rater":"bob","score":1},{"rater":"carol","score":0}] }, - { "runId": "claim-13", "range": 1.00, - "ratings": [{"rater":"alice","score":0},{"rater":"bob","score":0},{"rater":"carol","score":1}] } - // ...top 20 by range - ] - } -} -``` - -**Read first:** `kappa` and `icc`, which measure absolute agreement. -Pearson and Spearman measure correlation and can remain high when raters use different score levels. -When absolute agreement is low, review the largest disagreement cases before automating the rubric. - -**Use this when:** building per-rater LLM judges. Each rater's individual scores are the gold signal you calibrate against. Once a calibrated LLM matches the human ≥85%, you can auto-grade and escalate only the disagreement cases. - ---- - -## `lift`: paired-bootstrap statistical lift - -Populated when baseline + candidate candidates are present (auto-detected from two distinct `candidateId`s, or explicit via `baselineCandidateId` + `candidateCandidateId`). - -```jsonc -{ - "lift": { - "baselineMean": 0.58, - "candidateMean": 0.65, - "delta": 0.07, - "ci95": [0.04, 0.10], // bootstrap CI on the delta - "pValue": 0.0008, // paired t-test; null when the delta is a non-zero constant - "n": 40, // paired observations - "unpairedBaselineRuns": 2, - "unpairedCandidateRuns": 1, - "cohensD": 0.41, // paired Cohen's dz; null when delta variance is zero - "mde": 0.06, // min detectable effect at current n, 80% power - "requiredN": 38 // paired n needed at 80% power; null when dz is undefined - } -} -``` - -Rows pair only when `(experimentId, scenarioId, seed)` matches. -Missing `scenarioId` and duplicate identities fail loudly. -Unmatched rows are reported and excluded from paired statistics. - -**Decision rule:** -- `ci95[0] > threshold` → **SHIP.** Lower bound above your delta threshold means the lift is real at 95% confidence. -- `ci95[0] ≤ threshold < ci95[1]` → **INCONCLUSIVE.** Expand the corpus or wait for more data. -- `ci95[1] ≤ threshold` → **HOLD.** No evidence the candidate is better. - -The `recommendations` array surfaces exactly this decision (`kind: 'ship' | 'hold' | 'expand-corpus'`): that's what consumers should read. - -**Why bootstrap, not t-test alone:** paired bootstrap is distribution-free. Your judge scores are bounded in [0,1] and almost never normal; the bootstrap CI is the honest one. - ---- - -## `failureClasses`: canonical task-failure counts - -Populated when a run has a non-success `failureClass` or a measured task score below the failure threshold. -Runs without an explicit class are counted as `unknown`. -The optional `failureMode` remains domain-specific detail on the original run and is never used as a second grouping key. - -```jsonc -{ - "failureClasses": [ - { "failureClass": "bad_retrieval", "count": 9, "share": 0.28 }, - { "failureClass": "instruction_following", "count": 4, "share": 0.13 } - ] -} -``` - -Use this section to compare failure causes across products without a model call. -Use `failureClusters` when you need a semantic diagnosis within those classes. - ---- - -## `failureClusters`: grouped failure modes - -Populated when an `AnalystRegistry` is passed via `analyzeRuns({ analyst })`. The substrate runs each failed run through the registered analysts and groups findings by `analyst_id` / `area`. - -```jsonc -{ - "failureClusters": { - "totalFailures": 11, - "clusters": [ - { "id": "off-topic-drift", "name": "off-topic-drift", - "share": 0.45, "exemplars": ["run-12", "run-19", "run-33"] }, - { "id": "over-confidence", "name": "over-confidence", - "share": 0.27, "exemplars": ["run-3", "run-21"] }, - { "id": "format-mismatch", "name": "format-mismatch", - "share": 0.18, "exemplars": ["run-41", "run-44"] } - ] - } -} -``` - -**Read first:** the top cluster's `share`. If one cluster is > 40% of failures, fix that pattern before doing anything else. - -**Use this when:** triaging a regression. Failure clusters tell you "fix this kind of thing first." - -**To wire it:** register analysts in `AnalystRegistry`. See `src/analyst/registry.ts` and `src/analyst/kinds/index.ts` for the four built-in kinds (`failure-mode`, `improvement`, `knowledge-gap`, `knowledge-poisoning`). - ---- - -## `contamination`: canary check - -Populated when canary scenarios are passed via `analyzeRuns({ canaryScenarios })`. Each canary carries a sentinel string the agent should never emit; the report counts leaks. - -```jsonc -{ - "contamination": { - "leaks": 0, - "holdoutAuditPassed": true, - "details": [] - } -} -``` - -When `leaks > 0`: - -```jsonc -{ - "contamination": { - "leaks": 2, - "holdoutAuditPassed": false, - "details": [ - { "runId": "run-12", "canary": "xyz-secret-canary-123", "matched": "...the secret xyz-secret-canary-123 says..." } - ] - } -} -``` - -**When this fails:** your holdout corpus has leaked into training context. The `lift` number is **unreliable**. Investigate before shipping anything. - ---- - -## `outcomeCorrelation`: closing the loop on real outcomes - -Populated when `outcomeSignal: { metric, valueByRunId }` is supplied. - -```jsonc -{ - "outcomeCorrelation": { - "metric": "engagement_rate", - "n": 80, - "pearson": 0.72, // linear correlation - "spearman": 0.69, // rank correlation (robust to monotonic nonlinearity) - "rewardModel": { - "intercept": 0.04, - "slope": 1.93, - "r2": 0.52 // share of outcome variance the judge explains - } - } +import { analyzeRuns } from '@tangle-network/agent-eval/contract' +import type { RunRecord } from '@tangle-network/agent-eval' + +export async function compareCapturedRuns(runs: RunRecord[]) { + return analyzeRuns({ + runs, + split: 'holdout', + baselineCandidateId: 'baseline', + candidateCandidateId: 'candidate', + decisionThreshold: 0.02, + }) } ``` -This is the layer that says **"does my judge's taste actually predict the metric the business cares about?"** +Pass both candidate IDs to preserve the intended comparison direction, including regressions. +Without both IDs, the analyzer infers a comparison only when exactly two candidates exist. +It then treats the lower-scoring candidate as the baseline. +That inference cannot establish whether a specific change regressed. -**Read first:** `spearman`. If it's < 0.3 in absolute value, your judges are scoring something different from what wins downstream. Refit the judges (use the customer's downstream signal as gold) or change the rubric. +Use `summarizeExecution({ runs })` when traces contain runtime facts without task-quality labels. +It returns `execution` and `costProvenance` without interpreting release readiness. -**The reward model** is the simple linear `y = intercept + slope * composite`. Use it to: -- Predict the engagement of a new run from its composite score alone. -- Set a `composite` threshold for "must beat X to ship" based on the engagement equivalent. +The [offline example](../examples/analyze-existing-runs/) shows a complete call. +The [report types](../src/contract/insight-report.ts) and [analysis options](../src/contract/analyze-runs.ts) define the current API. ---- +## Match sections to evidence -## `release`: pass/warn/fail axes - -Always present. Roll-up across three axes: quality lift, contamination, composite distribution. +| Section | Input and scope | +|---|---| +| `n` | Number of validated input runs. | +| `execution` | Recorded durations, tokens, models, execution errors, and terminal outcomes. | +| `composite` | Finite scores from the selected split. | +| `perDimension`, `judges` | Recorded `outcome.judgeScores`; empty maps when absent. | +| `costQuality` | Known observed or estimated costs, their provenance, and candidate cost/quality points. | +| `lift` | Scored baseline/candidate rows sharing pairing identities. | +| `interRater` | Supplied `raterScores`, with at least two raters and jointly rated runs. | +| `failureClasses` | Explicit non-success classes or measured scores below the analyzer's failure threshold. | +| `failureClusters` | Findings from the supplied `AnalystRegistry` on failed runs. | +| `contamination` | Supplied `canaryScenarios`, searched within captured text outputs. | +| `outcomeCorrelation` | Supplied `outcomeSignal`, joined to at least three finite run scores. | +| `priorPeriodComparison` | Supplied `baselineRuns`, compared as an unpaired prior window. | +| `release`, `recommendations` | Diagnostic rules applied to the populated sections. | + +Optional sections are absent when their required inputs are unavailable. +Some supplied inputs can produce empty sections, such as no failure clusters among successful tasks. +An absent section does not establish that its check passed. + +`split: 'auto'` selects holdout if any run has a holdout score; otherwise it selects search. +Set `split` explicitly when you know which scores the report should use. +Inspect `composite.n`: it can be smaller than `n` when some runs have no score for that split. + +## Read execution and missing values + +`execution` describes what ran. +A successful process can still produce an incorrect task result. +A child tool error can also occur in a run whose root outcome is `succeeded`. + +`terminalOutcomes` reads `RunRecord.terminalOutcome`. +Missing terminal evidence counts as `unknown`. +`executionErrors` reads producer-reported error counts independently. +Its `fraction` uses `reportingRuns` as the denominator and becomes `null` when no run reports error telemetry. +The `byTerminalOutcome` table separates reported errors, reported zeroes, and unreported error telemetry for each terminal outcome. +It describes their co-occurrence; it does not establish recovery or causation. + +Optional token categories and queue time carry their own distribution counts. +For any `ScalarDistribution`, `n: 0` means no finite measurement was available. +Its mean, percentiles, standard deviation, minimum, and maximum are then `null`, with an empty histogram. +A measured zero has a positive count and a value of zero. + +Keep orchestration `aggregateUsage` separate from direct token usage. +An aggregate span can repeat usage already captured in model-call traces. +Adding both totals can count the same work twice. + +## Inspect quality and cost distributions + +`composite` describes the scored input corpus, including both candidates when both are supplied. +It is not the candidate's mean alone. +Use `lift.baselineMean` and `lift.candidateMean` for the paired comparison. +`composite.tailRuns` identifies the lowest-scoring runs for inspection. +Histogram peaks can suggest subgroups; inspect cases before attributing them to distinct agent behaviors. + +Judge details use this recorded shape: -```jsonc -{ - "release": { - "status": "pass", - "axes": [ - { "name": "quality-lift", "status": "pass", - "detail": "delta=0.070, CI95=[0.040, 0.100], n=40" }, - { "name": "contamination", "status": "pass", - "detail": "0 canary leak(s)" }, - { "name": "composite-distribution", "status": "pass", - "detail": "mean=0.683, p50=0.667, p95=1.000 over n=30" } - ], - "issues": [] - } +```ts +const judgeScores = { + perJudge: { + 'field-check': { accuracy: 0.75 }, + }, + perDimMean: { accuracy: 0.75 }, + composite: 0.75, } ``` -Overall `status` is `fail` if any axis fails; `warn` if any warn; `pass` otherwise. - -**Use this when:** wiring agent-eval into CI. A `status === 'pass'` from `analyzeRuns` on the candidate vs baseline is your green-light gate. - ---- - -## `recommendations`: the actionable layer +Store it as `RunRecord.outcome.judgeScores` alongside the relevant search or holdout score. +`perDimension` summarizes dimensions; `judges` reports per-judge counts and means. +Different judge means can reflect different coverage, scales, or criteria. +Compare shared cases before attributing a difference to miscalibration. -Always present. Read this first. +`costQuality.provenance` separates observed USD, estimated USD, and uncaptured costs. +Uncaptured rows are excluded from the cost distribution and Pareto calculation. +Read `knownFraction` and `costQuality.degraded` before comparing costs. +A frontier only compares the observed candidate points; it does not identify the best possible system. -```jsonc -{ - "recommendations": [ - { "priority": "critical", "kind": "ship", - "title": "Ship: lift 0.070 (95% CI 0.040..0.100)", - "detail": "Holdout lift exceeds threshold 0.02 with 95% bootstrap confidence (n=40, p=0.0008, d=0.41).", - "evidencePath": "lift" }, - { "priority": "high", "kind": "investigate", - "title": "Top failure cluster: off-topic-drift (45% of failures)", - "detail": "11 runs failed. The largest cluster groups 3 exemplars under 'off-topic-drift'.", - "evidencePath": "failureClusters.clusters[0]" } - ] -} -``` - -| `kind` | When emitted | -|---|---| -| `ship` | lift CI lower bound > threshold | -| `hold` | lift CI upper bound ≤ threshold | -| `expand-corpus` | lift CI straddles threshold: more data needed | -| `fix` | canary contamination detected | -| `recalibrate` | inter-rater κ < 0.5, OR outcome correlation < 0.3 | -| `investigate` | top failure cluster > some-share | +Campaign aggregates use `SeriesDistribution`; insight reports use `ScalarDistribution`. +The latter adds report fields such as histograms and optional run examples. +An empty campaign number series returns `null`; an empty report distribution retains its slot with `n: 0` and null statistics. +Both preserve the distinction between missing measurements and measured zeroes. -`evidencePath` points back into the report (`"lift"`, `"contamination"`, `"failureClusters.clusters[0]"`) so a UI can deep-link from each recommendation to its evidence. +## Interpret paired lift ---- +Rows pair on `(experimentId, scenarioId, seed)`. +Missing scenario IDs and duplicate identities within an arm fail validation. +Scored rows without a partner remain in `unpairedBaseline` and `unpairedCandidate` counts and are excluded from the paired statistics. +Unscored rows are also excluded; check the input and score counts separately. -## How `analyzeRuns` populates each section +For repeated tasks from one source, supply `independentUnitByScenarioId` as a map from every scored scenario ID to its independent unit. +The analyzer pairs runs first, averages matched scores within each declared unit, and weights units equally. +Raw score distributions and unmatched-run counts remain unchanged. +Repeated runs measure variation on those tasks; they do not create new independent tasks. -| Section | Required input | +| Lift field | Meaning | |---|---| -| `composite`, `perDimension`, `costQuality`, `release`, `recommendations` | `runs` | -| `judges` | `runs` with `outcome.judgeScores` | -| `interRater` | `raterScores` (≥ 2 raters jointly rated some runs) | -| `lift` | two distinct `candidateId`s in `runs` (or explicit baseline/candidate ids) | -| `failureClusters` | `analyst` registry passed in | -| `contamination` | `canaryScenarios` passed in | -| `outcomeCorrelation` | `outcomeSignal` passed in | - -All sections beyond the always-present ones are `T | undefined`, never empty objects. If a section is missing, your inputs didn't support it: the report is honest about that. +| `baselineMean`, `candidateMean`, `delta` | Paired means and candidate-minus-baseline difference after any unit aggregation. | +| `ci95` | Paired bootstrap interval for the mean difference. | +| `n` | Paired observations used for inference; independent units when declared. | +| `pairedRunN`, `independentUnitIds` | Raw matched count and unit IDs, present when units are declared. | +| `minimumRequired`, `decisionEligible` | Bootstrap sample floor and whether the count reaches it. | +| `pValue` | Paired t-test diagnostic; `null` for a nonzero constant difference. | +| `cohensD` | Paired Cohen's dz; `null` when difference variance is zero. | +| `mde` | Approximate detectable effect in standardized units at 80% power. | +| `requiredN` | Approximate sample size using the observed standardized effect; `null` when it cannot be estimated. | + +The analyzer's bootstrap decision floor is 20 paired observations. +Below it, a positive interval remains descriptive and the lift recommendation requests more evidence. +Reaching the floor only establishes sample-count eligibility. +A zero-width interval still cannot produce a lift-based ship recommendation. +An eligible, nonzero-width interval must exceed `decisionThreshold`, which defaults to `0.02` in score units. + +Bootstrap inference depends on representative, independent observations and adequate sample size. +It cannot repair selection bias, leaked final cases, or a miscalibrated judge. +Do not compare standardized `mde` directly with raw score lift. +Treat `requiredN` as an exploratory approximation, not a prospective power calculation for a target chosen before the study. + +## Use recommendations as diagnostics + +`recommendations` links findings to report sections through `evidencePath`. +Its `ship` kind can refer to lift or an improved prior-period metric. +Other findings can coexist with it, including a failed canary check. +Read the complete report before acting. + +`release` rolls up quality lift, canary matches, and composite score thresholds. +An unavailable axis is `not_evaluated` and makes the overall status at least `warn`. +The quality-lift axis uses positive lift; recommendation thresholds can differ. +These built-in thresholds are report heuristics, not your product's complete release policy. + +For automated promotion, use the campaign gate and inspect its contributing checks. +`selfImprove().gateDecision` comes from that gate. +See [concepts](./concepts.md#the-five-release-decisions) and the [held-out gate example](../examples/held-out-gate/). +A reusable claim can declare independent units and a practical effect. +Optional final-evidence tracking records fresh confirmation; see [evaluation integrity](./evaluation-integrity.md). + +## Investigate failures and disagreement + +`failureClasses` counts explicit non-success classes and scores below `0.5`. +A low-scoring run without a non-success class is counted as `unknown`. +Its `share` uses all input runs as the denominator. +Domain-specific `failureMode` stays on the original record. + +`failureClusters` runs registered analysts on those failed runs. +It groups findings by area, with analyst ID as fallback. +Each cluster's `share` counts affected failed runs, including those beyond the five displayed exemplars. +Multiple findings in one cluster count once per run. +A run can belong to several clusters, so cluster shares can sum above one. +Cluster shares use `totalFailures`, unlike the corpus denominator in `failureClasses`. +Empty findings can mean analysts skipped or failed; inspect registry logs and hooks when coverage is uncertain. +See the [custom analyst example](../examples/custom-trace-analyst/) for registration. + +`interRater` uses runs scored by every supplied rater. +Check `jointlyRated` before interpreting agreement over the broader corpus. +Kappa and ICC assess agreement; Pearson and Spearman assess correlation. +Review the largest disagreement cases and validate any judge changes on independent examples. +Choose acceptance thresholds for the decision's actual error costs. + +## Check canaries and downstream outcomes + +The canary check searches strings in `metadata.output`, falling back to `metadata.text`. +Other output layouts need conversion before analysis. +A match establishes that captured output contains a sentinel; investigate how it arrived there. +A zero-leak result does not prove isolation, especially when outputs were not captured. +The section does not report output-coverage counts. + +`outcomeCorrelation` joins finite `outcomeSignal.valueByRunId` values to run scores. +Its Pearson and Spearman values describe association in that supplied sample. +The linear `rewardModel` is fitted and evaluated on those same observations. +Validate it on separate data before using it to predict outcomes or set a release threshold. +Weak correlation can reflect noise, limited range, confounding, or the wrong rubric. +It does not identify the cause by itself. + +Use [outcome validity](./outcome-validity.md) for declared outcome directions, explicit exclusions, and association intervals. +Use `baselineRuns` for an unpaired prior-period comparison of available metrics. +Period differences can reflect traffic, task mix, or capture changes; they do not isolate the effect of a deployment. diff --git a/docs/outcome-validity.md b/docs/outcome-validity.md index 9a62bde22..70878665a 100644 --- a/docs/outcome-validity.md +++ b/docs/outcome-validity.md @@ -2,7 +2,7 @@ `rubricPredictiveValidity()` measures associations between rubric scores and observations from deployment. Declare the desired direction for each outcome before reading the results. -Higher rubric scores always mean better evaluated behavior. +Define rubric scores so higher values mean better evaluated behavior. ```ts import { @@ -35,6 +35,21 @@ Append deployment observations with the same `runId` and exact outcome metric ke The store accepts finite numbers, including zero. Omit an unmeasured metric instead of replacing it with zero. +An observation records when and where its outcomes were captured: + +```ts +const observation: DeploymentOutcome = { + runId: 'support-run-42', + capturedAt: Date.now(), + metrics: { success_rate: 1, failure_rate: 0 }, + labels: { cohort: 'support' }, + source: 'resolution-events-v1', +} +``` + +`capturedAt` uses epoch milliseconds. +The matching run supplies the rubric score; this observation supplies the measured deployment outcome. + The default reduction selects the latest finite observation of each requested metric. A newer row containing another metric cannot supply its value or erase an older observation. The `mean` and `max` reductions operate on that metric alone. @@ -127,7 +142,8 @@ Repair the source and retry the read. For trace data, `correlationStudy()` accepts outcome names and reports descriptive associations without a desired direction. It shares the metric reduction, bootstrap, and exclusion behavior. -Its optional capture window includes only observations captured after the run started. +It excludes observations captured before the run started. +`maxCaptureLagMs` optionally bounds the elapsed time from run start through outcome capture, including both endpoints. ## Inspect calibration by score range @@ -153,6 +169,7 @@ It reports each bin's count, mean score, mean outcome, and absolute gap. Both quantities must use comparable numerical scales for this difference to measure calibration. The `range` option clips scores before binning and retains every joined observation. +The reported calibration error then describes the clipped scores; retain that transformation with the result. Constant scores form one bin, so a consistently wrong predictor still receives a measured calibration error. Equal-frequency binning produces the requested number of bins, capped by the observation count. Bin counts differ by at most one, except when all scores share one value. diff --git a/docs/product-eval-adoption.md b/docs/product-eval-adoption.md index 99e8bd907..fe9699167 100644 --- a/docs/product-eval-adoption.md +++ b/docs/product-eval-adoption.md @@ -152,8 +152,7 @@ set with a signed note. ## Optimization -Use `runImprovementLoop()` when the system is a multi-step agent, not a -single prompt. +Use `runImprovementLoop()` when a `SurfaceProposer` supplies candidates for Agent Eval's search loop. Good optimization targets: diff --git a/docs/research-report-methodology.md b/docs/research-report-methodology.md index 3a0edef32..24f46d97b 100644 --- a/docs/research-report-methodology.md +++ b/docs/research-report-methodology.md @@ -60,14 +60,14 @@ In order: first match wins: | Quantity | Function | Source file | |---|---|---| -| Marginal CI on score mean | `confidenceInterval` | `statistics.ts` | -| Paired Cohen's dz vs comparator | `pairedCohensDz` | `statistics.ts` | -| Wilcoxon signed-rank (paired), exact at n ≤ 20 | `wilcoxonSignedRank` | `statistics.ts` | -| BH-FDR q-values | `benjaminiHochberg` | `statistics.ts` | -| Paired bootstrap CI on median delta | `pairedBootstrap` | `statistics.ts` | -| Smallest p a rank-test design can produce | `pFloor` on the result | `statistics.ts` | +| Marginal CI on score mean | `confidenceInterval` | [`descriptive.ts`](../src/statistics/descriptive.ts) | +| Paired Cohen's dz vs comparator | `pairedCohensDz` | [`effect-sizes.ts`](../src/statistics/effect-sizes.ts) | +| Wilcoxon signed-rank (paired), exact at n ≤ 20 | `wilcoxonSignedRank` | [`rank-tests.ts`](../src/statistics/rank-tests.ts) | +| BH-FDR q-values | `benjaminiHochberg` | [`multiplicity.ts`](../src/statistics/multiplicity.ts) | +| Paired bootstrap CI on median delta | `pairedBootstrap` | [`paired-tests.ts`](../src/statistics/paired-tests.ts) | +| Smallest p a rank-test design can produce | `pFloor` on the result | [`rank-tests.ts`](../src/statistics/rank-tests.ts) | | Bayesian-bootstrap Pr(Δ>0), Pr(Δ∈ROPE) | `bayesianBootstrapMeanSamples` | `summary-report.ts` (private) | -| Minimum detectable paired effect | `pairedMde` | `statistics.ts` | +| Minimum detectable paired effect | `pairedMde` | [`power-and-mde.ts`](../src/statistics/power-and-mde.ts) | | Run fingerprint | `hashJson(...)` | `pre-registration.ts` | The Pr(Δ>0) and Pr(Δ∈ROPE) summaries use Rubin's Bayesian bootstrap. diff --git a/docs/verdicts.md b/docs/verdicts.md index 6e97477db..788d91969 100644 --- a/docs/verdicts.md +++ b/docs/verdicts.md @@ -1,67 +1,94 @@ -# One verdict vocabulary +# Verdicts and certifications -Every verification path in this package lands in one type: `DefaultVerdict` (`src/verdict.ts`). -`valid` answers "did it pass", `score` answers "how well" in [0, 1], `scores` carries the per-dimension breakdown, and `certification` says WHO certified. +`DefaultVerdict` is the shared base type for validator results. +Campaign judges return `JudgeScore`, and release gates return `GateResult`; their fields and score scales differ. +See [release check results](./concepts.md#release-check-results) when interpreting a campaign decision. -A certification is the epistemics a bare `valid` + `score` pair cannot carry. -A kernel-checked proof and an LLM judge can produce the same `{ valid: true, score: 1 }`; the certification is what tells them apart: +In `DefaultVerdict`, `valid` reports whether the validator's pass criteria were met, and `score` is its aggregate in [0, 1]. +Optional `scores` and `notes` carry dimensions and explanation. +Optional `certification` records the verification strategy, checker identity, assumptions, and evidence digest. +A certification can accompany a failed or incomplete check. +It does not establish that the result passed or that every required measurement exists. -- `strategy` — which verification-strategy member vouches (the 10-member family with per-member failure modes: [docs/verification-strategies.md](./verification-strategies.md)); -- `checker` — the exact identity that ran, with version and content pins, so the check is re-runnable; -- `assumptions` — every step the certificate rests on that the checker did NOT verify, named one by one; -- `evidenceDigest` — sha-256 of the evidence artifact (`certificationEvidenceDigest`). +The certification fields are: -An absent certification is itself a statement: scored, but nothing vouches. -No producer fakes one — a closed gate, an unexecuted proof, or an unattested checker yields an uncertified verdict, never an invented certificate. +- `strategy`: the [verification strategy](./verification-strategies.md), with its documented failure mode. +- `checker`: a name, version, and optional dependency pins. +- `assumptions`: the producer's list of steps the checker did not verify. +- `evidenceDigest`: the evidence identity; `certificationEvidenceDigest()` hashes its JSON-serialized form with canonical JSON and SHA-256. + +These fields record the producer's claims. +Reproduction also requires access to the evidence, checker, dependencies, and execution environment. +An absent certification means no verification strategy is recorded for that verdict. ## Producers -Every verifier below returns a `DefaultVerdict` (usually a richer extension of it) with a produced certification. +These result types extend `DefaultVerdict`. +Their additional fields distinguish failed checks from incomplete or unmeasured work. -| Verifier | Verdict type | Strategy | Certifies | Assumptions it names | +| Producer | Result type | Strategy | Scope | Fields to inspect | | --- | --- | --- | --- | --- | -| `MultiLayerVerifier.run` (`src/multi-layer-verifier.ts`) | `VerificationReport` | `composite` | ordered layer pipeline blend | each skipped / errored / timed-out layer | -| `verifyCompletion` (`src/completion-verifier.ts`) | `CompletionVerdict` | the checker's own (`judge` for the LLM checker, `schema` for token recall) | task completion over produced state | lexical structural stage + the checker's attestation | -| `evaluateTraceContract` (`src/trace-contracts.ts`) | `ContractVerdict` | `invariant` | LTLf rules over a span sequence | array ordering without timestamps; custom predicate functions | -| `evaluateOracles` (`src/oracle.ts`) | `OracleReport` | `test` | declarative expected-outcome assertions | — (the oracle set is its own answer key) | -| `replayVerify` (`src/trajectory-replay/verify.ts`) | `ReplayVerdict` | `replication` | recorded failure reproduced (and fix vanished) under re-execution | unadjudicated prefix steps; truncated prefix; returncode-only signature | -| `verifyFindings` (`src/trajectory-replay/findings.ts`) | `VerifyFindingsRun` | `replication` | a batch of analyst findings under executed replay | not-replayable findings leave the denominator | -| `gradeRepairRow` (`src/trace-repair/grade.ts`) | `RepairRowResult` | `test` | a proposed repair against the row's held-out suite (pins: suite + policy digests) | vacuous reproduction gate; prefix divergence | -| `runEquivalenceCheck` + `equivalenceVerdict` (`src/verification-strategy.ts`, `src/verdict.ts`) | `DefaultVerdict` | the spec's member (`proof-kernel` in the pilot) | two blind formal statements are equivalent | arms self-declare blindness | - -Two producers certify conditionally, on purpose: - -- `gradeRepairRow` certifies only a `measured` outcome — a funnel gate that closed before the suite ran has nothing to vouch for. -- `verifyFindings` certifies only when at least one proof executed — a batch where nothing was replayable measured nothing. - -## Reading a score out of a judge - -A model judge emits one grade per dimension. Discrete grades tie: two candidates that both score `8` carry no ranking signal between them, and a best-of-N selection then picks arbitrarily. - -`llmJudge({ scoring })` chooses how the number is read: - -| `scoring` | What it reads | Requires | +| `MultiLayerVerifier.run()` | `VerificationReport` | `composite` | Ordered verification layers | `layers`, `allPass`, and optional `taskScore` | +| `verifyCompletion()` | `CompletionVerdict` | The supplied checker's strategy | Completion requirements matched against produced state | Requirement evidence, `correct`, and `unmeasured` | +| `evaluateTraceContract()` | `ContractVerdict` | `invariant` | Temporal rules over recorded spans | Rule results and assumptions about ordering and predicates | +| `evaluateOracles()` | `OracleReport` | `test` | Declared expected-outcome assertions | `results`, `passCount`, and `failCount` | +| `replayVerify()` | `ReplayVerdict` | `replication` | Failure reproduction and an optional fix under re-execution | Prefix fidelity, signature matches, and both execution arms | +| `verifyFindings()` | `VerifyFindingsRun` | `replication` | Analyst findings checked through replay | `executions`, `counts`, and individual verifications | +| `gradeRepairRow()` | `RepairRowResult` | `test` | A repair against the admitted row's held-out suite | The grade's outcome and funnel evidence | +| `equivalenceVerdict(record)` | `DefaultVerdict` | The record's strategy | Formal-statement equivalence checked by `runEquivalenceCheck()` | Whether the obligation was proved, refuted, or unresolved | + +Certification is conditional for several producers: + +- `verifyCompletion()` includes it when the supplied checker provides an attestation. + Inspect requirement evidence to see which checks were assessed. +- `equivalenceVerdict()` includes it for a proved or refuted obligation with an evidence digest. + A certified refutation has `valid: false`. +- `gradeRepairRow()` includes it only for a `measured` grade. +- `verifyFindings()` includes it only when at least one replay execution ran. + Non-replayable findings remain in `counts` and are excluded from the score's denominator. + A score of 1 can therefore coexist with `valid: false` when some findings were not replayable. + +Some result shapes use `score: 0` when no task measurement is available. +For `VerificationReport`, use the presence of `taskScore` to identify a complete task measurement; `blendedScore` can describe a partial panel. +An empty oracle set has certification metadata but returns `valid: false` and no executed oracle results. +Read these completeness fields before using scores as task labels or release evidence. + +## Reading a score from a model judge + +A tied dimension score supplies no ordering between candidates. +`llmJudge({ scoring })` controls how the dimension score is read: + +| `scoring` | Measurement | Requirement | |---|---|---| -| `{ method: 'sampled' }` (default) | the grade the model emitted | nothing | -| `{ method: 'expectation', whenUnavailable }` | the expected grade over the integer grades the model considered at the score token | `scale: 'ten'` and a provider that returns log probabilities | - -Expectation scoring asks the provider for `logprobs` with `top_logprobs`, finds the token that carried each dimension's grade, and averages the integer grades in that token's probability window, weighted by probability. Two answers that both sample `8` separate by how much mass sat on `7` versus `9`. +| `{ method: 'sampled' }` (default) | The emitted grade | A valid grade response | +| `{ method: 'expectation', whenUnavailable }` | A probability-weighted grade from returned token alternatives | `scale: 'ten'` and provider log probabilities | -It needs one integer in one token, which is why `scale: 'ten'` is required: a `[0,1]` float is several tokens, and no single position carries its distribution. A grade that did not land in exactly one token — a two-token `10` — is refused rather than approximated. +With `scale: 'ten'`, the model emits grades from 0 to 10; `llmJudge()` divides them by 10 before returning dimensions and composite. +Expectation scoring finds each grade's token, keeps valid integer alternatives, and renormalizes their returned probabilities before averaging. +Its distribution is limited to those returned alternatives. +Additional precision alone does not establish better calibration or ranking accuracy. -`whenUnavailable` decides what happens when the provider returns no log probabilities, or the grade spans tokens: +Each emitted grade must occupy one integer token for expectation scoring. +A split `10`, missing grade token, or unavailable log probabilities invokes `whenUnavailable`: -- `'fail'` throws, and the campaign records a failed cell. -- `'sampled'` reads the emitted grade instead. +- `'fail'` throws; the campaign records a judge failure. +- `'sampled'` uses the emitted grades for that judge result. -`JudgeScore.scoringMethod` reports what actually produced the number, so a declared expectation run that fell back reads `'sampled'` and stays auditable. `JudgeScore.distribution` carries the probability mass per grade, and is present only for an expectation score. Panels are unchanged: `ensembleJudge` consumes the composite either way. +`JudgeScore.scoringMethod` is present when `scoring` was explicitly configured and records the method used. +An expectation request that falls back reports `'sampled'`. +When `scoring` is omitted, sampled scoring is the default and this metadata field is absent. +`JudgeScore.distribution` is present only for expectation scoring and contains the normalized probabilities over returned integer alternatives. +`ensembleJudge()` consumes the resulting composite through its usual interface. -Whether a given endpoint returns `logprobs.content` is a property of that provider and model, not of this package. `LlmCallResult.logprobs` is `null` when the provider returned none — never an inferred distribution. See `evidence/records/judge-logprob-wire-support.json` for the current verification state of that wire behavior. +Provider and model support determines whether log probabilities are available. +`LlmCallResult.logprobs` is `null` when none were returned. +The [recorded wire check](../evidence/records/judge-logprob-wire-support.json) documents the endpoints and conditions that were inspected. -## Consuming a certification +## Interpreting certification limits -Read `certification.strategy`, then weigh the member's documented failure mode — `VERIFICATION_STRATEGIES[strategy].failureMode` carries it at runtime. -"Certified" is never one bit: a `judge` certificate is Goodhart-gameable, a `test` certificate covers only its suite, a `composite` certificate can hide which member carried the score. -The assumptions list is the honest remainder; an empty list is the producer's explicit claim that nothing was left unverified, not a default. +Read `VERIFICATION_STRATEGIES[strategy].failureMode` alongside the evidence and assumptions. +A judge can reward misleading output; tests cover their suite; a proof can establish the wrong formal statement for the intended task. +A composite certification requires inspection of its component results. +An empty assumptions list is the producer's declaration, not independent verification that no assumptions remain. -Related docs: [verification-strategies.md](./verification-strategies.md) (the family and the equivalence protocol), [trace-repair-grader.md](./trace-repair-grader.md), [trajectory-replay.md](./trajectory-replay.md). +See [verification strategies](./verification-strategies.md), [repair grading](./trace-repair-grader.md), and [trajectory replay](./trajectory-replay.md) for the corresponding execution contracts. diff --git a/examples/README.md b/examples/README.md index d60864251..ebf8390a8 100644 --- a/examples/README.md +++ b/examples/README.md @@ -1,15 +1,14 @@ # Examples -Every directory holds one runnable file and a README that answers three questions: when to use it, how to run it, and why it is built that way. -Every expected output printed in a README is the output that file produces. - Start with [`evaluate-a-change`](./evaluate-a-change/). It is the smallest complete path: cases in, scores out. -Run any offline example from the repository root: +Install and build from the repository root, then run the first offline example: ```sh -pnpm tsx examples/evaluate-a-change/index.ts +pnpm install +pnpm build +pnpm exec tsx examples/evaluate-a-change/index.ts ``` ## Measure A Change @@ -18,7 +17,7 @@ pnpm tsx examples/evaluate-a-change/index.ts |---|---|---| | Score one change on the same cases | [`evaluate-a-change`](./evaluate-a-change/) | Offline | | See the case grid before you pay for it | [`plan-before-you-spend`](./plan-before-you-spend/) | Offline | -| Wrap an existing agent | [`foreign-agent-quickstart`](./foreign-agent-quickstart/) | Offline, or an OpenAI-compatible endpoint | +| Wrap an existing agent | [`foreign-agent-quickstart`](./foreign-agent-quickstart/) | Offline | | Evaluate several attempts per case | [`multi-shot-optimization`](./multi-shot-optimization/) | Offline | | Apply a release rule without any search | [`held-out-gate`](./held-out-gate/) | Offline | | Load cases from folders on disk | [`eval-fixtures-quickstart`](./eval-fixtures-quickstart/) | Offline | @@ -36,25 +35,8 @@ pnpm tsx examples/evaluate-a-change/index.ts | Let another package own the text search | [`adapt-a-text-optimizer`](./adapt-a-text-optimizer/) | Offline | | Compare official GEPA and SkillOpt | [`compare-optimization-methods`](./compare-optimization-methods/) | Python optimizer packages and an LLM endpoint | -Run one official optimizer: - -```sh -OPTIMIZERS=gepa \ -LLM_BASE_URL=https://router.tangle.tools/v1 \ -LLM_API_KEY="$TANGLE_API_KEY" \ -GEPA_PRICE_IN_PER_M=0.4 \ -GEPA_PRICE_OUT_PER_M=1.6 \ -pnpm tsx examples/compare-optimization-methods/index.ts -``` - -Any OpenAI-compatible endpoint works; point `LLM_BASE_URL` at it and set `LLM_MODEL` to a model it serves. - -The optimizer's reflection calls run through a caller-owned execution owner. These examples supply their own, `_shared/openai-compatible-owner.ts`, built from `LLM_BASE_URL` and `LLM_API_KEY`. -Set `OPTIMIZER_EXECUTION_OWNER_MODULE` to route them through your own execution package instead. -Replace the example rates with the exact rates for your endpoint. -Use `OPTIMIZERS=skillopt` for SkillOpt, or `OPTIMIZERS=gepa,skillopt` for a shared comparison. -Set `GEPA_RECIPE` to run a composed GEPA recipe — `sequential`, `adaptive-sequential`, `best-of`, `vote`, or `omni` — instead of one engine run. -Read the [optimizer install instructions](./compare-optimization-methods/README.md) first. +Use the comparison guide for [installation](./compare-optimization-methods/README.md#install) and [running selected methods](./compare-optimization-methods/README.md#run). +It also documents endpoint settings, rates, execution owners, and GEPA recipes. ## Prove A Result @@ -70,17 +52,20 @@ Read the [optimizer install instructions](./compare-optimization-methods/README. |---|---|---| | Get a report from runs you already have | [`analyze-existing-runs`](./analyze-existing-runs/) | Offline | | Get cited findings out of a failed batch | [`custom-trace-analyst`](./custom-trace-analyst/) | Offline | -| Analyze human approvals and rejections | [`customer-feedback-loop`](./customer-feedback-loop/) | Offline | -| Analyze OpenTelemetry spans | [`customer-otel-traces`](./customer-otel-traces/) | Offline | +| [Analyze human approvals and rejections](../docs/customer-journeys.md#2-analyze-human-ratings) | [`customer-feedback-loop`](./customer-feedback-loop/) | Offline | +| [Analyze OpenTelemetry spans](../docs/customer-journeys.md#1-analyze-existing-traces) | [`customer-otel-traces`](./customer-otel-traces/) | Offline | ## Benchmarks And Training | Goal | Example | |---|---| | Run public benchmark adapters | [`benchmarks`](./benchmarks/) | +| Compare optimizers on AppWorld tasks | [`AppWorld`](./benchmarks/appworld/) | | Export supervised and preference rows | [`publish-rl-dataset`](./publish-rl-dataset/) | | Fine-tune through Prime Intellect | [`fine-tune-with-prime-rl`](./fine-tune-with-prime-rl/) | +The AppWorld comparison requires separate AppWorld and optimizer Python environments, plus an LLM endpoint. + The GSM8K comparison reads a local dataset file from `AGENT_EVAL_GSM8K_PATH`. Produce it from the GSM8K test split with Python and `datasets`: @@ -95,6 +80,7 @@ python -c "from datasets import load_dataset; import json; \ or, without Python, from the upstream source of the Hugging Face dataset: ```sh +mkdir -p ~/.cache/agent-eval curl -L https://raw.githubusercontent.com/openai/grade-school-math/master/grade_school_math/data/test.jsonl \ | jq -c '{id: ("gsm8k-test-" + (input_line_number | tostring)), question, answer}' \ > ~/.cache/agent-eval/gsm8k.jsonl diff --git a/examples/_shared/env.ts b/examples/_shared/env.ts index f9977423e..1bd680d5c 100644 --- a/examples/_shared/env.ts +++ b/examples/_shared/env.ts @@ -1,3 +1,5 @@ +import type { CustomTokenPricing } from '../../src/cost-ledger' + export function positiveIntegerEnv(name: string, fallback: number): number { const value = Number(process.env[name] || fallback) if (!Number.isSafeInteger(value) || value <= 0) { @@ -49,3 +51,25 @@ export function optionalNonNegativeNumberEnv(name: string): number | undefined { } return value } + +export function tokenPricingFromEnv(): CustomTokenPricing | undefined { + const inputUsdPerMillion = optionalNonNegativeNumberEnv('PRICE_IN_PER_M') + const outputUsdPerMillion = optionalNonNegativeNumberEnv('PRICE_OUT_PER_M') + const cachedInputUsdPerMillion = optionalNonNegativeNumberEnv('PRICE_CACHED_IN_PER_M') + const cacheWriteUsdPerMillion = optionalNonNegativeNumberEnv('PRICE_CACHE_WRITE_IN_PER_M') + if ((inputUsdPerMillion === undefined) !== (outputUsdPerMillion === undefined)) { + throw new Error('PRICE_IN_PER_M and PRICE_OUT_PER_M must be set together') + } + if (inputUsdPerMillion === undefined || outputUsdPerMillion === undefined) { + if (cachedInputUsdPerMillion !== undefined || cacheWriteUsdPerMillion !== undefined) { + throw new Error('Cache token rates require PRICE_IN_PER_M and PRICE_OUT_PER_M') + } + return undefined + } + return { + inputUsdPerMillion, + outputUsdPerMillion, + ...(cachedInputUsdPerMillion === undefined ? {} : { cachedInputUsdPerMillion }), + ...(cacheWriteUsdPerMillion === undefined ? {} : { cacheWriteUsdPerMillion }), + } +} diff --git a/examples/_shared/openai-compatible-owner.ts b/examples/_shared/openai-compatible-owner.ts index a0c2b15d8..6679bf89e 100644 --- a/examples/_shared/openai-compatible-owner.ts +++ b/examples/_shared/openai-compatible-owner.ts @@ -1,7 +1,7 @@ /** * Caller-owned execution for the metered optimizer-model path. * - * Agent Eval never executes a paid model. Its loopback proxy owns admission, + * The optimizer bridge delegates model execution. Its loopback proxy owns admission, * budgets, identity checks, response bounds, and cost-ledger recording, then * hands the exact admitted request to the package that owns execution. A * product built on agent-runtime supplies `profileOptimizerModelCall`, which @@ -9,7 +9,7 @@ * * This file is the minimal transport for a caller who has only an * OpenAI-compatible `/chat/completions` endpoint. It is example code on - * purpose: the credential lives with the caller, not inside the package. + * purpose: the caller supplies the credential to this execution owner. * Copy it into your own project and replace the transport with whatever * client you already use. It exposes the same endpoint two ways — as the * `ChatClient` every Agent Eval judge and worker takes, and as the diff --git a/examples/analyze-existing-runs/README.md b/examples/analyze-existing-runs/README.md index 771c237d8..11a928acd 100644 --- a/examples/analyze-existing-runs/README.md +++ b/examples/analyze-existing-runs/README.md @@ -1,27 +1,18 @@ -# Get A Report From Runs You Already Have +# Analyze captured runs -## When to use this +This example constructs twelve synthetic `RunRecord` rows for two candidates answering six shared cases. +`analyzeRuns()` returns score distributions, cost, paired lift, and recommendations without rerunning the agent. +No API key or model call is required. -Use this example when the runs already happened. -You have logs, a feedback table, or exported rows, and you must know what they say. -No agent is invoked and no model is called. +## Run -Use a different front door when you want to run the agent again: [`evaluate-a-change`](../evaluate-a-change/). - -## How to run it +From the repository root: ```sh -pnpm tsx examples/analyze-existing-runs/index.ts +pnpm install --frozen-lockfile +pnpm exec tsx examples/analyze-existing-runs/index.ts ``` -No API key is required. - -## What it does - -1. Twelve `RunRecord` rows describe two candidates answering the same six cases. -2. `analyzeRuns()` reads them and returns one `InsightReport`. -3. The report carries score distributions, paired lift with an interval, judge agreement, cost, failure clusters, contamination checks, and recommendations. - The output is: ```text @@ -33,23 +24,31 @@ paired n: 6 recommendations: 2 ``` -## Why it is built this way +The positive interval is descriptive at six pairs. +The report requests more evidence because this bootstrap comparison requires at least 20 paired observations for decision eligibility. +The example supplies no judge details, rater scores, analyst, canaries, or downstream outcomes. +Their optional insights are therefore unavailable. -`baselineCandidateId` and `candidateCandidateId` make the lift paired. -Rows match on `(experimentId, scenarioId, seed)`, so each case is compared with itself, not with the mean of the other arm. -A row that finds no partner stays visible in the result instead of being dropped. +## Use your own evidence -Every field of the report is defined in [`docs/insight-report.md`](../../docs/insight-report.md). +Replace the synthetic rows in [index.ts](./index.ts) with captured records. +For an installed package, import `analyzeRuns` from `@tangle-network/agent-eval/contract`. +Import `RunRecord` from the package root. -## Where the rows come from +Pass both `baselineCandidateId` and `candidateCandidateId` to declare the direction of the comparison. +Rows pair on `(experimentId, scenarioId, seed)`. +Unmatched scored rows remain visible in `lift.unpairedBaseline` and `lift.unpairedCandidate` and are excluded from the paired statistics. +Missing or duplicate pairing identities fail validation. -If your data is not already in `RunRecord` shape, convert it first: +Declare `independentUnitByScenarioId` when repeated cases share a task, document, user, or other sampling unit. +Repetitions do not create more independent tasks. +Inspect the [InsightReport guide](../../docs/insight-report.md) for denominators, missing data, and diagnostic limits. -| Source | Adapter | +| Existing data | Starting point | |---|---| -| Approvals and rejections in a table | `fromFeedbackTable` — see [`customer-feedback-loop`](../customer-feedback-loop/) | -| OpenTelemetry spans | `fromOtelSpans` — see [`customer-otel-traces`](../customer-otel-traces/) | - -## Next +| Approvals and rejections | [`fromFeedbackTable`](../customer-feedback-loop/) | +| OpenTelemetry spans | [`fromOtelSpans`](../customer-otel-traces/) | +| Captured execution without quality labels | `summarizeExecution({ runs })` from `/contract` | +| Failures requiring an analyst | [Custom trace analyst](../custom-trace-analyst/) | -- Cluster the failures with a trace analyst: [`custom-trace-analyst`](../custom-trace-analyst/). +Use [evaluate a change](../evaluate-a-change/) when you need new agent executions. diff --git a/examples/benchmarks/gsm8k/compare-optimization-methods.ts b/examples/benchmarks/gsm8k/compare-optimization-methods.ts index b521ef075..fe3630114 100644 --- a/examples/benchmarks/gsm8k/compare-optimization-methods.ts +++ b/examples/benchmarks/gsm8k/compare-optimization-methods.ts @@ -187,7 +187,7 @@ interface Artifact { text: string } -// The worker transport is caller code: Agent Eval holds no provider key. +// The execution owner binds the caller's endpoint and credential. const chat = openAiCompatibleChatClient({ baseUrl: BASE_URL, apiKey: API_KEY, diff --git a/examples/compare-optimization-methods/README.md b/examples/compare-optimization-methods/README.md index 765eae381..72dae9128 100644 --- a/examples/compare-optimization-methods/README.md +++ b/examples/compare-optimization-methods/README.md @@ -1,159 +1,130 @@ -# Compare Official GEPA And SkillOpt +# Compare GEPA and SkillOpt -This example runs official GEPA, official SkillOpt, or both against the same transaction-extraction task. -Each method receives five train cases and three selection cases. -Agent Eval evaluates selected prompts on six separate final cases after optimization finishes. +This example runs GEPA, SkillOpt, or both on transaction extraction. +Each method receives the same five train cases, three selection cases, starting prompt, worker, and deterministic field-matching judge. +Agent Eval evaluates selected prompts on six separate final cases after every method finishes search. -The worker calls an OpenAI-compatible endpoint. -The field-level judge is deterministic. +Worker and optimizer calls use a paid Chat Completions endpoint. +Six final cases demonstrate the integration; they do not establish population-level optimizer superiority. +The example includes the unchanged starting prompt as a baseline, but no direct-edit or simple-search method control. ## Install -Install the Node dependencies from the repository root: - -```sh -pnpm install -``` - -Install the Python bridge and the official optimizer packages: - -```sh -python -m pip install agent-eval-rpc -python -m pip install \ - "skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90" -python -m pip install \ - "gepa @ git+https://github.com/gepa-ai/gepa.git@f919db0a622e2e9f9204779b81fe00cc1b2d808f" \ - "litellm>=1.83.0,<1.92" \ - "tqdm>=4.66.1" \ - "cloudpickle>=3.0.0" \ - "datasets>=2.14.6" \ - "wandb" -``` - -Do not install `gepa[full]`; its MLflow server dependency is unpatched. - -From this repository, the locked equivalent is: +Run these commands from the repository root with Node, pnpm, Python, and uv installed: ```sh +pnpm install --frozen-lockfile cd clients/python uv sync --frozen --group skillopt-source --group gepa-source cd ../.. export OPTIMIZER_PYTHON="$PWD/clients/python/.venv/bin/python" ``` -## Run GEPA - -```sh -export LLM_BASE_URL=https://router.tangle.tools/v1 -export LLM_API_KEY="$TANGLE_API_KEY" -export LLM_MODEL=deepseek-v4-flash -export GEPA_PRICE_IN_PER_M=0.4 -export GEPA_PRICE_OUT_PER_M=1.6 +The lock selects the source versions tested by this bridge. +See the [Python guide](../../clients/python/README.md) for supported environments and dependency maintenance. -OPTIMIZERS=gepa pnpm tsx examples/compare-optimization-methods/index.ts -``` +## Configure the endpoint -Any OpenAI-compatible endpoint works; point `LLM_BASE_URL` at it and name a model it serves. +Export the following variables before running the script. +Use an endpoint that accepts this example's request fields and returns model identity and complete token usage. -Replace the example rates with the current exact endpoint rates. -GEPA uses `LLM_MODEL` by default. -Set `GEPA_MODEL` when reflection should use another model. -Optimizer reflection calls run through a default execution owner built from `LLM_BASE_URL` and `LLM_API_KEY`. -Set `OPTIMIZER_EXECUTION_OWNER_MODULE` to replace it with your own execution package. +| Variable | Purpose | +|---|---| +| `LLM_BASE_URL` | Chat Completions API prefix, such as `https://your-endpoint.example/v1`. | +| `LLM_API_KEY` | Key for the endpoint. | +| `LLM_MODEL` | Worker model served by the endpoint; defaults to `deepseek-v4-flash`. | +| `PRICE_IN_PER_M`, `PRICE_OUT_PER_M` | Worker USD rates per million tokens; supply both together. | +| `GEPA_PRICE_IN_PER_M`, `GEPA_PRICE_OUT_PER_M` | Reflection rates when GEPA uses different prices; otherwise inherit `PRICE_*`. | +| `SKILLOPT_PRICE_IN_PER_M`, `SKILLOPT_PRICE_OUT_PER_M` | Reflection/editing rates when SkillOpt uses different prices; otherwise inherit `PRICE_*`. | -### Choose a GEPA recipe +Use current rates for the actual endpoint and models. +Worker `PRICE_*` overrides are optional when the package already has suitable model pricing. +Each selected optimizer needs input and output rates, either its own or inherited worker rates. +Optional `PRICE_CACHED_IN_PER_M` and `PRICE_CACHE_WRITE_IN_PER_M` require the worker input/output pair. +The same cache suffixes are available under `GEPA_` and `SKILLOPT_`. -`GEPA_RECIPE` selects how GEPA composes engine runs. -The default, `engine`, is one budgeted run of the standard `gepa` engine. +## Run ```sh -GEPA_RECIPE=omni OPTIMIZERS=gepa pnpm tsx examples/compare-optimization-methods/index.ts +OPTIMIZERS=gepa pnpm exec tsx examples/compare-optimization-methods/index.ts ``` -| Kind | What runs | -|---|---| -| `engine` | One budgeted engine run. | -| `sequential` | Engines in order; the best result across stages is kept. | -| `adaptive-sequential` | Switch engines after a plateau, under one shared evaluation budget. | -| `best-of` | Independent engines; the highest selection score wins. | -| `vote` | Independent engines; GEPA's vote composition selects. | -| `omni` | Best-of exploration, then one continuation from its winner. | +```sh +OPTIMIZERS=skillopt pnpm exec tsx examples/compare-optimization-methods/index.ts +``` -The example splits `GEPA_MAX_EVALUATIONS` and `GEPA_MAX_PROPOSER_COST_USD` evenly across stages, so every recipe runs at the same total budget. -Every stage uses the standard `gepa` engine, which keeps the provider key outside Python. -The composed kinds require the tested GEPA source revision from the install step above; the published wheel supports `engine` only. -[`docs/campaign-proposers.md`](../../docs/campaign-proposers.md) documents recipes, other engines, budgets, and resuming. +```sh +OPTIMIZERS=gepa,skillopt pnpm exec tsx examples/compare-optimization-methods/index.ts +``` -## Run SkillOpt +Both methods are selected by default. +The default execution owner reads `LLM_BASE_URL` and `LLM_API_KEY` and keeps provider credentials out of the Python optimizer process. +Set `GEPA_MODEL` or `SKILLOPT_MODEL` when optimizer calls should use a different model from `LLM_MODEL`. -```sh -export LLM_BASE_URL=https://router.tangle.tools/v1 -export LLM_API_KEY="$TANGLE_API_KEY" -export LLM_MODEL=deepseek-v4-flash -export SKILLOPT_PRICE_IN_PER_M=0.4 -export SKILLOPT_PRICE_OUT_PER_M=1.6 +To supply your own execution owner, set `OPTIMIZER_EXECUTION_OWNER_MODULE` to an absolute path, file URL, or installed package specifier. +The module must export `createOptimizerExecutionOwner(model)` returning `{ call, callRef }`. +Avoid relative paths: dynamic imports resolve relative to the shared loader, not the repository root. +See [campaign proposers](../../docs/campaign-proposers.md#configure-gepa) for the execution callback contract. -OPTIMIZERS=skillopt pnpm tsx examples/compare-optimization-methods/index.ts -``` +## Choose a GEPA recipe -Set `SKILLOPT_PRICE_IN_PER_M` and `SKILLOPT_PRICE_OUT_PER_M` to the current exact rates for your endpoint before running SkillOpt. -The example passes SkillOpt's `openai_compatible` traffic through Agent Eval's local proxy and then through the execution owner. -By default that owner is this repository's example owner, `examples/_shared/openai-compatible-owner.ts`, built from `LLM_BASE_URL` and `LLM_API_KEY`. Agent Eval owns no model transport; on agent-runtime the production owner is `profileOptimizerModelCall`. -An `OPTIMIZER_EXECUTION_OWNER_MODULE` override must export `createOptimizerExecutionOwner(model)` and return `{ call, callRef }`. -Discovery uses this module boundary to execute the model through Runtime with one exact AgentProfile. -Set `SKILLOPT_MODEL` to use a different optimizer model. +`GEPA_RECIPE` defaults to `engine`. +The other recipes compose standard GEPA engines and require the source dependencies installed above. -## Compare Both +| Value | Behavior | +|---|---| +| `engine` | One bounded engine run. | +| `sequential` | Run stages in order and keep the best selection result. | +| `adaptive-sequential` | Switch after a plateau under one shared evaluation limit. | +| `best-of` | Run independent stages and keep the highest selection score. | +| `vote` | Use GEPA's vote composition across stages. | +| `omni` | Explore with best-of, then continue from its winner. | ```sh -OPTIMIZERS=gepa,skillopt \ -LLM_BASE_URL=https://router.tangle.tools/v1 \ -LLM_API_KEY="$TANGLE_API_KEY" \ -LLM_MODEL=deepseek-v4-flash \ -GEPA_PRICE_IN_PER_M=0.4 \ -GEPA_PRICE_OUT_PER_M=1.6 \ -SKILLOPT_PRICE_IN_PER_M=0.4 \ -SKILLOPT_PRICE_OUT_PER_M=1.6 \ -pnpm tsx examples/compare-optimization-methods/index.ts +GEPA_RECIPE=omni OPTIMIZERS=gepa pnpm exec tsx examples/compare-optimization-methods/index.ts ``` -The execution owner controls the optimizer endpoint and credentials; Agent Eval's proxy never receives them. -Replace all four example rates with the exact rates charged by that endpoint. +Composed recipes allocate evaluation and proposer limits across their stages. +Integer rounding can leave evaluation capacity unused: a two-stage recipe receives 16 evaluations per stage at the default ceiling of 33. +Equal configured ceilings do not imply equal realized evaluations, model calls, tokens, or spend. +The [implementation](./index.ts) records the selected recipe and limits with the result. + +## Control work and spend + +| Variable | Default | Purpose | +|---|---|---| +| `GEPA_MAX_EVALUATIONS` | SkillOpt core plan size, initially `33` | Maximum GEPA candidate-case evaluations. | +| `SKILLOPT_MAX_EVALUATIONS` | Core plan size, initially `33` | Maximum SkillOpt candidate-case evaluations. | +| `SKILLOPT_EPOCHS`, `SKILLOPT_BATCH_SIZE` | `2`, `2` | Trainer settings that determine the default evaluation plan. | +| `MAX_OPTIMIZER_MODEL_COST_USD` | `5` | Default optimizer model spend limit per method. | +| `GEPA_MAX_PROPOSER_COST_USD` | `5` | GEPA proposer ceiling, allocated across recipe stages. | +| `GEPA_MAX_MODEL_COST_USD`, `SKILLOPT_MAX_MODEL_COST_USD` | `MAX_OPTIMIZER_MODEL_COST_USD` | Model spend limits for each optimizer. | +| `GEPA_MAX_MODEL_REQUESTS`, `SKILLOPT_MAX_MODEL_REQUESTS` | `100` | Maximum model requests per optimizer. | +| `MAX_TOTAL_COST_USD` | `20` | Shared ledger ceiling for search and final evaluation across all methods. | +| `OPTIMIZATION_CONCURRENCY` | `1` | Methods allowed to search concurrently. | +| `LLM_MAX_TOKENS` | `400` | Worker output cap; allow extra headroom for reasoning models. | +| `CALL_TIMEOUT_MS` | `30000` | Worker timeout per call. | +| `OPTIMIZER_PYTHON` | `python` | Python executable containing the bridge and selected optimizers. | -## Controls +When both methods run, the script requires matching candidate-case evaluation limits. +Model request and spend limits share defaults but can be overridden separately; record any differences. +The shared whole-run ceiling can still exhaust before a later method or final evaluation finishes. +Inspect actual usage and completion before describing a comparison as matched on resources. +The [budget helper](../_shared/optimizer-model-budget.ts) exposes additional request-byte, response-byte, token, and timeout controls. -| Variable | Default | Meaning | -|---|---:|---| -| `OPTIMIZERS` | `gepa,skillopt` | Comma-separated methods to run. | -| `LLM_MODEL` | `deepseek-v4-flash` | Worker model; must be served by `LLM_BASE_URL`. | -| `LLM_MAX_TOKENS` | `400` | Output cap per worker call; raise it for reasoning models. | -| `OPTIMIZER_PYTHON` | `python` | Python executable containing the bridge and selected optimizers. | -| `OPTIMIZER_EXECUTION_OWNER_MODULE` | built-in OpenAI-compatible owner | Module exporting `createOptimizerExecutionOwner(model)`; use Runtime for Discovery. | -| `GEPA_MODEL` | `LLM_MODEL` | Endpoint model used by GEPA reflection. | -| `GEPA_MAX_EVALUATIONS` | SkillOpt core plan size | Maximum GEPA candidate-case calls. Must match SkillOpt when both run. | -| `GEPA_MAX_PROPOSER_COST_USD` | `5` | Maximum GEPA model spend inside one engine stage. | -| `GEPA_PRICE_IN_PER_M` | required | Exact GEPA input rate per million tokens. | -| `GEPA_PRICE_OUT_PER_M` | required | Exact GEPA output rate per million tokens. | -| `GEPA_MAX_MODEL_COST_USD` | `MAX_OPTIMIZER_MODEL_COST_USD` | GEPA model spend limit. | -| `GEPA_MAX_MODEL_REQUESTS` | `100` | Shared GEPA model request limit. | -| `SKILLOPT_MODEL` | `LLM_MODEL` | Model used by SkillOpt reflection and editing. | -| `SKILLOPT_EPOCHS` | `2` | SkillOpt training epochs. | -| `SKILLOPT_BATCH_SIZE` | `2` | SkillOpt train cases per step. | -| `SKILLOPT_MAX_EVALUATIONS` | core plan size | Maximum SkillOpt candidate-case calls. | -| `SKILLOPT_PRICE_IN_PER_M` | required | Exact optimizer-model input rate per million tokens. | -| `SKILLOPT_PRICE_OUT_PER_M` | required | Exact optimizer-model output rate per million tokens. | -| `SKILLOPT_MAX_MODEL_COST_USD` | `MAX_OPTIMIZER_MODEL_COST_USD` | SkillOpt optimizer-model spend limit. | -| `SKILLOPT_MAX_MODEL_REQUESTS` | `100` | SkillOpt optimizer-model request limit. | -| `MAX_OPTIMIZER_MODEL_COST_USD` | `5` | Equal optimizer-model spend limit per method. | -| `MAX_TOTAL_COST_USD` | `20` | Shared limit for all optimization and final-case spend. | -| `OPTIMIZATION_CONCURRENCY` | `1` | Methods allowed to optimize concurrently. | -| `BILLING_NOTE` | inferred | Billing context saved with the result. | -| `PRICE_SOURCE` | inferred | Source of the token prices saved with the result. | - -The result is written to `.evolve/compare-optimization-methods//comparison.json` and mirrored to `.evolve/compare-optimization-methods/latest.json`. -It includes every method's selected surface, final-case scores, paired lift interval, duration, cost status, run limits, token prices, upstream package revision, run identity, token usage, and source model configuration. -Optimizer model spend uses provider-reported billed cost when present. -Otherwise it is estimated from complete token usage and the configured token rates. -`accountingComplete` means every call was priced; it does not mean the total was reconciled to an invoice. -The run fails when the endpoint omits usage instead of publishing an incomplete comparison. -Set `BILLING_NOTE` and `PRICE_SOURCE` when declared token prices estimate subscription usage rather than actual billed dollars. +## Inspect the artifacts + +The script writes `.evolve/compare-optimization-methods//comparison.json` and mirrors it to `.evolve/compare-optimization-methods/latest.json`. +Raw campaign artifacts remain under the timestamped directory. +The summary contains selected surfaces, final scores, lift intervals, cost status, configured limits, optimizer provenance, and available token usage. +Read `comparison.pairwise` and each method's decision before interpreting its rank as evidence of a difference. +A positive descriptive interval can still be ineligible for promotion. + +Provider-reported billed cost takes precedence when present. +Otherwise complete usage and declared token rates produce an estimate. +`accountingComplete` means each call was priced; it does not establish reconciliation with an invoice. +Missing required usage fails the comparison. +Set `BILLING_NOTE` and `PRICE_SOURCE` to retain the origin and interpretation of supplied rates. + +Keep final cases outside optimizer closures, shared files, and prior search memory when adapting the example. +For repeated or grouped tasks, declare the independent unit and use optional fresh-evidence controls described in [evaluation integrity](../../docs/evaluation-integrity.md). diff --git a/examples/compare-optimization-methods/index.ts b/examples/compare-optimization-methods/index.ts index c62b44384..cca43534b 100644 --- a/examples/compare-optimization-methods/index.ts +++ b/examples/compare-optimization-methods/index.ts @@ -26,7 +26,7 @@ import { } from '../../src/campaign' import { assertRealBackend, summarizeBackendIntegrity } from '../../src/integrity/backend-integrity' import type { RunRecord } from '../../src/run-record' -import { optionalNonNegativeNumberEnv, positiveIntegerEnv, positiveNumberEnv } from '../_shared/env' +import { positiveIntegerEnv, positiveNumberEnv, tokenPricingFromEnv } from '../_shared/env' import { type Artifact, BASELINE_SURFACE, @@ -52,10 +52,7 @@ const MODEL = process.env.LLM_MODEL || 'deepseek-v4-flash' const OPTIMIZER_PYTHON = process.env.OPTIMIZER_PYTHON?.trim() || 'python' const GEPA_MODEL = process.env.GEPA_MODEL || MODEL const SKILLOPT_MODEL = process.env.SKILLOPT_MODEL || MODEL -const PRICE_IN_PER_M = optionalNonNegativeNumberEnv('PRICE_IN_PER_M') -const PRICE_CACHED_IN_PER_M = optionalNonNegativeNumberEnv('PRICE_CACHED_IN_PER_M') -const PRICE_CACHE_WRITE_IN_PER_M = optionalNonNegativeNumberEnv('PRICE_CACHE_WRITE_IN_PER_M') -const PRICE_OUT_PER_M = optionalNonNegativeNumberEnv('PRICE_OUT_PER_M') +const customTokenPricing = tokenPricingFromEnv() const CALL_TIMEOUT_MS = positiveIntegerEnv('CALL_TIMEOUT_MS', 30_000) // Reasoning models spend thinking tokens against this cap; raise it for // families that reason, or the worker returns truncated JSON. @@ -153,15 +150,6 @@ if (!API_KEY) { if (!BASE_URL) { throw new Error('Set LLM_BASE_URL (or TANGLE_ROUTER_URL) to an OpenAI-compatible endpoint.') } -if ((PRICE_IN_PER_M === undefined) !== (PRICE_OUT_PER_M === undefined)) { - throw new Error('PRICE_IN_PER_M and PRICE_OUT_PER_M must be set together') -} -if ( - PRICE_IN_PER_M === undefined && - (PRICE_CACHED_IN_PER_M !== undefined || PRICE_CACHE_WRITE_IN_PER_M !== undefined) -) { - throw new Error('Cache token rates require PRICE_IN_PER_M and PRICE_OUT_PER_M') -} if (SKILLOPT_MAX_EVALUATIONS < SKILLOPT_CORE_EVALUATIONS) { throw new Error( `SKILLOPT_MAX_EVALUATIONS must be at least ${SKILLOPT_CORE_EVALUATIONS} for this split and trainer plan`, @@ -184,19 +172,6 @@ assertMatchedMethodLimits( 'Candidate-case evaluation limits', ) -const customTokenPricing = - PRICE_IN_PER_M === undefined || PRICE_OUT_PER_M === undefined - ? undefined - : { - inputUsdPerMillion: PRICE_IN_PER_M, - ...(PRICE_CACHED_IN_PER_M === undefined - ? {} - : { cachedInputUsdPerMillion: PRICE_CACHED_IN_PER_M }), - ...(PRICE_CACHE_WRITE_IN_PER_M === undefined - ? {} - : { cacheWriteUsdPerMillion: PRICE_CACHE_WRITE_IN_PER_M }), - outputUsdPerMillion: PRICE_OUT_PER_M, - } const BILLING_NOTE = process.env.BILLING_NOTE?.trim() || (customTokenPricing @@ -214,7 +189,7 @@ const skillOptModelBudget = selectedNames.includes('skillopt') ? optimizerModelBudgetFromEnv('SKILLOPT', MAX_OPTIMIZER_MODEL_COST_USD, customTokenPricing) : undefined -// The worker transport is caller code: Agent Eval holds no provider key. +// The execution owner binds the caller's endpoint and credential. const chat = openAiCompatibleChatClient({ baseUrl: BASE_URL, apiKey: API_KEY, diff --git a/examples/distributed-driver/driver.ts b/examples/distributed-driver/driver.ts index ac4bfa187..65ccaf26f 100644 --- a/examples/distributed-driver/driver.ts +++ b/examples/distributed-driver/driver.ts @@ -135,8 +135,7 @@ async function main() { storage: inMemoryCampaignStorage(), runDir: `mem://distributed-coordinator-${Date.now()}`, maxConcurrency: 4, - // The demo worker is a stub that reports no usage. Keep the default - // 'assert' once workers meter real calls through ctx.cost.runPaidCall. + // The demo worker reports no paid usage. Set 'assert' for metered workers. expectUsage: 'off', cellPlacement: IS_MULTIREGION ? ({ scenario }) => { diff --git a/examples/evaluate-a-change/README.md b/examples/evaluate-a-change/README.md index 18bff8d41..a6dd3f814 100644 --- a/examples/evaluate-a-change/README.md +++ b/examples/evaluate-a-change/README.md @@ -9,9 +9,12 @@ Start here before any optimizer. ## How to run it ```sh -pnpm tsx examples/evaluate-a-change/index.ts +pnpm install +pnpm build +pnpm exec tsx examples/evaluate-a-change/index.ts ``` +Run these commands from the repository root. No API key is required. The agent and the judge are local functions. @@ -25,18 +28,22 @@ The agent and the judge are local functions. The output is: ```text -baseline: { 'ticket-id': { mean: 0, stdev: 0, ci95: [ 0, 0 ], n: 3 } } -candidate: { 'ticket-id': { mean: 1, stdev: 0, ci95: [ 1, 1 ], n: 3 } } +baseline: 0 +candidate: 1 ``` +The example prints each mean. +The returned aggregates also contain counts, intervals, and score distributions. + ## Why it is built this way The surface is the only value that changes between the two calls. -The cases, the agent, and the judge stay identical, so the score difference measures the change and nothing else. +The cases, agent, and judge stay fixed in this deterministic fixture. +The changed surface accounts for its score difference. +Comparisons of real agents also need to account for execution variability and missing evidence. `expectUsage: 'off'` is set because this agent makes no paid model calls. -The default is `'assert'`, which fails a run whose cells report no cost receipt. -Keep the default whenever real model calls happen: it is the check that stops an unmeasured run from reading as a free one. +Set `expectUsage: 'assert'` when connecting a paid agent so missing dispatch receipts become execution failures. ## Next diff --git a/examples/evaluate-a-change/index.ts b/examples/evaluate-a-change/index.ts index 01a81b208..239caf4f1 100644 --- a/examples/evaluate-a-change/index.ts +++ b/examples/evaluate-a-change/index.ts @@ -1,13 +1,13 @@ /** * The smallest complete evaluation: one surface, one judge, two scores. * - * Run with: pnpm tsx examples/evaluate-a-change/index.ts + * Run after building: pnpm exec tsx examples/evaluate-a-change/index.ts * * Everything here is offline. Replace `agent` with your product call and * `judge` with your real scoring function to point this at production. */ -import { defineAgentEval } from '../../src/contract' +import { defineAgentEval } from '@tangle-network/agent-eval/contract' interface SupportCase { id: string @@ -40,5 +40,5 @@ const candidate = await evalKit.evaluate({ surface: 'Answer politely and cite the ticket id.', }) -console.log('baseline: ', baseline.aggregates.byJudge) -console.log('candidate:', candidate.aggregates.byJudge) +console.log('baseline:', baseline.aggregates.byJudge['ticket-id']?.mean) +console.log('candidate:', candidate.aggregates.byJudge['ticket-id']?.mean) diff --git a/examples/foreign-agent-quickstart/README.md b/examples/foreign-agent-quickstart/README.md index df915d197..ab27f8b25 100644 --- a/examples/foreign-agent-quickstart/README.md +++ b/examples/foreign-agent-quickstart/README.md @@ -1,92 +1,113 @@ -# Use Agent Eval With An Existing Agent +# Evaluate an existing agent -Your agent does not need to use Tangle's runtime or sandbox. -Agent Eval only needs a function that accepts the candidate prompt or configuration plus one scenario, then returns the artifact to score. +Your agent does not need Tangle's runtime or sandbox. +Adapt its input and output to the `agent` callback of `defineAgentEval()`. +The [complete example](./index.ts) wraps a local support-agent fixture and compares two prompts. -## Install +## Run the example + +From the repository root, with Node.js 20.19 or newer and pnpm installed: ```sh -npm install @tangle-network/agent-eval +pnpm install --frozen-lockfile +pnpm build +pnpm exec tsx examples/foreign-agent-quickstart/index.ts ``` -## Run The Example +```text +baseline: 0.5 +candidate: 1 +``` -From this repository: +The fixture runs offline and needs no credentials. +It checks two policy answers for a required fact and citation. +These checks demonstrate the adapter; they do not establish general answer quality or a release decision. -```sh -pnpm tsx examples/foreign-agent-quickstart/index.ts +## Connect your agent + +Replace `existingAgent()` with your SDK, service, workflow, or local model. +The adapter passes the candidate configuration, scenario input, and cancellation signal. +It returns the artifact that the judge scores: + +```ts +agent: async (surface, scenario, ctx) => { + const result = await existingAgent({ + instructions: String(surface), + question: scenario.question, + signal: ctx.signal, + }) + return { text: result.message } +}, ``` -The default path is offline and deterministic. -Set `LLM_API_KEY` and `LLM_BASE_URL` to run the agent and judge through an OpenAI-compatible endpoint. -`LLM_MODEL` selects the model; the default is `deepseek-v4-flash`. +Keep answer keys in the judge's input, outside the agent request. +Replace the fixture's string checks with checks calibrated for your task. +Throw when execution cannot produce a valid artifact. +Eval records execution failures separately from measured task scores. -## Adapt Your Agent +## Record paid calls -Define four inputs: +The offline fixture uses `expectUsage: 'off'` because it makes no paid calls. +Set `expectUsage: 'assert'` when connecting a paid agent. +`evaluate()` otherwise defaults to a warning for missing dispatch receipts; `selfImprove()` defaults to `'assert'`. +Wrap each model call with `ctx.cost.runPaidCall()` so usage and failures enter the run's ledger. -1. `scenarios`: representative tasks with stable IDs. -2. `agent`: calls your existing SDK, service, workflow, or local model. -3. `judge`: returns dimension scores and a composite score from `0` to `1`. -4. `baselineSurface`: the prompt or configuration you use today. +For an OpenAI-compatible endpoint, use the maintained `ChatClient` and receipt helpers. +This factory replaces the example's `agent` callback and uses its `SupportCase` and `SupportAnswer` types: ```ts import { - defineAgentEval, - type JudgeConfig, - type Scenario, + costReceiptFromLlm, + costReceiptFromLlmError, + type CustomTokenPricing, + maximumChargeForLlmRequest, +} from '@tangle-network/agent-eval' +import type { + ChatClient, + DispatchContext, + MutableSurface, } from '@tangle-network/agent-eval/contract' -interface SupportScenario extends Scenario { - question: string +function paidAgent(chat: ChatClient, model: string, pricing: CustomTokenPricing) { + return async ( + surface: MutableSurface, + scenario: SupportCase, + ctx: DispatchContext, + ): Promise => { + const request = { + model, + messages: [ + { role: 'system' as const, content: String(surface) }, + { role: 'user' as const, content: scenario.question }, + ], + maxTokens: 1000, + } + const paid = await ctx.cost.runPaidCall({ + actor: 'support-agent', + model, + maximumCharge: maximumChargeForLlmRequest(request, { + customTokenPricing: pricing, + maximumAttempts: chat.maximumAttempts, + }), + execute: (signal, callId) => chat.chat(request, { signal, idempotencyKey: callId }), + receipt: (result) => costReceiptFromLlm(result, pricing), + receiptFromError: (error) => costReceiptFromLlmError(error, pricing), + }) + if (!paid.succeeded) throw paid.error + return { text: paid.value.content } + } } - -interface SupportAnswer { - text: string -} - -const judge: JudgeConfig = { - name: 'support-quality', - dimensions: [ - { key: 'correct', description: 'The answer resolves the question correctly' }, - ], - score: ({ artifact }) => { - const correct = artifact.text.includes('expected fact') ? 1 : 0 - return { dimensions: { correct }, composite: correct, notes: '' } - }, -} - -const evalKit = defineAgentEval({ - scenarios, - baselineSurface: 'Answer accurately and briefly.', - agent: async (surface, scenario, ctx) => ({ - text: await yourAgent({ - systemPrompt: String(surface), - question: scenario.question, - signal: ctx.signal, - }), - }), - judge, - expectUsage: 'off', -}) - -const baseline = await evalKit.evaluate() -const candidate = await evalKit.evaluate({ - surface: 'Answer accurately, briefly, and cite the relevant policy.', -}) ``` -Wrap top-level calls in your application's existing async entry point. -Use `evalKit.improve()` with a caller-owned `SurfaceProposer` when you are ready to generate and evaluate candidates automatically. -Use `compareOptimizationMethods()` when official GEPA or SkillOpt should own the complete search procedure. - -## Failure And Data Handling +Configure the client as shown in the [main README](../../README.md#configure-model-calls). +Pass current endpoint rates to `paidAgent()` and set `agent: paidAgent(chat, model, pricing)`. +The maximum charge reserves for output limits and all declared transport attempts before a capped call starts. +An SDK adapter must return actual usage, preserve missing usage as unknown, and expose its maximum attempts. +For a multi-call agent, meter each call inside the agent where its receipt is available. -- Throw when the agent or judge cannot produce a valid result. -- Pass `ctx.signal` to downstream calls so cancellation stops work already in progress. -- Use in-memory storage for ephemeral runs or filesystem storage for resumable local runs. -- No hosted service is required. -- Data leaves your process only through providers and exporters you configure. +## Next steps -The runnable implementation is [`index.ts`](./index.ts). -For lower-level campaign control, read [`docs/campaign-proposers.md`](../../docs/campaign-proposers.md). +Use filesystem campaign storage when runs must survive process restarts. +Use [`selfImprove()`](../selfimprove-quickstart/) when you need candidate generation and a final comparison. +Use [`compareOptimizationMethods()`](../compare-optimization-methods/) when an official optimizer owns the complete search procedure. +For campaign scheduling and custom release rules, read [campaign proposers](../../docs/campaign-proposers.md). diff --git a/examples/foreign-agent-quickstart/index.ts b/examples/foreign-agent-quickstart/index.ts index 0676eace3..1e79e5beb 100644 --- a/examples/foreign-agent-quickstart/index.ts +++ b/examples/foreign-agent-quickstart/index.ts @@ -1,318 +1,86 @@ /** - * Wrap an existing agent, evaluate it, and improve its prompt with a - * caller-owned SurfaceProposer. - * - * Run: pnpm tsx examples/foreign-agent-quickstart/index.ts + * Adapt an existing agent's input and output to one evaluation. + * Run after building: pnpm exec tsx examples/foreign-agent-quickstart/index.ts */ -// IN-REPO: relative imports so the example typechecks against the workspace. -// COPY-PASTE INTO YOUR OWN PROJECT: change this to -// import { ... } from '@tangle-network/agent-eval/contract' -// The public subpath exposes these names with the same shapes. -import { - type Dispatch, - defaultProductionGate, - inMemoryCampaignStorage, - type JudgeConfig, - type MutableSurface, - runEval, - runImprovementLoop, - type Scenario, - type SurfaceProposer, -} from '../../src/contract' +import { defineAgentEval, type Scenario } from '@tangle-network/agent-eval/contract' -// 1. Representative cases. +interface SupportCase extends Scenario { + question: string + expectedAnswer: string + policyUrl: string +} -interface MarketingScenario extends Scenario { - blurb: string - surface: 'landing-hero' | 'tweet' | 'email-subject' - audience: string +interface SupportAnswer { + text: string } -const scenarios: MarketingScenario[] = [ - { - id: 's1', - kind: 'marketing-rewrite', - blurb: 'We help teams ship software faster with AI.', - surface: 'landing-hero', - audience: 'engineering leaders', - tags: ['saas', 'developer-tools'], - }, - { - id: 's2', - kind: 'marketing-rewrite', - blurb: 'Our note-taking app uses machine learning to organize your thoughts.', - surface: 'tweet', - audience: 'consumer prosumers', - tags: ['consumer', 'productivity'], - }, +const scenarios: SupportCase[] = [ { - id: 's3', - kind: 'marketing-rewrite', - blurb: 'Track expenses, file taxes, get refunds. Powered by AI.', - surface: 'email-subject', - audience: 'small-business owners', - tags: ['fintech', 'tax'], + id: 'refund', + kind: 'support', + question: 'How long do I have to request a refund?', + expectedAnswer: '30 days', + policyUrl: 'https://support.example/refunds', }, { - id: 's4', - kind: 'marketing-rewrite', - blurb: 'Generate marketing copy with our AI agent, faster and on-brand.', - surface: 'landing-hero', - audience: 'marketing teams', - tags: ['meta', 'marketing-tools'], + id: 'cancel', + kind: 'support', + question: 'When does cancellation take effect?', + expectedAnswer: 'end of the billing period', + policyUrl: 'https://support.example/cancellation', }, ] -// ── 2. Your agent, wrapped as a Dispatch ──────────────────────────── - -interface MarketingArtifact { - rewrite: string - modelUsed: string -} - -const apiKey = process.env.LLM_API_KEY -const baseUrl = process.env.LLM_BASE_URL -const modelId = process.env.LLM_MODEL ?? 'deepseek-v4-flash' -if (apiKey && !baseUrl) { - throw new Error('LLM_API_KEY is set but LLM_BASE_URL is not; set both to go online.') -} - -const baselineSystemPrompt = `You are a senior copywriter. Rewrite the given product blurb for the surface (landing-hero / tweet / email-subject) and audience. One sentence for tweets and email subjects, two for landing hero. Be concrete, not generic; no AI slop ("revolutionary", "powerful", "seamless"). Lead with the value, not the technology.` - -const customProposer: SurfaceProposer = { - kind: 'marketing-constraints', - async propose({ currentSurface, populationSize }) { - const current = String(currentSurface) - return [ - { - surface: `${current}\nName the audience in the opening phrase.`, - label: 'name-audience', - rationale: 'Training outputs did not make the audience explicit.', - }, - { - surface: `${current}\nEnd with one explicit next step.`, - label: 'explicit-next-step', - rationale: 'Training outputs lacked a clear action.', - }, - ].slice(0, populationSize) - }, -} - -async function callLLM(system: string, user: string, signal: AbortSignal): Promise { - if (!apiKey) { - const audience = /Audience: ([^\n]+)/.exec(user)?.[1] ?? 'reader' - const blurb = /Blurb: ([^\n]+)/.exec(user)?.[1] ?? user - const prefix = system.includes('Name the audience') ? `For ${audience}: ` : '' - const suffix = system.includes('explicit next step') ? ' Try it today.' : '' - return `${prefix}${blurb}${suffix}` - } - const res = await fetch(`${baseUrl}/chat/completions`, { - method: 'POST', - headers: { 'Content-Type': 'application/json', Authorization: `Bearer ${apiKey}` }, - body: JSON.stringify({ - model: modelId, - messages: [ - { role: 'system', content: system }, - { role: 'user', content: user }, - ], - temperature: 0.7, - }), - signal, - }) - if (!res.ok) throw new Error(`LLM call failed: ${res.status} ${await res.text()}`) - const data = (await res.json()) as { choices: { message: { content: string } }[] } - return data.choices[0]?.message?.content ?? '' -} - -function buildDispatch(systemPrompt: string): Dispatch { - return async (scenario, ctx) => { - const user = `Surface: ${scenario.surface}\nAudience: ${scenario.audience}\nBlurb: ${scenario.blurb}` - const rewrite = await callLLM(systemPrompt, user, ctx.signal) - return { rewrite: rewrite.trim(), modelUsed: apiKey ? modelId : 'stub' } +// This local fixture stands in for your existing SDK or service. +// Its interface has no dependency on Agent Eval or the answer-key fields. +async function existingAgent(input: { + instructions: string + question: string + signal: AbortSignal +}): Promise<{ message: string }> { + input.signal.throwIfAborted() + const refund = input.question.toLowerCase().includes('refund') + const answer = refund + ? 'Request a refund within 30 days.' + : 'Cancellation takes effect at the end of the billing period.' + const url = refund ? 'https://support.example/refunds' : 'https://support.example/cancellation' + return { + message: input.instructions.includes('cite') ? `${answer} Policy: ${url}` : answer, } } -// 3. The scoring rule. - -const judge: JudgeConfig = { - name: 'marketing-quality', - dimensions: [ - { - key: 'hook_strength', - description: 'Opens with a concrete value claim, not a category description.', - }, - { - key: 'voice_match', - description: - 'Avoids AI slop ("revolutionary", "powerful", "seamless"); reads like a human wrote it.', - }, - { key: 'cta_clarity', description: 'Makes the next step obvious for the named audience.' }, - { - key: 'factual_grounding', - description: - 'Claims only what the blurb says or what is obviously true. No invented features.', - }, - ], - async score({ artifact, scenario, signal }) { - if (!apiKey) { - // Heuristic judge so the wiring is verifiable without an LLM key. - const text = artifact.rewrite.toLowerCase() - const slop = ['revolutionary', 'powerful', 'seamless', 'cutting-edge', 'next-gen'].filter( - (w) => text.includes(w), - ) - const surfaceTargets: Record = { - 'landing-hero': [60, 180], - tweet: [40, 140], - 'email-subject': [20, 80], - } - const [lo, hi] = surfaceTargets[scenario.surface] - const lenOk = artifact.rewrite.length >= lo && artifact.rewrite.length <= hi - const audienceHit = text.includes(scenario.audience.split(' ')[0]?.toLowerCase() ?? '') - const base = 0.5 - const slopPenalty = slop.length * 0.1 - const lenBonus = lenOk ? 0.2 : 0 - const audienceBonus = audienceHit ? 0.15 : 0 - const hook = Math.max(0, Math.min(1, base - slopPenalty + lenBonus + audienceBonus)) - const voice = Math.max(0, 1 - slopPenalty * 2) - const cta = audienceBonus > 0 ? 0.7 : 0.4 - const grounding = 0.7 - const dims = { - hook_strength: hook, - voice_match: voice, - cta_clarity: cta, - factual_grounding: grounding, - } - const composite = (hook + voice + cta + grounding) / 4 - return { - dimensions: dims, - composite, - notes: `heuristic: slop=${slop.length} lenOk=${lenOk} audience=${audienceHit}`, - } - } - const judgePrompt = `Score the rewrite below on 4 dimensions, 0.0 to 1.0. Return strict JSON: -{"hook_strength": n, "voice_match": n, "cta_clarity": n, "factual_grounding": n, "notes": "one sentence"} - -Dimensions: -- hook_strength: Opens with a concrete value claim, not a category description. -- voice_match: Avoids AI slop; human-sounding. -- cta_clarity: Makes the next step obvious for the named audience. -- factual_grounding: Claims only what the blurb says or what is obviously true. - -Surface: ${scenario.surface} -Audience: ${scenario.audience} -Original: ${scenario.blurb} -Rewrite: ${artifact.rewrite}` - const raw = await callLLM( - 'You are a strict copywriting judge. Respond with only JSON.', - judgePrompt, - signal, - ) - const match = raw.match(/\{[\s\S]*\}/) - if (!match) throw new Error(`Judge returned non-JSON: ${raw.slice(0, 200)}`) - const parsed = JSON.parse(match[0]) as { - hook_strength: number - voice_match: number - cta_clarity: number - factual_grounding: number - notes?: string - } - const dims = { - hook_strength: parsed.hook_strength, - voice_match: parsed.voice_match, - cta_clarity: parsed.cta_clarity, - factual_grounding: parsed.factual_grounding, - } - const composite = - (dims.hook_strength + dims.voice_match + dims.cta_clarity + dims.factual_grounding) / 4 - return { dimensions: dims, composite, notes: parsed.notes ?? '' } +const evalKit = defineAgentEval({ + scenarios, + baselineSurface: 'Answer the support question.', + agent: async (surface, scenario, ctx) => { + const result = await existingAgent({ + instructions: String(surface), + question: scenario.question, + signal: ctx.signal, + }) + return { text: result.message } }, -} - -// 4. Helpers. - -function meanComposite(result: { - aggregates: { byScenario: Record } -}): number { - const vs = Object.values(result.aggregates.byScenario).map((s) => s.meanComposite) - return vs.length === 0 ? 0 : vs.reduce((a, b) => a + b, 0) / vs.length -} - -// 5. Run the baseline and improvement loop. - -async function main() { - const storage = inMemoryCampaignStorage() - const runDir = `mem://quickstart-${Date.now()}` - - console.log('Baseline evaluation') - const baseline = await runEval({ - scenarios, - dispatch: buildDispatch(baselineSystemPrompt), - judges: [judge], - storage, - runDir, - dispatchRef: 'foreign-agent-baseline', - // This demo's callLLM does not report usage through ctx.cost.runPaidCall, - // so no cell can carry a receipt. Wire runPaidCall and keep the default - // 'assert' when you adapt it to a metered production agent. - expectUsage: 'off', - }) - const baselineScore = meanComposite(baseline) - console.log(`Baseline composite mean: ${baselineScore.toFixed(3)}`) - console.log( - `Cells executed: ${baseline.aggregates.cellsExecuted}, cost: $${baseline.aggregates.cost.totalCostUsd.toFixed(4)}`, - ) - - console.log('\nImprovement loop with a custom candidate generator') - const holdout = scenarios.slice(0, 2) - const train = scenarios.slice(2) - - const result = await runImprovementLoop({ - scenarios: train, - baselineSurface: baselineSystemPrompt, - dispatchWithSurface: async (surface, scenario, ctx) => { - const prompt = typeof surface === 'string' ? surface : JSON.stringify(surface) - return buildDispatch(prompt)(scenario, ctx) + judge: { + name: 'support-answer', + dimensions: [ + { key: 'correct', description: 'The answer contains the expected policy fact' }, + { key: 'cited', description: 'The answer links to the relevant policy' }, + ], + score: ({ artifact, scenario }) => { + const correct = artifact.text.includes(scenario.expectedAnswer) ? 1 : 0 + const cited = artifact.text.includes(scenario.policyUrl) ? 1 : 0 + return { dimensions: { correct, cited }, composite: (correct + cited) / 2, notes: '' } }, - proposer: customProposer, - judges: [judge], - populationSize: 2, - maxGenerations: 1, - holdoutScenarios: holdout, - gate: defaultProductionGate({ - holdoutScenarios: holdout, - deltaThreshold: 0.05, - }), - autoOnPromote: 'none', - storage, - runDir: `${runDir}/improve`, - dispatchRef: 'foreign-agent-with-surface', - expectUsage: 'off', - }) - - const winnerScore = meanComposite(result.winnerOnHoldout) - const baselineHoldoutScore = meanComposite(result.baselineOnHoldout) - const lift = winnerScore - baselineHoldoutScore - console.log(`Generations explored: ${result.generations.length}`) - console.log(`Gate decision: ${result.gateResult.decision}`) - console.log(`Holdout baseline: ${baselineHoldoutScore.toFixed(3)}`) - console.log( - `Holdout winner: ${winnerScore.toFixed(3)} (lift ${lift >= 0 ? '+' : ''}${lift.toFixed(3)})`, - ) - - const shipped: MutableSurface | null = - result.gateResult.decision === 'ship' ? result.winnerSurface : null - if (shipped) { - const prompt = typeof shipped === 'string' ? shipped : JSON.stringify(shipped, null, 2) - console.log(`\n--- Shipped prompt ---\n${prompt}\n`) - } else { - console.log( - `\nThe release rule returned ${result.gateResult.decision}. Revise the candidate logic or threshold before promotion.`, - ) - } -} + }, + // The local fixture makes no paid calls. Meter real calls and set 'assert'. + expectUsage: 'off', +}) -main().catch((err) => { - console.error(err) - process.exit(1) +const baseline = await evalKit.evaluate() +const candidate = await evalKit.evaluate({ + surface: 'Answer the support question and cite the relevant policy.', }) + +console.log('baseline:', baseline.aggregates.byJudge['support-answer']?.mean) +console.log('candidate:', candidate.aggregates.byJudge['support-answer']?.mean) diff --git a/examples/self-improve-optimizer/README.md b/examples/self-improve-optimizer/README.md index a1e611b9d..14ab7a580 100644 --- a/examples/self-improve-optimizer/README.md +++ b/examples/self-improve-optimizer/README.md @@ -1,98 +1,95 @@ -# Improve One Prompt With Official GEPA +# Improve one prompt with GEPA -This example calls `selfImprove()` with a `gepaOptimizationMethod()` method. -One call runs the complete path: GEPA searches on train and selection partitions, Agent Eval re-scores the selected prompt on a held-out split GEPA never received, and the promotion gate returns a release decision. +This example calls `selfImprove()` with `gepaOptimizationMethod()` to search a transaction-extraction prompt. +GEPA receives separate train and selection partitions. +Agent Eval evaluates its selected prompt on four final cases and returns a gate decision. +The field-matching judge is deterministic; worker and reflection calls use a paid model endpoint. -## When To Use It - -Use this path when one surface must get better and you want the search, the held-out re-score, and the release decision in one call. -Use [`compare-optimization-methods`](../compare-optimization-methods/) instead when two or more optimizers must be benchmarked against each other at equal budget. -Use [`selfimprove-quickstart`](../selfimprove-quickstart/) when your own code generates the candidates. +Use [the local quickstart](../selfimprove-quickstart/) to try the flow without credentials or paid calls. +Use [the method comparison](../compare-optimization-methods/) to compare multiple search procedures. ## Install -Install the Node dependencies from the repository root: - -```sh -pnpm install -``` - -Install the Python bridge and the published GEPA package: - -```sh -python -m pip install agent-eval-rpc -python -m pip install \ - "gepa==0.1.4" \ - "litellm>=1.83.0,<1.92" \ - "tqdm>=4.66.1" \ - "cloudpickle>=3.0.0" \ - "datasets>=2.14.6" \ - "wandb" -``` - -Do not install `gepa[full]`; its MLflow server dependency is unpatched. - -From this repository, the locked equivalent is: +Run these commands from the repository root with Node, pnpm, Python, and uv installed: ```sh +pnpm install --frozen-lockfile cd clients/python uv sync --frozen --group gepa-release cd ../.. export OPTIMIZER_PYTHON="$PWD/clients/python/.venv/bin/python" ``` -## Run +This installs the bridge from the checkout and the locked GEPA dependencies. +The standard engine used here works with the published GEPA package. +See the [Python guide](../../clients/python/README.md) for other environments and supported versions. + +## Configure and run + +Export the endpoint, key, worker model, and reflection token rates before running the script. +Use a Chat Completions endpoint that accepts this example's request fields and returns model identity and complete token usage. +Its base URL should end at the API prefix, such as `/v1`. +Select a model your endpoint serves; the script's default is listed below. + +| Variable | Default | Purpose | +|---|---|---| +| `LLM_BASE_URL` | required | Endpoint used by worker and reflection calls. | +| `LLM_API_KEY` | required | Key for that endpoint. | +| `LLM_MODEL` | `deepseek-v4-flash` | Worker model. | +| `GEPA_MODEL` | `LLM_MODEL` | Reflection model. | +| `PRICE_IN_PER_M`, `PRICE_OUT_PER_M` | package pricing | Worker input/output USD rates per million tokens; set both for an unlisted model or endpoint-specific prices. | +| `PRICE_CACHED_IN_PER_M`, `PRICE_CACHE_WRITE_IN_PER_M` | worker input rate | Optional worker cache-read and cache-write rates; require both worker input/output rates. | +| `GEPA_PRICE_IN_PER_M`, `GEPA_PRICE_OUT_PER_M` | required | Current reflection input/output USD rates per million tokens. | +| `LLM_MAX_TOKENS` | `400` | Output limit per worker call. | +| `CALL_TIMEOUT_MS` | `30000` | Worker and reflection owner timeout per call. | +| `GEPA_MAX_EVALUATIONS` | `12` | Maximum candidate-case evaluations during search. | +| `GEPA_MAX_PROPOSER_COST_USD` | `2` | Reflection spend limit for the GEPA stage. | +| `MAX_TOTAL_COST_USD` | `10` | Shared limit for search and final evaluation calls admitted through the ledger. | +| `OPTIMIZER_PYTHON` | `python` | Python executable containing the bridge and GEPA. | ```sh -LLM_BASE_URL=https://router.tangle.tools/v1 \ -LLM_API_KEY="$TANGLE_API_KEY" \ -GEPA_PRICE_IN_PER_M=0.4 \ -GEPA_PRICE_OUT_PER_M=1.6 \ -pnpm tsx examples/self-improve-optimizer/index.ts +pnpm exec tsx examples/self-improve-optimizer/index.ts ``` -Any OpenAI-compatible endpoint works; set `LLM_MODEL` to a model that endpoint serves. -Replace the two rates with the exact rates charged by your endpoint. -The script validates every required variable before it makes a paid call. +The script validates required environment values before search. +Provider compatibility, installed Python capabilities, and returned usage are checked on their execution paths. +For reasoning models, increase `LLM_MAX_TOKENS` enough to include reasoning and final JSON output. -## Why It Is Built This Way +The [shared budget helper](../_shared/optimizer-model-budget.ts) also accepts `GEPA_MAX_MODEL_REQUESTS` and `GEPA_MAX_MODEL_COST_USD`. +It exposes byte limits, token limits, timeouts, and separate cache rates. +The reflection model budget defaults to the stage spend limit. -- The ten cases live inline in `index.ts`; `selfImprove()` derives every partition from that one list, so GEPA can never see the held-out cases. -- The example execution owner (`_shared/openai-compatible-owner.ts`) supplies the metered model call GEPA reflection runs through; the provider key never reaches Agent Eval or the Python child, and each reflection call is metered against the declared budget. -- The judge is deterministic field matching, so a score change traces to prompt content, not judge noise. -- `budget.generations` stays unset because the external method owns its rounds. -- `assertRealBackend` fails the run when any cell lacks a real backend receipt. +## Understand the data and cost -## Cost +Ten inline cases feed every partition: four final cases and six cases divided between train and selection. +`selfImprove()` does not pass final cases to the optimizer's callbacks. +Keep them out of callback closures, shared files, and external optimizer memory when adapting this code. +These APIs do not provide process or filesystem isolation. -A default run makes roughly 20 to 40 worker calls and up to 12 GEPA candidate evaluations plus reflection calls. -With a flash-tier model at the example rates, expect $0.10 to $0.50. -Hard limits: `MAX_TOTAL_COST_USD` (default 10) caps the whole run and `GEPA_MAX_PROPOSER_COST_USD` (default 2) caps reflection spend. +The worker dispatch and reflection owner record model usage and cost receipts. +Provider-reported billed cost takes precedence when present; otherwise configured or package token prices produce estimates. +Worker pricing also reserves the maximum charge before a capped call starts. +An unlisted worker model therefore needs `PRICE_IN_PER_M` and `PRICE_OUT_PER_M`, even when its provider later reports billed cost. +Actual call counts and spend depend on optimizer behavior, retries, and token usage. +The limits above are ceilings, not expected costs. +The selected prompt is evaluated only after search finishes. -## Read The Result +The script checks captured worker receipts with `assertRealBackend(records, { allowMixed: false })`. +That check establishes the identity of recorded worker execution; it does not validate a model's task quality. +The provider key stays in the example's execution owner rather than being passed to the metered Python optimizer. -The script prints the gate decision, the held-out baseline and winner composites, the lift, the total spend, and the baseline-to-winner diff. -Method results use `mode: 'method'`; search evidence is in `raw.method` and optional `searchHistory`. -They do not contain native `raw.generations` or `generationsExplored`. -With deferred holdout, baseline and winner scores are `null` and no lift exists. -Four held-out cases are wiring-scale, not statistical evidence: a `need_more_work` decision at this size is the gate refusing to claim significance, not a failure. -Grow the case list and set `budget.reps` above 1 before treating the decision as a production threshold. +## Read the result -## Controls +The script prints the gate decision, final baseline and selected scores, lift, cost, and prompt diff. +Method results use `mode: 'method'`. +Search evidence lives in `raw.method` and optional `searchHistory`; method results have no native generation count. +A selected surface remains in `winner.surface` even when it is unchanged, worse on final cases, or held by the gate. -| Variable | Default | Meaning | -|---|---:|---| -| `LLM_API_KEY` | required | Key for the worker and optimizer endpoint. | -| `LLM_BASE_URL` | required | OpenAI-compatible endpoint. | -| `LLM_MODEL` | `deepseek-v4-flash` | Worker model. | -| `LLM_MAX_TOKENS` | `400` | Output cap per worker call; raise it for reasoning models. | -| `GEPA_MODEL` | `LLM_MODEL` | Reflection model. | -| `GEPA_PRICE_IN_PER_M` | required | Exact input rate per million tokens. | -| `GEPA_PRICE_OUT_PER_M` | required | Exact output rate per million tokens. | -| `GEPA_MAX_EVALUATIONS` | `12` | Maximum GEPA candidate-case calls. Keep it at or above the train partition size. | -| `GEPA_MAX_PROPOSER_COST_USD` | `2` | Maximum reflection spend. | -| `MAX_TOTAL_COST_USD` | `10` | Hard cap across the whole run. | -| `OPTIMIZER_PYTHON` | `python` | Python executable containing the bridge and GEPA. | -| `CALL_TIMEOUT_MS` | `30000` | Per-call timeout. | +Four final cases demonstrate integration and cannot support a broad claim that GEPA improves future tasks. +A `hold` decision can reflect insufficient evidence; inspect gate contributions before diagnosing a failure. +Add representative independent cases and calibrate the judge before using this example for release decisions. +Use repetitions to estimate variation on those cases, without counting them as new tasks. -The complete implementation is [`index.ts`](./index.ts). +See [campaign proposers](../../docs/campaign-proposers.md) for method contracts, costs, and optional final-evidence controls. +The complete implementation is [index.ts](./index.ts). +For an installed package, import `selfImprove` from `/contract` and `gepaOptimizationMethod` from `/campaign` under `@tangle-network/agent-eval`. diff --git a/examples/self-improve-optimizer/index.ts b/examples/self-improve-optimizer/index.ts index 6ea528723..760750556 100644 --- a/examples/self-improve-optimizer/index.ts +++ b/examples/self-improve-optimizer/index.ts @@ -24,7 +24,7 @@ import { gepaOptimizationMethod } from '../../src/campaign' import { selfImprove } from '../../src/contract' import { assertRealBackend, summarizeBackendIntegrity } from '../../src/integrity/backend-integrity' import type { RunRecord } from '../../src/run-record' -import { positiveIntegerEnv, positiveNumberEnv } from '../_shared/env' +import { positiveIntegerEnv, positiveNumberEnv, tokenPricingFromEnv } from '../_shared/env' import { type Artifact, type ExtractScenario, @@ -59,6 +59,7 @@ const LLM_MAX_TOKENS = positiveIntegerEnv('LLM_MAX_TOKENS', 400) const GEPA_MAX_EVALUATIONS = positiveIntegerEnv('GEPA_MAX_EVALUATIONS', 12) const GEPA_MAX_PROPOSER_COST_USD = positiveNumberEnv('GEPA_MAX_PROPOSER_COST_USD', 2) const MAX_TOTAL_COST_USD = positiveNumberEnv('MAX_TOTAL_COST_USD', 10) +const customTokenPricing = tokenPricingFromEnv() // Throws unless GEPA_PRICE_IN_PER_M and GEPA_PRICE_OUT_PER_M carry the exact // endpoint rates, so reflection spend is never a guessed zero. const gepaModelBudget = optimizerModelBudgetFromEnv('GEPA', GEPA_MAX_PROPOSER_COST_USD) @@ -139,11 +140,12 @@ const BASELINE_SURFACE = 'Extract the transaction info from the message as JSON. // ── Agent, judge, and the GEPA method ──────────────────────────────────── const records: RunRecord[] = [] -// The worker transport is caller code: Agent Eval holds no provider key. +// The execution owner binds the caller's endpoint and credential. const chat = openAiCompatibleChatClient({ baseUrl: BASE_URL, apiKey: API_KEY, model: MODEL, + pricing: customTokenPricing, maximumAttempts: 2, timeoutMs: CALL_TIMEOUT_MS, }) @@ -154,12 +156,13 @@ const worker = makeExtractionWorker({ timeoutMs: CALL_TIMEOUT_MS, maxTokens: LLM_MAX_TOKENS, experimentId: 'self-improve-optimizer', + customTokenPricing, }) // The execution owner is caller code: it wraps one OpenAI-compatible endpoint // as the metered model call every official optimizer requires. Agent Eval's // loopback proxy meters each reflection call against `budget`, and the -// provider key never reaches Agent Eval or the Python child. +// provider key never reaches the Python child. const optimizerCall = openAiCompatibleExecutionOwner({ baseUrl: BASE_URL, apiKey: API_KEY, diff --git a/examples/selfimprove-quickstart/README.md b/examples/selfimprove-quickstart/README.md index d0889563e..6583e924f 100644 --- a/examples/selfimprove-quickstart/README.md +++ b/examples/selfimprove-quickstart/README.md @@ -1,48 +1,55 @@ -# Improve A Prompt Automatically +# Improve a prompt with a local proposer -This example defines an agent, scenarios, a judge, a starting prompt, and a candidate generator once. -It then calls `defineAgentEval().improve()` to search for a better prompt and evaluate the winner on scenarios that were not used to generate candidates. +This example defines an agent, twelve synthetic cases, a judge, a starting prompt, and a candidate generator. +It calls `defineAgentEval().improve()` to search and evaluate the selected prompt on six held-out cases. +All three functions are deterministic and local. +No API key or model call is required. -Run it from the repository root: +## Run + +From the repository root: ```sh -pnpm tsx examples/selfimprove-quickstart/index.ts +pnpm install --frozen-lockfile +pnpm exec tsx examples/selfimprove-quickstart/index.ts ``` -No API key is required. -The agent, judge, and candidate generator are deterministic local functions. - -## What It Demonstrates - -1. Define twelve representative tasks. -2. Run the starting prompt on a training split. -3. Generate two candidate prompts. -4. Score every candidate with the same judge. -5. Evaluate the selected candidate on six held-back tasks. -6. Return the selected prompt, measured score change, cost, and release decision. - -The stable part of the output is: +The output includes: ```text -Release decision: ship +Release decision: hold Raw lift: +0.351 Generations explored: 1 Total cost: $0.000 ``` -This is a wiring example, not statistical evidence. -Only six tasks are held back, so do not use its release decision as a production threshold. +The candidate improves the fixture's score, but six continuous-score pairs do not meet the gate's inference floor. +The selected prompt remains available for inspection even when the gate holds it. +The result demonstrates search, scoring, result capture, and a refused promotion under insufficient evidence. +It does not establish performance on real tasks. -The release decision comes from the held-out promotion gate, not from the search score. -The gate holds a candidate that is byte-identical to the baseline, requires the bootstrap confidence interval on the paired holdout delta to clear the threshold, and refuses a candidate whose search-to-holdout gap says it won the optimizer but lost the exam. -[`held-out-gate`](../held-out-gate/) walks each check with a promoting and a refused candidate. +## Read the result -## Adapt It +The release decision comes from the held-out gate. +Its paired decision depends on the outcome type, sample size, practical effect, and configured checks. +An unchanged candidate also remains on hold. +Optional red-team, canary, and reward-hacking checks need their own configured inputs. +The [held-out gate example](../held-out-gate/) demonstrates these checks. + +The result uses `mode: 'proposer'` and records the native generation count. +See [campaign proposers](../../docs/campaign-proposers.md#read-an-improvement-result) for the method/proposer result distinction. + +## Adapt it - Replace `agent` with the product call you want to improve. -- Replace `judge.score` with a deterministic check or a calibrated model-based judge. -- Replace the synthetic candidate generator with your own `SurfaceProposer` or delegate candidate creation to agent-runtime. -- Pass an official GEPA or SkillOpt method to `selfImprove()` instead of wrapping either complete optimizer as a proposer; [`self-improve-optimizer`](../self-improve-optimizer/) is the runnable version. Use `compareOptimizationMethods()` only to benchmark methods against each other. -- Increase the task corpus and repetitions until the score can distinguish known-good from known-bad behavior. +- Replace `judge.score` with objective checks or a model judge calibrated on independent examples. +- Replace the synthetic generator with your own `SurfaceProposer`. +- Add representative independent tasks; use repetitions to measure variation within tasks. +- Keep final decision cases outside candidate generation and selection. +- Use [`selfImprove({ method })`](../self-improve-optimizer/) for a complete optimizer such as GEPA. + +Check that known good and known bad outputs receive the intended scores before starting a larger search. +Keep `expectUsage: 'off'` only for calls that have no paid usage. -The complete implementation is [`index.ts`](./index.ts). +The complete implementation is [index.ts](./index.ts). +For an installed package, import `defineAgentEval` from `@tangle-network/agent-eval/contract`. diff --git a/examples/selfimprove-quickstart/index.ts b/examples/selfimprove-quickstart/index.ts index e8a6952eb..7c25311f0 100644 --- a/examples/selfimprove-quickstart/index.ts +++ b/examples/selfimprove-quickstart/index.ts @@ -14,8 +14,8 @@ interface CopyScenario extends Scenario { brief: string } -// Twelve cases so the 50% holdout split keeps six: the gate cannot claim -// 95% significance on fewer paired holdout observations at this effect size. +// Six final cases keep the demo small. Continuous mean inference needs more +// independent pairs, so the selected improvement retains an inconclusive gate. const scenarios: CopyScenario[] = [ { id: 'launch', kind: 'copy', brief: 'announce a new pricing tier' }, { id: 'feature', kind: 'copy', brief: 'highlight a new collaboration feature' }, diff --git a/scripts/verify-package-exports.mjs b/scripts/verify-package-exports.mjs index 910163015..79cdeb8de 100644 --- a/scripts/verify-package-exports.mjs +++ b/scripts/verify-package-exports.mjs @@ -149,7 +149,18 @@ try { const readme = readFileSync(join(repoRoot, 'README.md'), 'utf8') const quickstart = readme.match(/## Quickstart[\s\S]*?```ts\n([\s\S]*?)\n```/)?.[1] if (!quickstart) throw new Error('README quickstart TypeScript block was not found') - writeFileSync(join(appDir, 'quickstart.ts'), `${quickstart}\n`) + writeFileSync(join(appDir, 'quickstart.ts'), `${quickstart} + for (const [result, expected] of [[baseline, 0], [candidate, 1]] as const) { + const distribution = result.aggregates.byJudge['ticket-id']?.distribution + if ( + !distribution || distribution.n !== 3 || distribution.sum !== expected * 3 || + distribution.min !== expected || distribution.p50 !== expected || + distribution.p90 !== expected || distribution.max !== expected + ) { + throw new Error('README quickstart lost its complete score distribution') + } + } + `) writeFileSync(join(appDir, 'package.json'), JSON.stringify({ type: 'module' })) writeFileSync( join(appDir, 'index.ts'), @@ -809,20 +820,9 @@ try { run(process.execPath, [join(appDir, 'dist', 'integrity-imports.js')], appDir) const quickstartOutput = run(process.execPath, [join(appDir, 'dist', 'quickstart.js')], appDir) const plainQuickstartOutput = quickstartOutput.replace(/\x1b\[[0-9;]*m/g, '') - // Whitespace-tolerant: Node's inspector wraps the aggregate across lines once - // it carries its distribution, so the shape of the break is not the contract. - if (!/'ticket-id':\s*\{\s*mean: 0,/.test(plainQuickstartOutput)) { - throw new Error(`README quickstart baseline output changed:\n${quickstartOutput}`) - } - if (!/'ticket-id':\s*\{\s*mean: 1,/.test(plainQuickstartOutput)) { - throw new Error(`README quickstart candidate output changed:\n${quickstartOutput}`) - } - // The published aggregate carries the spread, not only the mean. - if (!/distribution: \{ n: 3, min: 0, p50: 0, p90: 0, max: 0, sum: 0 \}/.test(plainQuickstartOutput)) { - throw new Error(`README quickstart baseline distribution is missing:\n${quickstartOutput}`) - } - if (!/distribution: \{ n: 3, min: 1, p50: 1, p90: 1, max: 1, sum: 3 \}/.test(plainQuickstartOutput)) { - throw new Error(`README quickstart candidate distribution is missing:\n${quickstartOutput}`) + const expectedQuickstartOutput = readme.match(/## Quickstart[\s\S]*?```text\n([\s\S]*?)\n```/)?.[1] + if (!expectedQuickstartOutput || plainQuickstartOutput.trim() !== expectedQuickstartOutput.trim()) { + throw new Error(`README quickstart output differs from its documented output:\n${quickstartOutput}`) } run( process.execPath, diff --git a/src/cli-config.ts b/src/cli-config.ts index a2f9abe62..899a5dcbe 100644 --- a/src/cli-config.ts +++ b/src/cli-config.ts @@ -1,12 +1,9 @@ /** * Provider configuration for the `agent-eval` binary. * - * This is the ONE place in the package that turns an environment credential - * into a model transport, and it exists only inside the binary. The `agent-eval` - * server is a deployed process whose caller is a JSON-RPC or HTTP client in - * another language, so it cannot be handed a `ChatClient`; it reads its own - * credential the way every server does. The library never does: a TypeScript - * consumer binds its own transport and agent-eval holds no provider key. + * The binary reads endpoint and credential values from its environment. + * TypeScript library callers bind a transport explicitly. The maintained + * HTTP transport accepts their credential and sends provider requests. */ import { type ChatClient, createChatClient } from './analyst/chat-client' diff --git a/src/contract/analyze-runs.ts b/src/contract/analyze-runs.ts index 653461437..4326b3622 100644 --- a/src/contract/analyze-runs.ts +++ b/src/contract/analyze-runs.ts @@ -987,29 +987,36 @@ async function computeFailureClusters( const failed = runs.filter((run) => isTaskFailure(run, split)) if (failed.length === 0) return { clusters: [], totalFailures: 0 } - const clusters = new Map() + const clusters = new Map() + const recordCluster = (key: string, runId: string) => { + const cluster = clusters.get(key) ?? { exemplars: [], runs: 0 } + cluster.runs += 1 + if (cluster.exemplars.length < 5 && !cluster.exemplars.includes(runId)) { + cluster.exemplars.push(runId) + } + clusters.set(key, cluster) + } for (const run of failed) { try { // AnalystRunInputs routes by field name: run-record analysts read // `runRecord`. Any other shape makes every analyst skip with // "missing input" and the clusters come back silently empty. const result = await analyst.run(run.runId, { runRecord: run }) - for (const finding of result.findings as AnalystFinding[]) { - const key = finding.area || finding.analyst_id || 'unclassified' - const c = clusters.get(key) ?? { exemplars: [], share: 0 } - if (c.exemplars.length < 5) c.exemplars.push(run.runId) - clusters.set(key, c) - } + // One run can produce several findings in one cluster; count it once. + const keys = new Set( + result.findings.map( + (finding: AnalystFinding) => finding.area || finding.analyst_id || 'unclassified', + ), + ) + for (const key of keys) recordCluster(key, run.runId) } catch { - const c = clusters.get('analyst-error') ?? { exemplars: [], share: 0 } - if (c.exemplars.length < 5) c.exemplars.push(run.runId) - clusters.set('analyst-error', c) + recordCluster('analyst-error', run.runId) } } const clusterList = [...clusters.entries()].map(([id, c]) => ({ id, name: id, - share: c.exemplars.length / failed.length, + share: c.runs / failed.length, exemplars: c.exemplars, })) clusterList.sort((a, b) => b.share - a.share) diff --git a/tests/contract-analyze-runs.test.ts b/tests/contract-analyze-runs.test.ts index ebf35c7c1..1f18b4f74 100644 --- a/tests/contract-analyze-runs.test.ts +++ b/tests/contract-analyze-runs.test.ts @@ -1433,27 +1433,29 @@ describe('analyzeRuns — failure clustering via the analyst registry', () => { ) }) - function failureRegistry(): AnalystRegistry { + function failureRegistry( + areas: (run: RunRecord) => readonly string[] = () => ['timeout'], + ): AnalystRegistry { const registry = new AnalystRegistry() registry.register({ id: 'failure-classifier', - description: 'tags every failed run with a timeout finding', + description: 'classifies failed runs for report aggregation', inputKind: 'run-record', cost: { kind: 'deterministic' }, version: '1', analyze: async (input) => { const run = input as RunRecord - return [ + return areas(run).map((area, index) => makeFinding({ analyst_id: 'failure-classifier', severity: 'major', - area: 'timeout', - claim: `run ${run.runId} timed out`, + area, + claim: `run ${run.runId}: ${area} finding ${index}`, evidence_refs: [], confidence: 1, subject: run.runId, }), - ] + ) }, }) return registry @@ -1475,6 +1477,72 @@ describe('analyzeRuns — failure clustering via the analyst registry', () => { expect(cluster.share).toBeCloseTo(1, 5) }) + it('counts every affected failure while displaying at most five exemplars', async () => { + const runs = Array.from({ length: 7 }, (_, index) => + makeRun({ id: `f-${index}`, candidate: 'c', composite: 0.2 }), + ) + + const report = await analyzeRuns({ runs, analyst: failureRegistry() }) + + expect(report.failureClusters).toEqual({ + totalFailures: 7, + clusters: [ + { + id: 'timeout', + name: 'timeout', + share: 1, + exemplars: ['f-0', 'f-1', 'f-2', 'f-3', 'f-4'], + }, + ], + }) + expect(InsightReportSchema.parse(report).failureClusters).toEqual(report.failureClusters) + }) + + it('counts multiple same-cluster findings once per run', async () => { + const report = await analyzeRuns({ + runs: [makeRun({ id: 'f-1', candidate: 'c', composite: 0.2 })], + analyst: failureRegistry(() => ['timeout', 'timeout', 'timeout']), + }) + + expect(report.failureClusters?.clusters).toEqual([ + { id: 'timeout', name: 'timeout', share: 1, exemplars: ['f-1'] }, + ]) + }) + + it('ranks overlapping clusters using failed runs as the denominator', async () => { + const analyzed: string[] = [] + const runs = [ + ...Array.from({ length: 7 }, (_, index) => + makeRun({ id: `f-${index}`, candidate: 'c', composite: 0.2 }), + ), + makeRun({ id: 'f-7', candidate: 'c', composite: 0.9, failureMode: 'parse' }), + makeRun({ id: 'ok-1', candidate: 'c', composite: 0.9 }), + makeRun({ id: 'ok-2', candidate: 'c', composite: 0.8 }), + ] + const report = await analyzeRuns({ + runs, + analyst: failureRegistry((run) => { + analyzed.push(run.runId) + const index = Number(run.runId.slice(2)) + return [...(index < 6 ? ['timeout'] : []), ...(index >= 5 ? ['parse'] : [])] + }), + }) + + expect(analyzed).toEqual(Array.from({ length: 8 }, (_, index) => `f-${index}`)) + expect(report.failureClusters).toEqual({ + totalFailures: 8, + clusters: [ + { + id: 'timeout', + name: 'timeout', + share: 6 / 8, + exemplars: ['f-0', 'f-1', 'f-2', 'f-3', 'f-4'], + }, + { id: 'parse', name: 'parse', share: 3 / 8, exemplars: ['f-5', 'f-6', 'f-7'] }, + ], + }) + }) + it('does not turn terminal execution failure or recovered child errors into task failures', async () => { const recovered = makeRun({ id: 'recovered',