Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,13 @@ runs need headroom. Set `maxBudgetUsd` to `3` or more for artifact scenarios.
A run that hits the cap fails with an explicit
`claude hit the $N max budget` error rather than a silent bad sample.

**Turn cap:** `--max-turns <n>` (or `"maxTurns"` in a scenario file, with a
per-case override) caps agentic turns per run on the claude-p runner. Unlike
a budget abort, a turn-capped run is a *measured outcome*: it is recorded and
scored as a failed run without grading its partial output — an answer the
agent stumbled into at the cap must not count as a cheap pass. Completion
runners (openai) are single-turn and ignore the cap.

## Runners

`--runner claude-p` (default) shells out to headless Claude Code and supports
Expand Down Expand Up @@ -285,6 +292,21 @@ Reading results:
model gets a warning on the summary — a pass on the wrong model validates
prompt logic, not production behavior.

### Raw per-run results

Summaries, receipts, and the ndjson report digest each run down to pass/cost —
the right altitude for verdicts, and the wrong one for token economics.
`--raw-out <dir>` (on `compare` and `measure`) writes each run's **full runner
result JSON** as `<scenario>_<arm>_<n>.json` the moment the run finishes: for
claude-p that includes `usage` (input, cache-creation, cache-read, and output
tokens), per-model `modelUsage`, and the result `subtype`. Cost-USD is a
weighted sum of those categories (cache reads are priced far below fresh
input), so experiments whose *question* is token behavior — did a prompt
change cut exploration tokens, or just shift reads into cache? — need the
categories, not the blend. Files are written per run, so a compare that dies
midway still leaves the completed runs' records. Cache-served baseline arms
ran in an earlier invocation and write nothing.

### Run history

`--report ndjson --report-out ./runs.ndjson` appends one record per scenario
Expand Down
15 changes: 13 additions & 2 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,8 @@ claude -p "<fixture work item>" \
--model <model> \
--output-format json \
--tools <tools> \
--max-budget-usd <amount>
--max-budget-usd <amount> \
[--max-turns <n>]
```

The spawned process runs with the per-run sandbox as its actual working
Expand All @@ -68,6 +69,12 @@ used as a substitute for sandboxing.
The runner captures final output text, `total_cost_usd`, turn count, duration,
and model usage keys.

A `maxTurns` cap (scenario field, per-case override, or `--max-turns`) is
passed through as `--max-turns`. A run that stops at the cap (result subtype
`error_max_turns`, on any exit code) is returned as a normal result flagged
`exhaustedTurns` — the engine records it and scores it as a failed run
without invoking the grader, so partial output cannot pass by accident.

### 3.2 `openai`

Sends one `POST <baseUrl>/chat/completions` request to any OpenAI-compatible
Expand Down Expand Up @@ -183,7 +190,11 @@ each run) recording the content hash of every prompt file tested plus the
verdict — content addressing instead of hand-maintained prompt_version
strings, so a consuming repo's CI can require a passing receipt for each
shipped prompt's current hash. Reports are history; receipts are current
state.
state. `--raw-out <dir>` (compare and measure) additionally writes each run's
full runner result JSON — token usage categories, modelUsage, subtype — one
file per completed run, for token-level analysis that the digested summaries
cannot support; files land as runs finish, so a partial invocation still
leaves its completed records.

`compare` scenarios assert nothing: they exist to report both arms' pass
rates and the delta, for comparisons (typically model-vs-model) where neither
Expand Down
43 changes: 40 additions & 3 deletions src/cli.ts
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ const runSpecs: FlagSpecs = {
tools: { arity: "one" },
"timeout-ms": { arity: "one" },
"max-budget-usd": { arity: "one" },
"max-turns": { arity: "one" },
"keep-sandbox": { arity: "none" },
"clean-sandbox": { arity: "none" },
};
Expand Down Expand Up @@ -63,10 +64,12 @@ const compareSpecs: FlagSpecs = {
tools: { arity: "one" },
"timeout-ms": { arity: "one" },
"max-budget-usd": { arity: "one" },
"max-turns": { arity: "one" },
"keep-sandbox": { arity: "none" },
report: { arity: "one" },
"report-out": { arity: "one" },
receipts: { arity: "one" },
"raw-out": { arity: "one" },
cache: { arity: "none" },
"cache-dir": { arity: "one" },
};
Expand Down Expand Up @@ -196,6 +199,7 @@ async function cmdRun(argv: string[]): Promise<void> {
tools,
timeoutMs: args.number("timeout-ms", DEFAULT_TIMEOUT_MS),
maxBudgetUsd: args.number("max-budget-usd", DEFAULT_MAX_BUDGET_USD),
maxTurns: maxTurnsFromArgs(args.has("max-turns") ? args.number("max-turns", 0) : undefined),
};

console.error(
Expand Down Expand Up @@ -233,8 +237,10 @@ const measureSpecs: FlagSpecs = {
tools: { arity: "one" },
"timeout-ms": { arity: "one" },
"max-budget-usd": { arity: "one" },
"max-turns": { arity: "one" },
"keep-sandbox": { arity: "none" },
receipts: { arity: "one" },
"raw-out": { arity: "one" },
};

async function cmdMeasure(argv: string[]): Promise<void> {
Expand All @@ -256,6 +262,7 @@ async function cmdMeasure(argv: string[]): Promise<void> {
runs: args.has("runs") ? args.number("runs", 0) : undefined,
timeoutMs: args.has("timeout-ms") ? args.number("timeout-ms", DEFAULT_TIMEOUT_MS) : undefined,
maxBudgetUsd: args.has("max-budget-usd") ? args.number("max-budget-usd", DEFAULT_MAX_BUDGET_USD) : undefined,
maxTurns: maxTurnsFromArgs(args.has("max-turns") ? args.number("max-turns", 0) : undefined),
mode: args.one("mode") ? modeFromString(args.one("mode")) : undefined,
tools: args.one("tools"),
addDirs: args.many("add-dir").length ? args.many("add-dir") : undefined,
Expand All @@ -264,11 +271,13 @@ async function cmdMeasure(argv: string[]): Promise<void> {
keepSandbox: args.has("keep-sandbox") ? true : undefined,
};

const rawOutDir = args.one("raw-out");
const config = loadCompareConfig(scenario, overrides, { singleArm: true });
const summary = await runMeasure({
config,
runner: armRunner(config.arms.baseline, config),
onProgress: (message) => console.error(`[promptdiff] ${message}`),
rawOut: rawOutDir === undefined ? undefined : { dir: rawOutDir },
});

const receiptsDir = args.one("receipts");
Expand Down Expand Up @@ -350,6 +359,7 @@ async function cmdCompare(argv: string[]): Promise<number> {
runs: args.has("runs") ? args.number("runs", 0) : undefined,
timeoutMs: args.has("timeout-ms") ? args.number("timeout-ms", DEFAULT_TIMEOUT_MS) : undefined,
maxBudgetUsd: args.has("max-budget-usd") ? args.number("max-budget-usd", DEFAULT_MAX_BUDGET_USD) : undefined,
maxTurns: maxTurnsFromArgs(args.has("max-turns") ? args.number("max-turns", 0) : undefined),
mode: args.one("mode") ? modeFromString(args.one("mode")) : undefined,
tools: args.one("tools"),
addDirs: args.many("add-dir").length ? args.many("add-dir") : undefined,
Expand All @@ -358,6 +368,7 @@ async function cmdCompare(argv: string[]): Promise<number> {
keepSandbox: args.has("keep-sandbox") ? true : undefined,
};

const rawOutDir = args.one("raw-out");
const config = loadCompareConfig(scenario, overrides);
const summary = await runCompare({
config,
Expand All @@ -367,6 +378,7 @@ async function cmdCompare(argv: string[]): Promise<number> {
},
onProgress: (message) => console.error(`[promptdiff] ${message}`),
cache,
rawOut: rawOutDir === undefined ? undefined : { dir: rawOutDir },
});

// History is appended before the exit code is decided — failed comparisons
Expand Down Expand Up @@ -395,6 +407,14 @@ function promptFromArgs(prompt: string | undefined, promptFile: string | undefin
throw new CliError(runUsage());
}

function maxTurnsFromArgs(value: number | undefined): number | undefined {
if (value === undefined) return undefined;
if (!Number.isInteger(value) || value < 1) {
throw new CliError("--max-turns must be an integer of at least 1");
}
return value;
}

function modeFromString(value: string | undefined): RunMode {
if (value === "text" || value === "artifact") return value;
throw new CliError("--mode must be either text or artifact");
Expand Down Expand Up @@ -475,6 +495,7 @@ function runUsage(): string {
" --runner <claude-p|openai> --base-url <url> --image <file>...",
" --var <name=value>... --sandbox <dir> --seed <dir>",
" --tools <tools|default|''> --timeout-ms <ms> --max-budget-usd <usd>",
" --max-turns <n> (claude-p: cap agentic turns per run)",
"",
"templates:",
" --var draft=./fixture.md binds {{draft}} in the agent, skills, and prompt;",
Expand Down Expand Up @@ -522,9 +543,20 @@ function compareUsage(): string {
" --baseline-model <m> --proposed-model <m>",
" --baseline-runner <r> --proposed-runner <r>",
" --mode <text|artifact> --tools <tools|default|''>",
" --timeout-ms <ms> --max-budget-usd <usd>",
" --timeout-ms <ms> --max-budget-usd <usd> --max-turns <n>",
" --report ndjson --report-out <file>",
" --cache [--cache-dir <dir>]",
" --raw-out <dir> --cache [--cache-dir <dir>]",
"",
"turn cap:",
" --max-turns <n> (or scenario \"maxTurns\", per-case override allowed) caps",
" agentic turns per run (claude-p). A run that hits the cap is scored as a",
" FAILED run without grading — partial output must not pass by accident.",
"",
"raw results:",
" --raw-out <dir> writes each run's full runner result JSON (token usage,",
" modelUsage, subtype) as <scenario>_<arm>_<n>.json, one file per completed",
" run — the token-level record that summaries and reports digest away.",
" Cache-served baseline arms ran earlier and write nothing here.",
"",
"caching:",
" --cache reuses recorded baseline-arm results (default dir .promptdiff/cache)",
Expand Down Expand Up @@ -631,11 +663,16 @@ function measureUsage(): string {
" --runner <claude-p|openai> --base-url <url> --runs <n>",
" --mode <text|artifact> --tools <tools|default|''>",
" --sandbox <dir> --seed <dir> --keep-sandbox",
" --timeout-ms <ms> --max-budget-usd <usd> --receipts <dir>",
" --timeout-ms <ms> --max-budget-usd <usd> --max-turns <n>",
" --receipts <dir> --raw-out <dir>",
"",
"--receipts <dir> writes one <scenario>.receipt.json per scenario with",
"per-file prompt hashes and the measured rates (verdict \"measured\").",
"",
"--max-turns <n> caps agentic turns per run (claude-p); a capped run is",
"scored as a failure without grading. --raw-out <dir> writes each run's",
"full runner result JSON (token usage included) as <scenario>_measure_<n>.json.",
"",
"Exit code is 0 whenever the runs complete — a measurement has no pass/fail.",
].join("\n");
}
Expand Down
5 changes: 5 additions & 0 deletions src/engine/cache.ts
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,8 @@ export interface CacheKeyInput {
arm: ArmConfig;
/** Effective run count for the case (case runs ?? config runs). */
runs: number;
/** Effective turn cap for the case (case maxTurns ?? config maxTurns), if any. */
maxTurns?: number;
/** Effective tools string for the case. */
tools: string;
/** Effective run mode for the case. */
Expand Down Expand Up @@ -53,6 +55,9 @@ export function buildCacheKey(input: CacheKeyInput): string {
runner: input.arm.runner,
baseUrl: input.arm.baseUrl ?? "none",
runs: input.runs,
// stableStringify drops undefined entries, so scenarios without a cap keep
// their pre-maxTurns cache keys.
maxTurns: input.maxTurns,
tools: input.tools,
mode: input.mode,
delivery: input.delivery,
Expand Down
64 changes: 51 additions & 13 deletions src/engine/compare.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,6 @@
import { createHash } from "node:crypto";
import { mkdirSync, writeFileSync } from "node:fs";
import { join } from "node:path";
import { assembleSystemPrompt } from "../prompt";
import type { Runner, RunnerRunOptions, RunResult } from "../types";
import { buildCacheKey, loadCachedArm, storeCachedArm } from "./cache";
Expand Down Expand Up @@ -59,10 +61,17 @@ export interface CompareRunOptions {
onProgress?: (message: string) => void;
/** Opt-in baseline-arm result cache; hits skip the baseline runs entirely. */
cache?: { dir: string };
/**
* Opt-in per-run raw persistence: each completed run's full runner result
* (token usage included) is written to this directory as it finishes.
* Cache-served baseline arms ran in an earlier invocation, so they write
* nothing here.
*/
rawOut?: { dir: string };
}

export async function runCompare(options: CompareRunOptions): Promise<CompareSummary> {
const { config, runners, onProgress, cache } = options;
const { config, runners, onProgress, cache, rawOut } = options;
// Validate each arm against its own runner — a mixed comparison fails on the
// violating arm before either arm's paid runs.
validateRunnerSupport(config, runners.baseline);
Expand All @@ -80,8 +89,8 @@ export async function runCompare(options: CompareRunOptions): Promise<CompareSum

for (const { evalCase, baselineSystem, proposedSystem } of renderedCases) {
onProgress?.(`scenario ${evalCase.name} (${evalCase.kind})`);
const baseline = await runBaselineArm(config, evalCase, baselineSystem, runners.baseline, cache, onProgress);
const proposed = await runArm(config, evalCase, "proposed", proposedSystem, runners.proposed, onProgress);
const baseline = await runBaselineArm(config, evalCase, baselineSystem, runners.baseline, cache, onProgress, rawOut);
const proposed = await runArm(config, evalCase, "proposed", proposedSystem, runners.proposed, onProgress, "proposed", rawOut);
cases.push({
name: evalCase.name,
kind: evalCase.kind,
Expand Down Expand Up @@ -277,6 +286,8 @@ export interface MeasureRunOptions {
config: CompareConfig;
runner: Runner;
onProgress?: (message: string) => void;
/** Opt-in per-run raw persistence; see CompareRunOptions.rawOut. */
rawOut?: { dir: string };
}

/**
Expand All @@ -286,7 +297,7 @@ export interface MeasureRunOptions {
* nonsense verdicts from sampling noise.
*/
export async function runMeasure(options: MeasureRunOptions): Promise<MeasureSummary> {
const { config, runner, onProgress } = options;
const { config, runner, onProgress, rawOut } = options;
validateRunnerSupport(config, runner);
assertJudgeGradersCalibrated(config);
const inline = config.delivery !== "install";
Expand All @@ -304,7 +315,7 @@ export async function runMeasure(options: MeasureRunOptions): Promise<MeasureSum
const cases: MeasureCaseSummary[] = [];
for (const { evalCase, systemPrompt } of rendered) {
onProgress?.(`scenario ${evalCase.name}`);
const result = await runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress, "measure");
const result = await runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress, "measure", rawOut);
cases.push({ name: evalCase.name, kind: evalCase.kind, result, promptSha256: sha256(systemPrompt) });
}

Expand Down Expand Up @@ -385,15 +396,17 @@ async function runBaselineArm(
runner: Runner,
cache: { dir: string } | undefined,
onProgress?: (message: string) => void,
rawOut?: { dir: string },
): Promise<ArmSummary> {
if (cache === undefined) {
return runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress);
return runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress, "baseline", rawOut);
}
const key = buildCacheKey({
systemPrompt,
casePrompt: evalCase.prompt,
arm: config.arms.baseline,
runs: evalCase.runs ?? config.runs,
maxTurns: evalCase.maxTurns ?? config.maxTurns,
tools: effectiveTools(config, evalCase),
mode: effectiveMode(config, evalCase),
delivery: config.delivery,
Expand All @@ -405,9 +418,12 @@ async function runBaselineArm(
const hit = loadCachedArm(cache.dir, key);
if (hit !== undefined) {
onProgress?.(` baseline: cache hit (${key.slice(0, 12)})`);
if (rawOut !== undefined) {
onProgress?.(" baseline: cache hit — no raw results to write");
}
return { ...hit, cached: true };
}
const summary = await runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress);
const summary = await runArm(config, evalCase, "baseline", systemPrompt, runner, onProgress, "baseline", rawOut);
storeCachedArm(cache.dir, key, summary);
return summary;
}
Expand All @@ -420,6 +436,7 @@ async function runArm(
runner: Runner,
onProgress?: (message: string) => void,
label: string = arm,
rawOut?: { dir: string },
): Promise<ArmSummary> {
const runs = evalCase.runs ?? config.runs;
const summaries: ArmRunSummary[] = [];
Expand All @@ -445,12 +462,20 @@ async function runArm(
}
}
const run = await runner.run(buildRunnerOptions(config, evalCase, arm, systemPrompt, sandbox.dir));
const grade = await gradeRun(evalCase.grader, {
run,
sandboxDir: sandbox.dir,
timeoutMs: config.timeoutMs,
maxBudgetUsd: config.maxBudgetUsd,
});
if (rawOut !== undefined) {
writeRawResult(rawOut.dir, evalCase.name, label, index + 1, run.raw);
}
// A capped run did not finish; grading its partial output would let an
// accidental-looking answer count as a cheap pass (and bill a judge call
// for a result that cannot stand either way).
const grade = run.exhaustedTurns
? { pass: false, message: `hit the ${evalCase.maxTurns ?? config.maxTurns}-turn cap before finishing` }
: await gradeRun(evalCase.grader, {
run,
sandboxDir: sandbox.dir,
timeoutMs: config.timeoutMs,
maxBudgetUsd: config.maxBudgetUsd,
});
summaries.push(toArmRunSummary(index + 1, run, grade, sandbox.dir));
} finally {
sandbox.cleanup();
Expand Down Expand Up @@ -488,9 +513,22 @@ function buildRunnerOptions(
tools: effectiveTools(config, evalCase),
timeoutMs: config.timeoutMs,
maxBudgetUsd: config.maxBudgetUsd,
maxTurns: evalCase.maxTurns ?? config.maxTurns,
};
}

/**
* One file per completed run, written as the run finishes — the full runner
* result (usage, modelUsage, subtype for claude-p) survives even if a later
* run aborts the invocation. Summaries deliberately keep only the digest;
* this is the escape hatch for token-level analysis.
*/
function writeRawResult(dir: string, caseName: string, label: string, runNumber: number, raw: unknown): void {
mkdirSync(dir, { recursive: true });
const safeCase = caseName.toLowerCase().replace(/[^a-z0-9_-]+/g, "-").replace(/^-+|-+$/g, "") || "scenario";
writeFileSync(join(dir, `${safeCase}_${label}_${runNumber}.json`), JSON.stringify(raw, null, 2) + "\n", "utf8");
}

function effectiveTools(config: CompareConfig, evalCase: EvalCaseConfig): string {
// Install delivery always needs tools (the Skill tool does the triggering),
// even when the grader is text-only and inline delivery would disable them.
Expand Down
Loading
Loading