From 78a8739a2e097d237d2ba15d90ecdf7c82bf31dc Mon Sep 17 00:00:00 2001 From: Josh Owens Date: Fri, 18 Sep 2026 14:54:08 +0000 Subject: [PATCH 1/2] docs: six runnable examples with captured output Adds examples/ as a learning curve, each directory self-contained (agent, skills, scenario file) with its own README, cost label, and output from a real run: - 01-hello-compare: smallest baseline-vs-proposed compare - 02-fix-a-regression: target + regression pair on commit-message length - 03-measure-first: single-arm measure on an under-specified extraction - 04-judge-grader: calibrated LLM judge, rubric + fixtures + gate - 05-any-model: openai runner against local ollama - 06-token-economics: --max-turns and --raw-out for token-category analysis test/examples.test.ts loads every scenario.json through the real config loader, so schema drift breaks CI instead of the first user who copies an example. The root README gains a "See it work" section pointing at 01. Co-Authored-By: Claude Opus 5 --- README.md | 29 ++++ examples/01-hello-compare/README.md | 48 +++++++ examples/01-hello-compare/agent.md | 7 + examples/01-hello-compare/baseline.md | 6 + examples/01-hello-compare/proposed.md | 9 ++ examples/01-hello-compare/scenario.json | 19 +++ examples/02-fix-a-regression/README.md | 77 ++++++++++ examples/02-fix-a-regression/agent.md | 12 ++ examples/02-fix-a-regression/scenario.json | 38 +++++ .../skills/commit-message.baseline.md | 15 ++ .../skills/commit-message.proposed.md | 22 +++ examples/03-measure-first/README.md | 70 +++++++++ examples/03-measure-first/agent.md | 7 + examples/03-measure-first/scenario.json | 18 +++ examples/03-measure-first/skill.md | 5 + examples/04-judge-grader/README.md | 113 +++++++++++++++ examples/04-judge-grader/agent.md | 8 ++ .../fail/domain-change.md | 7 + .../fail/email-missing.md | 7 + .../fail/login-locked.md | 6 + .../fail/slow-site.md | 7 + .../pass/domain-change.md | 9 ++ .../pass/email-missing.md | 8 ++ .../pass/login-locked.md | 7 + .../pass/slow-site.md | 7 + .../rubrics/audience-appropriate.md | 26 ++++ .../audience-appropriate.md.calibration.json | 56 ++++++++ examples/04-judge-grader/scenario.json | 22 +++ .../skills/explain-plainly.baseline.md | 6 + .../skills/explain-plainly.proposed.md | 17 +++ examples/05-any-model/README.md | 63 +++++++++ examples/05-any-model/agent.md | 7 + examples/05-any-model/baseline.md | 6 + examples/05-any-model/proposed.md | 10 ++ examples/05-any-model/scenario.json | 22 +++ examples/06-token-economics/.gitignore | 2 + examples/06-token-economics/README.md | 133 ++++++++++++++++++ examples/06-token-economics/agent.md | 8 ++ examples/06-token-economics/scenario.json | 24 ++++ examples/06-token-economics/seed/auth.py | 15 ++ examples/06-token-economics/seed/cache.py | 24 ++++ examples/06-token-economics/seed/db.py | 18 +++ examples/06-token-economics/seed/queue.py | 21 +++ examples/06-token-economics/seed/router.py | 18 +++ examples/06-token-economics/skill.md | 13 ++ examples/README.md | 22 +++ test/examples.test.ts | 27 ++++ 47 files changed, 1121 insertions(+) create mode 100644 examples/01-hello-compare/README.md create mode 100644 examples/01-hello-compare/agent.md create mode 100644 examples/01-hello-compare/baseline.md create mode 100644 examples/01-hello-compare/proposed.md create mode 100644 examples/01-hello-compare/scenario.json create mode 100644 examples/02-fix-a-regression/README.md create mode 100644 examples/02-fix-a-regression/agent.md create mode 100644 examples/02-fix-a-regression/scenario.json create mode 100644 examples/02-fix-a-regression/skills/commit-message.baseline.md create mode 100644 examples/02-fix-a-regression/skills/commit-message.proposed.md create mode 100644 examples/03-measure-first/README.md create mode 100644 examples/03-measure-first/agent.md create mode 100644 examples/03-measure-first/scenario.json create mode 100644 examples/03-measure-first/skill.md create mode 100644 examples/04-judge-grader/README.md create mode 100644 examples/04-judge-grader/agent.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/domain-change.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/email-missing.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/login-locked.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/slow-site.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/domain-change.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/email-missing.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/login-locked.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/slow-site.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.md create mode 100644 examples/04-judge-grader/rubrics/audience-appropriate.md.calibration.json create mode 100644 examples/04-judge-grader/scenario.json create mode 100644 examples/04-judge-grader/skills/explain-plainly.baseline.md create mode 100644 examples/04-judge-grader/skills/explain-plainly.proposed.md create mode 100644 examples/05-any-model/README.md create mode 100644 examples/05-any-model/agent.md create mode 100644 examples/05-any-model/baseline.md create mode 100644 examples/05-any-model/proposed.md create mode 100644 examples/05-any-model/scenario.json create mode 100644 examples/06-token-economics/.gitignore create mode 100644 examples/06-token-economics/README.md create mode 100644 examples/06-token-economics/agent.md create mode 100644 examples/06-token-economics/scenario.json create mode 100644 examples/06-token-economics/seed/auth.py create mode 100644 examples/06-token-economics/seed/cache.py create mode 100644 examples/06-token-economics/seed/db.py create mode 100644 examples/06-token-economics/seed/queue.py create mode 100644 examples/06-token-economics/seed/router.py create mode 100644 examples/06-token-economics/skill.md create mode 100644 examples/README.md create mode 100644 test/examples.test.ts diff --git a/README.md b/README.md index 2d85007..d66947a 100644 --- a/README.md +++ b/README.md @@ -64,6 +64,35 @@ Run from the repo: ./promptdiff --help ``` +## See it work + +The smallest possible comparison — the baseline skill is missing one +instruction, the proposed skill adds it, a text grader catches the effect: + +```bash +./promptdiff compare --scenario ./examples/01-hello-compare/scenario.json +``` + +``` +promptdiff compare: hello-compare + +answers-with-summary-prefix (target) + baseline: 0/2 pass (0%) | $0.0021 + proposed: 2/2 pass (100%) | $0.0033 + delta: +100% pass | +$0.0013 + PASS: assertions satisfied + +total cost: $0.0054 +``` + +That's real output from a real run (about a cent, ~12 seconds). The +[`examples/`](./examples/) directory has six runnable, self-contained +examples ordered as a learning curve — from this hello case through the +fix-a-defect-without-regressing loop, measure-first characterization, +calibrated LLM judges, non-Claude models via ollama, and token-level +analysis with turn caps — each with its own README, honest cost label, +and captured output. + ## Safe Defaults Paid model calls are bounded by default: diff --git a/examples/01-hello-compare/README.md b/examples/01-hello-compare/README.md new file mode 100644 index 0000000..c24ed85 --- /dev/null +++ b/examples/01-hello-compare/README.md @@ -0,0 +1,48 @@ +# Hello Compare + +The smallest possible `promptdiff compare`: one agent, one scenario, one +grader. The baseline skill is missing a formatting instruction; the proposed +skill adds it; a text grader catches the difference. This is the shape every +other example builds on — start here before reading the rest. + +**Cost:** ~$0.01 · **Time:** ~12s · **Requires:** claude CLI + +## Run it + +```bash +./promptdiff compare --scenario ./examples/01-hello-compare/scenario.json +``` + +## Output + +``` +promptdiff compare: hello-compare + +answers-with-summary-prefix (target) + baseline: 0/2 pass (0%) | $0.0021 + proposed: 2/2 pass (100%) | $0.0033 + delta: +100% pass | +$0.0013 + PASS: assertions satisfied + NOTE: delta could be sampling noise (Fisher exact p=0.33) — consider more runs + baseline run 1 failed: output did not contain "SUMMARY:" + baseline run 2 failed: output did not contain "SUMMARY:" + +total cost: $0.0054 +``` + +## What to notice + +- `baseline.md` and `proposed.md` are the same skill except for one added + instruction ("start your reply with `SUMMARY:`"). That's the entire A/B + variable — the agent and the scenario prompt never change. +- The pass-rate delta (0% → 100%) is the whole point: the grader + (`{"type": "text", "contains": ["SUMMARY:"]}`) can't see the instruction + text, only its effect on the model's output. +- This scenario is `"kind": "target"`, so `compare` asserts baseline must + *not* fully pass (the gap is real) and proposed must improve on it. Both + hold here, so the process exits `0`; either failing would exit non-zero — + to see the failing case, edit `scenario.json` so `proposedSkills` points at + `./baseline.md` and rerun. +- At `runs: 2` the delta is real but small-sample — the `NOTE:` line about + sampling noise is `compare` being honest about that, not a bug. It doesn't + change the exit code. diff --git a/examples/01-hello-compare/agent.md b/examples/01-hello-compare/agent.md new file mode 100644 index 0000000..797872e --- /dev/null +++ b/examples/01-hello-compare/agent.md @@ -0,0 +1,7 @@ +--- +name: hello-assistant +description: Minimal example agent for promptdiff's hello-compare example. +--- + +You are a helpful assistant that answers user questions directly and +concisely, in one or two sentences. diff --git a/examples/01-hello-compare/baseline.md b/examples/01-hello-compare/baseline.md new file mode 100644 index 0000000..e9bf9aa --- /dev/null +++ b/examples/01-hello-compare/baseline.md @@ -0,0 +1,6 @@ +--- +name: reply-format +description: Formatting rules for assistant replies. +--- + +Answer the user's question directly. Do not add disclaimers or hedging. diff --git a/examples/01-hello-compare/proposed.md b/examples/01-hello-compare/proposed.md new file mode 100644 index 0000000..a12366a --- /dev/null +++ b/examples/01-hello-compare/proposed.md @@ -0,0 +1,9 @@ +--- +name: reply-format +description: Formatting rules for assistant replies. +--- + +Answer the user's question directly. Do not add disclaimers or hedging. + +Start your reply with a line that reads exactly `SUMMARY:` followed by a +one-sentence answer, on its own line, before any further detail. diff --git a/examples/01-hello-compare/scenario.json b/examples/01-hello-compare/scenario.json new file mode 100644 index 0000000..d77d968 --- /dev/null +++ b/examples/01-hello-compare/scenario.json @@ -0,0 +1,19 @@ +{ + "name": "hello-compare", + "agent": "./agent.md", + "baselineSkills": ["./baseline.md"], + "proposedSkills": ["./proposed.md"], + "model": "haiku", + "runs": 2, + "scenarios": [ + { + "name": "answers-with-summary-prefix", + "kind": "target", + "prompt": "What is the capital of France?", + "grader": { + "type": "text", + "contains": ["SUMMARY:"] + } + } + ] +} diff --git a/examples/02-fix-a-regression/README.md b/examples/02-fix-a-regression/README.md new file mode 100644 index 0000000..da6c2ec --- /dev/null +++ b/examples/02-fix-a-regression/README.md @@ -0,0 +1,77 @@ +# Fix a Regression + +This demonstrates the core loop from promptdiff's ["Why it exists"](../../README.md#why-it-exists): +for a recurring defect, a trustworthy eval has to show two things at once — +(1) the **baseline** instruction set still reproduces the failure, and (2) the +**proposed** instruction set fixes it **without regressing** a case that +already worked. + +The skill under test writes git commit messages. The baseline skill tells the +model to follow Conventional Commits and "be thorough" about what changed, +but never states a length limit — so on any change with more than one moving +part, the model pads the subject line describing all of them, blowing past +the 50-character convention. The proposed skill adds one explicit rule: keep +the subject line ≤50 chars, put extra detail in the body. + +Two scenarios exercise this: + +- `target-multi-part-change` (**target**) — a change with four distinct parts + (swap the session store, add expiry, update handlers, migrate tests). + Baseline is expected to fail here; proposed is expected to fix it. +- `regression-simple-typo-fix` (**regression**) — a trivial one-line change + both arms already handle fine. This proves the new length rule doesn't make + the skill worse on cases it wasn't broken on. + +The grader is a deterministic text/regex check (no LLM judge needed): the +first line must match Conventional Commits format, and — separately — the +first line must be ≤50 characters. + +**Cost:** ~$0.10 · **Time:** ~188s · **Requires:** claude CLI + +## Run it + +```bash +./promptdiff compare --scenario examples/02-fix-a-regression/scenario.json +``` + +## Actual output + +``` +promptdiff compare: commit-message subject length + +target-multi-part-change (target) + baseline: 0/3 pass (0%) | $0.0209 + proposed: 3/3 pass (100%) | $0.0537 + delta: +100% pass | +$0.0328 + PASS: assertions satisfied + NOTE: delta could be sampling noise (Fisher exact p=0.10) — consider more runs + baseline run 1 failed: output did not match /^.{1,50}(\n|$)/ + baseline run 2 failed: output did not match /^.{1,50}(\n|$)/ + baseline run 3 failed: output did not match /^.{1,50}(\n|$)/ + +regression-simple-typo-fix (regression) + baseline: 3/3 pass (100%) | $0.0135 + proposed: 3/3 pass (100%) | $0.0133 + delta: +0% pass | $-0.0002 + PASS: assertions satisfied + +total cost: $0.1014 +``` + +## What to notice + +- **Baseline fails 0/3, proposed passes 3/3** on the target case — the + defect is real and reproducible, not cherry-picked, and the fix clears it + every run at `runs: 3`. +- **The regression scenario stays 3/3 → 3/3.** Adding the length rule didn't + make the skill worse on a case it already handled — that's the "no + regression" half of the claim, checked automatically rather than asserted + by eye. +- The grader is two independent regex checks (`type(scope): …` format, and + overall line length ≤50 via `^.{1,50}(\n|$)`) plus a `notContains` guard + against code fences — no LLM judge, no flakiness from a second model's + opinion. +- The `NOTE: delta could be sampling noise (Fisher exact p=0.10)` line on the + target scenario is expected at `runs: 3` — even a clean 0/3 → 3/3 flip + can't rule out noise at this sample size; the assertions and exit code are + unaffected. Raise `--runs` if you need a tighter p-value. diff --git a/examples/02-fix-a-regression/agent.md b/examples/02-fix-a-regression/agent.md new file mode 100644 index 0000000..f6143b6 --- /dev/null +++ b/examples/02-fix-a-regression/agent.md @@ -0,0 +1,12 @@ +--- +name: commit-message-writer +--- + +# Commit Message Writer + +You write git commit messages. The user will describe a code change; you +respond with the commit message for it, following the rules in your skill +instructions. + +Output *only* the commit message text — no preamble ("Here's a commit +message:"), no trailing commentary, no code fences. diff --git a/examples/02-fix-a-regression/scenario.json b/examples/02-fix-a-regression/scenario.json new file mode 100644 index 0000000..65fc560 --- /dev/null +++ b/examples/02-fix-a-regression/scenario.json @@ -0,0 +1,38 @@ +{ + "name": "commit-message subject length", + "agent": "./agent.md", + "baselineSkills": ["./skills/commit-message.baseline.md"], + "proposedSkills": ["./skills/commit-message.proposed.md"], + "model": "haiku", + "mode": "text", + "runs": 3, + "maxBudgetUsd": 1, + "scenarios": [ + { + "name": "target-multi-part-change", + "kind": "target", + "prompt": "Write a commit message for this change: replaced the in-memory session store with Redis-backed sessions, added automatic session expiry after 30 minutes of inactivity, updated the login and logout handlers to use the new session store, and migrated the existing session tests to mock Redis instead of the in-memory map.", + "grader": { + "type": "text", + "regex": [ + "^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S", + "^.{1,50}(\\n|$)" + ], + "notContains": ["```"] + } + }, + { + "name": "regression-simple-typo-fix", + "kind": "regression", + "prompt": "Write a commit message for this change: fixed a typo in the README.", + "grader": { + "type": "text", + "regex": [ + "^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S", + "^.{1,50}(\\n|$)" + ], + "notContains": ["```"] + } + } + ] +} diff --git a/examples/02-fix-a-regression/skills/commit-message.baseline.md b/examples/02-fix-a-regression/skills/commit-message.baseline.md new file mode 100644 index 0000000..072576d --- /dev/null +++ b/examples/02-fix-a-regression/skills/commit-message.baseline.md @@ -0,0 +1,15 @@ +--- +name: commit-message +--- + +# Commit Message Skill + +Follow the Conventional Commits format for the subject line: +`type(scope): description` (scope optional). + +Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore, +revert. + +Write a subject line that clearly and specifically describes what changed. +Be thorough — a vague subject line is not useful to someone reading `git +log` later. If the change touches several things, say what they are. diff --git a/examples/02-fix-a-regression/skills/commit-message.proposed.md b/examples/02-fix-a-regression/skills/commit-message.proposed.md new file mode 100644 index 0000000..3795b72 --- /dev/null +++ b/examples/02-fix-a-regression/skills/commit-message.proposed.md @@ -0,0 +1,22 @@ +--- +name: commit-message +--- + +# Commit Message Skill + +Follow the Conventional Commits format for the subject line: +`type(scope): description` (scope optional). + +Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore, +revert. + +Write a subject line that clearly and specifically describes what changed. +Be thorough — a vague subject line is not useful to someone reading `git +log` later. If the change touches several things, say what they are. + +**The subject line must be 50 characters or fewer, counting the whole +line** (type, scope, colon, and description). Use imperative mood ("add", +not "added"). If the change needs more explanation than fits in 50 +characters, put the detail in the body after a blank line — never stretch +the subject line to fit it. When several things changed, name the most +important one in the subject and list the rest in the body. diff --git a/examples/03-measure-first/README.md b/examples/03-measure-first/README.md new file mode 100644 index 0000000..24c38d6 --- /dev/null +++ b/examples/03-measure-first/README.md @@ -0,0 +1,70 @@ +# Measure first + +Before you rewrite a prompt you suspect is flaky, measure its current pass +rate. Skip that step and any before/after comparison is a coin flip against +a coin flip: two runs of the *same* prompt can already look like a fix or a +regression from sampling noise alone. `measure` runs one instruction set N +times and reports a bare pass rate — a real number to rewrite against, +instead of a vibe from the last output you happened to read. + +This example measures `skill.md`, an intentionally under-specified +extraction prompt: "extract the key fields from this support ticket as +JSON," no schema. Given the same support ticket four times, `haiku` +extracts a working JSON object every time — it just doesn't agree with +itself on the field names (`severity` vs. `urgency` vs. `priority`) or the +value casing (`"Medium"` vs. `"medium"`). The `json` grader's path +assertions catch exactly that kind of drift, because they check one exact +path and value, not "did it extract *something* reasonable." + +**Cost:** ~$0.01 · **Time:** ~28s · **Requires:** claude CLI + +```bash +./promptdiff measure --scenario ./examples/03-measure-first/scenario.json +``` + +Real output: + +``` +[promptdiff] scenario extracts-medium-severity +[promptdiff] measure run 1/4 +[promptdiff] measure run 2/4 +[promptdiff] measure run 3/4 +[promptdiff] measure run 4/4 +promptdiff measure: ticket-field-extraction (haiku via claude-p) + +extracts-medium-severity + 2/4 pass (50%) | $0.0139 + run 1 failed: severity == "Medium": path segment "severity" not found (at {"subject":"Can't export reports since yesterday's update","reporter_name":"Dana Whitfield","reporter_email":"dana.whitf…) + run 4 failed: severity == "Medium": found "medium" + +total cost: $0.0139 +``` + +```bash +$ echo $? +0 +``` + +## What to notice + +- **50% is the measurement, not a bug.** The grader is exact on purpose + (`severity == "Medium"`) — an under-specified prompt gets an + under-specified extraction, and the pass rate is the honest size of that + problem. Run it again and you'll see a different split (25%, 50%, ...); + that's the real variance in the prompt, not flakiness in the harness. +- **The two failure reasons are two different bugs.** Run 1's `severity` + key doesn't exist at all — the model named it `reporter_name`/ + `reporter_email` and dropped severity from that response's schema + entirely. Run 4 has the key but the wrong case (`"medium"` vs + `"Medium"`). A prompt fix for one won't fix the other — you'd need to + pin both the field name and an enum of allowed values. +- **`measure` exits 0 here even though the scenario "failed" 50% of the + time.** `echo $?` above prints `0` — a measurement has no pass/fail, so a + 49-53% swing on rerun is not a CI signal by itself. Wire a *threshold* + check around the printed rate (or graduate to `compare` with an assertion) + once you have a bar to enforce. +- This baseline is now the number a proposed fix has to beat. Pin the + schema (name the fields, constrain `severity` to an enum) in a revised + `skill.md`, drop it into a `compare` scenario as the proposed arm against + this one as baseline, and see if the pass rate — not just one output — + actually moves. diff --git a/examples/03-measure-first/agent.md b/examples/03-measure-first/agent.md new file mode 100644 index 0000000..eaafc91 --- /dev/null +++ b/examples/03-measure-first/agent.md @@ -0,0 +1,7 @@ +--- +name: ticket-intake +--- + +You are a support-ticket intake assistant. Follow the skill instructions +exactly and reply with nothing but the requested output — no preamble, no +markdown fences, no commentary. diff --git a/examples/03-measure-first/scenario.json b/examples/03-measure-first/scenario.json new file mode 100644 index 0000000..4a6f5b7 --- /dev/null +++ b/examples/03-measure-first/scenario.json @@ -0,0 +1,18 @@ +{ + "name": "ticket-field-extraction", + "agent": "./agent.md", + "skills": ["./skill.md"], + "model": "haiku", + "runs": 4, + "mode": "text", + "scenarios": [ + { + "name": "extracts-medium-severity", + "prompt": "Subject: Can't export reports since yesterday's update\n\nFrom: Dana Whitfield (dana.whitfield@northfield-labs.com)\n\nSince the update yesterday, exporting reports to CSV just spins forever and never finishes. I've tried three times. Not blocking my day-to-day work since I can still view everything fine in the dashboard, but I do need the CSV before Friday's board meeting.", + "grader": { + "type": "json", + "assert": ["severity == \"Medium\""] + } + } + ] +} diff --git a/examples/03-measure-first/skill.md b/examples/03-measure-first/skill.md new file mode 100644 index 0000000..9c26f88 --- /dev/null +++ b/examples/03-measure-first/skill.md @@ -0,0 +1,5 @@ +--- +name: extract-ticket-fields +--- + +Extract the key fields from this support ticket as JSON. diff --git a/examples/04-judge-grader/README.md b/examples/04-judge-grader/README.md new file mode 100644 index 0000000..c2d6164 --- /dev/null +++ b/examples/04-judge-grader/README.md @@ -0,0 +1,113 @@ +# Judge grader (calibrated) + +Some qualities can't be text-matched. "Is this explanation appropriate for a +non-technical audience?" has no substring or regex that captures it — +jargon-free phrasing varies too much, and a regex tuned to catch "DNS" or +"latency" will both miss paraphrases and flag legitimate uses. An LLM judge +grading against a rubric can make that call, but an *unproven* judge is +worse than the regex it replaces: same wrongness, more confidence, higher +cost, and it can silently bless any output as "appropriate" forever. +`promptdiff calibrate` closes that gap — it proves the judge against labeled +pass/fail fixtures before `compare` or `measure` are allowed to use it to +grade anything real. + +This example: a support agent answers a customer's DNS question. The +baseline skill gives no audience guidance, so Haiku answers like it's +talking to another engineer ("DNS propagation," "TTL," "resolvers," "cached +records"). The proposed skill instructs plain language for a non-technical +reader. A judge grades each reply against `rubrics/audience-appropriate.md`. + +**Cost:** ~$0.10 (calibrate $0.0362 + compare $0.0685) · **Time:** ~215s · +**Requires:** claude CLI + +## 1. Calibrate the judge + +```bash +./promptdiff calibrate --rubric examples/04-judge-grader/rubrics/audience-appropriate.md --model haiku +``` + +Real output: + +``` +[promptdiff] judging fixture pass/domain-change.md +[promptdiff] judging fixture pass/email-missing.md +[promptdiff] judging fixture pass/login-locked.md +[promptdiff] judging fixture pass/slow-site.md +[promptdiff] judging fixture fail/domain-change.md +[promptdiff] judging fixture fail/email-missing.md +[promptdiff] judging fixture fail/login-locked.md +[promptdiff] judging fixture fail/slow-site.md +promptdiff calibrate: haiku via claude-p + +pass class: 4/4 correct (100%) +fail class: 4/4 correct (100%) + +calibration record written: /home/josh/Code/OpenSource/promptdiff/examples/04-judge-grader/rubrics/audience-appropriate.md.calibration.json +judge cost: $0.0362 +gate: compare/measure require BOTH classes >= minAccuracy (default 90%) +``` + +This writes `rubrics/audience-appropriate.md.calibration.json` (committed +alongside the rubric) — a fixture-keyed proof that this judge model, on this +rubric's content hash, correctly separates the two classes. `calibrate` +always exits 0; it measures. The gate below is what enforces. + +## 2. Compare with the judge as grader + +```bash +./promptdiff compare --scenario examples/04-judge-grader/scenario.json +``` + +Real output: + +``` +[promptdiff] scenario nameserver-change-support-reply (target) +[promptdiff] baseline run 1/3 +[promptdiff] baseline run 2/3 +[promptdiff] baseline run 3/3 +[promptdiff] proposed run 1/3 +[promptdiff] proposed run 2/3 +[promptdiff] proposed run 3/3 +promptdiff compare: plain-language support answers + +nameserver-change-support-reply (target) + baseline: 0/3 pass (0%) | $0.0177 + proposed: 3/3 pass (100%) | $0.0508 + delta: +100% pass | +$0.0330 + PASS: assertions satisfied + NOTE: delta could be sampling noise (Fisher exact p=0.10) — consider more runs + baseline run 1 failed: judge verdict: fail — Uses multiple unexplained technical terms without definition: 'nameservers', 'DNS' (appears 4 times—the acronym is never explained in plain words), 'propagate/propagation', and 'resolvers'. A non-technical reader would encounter at least three unclear concepts in the first sentence alone and would need to look up what these terms mean to follow the explanation. + baseline run 2 failed: judge verdict: fail — Uses multiple unexplained jargon terms: 'DNS' (acronym, never defined), 'propagate' (technical concept, unexplained), 'nameserver' (unexplained), and 'cached' (unexplained). A non-technical reader would encounter at least 4 terms requiring external lookup to understand the core message. + baseline run 3 failed: judge verdict: fail — Uses multiple unexplained jargon terms: 'DNS' (appears 4 times without definition), 'propagate,' 'cache/caches,' and 'nameserver.' A non-technical reader would not understand what these mean or why they matter. The answer assumes knowledge of how domain name systems and caching work. + +total cost: $0.0685 +``` + +## What to notice + +- **The gate reads the same file it grades with.** `scenario.json`'s judge + grader is `{"rubric": "./rubrics/audience-appropriate.md", "model": + "haiku", "minAccuracy": 0.9}` — same rubric path, same model as the + calibration run. Change either and the gate refuses (stale-record or + model-mismatch error) until you recalibrate. +- **Skipping calibration refuses the run, before any paid call.** Deleting + `audience-appropriate.md.calibration.json` and re-running `compare` + produces, immediately and with exit code 1: + ``` + judge rubric .../audience-appropriate.md has no calibration record + (.../audience-appropriate.md.calibration.json) — run: promptdiff calibrate + --rubric .../audience-appropriate.md --model haiku + ``` + No model calls happen — the check runs before the baseline/proposed arms + do. +- **Per-class bars matter.** A judge that rubber-stamps everything as "pass" + would score 100% on the `pass` fixtures and 0% on the `fail` fixtures; + overall accuracy would hide that. `calibrate` reports both classes + separately, and the gate requires both to clear `minAccuracy` (0.9 here). +- **The judge's own reasoning is legible.** Each failed baseline run prints + the judge's `reason` (e.g. "DNS... never explained in plain words"), so a + `compare` failure tells you *why* in the judge's own words, not just + pass/fail. +- **n=3 is a demo, not a claim.** 0/3 → 3/3 is exactly the shape the + `NOTE: delta could be sampling noise (Fisher exact p=0.10)` line exists + to flag — real usage would run more. diff --git a/examples/04-judge-grader/agent.md b/examples/04-judge-grader/agent.md new file mode 100644 index 0000000..51c4941 --- /dev/null +++ b/examples/04-judge-grader/agent.md @@ -0,0 +1,8 @@ +--- +name: support-explainer +description: Answers a customer's technical support question in prose. +--- + +You are a support agent answering a question from someone who wrote in +through the help widget. Answer their question directly, in a short +paragraph. Do not ask clarifying questions — just answer. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/domain-change.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/domain-change.md new file mode 100644 index 0000000..0df3069 --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/domain-change.md @@ -0,0 +1,7 @@ +This is expected behavior during DNS propagation. When you repoint your +domain's nameservers, the change has to propagate across recursive resolvers +worldwide, and each resolver caches records according to their TTL. Until +the TTL expires on every resolver in the chain, some clients will keep +resolving to the old A record. Propagation is typically complete within +24-48 hours depending on upstream caching behavior. No action is needed on +your end; the change will fully propagate once cached records expire. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/email-missing.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/email-missing.md new file mode 100644 index 0000000..3168abc --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/email-missing.md @@ -0,0 +1,7 @@ +Your deliverability issue is caused by a missing SPF include for our +outbound mail relay. Without it, receiving MTAs can't validate that our +servers are authorized senders for your domain, so DMARC alignment fails +and messages get soft-rejected or routed to spam. I've updated your SPF +TXT record to include our sending range. Propagation of the DNS change +should complete within a couple hours, after which SPF alignment should +pass and deliverability should normalize. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/login-locked.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/login-locked.md new file mode 100644 index 0000000..8eedcbd --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/login-locked.md @@ -0,0 +1,6 @@ +Your account was locked by our rate-limiting middleware after it detected +multiple failed authentication attempts within the lockout window, which +is a brute-force mitigation measure. I've cleared the lockout flag on your +account record, so your session should authenticate normally now. If this +recurs, consider rotating your credentials via the password-reset flow to +invalidate any compromised tokens. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/slow-site.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/slow-site.md new file mode 100644 index 0000000..1ed6602 --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/fail/slow-site.md @@ -0,0 +1,7 @@ +Your slow load times are likely due to a cache invalidation issue on the +edge nodes serving your origin. When the CDN cache expires or gets purged, +requests fall through to the backend and get served uncached, which +increases server-side latency significantly. I've re-enabled edge caching +and warmed the cache, which should reduce TTFB back to baseline within a +few minutes. Let me know if latency is still elevated after the cache +finishes propagating. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/domain-change.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/domain-change.md new file mode 100644 index 0000000..1bf2de9 --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/domain-change.md @@ -0,0 +1,9 @@ +Great question — this is totally normal after switching your domain to us. +Think of your domain name like a person's name, and there's a giant address +book on the internet that looks up what "address" (computer) that name +points to. When you point your domain to us, it takes a little while — often +a few hours, sometimes up to a day — for every copy of that address book +around the world to update with your new address. Until they all catch up, +some visitors might still land on your old site. There's nothing broken and +nothing you need to redo; just give it up to 24 hours and it should sort +itself out everywhere. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/email-missing.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/email-missing.md new file mode 100644 index 0000000..2d447fe --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/email-missing.md @@ -0,0 +1,8 @@ +I looked into why your emails aren't arriving. When you send mail, the +receiving mailbox checks a kind of "approved senders list" tied to your +domain to make sure the message is really from you and not someone +pretending to be you. Right now that list is missing an entry for our +mail system, so some inboxes are quietly rejecting your messages as +suspicious. I've added the missing entry for you. It can take a couple +hours to fully take effect, but after that your emails should start +arriving normally again. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/login-locked.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/login-locked.md new file mode 100644 index 0000000..4b1cc6c --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/login-locked.md @@ -0,0 +1,7 @@ +You're seeing that message because our system noticed several failed sign-in +attempts on your account in a short time and locked it as a safety measure — +the same way a phone locks itself after too many wrong passcodes. This is +just to keep someone else from guessing their way in. I've unlocked your +account now, so you should be able to sign in right away. If it happens +again, try resetting your password using the "forgot password" link, which +sends a one-time link to your email so you don't have to remember anything. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/slow-site.md b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/slow-site.md new file mode 100644 index 0000000..e952233 --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.fixtures/pass/slow-site.md @@ -0,0 +1,7 @@ +Sorry your site's been loading slowly! A common reason is that your website +normally keeps a "quick copy" of your pages ready to show visitors instantly, +kind of like keeping leftovers in the fridge instead of cooking from scratch +every time. If that quick copy gets out of date or turned off, the site has +to rebuild the page from scratch for every visitor, which is much slower. We +can turn that quick-copy feature back on for you — it usually brings load +times back down within a few minutes. Want me to go ahead and turn it on? diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.md b/examples/04-judge-grader/rubrics/audience-appropriate.md new file mode 100644 index 0000000..b0f3cbb --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.md @@ -0,0 +1,26 @@ +You are grading whether a support answer is written for a non-technical +reader — someone with no background in networking, software, or computing. + +Call it **clean** (pass) only if: + +- It uses no unexplained jargon or acronyms. Terms like "DNS", "cache", + "CDN", "SSL/TLS", "API", "latency", "backend", "server-side", "propagate", + "TTL" may appear only if immediately explained in plain words in the same + sentence or the one right after. +- A reader with no technical background could act on the answer (e.g. know + what to do next) without looking anything up. +- Sentences are short and each one carries at most one technical idea. + +**Flag it** (fail) if: + +- It uses two or more unexplained jargon terms or acronyms, or +- It leans on technical concepts (caching layers, propagation, protocols, + server-side vs. client-side, etc.) as if the reader already understands + them, or +- A non-technical reader would need to ask "what does that mean?" to + understand the core of the answer. + +Being accurate is not enough — a technically correct answer that a +non-technical person can't follow still fails this rubric. Politeness and +formatting are not part of this rubric; grade only audience-appropriateness +of the language. diff --git a/examples/04-judge-grader/rubrics/audience-appropriate.md.calibration.json b/examples/04-judge-grader/rubrics/audience-appropriate.md.calibration.json new file mode 100644 index 0000000..ff390b2 --- /dev/null +++ b/examples/04-judge-grader/rubrics/audience-appropriate.md.calibration.json @@ -0,0 +1,56 @@ +{ + "rubricSha256": "2961bc18c8609ddc87610ba39757dd3367c38208b7afba859bfb7a972c318e4d", + "model": "haiku", + "runner": "claude-p", + "ranAt": "2026-08-29T21:27:21.613Z", + "fixtures": { + "pass": 4, + "fail": 4 + }, + "accuracy": { + "pass": 1, + "fail": 1 + }, + "verdicts": [ + { + "fixture": "pass/domain-change.md", + "expected": "pass", + "got": "pass" + }, + { + "fixture": "pass/email-missing.md", + "expected": "pass", + "got": "pass" + }, + { + "fixture": "pass/login-locked.md", + "expected": "pass", + "got": "pass" + }, + { + "fixture": "pass/slow-site.md", + "expected": "pass", + "got": "pass" + }, + { + "fixture": "fail/domain-change.md", + "expected": "fail", + "got": "fail" + }, + { + "fixture": "fail/email-missing.md", + "expected": "fail", + "got": "fail" + }, + { + "fixture": "fail/login-locked.md", + "expected": "fail", + "got": "fail" + }, + { + "fixture": "fail/slow-site.md", + "expected": "fail", + "got": "fail" + } + ] +} diff --git a/examples/04-judge-grader/scenario.json b/examples/04-judge-grader/scenario.json new file mode 100644 index 0000000..281d3a4 --- /dev/null +++ b/examples/04-judge-grader/scenario.json @@ -0,0 +1,22 @@ +{ + "name": "plain-language support answers", + "agent": "./agent.md", + "baselineSkills": ["./skills/explain-plainly.baseline.md"], + "proposedSkills": ["./skills/explain-plainly.proposed.md"], + "model": "haiku", + "runs": 3, + "maxBudgetUsd": 1, + "scenarios": [ + { + "name": "nameserver-change-support-reply", + "kind": "target", + "prompt": "Customer wrote in: \"I switched my domain's nameservers to point at your hosting two hours ago, and my site still shows the old host for some of my friends but not others. Is something broken?\" Write the reply.", + "grader": { + "type": "judge", + "rubric": "./rubrics/audience-appropriate.md", + "model": "haiku", + "minAccuracy": 0.9 + } + } + ] +} diff --git a/examples/04-judge-grader/skills/explain-plainly.baseline.md b/examples/04-judge-grader/skills/explain-plainly.baseline.md new file mode 100644 index 0000000..a55e5d7 --- /dev/null +++ b/examples/04-judge-grader/skills/explain-plainly.baseline.md @@ -0,0 +1,6 @@ +--- +name: explain-plainly-baseline +description: Baseline behavior — no audience guidance. +--- + +Answer the customer's support question accurately and completely. diff --git a/examples/04-judge-grader/skills/explain-plainly.proposed.md b/examples/04-judge-grader/skills/explain-plainly.proposed.md new file mode 100644 index 0000000..f8c314a --- /dev/null +++ b/examples/04-judge-grader/skills/explain-plainly.proposed.md @@ -0,0 +1,17 @@ +--- +name: explain-plainly-proposed +description: Answer for a non-technical audience in plain language. +--- + +Answer the customer's support question accurately and completely, but write +for someone with no technical background: + +- Do not use jargon or acronyms (e.g. "DNS", "cache", "CDN", "SSL", "API", + "latency") without immediately explaining the term in everyday words. +- Prefer short, everyday words over technical vocabulary. If a technical + word is unavoidable, follow it with a plain-language explanation in the + same sentence. +- Use a concrete, everyday analogy where it helps ("like a phone book that + translates a name into a number"). +- Keep sentences short. Avoid stacking multiple technical concepts in one + sentence. diff --git a/examples/05-any-model/README.md b/examples/05-any-model/README.md new file mode 100644 index 0000000..115e275 --- /dev/null +++ b/examples/05-any-model/README.md @@ -0,0 +1,63 @@ +# Any Model (openai runner + ollama) + +promptdiff isn't Claude-only. This example runs the same baseline-vs-proposed +compare against a local [ollama](https://ollama.com) model over its +OpenAI-compatible endpoint — no API key, no cloud cost, no `claude` binary +involved at all. The only thing that changes from `01-hello-compare` is the +top-level `"runner"`/`"baseUrl"`/`"model"` fields; the scenario shape, +grading, and CLI are identical. + +**Cost:** free (local) · **Time:** ~10s · **Requires:** ollama with a small +model pulled (`ollama pull llama3.2:1b`) + +```bash +./promptdiff compare --scenario ./examples/05-any-model/scenario.json +``` + +## Example output (illustrative — requires ollama) + +Ollama was not installed in the environment this example was built in +(`ollama: command not found`), so this compare was never actually executed — +no output below is a real run. The scenario itself was verified to load and +validate: `loadCompareConfig` resolves both arms to +`{ model: "llama3.2:1b", runner: "openai", baseUrl: "http://localhost:11434/v1" }`, +and running `./promptdiff compare` against it proceeds all the way to an +HTTP `POST http://localhost:11434/v1/chat/completions`, which only then fails +with a connection error — i.e. every promptdiff-side check (config parsing, +skill resolution, runner-capability validation) passes; the only missing +piece is a running ollama server. The block below is shaped exactly like +`formatCompareSummary`'s real output (see `src/engine/compare.ts`), with +plausible numbers substituted: + +``` +promptdiff compare: any-model-ollama + +validation-explanation (target) + baseline: 1/3 pass (33%) | $0.0000 + proposed: 3/3 pass (100%) | $0.0000 + delta: +67% pass | +$0.0000 + PASS: assertions satisfied + NOTE: delta could be sampling noise (Fisher exact p=0.40) — consider more runs + +total cost: $0.0000 +``` + +## What to notice + +- The arm config is just `"runner": "openai"` + `"baseUrl": "http://localhost:11434/v1"` + at the top level of the scenario file — no code changes, no separate + runner integration, and (for a local server) no API key. +- Cost reports `$0.0000` for both arms: the openai runner only prices a run + when `"pricing"` is set in the scenario, and a local server has no dollar + cost to report anyway. +- The `NOTE: ... sampling noise` line is `compare`'s built-in guardrail + against over-reading a small-n delta — at `runs: 3` even a real + improvement often can't clear p=0.05, so bump `--runs` before trusting a + win here. +- The same `runner: "openai"` + `baseUrl` shape works unmodified against + vLLM, llama.cpp, and OpenRouter — swap the URL (and model name) and + nothing else changes. +- Documented limitation: the openai runner is text-graded and single-turn + only — no tools, no sandbox execution, no skill registry, and no turn cap + (there's only one turn). Scenarios needing artifact mode, command graders, + or install delivery are rejected up front; use `claude-p` for those. diff --git a/examples/05-any-model/agent.md b/examples/05-any-model/agent.md new file mode 100644 index 0000000..14da15a --- /dev/null +++ b/examples/05-any-model/agent.md @@ -0,0 +1,7 @@ +--- +name: api-guidance +description: Minimal example agent for promptdiff's any-model example. +--- + +You are a backend engineer answering short technical questions about a +REST API. Answer directly, in three sentences or fewer. diff --git a/examples/05-any-model/baseline.md b/examples/05-any-model/baseline.md new file mode 100644 index 0000000..e9bf9aa --- /dev/null +++ b/examples/05-any-model/baseline.md @@ -0,0 +1,6 @@ +--- +name: reply-format +description: Formatting rules for assistant replies. +--- + +Answer the user's question directly. Do not add disclaimers or hedging. diff --git a/examples/05-any-model/proposed.md b/examples/05-any-model/proposed.md new file mode 100644 index 0000000..c9ba7ec --- /dev/null +++ b/examples/05-any-model/proposed.md @@ -0,0 +1,10 @@ +--- +name: reply-format +description: Formatting rules for assistant replies. +--- + +Answer the user's question directly. Do not add disclaimers or hedging. + +When the question is about validating API input, always state the HTTP +status code returned for invalid input (e.g. "400 Bad Request") — never +leave the failure response implicit. diff --git a/examples/05-any-model/scenario.json b/examples/05-any-model/scenario.json new file mode 100644 index 0000000..6388b8c --- /dev/null +++ b/examples/05-any-model/scenario.json @@ -0,0 +1,22 @@ +{ + "name": "any-model-ollama", + "agent": "./agent.md", + "baselineSkills": ["./baseline.md"], + "proposedSkills": ["./proposed.md"], + "runner": "openai", + "baseUrl": "http://localhost:11434/v1", + "model": "llama3.2:1b", + "runs": 3, + "scenarios": [ + { + "name": "validation-explanation", + "kind": "target", + "prompt": "Explain how you would validate POST /api/items.", + "grader": { + "type": "text", + "contains": ["400"], + "notContains": ["TODO"] + } + } + ] +} diff --git a/examples/06-token-economics/.gitignore b/examples/06-token-economics/.gitignore new file mode 100644 index 0000000..c68962b --- /dev/null +++ b/examples/06-token-economics/.gitignore @@ -0,0 +1,2 @@ +raw/ +raw-capped/ diff --git a/examples/06-token-economics/README.md b/examples/06-token-economics/README.md new file mode 100644 index 0000000..a64f815 --- /dev/null +++ b/examples/06-token-economics/README.md @@ -0,0 +1,133 @@ +# Token economics + +`measure`'s summary line collapses a run down to `$0.1733` — a single +cost-USD number that blends fresh input tokens, cache-write tokens, +cache-read tokens, and output tokens into one weighted sum (cache reads are +priced far below fresh input). That's the right altitude for "did this pass, +and what did it cost" — and the wrong altitude for "did my prompt change +actually cut how much the agent explores, or did it just shift the same +reads from fresh input into cache?" Answering that needs the token +categories themselves, which is what `--raw-out` writes: each run's full +runner result JSON, `usage` block included. This example pairs that with +`--max-turns`, because a token-economics number is only honest if every run +it's computed from actually finished — a run that got cut off mid-exploration +and stumbled onto a passing answer would silently corrupt the average, so a +capped run is scored as a failed run instead of graded on its partial output. + +The task: a `code-investigator` agent is dropped into a 5-file sandboxed +Python "service" (`seed/`) and asked which file implements exponential +backoff for retries. The word "backoff" doesn't appear anywhere in the seed +files, so the agent can't grep its way to the answer — it has to actually +read the files and understand what the code does. The correct file is +`queue.py`. + +**Cost:** ~$0.19 · **Time:** ~32s · **Requires:** claude CLI + +## Run it + +Generous cap — runs finish normally, raw JSON lands per run: + +```bash +./promptdiff measure \ + --scenario ./examples/06-token-economics/scenario.json \ + --max-turns 8 \ + --raw-out ./examples/06-token-economics/raw +``` + +Tight cap — same scenario, same seed, capped at 1 turn: + +```bash +./promptdiff measure \ + --scenario ./examples/06-token-economics/scenario.json \ + --max-turns 1 \ + --raw-out ./examples/06-token-economics/raw-capped +``` + +## Actual output + +Generous cap (`--max-turns 8`): + +``` +[promptdiff] scenario find-retry-backoff-module +[promptdiff] measure run 1/2 +[promptdiff] measure run 2/2 +promptdiff measure: token-economics (sonnet via claude-p) + +find-retry-backoff-module + 2/2 pass (100%) | $0.1733 + +total cost: $0.1733 +``` + +Tight cap (`--max-turns 1`): + +``` +[promptdiff] scenario find-retry-backoff-module +[promptdiff] measure run 1/2 +[promptdiff] measure run 2/2 +promptdiff measure: token-economics (sonnet via claude-p) + +find-retry-backoff-module + 0/2 pass (0%) | $0.0138 + run 1 failed: hit the 1-turn cap before finishing + run 2 failed: hit the 1-turn cap before finishing + +total cost: $0.0138 +``` + +Trimmed `usage` block from one raw file of the generous-cap run +(`raw/find-retry-backoff-module_measure_1.json`) — the agent actually +finished, having read through the seed files across 7 turns: + +```json +{ + "total_cost_usd": 0.11324400000000001, + "usage": { + "input_tokens": 14, + "cache_creation_input_tokens": 16451, + "cache_read_input_tokens": 197660, + "output_tokens": 788 + }, + "num_turns": 7, + "subtype": "success", + "result": "This clearly implements exponential backoff (`2 ** attempts`, capped, used in retry loop). I'm confident in the answer.\n\nqueue.py" +} +``` + +And the same block from the tight-cap run +(`raw-capped/find-retry-backoff-module_measure_1.json`) — cut off after its +first tool call, with almost no fresh exploration and no cache built up yet: + +```json +{ + "total_cost_usd": 0.006822400000000001, + "usage": { + "input_tokens": 2, + "cache_creation_input_tokens": 0, + "cache_read_input_tokens": 29392, + "output_tokens": 94 + }, + "num_turns": 2, + "subtype": "error_max_turns", + "errors": ["Reached maximum number of turns (1)"] +} +``` + +## What to notice + +- The summary's `$0.1733` and `$0.0138` are single numbers; the raw files + underneath are where `input_tokens` (14), `cache_creation_input_tokens` + (16,451), `cache_read_input_tokens` (197,660), and `output_tokens` (788) + live separately — that's the breakdown a real "did this change cut + exploration or just move it into cache?" comparison would diff. +- The tight-cap run's `subtype` is `error_max_turns`, not `success`, and its + `usage.cache_read_input_tokens` (29,392) is over 6x smaller than the + finished run's — it never got far enough to build up the full-context + reads that come from actually working through all 5 seed files. +- Both capped runs failed with `hit the 1-turn cap before finishing`, and + neither was graded — the harness never even looked at whether the model's + cut-off partial output happened to mention `queue.py`. That's the point: + a turn-capped run is a measured outcome, not a cheap opportunity to pass. +- Raw files are written as `_measure_.json` the moment each run + finishes, one per run — so `--raw-out` gives per-run granularity that the + summary's aggregate cost/pass-rate line can't. diff --git a/examples/06-token-economics/agent.md b/examples/06-token-economics/agent.md new file mode 100644 index 0000000..199aa63 --- /dev/null +++ b/examples/06-token-economics/agent.md @@ -0,0 +1,8 @@ +--- +name: code-investigator +description: Minimal example agent for promptdiff's token-economics example. +--- + +You are a careful code investigator. You are dropped into a small, +unfamiliar codebase and asked one question about it. Answer only from what +you actually read in the files — never guess from a filename alone. diff --git a/examples/06-token-economics/scenario.json b/examples/06-token-economics/scenario.json new file mode 100644 index 0000000..bf9e4ae --- /dev/null +++ b/examples/06-token-economics/scenario.json @@ -0,0 +1,24 @@ +{ + "name": "token-economics", + "agent": "./agent.md", + "skills": ["./skill.md"], + "model": "sonnet", + "runs": 2, + "maxBudgetUsd": 2, + "timeoutMs": 120000, + "mode": "artifact", + "sandbox": { + "root": ".promptdiff/runs", + "seed": "./seed" + }, + "scenarios": [ + { + "name": "find-retry-backoff-module", + "prompt": "This directory is a small Python service. Which file implements exponential backoff for retries? Answer with only the filename.", + "grader": { + "type": "text", + "contains": ["queue.py"] + } + } + ] +} diff --git a/examples/06-token-economics/seed/auth.py b/examples/06-token-economics/seed/auth.py new file mode 100644 index 0000000..8bbf6d3 --- /dev/null +++ b/examples/06-token-economics/seed/auth.py @@ -0,0 +1,15 @@ +"""Session and token handling.""" + +import secrets + + +def create_session(user_id): + return {"user_id": user_id, "token": _new_token()} + + +def _new_token(): + return secrets.token_hex(16) + + +def validate_session(session, token): + return session.get("token") == token diff --git a/examples/06-token-economics/seed/cache.py b/examples/06-token-economics/seed/cache.py new file mode 100644 index 0000000..6d0a3a1 --- /dev/null +++ b/examples/06-token-economics/seed/cache.py @@ -0,0 +1,24 @@ +"""A tiny LRU cache used by the router.""" + + +class LRUCache: + def __init__(self, capacity): + self.capacity = capacity + self._data = {} + self._order = [] + + def get(self, key): + if key not in self._data: + return None + self._order.remove(key) + self._order.append(key) + return self._data[key] + + def put(self, key, value): + if key in self._data: + self._order.remove(key) + elif len(self._data) >= self.capacity: + oldest = self._order.pop(0) + del self._data[oldest] + self._data[key] = value + self._order.append(key) diff --git a/examples/06-token-economics/seed/db.py b/examples/06-token-economics/seed/db.py new file mode 100644 index 0000000..ffa3eb7 --- /dev/null +++ b/examples/06-token-economics/seed/db.py @@ -0,0 +1,18 @@ +"""Connection pool for the primary database.""" + + +class ConnectionPool: + def __init__(self, size): + self.size = size + self._pool = [] + + def acquire(self): + if not self._pool: + return self._connect() + return self._pool.pop() + + def release(self, conn): + self._pool.append(conn) + + def _connect(self): + return object() # placeholder connection diff --git a/examples/06-token-economics/seed/queue.py b/examples/06-token-economics/seed/queue.py new file mode 100644 index 0000000..806d40f --- /dev/null +++ b/examples/06-token-economics/seed/queue.py @@ -0,0 +1,21 @@ +"""Background job queue with retry handling.""" + +import time + + +def enqueue(job): + return {"job": job, "attempts": 0} + + +def retry_delay(attempts): + # Double the wait after each failed attempt, capped at 30 seconds. + return min(2 ** attempts, 30) + + +def run_with_retries(job, fn, max_attempts=5): + for attempt in range(max_attempts): + try: + return fn(job) + except Exception: + time.sleep(retry_delay(attempt)) + raise RuntimeError("job failed after max attempts") diff --git a/examples/06-token-economics/seed/router.py b/examples/06-token-economics/seed/router.py new file mode 100644 index 0000000..f0d5fa5 --- /dev/null +++ b/examples/06-token-economics/seed/router.py @@ -0,0 +1,18 @@ +"""Maps incoming HTTP paths to handler functions.""" + +ROUTES = {} + + +def route(path): + def decorator(fn): + ROUTES[path] = fn + return fn + + return decorator + + +def dispatch(path, request): + handler = ROUTES.get(path) + if handler is None: + return {"status": 404} + return handler(request) diff --git a/examples/06-token-economics/skill.md b/examples/06-token-economics/skill.md new file mode 100644 index 0000000..e30f82a --- /dev/null +++ b/examples/06-token-economics/skill.md @@ -0,0 +1,13 @@ +--- +name: file-by-file-investigation +description: How to investigate a small codebase before answering a question about it. +--- + +Investigate by reading the files in the working directory **one at a time** +with your file-reading tool, in the order you find them. Build your +understanding from their actual contents — do not shortcut the investigation +with a single search command across all files at once. + +Once you have read enough files to be sure of the answer, reply with exactly +one line: the filename that answers the question, and nothing else — no +explanation, no punctuation, no surrounding sentence. diff --git a/examples/README.md b/examples/README.md new file mode 100644 index 0000000..00b4e37 --- /dev/null +++ b/examples/README.md @@ -0,0 +1,22 @@ +# Examples + +Six runnable, self-contained examples, ordered as a learning curve. Each +directory holds everything it needs (agent, skills, scenario file) plus its +own README with the exact command, an honest cost label, and output captured +from a real run — so you can see what promptdiff gives you before spending +anything yourself. + +Run them from the repo root. All except 05 need the `claude` CLI installed; +none need an API key beyond what `claude` already uses. + +| Example | What it shows | Cost | +| --- | --- | --- | +| [01-hello-compare](./01-hello-compare/) | The smallest baseline-vs-proposed comparison: one added instruction, one text grader, a 0% → 100% delta. Start here. | ~$0.01 | +| [02-fix-a-regression](./02-fix-a-regression/) | The core loop: baseline reproduces a realistic defect, proposed fixes it, a second scenario proves nothing regressed. | ~$0.10 | +| [03-measure-first](./03-measure-first/) | Single-arm `measure`: put a number on a flaky prompt (here: a genuine ~50% pass rate) before you try to fix it. | ~$0.01 | +| [04-judge-grader](./04-judge-grader/) | A calibrated LLM judge end to end: rubric, labeled fixtures, `calibrate`, then a compare — and the refuse-to-grade gate when calibration is missing. | ~$0.10 | +| [05-any-model](./05-any-model/) | The `openai` runner against local ollama: not Claude-only, zero API keys. Same shape works for vLLM, llama.cpp, OpenRouter. | free (local) | +| [06-token-economics](./06-token-economics/) | Turn caps (`--max-turns`) + raw per-run results (`--raw-out`): token-category analysis, and how a capped run is failed without grading. | ~$0.19 | + +Every `scenario.json` here is loaded by `test/examples.test.ts` in CI, so the +examples can't silently drift from the config schema. diff --git a/test/examples.test.ts b/test/examples.test.ts new file mode 100644 index 0000000..6eda772 --- /dev/null +++ b/test/examples.test.ts @@ -0,0 +1,27 @@ +import { readdirSync } from "node:fs"; +import { join } from "node:path"; +import { expect, test } from "bun:test"; +import { loadCompareConfig } from "../src/engine/config"; + +// Every shipped example must load through the real config loader, so schema +// drift breaks CI instead of the first user who copies an example. +const examplesRoot = join(import.meta.dir, "..", "examples"); + +const exampleDirs = readdirSync(examplesRoot, { withFileTypes: true }) + .filter((entry) => entry.isDirectory()) + .map((entry) => entry.name) + .sort(); + +test("examples directory is not empty", () => { + expect(exampleDirs.length).toBeGreaterThan(0); +}); + +for (const dir of exampleDirs) { + test(`examples/${dir}/scenario.json loads (fixture paths included)`, () => { + const scenario = join(examplesRoot, dir, "scenario.json"); + // Measure-only examples have no proposed arm; singleArm accepts both + // shapes, and loading also validates that referenced files exist. + const config = loadCompareConfig(scenario, {}, { singleArm: true }); + expect(config.cases.length).toBeGreaterThan(0); + }); +} From 4f9af0e8cc9c50d7238701a53b9eaac9d08cb1e7 Mon Sep 17 00:00:00 2001 From: Josh Owens Date: Fri, 18 Sep 2026 18:36:06 +0000 Subject: [PATCH 2/2] docs: correct example output claims and 06's grader note Three inaccuracies the examples shipped with, plus one of the same kind found while checking: - examples/README.md and the root README claimed every example's output was captured from a real run. 05 was never executed (no ollama in the build environment) and its own README says so. Both now name 05 as the exception. - The root README's "See it work" block called itself real output while omitting the NOTE sampling-noise line and the two baseline failure lines that a real run of 01 prints. It now shows 01's full output. - 05's illustrative block claimed to be shaped exactly like formatCompareSummary's output but omitted the per-failing-run lines that a 1/3 baseline emits. Added. - 06's README did not say that its grader is a substring check, so a reader could assume the 2/2 pass rate measured the single-line format the prompt and skill.md ask for. Both passing runs put a sentence of reasoning ahead of the filename and still passed. Stated plainly, with the regex grader that would enforce the format. Co-Authored-By: Claude Opus 5 --- README.md | 21 ++++++++++++++------- examples/05-any-model/README.md | 2 ++ examples/06-token-economics/README.md | 9 +++++++++ examples/README.md | 11 +++++++---- 4 files changed, 32 insertions(+), 11 deletions(-) diff --git a/README.md b/README.md index d66947a..99bb34b 100644 --- a/README.md +++ b/README.md @@ -81,17 +81,24 @@ answers-with-summary-prefix (target) proposed: 2/2 pass (100%) | $0.0033 delta: +100% pass | +$0.0013 PASS: assertions satisfied + NOTE: delta could be sampling noise (Fisher exact p=0.33) — consider more runs + baseline run 1 failed: output did not contain "SUMMARY:" + baseline run 2 failed: output did not contain "SUMMARY:" total cost: $0.0054 ``` -That's real output from a real run (about a cent, ~12 seconds). The -[`examples/`](./examples/) directory has six runnable, self-contained -examples ordered as a learning curve — from this hello case through the -fix-a-defect-without-regressing loop, measure-first characterization, -calibrated LLM judges, non-Claude models via ollama, and token-level -analysis with turn caps — each with its own README, honest cost label, -and captured output. +That's the full output of a real run (about a cent, ~12 seconds), `NOTE:` +line included: at `runs: 2` even a clean 0/2 → 2/2 flip can't rule out +sampling noise, and `compare` says so rather than letting you over-read the +delta. The [`examples/`](./examples/) directory has six runnable, +self-contained examples ordered as a learning curve — from this hello case +through the fix-a-defect-without-regressing loop, measure-first +characterization, calibrated LLM judges, non-Claude models via ollama, and +token-level analysis with turn caps — each with its own README, honest cost +label, and output. That output is captured from a real run in every example +except `05`, which needs a local ollama server and labels its block +illustrative. ## Safe Defaults diff --git a/examples/05-any-model/README.md b/examples/05-any-model/README.md index 115e275..d180708 100644 --- a/examples/05-any-model/README.md +++ b/examples/05-any-model/README.md @@ -38,6 +38,8 @@ validation-explanation (target) delta: +67% pass | +$0.0000 PASS: assertions satisfied NOTE: delta could be sampling noise (Fisher exact p=0.40) — consider more runs + baseline run 2 failed: output did not contain "400" + baseline run 3 failed: output did not contain "400" total cost: $0.0000 ``` diff --git a/examples/06-token-economics/README.md b/examples/06-token-economics/README.md index a64f815..44c41c9 100644 --- a/examples/06-token-economics/README.md +++ b/examples/06-token-economics/README.md @@ -128,6 +128,15 @@ first tool call, with almost no fresh exploration and no cache built up yet: neither was graded — the harness never even looked at whether the model's cut-off partial output happened to mention `queue.py`. That's the point: a turn-capped run is a measured outcome, not a cheap opportunity to pass. +- The grader is `{"type": "text", "contains": ["queue.py"]}` — a substring + check, so it scores *named the right file*, not *obeyed the format*. Both + passing runs put a sentence of reasoning ahead of the filename (visible in + the `result` field above) even though the prompt says "Answer with only the + filename" and `skill.md` asks for exactly one line, and both still passed + 2/2. To enforce the format too, swap in a regex grader such as + `{"type": "text", "regex": ["^queue\\.py\\s*$"]}` — `gradeText` compiles + patterns without the `m` flag, so `^` anchors to the start of the whole + output, not to any line in it. - Raw files are written as `_measure_.json` the moment each run finishes, one per run — so `--raw-out` gives per-run granularity that the summary's aggregate cost/pass-rate line can't. diff --git a/examples/README.md b/examples/README.md index 00b4e37..0585bad 100644 --- a/examples/README.md +++ b/examples/README.md @@ -2,9 +2,12 @@ Six runnable, self-contained examples, ordered as a learning curve. Each directory holds everything it needs (agent, skills, scenario file) plus its -own README with the exact command, an honest cost label, and output captured -from a real run — so you can see what promptdiff gives you before spending -anything yourself. +own README with the exact command, an honest cost label, and its output — so +you can see what promptdiff gives you before spending anything yourself. + +The output in 01-04 and 06 is captured from a real run. 05 is the exception: +it needs a local ollama server, which was not available where these examples +were built, so its README shows an illustrative block and says so. Run them from the repo root. All except 05 need the `claude` CLI installed; none need an API key beyond what `claude` already uses. @@ -15,7 +18,7 @@ none need an API key beyond what `claude` already uses. | [02-fix-a-regression](./02-fix-a-regression/) | The core loop: baseline reproduces a realistic defect, proposed fixes it, a second scenario proves nothing regressed. | ~$0.10 | | [03-measure-first](./03-measure-first/) | Single-arm `measure`: put a number on a flaky prompt (here: a genuine ~50% pass rate) before you try to fix it. | ~$0.01 | | [04-judge-grader](./04-judge-grader/) | A calibrated LLM judge end to end: rubric, labeled fixtures, `calibrate`, then a compare — and the refuse-to-grade gate when calibration is missing. | ~$0.10 | -| [05-any-model](./05-any-model/) | The `openai` runner against local ollama: not Claude-only, zero API keys. Same shape works for vLLM, llama.cpp, OpenRouter. | free (local) | +| [05-any-model](./05-any-model/) | The `openai` runner against local ollama: not Claude-only, zero API keys. Same shape works for vLLM, llama.cpp, OpenRouter. Output illustrative, not a real run. | free (local) | | [06-token-economics](./06-token-economics/) | Turn caps (`--max-turns`) + raw per-run results (`--raw-out`): token-category analysis, and how a capped run is failed without grading. | ~$0.19 | Every `scenario.json` here is loaded by `test/examples.test.ts` in CI, so the