-
Notifications
You must be signed in to change notification settings - Fork 0
docs: six runnable examples with captured output #40
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
2 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,48 @@ | ||
| # Hello Compare | ||
|
|
||
| The smallest possible `promptdiff compare`: one agent, one scenario, one | ||
| grader. The baseline skill is missing a formatting instruction; the proposed | ||
| skill adds it; a text grader catches the difference. This is the shape every | ||
| other example builds on — start here before reading the rest. | ||
|
|
||
| **Cost:** ~$0.01 · **Time:** ~12s · **Requires:** claude CLI | ||
|
|
||
| ## Run it | ||
|
|
||
| ```bash | ||
| ./promptdiff compare --scenario ./examples/01-hello-compare/scenario.json | ||
| ``` | ||
|
|
||
| ## Output | ||
|
|
||
| ``` | ||
| promptdiff compare: hello-compare | ||
|
|
||
| answers-with-summary-prefix (target) | ||
| baseline: 0/2 pass (0%) | $0.0021 | ||
| proposed: 2/2 pass (100%) | $0.0033 | ||
| delta: +100% pass | +$0.0013 | ||
| PASS: assertions satisfied | ||
| NOTE: delta could be sampling noise (Fisher exact p=0.33) — consider more runs | ||
| baseline run 1 failed: output did not contain "SUMMARY:" | ||
| baseline run 2 failed: output did not contain "SUMMARY:" | ||
|
|
||
| total cost: $0.0054 | ||
| ``` | ||
|
|
||
| ## What to notice | ||
|
|
||
| - `baseline.md` and `proposed.md` are the same skill except for one added | ||
| instruction ("start your reply with `SUMMARY:`"). That's the entire A/B | ||
| variable — the agent and the scenario prompt never change. | ||
| - The pass-rate delta (0% → 100%) is the whole point: the grader | ||
| (`{"type": "text", "contains": ["SUMMARY:"]}`) can't see the instruction | ||
| text, only its effect on the model's output. | ||
| - This scenario is `"kind": "target"`, so `compare` asserts baseline must | ||
| *not* fully pass (the gap is real) and proposed must improve on it. Both | ||
| hold here, so the process exits `0`; either failing would exit non-zero — | ||
| to see the failing case, edit `scenario.json` so `proposedSkills` points at | ||
| `./baseline.md` and rerun. | ||
| - At `runs: 2` the delta is real but small-sample — the `NOTE:` line about | ||
| sampling noise is `compare` being honest about that, not a bug. It doesn't | ||
| change the exit code. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,7 @@ | ||
| --- | ||
| name: hello-assistant | ||
| description: Minimal example agent for promptdiff's hello-compare example. | ||
| --- | ||
|
|
||
| You are a helpful assistant that answers user questions directly and | ||
| concisely, in one or two sentences. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,6 @@ | ||
| --- | ||
| name: reply-format | ||
| description: Formatting rules for assistant replies. | ||
| --- | ||
|
|
||
| Answer the user's question directly. Do not add disclaimers or hedging. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,9 @@ | ||
| --- | ||
| name: reply-format | ||
| description: Formatting rules for assistant replies. | ||
| --- | ||
|
|
||
| Answer the user's question directly. Do not add disclaimers or hedging. | ||
|
|
||
| Start your reply with a line that reads exactly `SUMMARY:` followed by a | ||
| one-sentence answer, on its own line, before any further detail. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,19 @@ | ||
| { | ||
| "name": "hello-compare", | ||
| "agent": "./agent.md", | ||
| "baselineSkills": ["./baseline.md"], | ||
| "proposedSkills": ["./proposed.md"], | ||
| "model": "haiku", | ||
| "runs": 2, | ||
| "scenarios": [ | ||
| { | ||
| "name": "answers-with-summary-prefix", | ||
| "kind": "target", | ||
| "prompt": "What is the capital of France?", | ||
| "grader": { | ||
| "type": "text", | ||
| "contains": ["SUMMARY:"] | ||
| } | ||
| } | ||
| ] | ||
| } |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,77 @@ | ||
| # Fix a Regression | ||
|
|
||
| This demonstrates the core loop from promptdiff's ["Why it exists"](../../README.md#why-it-exists): | ||
| for a recurring defect, a trustworthy eval has to show two things at once — | ||
| (1) the **baseline** instruction set still reproduces the failure, and (2) the | ||
| **proposed** instruction set fixes it **without regressing** a case that | ||
| already worked. | ||
|
|
||
| The skill under test writes git commit messages. The baseline skill tells the | ||
| model to follow Conventional Commits and "be thorough" about what changed, | ||
| but never states a length limit — so on any change with more than one moving | ||
| part, the model pads the subject line describing all of them, blowing past | ||
| the 50-character convention. The proposed skill adds one explicit rule: keep | ||
| the subject line ≤50 chars, put extra detail in the body. | ||
|
|
||
| Two scenarios exercise this: | ||
|
|
||
| - `target-multi-part-change` (**target**) — a change with four distinct parts | ||
| (swap the session store, add expiry, update handlers, migrate tests). | ||
| Baseline is expected to fail here; proposed is expected to fix it. | ||
| - `regression-simple-typo-fix` (**regression**) — a trivial one-line change | ||
| both arms already handle fine. This proves the new length rule doesn't make | ||
| the skill worse on cases it wasn't broken on. | ||
|
|
||
| The grader is a deterministic text/regex check (no LLM judge needed): the | ||
| first line must match Conventional Commits format, and — separately — the | ||
| first line must be ≤50 characters. | ||
|
|
||
| **Cost:** ~$0.10 · **Time:** ~188s · **Requires:** claude CLI | ||
|
|
||
| ## Run it | ||
|
|
||
| ```bash | ||
| ./promptdiff compare --scenario examples/02-fix-a-regression/scenario.json | ||
| ``` | ||
|
|
||
| ## Actual output | ||
|
|
||
| ``` | ||
| promptdiff compare: commit-message subject length | ||
|
|
||
| target-multi-part-change (target) | ||
| baseline: 0/3 pass (0%) | $0.0209 | ||
| proposed: 3/3 pass (100%) | $0.0537 | ||
| delta: +100% pass | +$0.0328 | ||
| PASS: assertions satisfied | ||
| NOTE: delta could be sampling noise (Fisher exact p=0.10) — consider more runs | ||
| baseline run 1 failed: output did not match /^.{1,50}(\n|$)/ | ||
| baseline run 2 failed: output did not match /^.{1,50}(\n|$)/ | ||
| baseline run 3 failed: output did not match /^.{1,50}(\n|$)/ | ||
|
|
||
| regression-simple-typo-fix (regression) | ||
| baseline: 3/3 pass (100%) | $0.0135 | ||
| proposed: 3/3 pass (100%) | $0.0133 | ||
| delta: +0% pass | $-0.0002 | ||
| PASS: assertions satisfied | ||
|
|
||
| total cost: $0.1014 | ||
| ``` | ||
|
|
||
| ## What to notice | ||
|
|
||
| - **Baseline fails 0/3, proposed passes 3/3** on the target case — the | ||
| defect is real and reproducible, not cherry-picked, and the fix clears it | ||
| every run at `runs: 3`. | ||
| - **The regression scenario stays 3/3 → 3/3.** Adding the length rule didn't | ||
| make the skill worse on a case it already handled — that's the "no | ||
| regression" half of the claim, checked automatically rather than asserted | ||
| by eye. | ||
| - The grader is two independent regex checks (`type(scope): …` format, and | ||
| overall line length ≤50 via `^.{1,50}(\n|$)`) plus a `notContains` guard | ||
| against code fences — no LLM judge, no flakiness from a second model's | ||
| opinion. | ||
| - The `NOTE: delta could be sampling noise (Fisher exact p=0.10)` line on the | ||
| target scenario is expected at `runs: 3` — even a clean 0/3 → 3/3 flip | ||
| can't rule out noise at this sample size; the assertions and exit code are | ||
| unaffected. Raise `--runs` if you need a tighter p-value. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,12 @@ | ||
| --- | ||
| name: commit-message-writer | ||
| --- | ||
|
|
||
| # Commit Message Writer | ||
|
|
||
| You write git commit messages. The user will describe a code change; you | ||
| respond with the commit message for it, following the rules in your skill | ||
| instructions. | ||
|
|
||
| Output *only* the commit message text — no preamble ("Here's a commit | ||
| message:"), no trailing commentary, no code fences. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,38 @@ | ||
| { | ||
| "name": "commit-message subject length", | ||
| "agent": "./agent.md", | ||
| "baselineSkills": ["./skills/commit-message.baseline.md"], | ||
| "proposedSkills": ["./skills/commit-message.proposed.md"], | ||
| "model": "haiku", | ||
| "mode": "text", | ||
| "runs": 3, | ||
| "maxBudgetUsd": 1, | ||
| "scenarios": [ | ||
| { | ||
| "name": "target-multi-part-change", | ||
| "kind": "target", | ||
| "prompt": "Write a commit message for this change: replaced the in-memory session store with Redis-backed sessions, added automatic session expiry after 30 minutes of inactivity, updated the login and logout handlers to use the new session store, and migrated the existing session tests to mock Redis instead of the in-memory map.", | ||
| "grader": { | ||
| "type": "text", | ||
| "regex": [ | ||
| "^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S", | ||
| "^.{1,50}(\\n|$)" | ||
| ], | ||
| "notContains": ["```"] | ||
| } | ||
| }, | ||
| { | ||
| "name": "regression-simple-typo-fix", | ||
| "kind": "regression", | ||
| "prompt": "Write a commit message for this change: fixed a typo in the README.", | ||
| "grader": { | ||
| "type": "text", | ||
| "regex": [ | ||
| "^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S", | ||
| "^.{1,50}(\\n|$)" | ||
| ], | ||
| "notContains": ["```"] | ||
| } | ||
| } | ||
| ] | ||
| } |
15 changes: 15 additions & 0 deletions
15
examples/02-fix-a-regression/skills/commit-message.baseline.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,15 @@ | ||
| --- | ||
| name: commit-message | ||
| --- | ||
|
|
||
| # Commit Message Skill | ||
|
|
||
| Follow the Conventional Commits format for the subject line: | ||
| `type(scope): description` (scope optional). | ||
|
|
||
| Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore, | ||
| revert. | ||
|
|
||
| Write a subject line that clearly and specifically describes what changed. | ||
| Be thorough — a vague subject line is not useful to someone reading `git | ||
| log` later. If the change touches several things, say what they are. |
22 changes: 22 additions & 0 deletions
22
examples/02-fix-a-regression/skills/commit-message.proposed.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,22 @@ | ||
| --- | ||
| name: commit-message | ||
| --- | ||
|
|
||
| # Commit Message Skill | ||
|
|
||
| Follow the Conventional Commits format for the subject line: | ||
| `type(scope): description` (scope optional). | ||
|
|
||
| Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore, | ||
| revert. | ||
|
|
||
| Write a subject line that clearly and specifically describes what changed. | ||
| Be thorough — a vague subject line is not useful to someone reading `git | ||
| log` later. If the change touches several things, say what they are. | ||
|
|
||
| **The subject line must be 50 characters or fewer, counting the whole | ||
| line** (type, scope, colon, and description). Use imperative mood ("add", | ||
| not "added"). If the change needs more explanation than fits in 50 | ||
| characters, put the detail in the body after a blank line — never stretch | ||
| the subject line to fit it. When several things changed, name the most | ||
| important one in the subject and list the rest in the body. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,70 @@ | ||
| # Measure first | ||
|
|
||
| Before you rewrite a prompt you suspect is flaky, measure its current pass | ||
| rate. Skip that step and any before/after comparison is a coin flip against | ||
| a coin flip: two runs of the *same* prompt can already look like a fix or a | ||
| regression from sampling noise alone. `measure` runs one instruction set N | ||
| times and reports a bare pass rate — a real number to rewrite against, | ||
| instead of a vibe from the last output you happened to read. | ||
|
|
||
| This example measures `skill.md`, an intentionally under-specified | ||
| extraction prompt: "extract the key fields from this support ticket as | ||
| JSON," no schema. Given the same support ticket four times, `haiku` | ||
| extracts a working JSON object every time — it just doesn't agree with | ||
| itself on the field names (`severity` vs. `urgency` vs. `priority`) or the | ||
| value casing (`"Medium"` vs. `"medium"`). The `json` grader's path | ||
| assertions catch exactly that kind of drift, because they check one exact | ||
| path and value, not "did it extract *something* reasonable." | ||
|
|
||
| **Cost:** ~$0.01 · **Time:** ~28s · **Requires:** claude CLI | ||
|
|
||
| ```bash | ||
| ./promptdiff measure --scenario ./examples/03-measure-first/scenario.json | ||
| ``` | ||
|
|
||
| Real output: | ||
|
|
||
| ``` | ||
| [promptdiff] scenario extracts-medium-severity | ||
| [promptdiff] measure run 1/4 | ||
| [promptdiff] measure run 2/4 | ||
| [promptdiff] measure run 3/4 | ||
| [promptdiff] measure run 4/4 | ||
| promptdiff measure: ticket-field-extraction (haiku via claude-p) | ||
|
|
||
| extracts-medium-severity | ||
| 2/4 pass (50%) | $0.0139 | ||
| run 1 failed: severity == "Medium": path segment "severity" not found (at {"subject":"Can't export reports since yesterday's update","reporter_name":"Dana Whitfield","reporter_email":"dana.whitf…) | ||
| run 4 failed: severity == "Medium": found "medium" | ||
|
|
||
| total cost: $0.0139 | ||
| ``` | ||
|
|
||
| ```bash | ||
| $ echo $? | ||
| 0 | ||
| ``` | ||
|
|
||
| ## What to notice | ||
|
|
||
| - **50% is the measurement, not a bug.** The grader is exact on purpose | ||
| (`severity == "Medium"`) — an under-specified prompt gets an | ||
| under-specified extraction, and the pass rate is the honest size of that | ||
| problem. Run it again and you'll see a different split (25%, 50%, ...); | ||
| that's the real variance in the prompt, not flakiness in the harness. | ||
| - **The two failure reasons are two different bugs.** Run 1's `severity` | ||
| key doesn't exist at all — the model named it `reporter_name`/ | ||
| `reporter_email` and dropped severity from that response's schema | ||
| entirely. Run 4 has the key but the wrong case (`"medium"` vs | ||
| `"Medium"`). A prompt fix for one won't fix the other — you'd need to | ||
| pin both the field name and an enum of allowed values. | ||
| - **`measure` exits 0 here even though the scenario "failed" 50% of the | ||
| time.** `echo $?` above prints `0` — a measurement has no pass/fail, so a | ||
| 49-53% swing on rerun is not a CI signal by itself. Wire a *threshold* | ||
| check around the printed rate (or graduate to `compare` with an assertion) | ||
| once you have a bar to enforce. | ||
| - This baseline is now the number a proposed fix has to beat. Pin the | ||
| schema (name the fields, constrain `severity` to an enum) in a revised | ||
| `skill.md`, drop it into a `compare` scenario as the proposed arm against | ||
| this one as baseline, and see if the pass rate — not just one output — | ||
| actually moves. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,7 @@ | ||
| --- | ||
| name: ticket-intake | ||
| --- | ||
|
|
||
| You are a support-ticket intake assistant. Follow the skill instructions | ||
| exactly and reply with nothing but the requested output — no preamble, no | ||
| markdown fences, no commentary. |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
suggestion · Arithmetic error in documented sample output: +$0.0013 should be +$0.0012
The captured output block shows baseline $0.0021 and proposed $0.0033, so the delta line should read +$0.0012, not +$0.0013. The same block is duplicated in examples/01-hello-compare/README.md (line 24). Since the README explicitly claims the output is 'captured from a real run', an internally inconsistent cost figure undermines that claim.