Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 36 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,42 @@ Run from the repo:
./promptdiff --help
```

## See it work

The smallest possible comparison — the baseline skill is missing one
instruction, the proposed skill adds it, a text grader catches the effect:

```bash
./promptdiff compare --scenario ./examples/01-hello-compare/scenario.json
```

```
promptdiff compare: hello-compare

answers-with-summary-prefix (target)
baseline: 0/2 pass (0%) | $0.0021
proposed: 2/2 pass (100%) | $0.0033
delta: +100% pass | +$0.0013

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggestion · Arithmetic error in documented sample output: +$0.0013 should be +$0.0012

The captured output block shows baseline $0.0021 and proposed $0.0033, so the delta line should read +$0.0012, not +$0.0013. The same block is duplicated in examples/01-hello-compare/README.md (line 24). Since the README explicitly claims the output is 'captured from a real run', an internally inconsistent cost figure undermines that claim.

Suggested change
delta: +100% pass | +$0.0013
Correct the delta to +$0.0012 in both README.md and examples/01-hello-compare/README.md, or re-capture the block from an actual run so the numbers are self-consistent.

PASS: assertions satisfied
NOTE: delta could be sampling noise (Fisher exact p=0.33) — consider more runs
baseline run 1 failed: output did not contain "SUMMARY:"
baseline run 2 failed: output did not contain "SUMMARY:"

total cost: $0.0054
```

That's the full output of a real run (about a cent, ~12 seconds), `NOTE:`
line included: at `runs: 2` even a clean 0/2 → 2/2 flip can't rule out
sampling noise, and `compare` says so rather than letting you over-read the
delta. The [`examples/`](./examples/) directory has six runnable,
self-contained examples ordered as a learning curve — from this hello case
through the fix-a-defect-without-regressing loop, measure-first
characterization, calibrated LLM judges, non-Claude models via ollama, and
token-level analysis with turn caps — each with its own README, honest cost
label, and output. That output is captured from a real run in every example
except `05`, which needs a local ollama server and labels its block
illustrative.

## Safe Defaults

Paid model calls are bounded by default:
Expand Down
48 changes: 48 additions & 0 deletions examples/01-hello-compare/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Hello Compare

The smallest possible `promptdiff compare`: one agent, one scenario, one
grader. The baseline skill is missing a formatting instruction; the proposed
skill adds it; a text grader catches the difference. This is the shape every
other example builds on — start here before reading the rest.

**Cost:** ~$0.01 · **Time:** ~12s · **Requires:** claude CLI

## Run it

```bash
./promptdiff compare --scenario ./examples/01-hello-compare/scenario.json
```

## Output

```
promptdiff compare: hello-compare

answers-with-summary-prefix (target)
baseline: 0/2 pass (0%) | $0.0021
proposed: 2/2 pass (100%) | $0.0033
delta: +100% pass | +$0.0013
PASS: assertions satisfied
NOTE: delta could be sampling noise (Fisher exact p=0.33) — consider more runs
baseline run 1 failed: output did not contain "SUMMARY:"
baseline run 2 failed: output did not contain "SUMMARY:"

total cost: $0.0054
```

## What to notice

- `baseline.md` and `proposed.md` are the same skill except for one added
instruction ("start your reply with `SUMMARY:`"). That's the entire A/B
variable — the agent and the scenario prompt never change.
- The pass-rate delta (0% → 100%) is the whole point: the grader
(`{"type": "text", "contains": ["SUMMARY:"]}`) can't see the instruction
text, only its effect on the model's output.
- This scenario is `"kind": "target"`, so `compare` asserts baseline must
*not* fully pass (the gap is real) and proposed must improve on it. Both
hold here, so the process exits `0`; either failing would exit non-zero —
to see the failing case, edit `scenario.json` so `proposedSkills` points at
`./baseline.md` and rerun.
- At `runs: 2` the delta is real but small-sample — the `NOTE:` line about
sampling noise is `compare` being honest about that, not a bug. It doesn't
change the exit code.
7 changes: 7 additions & 0 deletions examples/01-hello-compare/agent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
name: hello-assistant
description: Minimal example agent for promptdiff's hello-compare example.
---

You are a helpful assistant that answers user questions directly and
concisely, in one or two sentences.
6 changes: 6 additions & 0 deletions examples/01-hello-compare/baseline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
---
name: reply-format
description: Formatting rules for assistant replies.
---

Answer the user's question directly. Do not add disclaimers or hedging.
9 changes: 9 additions & 0 deletions examples/01-hello-compare/proposed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
name: reply-format
description: Formatting rules for assistant replies.
---

Answer the user's question directly. Do not add disclaimers or hedging.

Start your reply with a line that reads exactly `SUMMARY:` followed by a
one-sentence answer, on its own line, before any further detail.
19 changes: 19 additions & 0 deletions examples/01-hello-compare/scenario.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
{
"name": "hello-compare",
"agent": "./agent.md",
"baselineSkills": ["./baseline.md"],
"proposedSkills": ["./proposed.md"],
"model": "haiku",
"runs": 2,
"scenarios": [
{
"name": "answers-with-summary-prefix",
"kind": "target",
"prompt": "What is the capital of France?",
"grader": {
"type": "text",
"contains": ["SUMMARY:"]
}
}
]
}
77 changes: 77 additions & 0 deletions examples/02-fix-a-regression/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# Fix a Regression

This demonstrates the core loop from promptdiff's ["Why it exists"](../../README.md#why-it-exists):
for a recurring defect, a trustworthy eval has to show two things at once —
(1) the **baseline** instruction set still reproduces the failure, and (2) the
**proposed** instruction set fixes it **without regressing** a case that
already worked.

The skill under test writes git commit messages. The baseline skill tells the
model to follow Conventional Commits and "be thorough" about what changed,
but never states a length limit — so on any change with more than one moving
part, the model pads the subject line describing all of them, blowing past
the 50-character convention. The proposed skill adds one explicit rule: keep
the subject line ≤50 chars, put extra detail in the body.

Two scenarios exercise this:

- `target-multi-part-change` (**target**) — a change with four distinct parts
(swap the session store, add expiry, update handlers, migrate tests).
Baseline is expected to fail here; proposed is expected to fix it.
- `regression-simple-typo-fix` (**regression**) — a trivial one-line change
both arms already handle fine. This proves the new length rule doesn't make
the skill worse on cases it wasn't broken on.

The grader is a deterministic text/regex check (no LLM judge needed): the
first line must match Conventional Commits format, and — separately — the
first line must be ≤50 characters.

**Cost:** ~$0.10 · **Time:** ~188s · **Requires:** claude CLI

## Run it

```bash
./promptdiff compare --scenario examples/02-fix-a-regression/scenario.json
```

## Actual output

```
promptdiff compare: commit-message subject length

target-multi-part-change (target)
baseline: 0/3 pass (0%) | $0.0209
proposed: 3/3 pass (100%) | $0.0537
delta: +100% pass | +$0.0328
PASS: assertions satisfied
NOTE: delta could be sampling noise (Fisher exact p=0.10) — consider more runs
baseline run 1 failed: output did not match /^.{1,50}(\n|$)/
baseline run 2 failed: output did not match /^.{1,50}(\n|$)/
baseline run 3 failed: output did not match /^.{1,50}(\n|$)/

regression-simple-typo-fix (regression)
baseline: 3/3 pass (100%) | $0.0135
proposed: 3/3 pass (100%) | $0.0133
delta: +0% pass | $-0.0002
PASS: assertions satisfied

total cost: $0.1014
```

## What to notice

- **Baseline fails 0/3, proposed passes 3/3** on the target case — the
defect is real and reproducible, not cherry-picked, and the fix clears it
every run at `runs: 3`.
- **The regression scenario stays 3/3 → 3/3.** Adding the length rule didn't
make the skill worse on a case it already handled — that's the "no
regression" half of the claim, checked automatically rather than asserted
by eye.
- The grader is two independent regex checks (`type(scope): …` format, and
overall line length ≤50 via `^.{1,50}(\n|$)`) plus a `notContains` guard
against code fences — no LLM judge, no flakiness from a second model's
opinion.
- The `NOTE: delta could be sampling noise (Fisher exact p=0.10)` line on the
target scenario is expected at `runs: 3` — even a clean 0/3 → 3/3 flip
can't rule out noise at this sample size; the assertions and exit code are
unaffected. Raise `--runs` if you need a tighter p-value.
12 changes: 12 additions & 0 deletions examples/02-fix-a-regression/agent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
name: commit-message-writer
---

# Commit Message Writer

You write git commit messages. The user will describe a code change; you
respond with the commit message for it, following the rules in your skill
instructions.

Output *only* the commit message text — no preamble ("Here's a commit
message:"), no trailing commentary, no code fences.
38 changes: 38 additions & 0 deletions examples/02-fix-a-regression/scenario.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
{
"name": "commit-message subject length",
"agent": "./agent.md",
"baselineSkills": ["./skills/commit-message.baseline.md"],
"proposedSkills": ["./skills/commit-message.proposed.md"],
"model": "haiku",
"mode": "text",
"runs": 3,
"maxBudgetUsd": 1,
"scenarios": [
{
"name": "target-multi-part-change",
"kind": "target",
"prompt": "Write a commit message for this change: replaced the in-memory session store with Redis-backed sessions, added automatic session expiry after 30 minutes of inactivity, updated the login and logout handlers to use the new session store, and migrated the existing session tests to mock Redis instead of the in-memory map.",
"grader": {
"type": "text",
"regex": [
"^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S",
"^.{1,50}(\\n|$)"
],
"notContains": ["```"]
}
},
{
"name": "regression-simple-typo-fix",
"kind": "regression",
"prompt": "Write a commit message for this change: fixed a typo in the README.",
"grader": {
"type": "text",
"regex": [
"^(feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert)(\\([^)]+\\))?!?: \\S",
"^.{1,50}(\\n|$)"
],
"notContains": ["```"]
}
}
]
}
15 changes: 15 additions & 0 deletions examples/02-fix-a-regression/skills/commit-message.baseline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
---
name: commit-message
---

# Commit Message Skill

Follow the Conventional Commits format for the subject line:
`type(scope): description` (scope optional).

Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore,
revert.

Write a subject line that clearly and specifically describes what changed.
Be thorough — a vague subject line is not useful to someone reading `git
log` later. If the change touches several things, say what they are.
22 changes: 22 additions & 0 deletions examples/02-fix-a-regression/skills/commit-message.proposed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
---
name: commit-message
---

# Commit Message Skill

Follow the Conventional Commits format for the subject line:
`type(scope): description` (scope optional).

Valid types: feat, fix, docs, style, refactor, perf, test, build, ci, chore,
revert.

Write a subject line that clearly and specifically describes what changed.
Be thorough — a vague subject line is not useful to someone reading `git
log` later. If the change touches several things, say what they are.

**The subject line must be 50 characters or fewer, counting the whole
line** (type, scope, colon, and description). Use imperative mood ("add",
not "added"). If the change needs more explanation than fits in 50
characters, put the detail in the body after a blank line — never stretch
the subject line to fit it. When several things changed, name the most
important one in the subject and list the rest in the body.
70 changes: 70 additions & 0 deletions examples/03-measure-first/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Measure first

Before you rewrite a prompt you suspect is flaky, measure its current pass
rate. Skip that step and any before/after comparison is a coin flip against
a coin flip: two runs of the *same* prompt can already look like a fix or a
regression from sampling noise alone. `measure` runs one instruction set N
times and reports a bare pass rate — a real number to rewrite against,
instead of a vibe from the last output you happened to read.

This example measures `skill.md`, an intentionally under-specified
extraction prompt: "extract the key fields from this support ticket as
JSON," no schema. Given the same support ticket four times, `haiku`
extracts a working JSON object every time — it just doesn't agree with
itself on the field names (`severity` vs. `urgency` vs. `priority`) or the
value casing (`"Medium"` vs. `"medium"`). The `json` grader's path
assertions catch exactly that kind of drift, because they check one exact
path and value, not "did it extract *something* reasonable."

**Cost:** ~$0.01 · **Time:** ~28s · **Requires:** claude CLI

```bash
./promptdiff measure --scenario ./examples/03-measure-first/scenario.json
```

Real output:

```
[promptdiff] scenario extracts-medium-severity
[promptdiff] measure run 1/4
[promptdiff] measure run 2/4
[promptdiff] measure run 3/4
[promptdiff] measure run 4/4
promptdiff measure: ticket-field-extraction (haiku via claude-p)

extracts-medium-severity
2/4 pass (50%) | $0.0139
run 1 failed: severity == "Medium": path segment "severity" not found (at {"subject":"Can't export reports since yesterday's update","reporter_name":"Dana Whitfield","reporter_email":"dana.whitf…)
run 4 failed: severity == "Medium": found "medium"

total cost: $0.0139
```

```bash
$ echo $?
0
```

## What to notice

- **50% is the measurement, not a bug.** The grader is exact on purpose
(`severity == "Medium"`) — an under-specified prompt gets an
under-specified extraction, and the pass rate is the honest size of that
problem. Run it again and you'll see a different split (25%, 50%, ...);
that's the real variance in the prompt, not flakiness in the harness.
- **The two failure reasons are two different bugs.** Run 1's `severity`
key doesn't exist at all — the model named it `reporter_name`/
`reporter_email` and dropped severity from that response's schema
entirely. Run 4 has the key but the wrong case (`"medium"` vs
`"Medium"`). A prompt fix for one won't fix the other — you'd need to
pin both the field name and an enum of allowed values.
- **`measure` exits 0 here even though the scenario "failed" 50% of the
time.** `echo $?` above prints `0` — a measurement has no pass/fail, so a
49-53% swing on rerun is not a CI signal by itself. Wire a *threshold*
check around the printed rate (or graduate to `compare` with an assertion)
once you have a bar to enforce.
- This baseline is now the number a proposed fix has to beat. Pin the
schema (name the fields, constrain `severity` to an enum) in a revised
`skill.md`, drop it into a `compare` scenario as the proposed arm against
this one as baseline, and see if the pass rate — not just one output —
actually moves.
7 changes: 7 additions & 0 deletions examples/03-measure-first/agent.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
name: ticket-intake
---

You are a support-ticket intake assistant. Follow the skill instructions
exactly and reply with nothing but the requested output — no preamble, no
markdown fences, no commentary.
Loading
Loading