Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 34 additions & 31 deletions .claude/skills/eval-architect/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,44 +1,47 @@
---
name: eval-architect
description: Build a measurement that scores an agent's REAL deliverable — not a proxy — for a product you've never seen before. Use when scaffolding or repairing the eval an Improve loop optimizes against. Get this wrong and every downstream optimization perfects a fiction.
description: Build or repair evaluations that score the agent's actual deliverable through the production path, with independent controls and visible missing evidence.
---

# Eval Architect — measure the real deliverable
# Eval architect

You are building the measurement an improvement loop will optimize against. **The loop optimizes whatever you measure.** If you measure the wrong thing, the loop perfects the wrong thing — confidently, expensively, and invisibly. The measurement is the product. Everything else in the Improve stack is downstream of getting this right.
Build the measurement around the product's required outcome and actual execution path.
Reuse maintained Eval contracts and existing product checks before creating another scorer.

This skill is held by the agent that *builds* the eval (often a delegated coding agent). Pair it with `measurement-validation` (the gate that proves your eval is sound before anyone spends money on it).
## Locate the deliverable

## The cardinal question
1. Inspect real runs to find the output channel: replies, validated tool calls, persisted artifacts, application state, or rendered UI.
Score the channel that carries the required outcome.
2. Define the completion boundary for the task.
For work that accumulates across turns, evaluate the completed artifact and retain intermediate evidence needed to explain failures.
3. Trace every consumer of the score, including completion checks, optimization selection, and release decisions.
When the output channel changes, update every affected consumer.

**Where does this agent's deliverable actually land?** Prose in the reply? Validated tool calls? Persisted artifacts (vault docs, DB rows)? A PR? A rendered UI? Find out by inspecting *real runs* — never by assuming it's the chat text.
## Build the checks

> Worked failure (legal-agent, this is why the skill exists): the eval scored the assistant's chat prose. A tool-migration moved the deliverable into `submit_proposal` calls + vault docs, leaving the prose empty. Every scorer reading prose silently collapsed to ~0. The loop would have optimized an empty string. The deliverable had *moved* and the measurement didn't follow it.
1. Map each requirement to observable evidence and an explicit failure condition.
Use answer keys when available; otherwise use independently justified constraints, executable checks, or calibrated judgment.
Keep unsupported requirements and missing evidence visible.
2. Establish a simple baseline through the same entrypoint as the candidate.
Investigate surprising scores instead of assuming either the scorer or the agent caused them.
3. Separate training, candidate selection, and final comparison evidence where the improvement claim requires those partitions.
Preserve scenario identities and shared source-unit mappings across baseline and candidate runs.
4. Define critical failure checks separately from aggregate quality.
A favorable composite must not erase a failure that violates the product's requirements.
5. Preserve scorer identity, actual cost receipts, execution failures, and diagnostic artifacts.
Change `judgeVersion` when an ensemble scorer's configuration changes.

## Invariant (non-negotiable — violate these and the loop is a slot machine)
## Prove the measurement

1. **Score the produced artifact, not the conversation.** Locate the real output channel and score *that*.
2. **For accumulating-artifact agents, score the CONVERGED multi-shot artifact, not turn 1.** Most real agents build their deliverable over several turns. Define a convergence criterion (e.g. the artifact stops growing for N shots) and score the converged state.
3. **A held-out split exists and is never trained on.** No held-out → no honest gate → no trustworthy lift.
4. **Every requirement has gold the scorer matches against, from real records — never fabricated.** A requirement with no gold means there is nothing to verify; fail loud, do not pass-by-default. A fluent hallucination that produced nothing must score 0, not 0.9.
Run known positive and negative examples through the complete scoring path.
Perturb a real deliverable so required behavior improves or regresses, then check that the score detects each change.
Confirm that absent output and evaluator failure remain distinguishable from measured low quality.
Report case coverage, detectable failures, uncertainty, and any requirements the evaluation cannot assess.
A training gain without a final-comparison gain needs diagnosis; it does not identify the cause by itself.

## Judgment (figure this out per product — the agentic core)
## Then consider

- What *is* the deliverable here, and where does it persist? Read the runtime events / tool calls / storage, not the transcript.
- What is the convergence criterion for this agent's artifact? When has it stopped accumulating?
- What gold defines "correct" for each requirement, and where does it come from (real records, never invented figures)?
- Which dimensions matter, and what are their weights? What is the one dimension that, if it regresses, kills the deal regardless of the composite (safety, hallucination, the regulated invariant)?

## Self-test (prove the metric works before trusting it)

- **Baseline sanity:** run it. Is the score non-zero and plausible for a competent agent? A near-zero baseline usually means you're scoring the wrong channel, not that the agent is terrible.
- **The mutation test (the one that catches the empty-string bug):** hand-edit the produced artifact to be *obviously better* and *obviously worse*. Does the score move in the right direction and magnitude? A metric that doesn't move under obvious changes is measuring the wrong thing.
- **Audit EVERY scoring surface together.** Completion, quality, and the optimizer's own scorer all read *something*. When the deliverable's channel moves, all of them that read the old channel silently zero. (Session: completion + quality were fixed; the optimizer's own scorer was missed and only found by tracing. Three surfaces — enumerate them, don't assume one.)

## Evolves-by

When a later optimization shows lift on *training* but none on *held-out*, your eval was overfittable or gameable — add the gap it missed as a new judgment rule. The architect's judgment surface is itself optimized by the meta-eval *"did evals built this way yield real held-out lift, no critical regression?"* See `skill-evolution`.

## Fleet as dogfood

legal / tax / gtm / creative / insurance each put their deliverable in a *different* channel — filings, forms, published copy, rendered artifacts, routed proposals. The skill is general precisely because it forces you to *locate* the channel for the product in front of you rather than hardcode "the reply text."
| Condition | Skill |
|---|---|
| The evaluation path executes and needs calibration or comparison checks | `measurement-validation` with the baseline and control results |
| The validated measurement supports a candidate search | `surface-evolution` with the target surface, evidence, and resource limits |
87 changes: 46 additions & 41 deletions .claude/skills/improve-conductor/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,45 +1,50 @@
---
name: improve-conductor
description: The user-facing controller for the Improve button. Decide whether a request is improvable, translate a dollar budget into a run, read the verdict honestly, and promote or refuse with a reason. Never promise a lift you cannot measure.
description: Drive a requested optimization from its target and resource limits through measured candidate selection, evidence review, and the product's promotion policy.
---

# Improve Conductor — own the user's trust

You are the agent the user talks to when they click **Improve**. You do not build the eval or run the loop yourself — you decide *whether* to, *how much* to spend, and *what to tell the user about the result*. The product you are protecting is **trust**, not lift. You would rather say "I could not prove an improvement — here is what another $X buys" than ship noise and call it a win.

You delegate the building to an agent holding `eval-architect` + `surface-evolution`; you both share `measurement-validation` as the honesty contract.

## Invariant (non-negotiable)

1. **Never promise or report a lift you cannot measure with valid paired evidence.** Surface the honest verdict: `ship` / `hold` / `need-more-data` / `invalid`. A paired, significant lift is still not shippable until a footprint-matched placebo shows the gain comes from the CONTENT, not from added prompt/mount footprint — the substrate's `neutralizationGate` / `runImprovementLoop({ neutralize })` (`@tangle-network/agent-eval/campaign`). A lift that a neutralized twin reproduces is footprint, not improvement — refuse it. Route this check through `measurement-validation`. "Invalid" (incomplete or unpaired evidence) is a first-class outcome — say it plainly, never paper over it with a survivor-mean number.
2. **Refuse below the data threshold, and say why** — "I have N real outcomes; I won't optimize below M. Here's how to get to M." A refusal with a reason builds more trust than a fabricated win.
3. **Route correctly.** Improvable by surface-tuning → dispatch `surface-evolution`. Needs a new capability or architecture → escalate and say so; don't pretend tuning will fix a structural gap.
4. **No optimization spend before the target is confirmed and the measurement is real.** If there is no improvement infrastructure yet, you do NOT improvise a metric and start spending — you dispatch `eval-bootstrap` to BUILD a validated, externally-grounded harness first. The gate between "build the apparatus" and "spend optimizing" is yours to hold.

## Cold start — no infrastructure yet

The most dangerous request is "improve this" for a product with no eval. The wrong move is to invent a metric and start a loop — you'll perfect a proxy and report a fake win. The right move is a strict two-step you orchestrate:

1. **Frame + build (no spend):** confirm with the user *what "better" means* — the thing they'd reject a draft over, tied to a product-value claim — then dispatch `eval-bootstrap` (often a delegated agent-runtime build loop) to construct a harness grounded in **external truth**, exiting only when `measurement-validation` passes. The improver is a *builder* here, not a tuner.
2. **Then optimize (spend):** only once the harness is validated, dispatch `surface-evolution` against it.

Never let the user believe step 2 happened when only a toy of step 1 did. If you can't yet build a real measurement (no grounding, target unclear), say so and ask for what you need — that's the honest move, not a loop against an invented number.

## Judgment (figure this out per request)

- Is this a surface-tuning problem or an architectural one? (If the agent literally cannot do the task, no prompt edit fixes it.)
- Translate the user's dollars into a run: more spend = wider candidate search + more reps = tighter CI + higher chance of clearing the gate. $0.20 ≈ one quick generation on a couple scenarios; $50 ≈ multi-generation search with a held-out gate that can actually reach significance.
- When to stop: threshold met, plateaued, or budget exhausted — and report which.

## Self-test

- **Before spending,** you can state out loud: the metric, its variance, the threshold, the held-out set, and what this budget buys. If you can't, you're not ready to charge for the click.
- **After,** you report the gated lift with its CI and the decision's *reason*. If the run came back `invalid` (a cell errored, evidence unpaired), you tell the user that and offer the re-run — you do not quote the broken number.

## Evolves-by

User accept/reject of promotions; spend→lift efficiency; the rate of `invalid` runs. A rising invalid rate is a signal the measurement or the infra needs hardening — route it back to `measurement-validation` / `eval-architect`, don't absorb it silently. See `skill-evolution`.

## Why this is calibrated, not timid

A naive Improve button maximizes the displayed number and tells the user "improved +47%". The disciplined one, faced with the same +47, checks the evidence, finds it unpaired, and says "I found a promising candidate but can't yet prove it beats baseline — $X more will confirm it." The second one is the one people pay for twice.
# Improve conductor

Own the requested outcome, actual spend, and the decision supported by the evidence.
Distinguish building an evaluation, searching for candidates, and proving an improvement.

## Establish the work

1. Recover the target, required outcome, existing evaluation, and authorization from the user's request and product context.
Ask only for consequential information that the available evidence cannot resolve.
2. Inspect the current baseline and the system's failure cases.
Choose a surface change, capability change, or architecture experiment according to the observed limitation.
3. Check that the evaluation can detect the required behavior through the production entrypoint.
If it cannot, build that measurement before making improvement claims.
Record measurement work as measurement work, including its actual cost.
4. Set resource limits and a stopping rule for the selected experiment.
Estimate cost from the actual execution path and retain uncertainty about additional calls, retries, and candidate evaluations.
More spend does not guarantee a useful candidate or a conclusive result.

## Run and decide

1. Use the maintained optimization method and execution path already available to the product.
Preserve training, selection, and final-comparison boundaries along with the registered observation units.
2. Retain candidate artifacts, scorer identity, paired raw evidence, failures, and complete attempt costs.
Missing usage remains an explicit accounting gap.
3. Read the shared deciding statistic and the producer's actual verdict.
Distinguish missing evidence, a measured failure, an inconclusive comparison, and a result that meets the release policy.
4. Investigate surprising gains and null results with controls that isolate the suspected mechanism.
Use a footprint control when the claim concerns content versus added context; it is not a universal release prerequisite.
5. Promote only through the product's authorized decision path after its required checks pass.
A promising exploratory result can justify another scoped experiment without establishing an improvement.

## Report

State what changed, the baseline comparison, deciding interval, independent-unit count, actual costs, and the verdict's reasons.
Explain whether the run stopped because it reached its registered criterion, exhausted its budget, or could not capture valid evidence.
A proposed follow-up must state which uncertainty it could resolve; extra budget alone does not promise confirmation.

## Then consider

| Condition | Skill |
|---|---|
| The product has no usable evaluation path | `eval-bootstrap` with the observed deliverable and missing checks |
| A scorer or output-channel defect prevents assessment | `eval-architect` with the failed case and execution evidence |
| Existing measurements need calibration or comparison review | `measurement-validation` with the baseline and retained results |
| The supported next experiment changes an existing surface | `surface-evolution` with the target, acceptance criteria, and resource limits |
Loading