Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,10 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval-

### Changed

- README and example guides include verified execution commands, public imports, and explicit limits on fixture results and release evidence.
The existing-agent quickstart is offline; its guide shows how to meter paid calls with the maintained transport and receipt helpers.
- The single-optimizer example accepts the same worker `PRICE_*` settings as the method-comparison example.
Both use one parser to validate endpoint rates before execution.
- **Breaking:** Root `Scenario`, `JudgeScore`, and `GateDecision` now match `/contract`.
Product workflows use `ProductScenario`, `DimensionJudgeScore`, and `HeldOutGateDecision`.
- **Breaking:** Current canonical envelopes and algorithm identifiers are required for seals, attestations, and profile identities.
Expand Down Expand Up @@ -42,6 +46,8 @@ All notable changes to `@tangle-network/agent-eval` and its sibling `agent-eval-

### Fixed

- Failure-cluster shares count all affected failed runs independently of the five displayed examples.
Multiple findings in the same cluster count once per run.
- Outcome queries select the latest finite requested metric instead of an unrelated latest observation.
Outcome-store corruption and unavailable evidence remain visible failures.
- Calibration preserves clipped observations, measures constant predictors, and honors the requested bin count.
Expand Down
51 changes: 28 additions & 23 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,14 +13,16 @@ It records outputs, failures, costs, and evidence for each comparison.

## Install

Use Node.js 20.19 or newer.

```sh
pnpm add @tangle-network/agent-eval
```

## Quickstart

This complete example runs offline.
Replace the agent and judge with your product functions when it works.
Save it as `eval.mts`.

```ts
import { defineAgentEval } from '@tangle-network/agent-eval/contract'
Expand Down Expand Up @@ -50,11 +52,25 @@ const evalKit = defineAgentEval<SupportCase, string>({
expectUsage: 'off',
})

console.log((await evalKit.evaluate()).aggregates.byJudge)
console.log(
(await evalKit.evaluate({ surface: 'Answer politely and cite the ticket id.' })).aggregates
.byJudge,
)
const baseline = await evalKit.evaluate()
const candidate = await evalKit.evaluate({
surface: 'Answer politely and cite the ticket id.',
})

console.log('baseline:', baseline.aggregates.byJudge['ticket-id']?.mean)
console.log('candidate:', candidate.aggregates.byJudge['ticket-id']?.mean)
```

Run it with a TypeScript runner:

```sh
pnpm add --save-dev tsx
pnpm exec tsx eval.mts
```

```text
baseline: 0
candidate: 1
```

The baseline scores `0`; the candidate scores `1` on all three cases.
Expand All @@ -66,8 +82,9 @@ A **surface** is the prompt, skill, or configuration being changed.
A **judge** scores the agent's result.

`expectUsage: 'off'` applies because this example makes no paid calls.
Keep the default, `'assert'`, for model calls so missing cost receipts fail visibly.
Set `expectUsage: 'assert'` for paid agents so missing dispatch receipts become execution failures.
The [runnable example](./examples/evaluate-a-change/) uses the same evaluation.
The [existing-agent example](./examples/foreign-agent-quickstart/) shows how to connect your agent and record model usage.

## Choose a workflow

Expand Down Expand Up @@ -164,29 +181,17 @@ The [benchmark-book review](./docs/design/mlbenchmarks-book-review.md) records t

```sh
pnpm install
pnpm build
pnpm typecheck
pnpm typecheck:examples
pnpm typecheck:scripts
pnpm lint
pnpm test
pnpm build
pnpm verify:package
```

Python compatibility tests use the locked dependencies:

```sh
cd clients/python
uv sync --frozen --extra dev --group gepa-release
AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \
uv run --frozen --extra dev --group gepa-release \
pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py

uv sync --frozen --extra dev --group skillopt-source --group gepa-source
uv run --frozen pytest

uv sync --frozen --extra dev --extra dspy
uv run --frozen pytest tests/test_dspy_metric.py
```
Build before checking examples because they resolve the package's generated declarations.
The [Python development guide](./clients/python/README.md#development) gives the locked commands for each optimizer environment.

## License

Expand Down
81 changes: 40 additions & 41 deletions clients/python/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ The Node package owns rubric execution, model calls, and scoring.

## Install

Python 3.10 or newer and Node.js 20 or newer are required.
Python 3.10 or newer and Node.js 20.19 or newer are required.
Install matching package versions:

```sh
Expand All @@ -19,15 +19,17 @@ Configure an OpenAI-compatible model endpoint for judge calls:
```sh
export AGENT_EVAL_LLM_BASE_URL=https://api.openai.com/v1
export AGENT_EVAL_LLM_API_KEY="$YOUR_API_KEY"
export AGENT_EVAL_LLM_MODEL=gpt-4.1-mini
export AGENT_EVAL_LLM_MODEL="$YOUR_MODEL_ID"
```

`OPENAI_BASE_URL`, `OPENAI_API_KEY`, and `OPENAI_MODEL` are also accepted.
The endpoint receives the content, rubric, and context passed to `client.judge()`.

The `agent-eval` binary is the only part of the package that reads a provider credential.
It is a server process, so it configures its own endpoint the way every server does; the TypeScript library holds no key and executes no paid model.
Without both a base URL and a key, `judge()` fails with `llm_not_configured` instead of calling an unintended endpoint.
The CLI resolves provider settings from environment variables.
An `OPENAI_API_KEY` or `TANGLE_API_KEY` can select that provider's default endpoint.
Set `AGENT_EVAL_LLM_BASE_URL` and `AGENT_EVAL_LLM_API_KEY` to make the route explicit.
TypeScript callers supply their transport, endpoint, and credentials directly.
Without a resolved endpoint and credential, `judge()` fails with `llm_not_configured`.

## Judge Content

Expand Down Expand Up @@ -102,11 +104,11 @@ for rubric in Client().list_rubrics().rubrics:
## Client Options

```python
Client(
base_url: str | None = None,
cli_path: str | None = None,
transport: "auto" | "http" | "subprocess" = "auto",
timeout_s: float = 200.0,
client = Client(
base_url="http://127.0.0.1:5005",
cli_path="agent-eval",
transport="auto",
timeout_s=200.0,
)
```

Expand Down Expand Up @@ -178,10 +180,11 @@ python -m pip install \
"wandb"
```

From an Agent Eval source checkout:
From `clients/python` in an Agent Eval source checkout, choose the required environment:

```sh
uv sync --frozen --group gepa-release
# For source-only engines or compositions, use this instead:
uv sync --frozen --group gepa-source
```

Expand Down Expand Up @@ -214,7 +217,7 @@ python -m pip install \
"skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90"
```

From an Agent Eval source checkout, install the locked package with:
From `clients/python` in an Agent Eval source checkout, install the locked package with:

```sh
uv sync --frozen --group skillopt-source
Expand Down Expand Up @@ -248,6 +251,8 @@ python -m pip install "agent-eval-rpc[dspy]"
```

```python
import os

import dspy

from agent_eval_rpc import DspyJudgeMetric
Expand All @@ -257,14 +262,15 @@ metric = DspyJudgeMetric(rubric_name="answer-quality")

gepa = dspy.GEPA(
metric=metric.feedback,
reflection_lm=dspy.LM("openai/gpt-4.1-mini"),
reflection_lm=dspy.LM(os.environ["DSPY_REFLECTION_MODEL"]),
max_metric_calls=100,
)
optimized = gepa.compile(program, trainset=train, valset=selection)

mipro = dspy.MIPROv2(metric=metric, auto="light")
```

Set `DSPY_REFLECTION_MODEL` to your configured DSPy model identifier, including its provider prefix.
Use `metric.feedback` for `dspy.GEPA`.
It returns `dspy.Prediction(score=..., feedback=...)` with dimension scores, failure modes, wins, and rationale.
Use the metric object directly for MIPROv2, SIMBA, bootstrap, and evaluation APIs that expect a number.
Expand Down Expand Up @@ -301,17 +307,12 @@ import { analyzeTraces } from '@tangle-network/agent-eval/traces'

type ModelOwner = Pick<
DspyRlmTraceEngineOptions,
'call' | 'callRef' | 'recordExecution'
'call' | 'callRef' | 'recordExecution' | 'model' | 'pricing'
>

export async function analyzeRun(modelOwner: ModelOwner) {
const engine = createDspyRlmTraceEngine({
...modelOwner,
model: 'deepseek-v4-flash',
pricing: {
inputUsdPerMillion: 3,
outputUsdPerMillion: 15,
},
runner: { command: '.venv/bin/python' },
})

Expand All @@ -324,21 +325,9 @@ export async function analyzeRun(modelOwner: ModelOwner) {

See [Trace Analysis](../../docs/trace-analysis.md) for custom definitions, limits, result fields, and the public quality benchmark.

DSPy 3.2.1 pins GEPA 0.0.27.
The general Optimize Anything bridge uses GEPA 0.1.4, so repository checks install them in separate environments:

```sh
uv sync --frozen --extra dev --group gepa-release
AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \
uv run --frozen --extra dev --group gepa-release \
pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py

uv sync --frozen --extra dev --group skillopt-source --group gepa-source
uv run --frozen pytest

uv sync --frozen --extra dev --extra dspy
uv run --frozen pytest tests/test_dspy_metric.py
```
The caller supplies the model identifier and endpoint rates with its execution callbacks.
DSPy and the Optimize Anything bridge require different GEPA versions.
The [development commands](#development) select each locked environment separately.

The bridge records the installed upstream package version and source revision with each run.
SkillOpt and a direct GEPA engine can restore official state only when the package revision, settings, starting candidate, described data, evaluation ID, and seed match.
Expand Down Expand Up @@ -372,19 +361,29 @@ print(version.version, version.wire_version)

## Development

From the repository root, build Node before running cross-language tests:

```sh
pnpm install --frozen-lockfile
pnpm build
cd clients/python
pip install -e ".[dev]"
pytest
```

Run the cross-language tests after building the Node package:
Run each compatibility suite with its locked dependencies:

```sh
cd ../..
pnpm build
cd clients/python
pytest
uv sync --frozen --extra dev --group gepa-release
AGENT_EVAL_EXPECT_GEPA_RELEASE=1 \
uv run --frozen --extra dev --group gepa-release \
pytest tests/test_gepa_release_compatibility.py tests/test_gepa_bridge.py

uv sync --frozen --extra dev --group skillopt-source --group gepa-source
uv run --frozen --extra dev --group skillopt-source --group gepa-source pytest

uv sync --frozen --extra dev --extra dspy
uv run --frozen --extra dev --extra dspy pytest tests/test_dspy_metric.py
```

Keep the same extras and groups on `uv sync` and `uv run`.
Each `uv sync` switches the local environment to that optimizer's required dependency set.
The runnable Python example is [`examples/judge_anti_slop.py`](./examples/judge_anti_slop.py).
Loading