feat(evals): add first agent tool-use evaluation suite - #8409
sudoKrishna wants to merge 6 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
| return async () => { | ||
| const turn = scenario.script[turnIndex] |
There was a problem hiding this comment.
Tool feedback goes unchecked The scripted model emits its next hardcoded turn without reading the messages it receives. If the loop stops forwarding a tool result or error, the retrieval, dependent-planning, and recovery cases can still produce their expected answers and pass. Check the tool messages received on later turns so these cases can catch that regression.
Knowledge Base Used: Agent execution and sandbox tasks
| const actualSequence = toolCalls.map((call) => call.name) | ||
| const checks: EvalCheck[] = [] |
There was a problem hiding this comment.
Tool arguments are not checked The harness records the arguments sent to each tool, but scoring checks only tool names, counts, iterations, and final content. A call to the right tool with missing or incorrect arguments can therefore pass, limiting the suite’s ability to catch tool-use regressions. Add argument expectations to the relevant scenarios.
|
|
||
| const logger = createLogger('AgentToolUseEval') | ||
|
|
||
| interface CapturedToolCall { |
There was a problem hiding this comment.
Relative imports violate app requirement This file imports
./types, and report.ts and scenarios.ts do the same. The Sim app’s import directive requires absolute imports and prohibits relative imports. Use the @/evals/agent-tool-use/types alias in all three files; this repository requirement must be satisfied before merging.
Context Used: Import patterns for the Sim application (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor, with only executeProviderRequest mocked at the provider boundary. This covers agent-block input wiring, variable resolution from Start outputs, and executor run/error handling, which the direct loop harness cannot see. - executor-harness.ts: workflow builder + runExecutorScenario - shares the scorer (scoreExpectations) and report with the loop suite - two scenarios: Start->Agent output, and <start.message> resolution - README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the Agent block has retry enabled, and the executor replays it. The run must complete with the second response. Verifies providerCalls === 2, and fails without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the Agent block has a fallback model, and the handler serves the answer from gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails without the fallback row (checked locally: got gpt-4o, run errored).
Summary
Adds a deterministic evaluation layer for agent tool use, then extends it to the full executor. Scenarios script the real OpenAI-compatible streaming tool loop and run real Start → Agent workflows through
DAGExecutor; only the provider boundary is mocked. Runs in CI with no provider key, and against a live model on demand.Closes #8408
Closes #8422
What changed
apps/sim/evals/agent-tool-use/— scenario contract, scoring, JSON + Markdown reportscenarios.ts— 8 tool-loop scenariosharness.ts— drives the real tool loop, pluggable model transportexecutor-harness.ts— drives realDAGExecutorStart → Agent workflowslive.ts+agent-tool-use.live.test.ts— live model runs (DeepSeek first)apps/sim/evals/README.md— how it works and how to add a caseapps/sim/package.json—test:evalsandtest:evals:liveCoverage
providers/openai-compat/streaming-tool-loop.tsDAGExecutor+AgentBlockHandler<start.field>resolution, block retry, model fallback)How it works
The model is scripted (or live); the code under test is real. The tool loop suite stubs tools, so it isolates model decisions from API flakiness. The executor suite mocks only
executeProviderRequest, so agent-block input wiring, variable resolution, and run/error handling are real. Both share one report.How to run
Live:
Opt-in only, never in CI. Each scenario runs
EVAL_TRIALStimes (default 3) and the report carries pass rates.EVAL_MIN_PASS_RATEturns a floor into a gate.Test plan
bun run test:evals→ 12/12 pass, one report with loop + executor rowsbun run check:test-patternspassesbun run type-check— run in CI (local run OOMs in the author's environment)