Skip to content

feat(evals): add first agent tool-use evaluation suite - #8409

Open
sudoKrishna wants to merge 6 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals
Open

sudoKrishna wants to merge 6 commits into
simstudioai:mainfrom
sudoKrishna:feat/agent-tool-use-evals

Conversation

@sudoKrishna

@sudoKrishna sudoKrishna commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Adds a deterministic evaluation layer for agent tool use, then extends it to the full executor. Scenarios script the real OpenAI-compatible streaming tool loop and run real Start → Agent workflows through DAGExecutor; only the provider boundary is mocked. Runs in CI with no provider key, and against a live model on demand.

Closes #8408
Closes #8422

What changed

  • apps/sim/evals/agent-tool-use/ — scenario contract, scoring, JSON + Markdown report
    • scenarios.ts — 8 tool-loop scenarios
    • harness.ts — drives the real tool loop, pluggable model transport
    • executor-harness.ts — drives real DAGExecutor Start → Agent workflows
    • live.ts + agent-tool-use.live.test.ts — live model runs (DeepSeek first)
  • apps/sim/evals/README.md — how it works and how to add a case
  • apps/sim/package.json — test:evals and test:evals:live

Coverage

Harness Subsystem under test Cases
Tool loop providers/openai-compat/streaming-tool-loop.ts 8 (selection, planning, retrieval, recovery)
Executor DAGExecutor + AgentBlockHandler 4 (Start→Agent run, <start.field> resolution, block retry, model fallback)

How it works

The model is scripted (or live); the code under test is real. The tool loop suite stubs tools, so it isolates model decisions from API flakiness. The executor suite mocks only executeProviderRequest, so agent-block input wiring, variable resolution, and run/error handling are real. Both share one report.

How to run

cd apps/sim
bun run test:evals          # deterministic, writes test-results/evals/agent-tool-use.{json,md}

Live:

DEEPSEEK_API_KEY=... bun run test:evals:live

Opt-in only, never in CI. Each scenario runs EVAL_TRIALS times (default 3) and the report carries pass rates. EVAL_MIN_PASS_RATE turns a floor into a gate.

Test plan

  • bun run test:evals → 12/12 pass, one report with loop + executor rows
  • Negative check: intentionally broke an expectation → suite failed, then reverted
  • Live DeepSeek: first run 61% (brittle assertions caught), then 100% across 30 trials
  • bun run check:test-patterns passes
  • Eval files type-check against the real loop/executor signatures
  • Full bun run type-check — run in CI (local run OOMs in the author's environment)

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 09:51
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 29, 2026 9:51am UTC

Request Review

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation suite for the agent tool-use loop.

The evaluation suite has no identified runtime blocker, but its import paths must satisfy the repository requirement before merging.

Findings

  1. P2 Tool feedback goes unchecked ▶
  2. P2 Tool arguments are not checked ▶
  3. P2 Relative imports violate app requirement ▶

Summary

The PR adds eight deterministic scenarios that drive the production OpenAI-compatible streaming tool loop, score outcomes, and write JSON and Markdown reports.

  • The suite exercises dispatch and error accounting without a provider key.
  • Its scripted model does not inspect tool feedback, and scoring does not check dispatched arguments, limiting the regressions it can detect.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Scenario script] --> B[Scripted model turns]
  B --> C[Production streaming tool loop]
  C --> D[Stub tool results]
  D --> C
  C --> E[Scoring]
  E --> F[JSON and Markdown reports]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add agent tool-use evaluati..."

Comment on lines +110 to +111
return async () => {
const turn = scenario.script[turnIndex]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Tool feedback goes unchecked The scripted model emits its next hardcoded turn without reading the messages it receives. If the loop stops forwarding a tool result or error, the retrieval, dependent-planning, and recovery cases can still produce their expected answers and pass. Check the tool messages received on later turns so these cases can catch that regression.

Knowledge Base Used: Agent execution and sandbox tasks

Comment on lines +195 to +196
const actualSequence = toolCalls.map((call) => call.name)
const checks: EvalCheck[] = []

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Tool arguments are not checked The harness records the arguments sent to each tool, but scoring checks only tool names, counts, iterations, and final content. A call to the right tool with missing or incorrect arguments can therefore pass, limiting the suite’s ability to catch tool-use regressions. Add argument expectations to the relevant scenarios.


const logger = createLogger('AgentToolUseEval')

interface CapturedToolCall {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate app requirement This file imports ./types, and report.ts and scenarios.ts do the same. The Sim app’s import directive requires absolute imports and prohibits relative imports. Use the @/evals/agent-tool-use/types alias in all three files; this repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).

This branch was previously deployed

1 inactive (outdated) deployment
Preview — 86c79d78 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): run agent scenarios through the DAGExecutor feat(evals): add first agent tool-use evaluation suite

1 participant