Skip to content

feat(evals): record and replay live model transcripts - #8425

Open
sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-record-replay
Open

sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-record-replay

Conversation

@sudoKrishna

@sudoKrishna sudoKrishna commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Record a live model run once, replay it forever through the real tool loop with
no key and no network. This turns today's nondeterministic live finding into a
durable, deterministic regression test, and unblocks suites (context, subagents)
that need many real model turns.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8424

What changed

  • replay.ts — createRecordingCompletion (wraps a live completion, captures
    each call's chunks) and createReplayCompletion (feeds them back), plus
    fixture read/write/list helpers
  • agent-tool-use.replay.test.ts — replays every committed fixture through
    createOpenAICompatStreamingToolLoopStream with the same scoring; skips until
    a fixture exists
  • replay.test.ts — key-free unit coverage of chunk round-trip and fixture I/O
  • agent-tool-use.live.test.ts — EVAL_RECORD=1 records the first trial
  • test:evals:record script; README documents recording, replaying, refreshing

How to record and replay

cd apps/sim
DEEPSEEK_API_KEY=... bun run test:evals:record   # writes fixtures/<scenario>.json
bun run test:evals                               # replays them, no key

Fixtures hold the raw streamed chunks per model call, so a diff shows a behavior
change exactly as the model produced it. They are committed and reviewed like
snapshots.

Test plan

  • bun run test:evals → 16/16 (12 harness + 4 replay primitives), replay
    suite skips while no fixtures exist
  • Replay path validated end to end with a crafted fixture (single-tool-lookup
    replayed through the real loop and passed), then the synthetic fixture was
    removed
  • replay.ts type-checks against the real loop signature
  • bun run check:test-patterns passes
  • Record real fixtures with DeepSeek and commit them (follow-up)
  • Full bun run type-check — run in CI

Follow-up

Once real fixtures are committed, the context/memory and subagent suites can
replay real multi-turn transcripts instead of scripting every turn.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Record a live run once, replay it forever through the real tool loop with no
key. EVAL_RECORD=1 wraps the live completion and writes each model call's
streamed chunks to fixtures/<scenario>.json; agent-tool-use.replay.test.ts
feeds them back through createOpenAICompatStreamingToolLoopStream and scores
them with the same checks.

- replay.ts: recording/replay completions + fixture I/O
- replay.test.ts: chunk round-trip and fixture I/O (key-free)
- live test records on EVAL_RECORD=1; test:evals:record script
- replay suite skips until a fixture exists; README documents the loop
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 19:45
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation framework for the agent tool-use loop.

The PR should not merge until recording rejects incomplete first trials and the explicit import requirement is satisfied.

Findings

  1. P1 Failed trials produce fixtures ▶
  2. P2 Unmatched fixtures disappear silently ▶
  3. P2 Stale fixtures can pass ▶
  4. P2 Relative imports violate app rules ▶

Summary

The PR adds scripted and live agent-tool-use evals, a Start-to-Agent executor harness, and raw model-stream recording and offline replay. Recording needs to reject incomplete first trials before their fixtures become replay inputs.

  • Replay should detect unmatched or stale fixtures rather than silently losing coverage.
  • New local imports need to follow the app’s absolute-import requirement.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Live[Live completion] --> Recorder[Record streamed turns]
  Recorder --> Loop[Production tool loop]
  Loop --> Stub[Stubbed tool results]
  Recorder --> Fixture[Scenario fixture]
  Fixture --> Replay[Replay completion]
  Replay --> Loop
  Loop --> Checks[Scenario checks]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): record and replay live mode..."

Comment on lines +82 to +89
if (RECORD && trial === 0) {
fixturesToWrite.push({
scenarioId: scenario.id,
model: MODEL,
recordedAt: new Date().toISOString(),
turns: recordedTurns,
})
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Failed trials produce fixtures When the first live trial fails during a model stream, the recorder keeps only completed turns, but this code still writes them as a fixture. Because the default pass-rate floor is zero, the record command can succeed while producing an empty or truncated fixture that fails the offline replay suite. Write a fixture only after a complete, passing trial.

Comment on lines +33 to +36
for (const fixture of listReplayFixtures(FIXTURES_DIR)) {
const scenario = AGENT_TOOL_USE_SCENARIOS.find((candidate) => candidate.id === fixture.scenarioId)
if (scenario) replayCases.push({ id: fixture.scenarioId, fixture, scenario })
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unmatched fixtures disappear silently If a scenario is renamed or removed, its committed fixture is discarded here without a warning. When no matching fixtures remain, the replay suite skips entirely and CI stays green despite running no recorded cases. Fail on unmatched fixture IDs so lost replay coverage is visible.

Comment on lines +72 to +75
let index = 0
return async () => {
const turn = turns[index]
index += 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Stale fixtures can pass Replay returns chunks by turn number without checking the current model request. If a scenario’s prompt or available tools change, an old fixture can still satisfy the output checks, leaving the replay green without testing the changed request. Record and check which request each fixture belongs to.

Comment on lines +8 to +14
import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness'
import type {
AgentToolUseExpectations,
AgentToolUseResult,
EvalCategory,
EvalToolInvocation,
} from './types'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate app rules These neighboring-module imports use relative paths, contrary to the Sim app directive: “Always use absolute imports. Never use relative imports.” The same pattern appears in harness.ts, report.ts, and scenarios.ts. Use the @/evals/agent-tool-use/... alias in these files; this repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): record and replay live model transcripts

1 participant