Skip to content

feat(evals): add agent context eval suite - #8428

Open
sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-agent-context
Open

sudoKrishna wants to merge 7 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-agent-context

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Evaluate what the agent actually sends the model from conversation memory. The
suite drives the real Agent block through the DAGExecutor with memory on,
stubs the memory read per conversation id, and asserts the assembled provider
request: prior history, then the new user prompt, system prompt preserved, and
the right conversation id.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8427

What changed

  • agent-context/scenarios.ts — context cases (prior turns, conversation
    isolation)
  • agent-tool-use/executor-harness.ts — a agent.memory seam plus checks:
    memory-history-in-request, memory-before-user-prompt,
    system-prompt-in-request, conversation-id
  • test:evals:context script; README documents the suite

How to run

cd apps/sim
bun run test:evals:context   # writes test-results/evals/agent-context.{json,md}

The suite is also collected by the normal bun run test.

Test plan

  • bun run test:evals:context → 2/2, report written with the assembly checks
  • Negative check: forced the memory read to return the wrong conversation →
    memory-history-in-request failed, then reverted
  • bun run test:evals still passes 12/12
  • bun run check:test-patterns passes
  • Changed files type-check against the real executor/handler signatures
  • Full bun run type-check — run in CI

Follow-up

Memory windowing (sliding window size/tokens) and the retrieval tool are covered
by their unit tests; the next eval layer is an end-to-end memory-window case and
the subagent/orchestration suites.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Drive the Agent block through the executor with conversation memory on. The
memory read is stubbed per conversation id, so the provider request shows what
the handler assembled: prior history, then the new prompt, system prompt
preserved, correct conversation id. A wrong id surfaces as missing history and
fails (checked locally).

- agent-context/scenarios.ts: two context scenarios
- executor-harness.ts: memory seam + assembly/isolation checks
- test:evals:context script; README documents the suite
@sudoKrishna
sudoKrishna requested a review from a team as a code owner September 29, 2026 20:04
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation test suites for agent tool handling.

The PR has no identified production failure, but the explicit import requirement should be met before merging; the eval checks also need tightening to support their coverage claims.

Findings

  1. P2 History checks miss wrong roles ▶
  2. P2 Scripted turns ignore tool feedback ▶
  3. P2 Extra live calls appear successful ▶
  4. P2 Relative imports violate convention ▶
  5. P2 Context reports have wrong label ▶

Summary

The PR adds scripted tool-loop and executor evals, a live-model option, and an executor-based conversation-context suite.

  • Context checks do not verify history roles or full ordering.
  • Tool-feedback coverage and live-trial result attribution need stronger assertions.
  • Context reports use the tool-use suite label.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  S[Context scenario] --> E[DAGExecutor]
  E --> M[Stubbed memory read]
  M --> A[Agent handler]
  A --> P[Captured provider request]
  P --> C[History and prompt checks]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add agent context eval suit..."

Comment on lines +254 to +255
const missingHistory = memory.history.filter(
(message) => !contents.some((content) => content.includes(message.content))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 History checks miss wrong roles The check finds each history string anywhere in the request, while the ordering check compares only the last matching positions. Swapping the prior user and assistant messages, reversing their order, or merging them into one message would still pass. The suite could therefore accept a provider request with the wrong conversation history. Assert the expected roles and message order.

Comment on lines +124 to +127
function createScriptedCompletion(scenario: AgentToolUseScenario): OpenAICompatCreateCompletion {
let turnIndex = 0
return async () => {
const turn = scenario.script[turnIndex]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Scripted turns ignore tool feedback The scripted completion returns the next preset turn without reading the request. If the loop stops sending a tool result or error to the model, the retrieval and recovery cases still produce their prewritten retries and answers and can pass. Inspect the next request's tool messages so these cases verify the feedback they claim to cover.

toolsMockFns.mockExecuteTool.mockImplementation(
async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => {
const startedAt = Date.now()
const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Extra live calls appear successful If the live model calls a tool more often than the scripted transcript does, its queued results run out and this fallback reports an empty result as a success. That inflates successful-call counts and gives the model fabricated feedback, making live results less reliable. Handle live calls independently of scripted queue lengths.

import { DAGExecutor } from '@/executor/execution/executor'
import { memoryService } from '@/executor/handlers/agent/memory'
import type { SerializedBlock, SerializedWorkflow } from '@/serializer/types'
import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate convention This new file imports sibling modules through relative paths, contrary to the Sim app's directive to use absolute imports. The same pattern appears in harness.ts and report.ts. Use the @/evals/agent-tool-use/... aliases throughout; this repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

providersUtilsMockFns.mockGetProviderFromModel.mockReturnValue('mock-provider')
})

afterAll(() => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Context reports have wrong label The context suite uses a report writer that always labels the JSON suite and Markdown heading as agent-tool-use. As a result, the context-only report identifies itself as the wrong suite, misleading readers and consumers that group reports by suite.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evals): add agent context eval suite

1 participant