Skip to content

feat(bench): add toolcall_formats benchmark - #133

Merged
abhinav-pola merged 3 commits into
mainfrom
devin/1790984183-toolcall-formats
Oct 4, 2026
Merged

abhinav-pola merged 3 commits into
mainfrom
devin/1790984183-toolcall-formats

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR

Ports the internal toolcall_formats serving-stack benchmark from openrouter-web into the public harness, including the adversarial cases and scorer changes from OpenRouterTeam/openrouter-web#48802 (that PR doesn't need to merge).

What changed?

  • src/benchmarks/toolcall-formats/: cases.ts, json-schema.ts, scorer.ts (ported; monorepo libraries replaced with Either and internal/guards), benchmark.ts (dataset, solver, and defineSingleTurnBenchmark), colocated tests, and a README.
  • 27 scenarios × 5 optional-field encodings (omitted, anyof_null, anyof_null_strict, type_array_null, type_array_null_strict) = 135 samples. Sample ids are stable: toolcall_formats-<scenario>-<variant>.
  • 6 of the scenarios are adversarial: {read,bash,ticket}_omit_requested and {read,bash,ticket}_wrong_type_requested.
  • Solver request:
    • one system message, then the user prompt;
    • tools: entry.tools;
    • temperature: 0, fixed through FixedTemperatureBenchmarkBaseSchema;
    • the usual inference and routing overrides (providerOnly, endpointId, allowFallbacks, and the rest) are forwarded.
  • strict is always sent explicitly as strict: variant.strict, so non-strict variants don't depend on how the API treats a missing flag.
  • Scoring checks, in this order:
    1. every call validates against the exact schema that was sent;
    2. the call count is right;
    3. the values the prompt states match (partial object match). Keys the prompt leaves unset only have to pass the schema.
  • Failure classes: no_tool_call, unknown_tool, invalid_json, schema_violation, call_count, wrong_value.
  • Wiring: config schema and union/options entry, TOOLCALL_FORMATS_META, registry entry, and the CLI fixed-temperature case.

Why?

Debugging GLM-5.3 on Morph showed that Morph doesn't constrain tool arguments to the schema. When the model drops required keys or sends "50" into number slots, Morph passes those calls through, and this benchmark catches that. It's moving to the public repo per the Slack thread.

How to test

bun test src/benchmarks/toolcall-formats
bun run bench --benchmark toolcall_formats --model z-ai/glm-5.3 --reasoning-effort low --concurrency 8 \
  --solver-config '{"providerOnly":["morph"],"allowFallbacks":false}'

Benchmark impact

New benchmark. I ran the full 135 samples once on this branch (z-ai/glm-5.3, low effort, pinned with allowFallbacks: false):

  • Morph: 113/135. All 22 failures are schema_violation:
    • omit cases: required keys missing, including on strict variants;
    • wrong-type cases: limit: "50", timeout: "30", blocking: "true".
  • Wafer: 115/135. All 20 failures are schema_violation:
    • mostly missing keys on the non-strict nullable variants;
    • "null" strings in boolean|null and enum slots on the ticket and weather cases.

These numbers come through the Responses API path, so they differ from the internal chat-completions run (Wafer 124, Morph 112 there). Wafer misses more non-strict nullable variants here.

Reviewer focus

  • Non-strict variants now send strict: false. The internal version left the flag out.
  • The solver looks up each sample's tools by sample.id.

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented
  • Benchmark changes document dataset provenance and licensing
  • No credentials, private results, or restricted dataset contents are included
  • Documentation is updated where needed

Searched existing PRs ("toolcall", "tool call formats"); none matched.

Link to Devin session: https://openrouter.devinenterprise.com/sessions/257c4c8842f94f188a26b5a4e4772390
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/257c4c8842f94f188a26b5a4e4772390?variant=devin
Requested by: @johnpyp


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Original prompt from John

SYSTEM:
<latest_message>
John Paul Penaloza (U0BGKKT3AES) [ts=1790920778.808839]: @Devin check why specifically GLM-5.3 is failing tool calls on Morph? you can try running the tool format internal bench locally
</latest_message>

=== BEGIN THREAD HISTORY (in #agents-ecosystem) ===
John Paul Penaloza (U0BGKKT3AES) [ts=1790920778.808839]: @Devin check why specifically GLM-5.3 is failing tool calls on Morph? you can try running the tool format internal bench locally
=== END THREAD HISTORY ===
Channel ID: C0BU53A7VEH
Thread URL: https://openrouter.slack.com/archives/C0BU53A7VEH/p1790920778808839?thread_ts=1790920778.808839&amp;cid=C0BU53A7VEH

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Devin Review

@abhinav-pola
abhinav-pola merged commit d22d57b into main Oct 4, 2026
5 checks passed
@abhinav-pola
abhinav-pola deleted the devin/1790984183-toolcall-formats branch October 4, 2026 19:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant