You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ports the internal toolcall_formats serving-stack benchmark from openrouter-web into the public harness, including the adversarial cases and scorer changes from OpenRouterTeam/openrouter-web#48802 (that PR doesn't need to merge).
What changed?
src/benchmarks/toolcall-formats/: cases.ts, json-schema.ts, scorer.ts (ported; monorepo libraries replaced with Either and internal/guards), benchmark.ts (dataset, solver, and defineSingleTurnBenchmark), colocated tests, and a README.
Wiring: config schema and union/options entry, TOOLCALL_FORMATS_META, registry entry, and the CLI fixed-temperature case.
Why?
Debugging GLM-5.3 on Morph showed that Morph doesn't constrain tool arguments to the schema. When the model drops required keys or sends "50" into number slots, Morph passes those calls through, and this benchmark catches that. It's moving to the public repo per the Slack thread.
How to test
bun test src/benchmarks/toolcall-formats
bun run bench --benchmark toolcall_formats --model z-ai/glm-5.3 --reasoning-effort low --concurrency 8 \
--solver-config '{"providerOnly":["morph"],"allowFallbacks":false}'
Benchmark impact
New benchmark. I ran the full 135 samples once on this branch (z-ai/glm-5.3, low effort, pinned with allowFallbacks: false):
Morph: 113/135. All 22 failures are schema_violation:
omit cases: required keys missing, including on strict variants;
Wafer: 115/135. All 20 failures are schema_violation:
mostly missing keys on the non-strict nullable variants;
"null" strings in boolean|null and enum slots on the ticket and weather cases.
These numbers come through the Responses API path, so they differ from the internal chat-completions run (Wafer 124, Morph 112 there). Wafer misses more non-strict nullable variants here.
Reviewer focus
Non-strict variants now send strict: false. The internal version left the flag out.
The solver looks up each sample's tools by sample.id.
Checklist
Tests cover changed behavior
Public API or configuration changes are backward compatible, or the break is documented
Benchmark changes document dataset provenance and licensing
No credentials, private results, or restricted dataset contents are included
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
Disable automatic comment, CI, and merge conflict monitoring
Original prompt from John
SYSTEM:
<latest_message>
John Paul Penaloza (U0BGKKT3AES) [ts=1790920778.808839]: @Devin check why specifically GLM-5.3 is failing tool calls on Morph? you can try running the tool format internal bench locally
</latest_message>
The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
Ports the internal
toolcall_formatsserving-stack benchmark from openrouter-web into the public harness, including the adversarial cases and scorer changes from OpenRouterTeam/openrouter-web#48802 (that PR doesn't need to merge).What changed?
src/benchmarks/toolcall-formats/:cases.ts,json-schema.ts,scorer.ts(ported; monorepo libraries replaced withEitherandinternal/guards),benchmark.ts(dataset, solver, anddefineSingleTurnBenchmark), colocated tests, and a README.omitted,anyof_null,anyof_null_strict,type_array_null,type_array_null_strict) = 135 samples. Sample ids are stable:toolcall_formats-<scenario>-<variant>.{read,bash,ticket}_omit_requestedand{read,bash,ticket}_wrong_type_requested.tools: entry.tools;temperature: 0, fixed throughFixedTemperatureBenchmarkBaseSchema;providerOnly,endpointId,allowFallbacks, and the rest) are forwarded.strictis always sent explicitly asstrict: variant.strict, so non-strict variants don't depend on how the API treats a missing flag.no_tool_call,unknown_tool,invalid_json,schema_violation,call_count,wrong_value.TOOLCALL_FORMATS_META, registry entry, and the CLI fixed-temperature case.Why?
Debugging GLM-5.3 on Morph showed that Morph doesn't constrain tool arguments to the schema. When the model drops required keys or sends
"50"intonumberslots, Morph passes those calls through, and this benchmark catches that. It's moving to the public repo per the Slack thread.How to test
Benchmark impact
New benchmark. I ran the full 135 samples once on this branch (z-ai/glm-5.3, low effort, pinned with
allowFallbacks: false):schema_violation:limit: "50",timeout: "30",blocking: "true".schema_violation:"null"strings inboolean|nulland enum slots on the ticket and weather cases.These numbers come through the Responses API path, so they differ from the internal chat-completions run (Wafer 124, Morph 112 there). Wafer misses more non-strict nullable variants here.
Reviewer focus
strict: false. The internal version left the flag out.sample.id.Checklist
Searched existing PRs ("toolcall", "tool call formats"); none matched.
Link to Devin session: https://openrouter.devinenterprise.com/sessions/257c4c8842f94f188a26b5a4e4772390
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/257c4c8842f94f188a26b5a4e4772390?variant=devin
Requested by: @johnpyp