You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds max to PINNED_REASONING_EFFORTS, so benchmark configs can run models such as openai/gpt-6-luna and deepseek/deepseek-v4.1-flash at their max effort.
What changed?
src/harness/constants.ts: max is now the first entry in PINNED_REASONING_EFFORTS. That also adds it to REASONING_EFFORTS, the reasoningEffort / userReasoningEffort config schemas, and the CLI --reasoning-effort check.
reasoningRequestFor("max") returns { effort: "max" }. The SDK's request effort enum already includes max. A test covers it.
Why?
The model catalog lists max for GPT-6 Luna and DeepSeek V4.1 Flash, but configs that use it fail schema validation. DeepSeek V4.1 Flash supports only max/high/low, so sending xhigh runs it at high. The agent CLI efforts (ORI_REASONING_EFFORTS) already include max, and so does toOriReasoningEffort's output type, so the agentic suites now receive it unchanged.
How to test
bun test src/harness/constants.test.ts
bun run src/cli/index.ts --benchmark gpqa_diamond --model openai/gpt-6-luna --reasoning-effort max --limit 1
The second command now passes argument validation and sends reasoning: { effort: "max" }.
Benchmark impact
No change to existing runs or their defaults. max is a new opt-in value.
Reviewer focus
Whether any consumer of ReasoningEffort should reject max. I found none: toOriReasoningEffort already maps into a superset that includes max.
Checklist
Tests cover changed behavior
Public API or configuration changes are backward compatible, or the break is documented
Benchmark changes document dataset provenance and licensing
No credentials, private results, or restricted dataset contents are included
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
Disable automatic comment, CI, and merge conflict monitoring
Original prompt from Ayush
SYSTEM:
<latest_message>
Ayush Patel (U0B8L6RNMA9) [ts=1791424352.095979]: @Devin !router_benchmark run these 3 models at these reasoning levels (single model baselines) on these 5 benchmarks. mimo v2.6pro should be at highest thinking level whatever it is
</latest_message>
=== BEGIN THREAD HISTORY (in #agents-benchmarks) ===
Ayush Patel (U0B8L6RNMA9) [ts=1791424343.837149]: !router_benchmark run these 3 models at these reasoning levels (single model baselines) on these 5 benchmarks. mimo v2.6pro should be at highest thinking level whatever it is
no acu,time,spend limit. no smoke, start full 1 epoch each right away.
The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
Adds
maxtoPINNED_REASONING_EFFORTS, so benchmark configs can run models such asopenai/gpt-6-lunaanddeepseek/deepseek-v4.1-flashat theirmaxeffort.What changed?
src/harness/constants.ts:maxis now the first entry inPINNED_REASONING_EFFORTS. That also adds it toREASONING_EFFORTS, thereasoningEffort/userReasoningEffortconfig schemas, and the CLI--reasoning-effortcheck.reasoningRequestFor("max")returns{ effort: "max" }. The SDK's request effort enum already includesmax. A test covers it.Why?
The model catalog lists
maxfor GPT-6 Luna and DeepSeek V4.1 Flash, but configs that use it fail schema validation. DeepSeek V4.1 Flash supports onlymax/high/low, so sendingxhighruns it athigh. The agent CLI efforts (ORI_REASONING_EFFORTS) already includemax, and so doestoOriReasoningEffort's output type, so the agentic suites now receive it unchanged.How to test
bun test src/harness/constants.test.ts bun run src/cli/index.ts --benchmark gpqa_diamond --model openai/gpt-6-luna --reasoning-effort max --limit 1The second command now passes argument validation and sends
reasoning: { effort: "max" }.Benchmark impact
No change to existing runs or their defaults.
maxis a new opt-in value.Reviewer focus
ReasoningEffortshould rejectmax. I found none:toOriReasoningEffortalready maps into a superset that includesmax.Checklist
Link to Devin session: https://openrouter.devinenterprise.com/sessions/116e7df8fcf449b4b19b1919ffe3cec5
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/116e7df8fcf449b4b19b1919ffe3cec5?variant=devin
Requested by: @ayush-or