Skip to content

feat: accept max reasoning effort - #137

Merged
ayush-or merged 1 commit into
mainfrom
devin/1791425200-max-reasoning-effort
Oct 8, 2026
Merged

ayush-or merged 1 commit into
mainfrom
devin/1791425200-max-reasoning-effort

Conversation

@ayush-or

@ayush-or ayush-or commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

TL;DR

Adds max to PINNED_REASONING_EFFORTS, so benchmark configs can run models such as openai/gpt-6-luna and deepseek/deepseek-v4.1-flash at their max effort.

What changed?

  • src/harness/constants.ts: max is now the first entry in PINNED_REASONING_EFFORTS. That also adds it to REASONING_EFFORTS, the reasoningEffort / userReasoningEffort config schemas, and the CLI --reasoning-effort check.
  • reasoningRequestFor("max") returns { effort: "max" }. The SDK's request effort enum already includes max. A test covers it.

Why?

The model catalog lists max for GPT-6 Luna and DeepSeek V4.1 Flash, but configs that use it fail schema validation. DeepSeek V4.1 Flash supports only max/high/low, so sending xhigh runs it at high. The agent CLI efforts (ORI_REASONING_EFFORTS) already include max, and so does toOriReasoningEffort's output type, so the agentic suites now receive it unchanged.

How to test

bun test src/harness/constants.test.ts
bun run src/cli/index.ts --benchmark gpqa_diamond --model openai/gpt-6-luna --reasoning-effort max --limit 1

The second command now passes argument validation and sends reasoning: { effort: "max" }.

Benchmark impact

No change to existing runs or their defaults. max is a new opt-in value.

Reviewer focus

  • Whether any consumer of ReasoningEffort should reject max. I found none: toOriReasoningEffort already maps into a superset that includes max.

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented
  • Benchmark changes document dataset provenance and licensing
  • No credentials, private results, or restricted dataset contents are included
  • Documentation is updated where needed

Link to Devin session: https://openrouter.devinenterprise.com/sessions/116e7df8fcf449b4b19b1919ffe3cec5
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/116e7df8fcf449b4b19b1919ffe3cec5?variant=devin
Requested by: @ayush-or

@ayush-or
ayush-or requested a review from a team as a code owner October 8, 2026 02:06
@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Original prompt from Ayush

SYSTEM:
<latest_message>
Ayush Patel (U0B8L6RNMA9) [ts=1791424352.095979]: @Devin !router_benchmark run these 3 models at these reasoning levels (single model baselines) on these 5 benchmarks. mimo v2.6pro should be at highest thinking level whatever it is
</latest_message>

=== BEGIN THREAD HISTORY (in #agents-benchmarks) ===
Ayush Patel (U0B8L6RNMA9) [ts=1791424343.837149]: !router_benchmark run these 3 models at these reasoning levels (single model baselines) on these 5 benchmarks. mimo v2.6pro should be at highest thinking level whatever it is
no acu,time,spend limit. no smoke, start full 1 epoch each right away.

ATTACHMENT:"https://openrouter.devinenterprise.com/attachments/b4d58163-f1db-4112-b959-ddb4aacb23b6/image.png"

ATTACHMENT:"https://openrouter.devinenterprise.com/attachments/112fbfad-67b2-430b-af7f-e9edbfa52e1a/image.png"

Ayush Patel (U0B8L6RNMA9) [ts=1791424352.095979]: @Devin !router_benchmark run these 3 models at these reasoning levels (single model baselines) on these 5 benchmarks. mimo v2.6pro should be at highest thinking level whatever it is
=== END THREAD HISTORY ===
Channel ID: C0BAKP8P5C3
Thread URL: https://openrouter.slack.com/archives/C0BAKP8P5C3/p1791424343837149?thread_ts=1791424343.837149&amp;cid=C0BAKP8P5C3

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

@playbook:playbook-67bd60e6265449979d13f557104b2ef7

@ayush-or
ayush-or merged commit bd5c6d0 into main Oct 8, 2026
4 checks passed
@ayush-or
ayush-or deleted the devin/1791425200-max-reasoning-effort branch October 8, 2026 02:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant