Hey! I run Nonobench, a benchmark that has LLMs solve nonogram puzzles through OpenRouter. A reader spotted something odd in our Qwen3.8 Max results, and I can't answer it from the API alone.
What I see
/api/v1/models lists supported_efforts: ["xhigh", "high", "medium", "low", "minimal"] for qwen/qwen3.8-max-0902.
- Qwen's own docs list three levels for Qwen3.8-Max:
low, medium and xhigh (default).
- Your reasoning docs say unsupported levels map "to the nearest supported level". But the model lists all five as supported, and the Alibaba endpoint has
reasoning: null, so I can't tell which levels are native and which are mapped.
What our runs suggest
Same 30 puzzles, one attempt each, all served by Alibaba. Average tokens per puzzle:
| Sent |
5x5 |
10x10 |
15x15 |
Solved |
| minimal |
3.1k |
36.7k |
39.9k |
18/30 |
| low |
3.1k |
34.4k |
45.1k |
18/30 |
| medium |
4.2k |
38.5k |
38.3k |
14/30 |
| high |
4.7k |
30.1k |
55.9k |
18/30 |
| xhigh |
5.6k |
32.6k |
58.6k |
18/30 |
That looks like three settings, not five: minimal → low, and high → xhigh. high sits between medium and xhigh, so "nearest" could go either way.
Questions
- Which native effort do
minimal and high map to for this model?
- Could
supported_efforts list only native levels, or flag the mapped ones? Then anyone running effort ladders through OpenRouter knows which levels are real before paying for them.
- Is there a way to see the native effort you forwarded, e.g. in the generation metadata?
Happy to share generation IDs if that helps. Thanks for looking!
Hey! I run Nonobench, a benchmark that has LLMs solve nonogram puzzles through OpenRouter. A reader spotted something odd in our Qwen3.8 Max results, and I can't answer it from the API alone.
What I see
/api/v1/modelslistssupported_efforts: ["xhigh", "high", "medium", "low", "minimal"]forqwen/qwen3.8-max-0902.low,mediumandxhigh(default).reasoning: null, so I can't tell which levels are native and which are mapped.What our runs suggest
Same 30 puzzles, one attempt each, all served by Alibaba. Average tokens per puzzle:
That looks like three settings, not five: minimal → low, and high → xhigh.
highsits betweenmediumandxhigh, so "nearest" could go either way.Questions
minimalandhighmap to for this model?supported_effortslist only native levels, or flag the mapped ones? Then anyone running effort ladders through OpenRouter knows which levels are real before paying for them.Happy to share generation IDs if that helps. Thanks for looking!