Skip to content

claude-opus-latest alias silently serves opus-5 / 4.8 (billed at their higher price) when the target has no compatible endpoint; model: "unknown" on the web_search path #601

Description

@wukun2005-gif

Summary

Calling ~anthropic/claude-opus-latest with a request shape the current target model (anthropic/claude-opus-5.5) cannot serve causes OpenRouter to silently walk down the family to older models (anthropic/claude-opus-5, then anthropic/claude-4.8-opus) and bill at those models' higher list prices, instead of returning an error as documented.

This behavior is not documented on the Latest Model Resolution page, which only documents parameter remapping for reasoning fields and explicitly states there is "no built-in way to … roll back through the alias".

A second, related defect: on the Anthropic native web_search_20250305 server-tool path, the response model field came back as "unknown", violating the documented "transparent reporting" contract.


Expected behavior (per OpenRouter docs)

From Latest Model Resolution:

Target selection: The newest visible model in that family is selected.

Transparent reporting: The response's model field reports the concrete model that served the request … so you can always tell which version answered any given call.

If a family has no eligible model available, the request returns an error rather than falling back to something unrelated.

Compatibility contract: … Instead of rejecting the request, OpenRouter remaps unsupported reasoning parameters to the nearest supported value. (remap only — no model switching is described)

Only latest: The router always resolves to the newest eligible model. There's no built-in way to pin to "second newest" or to roll back through the alias. To downgrade, switch to a concrete slug.

From Model Fallbacks:

Pricing: Requests are priced using the model that was ultimately used, which will be returned in the model attribute of the response body.

Model-layer fallback is opt-in via the models array (Messages API: fallbacks). Neither was sent in any request below.

From the Opus 5.5 Migration Guide:

OpenRouter's catalog records this per endpoint, so a forced tool_choice on anthropic/claude-opus-5.5 fails at routing with no compatible endpoint rather than reaching the provider.


Actual behavior

GET /api/v1/models reports alias_target = {"slug": "anthropic/claude-opus-5.5"} for ~anthropic/claude-opus-latest.

Deterministic reproduction (three shapes, identical result every run)

// T1 — concrete slug, forced tool_choice (no thinking/effort fields)
{ "model": "anthropic/claude-opus-5.5", "max_tokens": 16,
  "tools": [{ "name": "get_weather", "description": "get weather",
              "input_schema": { "type": "object", "properties": { "city": { "type": "string" } }, "required": ["city"] } }],
  "tool_choice": { "type": "tool", "name": "get_weather" },
  "messages": [{ "role": "user", "content": "weather in Paris?" }] }

→ 400 tool_choice: type "tool" and "any" are not supported for this model. (routing-stage, no provider reached)

// T2 — same, on the alias (no thinking/effort)          → 200, model = anthropic/claude-opus-5
// T3 — alias + thinking:{type:"disabled"} + output_config.effort:"max"
//                                                          → 200, model = anthropic/claude-4.8-opus
// T4 — alias + thinking:{type:"disabled"} + output_config.effort:"high"
//                                                          → 200, model = anthropic/claude-opus-5
Test Request shape Response model Charged
T1 5.5 + forced tool_choice — (400) $0
T2 alias + forced tool_choice anthropic/claude-opus-5 $0.0028215
T3 alias + forced + disabled + effort:max anthropic/claude-4.8-opus $0.00283635
T4 alias + forced + disabled + effort:high anthropic/claude-opus-5 $0.0028215

The chain is fully explained by upstream rules:

  1. Opus 5.5 does not accept forced tool_use (Anthropic retired it on 5.5) → target has no compatible endpoint.
  2. Opus 5 accepts forced tool use, but rejects thinking:{type:"disabled"} at effort max/xhigh (Anthropic: "accepted only when the effort level is high or below … enforced on every request"; AWS Bedrock documents the identical rule) → 400 on all providers → continues down to 4.8.
  3. Opus 4.8 accepts disabled at any effort (per Anthropic's per-model table) → 200, and is billed at Opus 4.8 pricing.

Requested vs billed on live traffic

Generation records whose router field is ~anthropic/claude-opus-latest but which were served/billed on older models (Anthropic native web_search_20250305 + forced tool_choice shape):

generation_id model served prompt/completion charged
gen-1790470174-wUnBCgQ1CmEByYkK0OOP anthropic/claude-4.8-opus-20260528 1063 / 39 $0.007515337
gen-1790470182-XN9EREvitWXVMaLutwNK anthropic/claude-4.8-opus-20260528 1053 (1041 cached) / 50 $0.001812195
gen-1790470190-XDRnKwN5InjN5lVS75bk anthropic/claude-opus-5-20260723 13947 / 110 $0.081660150
gen-1790470174-pGllxHgoR10I3rsDPOv9 anthropic/claude-opus-5-20260723 0 / 0 (400 leg) $0
gen-1790470182-ckcLzlUw3NTL9dF8ewnE anthropic/claude-opus-5-20260723 0 / 0 (400 leg) $0

provider_responses on gen-1790470174-wUnBCgQ1CmEByYkK0OOP shows the cross-model move explicitly — no models/fallbacks was sent:

anthropic/claude-opus-5-20260723  Claude Platform on AWS   400
anthropic/claude-opus-5-20260723  Amazon Bedrock           400
anthropic/claude-opus-5-20260723  Anthropic                400
anthropic/claude-opus-5-20260723  Google                   400
anthropic/claude-opus-5-20260723  Azure                    400
anthropic/claude-4.8-opus-20260528 Claude Platform on AWS  200   ← billed

model: "unknown" on the server web_search path

Two responses to the same shape (alias + tools:[{"type":"web_search_20250305",...}] + tool_choice named tool) returned "model": "unknown" while still being billed:

  • gen-1790477611-CX3ze4JGf3xmVNRV24CE — cost 0.0300415075 (its own generation record is a 0-token 400 leg)
  • gen-1790477745-kOcGILbC8D8gkkVqV0s — cost 0.021007015 (the id returned in the response 404s on GET /api/v1/generation, so the billed leg cannot be audited at all)

Requests without the server tool (T2–T4) do report model correctly, so this appears specific to the server-tool execution path.


The providers are not at fault

Both providers named in the trace behave exactly as their published documentation requires:

  1. Availability was fine. GET /api/v1/models/*/endpoints at the time showed for all 11 endpoints of each model: status = 0, uptime_last_1d 99.72–100%, uptime_last_5m 100. All three models (5.5, 5, 4.8) are served by the same five providers (Amazon Bedrock, Claude Platform on AWS, Anthropic, Google, Azure) — so this is not a capacity, RPS, or outage issue, and there is no provider-availability gap for 5.5.

  2. It is a static capability flag, not load. Every endpoint of anthropic/claude-opus-5.5 declares
    supports_tool_choice = {"none": true, "auto": true, "required": false, "function": false},
    while every endpoint of anthropic/claude-opus-5 and anthropic/claude-4.8-opus declares required: true, function: true.
    A busy/rate-limited/down endpoint yields 429/5xx and provider failover within the same model — it never changes the model. Here the model changed, so the trigger was capability filtering.

  3. Both AWS providers reject disabled+max on Opus 5 as documented. Amazon Bedrock's Adaptive thinking page: "Effort cap when thinking is disabled (Claude Opus 5) … Requests with xhigh or max effort combined with disabled thinking will return an invalid_request_error." Anthropic's Thinking troubleshooting states the same rule and that it is "enforced on each request". All five providers returned the identical 400 for Opus 5, which is model-level, not an AWS implementation difference. The 200 from Opus 4.8 is likewise the documented behavior.

  4. The one documented Bedrock-specific quirk (InvokeModel rejecting per-message output_config.effort) is already handled by OpenRouter per your migration guide and is unrelated to this report.


Billing impact

Opus 5 / 4.8 are listed at $5 / $25 per M with $0.50 / $6.25 cache read/write, versus Opus 5.5's $4 / $20 with $0.20 / $5 — i.e. +25% on input and output, +150% on cache reads for work that was requested against the 5.5 alias.

Recomputing each affected generation at 5.5 rates with OpenRouter's applied discount reproduces the charged amounts to 9 decimal places, so the delta is exact:

generation charged at Opus 5.5 rates overcharge
gen-1790470174-wUnBCgQ1CmEByYkK0OOP $0.007515337 $0.006012270 $0.001503067
gen-1790470182-XN9EREvitWXVMaLutwNK $0.001812195 $0.001243638 $0.000568557
gen-1790470190-XDRnKwN5InjN5lVS75bk $0.081660150 $0.067308120 $0.014352030
total $0.090987682 $0.074564028 $0.016423654 (+22%)

Plus the two model: "unknown" responses totalling $0.05104852 whose billed model cannot be determined from the response or the generation endpoint.

Request: credit $0.016424 for the three attributable generations (and the $0.051049 once the billed models behind the unknown responses are identified).


What we'd like fixed

  1. Either document the family walk-down on the Latest Model Resolution page (it currently says the opposite: newest model only, no rollback, error rather than falling back), or stop doing it — return the routing error when the alias target has no compatible endpoint.
  2. Price the downgrade at the requested model's rate if the walk-down is intentional: being silently served by a 25%-more-expensive older model because the caller's request shape was incompatible with the target is a surprising way to pay more.
  3. Return the concrete model in model on the server web_search path instead of "unknown", and make the response id resolvable via GET /api/v1/generation so billed legs are auditable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions