Skip to content

envelope: profile at the declared maximum input, fitted by the engine's own tokenizer - #109

Merged
BANADDA merged 2 commits into
mainfrom
measure-at-declared-max
Sep 30, 2026
Merged

BANADDA merged 2 commits into
mainfrom
measure-at-declared-max

Conversation

@BANADDA

@BANADDA BANADDA commented Sep 30, 2026

Copy link
Copy Markdown
Member

Ship at a round boundary only, announced a round ahead. This changes admission for live arenas.

The bug

max_input_prompt sized the latency and memory probe at 4 characters per token. Its probe words tokenise at about 8.6, so every probe came out at 46 to 49% of the declared maximum input. Peak memory is sampled during the same requests, so memory was understated too, not just latency. And the certificate's input_at_peak recorded the declared input rather than what was fed, so certificates claimed a measurement at full size that was made at half.

Measured with the Qwen tokenizer:

declared probe tokens share
512 249 49%
1,024 484 47%
2,048 956 47%
4,096 1,901 46%

The fix

Engines report how many tokens a request actually feeds them (input_tokens, on GGUF and ONNX, chat template and decisions included). After load, the profiler binary searches the probe length so the input fills the window: the declared tokens less the output budget for generation, since the context must also hold the answer, and the full declared tokens for decisions, which generate nothing. The fitting happens outside the cold start timing. input_at_peak now records fed_tokens and declared_tokens.

Engines without a counter keep today's behaviour.

Verified

Qwen3-0.6B, 2,048 declared, 256 output budget: the old probe fed 951 tokens (46%), the fitted probe feeds 1,791, and generation at that size completes all 256 output tokens.

Before it ships

Live certified systems must be re-measured on the reference device at the fitted size, because some will cross their class ceilings once they are measured at the input they declared.

🤖 Generated with Claude Code

# Conflicts:
#	microtensor/harness/engines/gguf.py
@BANADDA
BANADDA merged commit 9d8c03e into main Sep 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant