envelope: profile at the declared maximum input, fitted by the engine's own tokenizer - #109
Merged
Merged
Conversation
# Conflicts: # microtensor/harness/engines/gguf.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ship at a round boundary only, announced a round ahead. This changes admission for live arenas.
The bug
max_input_promptsized the latency and memory probe at 4 characters per token. Its probe words tokenise at about 8.6, so every probe came out at 46 to 49% of the declared maximum input. Peak memory is sampled during the same requests, so memory was understated too, not just latency. And the certificate'sinput_at_peakrecorded the declared input rather than what was fed, so certificates claimed a measurement at full size that was made at half.Measured with the Qwen tokenizer:
The fix
Engines report how many tokens a request actually feeds them (
input_tokens, on GGUF and ONNX, chat template and decisions included). After load, the profiler binary searches the probe length so the input fills the window: the declared tokens less the output budget for generation, since the context must also hold the answer, and the full declared tokens for decisions, which generate nothing. The fitting happens outside the cold start timing.input_at_peaknow recordsfed_tokensanddeclared_tokens.Engines without a counter keep today's behaviour.
Verified
Qwen3-0.6B, 2,048 declared, 256 output budget: the old probe fed 951 tokens (46%), the fitted probe feeds 1,791, and generation at that size completes all 256 output tokens.
Before it ships
Live certified systems must be re-measured on the reference device at the fitted size, because some will cross their class ceilings once they are measured at the input they declared.
🤖 Generated with Claude Code