Skip to content

Next round boundary: latency and memory are measured at your full declared input #110

Description

@BANADDA

From the next round boundary, admission profiles every artifact at the maximum input it declares. Today it does not.

What changes. The profiling probe was sized by a character estimate and came out at about 47% of the declared max_input tokens. From the boundary, each engine counts its own input tokens and the probe is fitted to fill the window: the declared tokens less the output budget for generation tracks, and the full declared tokens for decision tracks. Peak memory is sampled during the same requests, so it rises too.

What this means for you. Your p95 latency and peak RSS will be measured on roughly twice the input they are measured on today. If your artifact sits close to its class ceiling, it may no longer fit.

What to do. Either declare a max_input your artifact can serve inside the class ceilings, or check it now: mt miner selfcheck will use the new fitted probe once this ships.

Live certified systems are re-measured at the boundary on the reference device. Operators of any system that would fail are told before it takes effect.

Change: #109.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions