Skip to content

scoring: per part breakdown of full system quality - #139

Merged
BANADDA merged 1 commit into
mainfrom
system-breakdown
Oct 1, 2026
Merged

BANADDA merged 1 commit into
mainfrom
system-breakdown

Conversation

@BANADDA

@BANADDA BANADDA commented Oct 1, 2026

Copy link
Copy Markdown
Member

Splits each certified full system's result into its parts so miners can see which part to improve. Information only: pay still follows the whole system's place on the frontier.

Part Measured as Cost
Small model quality alone minus the arena's quality_floor (untrained base model under the harness) none, from traces
Harness on the verification sample: the archived small model with the plain task prompt vs the same model through the harness; plus the output step's gain on every task one plain prompt generation per sampled task
Router and escalation escalation gain, rescues, waste, misses, precision, escalation spend per rescue none, from traces

quality = small_quality + escalation_gain + output_gain holds exactly.

Isolation. The harness probe runs in its own jailed call after certification. If it fails or runs out of budget, the harness part reads "not measured"; certification and scores are untouched.

Where it shows. The report's calibration.breakdown (the server's escalations summary picks it up), mt miner status, mt miner simulate (router and output step parts), the published card, and a new section in docs/system_miner.md.

Floor reaches the validator through the anchored arena config (RoundBudget.quality_floor).

🤖 Generated with Claude Code

@BANADDA
BANADDA merged commit 842c3e1 into main Oct 1, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant