The full battery against qwen3: 89% gradeable, and the edge probe behaves - #31
Merged
Conversation
All four domains, all three grading modes, 40 questions:
exact+text 32/36 (89%) computed-but-unused 1 errors 0
units 8/8 every conversion group perfect -- mass, length, time, digital
logic 8/8 pigeonhole, inclusion-exclusion, handshakes
arithmetic 16/20 misses in comparison, division_with_remainder, unit_chain
consistency (never counted as correctness):
anchored 2/2 steady facts any grounded model should hold
invented 0/2 steady places assembled from an RNG
The consistency rows are the result worth pinning hardest. The knowledge-edge
probe, run as a battery, behaved exactly as designed: steady where knowledge
exists, unsteady where nothing true can be known -- and an invented place
answered STEADILY would have been manufactured confidence, so the zero is the
honest signal. Pinned in both directions, with the caveat travelling in the
data: a consistently wrong answer is invisible to this grading.
One oddity recorded rather than chased: unit_chain (the arithmetic domain's
phrasing of seconds-in-N-weeks) went 0/2 while the units domain's conversion
questions went 8/8. Same underlying task, different phrasing, n=2 -- the noise
floor at this sample size has been measured at 28 points, so no conclusion is
drawn.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First run of all four domains and all three grading modes, 40 questions against
qwen3:4b-instruct-2507.The result worth pinning hardest
The consistency domain — the knowledge-edge probe run as a battery — behaved exactly as designed:
An invented place answered steadily would be manufactured confidence, so the zero is the success. Pinned in both directions, with the caveat travelling in the data itself: a consistently wrong answer is invisible to this grading, and these rows are never mixed into correctness.
One oddity recorded rather than chased:
unit_chain(arithmetic phrasing) 0/2 while the units domain's conversions went 8/8 — same task, different phrasing, n=2, and the measured noise floor at this size is 28 points. No conclusion drawn.4 new pinned claims; 415 tests.