Skip to content

The full battery against qwen3: 89% gradeable, and the edge probe behaves - #31

Merged
punnerud merged 1 commit into
mainfrom
measure/full-battery
Aug 12, 2026
Merged

The full battery against qwen3: 89% gradeable, and the edge probe behaves#31
punnerud merged 1 commit into
mainfrom
measure/full-battery

Conversation

@punnerud

Copy link
Copy Markdown
Owner

First run of all four domains and all three grading modes, 40 questions against qwen3:4b-instruct-2507.

result
exact+text 32/36 (89%) — computed-but-unused 1, errors 0
units 8/8 — every conversion group perfect
logic 8/8 — pigeonhole, inclusion-exclusion, handshakes
arithmetic 16/20

The result worth pinning hardest

The consistency domain — the knowledge-edge probe run as a battery — behaved exactly as designed:

group steady meaning
anchored (capital of France…) 2/2 held without wobbling
invented (RNG-assembled places) 0/2 unsteady — the honest signal

An invented place answered steadily would be manufactured confidence, so the zero is the success. Pinned in both directions, with the caveat travelling in the data itself: a consistently wrong answer is invisible to this grading, and these rows are never mixed into correctness.

One oddity recorded rather than chased: unit_chain (arithmetic phrasing) 0/2 while the units domain's conversions went 8/8 — same task, different phrasing, n=2, and the measured noise floor at this size is 28 points. No conclusion drawn.

4 new pinned claims; 415 tests.

All four domains, all three grading modes, 40 questions:

  exact+text   32/36 (89%)   computed-but-unused 1   errors 0

  units        8/8    every conversion group perfect -- mass, length, time, digital
  logic        8/8    pigeonhole, inclusion-exclusion, handshakes
  arithmetic  16/20   misses in comparison, division_with_remainder, unit_chain

  consistency (never counted as correctness):
    anchored   2/2 steady    facts any grounded model should hold
    invented   0/2 steady    places assembled from an RNG

The consistency rows are the result worth pinning hardest. The knowledge-edge
probe, run as a battery, behaved exactly as designed: steady where knowledge
exists, unsteady where nothing true can be known -- and an invented place
answered STEADILY would have been manufactured confidence, so the zero is the
honest signal. Pinned in both directions, with the caveat travelling in the
data: a consistently wrong answer is invisible to this grading.

One oddity recorded rather than chased: unit_chain (the arithmetic domain's
phrasing of seconds-in-N-weeks) went 0/2 while the units domain's conversion
questions went 8/8. Same underlying task, different phrasing, n=2 -- the noise
floor at this sample size has been measured at 28 points, so no conclusion is
drawn.
@punnerud
punnerud merged commit f0c4d38 into main Aug 12, 2026
6 checks passed
@punnerud
punnerud deleted the measure/full-battery branch August 12, 2026 15:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant