fix(datasets): anchored grade extraction + surface judge failures in generic llmjudge - #2598
YuhaoLin2005 wants to merge 2 commits into
Conversation
…generic llmjudge _generic_llmjudge_postprocess took the first A/B anywhere in the judge's free-text reply, so reasoning letters silently overrode the real grade. Only an explicitly anchored grade (cue word + connector + letter: 'grade is B', 'grade: B', 'final answer = A', 'grade of A') is now trusted; everything else keeps the legacy loose scan (zero recall regression). get_final_results now surfaces judge failures via judge_error_count and a per-detail judge_error flag, while accuracy/not_attempted_count keep their historical semantics (unknown still counts as not-attempted).
The anchored grade pattern shipped in this branch had defects that
silently changed reported scores:
- re.IGNORECASE applied to the ([AB]) capture class, so the indefinite
article was read as the grade: "Grade: ambiguous, but leaning
strongly B." scored A (correct) where upstream scored B.
- "final answer" was a cue, but in a judge reply it names the
candidate's answer, not the grade.
- .search() returned the first anchor, so a stale or quoted earlier
grade beat the judge's final verdict.
- no word boundary after the letter, so the first letter of the next
word became the grade ("Verdict: Based on ..." -> B).
- when a cue opened a verdict the payload did not close, the fallback
re-scanned loosely and read the A out of the word "GRADE" itself.
- the anchor hardcoded [AB], ignoring the caller's true_tag/false_tag,
so a caller configured with e.g. 'A+' had correct samples silently
scored as not-attempted.
The pattern is now built from the caller's tags, matched
case-sensitively for the letter only, requires a standalone token,
and takes the last anchor so a self-correcting judge is read
correctly. The fallback no longer accepts a letter that is part of a
word. detail['judge_error'] is now present on every row so the details
table is no longer ragged. Comments now state plainly that accuracy
still counts judge failures in its denominator and that
<metric>_given_attempted is the figure excluding them.
Tests: 29 cases, each pinning one of the above.
Anchored grade extraction: 8 defects found and fixed before reviewI ran an adversarial review panel over the anchored-grade regex in this branch and reproduced every finding against real CPython Confirmed regressions (each reproduced, with the wrong value it produced)
Root causes, in order of severity:
What changed
Also fixed: a quadratic backtracking path
Verification
One design note for reviewers
|
Motivation
_generic_llmjudge_postprocessscans the judge's free-text reply for the firstA/Banywhere in it. Judges routinely justify their verdict before stating it, so a reasoning sentence that merely contains a letter silently overrides the real grade ("As a judge, I considered the evidence carefully. The final grade is B." parses asA). Separately, when a judge fails to emit a usable grade (API error, truncation, no letter in the reply), the result degrades to'unknown', whichget_final_resultsfolds into the accuracy denominator — indistinguishable from a genuine wrong answer. Related upstream reports: #1232, #2522, #2392.The fix (two parts)
1. Anchored grade extraction (no regression)
Only a grade that is explicitly anchored is trusted: a cue word (
grade,verdict,final answer) immediately followed by a connector (is/was/be/of/:/=/:=) and the letter. Everything else keeps the existing loose scan, so every judge reply that previously parsed still parses (zero recall regression, no silently-dropped samples).Deliberate exclusions, each with a regression test:
answeris not a cue — "the answer is A" names the option, not the grade;2. Judge failures are surfaced, not absorbed
get_final_resultsnow counts judge failures separately (judge_error_count, plus ajudge_errorflag on each failing detail), whileaccuracy,accuracy_given_attempted,not_attempted_count, and friends keep their exact historical semantics (unknownstill counts as not-attempted). A judge failure can no longer be silently read as a wrong answer.Tests
tests/datasets/test_generic_llmjudge_postprocess.py— 13 cases covering anchored extraction, the anti-false-anchor contract, and the aggregation surface. Followstests/TESTING_GUIDE.mdconventions.Impact on existing results
This fix changes how judge replies that contain an anchored grade are parsed:
previously such replies could be mis-parsed as the first letter appearing in the
judge's prose (often wrong); now the anchored grade wins. Replies with a single
letter, or with no conflicting prose letter, are parsed exactly as before, so
only previously-wrong results change — in the correct direction.
accuracy,accuracy_given_attempted, andnot_attempted_countkeep their exacthistorical semantics;
judge_error_countis a new, purely additive field.Out of scope
continue to follow the legacy scan — adding them would reintroduce the
qualifier false-positive the connector requirement exists to prevent.
MedXpertQA.py,medmcqa.py,supergpqa/supergpqa.py, andatlas/evaluation.py(each withits own
_generic_llmjudge_postprocess/get_final_results). This PR keepsthe diff focused on the canonical
generic.py; a follow-up could apply thesame anchored-extraction pattern to the copies.