Reported to the authors by email on 20 July 2026; posting here at Dr. Zhang's suggestion. The authors responded promptly and noted that the prompt was written in early 2024 to enforce strict alignment with the clinical reference labels, using the judge models available at the time, and that the project is not currently under active development. This is filed as a limitation of the evaluation rubric — the full clinical and methodological write-up is attached below.
Summary
The Report Generation LLM-as-a-judge rubric (Supplementary Fig. 9) contains a single exclusionary clause that makes the evaluator structurally unable to penalise content the model invents. The clause is identical across all three versions of the work: arXiv:2410.19008v1, ICLR 2025 submission 10990, and the version of record in npj Digital Medicine 9, 349 (2026).
The clause:
"Only focus on information present in the ground truth report, identifying any mistakes."
Mechanism
The judge is directed to compare the generated report against the cardiologist's reference and mark what is missing. It is never directed to check the reverse — whether the report contains claims the cardiologist never made. Omissions are caught; inventions are not. The metric grades recall and reports it as report quality.
The counterargument that "identifying any mistakes" covers fabrication does not hold: all three scoring anchors are defined relative to what the ground truth contains (e.g. Diagnosis 10 = "all key diagnoses correctly identified with no errors or omissions").
Two cases from Supplementary Fig. 11 (single report, scored 100/100)
-
Fabricated lead enumeration. The reference reports high voltages in "limb leads," unqualified. The generated report specifies "leads I, II, and III" — an enumeration with no source in the reference and not derivable without measurement. It omits aVL, the limb lead that carries the standard LVH voltage criteria. Within the same output the model reproduces the ST-depression lead list (I, II, aVL, V5, V6) exactly as the reference gives it, so it copies lead lists faithfully where a source exists and generates one where none exists. Judge: Waveform 10/10, "descriptions match the ground truth report precisely."
-
Diagnostic over-call. The reference records "non-specific but consistent with myocardial changes." The generated report drops the "non-specific" qualifier and lists myocardial ischemia in its diagnosis line. Judge: Diagnosis 10/10.
Control: the clause is the cause
The ECG Arena rubric (Supplementary Fig. 10) omits this clause. There, the same GPT-4o judge did penalise PULSE for an analogous addition — scoring it 5 on Accuracy on the grounds that its LVH mention "is not part of the ground truth answer." Same judge model, same failure type, opposite outcome. The difference is the clause.
Proposed fix
Minimal: delete the exclusionary clause from Supplementary Fig. 9.
Thorough: add a fabrication-penalty dimension, scoped narrowly so it does not punish valid clinical elaboration. Penalise unsupported assertions about the tracing (hallucinated lead localisations, measurements, or findings claimed as directly observed) and diagnoses stated at higher certainty or acuity than the reference supports; leave general clinical explanation intact.
Confirming test: one prompt edit plus a re-run of the 500-case PTB-XL Report set. The hypothesis predicts scores fall for reports containing unsourced specifics while holding steady for reports that only omit findings.
Note on the human-correlation validation
The published human–judge correlation (Pearson r = 0.919–0.934) establishes that the judge ranks reports as clinicians do, but does not address this limitation: concordance between two evaluators both assessing against reference content cannot establish that either detects content the reference lacks.
Full write-up — clinical severity rationales, the decoupled judge rationales in the Arena case study, and the acceptance tests:
DiVi-Clinical-Labs_PULSE-ECGBench_Rubric-Review_v2_2026-07-20 (1).pdf
Analysis by Dr. Divyanshu Mishra, MBBS (Dr. DiVi Clinical Labs). No model inference was performed; all findings derive from the published figures and prompts.
Reported to the authors by email on 20 July 2026; posting here at Dr. Zhang's suggestion. The authors responded promptly and noted that the prompt was written in early 2024 to enforce strict alignment with the clinical reference labels, using the judge models available at the time, and that the project is not currently under active development. This is filed as a limitation of the evaluation rubric — the full clinical and methodological write-up is attached below.
Summary
The Report Generation LLM-as-a-judge rubric (Supplementary Fig. 9) contains a single exclusionary clause that makes the evaluator structurally unable to penalise content the model invents. The clause is identical across all three versions of the work: arXiv:2410.19008v1, ICLR 2025 submission 10990, and the version of record in npj Digital Medicine 9, 349 (2026).
The clause:
"Only focus on information present in the ground truth report, identifying any mistakes."
Mechanism
The judge is directed to compare the generated report against the cardiologist's reference and mark what is missing. It is never directed to check the reverse — whether the report contains claims the cardiologist never made. Omissions are caught; inventions are not. The metric grades recall and reports it as report quality.
The counterargument that "identifying any mistakes" covers fabrication does not hold: all three scoring anchors are defined relative to what the ground truth contains (e.g. Diagnosis 10 = "all key diagnoses correctly identified with no errors or omissions").
Two cases from Supplementary Fig. 11 (single report, scored 100/100)
Fabricated lead enumeration. The reference reports high voltages in "limb leads," unqualified. The generated report specifies "leads I, II, and III" — an enumeration with no source in the reference and not derivable without measurement. It omits aVL, the limb lead that carries the standard LVH voltage criteria. Within the same output the model reproduces the ST-depression lead list (I, II, aVL, V5, V6) exactly as the reference gives it, so it copies lead lists faithfully where a source exists and generates one where none exists. Judge: Waveform 10/10, "descriptions match the ground truth report precisely."
Diagnostic over-call. The reference records "non-specific but consistent with myocardial changes." The generated report drops the "non-specific" qualifier and lists myocardial ischemia in its diagnosis line. Judge: Diagnosis 10/10.
Control: the clause is the cause
The ECG Arena rubric (Supplementary Fig. 10) omits this clause. There, the same GPT-4o judge did penalise PULSE for an analogous addition — scoring it 5 on Accuracy on the grounds that its LVH mention "is not part of the ground truth answer." Same judge model, same failure type, opposite outcome. The difference is the clause.
Proposed fix
Minimal: delete the exclusionary clause from Supplementary Fig. 9.
Thorough: add a fabrication-penalty dimension, scoped narrowly so it does not punish valid clinical elaboration. Penalise unsupported assertions about the tracing (hallucinated lead localisations, measurements, or findings claimed as directly observed) and diagnoses stated at higher certainty or acuity than the reference supports; leave general clinical explanation intact.
Confirming test: one prompt edit plus a re-run of the 500-case PTB-XL Report set. The hypothesis predicts scores fall for reports containing unsourced specifics while holding steady for reports that only omit findings.
Note on the human-correlation validation
The published human–judge correlation (Pearson r = 0.919–0.934) establishes that the judge ranks reports as clinicians do, but does not address this limitation: concordance between two evaluators both assessing against reference content cannot establish that either detects content the reference lacks.
Full write-up — clinical severity rationales, the decoupled judge rationales in the Arena case study, and the acceptance tests:
DiVi-Clinical-Labs_PULSE-ECGBench_Rubric-Review_v2_2026-07-20 (1).pdf
Analysis by Dr. Divyanshu Mishra, MBBS (Dr. DiVi Clinical Labs). No model inference was performed; all findings derive from the published figures and prompts.