This dataset contains 27,540 LLM judgments from a blind peer matrix evaluation of 55 frontier language models across 286 evaluations and 198 unique questions in 9 category pools. Of the 27,540: 23,356 parsed successfully and 22,252 carry a usable score (the analysis set); 2,781 are intentional self-exclusions (the matrix diagonal) and 1,403 are judge failures (parse/API errors). To our knowledge, this is the largest publicly available multi-judge LLM evaluation dataset with full provenance. See scripts/count_reconciliation.py for the exact breakdown.
Paper: The Multivac: Blind Peer Matrix Evaluation of Frontier Language Models
multivac-evaluation/
├── README.md # This file
├── LICENSE # MIT License
├── DATASHEET.md # NeurIPS Datasheet for Datasets
├── CITATION.cff # Citation metadata
├── evaluation_framework/
│ ├── multivac.py # Core evaluation engine
│ ├── extract_multivac_data.py # Data extraction pipeline
│ ├── statistical_analysis.py # Statistical tests
│ ├── questions.py # Wave 1 question bank
│ └── questions_wave2_15032026.py # Wave 2 question bank
├── data/
│ ├── peer_matrix/ # 286 EVAL-* folders
│ └── head_to_head/ # H2H batch folders
└── paper_tables/ # Pre-computed tables from the paper
| Metric | Value |
|---|---|
| Peer matrix evaluations | 286 |
| Unique questions | 198 |
| Total judgments | 27,540 |
| Parsed judgments | 23,356 |
| Usable-scored (analysis set) | 22,252 |
| Self-excluded (diagonal) | 2,781 |
| Judge failures (parse/API) | 1,403 |
| Unique models | 55 |
| Vendor families | 17 |
| Category pools | 9 |
| H2H questions (recorded) | 185 |
| Date range | Feb–Apr 2026 |
- No category has a statistically separated winner — in all nine pools the top model beats second place in only 50–58% of matched comparisons, with the 95% interval spanning 50%. Four distinct models lead the six primary pools by point estimate, but none separates from its runner-up (
scripts/make_finding1.py) - A same-vendor scoring premium survives correction for only 2 of 8 families — under a within-response fixed-effects model with standard errors clustered two-way by judge and question, Anthropic (+0.41) is robust (worst-case leave-one-judge-out p < 10⁻⁴) and MiniMax (+0.40) is significant under the primary specification (p = 0.001) but judge-dependent (its two-way significance does not survive leave-one-judge-out, worst-case p ≈ 0.02); Qwen (+0.56) does not survive the two-way clustering and is not counted. The large naive estimates — including Mistral −1.02 and Google −0.59 — are artifacts of judge leniency and respondent quality, not a same-vendor premium, and do not survive controls. The two surviving premiums are an upper bound on favoritism, not a measurement of it: the estimator cannot separate preferential scoring from a sibling judge parsing same-distribution output more accurately. See
paper_tables/FOUR_CELL_DECOMPOSITION_FINDINGS.mdandWITHIN_RESPONSE_FINDINGS.md. (Supersedes the earlier "significant bias in all families" claim.) - The "overall" champion is an artifact of the aggregation — naive mean, leniency-adjusted Bradley–Terry, and exposure-restricted BT each crown a different model. Among broad-participation generalists GPT-5.4 holds rank 1 in 99% of question-clustered resamples, with a 54% per-comparison edge over the runner-up (
scripts/make_finding1.py). (An earlier bootstrap on the naive-mean leaderboard — where rank 1 is Grok 4.1 Fast — gave p = 0.266/0.073/0.071 vs ranks 2–4; that aggregate is leniency-confounded and is superseded by the BT rankings.) - Judge disagreement is category-dependent — code σ=1.27 vs meta-alignment σ=0.63 (ratio 2.01×;
scripts/category_disagreement.py) - Overall inter-annotator agreement: Krippendorff's α = 0.618 (
evaluation_framework/statistical_analysis.py)
- Model selection: Category-specific rankings for task-appropriate model choice
- DPO/RLHF training data: 90 scored judgments per evaluation, convertible to preference pairs, at <$0.01/sample
- Judge bias research: Analysis of systematic judge behavior and family bias
- Evaluation methodology research: Comparing single-judge vs multi-judge approaches
- Safety monitoring: Tracking alignment behavior across model versions
- All judgments are from LLMs, not humans. Shared biases across models are not captured.
- Question selection reflects one author's judgment of what constitutes a good evaluation prompt.
- Model participation is non-uniform: core-pool models appear in 163–238 evaluations, focused-batch models in as few as 1.
- See §6 of the paper for full limitations discussion.
@article{darji2026multivac,
title={The Multivac: Blind Peer Matrix Evaluation of Frontier Language Models},
author={Darji, Yash},
journal={arXiv preprint arXiv:TODO},
year={2026}
}MIT License. See LICENSE for details.
- Platform: app.themultivac.com
- Discord: discord.gg/QvVTPCxH
- ORCID: 0009-0009-6895-842X