PhysElite is a bilingual, multimodal benchmark for evaluating Olympiad-level physics reasoning in large language models. It includes 10K+ problems, with visual diagrams, Chinese-English solution derivations, and final answers.
The paper evaluates 18 open-source and closed-source multimodal large language models. The strongest evaluated model achieves 33.7% answer accuracy, highlighting the challenge of advanced physics reasoning. Step-level evaluation further examines where models make mistakes during their solutions.
The benchmark data is hosted on Hugging Face. Its dataset viewer shows the available fields and lets you inspect examples before downloading. You can load it with the Hugging Face datasets library:
from datasets import load_dataset
physelite = load_dataset("physelite/PhysElite")
print(physelite)If you use PhysElite in your research, please cite:
@misc{xu2026physelite,
title = {PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?},
author = {Ruoran Xu and Wending Gao and Liyunfeng Chen and Aixin Shi and Haoyu Cheng and Zixiang Fang and Yiqiang Zou and Qiufeng Wang},
year = {2026},
eprint = {2608.25097},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.25097}
}