diff --git a/benchmarks/benchmark_bigbang_proton.ipynb b/benchmarks/benchmark_bigbang_proton.ipynb new file mode 100644 index 00000000..267831ab --- /dev/null +++ b/benchmarks/benchmark_bigbang_proton.ipynb @@ -0,0 +1,132 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "# BigBang-Proton — matbench_mp_e_form Benchmark\n", + "\n", + "**Paper:** Wu, Liu, He, Xin, Wu et al. *BigBang-Proton Technical Report: Next-Word-Prediction is Scientific Multitask Learner*, arXiv:2510.00129 (2025). https://arxiv.org/abs/2510.00129\n", + "\n", + "**Resources:**\n", + "- Code: https://github.com/supersymmetry-technologies/BigBang-Proton\n", + "- Model: https://huggingface.co/SuperSymmetryTechnologies/BigBang-Proton\n", + "\n", + "---\n", + "\n", + "## Algorithm\n", + "\n", + "BigBang-Proton is a 1.5B-parameter unified sequence-based architecture for auto-regressive language modeling, pretrained **from scratch** on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientific multi-task learner. Three fundamental innovations:\n", + "\n", + "1. **Theory-Experiment Learning**: aligns large-scale numerical experimental data with theoretical text corpora;\n", + "2. **Binary Patch Encoding**: replaces BPE tokenization, transforming textual/numerical/symbolic scientific data into byte sequences (259-symbol vocabulary);\n", + "3. **Monte Carlo Attention**: substitutes traditional transformer attention.\n", + "\n", + "The base model is pretrained end-to-end by next-word prediction on multidisciplinary datasets — quark jet tagging (particle physics), inter-atomic potential simulations (materials science), genome and protein sequence structure prediction, lake water quality (spatio-temporal sensor) prediction, and up-to-50-digit arithmetic operations — mixed with general textual corpus (SlimPajama).\n", + "\n", + "**Adaptation to matbench_mp_e_form** (formation energy per atom, eV/atom):\n", + "\n", + "1. Supervised fine-tuning (SFT) of the base model on the MPTRJ (Materials Project + JARVIS) datasets -> BigBang-Proton-MPTRJ;\n", + "2. Each crystal structure is serialized into a compact deterministic text representation ending with the fixed prompt *\"The formation energy per atom of the lattice is:\"*;\n", + "3. The model autoregressively generates the formation-energy value as a byte string (e.g. `-1.23456 eV.`), parsed back to float;\n", + "4. Standard matbench protocol: one model fine-tuned per fold on the official train/val split, starting from BigBang-Proton-MPTRJ weights;\n", + "5. Test-time decoding is greedy (argmax) with temperature-1e-4 fallback sampling for non-digit tokens; per fold, best-loss and best-acc checkpoints are averaged.\n", + "\n", + "**Results (official test folds, MAE eV/atom):** fold_0 0.0914, fold_1 0.0896, fold_2 0.0887, fold_3 0.0934, fold_4 0.0871; 5-fold mean 0.0900." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Recording the benchmark results\n", + "\n", + "The complete per-fold predictions are provided in `results.json.gz` (this package). The cell below re-records them through the official `MatbenchBenchmark` flow (`get_test_data` -> `record` -> `to_file`), exactly as described in the matbench documentation. This reproduces the submitted `results.json.gz` from the saved prediction arrays." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import gzip\n", + "import json\n", + "import numpy as np\n", + "from matbench.bench import MatbenchBenchmark\n", + "\n", + "# Saved per-fold prediction arrays (test.jsonl order == official test order)\n", + "PRED_FILES = {\n", + " 0: \"preds_fold0.npy\",\n", + " 1: \"preds_fold1.npy\",\n", + " 2: \"preds_fold2.npy\",\n", + " 3: \"preds_fold3.npy\",\n", + " 4: \"preds_fold4.npy\",\n", + "}\n", + "\n", + "mb = MatbenchBenchmark(autoload=False, subset=[\"matbench_mp_e_form\"])\n", + "task = list(mb.tasks)[0]\n", + "task.load() # downloads the official dataset on first run\n", + "\n", + "for fold in task.folds:\n", + " test_inputs = task.get_test_data(fold, include_target=False)\n", + " preds = np.load(PRED_FILES[fold]).astype(float)\n", + " assert len(preds) == len(test_inputs), f\"fold{fold}: length mismatch\"\n", + " task.record(fold, preds) # predictions are floats (eV/atom)\n", + " print(f\"fold{fold}: recorded {len(preds)} predictions\")\n", + "\n", + "mb.to_file(\"results.json.gz\")\n", + "print(\"saved results.json.gz\")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Verification\n", + "\n", + "Reload the submission file through the official loader path to confirm it validates (5 folds, 132,752 predictions)." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "with gzip.open(\"results.json.gz\", \"rt\", encoding=\"utf-8\") as f:\n", + " d = json.load(f)\n", + "mb2 = MatbenchBenchmark.from_dict(d) # official load path\n", + "for t in mb2.tasks:\n", + " for f in t.folds:\n", + " n = len(t.predictions[f])\n", + " print(f\"{t.dataset_name} fold{f}: {n} predictions OK\")\n", + "print(\"total:\", sum(len(t.predictions[f]) for t in mb2.tasks for f in t.folds))" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Notes\n", + "\n", + "- One model is fine-tuned per fold on the official train split (standard matbench protocol); all folds start from the same BigBang-Proton-MPTRJ weights.\n", + "- Full training/inference code lives in the BigBang-Proton materials repository; this notebook provides the matbench integration.\n", + "- Requirements: python>=3.8, numpy, matbench>=0.6. See `info.json` for details." + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "Python 3", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "version": "3" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/benchmarks/info.json b/benchmarks/info.json new file mode 100644 index 00000000..114f6067 --- /dev/null +++ b/benchmarks/info.json @@ -0,0 +1,10 @@ +{ + "authors": "Hengkui Wu, Liujiang Liu, Jihua He, Yingsi Xin, Hengyuan Wu", + "algorithm": "BigBang-Proton", + "algorithm_long": "BigBang-Proton is a unified sequence-based architecture for auto-regressive language modeling, pretrained from scratch on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientific multi-task learner. It incorporates three fundamental innovations compared to mainstream general-purpose LLMs: (1) the Theory-Experiment Learning paradigm aligns large-scale numerical experimental data with theoretical text corpora; (2) Binary Patch Encoding replaces BPE tokenization, transforming textual, numerical and symbolic scientific data into byte sequences (a 259-symbol vocabulary: 256 byte values plus 3 special tokens); (3) Monte Carlo Attention substitutes traditional transformer attention, computing attention over patches while exponentially expanding effective context length with layer depth. Multidisciplinary datasets \u2014 quark jet tagging from particle collision experiments, inter-atomic potential simulations from materials science, genome and protein sequence structure prediction, lake water quality (spatio-temporal sensor) prediction, and up-to-50-digit arithmetic operations \u2014 are curated under the Theory-Experiment paradigm and, mixed with general textual corpus (SlimPajama), concatenated into a unified byte sequence for end-to-end next-word-prediction pretraining, jointly training all specialized tasks in a single base model.\n\nFor matbench_mp_e_form, we perform supervised fine-tuning (SFT) of the base model on the MPTRJ (Materials Project + JARVIS) datasets, obtaining BigBang-Proton-MPTRJ, a materials-adapted version. Each crystal structure is serialized into a compact, deterministic text representation (atomic numbers, Cartesian positions, lattice matrix, composition, and a fixed prompt ending with \u201cThe formation energy per atom of the lattice is:\u201d). The model autoregressively generates the formation energy value (eV/atom) as a byte string (e.g. \u201c-1.23456 eV.\u201d), which is parsed back into a float. Following the standard matbench protocol, one model is fine-tuned per fold on the official train/val split, starting from the BigBang-Proton-MPTRJ weights. Generation at test time is greedy (argmax), with temperature-1e-4 fallback sampling for non-digit tokens, identical to validation-time decoding. For each fold, the predictions of two checkpoints (best validation loss and best validation token accuracy) are averaged.\n\nResults on the official test folds (MAE, eV/atom): fold_0 0.0914, fold_1 0.0896, fold_2 0.0887, fold_3 0.0934, fold_4 0.0871; 5-fold mean 0.0900.", + "bibtex_refs": "@article{Dunn2020,\n doi = {10.1038/s41524-020-00406-3},\n url = {https://doi.org/10.1038/s41524-020-00406-3},\n year = {2020},\n month = sep,\n publisher = {Springer Science and Business Media {LLC}},\n volume = {6},\n number = {1},\n author = {Alexander Dunn and Qi Wang and Alex Ganose and Daniel Dopp and Anubhav Jain},\n title = {Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm},\n journal = {npj Computational Materials}\n}\n@techreport{wu2025bigbangproton,\n title = {BigBang-Proton Technical Report: Next-Word-Prediction is Scientific Multitask Learner},\n author = {Wu, Hengkui and Liu, Liujiang and He, Jihua and Xin, Yingsi and Wu, Hengyuan and others},\n institution = {SuperSymmetry Technologies},\n year = {2025},\n howpublished = {arXiv preprint arXiv:2510.00129},\n doi = {10.48550/arXiv.2510.00129},\n url = {https://arxiv.org/abs/2510.00129}\n}", + "notes": "Base model: BigBang-Proton pretrained on multi-disciplinary scientific task datasets (particle physics, material physics, sensor spatio-temporal data, DNA/RNA/protein) mixed with SlimPajama general language. Materials adaptation: fine-tuned on MPTRJ (Materials Project + JARVIS) datasets to obtain BigBang-Proton-MPTRJ. Following the standard matbench protocol, one model is trained per fold on the official train split: five fold-specific fine-tuned checkpoints, all starting from the BigBang-Proton-MPTRJ weights. Test-time ensemble: best_loss + best_acc checkpoints averaged. Predictions are 4-decimal floats recovered from inference logs. Reference MAE on this task (Automatminer/MEGNet v0.2.2): 0.0327 eV/atom.\nLinks: paper https://arxiv.org/abs/2510.00129 ; code https://github.com/supersymmetry-technologies/BigBang-Proton ; model https://huggingface.co/SuperSymmetryTechnologies/BigBang-Proton", + "requirements": { + "python": ["matbench>=0.6", "torch", "numpy", "pymatgen"] + } +} diff --git a/benchmarks/results.json.gz b/benchmarks/results.json.gz new file mode 100644 index 00000000..b79ae422 Binary files /dev/null and b/benchmarks/results.json.gz differ