Multiverse-core leaderboard
25 estimators on 57 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-24
| # | Estimator | Accuracy | Balanced accuracy | AUROC | F1 | Log loss ↓ | Sensitivity | Specificity | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | Rank | Score | Rank | Score | Rank | Score | Rank | Score | Rank | Score | Rank | Score | Rank | ||
| 1 | HC2 | 0.7938 | 7.40 | 0.7543 | 8.48 | 0.8951 | 5.95 | 0.7378 | 8.11 | 0.5316 | 5.81 | 0.7570 | 8.50 | 0.7987 | 7.11 |
| 2 | MRHydra | 0.7838 | 8.77 | 0.7529 | 8.61 | 0.8074 | 17.72 | 0.7336 | 8.65 | 7.7909 | 21.54 | 0.7609 | 9.07 | 0.7828 | 9.27 |
| 3 | RDST | 0.7752 | 9.18 | 0.7396 | 9.46 | 0.7968 | 17.94 | 0.7179 | 9.64 | 8.1012 | 21.61 | 0.7289 | 10.29 | 0.7891 | 9.00 |
| 4 | RIST | 0.7742 | 9.82 | 0.7430 | 10.46 | 0.8677 | 8.32 | 0.7213 | 10.66 | 0.6123 | 8.12 | 0.7469 | 11.11 | 0.7724 | 10.69 |
| 5 | CIF | 0.7775 | 10.16 | 0.7463 | 10.31 | 0.8843 | 8.35 | 0.7293 | 10.18 | 0.6414 | 9.11 | 0.7521 | 11.06 | 0.7767 | 10.39 |
| 6 | DrCIF | 0.7722 | 10.37 | 0.7402 | 10.90 | 0.8734 | 8.77 | 0.7216 | 11.24 | 0.6413 | 9.09 | 0.7437 | 11.71 | 0.7718 | 11.18 |
| 7 | Arsenal | 0.7709 | 10.80 | 0.7336 | 10.97 | 0.8469 | 13.25 | 0.7145 | 11.00 | 3.6382 | 17.64 | 0.7340 | 11.18 | 0.7792 | 10.65 |
| 8 | QUANT | 0.7651 | 10.92 | 0.7362 | 10.90 | 0.8678 | 8.32 | 0.7177 | 10.98 | 0.7263 | 7.63 | 0.7506 | 11.00 | 0.7554 | 12.10 |
| 9 | ROCKET | 0.7701 | 11.04 | 0.7345 | 10.81 | 0.7927 | 18.79 | 0.7134 | 11.28 | 8.2862 | 22.41 | 0.7315 | 11.74 | 0.7781 | 11.31 |
| 10 | LITETime-MV | 0.7483 | 11.24 | 0.7260 | 10.21 | 0.8493 | 9.90 | 0.6851 | 10.20 | 1.3395 | 13.12 | 0.7129 | 9.98 | 0.7668 | 11.25 |
| 11 | STSF | 0.7699 | 11.41 | 0.7447 | 10.86 | 0.8700 | 10.23 | 0.7145 | 11.54 | 0.6742 | 8.40 | 0.7379 | 11.98 | 0.7827 | 12.22 |
| 12 | H-InceptionTime | 0.7379 | 11.89 | 0.7137 | 11.15 | 0.8434 | 10.39 | 0.6837 | 11.00 | 1.3662 | 13.67 | 0.7192 | 10.89 | 0.7389 | 12.55 |
| 13 | Catch22 | 0.7539 | 12.90 | 0.7224 | 13.24 | 0.8680 | 10.83 | 0.7027 | 13.46 | 0.6972 | 10.49 | 0.7290 | 13.72 | 0.7507 | 13.47 |
| 14 | ConvTran | 0.7446 | 13.12 | 0.7110 | 12.88 | 0.8520 | 10.98 | 0.6846 | 12.74 | 0.9070 | 10.11 | 0.7219 | 12.75 | 0.7356 | 13.95 |
| 15 | PatchMTSC | 0.7443 | 13.31 | 0.6959 | 13.75 | 0.8247 | 12.53 | 0.6682 | 13.30 | 0.7905 | 9.53 | 0.6996 | 12.92 | 0.7381 | 13.59 |
| 16 | DisjointCNN | 0.7224 | 13.64 | 0.6975 | 12.90 | 0.8239 | 11.61 | 0.6611 | 13.16 | 2.1344 | 15.18 | 0.6814 | 12.43 | 0.7299 | 12.82 |
| 17 | STC | 0.7518 | 13.86 | 0.7106 | 14.43 | 0.8637 | 11.74 | 0.6853 | 14.41 | 0.6428 | 9.77 | 0.7104 | 14.54 | 0.7599 | 13.31 |
| 18 | TSF | 0.7419 | 14.15 | 0.7138 | 13.65 | 0.8558 | 12.24 | 0.6926 | 13.88 | 0.9914 | 10.75 | 0.7141 | 14.51 | 0.7498 | 14.26 |
| 19 | TDE | 0.7272 | 14.94 | 0.6843 | 15.06 | 0.8379 | 12.71 | 0.6601 | 14.49 | 0.8450 | 11.46 | 0.6919 | 14.14 | 0.7304 | 14.04 |
| 20 | TS2Vec | 0.7192 | 15.41 | 0.6790 | 15.45 | 0.7995 | 15.32 | 0.6562 | 15.18 | 0.7869 | 11.58 | 0.6907 | 15.28 | 0.7125 | 15.47 |
| 21 | Summary | 0.6936 | 16.52 | 0.6614 | 16.04 | 0.8194 | 15.72 | 0.6393 | 16.16 | 0.9505 | 13.39 | 0.6648 | 16.39 | 0.6979 | 17.08 |
| 22 | XCM | 0.6691 | 16.80 | 0.6359 | 17.04 | 0.7915 | 14.94 | 0.5859 | 16.22 | 2.1699 | 16.28 | 0.6299 | 15.33 | 0.6769 | 15.62 |
| 23 | TimesURL | 0.6975 | 17.23 | 0.6565 | 17.21 | 0.7833 | 18.85 | 0.6206 | 17.59 | 0.9962 | 15.95 | 0.6456 | 17.52 | 0.6998 | 16.65 |
| 24 | TimesNet | 0.7001 | 17.41 | 0.6647 | 16.95 | 0.8209 | 16.04 | 0.6399 | 17.25 | 1.2957 | 15.30 | 0.6804 | 16.47 | 0.6919 | 17.68 |
| 25 | Dummy | 0.3748 | 22.70 | 0.3072 | 23.29 | 0.5000 | 23.56 | 0.1836 | 22.68 | 1.3810 | 17.09 | 0.3248 | 20.49 | 0.3774 | 19.34 |
Average score and average rank over the 57 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.
Missing results
- MRHydra — Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
- RDST — Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738)); USCActivity (Segmentation fault (core dumped))
- ConvTran — Alzheimers, EigenWorms, PhotoStimulation (CUDA out of memory)
- TS2Vec — Locust2022, Tiselac, USCActivity (Timed out at 60 hours)
Scoring uses the 57 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.
Notes on listed estimators
- XCM — run at fixed parameters, a single fit at window 0.8 with batch 32, not the per-dataset cross-validated search over window and batch size that the paper describes. The search was run and did not pay: across the 65 shared datasets it was 0.017 mean accuracy worse, 31 wins to 31 with 3 ties, Wilcoxon p = 0.63. On the 14 datasets where the search selected 0.8, the window used here, the two runs still differed by 0.11 mean absolute accuracy and by as much as 0.48, so at one resample XCM's run-to-run variance is larger than the effect the search is tuning for.
Datasets not included
- AustraliaRainfall_disc — 112186 cases, and three estimators fail on it for reasons compute cannot fix. RDST and ROCKET hit LAPACK integer overflow in RidgeClassifierCV's SVD, aeon issue 3738, after 14 and 12 attempts; MRHydra exhausted 128 GB over 13.
- PenDigits — the series are length 8 and MRHydra requires at least 9, so the dataset cannot complete while MRHydra is a column.
- BenzeneConcentration_disc — removed from Multiverse-core. The results here are on version 1, whose PT08.S2 channel is a deterministic function of the target; version 2 drops it, on the advice of the original UCI Air Quality donors: https://zenodo.org/records/21871727.
Held out of the collection rather than reported as missing, because no scheduling closes them. Results that do exist for them remain in the repository.
Estimators not listed
- LiteTIME — LITE is a univariate architecture. The multivariate variant of the same method is listed here as LITETime-MV.
- FreshPRINCE — cannot complete the archive at the memory available: recorded OOM at 128 GB after eight attempts each on FaceDetection, FordChallenge and Skoda, and 38 on Tiselac.
- 1NN-DTW — cannot complete the archive within the walltime available: exceeded the limit on BIDMC32HR_disc, with no result recorded for BIDMC32SpO2_disc.
- MUSE — cannot complete the archive as implemented: on STEW and Skoda its word bag outgrows the 32-bit sparse indices scikit-learn's classifier accepts, which no memory allocation fixes, and on USCActivity its chi2 selection needs 287 GiB. MotorImagery exhausted 16 GB and awaits a rerun at the configuration of its other results.
- CIF-500 — the 500-tree configuration run for the Multiverse archive paper's full-archive benchmark, where it is reported as CIF. This table reports the default CIF; the full-archive leaderboard reports CIF-500.
- DrCIF-500 — the 500-tree configuration HC2 uses internally, run for the Multiverse archive paper's full-archive benchmark, where it is reported as DrCIF. This table reports the default DrCIF; the full-archive leaderboard reports DrCIF-500.
- DisjointCNN-Aeon — aeon's implementation applies a Permute after the final block, so its pooling reduces the wrong axes and the classifier head receives one feature instead of 64 (aeon issue #3775). Held as evidence for that issue; the port of the same method reports as DisjointCNN.
Their results remain in the repository under results/multiverse/. Removing an estimator that cannot finish the archive returns the datasets it alone was missing to every other estimator, which is why the scored count above is larger than the number of datasets any single run completed.
Reproducing this page
from multiverse.experiments.tables import leaderboard
+Multiverse-core leaderboard
25 estimators on 57 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-24
# Estimator Accuracy Balanced accuracy AUROC F1 Log loss ↓ Sensitivity Specificity Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank 1 HC2 0.7938 7.40 0.7543 8.48 0.8951 5.95 0.7378 8.11 0.5316 5.81 0.7570 8.50 0.7987 7.11 2 MRHydra 0.7838 8.77 0.7529 8.61 0.8074 17.72 0.7336 8.65 7.7909 21.54 0.7609 9.07 0.7828 9.27 3 RDST 0.7752 9.18 0.7396 9.46 0.7968 17.94 0.7179 9.64 8.1012 21.61 0.7289 10.29 0.7891 9.00 4 RIST 0.7742 9.82 0.7430 10.46 0.8677 8.32 0.7213 10.66 0.6123 8.12 0.7469 11.11 0.7724 10.69 5 CIF 0.7775 10.16 0.7463 10.31 0.8843 8.35 0.7293 10.18 0.6414 9.11 0.7521 11.06 0.7767 10.39 6 DrCIF 0.7722 10.37 0.7402 10.90 0.8734 8.77 0.7216 11.24 0.6413 9.09 0.7437 11.71 0.7718 11.18 7 Arsenal 0.7709 10.80 0.7336 10.97 0.8469 13.25 0.7145 11.00 3.6382 17.64 0.7340 11.18 0.7792 10.65 8 QUANT 0.7651 10.92 0.7362 10.90 0.8678 8.32 0.7177 10.98 0.7263 7.63 0.7506 11.00 0.7554 12.10 9 ROCKET 0.7701 11.04 0.7345 10.81 0.7927 18.79 0.7134 11.28 8.2862 22.41 0.7315 11.74 0.7781 11.31 10 LITETime-MV 0.7483 11.24 0.7260 10.21 0.8493 9.90 0.6851 10.20 1.3395 13.12 0.7129 9.98 0.7668 11.25 11 STSF 0.7699 11.41 0.7447 10.86 0.8700 10.23 0.7145 11.54 0.6742 8.40 0.7379 11.98 0.7827 12.22 12 H-InceptionTime 0.7379 11.89 0.7137 11.15 0.8434 10.39 0.6837 11.00 1.3662 13.67 0.7192 10.89 0.7389 12.55 13 Catch22 0.7539 12.90 0.7224 13.24 0.8680 10.83 0.7027 13.46 0.6972 10.49 0.7290 13.72 0.7507 13.47 14 ConvTran 0.7446 13.12 0.7110 12.88 0.8520 10.98 0.6846 12.74 0.9070 10.11 0.7219 12.75 0.7356 13.95 15 PatchMTSC 0.7443 13.31 0.6959 13.75 0.8247 12.53 0.6682 13.30 0.7905 9.53 0.6996 12.92 0.7381 13.59 16 DisjointCNN 0.7224 13.64 0.6975 12.90 0.8239 11.61 0.6611 13.16 2.1344 15.18 0.6814 12.43 0.7299 12.82 17 STC 0.7518 13.86 0.7106 14.43 0.8637 11.74 0.6853 14.41 0.6428 9.77 0.7104 14.54 0.7599 13.31 18 TSF 0.7419 14.15 0.7138 13.65 0.8558 12.24 0.6926 13.88 0.9914 10.75 0.7141 14.51 0.7498 14.26 19 TDE 0.7272 14.94 0.6843 15.06 0.8379 12.71 0.6601 14.49 0.8450 11.46 0.6919 14.14 0.7304 14.04 20 TS2Vec 0.7192 15.41 0.6790 15.45 0.7995 15.32 0.6562 15.18 0.7869 11.58 0.6907 15.28 0.7125 15.47 21 Summary 0.6936 16.52 0.6614 16.04 0.8194 15.72 0.6393 16.16 0.9505 13.39 0.6648 16.39 0.6979 17.08 22 XCM 0.6691 16.80 0.6359 17.04 0.7915 14.94 0.5859 16.22 2.1699 16.28 0.6299 15.33 0.6769 15.62 23 TimesURL 0.6975 17.23 0.6565 17.21 0.7833 18.85 0.6206 17.59 0.9962 15.95 0.6456 17.52 0.6998 16.65 24 TimesNet 0.7001 17.41 0.6647 16.95 0.8209 16.04 0.6399 17.25 1.2957 15.30 0.6804 16.47 0.6919 17.68 25 Dummy 0.3748 22.70 0.3072 23.29 0.5000 23.56 0.1836 22.68 1.3810 17.09 0.3248 20.49 0.3774 19.34
Average score and average rank over the 57 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.
Missing results
- MRHydra — Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
- RDST — Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738)); USCActivity (Segmentation fault (core dumped))
- ConvTran — Alzheimers, EigenWorms, PhotoStimulation (CUDA out of memory)
- TS2Vec — Locust2022, Tiselac, USCActivity (Timed out at 60 hours)
Scoring uses the 57 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.
Notes on listed estimators
- XCM — run at fixed parameters, a single fit at window 0.8 with batch 32, not the per-dataset cross-validated search over window and batch size that the paper describes. The search was run and did not pay: across the 65 shared datasets it was 0.017 mean accuracy worse, 31 wins to 31 with 3 ties, Wilcoxon p = 0.63. On the 14 datasets where the search selected 0.8, the window used here, the two runs still differed by 0.11 mean absolute accuracy and by as much as 0.48, so at one resample XCM's run-to-run variance is larger than the effect the search is tuning for.
Datasets not included
- AustraliaRainfall_disc — 112186 cases, and three estimators fail on it for reasons compute cannot fix. RDST and ROCKET hit LAPACK integer overflow in RidgeClassifierCV's SVD, aeon issue 3738, after 14 and 12 attempts; MRHydra exhausted 128 GB over 13.
- PenDigits — the series are length 8 and MRHydra requires at least 9, so the dataset cannot complete while MRHydra is a column.
- InsectWingbeat — 25,000 training cases of 200 channels. Only the Dummy baseline has a result on it: none of the classifiers in the Multiverse archive paper finished it within the resource limits there, so it cannot enter a table that needs every estimator.
- BenzeneConcentration_disc — removed from Multiverse-core. The results here are on version 1, whose PT08.S2 channel is a deterministic function of the target; version 2 drops it, on the advice of the original UCI Air Quality donors: https://zenodo.org/records/21871727.
Held out of the collection rather than reported as missing, because no scheduling closes them. Results that do exist for them remain in the repository.
Estimators not listed
- LiteTIME — LITE is a univariate architecture. The multivariate variant of the same method is listed here as LITETime-MV.
- FreshPRINCE — cannot complete the archive at the memory available: recorded OOM at 128 GB after eight attempts each on FaceDetection, FordChallenge and Skoda, and 38 on Tiselac.
- 1NN-DTW — cannot complete the archive within the walltime available: exceeded the limit on BIDMC32HR_disc, with no result recorded for BIDMC32SpO2_disc.
- MUSE — cannot complete the archive as implemented: on STEW and Skoda its word bag outgrows the 32-bit sparse indices scikit-learn's classifier accepts, which no memory allocation fixes, and on USCActivity its chi2 selection needs 287 GiB. MotorImagery exhausted 16 GB and awaits a rerun at the configuration of its other results.
- CIF-500 — the 500-tree configuration run for the Multiverse archive paper's full-archive benchmark, where it is reported as CIF. This table reports the default CIF; the full-archive leaderboard reports CIF-500.
- DrCIF-500 — the 500-tree configuration HC2 uses internally, run for the Multiverse archive paper's full-archive benchmark, where it is reported as DrCIF. This table reports the default DrCIF; the full-archive leaderboard reports DrCIF-500.
- DisjointCNN-Aeon — aeon's implementation applies a Permute after the final block, so its pooling reduces the wrong axes and the classifier head receives one feature instead of 64 (aeon issue #3775). Held as evidence for that issue; the port of the same method reports as DisjointCNN.
Their results remain in the repository under results/multiverse/. Removing an estimator that cannot finish the archive returns the datasets it alone was missing to every other estimator, which is why the scored count above is larger than the number of datasets any single run completed.
Reproducing this page
from multiverse.experiments.tables import leaderboard
leaderboard(
datasets=[...] # 57 datasets,
diff --git a/results/multiverse/leaderboard_uea.html b/results/multiverse/leaderboard_uea.html
index 0a47367..158c19b 100644
--- a/results/multiverse/leaderboard_uea.html
+++ b/results/multiverse/leaderboard_uea.html
@@ -60,7 +60,7 @@
details { margin-top: .6rem; }
summary { cursor: pointer; color: var(--accent); }
code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .9em; }
-UEA leaderboard
25 estimators on 24 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-24
# Estimator Accuracy Balanced accuracy AUROC F1 Log loss ↓ Sensitivity Specificity Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank 1 HC2 0.7617 6.65 0.7412 7.56 0.8752 6.77 0.7372 7.71 0.6681 5.88 0.7429 8.21 0.7655 6.52 2 RDST 0.7407 8.27 0.7250 8.56 0.8098 18.33 0.7197 8.54 9.3448 20.83 0.7212 9.06 0.7512 8.29 3 MRHydra 0.7462 8.77 0.7332 8.62 0.8145 18.40 0.7371 8.29 9.1480 20.88 0.7526 8.33 0.7315 8.85 4 Arsenal 0.7282 9.38 0.7103 9.96 0.8375 14.60 0.7084 9.67 5.3575 17.42 0.7084 10.31 0.7380 8.73 5 H-InceptionTime 0.7208 9.56 0.7214 9.08 0.8605 8.33 0.6962 9.79 1.5037 11.29 0.7046 9.71 0.7323 10.08 6 ROCKET 0.7265 9.65 0.7101 10.04 0.8004 19.40 0.7084 10.31 9.8595 21.96 0.7084 10.96 0.7341 9.50 7 RIST 0.7374 10.04 0.7226 10.25 0.8660 8.69 0.7261 10.44 0.7933 9.71 0.7370 10.79 0.7292 10.46 8 CIF 0.7475 10.42 0.7334 10.58 0.8742 8.75 0.7356 10.35 0.8410 10.88 0.7550 10.67 0.7307 10.90 9 LITETime-MV 0.7048 10.96 0.7040 10.17 0.8505 8.58 0.6769 9.77 1.4815 10.50 0.6894 9.67 0.7181 10.65 10 DrCIF 0.7328 11.10 0.7199 11.08 0.8639 9.92 0.7193 11.98 0.8386 11.04 0.7324 11.79 0.7250 12.10 11 QUANT 0.7245 12.54 0.7136 12.29 0.8803 8.19 0.7161 12.42 0.7978 8.75 0.7383 11.88 0.7035 13.52 12 DisjointCNN 0.6943 12.62 0.6961 11.69 0.8387 9.71 0.6627 12.54 1.8839 11.92 0.6830 11.35 0.7044 11.79 13 STSF 0.7309 12.71 0.7191 12.65 0.8703 11.19 0.6914 13.17 0.8255 9.38 0.6982 13.50 0.7556 12.29 14 PatchMTSC 0.7096 13.71 0.6977 13.40 0.8549 10.60 0.6906 13.06 0.7619 7.50 0.7220 12.94 0.6874 14.58 15 TS2Vec 0.6990 13.96 0.6839 14.27 0.8334 13.85 0.6853 13.71 0.8820 11.71 0.7088 13.77 0.6783 14.31 16 TDE 0.7026 14.19 0.6818 14.27 0.8386 13.08 0.6745 13.94 1.1277 12.42 0.6877 14.10 0.7016 13.79 17 ConvTran 0.6915 14.23 0.6791 13.81 0.8493 11.44 0.6784 13.35 0.8090 9.75 0.7032 13.79 0.6725 15.06 18 TSF 0.7183 14.50 0.7052 14.21 0.8600 11.81 0.6901 14.04 0.9018 10.96 0.6962 14.48 0.7324 14.10 19 STC 0.7224 14.62 0.7004 15.04 0.8722 11.33 0.7000 14.75 0.8052 10.25 0.7138 15.29 0.7157 13.85 20 Catch22 0.6945 14.98 0.6800 15.48 0.8443 12.73 0.6836 15.19 0.9690 13.46 0.7021 14.83 0.6762 15.44 21 XCM 0.6356 15.98 0.6204 16.52 0.8184 13.79 0.5899 16.69 1.3507 13.75 0.6172 15.44 0.6410 15.75 22 TimesURL 0.6745 17.04 0.6600 16.85 0.8170 18.29 0.6473 17.12 1.3326 17.46 0.6612 17.62 0.6703 16.27 23 TimesNet 0.6582 17.92 0.6505 17.21 0.8275 16.15 0.6400 17.67 1.1485 13.54 0.6648 17.17 0.6481 18.77 24 Summary 0.6431 18.06 0.6314 17.98 0.8181 17.38 0.6179 17.79 1.2745 16.00 0.6269 17.62 0.6521 17.75 25 Dummy 0.2286 23.15 0.2106 23.42 0.5000 23.69 0.1044 22.71 1.8615 17.79 0.2193 21.71 0.2193 21.62
Average score and average rank over the 24 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.
Missing results
- CIF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- DrCIF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- DisjointCNN — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- STSF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- TS2Vec — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- ConvTran — EigenWorms (CUDA out of memory)
- TSF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- XCM — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- Summary — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
Scoring uses the 24 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.
1 requested dataset(s) have no results from any estimator: InsectWingbeat.
Notes on listed estimators
- XCM — run at fixed parameters, a single fit at window 0.8 with batch 32, not the per-dataset cross-validated search over window and batch size that the paper describes. The search was run and did not pay: across the 65 shared datasets it was 0.017 mean accuracy worse, 31 wins to 31 with 3 ties, Wilcoxon p = 0.63. On the 14 datasets where the search selected 0.8, the window used here, the two runs still differed by 0.11 mean absolute accuracy and by as much as 0.48, so at one resample XCM's run-to-run variance is larger than the effect the search is tuning for.
Datasets not included
- AustraliaRainfall_disc — 112186 cases, and three estimators fail on it for reasons compute cannot fix. RDST and ROCKET hit LAPACK integer overflow in RidgeClassifierCV's SVD, aeon issue 3738, after 14 and 12 attempts; MRHydra exhausted 128 GB over 13.
- PenDigits — the series are length 8 and MRHydra requires at least 9, so the dataset cannot complete while MRHydra is a column.
- BenzeneConcentration_disc — removed from Multiverse-core. The results here are on version 1, whose PT08.S2 channel is a deterministic function of the target; version 2 drops it, on the advice of the original UCI Air Quality donors: https://zenodo.org/records/21871727.
Held out of the collection rather than reported as missing, because no scheduling closes them. Results that do exist for them remain in the repository.
Estimators not listed
- LiteTIME — LITE is a univariate architecture. The multivariate variant of the same method is listed here as LITETime-MV.
- FreshPRINCE — cannot complete the archive at the memory available: recorded OOM at 128 GB after eight attempts each on FaceDetection, FordChallenge and Skoda, and 38 on Tiselac.
- 1NN-DTW — cannot complete the archive within the walltime available: exceeded the limit on BIDMC32HR_disc, with no result recorded for BIDMC32SpO2_disc.
- MUSE — cannot complete the archive as implemented: on STEW and Skoda its word bag outgrows the 32-bit sparse indices scikit-learn's classifier accepts, which no memory allocation fixes, and on USCActivity its chi2 selection needs 287 GiB. MotorImagery exhausted 16 GB and awaits a rerun at the configuration of its other results.
- CIF-500 — the 500-tree configuration run for the Multiverse archive paper's full-archive benchmark, where it is reported as CIF. This table reports the default CIF; the full-archive leaderboard reports CIF-500.
- DrCIF-500 — the 500-tree configuration HC2 uses internally, run for the Multiverse archive paper's full-archive benchmark, where it is reported as DrCIF. This table reports the default DrCIF; the full-archive leaderboard reports DrCIF-500.
- DisjointCNN-Aeon — aeon's implementation applies a Permute after the final block, so its pooling reduces the wrong axes and the classifier head receives one feature instead of 64 (aeon issue #3775). Held as evidence for that issue; the port of the same method reports as DisjointCNN.
Their results remain in the repository under results/multiverse/. Removing an estimator that cannot finish the archive returns the datasets it alone was missing to every other estimator, which is why the scored count above is larger than the number of datasets any single run completed.
Reproducing this page
from multiverse.experiments.tables import leaderboard
+UEA leaderboard
25 estimators on 24 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-24
# Estimator Accuracy Balanced accuracy AUROC F1 Log loss ↓ Sensitivity Specificity Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank Score Rank 1 HC2 0.7617 6.65 0.7412 7.56 0.8752 6.77 0.7372 7.71 0.6681 5.88 0.7429 8.21 0.7655 6.52 2 RDST 0.7407 8.27 0.7250 8.56 0.8098 18.33 0.7197 8.54 9.3448 20.83 0.7212 9.06 0.7512 8.29 3 MRHydra 0.7462 8.77 0.7332 8.62 0.8145 18.40 0.7371 8.29 9.1480 20.88 0.7526 8.33 0.7315 8.85 4 Arsenal 0.7282 9.38 0.7103 9.96 0.8375 14.60 0.7084 9.67 5.3575 17.42 0.7084 10.31 0.7380 8.73 5 H-InceptionTime 0.7208 9.56 0.7214 9.08 0.8605 8.33 0.6962 9.79 1.5037 11.29 0.7046 9.71 0.7323 10.08 6 ROCKET 0.7265 9.65 0.7101 10.04 0.8004 19.40 0.7084 10.31 9.8595 21.96 0.7084 10.96 0.7341 9.50 7 RIST 0.7374 10.04 0.7226 10.25 0.8660 8.69 0.7261 10.44 0.7933 9.71 0.7370 10.79 0.7292 10.46 8 CIF 0.7475 10.42 0.7334 10.58 0.8742 8.75 0.7356 10.35 0.8410 10.88 0.7550 10.67 0.7307 10.90 9 LITETime-MV 0.7048 10.96 0.7040 10.17 0.8505 8.58 0.6769 9.77 1.4815 10.50 0.6894 9.67 0.7181 10.65 10 DrCIF 0.7328 11.10 0.7199 11.08 0.8639 9.92 0.7193 11.98 0.8386 11.04 0.7324 11.79 0.7250 12.10 11 QUANT 0.7245 12.54 0.7136 12.29 0.8803 8.19 0.7161 12.42 0.7978 8.75 0.7383 11.88 0.7035 13.52 12 DisjointCNN 0.6943 12.62 0.6961 11.69 0.8387 9.71 0.6627 12.54 1.8839 11.92 0.6830 11.35 0.7044 11.79 13 STSF 0.7309 12.71 0.7191 12.65 0.8703 11.19 0.6914 13.17 0.8255 9.38 0.6982 13.50 0.7556 12.29 14 PatchMTSC 0.7096 13.71 0.6977 13.40 0.8549 10.60 0.6906 13.06 0.7619 7.50 0.7220 12.94 0.6874 14.58 15 TS2Vec 0.6990 13.96 0.6839 14.27 0.8334 13.85 0.6853 13.71 0.8820 11.71 0.7088 13.77 0.6783 14.31 16 TDE 0.7026 14.19 0.6818 14.27 0.8386 13.08 0.6745 13.94 1.1277 12.42 0.6877 14.10 0.7016 13.79 17 ConvTran 0.6915 14.23 0.6791 13.81 0.8493 11.44 0.6784 13.35 0.8090 9.75 0.7032 13.79 0.6725 15.06 18 TSF 0.7183 14.50 0.7052 14.21 0.8600 11.81 0.6901 14.04 0.9018 10.96 0.6962 14.48 0.7324 14.10 19 STC 0.7224 14.62 0.7004 15.04 0.8722 11.33 0.7000 14.75 0.8052 10.25 0.7138 15.29 0.7157 13.85 20 Catch22 0.6945 14.98 0.6800 15.48 0.8443 12.73 0.6836 15.19 0.9690 13.46 0.7021 14.83 0.6762 15.44 21 XCM 0.6356 15.98 0.6204 16.52 0.8184 13.79 0.5899 16.69 1.3507 13.75 0.6172 15.44 0.6410 15.75 22 TimesURL 0.6745 17.04 0.6600 16.85 0.8170 18.29 0.6473 17.12 1.3326 17.46 0.6612 17.62 0.6703 16.27 23 TimesNet 0.6582 17.92 0.6505 17.21 0.8275 16.15 0.6400 17.67 1.1485 13.54 0.6648 17.17 0.6481 18.77 24 Summary 0.6431 18.06 0.6314 17.98 0.8181 17.38 0.6179 17.79 1.2745 16.00 0.6269 17.62 0.6521 17.75 25 Dummy 0.2286 23.15 0.2106 23.42 0.5000 23.69 0.1044 22.71 1.8615 17.79 0.2193 21.71 0.2193 21.62
Average score and average rank over the 24 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.
Missing results
- CIF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- DrCIF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- DisjointCNN — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- STSF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- TS2Vec — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- ConvTran — EigenWorms (CUDA out of memory)
- TSF — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- XCM — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
- Summary — BasicMotions, FingerMovements, SelfRegulationSCP2 (not run outside Multiverse-core)
Scoring uses the 24 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.
Notes on listed estimators
- XCM — run at fixed parameters, a single fit at window 0.8 with batch 32, not the per-dataset cross-validated search over window and batch size that the paper describes. The search was run and did not pay: across the 65 shared datasets it was 0.017 mean accuracy worse, 31 wins to 31 with 3 ties, Wilcoxon p = 0.63. On the 14 datasets where the search selected 0.8, the window used here, the two runs still differed by 0.11 mean absolute accuracy and by as much as 0.48, so at one resample XCM's run-to-run variance is larger than the effect the search is tuning for.
Datasets not included
- AustraliaRainfall_disc — 112186 cases, and three estimators fail on it for reasons compute cannot fix. RDST and ROCKET hit LAPACK integer overflow in RidgeClassifierCV's SVD, aeon issue 3738, after 14 and 12 attempts; MRHydra exhausted 128 GB over 13.
- PenDigits — the series are length 8 and MRHydra requires at least 9, so the dataset cannot complete while MRHydra is a column.
- InsectWingbeat — 25,000 training cases of 200 channels. Only the Dummy baseline has a result on it: none of the classifiers in the Multiverse archive paper finished it within the resource limits there, so it cannot enter a table that needs every estimator.
- BenzeneConcentration_disc — removed from Multiverse-core. The results here are on version 1, whose PT08.S2 channel is a deterministic function of the target; version 2 drops it, on the advice of the original UCI Air Quality donors: https://zenodo.org/records/21871727.
Held out of the collection rather than reported as missing, because no scheduling closes them. Results that do exist for them remain in the repository.
Estimators not listed
- LiteTIME — LITE is a univariate architecture. The multivariate variant of the same method is listed here as LITETime-MV.
- FreshPRINCE — cannot complete the archive at the memory available: recorded OOM at 128 GB after eight attempts each on FaceDetection, FordChallenge and Skoda, and 38 on Tiselac.
- 1NN-DTW — cannot complete the archive within the walltime available: exceeded the limit on BIDMC32HR_disc, with no result recorded for BIDMC32SpO2_disc.
- MUSE — cannot complete the archive as implemented: on STEW and Skoda its word bag outgrows the 32-bit sparse indices scikit-learn's classifier accepts, which no memory allocation fixes, and on USCActivity its chi2 selection needs 287 GiB. MotorImagery exhausted 16 GB and awaits a rerun at the configuration of its other results.
- CIF-500 — the 500-tree configuration run for the Multiverse archive paper's full-archive benchmark, where it is reported as CIF. This table reports the default CIF; the full-archive leaderboard reports CIF-500.
- DrCIF-500 — the 500-tree configuration HC2 uses internally, run for the Multiverse archive paper's full-archive benchmark, where it is reported as DrCIF. This table reports the default DrCIF; the full-archive leaderboard reports DrCIF-500.
- DisjointCNN-Aeon — aeon's implementation applies a Permute after the final block, so its pooling reduces the wrong axes and the classifier head receives one feature instead of 64 (aeon issue #3775). Held as evidence for that issue; the port of the same method reports as DisjointCNN.
Their results remain in the repository under results/multiverse/. Removing an estimator that cannot finish the archive returns the datasets it alone was missing to every other estimator, which is why the scored count above is larger than the number of datasets any single run completed.
Reproducing this page
from multiverse.experiments.tables import leaderboard
leaderboard(
datasets=[...] # 24 datasets,