From f43fe3fb597792589b915a02215c3f77b4050ed7 Mon Sep 17 00:00:00 2001 From: Tony Bagnall Date: Sun, 6 Sep 2026 10:49:58 +0100 Subject: [PATCH 1/7] Withhold LiteTIME, FreshPRINCE and 1NN-DTW from the tables LiteTIME is the univariate LITE architecture; the multivariate variant of the same method already reports as LITETime-MV. FreshPRINCE and 1NN-DTW cannot finish the archive on the resources available, and because scoring uses the datasets every estimator completed, one estimator's gaps are subtracted from everyone's table. Removing the three returns five datasets to the scored set: BIDMC32HR_disc, BIDMC32SpO2_disc, FaceDetection, FordChallenge and Skoda. The Multiverse-core leaderboard goes from 27 estimators on 51 datasets to 24 on 56. The four withheld estimators are now named on every page under "Estimators not listed", each with its reason, rather than being absent without explanation. The exclusion list and those reasons are one structure, WITHHELD_ESTIMATORS, so the page cannot drift from what main() actually excludes. Co-Authored-By: Claude Opus 5 --- README.md | 55 ++++++++++---------- docs/leaderboard.md | 55 ++++++++++---------- multiverse/experiments/tables.py | 69 +++++++++++++++++++++---- results/multiverse/datasets.html | 2 +- results/multiverse/leaderboard.html | 4 +- results/multiverse/leaderboard_uea.html | 4 +- 6 files changed, 116 insertions(+), 73 deletions(-) diff --git a/README.md b/README.md index 8ee3cdc..d0cdbee 100644 --- a/README.md +++ b/README.md @@ -40,35 +40,32 @@ The current paper version describes: | # | Estimator | Accuracy rank | Accuracy | Balanced accuracy | AUROC | F1 | Log loss ↓ | Sensitivity | Specificity | |---|---|---|---|---|---|---|---|---|---| -| 1 | HC2 | **8.40** | **0.7887** | 0.7541 | **0.9000** | 0.7346 | **0.5440** | 0.7547 | **0.7910** | -| 2 | MRHydra | 9.21 | 0.7810 | **0.7579** | 0.8130 | **0.7368** | 7.8942 | **0.7715** | 0.7718 | -| 3 | RDST | 10.10 | 0.7707 | 0.7372 | 0.7963 | 0.7105 | 8.2660 | 0.7236 | 0.7833 | -| 4 | RIST | 10.75 | 0.7693 | 0.7422 | 0.8755 | 0.7221 | 0.6294 | 0.7504 | 0.7613 | -| 5 | DrCIF | 10.87 | 0.7721 | 0.7454 | 0.8821 | 0.7248 | 0.6558 | 0.7490 | 0.7669 | -| 6 | CIF | 11.07 | 0.7756 | 0.7497 | 0.8920 | 0.7288 | 0.6497 | 0.7536 | 0.7714 | -| 7 | FreshPRINCE | 11.08 | 0.7717 | 0.7516 | 0.8752 | 0.7293 | 0.6075 | 0.7515 | 0.7731 | -| 8 | Arsenal | 11.42 | 0.7654 | 0.7340 | 0.8471 | 0.7092 | 3.9265 | 0.7337 | 0.7696 | -| 9 | QUANT | 11.65 | 0.7693 | 0.7486 | 0.8839 | 0.7262 | 0.6238 | 0.7616 | 0.7539 | -| 10 | LITETime-MV | 11.97 | 0.7476 | 0.7312 | 0.8518 | 0.6875 | 1.3300 | 0.7200 | 0.7600 | -| 11 | ROCKET | 12.01 | 0.7661 | 0.7345 | 0.7955 | 0.7080 | 8.4299 | 0.7282 | 0.7724 | -| 12 | STSF | 12.60 | 0.7698 | 0.7503 | 0.8813 | 0.7155 | 0.6493 | 0.7439 | 0.7790 | -| 13 | H-InceptionTime | 12.71 | 0.7375 | 0.7205 | 0.8506 | 0.6897 | 1.3334 | 0.7303 | 0.7333 | -| 14 | LiteTIME | 13.22 | 0.7308 | 0.7122 | 0.8402 | 0.6746 | 1.4921 | 0.7199 | 0.7291 | -| 15 | DisjointCNN | 13.63 | 0.7286 | 0.7061 | 0.8354 | 0.6688 | 1.9705 | 0.6889 | 0.7368 | -| 16 | ConvTran | 14.37 | 0.7430 | 0.7139 | 0.8606 | 0.6882 | 0.8300 | 0.7289 | 0.7295 | -| 17 | Catch22 | 14.50 | 0.7442 | 0.7203 | 0.8703 | 0.6996 | 0.7238 | 0.7337 | 0.7326 | -| 18 | PatchMTSC | 14.51 | 0.7395 | 0.6934 | 0.8288 | 0.6660 | 0.7748 | 0.6985 | 0.7300 | -| 19 | STC | 15.19 | 0.7516 | 0.7188 | 0.8748 | 0.7004 | 0.6447 | 0.7264 | 0.7496 | -| 20 | TSF | 15.36 | 0.7484 | 0.7257 | 0.8747 | 0.6952 | 0.7335 | 0.7179 | 0.7565 | -| 21 | TS2Vec | 15.87 | 0.7212 | 0.6849 | 0.8082 | 0.6588 | 0.7326 | 0.6980 | 0.7100 | -| 22 | TDE | 15.93 | 0.7230 | 0.6823 | 0.8383 | 0.6441 | 0.8859 | 0.6786 | 0.7301 | -| 23 | Summary | 18.54 | 0.6814 | 0.6586 | 0.8263 | 0.6294 | 0.9251 | 0.6661 | 0.6787 | -| 24 | TimesNet | 18.86 | 0.6971 | 0.6688 | 0.8280 | 0.6390 | 1.1785 | 0.6850 | 0.6838 | -| 25 | TimesURL | 18.95 | 0.6916 | 0.6563 | 0.7931 | 0.6084 | 1.0193 | 0.6379 | 0.6914 | -| 26 | 1NN-DTW | 20.47 | 0.6672 | 0.6457 | 0.7214 | 0.6193 | 11.9949 | 0.6584 | 0.6584 | -| 27 | Dummy | 24.77 | 0.3538 | 0.2991 | 0.5000 | 0.1537 | 1.4284 | 0.2911 | 0.3695 | - -Average over the 51 Multiverse-core datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold. +| 1 | HC2 | **7.45** | **0.7917** | **0.7557** | **0.8935** | **0.7302** | **0.5350** | 0.7469 | **0.7998** | +| 2 | MRHydra | 8.89 | 0.7794 | 0.7520 | 0.8040 | 0.7266 | 7.9526 | **0.7577** | 0.7768 | +| 3 | RDST | 9.29 | 0.7729 | 0.7386 | 0.7928 | 0.7075 | 8.1867 | 0.7172 | 0.7902 | +| 4 | RIST | 9.60 | 0.7744 | 0.7451 | 0.8679 | 0.7174 | 0.6150 | 0.7416 | 0.7743 | +| 5 | CIF | 9.97 | 0.7770 | 0.7487 | 0.8842 | 0.7246 | 0.6442 | 0.7459 | 0.7780 | +| 6 | DrCIF | 9.97 | 0.7731 | 0.7433 | 0.8745 | 0.7189 | 0.6430 | 0.7400 | 0.7736 | +| 7 | QUANT | 10.43 | 0.7668 | 0.7404 | 0.8694 | 0.7166 | 0.7285 | 0.7491 | 0.7570 | +| 8 | LITETime-MV | 10.70 | 0.7511 | 0.7320 | 0.8503 | 0.6878 | 1.3004 | 0.7167 | 0.7660 | +| 9 | Arsenal | 10.75 | 0.7663 | 0.7335 | 0.8419 | 0.7061 | 3.6444 | 0.7266 | 0.7752 | +| 10 | ROCKET | 10.83 | 0.7690 | 0.7362 | 0.7925 | 0.7065 | 8.3274 | 0.7228 | 0.7798 | +| 11 | STSF | 11.01 | 0.7727 | 0.7508 | 0.8723 | 0.7164 | 0.6685 | 0.7412 | 0.7845 | +| 12 | H-InceptionTime | 11.38 | 0.7421 | 0.7208 | 0.8447 | 0.6853 | 1.3448 | 0.7227 | 0.7436 | +| 13 | ConvTran | 12.54 | 0.7490 | 0.7177 | 0.8529 | 0.6862 | 0.8826 | 0.7234 | 0.7419 | +| 14 | DisjointCNN | 12.69 | 0.7296 | 0.7057 | 0.8246 | 0.6641 | 2.1100 | 0.6815 | 0.7431 | +| 15 | PatchMTSC | 13.01 | 0.7454 | 0.6990 | 0.8250 | 0.6671 | 0.7818 | 0.6981 | 0.7397 | +| 16 | Catch22 | 13.07 | 0.7463 | 0.7177 | 0.8605 | 0.6929 | 0.7068 | 0.7229 | 0.7420 | +| 17 | TSF | 13.67 | 0.7426 | 0.7175 | 0.8571 | 0.6896 | 0.9987 | 0.7095 | 0.7521 | +| 18 | STC | 13.68 | 0.7507 | 0.7137 | 0.8624 | 0.6803 | 0.6468 | 0.7036 | 0.7611 | +| 19 | TDE | 14.60 | 0.7251 | 0.6834 | 0.8339 | 0.6446 | 0.8524 | 0.6759 | 0.7349 | +| 20 | TS2Vec | 14.79 | 0.7201 | 0.6809 | 0.7994 | 0.6527 | 0.7835 | 0.6879 | 0.7144 | +| 21 | Summary | 16.37 | 0.6845 | 0.6570 | 0.8113 | 0.6299 | 0.9645 | 0.6621 | 0.6852 | +| 22 | TimesNet | 16.67 | 0.7020 | 0.6709 | 0.8218 | 0.6396 | 1.2889 | 0.6810 | 0.6934 | +| 23 | TimesURL | 16.79 | 0.6950 | 0.6562 | 0.7831 | 0.6052 | 0.9868 | 0.6312 | 0.7016 | +| 24 | Dummy | 21.87 | 0.3709 | 0.3067 | 0.5000 | 0.1626 | 1.3928 | 0.2987 | 0.3880 | + +Average over the 56 Multiverse-core datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold. Rebuilt with `python -m multiverse.experiments.tables`, which also writes a sortable diff --git a/docs/leaderboard.md b/docs/leaderboard.md index 2318f21..6614cc8 100644 --- a/docs/leaderboard.md +++ b/docs/leaderboard.md @@ -109,35 +109,32 @@ inferred from the ranking. | # | Estimator | Accuracy rank | Accuracy | Balanced accuracy | AUROC | F1 | Log loss ↓ | Sensitivity | Specificity | |---|---|---|---|---|---|---|---|---|---| -| 1 | HC2 | **7.37** | **0.7665** | **0.7452** | 0.8823 | 0.7411 | **0.6692** | 0.7470 | **0.7703** | -| 2 | RDST | 8.80 | 0.7459 | 0.7294 | 0.8179 | 0.7243 | 9.1587 | 0.7263 | 0.7560 | -| 3 | MRHydra | 9.20 | 0.7523 | 0.7388 | 0.8236 | **0.7432** | 8.9285 | 0.7599 | 0.7360 | -| 4 | Arsenal | 10.26 | 0.7321 | 0.7134 | 0.8439 | 0.7116 | 5.5624 | 0.7119 | 0.7417 | -| 5 | ROCKET | 10.28 | 0.7317 | 0.7146 | 0.8089 | 0.7130 | 9.6705 | 0.7133 | 0.7393 | -| 6 | RIST | 10.57 | 0.7433 | 0.7278 | 0.8755 | 0.7325 | 0.7983 | 0.7454 | 0.7322 | -| 7 | H-InceptionTime | 10.67 | 0.7223 | 0.7230 | 0.8653 | 0.6967 | 1.5030 | 0.7053 | 0.7345 | -| 8 | CIF | 11.33 | 0.7525 | 0.7378 | 0.8825 | 0.7400 | 0.8488 | **0.7604** | 0.7349 | -| 9 | FreshPRINCE | 11.74 | 0.7422 | 0.7281 | 0.8796 | 0.7239 | 0.7764 | 0.7313 | 0.7457 | -| 10 | DrCIF | 11.83 | 0.7386 | 0.7252 | 0.8734 | 0.7246 | 0.8458 | 0.7384 | 0.7303 | -| 11 | LITETime-MV | 12.00 | 0.7073 | 0.7064 | 0.8568 | 0.6779 | 1.4779 | 0.6905 | 0.7218 | -| 12 | LiteTIME | 12.50 | 0.7087 | 0.7019 | 0.8576 | 0.6751 | 1.6854 | 0.6985 | 0.7204 | -| 13 | DisjointCNN | 13.20 | 0.7011 | 0.7030 | 0.8510 | 0.6704 | 1.7943 | 0.6938 | 0.7070 | -| 14 | QUANT | 13.78 | 0.7285 | 0.7171 | **0.8888** | 0.7195 | 0.8041 | 0.7421 | 0.7074 | -| 15 | STSF | 14.24 | 0.7345 | 0.7223 | 0.8774 | 0.6934 | 0.8338 | 0.7007 | 0.7600 | -| 16 | TS2Vec | 14.78 | 0.7070 | 0.6913 | 0.8470 | 0.6917 | 0.8902 | 0.7150 | 0.6877 | -| 17 | TDE | 15.15 | 0.7079 | 0.6862 | 0.8484 | 0.6775 | 1.1475 | 0.6897 | 0.7095 | -| 18 | PatchMTSC | 15.37 | 0.7110 | 0.6986 | 0.8601 | 0.6899 | 0.7670 | 0.7192 | 0.6928 | -| 19 | ConvTran | 16.11 | 0.6931 | 0.6801 | 0.8552 | 0.6793 | 0.8155 | 0.7049 | 0.6736 | -| 20 | STC | 16.15 | 0.7265 | 0.7036 | 0.8803 | 0.7035 | 0.8124 | 0.7186 | 0.7184 | -| 21 | TSF | 16.17 | 0.7214 | 0.7076 | 0.8671 | 0.6917 | 0.9127 | 0.6977 | 0.7365 | -| 22 | Catch22 | 16.30 | 0.7006 | 0.6854 | 0.8557 | 0.6897 | 0.9814 | 0.7096 | 0.6802 | -| 23 | 1NN-DTW | 17.61 | 0.6848 | 0.6759 | 0.7785 | 0.6702 | 11.3600 | 0.6720 | 0.6879 | -| 24 | TimesURL | 18.24 | 0.6809 | 0.6658 | 0.8290 | 0.6539 | 1.3557 | 0.6698 | 0.6738 | -| 25 | Summary | 19.48 | 0.6477 | 0.6355 | 0.8295 | 0.6206 | 1.3000 | 0.6291 | 0.6589 | -| 26 | TimesNet | 20.09 | 0.6584 | 0.6504 | 0.8332 | 0.6386 | 1.1641 | 0.6628 | 0.6504 | -| 27 | Dummy | 24.78 | 0.2168 | 0.1980 | 0.5000 | 0.0800 | 1.9123 | 0.1853 | 0.2288 | - -Average over the 23 UEA datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold. +| 1 | HC2 | **6.48** | **0.7617** | **0.7412** | 0.8752 | **0.7372** | **0.6681** | 0.7429 | **0.7655** | +| 2 | RDST | 8.08 | 0.7407 | 0.7250 | 0.8098 | 0.7197 | 9.3448 | 0.7212 | 0.7512 | +| 3 | MRHydra | 8.50 | 0.7462 | 0.7332 | 0.8145 | 0.7371 | 9.1480 | 0.7526 | 0.7315 | +| 4 | Arsenal | 9.08 | 0.7282 | 0.7103 | 0.8375 | 0.7084 | 5.3575 | 0.7084 | 0.7380 | +| 5 | ROCKET | 9.38 | 0.7265 | 0.7101 | 0.8004 | 0.7084 | 9.8595 | 0.7084 | 0.7341 | +| 6 | H-InceptionTime | 9.42 | 0.7208 | 0.7214 | 0.8605 | 0.6962 | 1.5037 | 0.7046 | 0.7323 | +| 7 | RIST | 9.73 | 0.7374 | 0.7226 | 0.8660 | 0.7261 | 0.7933 | 0.7370 | 0.7292 | +| 8 | CIF | 10.15 | 0.7475 | 0.7334 | 0.8742 | 0.7356 | 0.8410 | **0.7550** | 0.7307 | +| 9 | LITETime-MV | 10.67 | 0.7048 | 0.7040 | 0.8505 | 0.6769 | 1.4815 | 0.6894 | 0.7181 | +| 10 | DrCIF | 10.73 | 0.7328 | 0.7199 | 0.8639 | 0.7193 | 0.8386 | 0.7324 | 0.7250 | +| 11 | QUANT | 12.17 | 0.7245 | 0.7136 | **0.8803** | 0.7161 | 0.7978 | 0.7383 | 0.7035 | +| 12 | DisjointCNN | 12.27 | 0.6943 | 0.6961 | 0.8387 | 0.6627 | 1.8839 | 0.6830 | 0.7044 | +| 13 | STSF | 12.33 | 0.7309 | 0.7191 | 0.8703 | 0.6914 | 0.8255 | 0.6982 | 0.7556 | +| 14 | PatchMTSC | 13.31 | 0.7096 | 0.6977 | 0.8549 | 0.6906 | 0.7619 | 0.7220 | 0.6874 | +| 15 | TS2Vec | 13.58 | 0.6990 | 0.6839 | 0.8334 | 0.6853 | 0.8820 | 0.7088 | 0.6783 | +| 16 | TDE | 13.81 | 0.7026 | 0.6818 | 0.8386 | 0.6745 | 1.1277 | 0.6877 | 0.7016 | +| 17 | ConvTran | 13.85 | 0.6915 | 0.6791 | 0.8493 | 0.6784 | 0.8090 | 0.7032 | 0.6725 | +| 18 | TSF | 14.12 | 0.7183 | 0.7052 | 0.8600 | 0.6901 | 0.9018 | 0.6962 | 0.7324 | +| 19 | STC | 14.29 | 0.7224 | 0.7004 | 0.8722 | 0.7000 | 0.8052 | 0.7138 | 0.7157 | +| 20 | Catch22 | 14.50 | 0.6945 | 0.6800 | 0.8443 | 0.6836 | 0.9690 | 0.7021 | 0.6762 | +| 21 | TimesURL | 16.48 | 0.6745 | 0.6600 | 0.8170 | 0.6473 | 1.3326 | 0.6612 | 0.6703 | +| 22 | TimesNet | 17.35 | 0.6582 | 0.6505 | 0.8275 | 0.6400 | 1.1485 | 0.6648 | 0.6481 | +| 23 | Summary | 17.48 | 0.6431 | 0.6314 | 0.8181 | 0.6179 | 1.2745 | 0.6269 | 0.6521 | +| 24 | Dummy | 22.23 | 0.2286 | 0.2106 | 0.5000 | 0.1044 | 1.8615 | 0.2193 | 0.2193 | + +Average over the 24 UEA datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold. Sortable version with per-metric ranks: diff --git a/multiverse/experiments/tables.py b/multiverse/experiments/tables.py index d00508b..9784dca 100644 --- a/multiverse/experiments/tables.py +++ b/multiverse/experiments/tables.py @@ -344,6 +344,51 @@ def _missing_by_estimator(frames, estimators, common, datasets): return missing +# Estimators held out of the published tables, with the reason each is held out. +# They are named on every page rather than quietly dropped, since an unexplained +# absence is the thing this archive exists to argue against. +WITHHELD_ESTIMATORS = { + "LiteTIME": ( + "LITE is a univariate architecture. The multivariate variant of the same " + "method is listed here as LITETime-MV" + ), + "FreshPRINCE": ( + "cannot complete the archive at the memory available: recorded OOM at " + "128 GB after eight attempts each on FaceDetection, FordChallenge and " + "Skoda, and 38 on Tiselac" + ), + "1NN-DTW": ( + "cannot complete the archive within the walltime available: exceeded the " + "limit on BIDMC32HR_disc, with no result recorded for BIDMC32SpO2_disc" + ), + "DisjointCNN-Aeon": ( + "aeon's implementation applies a Permute after the final block, so its " + "pooling reduces the wrong axes and the classifier head receives one " + "feature instead of 64 (aeon issue #3775). Held as evidence for that " + "issue; the port of the same method reports as DisjointCNN" + ), +} + + +def _withheld_html() -> str: + """Name the estimators kept out of the table, and why.""" + if not WITHHELD_ESTIMATORS: + return "" + items = "".join( + f"
  • {escape(name)} — {escape(reason)}.
  • " + for name, reason in WITHHELD_ESTIMATORS.items() + ) + return ( + "

    Estimators not listed

    " + f'' + '

    Their results remain in the repository under ' + "results/multiverse/. Removing an estimator that cannot " + "finish the archive returns the datasets it alone was missing to every " + "other estimator, which is why the scored count above is larger than " + "the number of datasets any single run completed.

    " + ) + + def _excluded_html(missing, reasons, common, dropped) -> str: """Render a one-line summary of what each estimator is missing.""" lookup = { @@ -711,6 +756,7 @@ def leaderboard( ) parts.append(_excluded_html(missing, reasons, common, dropped)) + parts.append(_withheld_html()) parts.append( _snippet_html( datasets_expr if datasets_expr is not None else _describe_datasets(datasets), @@ -1150,20 +1196,23 @@ def main() -> None: """Build the Multiverse-core leaderboard. Uses every estimator with results in the repository, including the Dummy - baseline, over the Multiverse-core datasets all of them have results for. - - DisjointCNN-Aeon is held back. Those results are around 20 accuracy points - below the authors' published numbers on all 23 shared datasets, because - aeon's network applies a Permute after the final block and its pooling then - reduces the wrong axes, leaving the classifier head one feature instead of - 64 (aeon issue #3775). They are kept as evidence for that issue rather than - deleted, but listing them would read as a claim about the method. The - Multiverse port of the same method reports under DisjointCNN. + baseline, over the Multiverse-core datasets all of them have results for, + except those named in WITHHELD_ESTIMATORS. That set is rendered onto every + page by _withheld_html, so an omission is stated rather than inferred. + + Three of the four are held back because they cannot finish the archive, and + scoring on the intersection makes an estimator's gaps everyone's: LiteTIME + is univariate, with LITETime-MV the multivariate variant of the same method, + while FreshPRINCE and 1NN-DTW exhaust the available memory and walltime + respectively. Removing them returns five datasets to the scored set. The + fourth, DisjointCNN-Aeon, completes the archive but scores around 20 + accuracy points below the published numbers because of aeon issue #3775; it + is kept as evidence for that issue, and the port reports as DisjointCNN. """ from aeon.datasets.tsc_datasets import UEA, multiverse_core datasets = sorted(multiverse_core) - estimators = available_estimators(exclude=("DisjointCNN-Aeon",)) + estimators = available_estimators(exclude=tuple(WITHHELD_ESTIMATORS)) print(f"estimators: {', '.join(estimators)}") path = leaderboard( diff --git a/results/multiverse/datasets.html b/results/multiverse/datasets.html index f024c1e..01a1239 100644 --- a/results/multiverse/datasets.html +++ b/results/multiverse/datasets.html @@ -60,7 +60,7 @@ details { margin-top: .6rem; } summary { cursor: pointer; color: var(--accent); } code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .9em; } -tr.nosignal td { background: rgba(214, 158, 46, .16); }tr.saturated td { background: rgba(56, 161, 105, .14); }

    Multiverse-core datasets: accuracy

    66 datasets · accuracy · best of up to 26 estimators against the Dummy baseline · built 2026-09-05

    DatasetDummyMedianBestBest estimatorGain over dummySpreadEstimators
    KINECAL-QSEO0.94120.94120.9412Arsenal0.00000.117626
    BIDMC32SpO2_disc0.71530.66190.7203ROCKET0.00500.190124
    Locust20220.91120.90820.9206MRHydra0.00940.045225
    Heartbeat0.72200.74390.7854CIF0.06340.126826
    HouseholdPowerConsumption2_disc0.72160.76750.7872DisjointCNN0.06560.141426
    AutomotiveRoadTrials0.75320.79220.8442CIF0.09090.233826
    AustraliaRainfall_disc0.68600.77310.7808LITETime-MV0.09480.088117
    EyesOpenShut0.50000.50000.5952STSF0.09520.190526
    MotorImagery0.50000.50500.6000FreshPRINCE0.10000.140026
    BeijingPM10Quality_disc0.71120.82420.8417FreshPRINCE0.13050.107026
    Alzheimers0.41860.37210.5581MRHydra0.13950.302325
    AppliancesEnergy_disc0.80950.82140.9524FreshPRINCE0.14290.452426
    EmoPain0.78310.84080.92681NN-DTW0.14370.242320
    PhotoStimulation0.41670.38890.5833ROCKET0.16670.388925
    FaceDetection0.50000.62830.6850H-InceptionTime0.18500.170525
    BeijingPM25Quality_disc0.69770.87690.8879ConvTran0.19020.128826
    AtrialFibrillation0.33330.26670.5333TS2Vec0.20000.466726
    HouseholdPowerConsumption1_disc0.77840.91180.9825FreshPRINCE0.20410.218726
    LowCost0.50000.63170.7300TSF0.23000.248326
    BoneProbAgeGroup0.47640.64380.7124H-InceptionTime0.23600.193326
    StandWalkJump0.33330.40000.6000MRHydra0.26670.400026
    CrowdSourced0.50020.71430.7734LITETime-MV0.27320.176826
    BenzeneConcentration_disc0.68970.82010.9768STSF0.28700.577026
    FordChallenge0.62320.87900.9360QUANT0.31280.312825
    BIDMC32HR_disc0.65070.80120.9637RIST0.31300.635324
    STEW0.50000.73600.8385Arsenal0.33850.210024
    BoneIntensitiesAgeGroup0.47640.79100.8202HC20.34380.296626
    PhonemeSpectra0.02560.27890.3746H-InceptionTime0.34890.291126
    KERAAL-RTK0.57140.78570.9286HC20.35710.571426
    LSST0.31510.62900.7040FreshPRINCE0.38890.480926
    HandMovementDirection0.20270.41220.6081TSF0.40540.418926
    DuckDuckGeese0.20000.47000.6400H-InceptionTime0.44000.480026
    SelfRegulationSCP10.50170.85320.9454MRHydra0.44370.208226
    IEEEPPG_disc0.26050.43980.7078ConvTran0.44730.438326
    AsphaltRegularityCoordinates0.50730.98140.9947H-InceptionTime0.48740.291626
    MindReading0.23120.52830.7243LITETime-MV0.49310.385926
    EthanolConcentration0.25100.42970.7490STC0.49810.532326
    UIPRMD-DS-C0.50000.83331.0000Catch220.50000.388926
    WISDM0.36640.86580.8965MRHydra0.53000.131026
    Blink0.44440.98671.0000Arsenal0.55560.428926
    EigenWorms0.41980.86260.9771MRHydra0.55730.557325
    KIMORE-PR-C0.14290.42860.7143LITETime-MV0.57140.571426
    AsphaltObstaclesCoordinates0.28390.82100.8670MRHydra0.58310.289026
    CounterMovementJump0.33520.74860.9274Arsenal0.59220.458126
    Handwriting0.03760.38940.6529H-InceptionTime0.61530.478826
    USCActivity0.11380.69360.7354LITETime-MV0.62160.137921
    RacketSports0.28290.87830.9079RDST0.62500.125026
    UCDHE-Rowing-MC0.20450.73410.8295PatchMTSC0.62500.338626
    IRDS-SFL0.20690.79310.8621RDST0.65520.448326
    Skoda0.23560.94660.9646H-InceptionTime0.72900.119925
    Epilepsy0.26810.98191.0000HC20.73190.101426
    Tiselac0.06280.81360.8373STSF0.77450.204420
    MotionSenseHAR0.20380.98681.0000DrCIF0.79620.101926
    NATOPS0.16670.89170.9667LITETime-MV0.80000.155626
    UCIActivity0.19160.97610.9983LITETime-MV0.80670.174926
    UWaveGestureLibrary0.12500.90940.9406Arsenal0.81560.553126
    ERing0.16670.93330.9963MRHydra0.82960.237026
    PEMS-SF0.11560.86991.0000CIF0.88440.317926
    PenDigits0.10380.97780.9911H-InceptionTime0.88740.234424
    SpokenArabicDigits0.10000.98020.9945DisjointCNN0.89450.129126
    Libras0.06670.88890.9722RIST0.90560.338926
    JapaneseVowels0.08380.96620.9946LiteTIME0.91080.208126
    Cricket0.08330.97921.00001NN-DTW0.91670.069426
    CharacterTrajectories0.06480.98960.9958H-InceptionTime0.93110.044626
    TactileTextureRecognition0.05140.99851.0000H-InceptionTime0.94860.168926
    ArticularyWordRecognition0.04000.98170.9933Arsenal0.95330.050026

    One row per dataset. Dummy is the no-skill floor. Median, best and spread are over the other estimators, so the baseline cannot flatter them. Gain over dummy is best minus dummy, how much skill was found at all; spread is best minus worst, how much the choice of estimator mattered. The two answer different questions, and a single range would conflate them.

    3 of 66 datasets gained 0.05 or less over the baseline (shaded amber) and 15 have a best of 0.99 or more (shaded green). Both separate estimators poorly, for opposite reasons. Best is a maximum over many estimators, so it is optimistic by construction: read it as what the archive can currently do on a problem, not as what any one method delivers.