diff --git a/README.md b/README.md index 0bad635..c2153f2 100644 --- a/README.md +++ b/README.md @@ -40,30 +40,31 @@ The current paper version describes: | # | Estimator | Accuracy rank | Accuracy | Balanced accuracy | AUROC | F1 | Log loss ↓ | Sensitivity | Specificity | |---|---|---|---|---|---|---|---|---|---| -| 1 | HC2 | **7.86** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** | -| 2 | MRHydra | 8.31 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 | -| 3 | RDST | 9.16 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 | -| 4 | RIST | 9.76 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 | -| 5 | DrCIF | 10.07 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 | -| 6 | FreshPRINCE | 10.10 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 | -| 7 | CIF | 10.13 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 | -| 8 | QUANT | 10.55 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 | -| 9 | Arsenal | 10.70 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 | -| 10 | ROCKET | 10.82 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 | -| 11 | LITETime-MV | 11.14 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 | -| 12 | STSF | 11.55 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 | -| 13 | H-InceptionTime | 11.63 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 | -| 14 | LiteTIME | 12.34 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 | -| 15 | PatchMTSC | 13.12 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 | -| 16 | ConvTran | 13.14 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 | -| 17 | Catch22 | 13.15 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 | -| 18 | STC | 13.95 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 | -| 19 | TSF | 14.00 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 | -| 20 | TDE | 14.63 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 | -| 21 | Summary | 16.66 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 | -| 22 | TimesURL | 16.95 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 | -| 23 | 1NN-DTW | 18.45 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 | -| 24 | Dummy | 21.81 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 | +| 1 | HC2 | **8.06** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** | +| 2 | MRHydra | 8.51 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 | +| 3 | RDST | 9.35 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 | +| 4 | RIST | 9.97 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 | +| 5 | DrCIF | 10.25 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 | +| 6 | FreshPRINCE | 10.30 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 | +| 7 | CIF | 10.37 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 | +| 8 | QUANT | 10.81 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 | +| 9 | Arsenal | 10.96 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 | +| 10 | ROCKET | 11.09 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 | +| 11 | LITETime-MV | 11.50 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 | +| 12 | STSF | 11.79 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 | +| 13 | H-InceptionTime | 11.93 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 | +| 14 | LiteTIME | 12.68 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 | +| 15 | PatchMTSC | 13.38 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 | +| 16 | ConvTran | 13.39 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 | +| 17 | Catch22 | 13.42 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 | +| 18 | STC | 14.25 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 | +| 19 | TSF | 14.26 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 | +| 20 | TDE | 15.07 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 | +| 21 | Summary | 17.18 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 | +| 22 | TimesNet | 17.26 | 0.7013 | 0.6659 | 0.8275 | 0.6281 | 1.1613 | 0.6726 | 0.6898 | +| 23 | TimesURL | 17.46 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 | +| 24 | 1NN-DTW | 19.08 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 | +| 25 | Dummy | 22.68 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 | Average over the 52 Multiverse-core datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold. @@ -74,6 +75,14 @@ version with per-metric ranks to ([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/leaderboard.html), since GitHub shows HTML as source). Missing results, and why, are listed on that page. +The same command writes a per-dataset view to +[`results/multiverse/datasets.html`](results/multiverse/datasets.html) +([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/datasets.html)), +which turns the question around: for each dataset it gives the Dummy floor, the median +and best over the other estimators, which estimator was best, how much the best gained +over Dummy, and how far apart the estimators were. It is sorted by that gain, so the +problems where nothing yet beats the baseline come first. + This repository aims to make it easier to: - load Multiverse datasets through `aeon` @@ -91,6 +100,10 @@ This repository aims to make it easier to: · Leaderboard · + Runtime + · + Memory + · Evaluation · Classifiers diff --git a/docs/classifiers.md b/docs/classifiers.md index 8c0f4b7..72e461c 100644 --- a/docs/classifiers.md +++ b/docs/classifiers.md @@ -114,14 +114,50 @@ Mathematics, 9(23), 2021. This is the only Keras port here, following the authors, so it needs `tensorflow` rather than `torch`. Both are in the `deep-learning` extra. -The authors tune `window_size` per dataset. Their results table carries a `Win_pct` -column spread over a five point grid: 20, 40 and 60 on five datasets each, 80 on -thirteen, and 100 on two. The default here is **0.8**, the value they use most often. -The 0.2 in their `config.yml` is the worked example for BasicMotions, not a default. +### Tuning, and what we report + +**The XCM results in this repository follow the authors' protocol.** Section 4.3 sets +`window_size` and `batch_size` per dataset "by grid search based on the best average +accuracy following a stratified 5-fold cross-validation on the training set", over +windows {0.2, 0.4, 0.6, 0.8, 1.0} and batches {1, 8, 32}. Selection never touches the +test data, so the published figures are tuned but not leaked, and neither are ours. + +The reported run searches the window on that grid and holds batch size at 32. That is +the one departure, and it is a cost decision rather than a modelling one: batch 1 takes +roughly 32 times the gradient steps, which would turn a day of GPU time into about 900 +hours, for a value the published table selects on 4 of 30 datasets. + +Both parameters accept a sequence, which triggers the search; a scalar fits once. The +class default is a single fit at **0.8**, the modal published window, because a default +should be cheap, but `XCM` in `tsml-eval` supplies the grid, and `XCM-Fixed` is the +single-fit variant kept for comparison: + +```python +from multiverse.classification import XCMClassifier +from multiverse.classification._xcm import PAPER_WINDOW_SIZES, PAPER_BATCH_SIZES + +XCMClassifier(window_size=PAPER_WINDOW_SIZES) # window only +XCMClassifier(window_size=PAPER_WINDOW_SIZES, batch_size=PAPER_BATCH_SIZES) # full grid +``` + +The selected values are on `window_fraction_` and `batch_size_`, and every grid point's +mean and per-fold accuracy on `cv_results_`. + +Cost is the reason the class default is a single fit. A fixed-window pass over +Multiverse-core took 1.2 GPU-hours in total; searching the window is five candidates +over five folds, about 20 times that. + +That earlier fixed-window pass is what motivated the change. It averaged 0.699 across +the 23 datasets shared with the paper's table against their 0.761, and the paper itself +reports a mean relative accuracy drop of 7.0% +/- 1.3% from using a suboptimal window, +which is the size of the gap observed. Reporting a fixed window would have measured a +configuration the authors never used. Because `window_size` is a fraction, the kernel grows with the series, and 0.8 of -EigenWorms' 17984 points would be a 14387 point kernel. `max_window` bounds the kernel -at 100 points, and it is floored at 1 for very short series. +EigenWorms' 17984 points is a 14387 point kernel. `max_window` bounds it at 100 points +and floors it at 1 for very short series. That bound is ours, not the authors': they run +kernels of this order, 40% of EigenWorms being 7193 points. Set `max_window=None` to +reproduce them, and expect the memory cost to follow. ## Notes on the ports diff --git a/docs/memory.md b/docs/memory.md new file mode 100644 index 0000000..00e2070 --- /dev/null +++ b/docs/memory.md @@ -0,0 +1,92 @@ +# Memory + +**Coming soon.** This page will hold memory comparisons across the Multiverse +estimators. + +**We have not yet structured an experiment to compare memory.** As with +[runtime](runtime.md), every run behind the results in this repository was set up to +measure predictive performance. A job was given whatever memory ceiling got it to +finish, on whatever node was free, with whatever core count came with the partition, +and those were never held constant across estimators because nothing depended on it. +The peak figures that came out are a by-product of that, not a measurement anyone +designed. + +So nothing is published here yet, deliberately. The rest of this page records what a +memory comparison would have to fix, and why the figures we already hold cannot stand +in for one. + +## What is recorded + +`tsml-eval` writes one `memory_usage` value per classifier, dataset and resample: the +peak memory observed during `fit`. It is in the raw prediction files, and this +repository's ingest brings across only the accuracy-style measures, so it has not been +carried over. + +One number, host side, fit only. + +## Why the figures we hold cannot stand in + +Each of these is a condition a memory experiment would have to fix, and that these runs +left free. + +**Peak process memory is not the model's memory.** It includes the interpreter, every +imported library, the loaded dataset and any transient copies made along the way. +Importing TensorFlow or PyTorch alone accounts for a large fixed cost before a single +weight is allocated, so a small model in a heavy framework can report more than a large +model in a light one. Without subtracting a per-framework baseline the figure mostly +ranks frameworks. + +**The dataset can dominate the model.** Multiverse contains problems like EigenWorms at +17,984 timepoints and FaceDetection at 5,890 training cases. For an estimator with a +small parameter count, most of the peak is the data and its copies, which says +something about the problem rather than the method. + +**Host and device memory are different quantities, and only one is recorded.** A model +doing its work on a GPU can look inexpensive by peak host memory while occupying tens of +gigabytes of device memory that nothing here measures. The two are not interchangeable +and cannot be added. + +**Framework allocators do not report what the model needs.** PyTorch's caching allocator +and TensorFlow's default of reserving most of the visible GPU both hold memory they are +not using, so a naive reading measures the allocator's policy rather than the model's +requirement. Getting a meaningful device figure means enabling TensorFlow's memory +growth and reading PyTorch's allocated rather than reserved totals, neither of which +these runs did. + +**Failures censor the measurement.** Where an estimator ran out of memory we do not have +a peak, we have a lower bound and a ceiling. Both kinds appear in +`results/multiverse/missing_results.csv`: ConvTran hit CUDA out of memory on Alzheimers, +EigenWorms and PhotoStimulation, a device-side limit; FreshPRINCE hit OOM at 128 GB on +FaceDetection, FordChallenge and Skoda after eight attempts, a host-side one. Those are +the cases where memory mattered most, and they are exactly the cases with no number. A +table built only from successful runs is a survivorship-biased view of memory use. + +**Memory scales with the resources granted.** The classical ensembles allocate per +thread, so their peak moves with the cores allocated, which varied by partition. Our +controllers also escalate a job's request from 64 GB to 128 GB after a failure, so +different runs of the same estimator saw different ceilings. + +**Peak is timing-dependent.** Python's peak resident memory depends on when garbage +collection happens to run and on whether an allocator returned pages to the operating +system. Repeats of an identical run differ for reasons that have nothing to do with the +method. + +## What a fair comparison would need + +- All compared estimators on the same node, with cores and the memory ceiling fixed and + recorded. +- A per-framework baseline measured and subtracted, so the figure is the model's cost + rather than the cost of importing its library. +- Host and device peaks reported separately, never summed, with TensorFlow memory growth + enabled and PyTorch read via its allocated totals. +- `predict` measured as well as `fit`. Deployment cost is a separate question from + training cost, and the ranking is not the same on both. +- Failures reported alongside successes, as censored observations with the ceiling that + was in force, rather than dropped. +- Repeated runs, since peak varies between identical repeats. +- Asymptotic space complexity in the number of cases, series length and channels stated + beside the measured peaks, so a reader can tell whether a figure will hold at a + different scale. + +Until most of that is in place, this page stays empty. For the same reason the +[leaderboard](leaderboard.md) carries no memory column. diff --git a/docs/runtime.md b/docs/runtime.md new file mode 100644 index 0000000..c259acd --- /dev/null +++ b/docs/runtime.md @@ -0,0 +1,87 @@ +# Runtime + +**Coming soon.** This page will hold runtime comparisons across the Multiverse +estimators. Memory is a separate page, [memory](memory.md), for the same reasons and a +few of its own. + +**We have not yet structured an experiment to compare runtime.** Every run behind the +results in this repository was set up to measure predictive performance. Which partition +a classifier was queued on, how many cores it was given, how many epochs it trained for, +whether a job was retried at a higher memory ceiling: all of those were chosen to get +accurate results out at a reasonable cost, and none were held constant across estimators +because nothing depended on it. The timings that came out are a by-product of that, not +a measurement anyone designed. + +So nothing is published here yet, deliberately. Comparing runtime needs its own +experiment, with the conditions below fixed in advance, and we have not run one. The +rest of this page records what those conditions are, and why the figures we already hold +cannot stand in for them. + +## The measurements exist + +Every run already records timings. `tsml-eval` writes, per classifier, dataset and +resample: + +- `fit_time` and `predict_time`, in seconds; +- `benchmark_time`, the time that machine took to sort 1,000 seeded random arrays of + 20,000 elements; +- `memory_usage`, the peak memory during `fit`, which [memory](memory.md) covers. + +They are in the raw prediction files. What this repository ingests under `results/` is +only the accuracy-style measures, one file per metric, so the timings have not been +brought across yet. That is a small piece of work; the reason it has not been done is +below, not the effort. + +## Why the figures we hold cannot stand in + +Each of these is a condition a timing experiment would have to fix, and that these runs +left free. + +**The runs are spread across different hardware.** Multiverse results have been produced +on GPU partitions with H200 and A100 cards and on CPU-only nodes, with different core +counts. GPU jobs in our configurations are allocated two CPUs each. A fit time from one +partition and a fit time from another are two different measurements that happen to share +a unit. + +**GPU and CPU methods are not on one axis.** For the deep learners nearly all the work is +on the accelerator and the host CPU mostly feeds batches; for the classical ensembles +there is no accelerator at all and the time scales with the cores allocated. Comparing +them measures the hardware at least as much as the algorithm, and the ratio moves when +either side changes. A statement like "X is 40 times faster than Y" is, in this setting, +a statement about a purchasing decision. + +**Wall-clock contains things that are not the algorithm.** Queueing, data loading, and +retries: our controllers escalate a job's memory request from 64 GB to 128 GB after a +failure, so an elapsed time can include a dead run at the lower ceiling. + +**Training time is a hyperparameter, not a property of a method.** A deep learner's fit +time is close to linear in the number of epochs, and the epoch count is a choice. Two +faithful ports of the same paper can differ several-fold on time because the authors +picked 500 epochs and the toolkit's default is 2000. Early stopping and best-epoch +selection move it again. None of that is a fact about the architecture. + +**Which device a run actually used is not reliably recorded.** In the version of the +experiment tooling used for these runs, the device description inspects TensorFlow only, +so a PyTorch estimator reports CPU whether or not it ran on a GPU. Any timing table built +from those records has to have its device column reconstructed from the job +configuration rather than trusted as written. + +## What a fair comparison would need + +- The compared estimators run on the same hardware, or CPU timings normalised by + `benchmark_time`, which exists for exactly this purpose. There is no equivalent + normaliser for GPU work. +- Thread and core counts fixed and recorded, since the classical methods scale with them. +- `fit` and `predict` reported separately. They answer different questions: fit time is + the cost of research, predict time is the cost of deployment, and the ranking is not + the same on both. +- Repeated runs. Timings vary far more between repeats than accuracy does, especially on + shared nodes. +- Asymptotic complexity in the number of cases, series length and channels reported + beside the measured times, so a reader can tell whether a result will hold at a + different scale. +- The device stated per run, from the job configuration. + +Until most of that is in place, this page stays empty. For the same reason the +[leaderboard](leaderboard.md) carries no fit time or predict time columns, and +[memory](memory.md) is empty too. diff --git a/multiverse/classification/__init__.py b/multiverse/classification/__init__.py index cb5dd87..c97a596 100644 --- a/multiverse/classification/__init__.py +++ b/multiverse/classification/__init__.py @@ -5,6 +5,7 @@ __all__ = [ "ConvTranClassifier", + "DisjointCNNClassifier", "PatchMTSCClassifier", "TimesNetClassifier", "TS2VecClassifier", @@ -13,6 +14,7 @@ ] from multiverse.classification._convtran import ConvTranClassifier +from multiverse.classification._disjoint_cnn import DisjointCNNClassifier from multiverse.classification._patchmtsc import PatchMTSCClassifier from multiverse.classification._timesnet import TimesNetClassifier from multiverse.classification._ts2vec import TS2VecClassifier diff --git a/multiverse/classification/_disjoint_cnn.py b/multiverse/classification/_disjoint_cnn.py new file mode 100644 index 0000000..eb4af7e --- /dev/null +++ b/multiverse/classification/_disjoint_cnn.py @@ -0,0 +1,389 @@ +"""Disjoint-CNN classifier for aeon. + +Adapted from the authors' Disjoint-CNN implementation: +https://github.com/Navidfoumani/Disjoint-CNN + +Disjoint-CNN factorises a multivariate convolution into two disjoint steps: a +temporal convolution applied within each channel, then a spatial convolution +across channels, with a non-linearity between them. Stacking these 1+1D blocks +is the paper's alternative to a single joint convolution over both axes. + +This exists alongside ``aeon.classification.deep_learning.DisjointCNNClassifier`` +deliberately. That implementation scores far below the authors' published +numbers, 20.3 accuracy points below on the 23 shared UEA datasets in our runs +(aeon issue #3775). This port follows the authors' own training procedure rather +than aeon's, so the two can be run against each other and the gap attributed. +The training differences are recorded under "Deviations from aeon" below; the +data pipeline is not one of them, since the authors pass ``normalise=False`` in +``Main.py`` and so train on raw series, as aeon does. + +This wrapper is designed for aeon and therefore assumes input X is a 3D NumPy +array with shape (n_cases, n_channels, n_timepoints). The original expects +(n_cases, n_timepoints, n_channels, 1), so X is transposed and expanded +internally. + +The original source is distributed under the MIT License. + +MIT License + +Copyright (c) 2021 Navid Mohammadi Foumani + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. +""" + +from __future__ import annotations + +__maintainer__ = ["TonyBagnall"] +__all__ = ["DisjointCNNClassifier"] + +import math + +import numpy as np +from aeon.classification import BaseClassifier +from sklearn.utils import check_random_state + +# Temporal kernel length per block, from the authors' DCNN_2L, DCNN_3L and +# DCNN_4L. The kernel shortens as the stack deepens. +KERNEL_SIZES = {2: (8, 5), 3: (8, 5, 3), 4: (8, 5, 5, 3)} +# Pool size after the last block, which also differs per variant. +POOL_SIZES = {2: 3, 3: 3, 4: 5} + + +def _build_disjoint_cnn(n_timepoints, n_channels, n_classes, n_layers, n_filters): + """Build the authors' Disjoint-CNN network. + + A transcription of ``classifiers/DCNN_{2,3,4}L.py``. Each block is a + temporal convolution over time within a channel, then a spatial convolution + that spans every channel at once, then a permute that puts the filter axis + back where the next block's spatial convolution expects it. + + The authors' arrays are (case, time, channel), from + ``utils/data_loader.py::process_ts_data``, so time is axis 1 and the channel + axis is the one the spatial convolution spans entirely. + """ + from tensorflow.keras.layers import ( + BatchNormalization, + Conv2D, + Dense, + ELU, + GlobalAveragePooling2D, + Input, + MaxPooling2D, + Permute, + ) + from tensorflow.keras.models import Model + + kernels = KERNEL_SIZES[n_layers] + input_layer = Input((n_timepoints, n_channels, 1)) + + x = input_layer + for block, kernel in enumerate(kernels): + # Temporal: (kernel, 1) slides along time inside each channel. + x = Conv2D( + n_filters, + (kernel, 1), + strides=1, + padding="same", + kernel_initializer="he_uniform", + )(x) + x = BatchNormalization()(x) + x = ELU(alpha=1.0)(x) + + # Spatial: (1, width) spans the whole channel axis in one filter, so + # padding is valid and the axis collapses to length 1. + width = int(x.shape[2]) + x = Conv2D( + n_filters, + (1, width), + strides=1, + padding="valid", + kernel_initializer="he_uniform", + )(x) + x = BatchNormalization()(x) + x = ELU(alpha=1.0)(x) + + # The authors permute after every block except the last. + if block < len(kernels) - 1: + x = Permute((1, 3, 2))(x) + + x = MaxPooling2D(pool_size=(POOL_SIZES[n_layers], 1), strides=None, padding="valid")(x) + x = GlobalAveragePooling2D()(x) + output_layer = Dense(n_classes, activation="softmax")(x) + return Model(inputs=input_layer, outputs=output_layer) + + +def _class_weights(one_hot, mu=2.0): + """Reproduce ``utils/classifier_tools.py::create_class_weight``. + + Weight for a class is ``log(mu * total / count)``, floored at 1.0, so rare + classes are upweighted and common ones are never downweighted below 1. + """ + counts = one_hot.sum(axis=0) + total = counts.sum() + weights = {} + for index, count in enumerate(counts): + score = math.log(mu * total / float(count)) if count > 0 else 1.0 + weights[index] = score if score > 1.0 else 1.0 + return weights + + +class DisjointCNNClassifier(BaseClassifier): + """Disjoint-CNN, following the authors' training procedure. + + A stack of 1+1D blocks: a temporal convolution within each channel, then a + spatial convolution across all channels, with ELU and batch normalisation + between. The stack is max pooled, globally average pooled and classified. + + Parameters + ---------- + n_layers : int, default=4 + Number of disjoint blocks, 2, 3 or 4, selecting the authors' DCNN_2L, + DCNN_3L or DCNN_4L. The temporal kernel lengths and the final pool size + follow the variant. + n_filters : int, default=64 + Filters in every convolution. + n_epochs : int, default=500 + Training epochs. The authors' ``Main.py`` sets 500. + batch_size : int, default=8 + Upper bound on the batch size. The batch actually used is + ``min(n_cases // 10, batch_size)``, the authors' rule, so small + collections train with very small batches. + validation_size : float, default=0.1 + Fraction of the training collection sampled to monitor validation loss. + The authors sample this from the training set *with replacement* and do + not hold it out, so it overlaps the training data. Reproduced because + the learning rate schedule and the retained epoch both depend on it. + Set to 0 to monitor training loss instead, which is what aeon does. + use_class_weights : bool, default=True + Whether to weight the loss by class, as the authors do. aeon does not. + verbose : bool, default=False + Whether Keras prints training progress. + random_state : int, RandomState instance or None, default=None + Seed controlling weight initialisation, batch shuffling and the + validation sample. + + Attributes + ---------- + model_ : keras.Model + The fitted network, with the retained epoch's weights. + history_ : dict + Keras training history, one entry per metric. + batch_size_ : int + Batch size actually used, after the authors' rule. + class_weights_ : dict or None + Weights applied to the loss, or None when disabled. + n_channels_ : int + Number of channels seen in ``fit``. + n_timepoints_ : int + Series length seen in ``fit``. + classes_ : np.ndarray + Class labels, from ``BaseClassifier``. + n_classes_ : int + Number of classes, from ``BaseClassifier``. + + Notes + ----- + Deviations from ``aeon.classification.deep_learning.DisjointCNNClassifier``, + all of them places where aeon departs from the authors' code, and the + candidates for the published-versus-obtained gap in aeon issue #3775: + + - **Class weighting.** The authors weight the loss per class; aeon does not. + - **Batch size.** The authors use ``min(n_cases // 10, 8)``; aeon defaults + to a flat 16, since ``use_mini_batch_size`` is False. + - **Epochs.** The authors train for 500; aeon defaults to 2000. + - **What is monitored.** The authors monitor validation loss for both the + learning rate schedule and the retained epoch; aeon monitors training + loss, having no validation split. + + The architecture here is the clean stack, matching aeon, so that a + difference between the two is attributable to the training procedure alone. + It is worth recording that the authors' own ``DCNN_4L.py`` does not build + that stack: in the third block the spatial convolution is applied to + ``conv2``, the second block's output, rather than to the third block's + temporal convolution, which is therefore computed and discarded. It reads + like a copy-paste slip, but it is the graph their published numbers came + from, so replicating those exactly would need it. ``DCNN_2L`` and + ``DCNN_3L`` are unaffected. + + References + ---------- + .. [1] Foumani, S. N. M., Tan, C. W. and Salehi, M. "Disjoint-CNN for + Multivariate Time Series Classification." ICDMW, 2021. + + Examples + -------- + >>> from aeon.testing.data_generation import make_example_3d_numpy + >>> from multiverse.classification import DisjointCNNClassifier + >>> X, y = make_example_3d_numpy(n_cases=8, n_channels=2, n_timepoints=20) + >>> clf = DisjointCNNClassifier(n_epochs=2) # doctest: +SKIP + >>> clf.fit(X, y) # doctest: +SKIP + """ + + _tags = { + "X_inner_type": "numpy3D", + "capability:multivariate": True, + "capability:unequal_length": False, + "algorithm_type": "deeplearning", + "non_deterministic": True, + "python_dependencies": "tensorflow", + } + + def __init__( + self, + n_layers: int = 4, + n_filters: int = 64, + n_epochs: int = 500, + batch_size: int = 8, + validation_size: float = 0.1, + use_class_weights: bool = True, + verbose: bool = False, + random_state=None, + ): + self.n_layers = n_layers + self.n_filters = n_filters + self.n_epochs = n_epochs + self.batch_size = batch_size + self.validation_size = validation_size + self.use_class_weights = use_class_weights + self.verbose = verbose + self.random_state = random_state + super().__init__() + + def _validate_parameters(self) -> None: + """Check constructor parameters before any work is done.""" + if self.n_layers not in KERNEL_SIZES: + raise ValueError( + f"n_layers must be one of {sorted(KERNEL_SIZES)}, got {self.n_layers}" + ) + for name in ["n_filters", "n_epochs", "batch_size"]: + value = getattr(self, name) + if not isinstance(value, int) or value <= 0: + raise ValueError(f"{name} must be a positive integer") + if not 0 <= self.validation_size < 1: + raise ValueError("validation_size must be in [0, 1)") + + @staticmethod + def _to_original_layout(X: np.ndarray) -> np.ndarray: + """Convert an aeon collection to the authors' (case, time, channel, 1).""" + series = np.transpose(np.asarray(X, dtype=np.float32), (0, 2, 1)) + return series[..., np.newaxis] + + def _fit(self, X: np.ndarray, y): + self._validate_parameters() + + import tensorflow as tf + + rng = check_random_state(self.random_state) + seed = int(rng.randint(np.iinfo(np.int32).max)) + tf.keras.utils.set_random_seed(seed) + + self.n_channels_, self.n_timepoints_ = X.shape[1], X.shape[2] + + encoded_y = np.asarray( + [self._class_dictionary[label] for label in y], dtype=np.int64 + ) + one_hot = np.eye(self.n_classes_, dtype=np.float32)[encoded_y] + + # The authors' rule, which can floor at zero on tiny collections. + self.batch_size_ = max(1, min(X.shape[0] // 10, self.batch_size)) + + self.model_ = _build_disjoint_cnn( + self.n_timepoints_, + self.n_channels_, + self.n_classes_, + self.n_layers, + self.n_filters, + ) + self.model_.compile( + loss="categorical_crossentropy", + optimizer=tf.keras.optimizers.Adam(), + metrics=["accuracy"], + ) + + data = self._to_original_layout(X) + validation_data = None + monitor = "loss" + if self.validation_size: + # Sampled from the training set with replacement and left in it, + # exactly as the authors do. + size = max(1, int(X.shape[0] * self.validation_size)) + index = rng.randint(0, X.shape[0], size) + validation_data = (data[index], one_hot[index]) + monitor = "val_loss" + + self.class_weights_ = ( + _class_weights(one_hot) if self.use_class_weights else None + ) + + callbacks = [ + tf.keras.callbacks.ReduceLROnPlateau( + monitor=monitor, factor=0.5, patience=50, min_lr=0.0001 + ), + # The authors checkpoint to disk and reload; restoring the best + # weights in memory is the same selection without the file. + tf.keras.callbacks.EarlyStopping( + monitor=monitor, + patience=self.n_epochs, + restore_best_weights=True, + ), + ] + + history = self.model_.fit( + data, + one_hot, + validation_data=validation_data, + class_weight=self.class_weights_, + epochs=self.n_epochs, + batch_size=self.batch_size_, + verbose=1 if self.verbose else 0, + callbacks=callbacks, + ) + self.history_ = history.history + return self + + def _check_shape(self, X: np.ndarray) -> None: + if X.shape[1] != self.n_channels_: + raise ValueError( + f"X has {X.shape[1]} channels, but the classifier was fitted " + f"with {self.n_channels_}." + ) + if X.shape[2] != self.n_timepoints_: + raise ValueError( + f"X has length {X.shape[2]}, but the classifier was fitted with " + f"length {self.n_timepoints_}." + ) + + def _predict_proba(self, X: np.ndarray) -> np.ndarray: + self._check_shape(X) + return self.model_.predict(self._to_original_layout(X), verbose=0) + + def _predict(self, X: np.ndarray): + return self.classes_[np.argmax(self._predict_proba(X), axis=1)] + + @classmethod + def _get_test_params(cls, parameter_set: str = "default") -> dict: + """Return a small parameter set for aeon estimator checks.""" + return { + "n_layers": 2, + "n_filters": 4, + "n_epochs": 2, + "batch_size": 4, + "validation_size": 0.0, + "random_state": 0, + } diff --git a/multiverse/classification/_xcm.py b/multiverse/classification/_xcm.py index 3254d4f..2254e57 100644 --- a/multiverse/classification/_xcm.py +++ b/multiverse/classification/_xcm.py @@ -66,6 +66,17 @@ from sklearn.utils import check_random_state +#: The authors' grid, from §4.3: window size as a fraction of series length, and +#: batch size. They search the product of the two per dataset. +PAPER_WINDOW_SIZES = (0.2, 0.4, 0.6, 0.8, 1.0) +PAPER_BATCH_SIZES = (1, 8, 32) + + +def _as_grid(value): + """Return a parameter as a list of candidates, scalar or sequence alike.""" + return list(value) if isinstance(value, (list, tuple)) else [value] + + def _build_xcm(n_timepoints, n_channels, n_classes, window_size, n_filters): """Build the authors' XCM network. @@ -145,25 +156,38 @@ class XCMClassifier(BaseClassifier): Parameters ---------- - window_size : float, default=0.8 + window_size : float or sequence of float, default=0.8 Length of the convolution kernels along time, as a fraction of the - series length. The authors tune this per dataset over - {0.2, 0.4, 0.6, 0.8, 1.0}; 0.8 is the value their results table uses - most often, on 13 of 30 datasets. The 0.2 in their ``config.yml`` is - the worked example for BasicMotions, not a default. - max_window : int, default=100 - Upper bound on the kernel length in points. Because ``window_size`` is - a fraction, the kernel grows with the series: 0.8 of EigenWorms' 17984 - points would be a 14387 point kernel. The bound keeps long series - tractable. The kernel is also floored at one point, since - ``int(window_size * n)`` is zero for very short series. + series length. Pass a single value to use it directly, or a sequence to + select from it by cross-validation as the authors do. Their grid is + :data:`PAPER_WINDOW_SIZES`, {0.2, 0.4, 0.6, 0.8, 1.0}; the default 0.8 + is the value their results table settles on most often, 13 of 30 + datasets. The 0.2 in their ``config.yml`` is the worked example for + BasicMotions, not a default. + max_window : int or None, default=100 + Upper bound on the kernel length in points, or None for no bound. + Because ``window_size`` is a fraction, the kernel grows with the + series: 0.8 of EigenWorms' 17984 points is a 14387 point kernel, and + the authors do run kernels of that order, 40% of EigenWorms being 7193 + points. The bound is ours, not theirs, and keeps long series tractable; + None reproduces their behaviour. The kernel is also floored at one + point, since ``int(window_size * n)`` is zero for very short series. n_filters : int, default=128 Number of filters in each convolution. n_epochs : int, default=100 Training epochs. The authors train for a fixed number with no early stopping. - batch_size : int, default=32 - Training batch size. + batch_size : int or sequence of int, default=32 + Training batch size, tuned alongside ``window_size`` when a sequence is + given. The authors' grid is :data:`PAPER_BATCH_SIZES`, {1, 8, 32}, but + 32 is their choice on 26 of 30 datasets and batch 1 costs roughly 32 + times the gradient steps, so tuning the window alone recovers most of + the benefit for a small fraction of the compute. + cv_folds : int, default=5 + Folds in the stratified cross-validation used to select parameters, + following the authors' five. Reduced automatically when a class has + fewer members than this, and selection is skipped entirely when the + rarest class appears once. verbose : bool, default=False Whether Keras prints training progress. random_state : int, RandomState instance or None, default=None @@ -177,6 +201,14 @@ class XCMClassifier(BaseClassifier): Keras training history, one entry per metric. window_size_ : int Kernel length along time actually used, in points. + window_fraction_ : float + Fraction selected, before conversion to points and any ``max_window`` + bound. + batch_size_ : int + Batch size actually used. + cv_results_ : list of dict + One entry per grid point, with its mean and per-fold accuracy. Empty + when no search was run. n_channels_ : int Number of channels seen in ``fit``. n_timepoints_ : int @@ -199,6 +231,11 @@ class XCMClassifier(BaseClassifier): >>> X, y = make_example_3d_numpy(n_cases=8, n_channels=2, n_timepoints=20) >>> clf = XCMClassifier(n_epochs=2) # doctest: +SKIP >>> clf.fit(X, y) # doctest: +SKIP + + Selecting the window by cross-validation, as the paper does: + + >>> from multiverse.classification._xcm import PAPER_WINDOW_SIZES + >>> clf = XCMClassifier(window_size=PAPER_WINDOW_SIZES) # doctest: +SKIP """ _tags = { @@ -212,11 +249,12 @@ class XCMClassifier(BaseClassifier): def __init__( self, - window_size: float = 0.8, - max_window: int = 100, + window_size=0.8, + max_window: int | None = 100, n_filters: int = 128, n_epochs: int = 100, - batch_size: int = 32, + batch_size=32, + cv_folds: int = 5, verbose: bool = False, random_state=None, ): @@ -225,18 +263,29 @@ def __init__( self.n_filters = n_filters self.n_epochs = n_epochs self.batch_size = batch_size + self.cv_folds = cv_folds self.verbose = verbose self.random_state = random_state super().__init__() def _validate_parameters(self) -> None: """Check constructor parameters before any work is done.""" - if not 0 < self.window_size <= 1: - raise ValueError("window_size must be in (0, 1]") - for name in ["max_window", "n_filters", "n_epochs", "batch_size"]: + for window in _as_grid(self.window_size): + if not 0 < window <= 1: + raise ValueError("window_size must be in (0, 1]") + for batch in _as_grid(self.batch_size): + if not isinstance(batch, int) or batch <= 0: + raise ValueError("batch_size must be a positive integer") + for name in ["n_filters", "n_epochs"]: value = getattr(self, name) if not isinstance(value, int) or value <= 0: raise ValueError(f"{name} must be a positive integer") + if self.max_window is not None and ( + not isinstance(self.max_window, int) or self.max_window <= 0 + ): + raise ValueError("max_window must be a positive integer or None") + if not isinstance(self.cv_folds, int) or self.cv_folds < 2: + raise ValueError("cv_folds must be an integer of at least 2") @staticmethod def _to_original_layout(X: np.ndarray) -> np.ndarray: @@ -244,47 +293,137 @@ def _to_original_layout(X: np.ndarray) -> np.ndarray: series = np.transpose(np.asarray(X, dtype=np.float32), (0, 2, 1)) return series[..., np.newaxis] - def _fit(self, X: np.ndarray, y): - self._validate_parameters() - - import tensorflow as tf + def _kernel_length(self, window_size: float) -> int: + """Kernel length along time in points, for a fraction of the series. - rng = check_random_state(self.random_state) - seed = int(rng.randint(np.iinfo(np.int32).max)) - tf.keras.utils.set_random_seed(seed) + The authors compute ``int(window_size * n)``, which is zero for very + short series and unbounded for long ones. The floor at one point is + needed for the first; ``max_window`` bounds the second, and is ours, + not theirs, so setting it to None reproduces their behaviour. + """ + length = max(1, int(window_size * self.n_timepoints_)) + return length if self.max_window is None else min(length, self.max_window) - self.n_channels_, self.n_timepoints_ = X.shape[1], X.shape[2] - # The authors compute the kernel length as int(window_size * n), which - # is zero for short series and unboundedly large for long ones. - self.window_size_ = min( - max(1, int(self.window_size * self.n_timepoints_)), self.max_window - ) + def _fit_model(self, X, one_hot, window_size, batch_size, epochs): + """Build, compile and train one network. Returns the fitted model.""" # _build_xcm takes a fraction, as the authors' function does, so convert # back. The half point guards against int() rounding the fraction down. - effective = (self.window_size_ + 0.5) / self.n_timepoints_ - - self.model_ = _build_xcm( + effective = (self._kernel_length(window_size) + 0.5) / self.n_timepoints_ + model = _build_xcm( self.n_timepoints_, self.n_channels_, self.n_classes_, effective, self.n_filters, ) - self.model_.compile( + model.compile( optimizer="adam", loss="categorical_crossentropy", metrics=["accuracy"] ) + history = model.fit( + self._to_original_layout(X), + one_hot, + epochs=epochs, + batch_size=batch_size, + verbose=1 if self.verbose else 0, + ) + return model, history + + def _select_parameters(self, X, encoded_y, one_hot, grid, seed): + """Choose window and batch size by the authors' cross-validation. + + Section 4.3: "hyperparameters were set by grid search based on the best + average accuracy following a stratified 5-fold cross-validation on the + training set". The selection therefore never sees the test data. + + Folds with fewer members than ``cv_folds`` cannot be stratified, so the + number of splits is reduced to the smallest class count, and a class + appearing once leaves nothing to select on, in which case the first + candidate is taken. + """ + from sklearn.model_selection import StratifiedKFold + + counts = np.bincount(encoded_y, minlength=self.n_classes_) + splits = min(self.cv_folds, int(counts[counts > 0].min())) + if splits < 2: + self.cv_results_ = [] + return grid[0] + + folds = list( + StratifiedKFold( + n_splits=splits, shuffle=True, random_state=seed + ).split(X, encoded_y) + ) + + self.cv_results_ = [] + for window_size, batch_size in grid: + scores = [] + for train_index, test_index in folds: + model, _ = self._fit_model( + X[train_index], + one_hot[train_index], + window_size, + batch_size, + self.n_epochs, + ) + predicted = model.predict( + self._to_original_layout(X[test_index]), verbose=0 + ).argmax(axis=1) + scores.append(float((predicted == encoded_y[test_index]).mean())) + del model + self.cv_results_.append( + { + "window_size": window_size, + "batch_size": batch_size, + "mean_accuracy": float(np.mean(scores)), + "fold_accuracies": scores, + } + ) + if self.verbose: + print( + f"window_size={window_size} batch_size={batch_size}: " + f"{np.mean(scores):.4f}" + ) + + best = max(self.cv_results_, key=lambda r: r["mean_accuracy"]) + return best["window_size"], best["batch_size"] + + def _fit(self, X: np.ndarray, y): + self._validate_parameters() + + import tensorflow as tf + + rng = check_random_state(self.random_state) + seed = int(rng.randint(np.iinfo(np.int32).max)) + tf.keras.utils.set_random_seed(seed) + + self.n_channels_, self.n_timepoints_ = X.shape[1], X.shape[2] encoded_y = np.asarray( [self._class_dictionary[label] for label in y], dtype=np.int64 ) one_hot = np.eye(self.n_classes_, dtype=np.float32)[encoded_y] - history = self.model_.fit( - self._to_original_layout(X), - one_hot, - epochs=self.n_epochs, - batch_size=self.batch_size, - verbose=1 if self.verbose else 0, + grid = [ + (window, batch) + for window in _as_grid(self.window_size) + for batch in _as_grid(self.batch_size) + ] + if len(grid) > 1: + window_size, batch_size = self._select_parameters( + X, encoded_y, one_hot, grid, seed + ) + else: + window_size, batch_size = grid[0] + self.cv_results_ = [] + + self.window_size_ = self._kernel_length(window_size) + self.batch_size_ = batch_size + self.window_fraction_ = window_size + + # The authors refit on the whole training set once the parameters are + # chosen, which is what produces their reported result. + self.model_, history = self._fit_model( + X, one_hot, window_size, batch_size, self.n_epochs ) self.history_ = history.history return self diff --git a/multiverse/classification/tests/test_ported_classifiers.py b/multiverse/classification/tests/test_ported_classifiers.py index e421f7a..b821c56 100644 --- a/multiverse/classification/tests/test_ported_classifiers.py +++ b/multiverse/classification/tests/test_ported_classifiers.py @@ -14,6 +14,7 @@ from multiverse.classification import ( ConvTranClassifier, + DisjointCNNClassifier, PatchMTSCClassifier, TimesNetClassifier, TimesURLClassifier, @@ -67,6 +68,14 @@ "device": "cpu", "random_state": 0, }, + DisjointCNNClassifier: { + "n_layers": 2, + "n_filters": 4, + "n_epochs": 2, + "batch_size": 4, + "validation_size": 0.0, + "random_state": 0, + }, XCMClassifier: { "window_size": 0.2, "n_filters": 4, diff --git a/multiverse/classification/tests/test_xcm_tuning.py b/multiverse/classification/tests/test_xcm_tuning.py new file mode 100644 index 0000000..cd9245d --- /dev/null +++ b/multiverse/classification/tests/test_xcm_tuning.py @@ -0,0 +1,83 @@ +"""Tests for XCM's cross-validated parameter search. + +The authors set window size and batch size per dataset by grid search over a +stratified five-fold cross-validation of the training set (§4.3). These check +the search runs when asked, is skipped when not, and never sees test data. +""" + +import pytest + +pytest.importorskip("tensorflow") + +from aeon.testing.data_generation import make_example_3d_numpy + +from multiverse.classification import XCMClassifier + + +def test_xcm_scalar_parameters_run_no_search(): + """A scalar window and batch fit once, with no cross-validation.""" + X, y = make_example_3d_numpy( + n_cases=20, n_channels=2, n_timepoints=40, n_labels=2, random_state=0 + ) + clf = XCMClassifier(n_epochs=1, n_filters=2, random_state=0).fit(X, y) + assert clf.cv_results_ == [] + assert clf.batch_size_ == 32 + # 0.8 of 40 points, under the 100 point bound + assert clf.window_size_ == 32 + assert clf.window_fraction_ == 0.8 + + +def test_xcm_tunes_over_a_window_grid(): + """A sequence triggers the authors' search and records every grid point.""" + X, y = make_example_3d_numpy( + n_cases=20, n_channels=2, n_timepoints=40, n_labels=2, random_state=0 + ) + clf = XCMClassifier( + window_size=[0.2, 1.0], n_epochs=1, n_filters=2, cv_folds=2, random_state=0 + ).fit(X, y) + assert len(clf.cv_results_) == 2 + assert {r["window_size"] for r in clf.cv_results_} == {0.2, 1.0} + assert all(len(r["fold_accuracies"]) == 2 for r in clf.cv_results_) + # the fitted fraction is whichever scored best + best = max(clf.cv_results_, key=lambda r: r["mean_accuracy"]) + assert clf.window_fraction_ == best["window_size"] + + +def test_xcm_tunes_window_and_batch_together(): + """Both sequences give the product of the two grids.""" + X, y = make_example_3d_numpy( + n_cases=20, n_channels=2, n_timepoints=40, n_labels=2, random_state=0 + ) + clf = XCMClassifier( + window_size=[0.2, 1.0], batch_size=[8, 16], n_epochs=1, n_filters=2, + cv_folds=2, random_state=0, + ).fit(X, y) + assert len(clf.cv_results_) == 4 + assert clf.batch_size_ in {8, 16} + + +def test_xcm_max_window_none_leaves_the_kernel_unbounded(): + """None reproduces the authors' behaviour of not bounding the kernel.""" + X, y = make_example_3d_numpy( + n_cases=10, n_channels=2, n_timepoints=300, n_labels=2, random_state=0 + ) + bounded = XCMClassifier(n_epochs=1, n_filters=2, random_state=0).fit(X, y) + assert bounded.window_size_ == 100 # capped + + unbounded = XCMClassifier( + max_window=None, n_epochs=1, n_filters=2, random_state=0 + ).fit(X, y) + assert unbounded.window_size_ == 240 # 0.8 of 300 + + +def test_xcm_rejects_bad_grids(): + """Every candidate is validated, not just the first.""" + X, y = make_example_3d_numpy( + n_cases=10, n_channels=2, n_timepoints=20, n_labels=2, random_state=0 + ) + with pytest.raises(ValueError, match="window_size"): + XCMClassifier(window_size=[0.5, 1.5]).fit(X, y) + with pytest.raises(ValueError, match="batch_size"): + XCMClassifier(batch_size=[8, 0]).fit(X, y) + with pytest.raises(ValueError, match="cv_folds"): + XCMClassifier(window_size=[0.2, 0.8], cv_folds=1).fit(X, y) diff --git a/multiverse/experiments/tables.py b/multiverse/experiments/tables.py index bcef1e6..ec14367 100644 --- a/multiverse/experiments/tables.py +++ b/multiverse/experiments/tables.py @@ -24,6 +24,9 @@ __maintainer__ = ["TonyBagnall"] __all__ = [ "leaderboard", + "dataset_summary", + "dataset_page", + "dataset_markdown", "leaderboard_markdown", "write_markdown_table", "available_estimators", @@ -42,6 +45,12 @@ DEFAULT_RESULTS_DIR = Path(__file__).resolve().parents[2] / "results" / "multiverse" +#: A dataset whose best estimator gains no more than this over the baseline +#: shows little signal; one whose best reaches SATURATED_BEST is solved. +#: Both separate estimators poorly, for opposite reasons. +NO_SIGNAL_GAIN = 0.05 +SATURATED_BEST = 0.99 + #: Metrics where a smaller value is a better result. Anything not listed here is #: treated as higher-is-better. LOWER_IS_BETTER = {"logloss"} @@ -846,6 +855,293 @@ def write_markdown_table( return True +def dataset_summary( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", +) -> pd.DataFrame: + """Summarise one metric per dataset rather than per estimator. + + The leaderboard answers "which estimator is best"; this answers "what is + this dataset worth", which is the question an archive has to keep asking of + its own problems. + + Unlike :func:`leaderboard` this does not restrict to the datasets every + estimator has results for. A dataset only some estimators finished is still + informative, so every dataset is kept and ``estimators`` records how many + contributed. + + Parameters + ---------- + datasets : list of str + Datasets to include, in any order. + estimators : list of str + Estimator names, matching their results directories. + metric : str + Metric to summarise. + results_dir : Path or str + Directory holding one sub-directory per estimator. + baseline : str + Estimator treated as the no-skill floor, excluded from best, worst, + median and spread. Pass None to keep it in. + + Returns + ------- + pd.DataFrame + One row per dataset, indexed by dataset name, with the baseline score, + the median, best and worst over the remaining estimators, the estimator + achieving the best, the gain over the baseline, the spread, and the + number of estimators contributing. + """ + frames = _load_all(list(estimators), [metric], results_dir) + if metric not in frames: + raise ValueError(f"no results for metric {metric!r}") + scores = frames[metric].reindex(sorted(dict.fromkeys(datasets))) + + others = scores.drop(columns=[baseline], errors="ignore") + lower_better = metric in LOWER_IS_BETTER + best = others.min(axis=1) if lower_better else others.max(axis=1) + worst = others.max(axis=1) if lower_better else others.min(axis=1) + + # idxmin/idxmax raise on an all-NaN row, so only ask where there is a value. + winner = pd.Series(pd.NA, index=others.index, dtype=object) + present = others.notna().any(axis=1) + if present.any(): + rows = others[present] + winner[present] = rows.idxmin(axis=1) if lower_better else rows.idxmax(axis=1) + + summary = pd.DataFrame( + { + "baseline": scores[baseline] if baseline in scores else np.nan, + "median": others.median(axis=1), + "best": best, + "worst": worst, + "best_estimator": winner, + "estimators": others.notna().sum(axis=1), + } + ) + # The results files carry a "Resamples:" header, which pandas takes as the + # index name and would otherwise appear as the first column heading. + summary.index.name = "dataset" + # Gain is how much skill the best estimator found beyond the baseline; + # spread is how much the choice of estimator mattered. Reporting a single + # range would conflate the two, and since the baseline is almost always the + # weakest entry that range would just restate the gain. + summary["gain"] = ( + summary["baseline"] - summary["best"] + if lower_better + else summary["best"] - summary["baseline"] + ) + summary["spread"] = ( + summary["worst"] - summary["best"] + if lower_better + else summary["best"] - summary["worst"] + ) + return summary + + +def _dataset_table_html(summary, decimals) -> str: + """Render the per-dataset table, one row per dataset.""" + head = ( + '' + "Dataset" + 'Dummy' + 'Median' + 'Best' + '' + "Best estimator" + 'Gain over dummy' + 'Spread' + 'Estimators' + ) + + def number(value): + return "—" if pd.isna(value) else f"{value:.{decimals}f}" + + rows = [] + for dataset, row in summary.iterrows(): + # Two flags worth seeing at a glance: nothing beat the baseline by much, + # and everything solves it. Both make a dataset weak at separating + # estimators, for opposite reasons. + classes = [] + if pd.notna(row["gain"]) and row["gain"] <= NO_SIGNAL_GAIN: + classes.append("nosignal") + if pd.notna(row["best"]) and row["best"] >= SATURATED_BEST: + classes.append("saturated") + attribute = f' class="{" ".join(classes)}"' if classes else "" + winner = ( + "—" + if pd.isna(row["best_estimator"]) + else escape(str(row["best_estimator"])) + ) + rows.append( + f"{escape(str(dataset))}" + f'{number(row["baseline"])}' + f'{number(row["median"])}' + f'{number(row["best"])}' + f"{winner}" + f'{number(row["gain"])}' + f'{number(row["spread"])}' + f'{int(row["estimators"])}' + ) + + return ( + '
' + f"{head}" + f"{''.join(rows)}
" + ) + + +def dataset_page( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", + sort_by: str = "gain", + output_path: Path | str | None = None, + title: str | None = None, + decimals: int = 4, +) -> Path: + """Write a self-contained page summarising one metric per dataset. + + Parameters + ---------- + datasets : list of str + Datasets to include. + estimators : list of str + Estimator names, matching their results directories. + metric : str + Metric to summarise. + results_dir : Path or str + Directory holding one sub-directory per estimator. + baseline : str + Estimator treated as the no-skill floor. + sort_by : str + Column to order rows by, one of the columns of + :func:`dataset_summary`. The default puts the datasets where the best + estimator gained least over the baseline at the top, because those are + the ones worth looking at. + output_path : Path or str, optional + Where to write. Defaults to ``datasets.html`` beside the results. + title : str, optional + Page heading. + decimals : int + Decimal places for scores. + + Returns + ------- + Path + The file written. + """ + summary = dataset_summary(datasets, estimators, metric, results_dir, baseline) + if sort_by not in summary.columns: + raise ValueError(f"sort_by={sort_by!r} is not a column") + ascending = sort_by not in {"median", "best", "worst", "spread", "estimators"} + summary = summary.sort_values(sort_by, ascending=ascending, na_position="last") + + label = METRIC_LABELS.get(metric, metric) + title = title or f"Multiverse datasets: {label.lower()}" + scored = int((summary["estimators"] > 0).sum()) + no_signal = int(summary["gain"].le(NO_SIGNAL_GAIN).sum()) + saturated = int(summary["best"].ge(SATURATED_BEST).sum()) + + parts = [ + f"

{escape(title)}

", + f'

{scored} datasets · {escape(label.lower())}' + f" · best of up to {int(summary['estimators'].max())} estimators" + f" against the {escape(baseline)} baseline · built " + f"{date.today().isoformat()}

", + _dataset_table_html(summary, decimals), + '

One row per dataset. Dummy is the ' + "no-skill floor. Median, best and " + "spread are over the other estimators, so the baseline " + "cannot flatter them. Gain over dummy is best minus " + "dummy, how much skill was found at all; spread is " + "best minus worst, how much the choice of estimator mattered. The two " + "answer different questions, and a single range would conflate them.

", + f'

{no_signal} of {scored} datasets gained ' + f"{NO_SIGNAL_GAIN:.2f} or less over the baseline (shaded amber) and " + f"{saturated} have a best of {SATURATED_BEST:.2f} or more (shaded " + "green). Both separate estimators poorly, for opposite reasons. Best is " + "a maximum over many estimators, so it is optimistic by construction: " + "read it as what the archive can currently do on a problem, not as what " + "any one method delivers.

", + ] + + page = ( + '' + '' + f"{escape(title)}" + f"
{''.join(parts)}
" + f"" + ) + + output_path = ( + Path(results_dir) / "datasets.html" + if output_path is None + else Path(output_path) + ) + output_path.parent.mkdir(parents=True, exist_ok=True) + output_path.write_text(page, encoding="utf-8") + return output_path + + +def dataset_markdown( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", + sort_by: str = "gain", + decimals: int = 4, +) -> str: + """Return the per-dataset summary as a Markdown table.""" + summary = dataset_summary(datasets, estimators, metric, results_dir, baseline) + ascending = sort_by not in {"median", "best", "worst", "spread", "estimators"} + summary = summary.sort_values(sort_by, ascending=ascending, na_position="last") + + header = [ + "Dataset", "Dummy", "Median", "Best", "Best estimator", + "Gain over dummy", "Spread", "Estimators", + ] + rows = ["| " + " | ".join(header) + " |", "|" + "---|" * len(header)] + + def number(value): + return "—" if pd.isna(value) else f"{value:.{decimals}f}" + + for dataset, row in summary.iterrows(): + winner = ( + "—" + if pd.isna(row["best_estimator"]) + else str(row["best_estimator"]) + ) + cells = [ + str(dataset), + number(row["baseline"]), + number(row["median"]), + f'**{number(row["best"])}**', + winner, + number(row["gain"]), + number(row["spread"]), + str(int(row["estimators"])), + ] + rows.append("| " + " | ".join(cells) + " |") + + rows.append("") + rows.append( + f"{METRIC_LABELS.get(metric, metric)} per dataset. Median, best and " + f"spread are over the estimators other than {baseline}. Gain over dummy " + "is best minus dummy; spread is best minus worst." + ) + return "\n".join(rows) + + def main() -> None: """Build the Multiverse-core leaderboard. @@ -872,6 +1168,14 @@ def main() -> None: ) print(f"wrote {path}") + datasets_path = dataset_page( + datasets, + estimators, + metric="accuracy", + title="Multiverse-core datasets: accuracy", + ) + print(f"wrote {datasets_path}") + table = leaderboard_markdown(datasets, estimators, sort_by="accuracy") readme = Path(__file__).resolve().parents[2] / "README.md" if write_markdown_table(readme, table): diff --git a/multiverse/experiments/tests/test_tables.py b/multiverse/experiments/tests/test_tables.py index e6f6e0b..b580ec2 100644 --- a/multiverse/experiments/tests/test_tables.py +++ b/multiverse/experiments/tests/test_tables.py @@ -12,6 +12,9 @@ LOWER_IS_BETTER, METRIC_LABELS, available_estimators, + dataset_markdown, + dataset_page, + dataset_summary, leaderboard, leaderboard_markdown, load_metric, @@ -222,3 +225,71 @@ def test_write_markdown_table_without_markers(tmp_path): path.write_text("no markers here\n", encoding="utf-8") assert not write_markdown_table(path, "new", marker="T") assert path.read_text(encoding="utf-8") == "no markers here\n" + + +def test_dataset_summary_excludes_the_baseline(results_dir): + """Best, median and spread ignore the baseline; gain is measured against it.""" + summary = dataset_summary( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], baseline="Carol", + results_dir=results_dir, + ) + # d1: Alice 0.9, Bob 0.6, baseline Carol 0.3 + assert summary.at["d1", "best"] == pytest.approx(0.9) + assert summary.at["d1", "best_estimator"] == "Alice" + assert summary.at["d1", "worst"] == pytest.approx(0.6) + assert summary.at["d1", "baseline"] == pytest.approx(0.3) + assert summary.at["d1", "gain"] == pytest.approx(0.6) + assert summary.at["d1", "spread"] == pytest.approx(0.3) + assert summary.at["d1", "estimators"] == 2 + + +def test_dataset_summary_keeps_partly_covered_datasets(results_dir): + """A dataset only some estimators finished is kept, with a lower count. + + This is where it differs from the leaderboard, which has to drop d3 to keep + the averages comparable. + """ + summary = dataset_summary( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], results_dir=results_dir + ) + assert list(summary.index) == ["d1", "d2", "d3"] + # Carol has no d3, and is not the baseline here, so only two contribute + assert summary.at["d3", "estimators"] == 2 + assert summary.at["d1", "estimators"] == 3 + + +def test_dataset_summary_lower_is_better_flips_best_and_gain(results_dir): + """For log loss the best score is the smallest, and gain stays positive.""" + summary = dataset_summary( + ["d1"], ["Alice", "Bob", "Carol"], metric="logloss", baseline="Carol", + results_dir=results_dir, + ) + # stored as 1 - score, so Alice 0.1, Bob 0.4, Carol 0.7 + assert summary.at["d1", "best"] == pytest.approx(0.1) + assert summary.at["d1", "best_estimator"] == "Alice" + assert summary.at["d1", "gain"] == pytest.approx(0.6) + assert summary.at["d1", "spread"] == pytest.approx(0.3) + + +def test_dataset_page_is_written_and_sorted_by_gain(results_dir, tmp_path): + """The page lists every dataset, worst gain first.""" + path = dataset_page( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], baseline="Carol", + results_dir=results_dir, output_path=tmp_path / "datasets.html", + ) + html = path.read_text(encoding="utf-8") + for dataset in ["d1", "d2", "d3"]: + assert f"{dataset}" in html + # d3 has no baseline score, so its gain is NaN and it sorts last + assert html.index("d2") < html.index("d3") + assert "Gain over dummy" in html + + +def test_dataset_markdown_marks_the_best(results_dir): + """The Markdown table bolds the best score and names the estimator.""" + table = dataset_markdown( + ["d1"], ["Alice", "Bob", "Carol"], baseline="Carol", results_dir=results_dir + ) + assert "| d1 |" in table + assert "**0.9000**" in table + assert "Alice" in table diff --git a/results/multiverse/TimesNet/TimesNet_accuracy.csv b/results/multiverse/TimesNet/TimesNet_accuracy.csv new file mode 100644 index 0000000..8dcefce --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_accuracy.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.37209302325581395 +AppliancesEnergy_disc,0.7857142857142857 +ArticularyWordRecognition,0.9466666666666667 +AsphaltObstaclesCoordinates,0.7544757033248082 +AsphaltRegularityCoordinates,0.9440745672436751 +AtrialFibrillation,0.26666666666666666 +AustraliaRainfall_disc,0.7796634845365114 +AutomotiveRoadTrials,0.7792207792207793 +BIDMC32HR_disc,0.5839933305543976 +BIDMC32SpO2_disc,0.7061275531471446 +BeijingPM10Quality_disc,0.8181458003169572 +BeijingPM25Quality_disc,0.8645007923930269 +BenzeneConcentration_disc,0.9213635483246174 +Blink,0.7333333333333333 +BoneIntensitiesAgeGroup,0.7393258426966293 +BoneProbAgeGroup,0.5370786516853933 +CharacterTrajectories,0.9811977715877437 +CounterMovementJump,0.6983240223463687 +Cricket,0.9444444444444444 +CrowdSourced,0.7116131309565831 +DuckDuckGeese,0.42 +ERing,0.7592592592592593 +EigenWorms,0.4732824427480916 +Epilepsy,0.8985507246376812 +EthanolConcentration,0.2965779467680608 +EyesOpenShut,0.5952380952380952 +FaceDetection,0.6535187287173666 +FordChallenge,0.8897160187482768 +HandMovementDirection,0.5675675675675675 +Handwriting,0.17411764705882352 +Heartbeat,0.6829268292682927 +HouseholdPowerConsumption1_disc,0.9256559766763849 +HouseholdPowerConsumption2_disc,0.6895043731778425 +IEEEPPG_disc,0.3825301204819277 +IRDS-SFL,0.7758620689655172 +JapaneseVowels,0.9783783783783784 +KERAAL-RTK,0.5 +KIMORE-PR-C,0.2857142857142857 +KINECAL-QSEO,0.9411764705882353 +LSST,0.5762368207623683 +Libras,0.75 +Locust2022,0.9127008184298272 +LowCost,0.6333333333333333 +MindReading,0.39203675344563554 +MotionSenseHAR,0.8981132075471698 +MotorImagery,0.52 +NATOPS,0.8333333333333334 +PEMS-SF,0.7283236994219653 +PenDigits,0.9745568896512292 +PhonemeSpectra,0.1258574410975246 +PhotoStimulation,0.3888888888888889 +RacketSports,0.8355263157894737 +STEW,0.7365319865319865 +SelfRegulationSCP1,0.8191126279863481 +Skoda,0.9267775026529891 +SpokenArabicDigits,0.9813551614370168 +StandWalkJump,0.26666666666666666 +TactileTextureRecognition,0.8311306901615272 +Tiselac,0.7634984833164813 +UCDHE-Rowing-MC,0.7204545454545455 +UCIActivity,0.9779258642232403 +UIPRMD-DS-C,0.7222222222222222 +USCActivity,0.6661603888213852 +UWaveGestureLibrary,0.790625 +WISDM,0.8516515356384006 diff --git a/results/multiverse/TimesNet/TimesNet_auroc.csv b/results/multiverse/TimesNet/TimesNet_auroc.csv new file mode 100644 index 0000000..04d227f --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_auroc.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.5601764234161989 +AppliancesEnergy_disc,0.463235294117647 +ArticularyWordRecognition,0.9968055555555556 +AsphaltObstaclesCoordinates,0.9242467401525889 +AsphaltRegularityCoordinates,0.9878768532311839 +AtrialFibrillation,0.3466666666666667 +AustraliaRainfall_disc,0.8625349241770757 +AutomotiveRoadTrials,0.8312159709618875 +BIDMC32HR_disc,0.6773794380553011 +BIDMC32SpO2_disc,0.4833702778431914 +BeijingPM10Quality_disc,0.8749904474783253 +BeijingPM25Quality_disc,0.9309040794318134 +BenzeneConcentration_disc,0.9749213020371545 +Blink,0.97954 +BoneIntensitiesAgeGroup,0.892701881911134 +BoneProbAgeGroup,0.6959173444706285 +CharacterTrajectories,0.9990348617048694 +CounterMovementJump,0.8642332910817959 +Cricket,0.9995791245791246 +CrowdSourced,0.7412263913974378 +DuckDuckGeese,0.7205000000000001 +ERing,0.9686090534979424 +EigenWorms,0.718388522476044 +Epilepsy,0.9862550380827751 +EthanolConcentration,0.5745624783473889 +EyesOpenShut,0.7165532879818594 +FaceDetection,0.6969903795733101 +FordChallenge,0.9407409080023597 +HandMovementDirection,0.7984681908410722 +Handwriting,0.6858361166766375 +Heartbeat,0.7318634423897582 +HouseholdPowerConsumption1_disc,0.9695901816991374 +HouseholdPowerConsumption2_disc,0.6116172978480137 +IEEEPPG_disc,0.5765685291450916 +IRDS-SFL,0.8786231884057971 +JapaneseVowels,0.9993631911934565 +KERAAL-RTK,0.9166666666666667 +KIMORE-PR-C,0.5 +KINECAL-QSEO,0.5625 +LSST,0.853268607858199 +Libras,0.971924603174603 +Locust2022,0.8016958120164678 +LowCost,0.7135999999999999 +MindReading,0.7258524604266321 +MotionSenseHAR,0.9868619219805685 +MotorImagery,0.4888 +NATOPS,0.9690000000000001 +PEMS-SF,0.9488123159450785 +PenDigits,0.999078948440171 +PhonemeSpectra,0.737837970891607 +PhotoStimulation,0.4712067562067562 +RacketSports,0.9445903678347245 +STEW,0.8172149305122556 +SelfRegulationSCP1,0.953499207902339 +Skoda,0.9935043476694889 +SpokenArabicDigits,0.9997081716255138 +StandWalkJump,0.5199999999999999 +TactileTextureRecognition,0.9936774041312487 +Tiselac,0.9380726632403095 +UCDHE-Rowing-MC,0.9353138674388676 +UCIActivity,0.9981610285888065 +UIPRMD-DS-C,0.8703703703703703 +USCActivity,0.9506426879444319 +UWaveGestureLibrary,0.9689062500000001 +WISDM,0.9484526588220853 diff --git a/results/multiverse/TimesNet/TimesNet_balacc.csv b/results/multiverse/TimesNet/TimesNet_balacc.csv new file mode 100644 index 0000000..100692b --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_balacc.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.33164983164983164 +AppliancesEnergy_disc,0.4852941176470588 +ArticularyWordRecognition,0.9466666666666668 +AsphaltObstaclesCoordinates,0.7521205393470547 +AsphaltRegularityCoordinates,0.9441796126835498 +AtrialFibrillation,0.26666666666666666 +AustraliaRainfall_disc,0.4140487445355887 +AutomotiveRoadTrials,0.588021778584392 +BIDMC32HR_disc,0.4709239620965495 +BIDMC32SpO2_disc,0.538099345749419 +BeijingPM10Quality_disc,0.7747954805109453 +BeijingPM25Quality_disc,0.8457038067403321 +BenzeneConcentration_disc,0.8871910848591745 +Blink,0.7595000000000001 +BoneIntensitiesAgeGroup,0.7920760052555407 +BoneProbAgeGroup,0.45603264482193634 +CharacterTrajectories,0.9797871384824239 +CounterMovementJump,0.6987758945386066 +Cricket,0.9444444444444445 +CrowdSourced,0.711667045440953 +DuckDuckGeese,0.42000000000000004 +ERing,0.7592592592592592 +EigenWorms,0.2979108813891423 +Epilepsy,0.8982511923688394 +EthanolConcentration,0.29545454545454547 +EyesOpenShut,0.5952380952380952 +FaceDetection,0.6535187287173667 +FordChallenge,0.8783134492990511 +HandMovementDirection,0.5892857142857143 +Handwriting,0.16666662298642781 +Heartbeat,0.6401730678046468 +HouseholdPowerConsumption1_disc,0.8643621873914809 +HouseholdPowerConsumption2_disc,0.4565518195655182 +IEEEPPG_disc,0.4120209782526767 +IRDS-SFL,0.8125 +JapaneseVowels,0.9787152311290241 +KERAAL-RTK,0.5625 +KIMORE-PR-C,0.5833333333333334 +KINECAL-QSEO,0.5 +LSST,0.4061515272163288 +Libras,0.75 +Locust2022,0.5177733270660045 +LowCost,0.6333333333333333 +MindReading,0.3821385274639657 +MotionSenseHAR,0.85686005458408 +MotorImagery,0.52 +NATOPS,0.8333333333333334 +PEMS-SF,0.7323219775393687 +PenDigits,0.9744859719656194 +PhonemeSpectra,0.1258585008243011 +PhotoStimulation,0.3111111111111111 +RacketSports,0.8468023255813953 +STEW,0.7365319865319866 +SelfRegulationSCP1,0.8195881092162892 +Skoda,0.9179357910721214 +SpokenArabicDigits,0.9813594852635947 +StandWalkJump,0.26666666666666666 +TactileTextureRecognition,0.8295600900292777 +Tiselac,0.5977147607084359 +UCDHE-Rowing-MC,0.7287063492063492 +UCIActivity,0.9784012153274504 +UIPRMD-DS-C,0.7222222222222222 +USCActivity,0.6668699118002607 +UWaveGestureLibrary,0.790625 +WISDM,0.5384253425988083 diff --git a/results/multiverse/TimesNet/TimesNet_f1.csv b/results/multiverse/TimesNet/TimesNet_f1.csv new file mode 100644 index 0000000..b55a80b --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_f1.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.3083127164769916 +AppliancesEnergy_disc,0.0 +ArticularyWordRecognition,0.9447295291707192 +AsphaltObstaclesCoordinates,0.7498447599468028 +AsphaltRegularityCoordinates,0.9436997319034852 +AtrialFibrillation,0.21666666666666667 +AustraliaRainfall_disc,0.7637884682119469 +AutomotiveRoadTrials,0.32 +BIDMC32HR_disc,0.5585945105904826 +BIDMC32SpO2_disc,0.2227122381477398 +BeijingPM10Quality_disc,0.6810284920083391 +BeijingPM25Quality_disc,0.7807692307692308 +BenzeneConcentration_disc,0.8628378378378379 +Blink,0.7683397683397684 +BoneIntensitiesAgeGroup,0.7292711888413573 +BoneProbAgeGroup,0.511495225510458 +CharacterTrajectories,0.9812348407458454 +CounterMovementJump,0.7074097798246634 +Cricket,0.9435536685536685 +CrowdSourced,0.749770290964778 +DuckDuckGeese,0.41054342260513366 +ERing,0.7404213649857618 +EigenWorms,0.35942968120789176 +Epilepsy,0.8924464174838986 +EthanolConcentration,0.19858702091342648 +EyesOpenShut,0.711864406779661 +FaceDetection,0.6722147651006711 +FordChallenge,0.8504113687359761 +HandMovementDirection,0.5618750618750619 +Handwriting,0.1330347969490157 +Heartbeat,0.4881889763779528 +HouseholdPowerConsumption1_disc,0.9257408724640781 +HouseholdPowerConsumption2_disc,0.652006321163392 +IEEEPPG_disc,0.36228767112726407 +IRDS-SFL,0.6176470588235294 +JapaneseVowels,0.9784002492198859 +KERAAL-RTK,0.631578947368421 +KIMORE-PR-C,0.2857142857142857 +KINECAL-QSEO,0.0 +LSST,0.5394750991542188 +Libras,0.7420271537372987 +Locust2022,0.07096774193548387 +LowCost,0.6518987341772152 +MindReading,0.37141073070989444 +MotionSenseHAR,0.8997451310063029 +MotorImagery,0.5636363636363636 +NATOPS,0.8313231547362717 +PEMS-SF,0.7236481018309302 +PenDigits,0.9744994689537161 +PhonemeSpectra,0.12036859740669961 +PhotoStimulation,0.23333333333333334 +RacketSports,0.8309665853150886 +STEW,0.7163141993957703 +SelfRegulationSCP1,0.8408408408408409 +Skoda,0.9265318723123889 +SpokenArabicDigits,0.9814030544189566 +StandWalkJump,0.2545454545454546 +TactileTextureRecognition,0.8283142005725257 +Tiselac,0.7566452965306562 +UCDHE-Rowing-MC,0.7202788045572184 +UCIActivity,0.9778189209589807 +UIPRMD-DS-C,0.6153846153846154 +USCActivity,0.6554462665298106 +UWaveGestureLibrary,0.7704187365622013 +WISDM,0.844569181289972 diff --git a/results/multiverse/TimesNet/TimesNet_logloss.csv b/results/multiverse/TimesNet/TimesNet_logloss.csv new file mode 100644 index 0000000..5cd770e --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_logloss.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,1.248236038572558 +AppliancesEnergy_disc,0.5412785260443004 +ArticularyWordRecognition,0.8492526636352818 +AsphaltObstaclesCoordinates,0.6551940357196788 +AsphaltRegularityCoordinates,0.1688360559340955 +AtrialFibrillation,1.1981288914883097 +AustraliaRainfall_disc,0.5103721333391383 +AutomotiveRoadTrials,0.4623138921081696 +BIDMC32HR_disc,6.590572726578991 +BIDMC32SpO2_disc,3.972382463186422 +BeijingPM10Quality_disc,0.5895476296760733 +BeijingPM25Quality_disc,0.45336789876064526 +BenzeneConcentration_disc,0.3598162728366507 +Blink,0.6506866480407071 +BoneIntensitiesAgeGroup,0.9006484097441354 +BoneProbAgeGroup,0.9204488497638516 +CharacterTrajectories,0.09339775284240061 +CounterMovementJump,1.2205767845095714 +Cricket,0.18759665772286246 +CrowdSourced,3.7190997533835852 +DuckDuckGeese,1.488468220036912 +ERing,1.1844290967923863 +EigenWorms,2.103077859528844 +Epilepsy,0.28349653897231647 +EthanolConcentration,1.8423970303040027 +EyesOpenShut,0.665966461485736 +FaceDetection,0.7907227657391813 +FordChallenge,0.4523257742272296 +HandMovementDirection,1.1498558573304878 +Handwriting,3.0647765836939316 +Heartbeat,0.5865046912663875 +HouseholdPowerConsumption1_disc,0.3529202338987952 +HouseholdPowerConsumption2_disc,1.907156836363749 +IEEEPPG_disc,3.591613029958346 +IRDS-SFL,0.6226220421832996 +JapaneseVowels,0.0889182589816358 +KERAAL-RTK,1.1093077996849414 +KIMORE-PR-C,2.0333026937729164 +KINECAL-QSEO,0.4393486385757983 +LSST,1.433947541079781 +Libras,0.813702952447647 +Locust2022,0.28520599640093425 +LowCost,1.0222239818187968 +MindReading,3.015668150293372 +MotionSenseHAR,1.6326868318251477 +MotorImagery,1.045675160199672 +NATOPS,0.3580698083671133 +PEMS-SF,0.7503315018015176 +PenDigits,0.0988062312152402 +PhonemeSpectra,7.437383773494004 +PhotoStimulation,1.3357289698174262 +RacketSports,0.4435775543274873 +STEW,1.29413291863661 +SelfRegulationSCP1,0.3899661776899016 +Skoda,0.2711256899608417 +SpokenArabicDigits,0.09320711216230121 +StandWalkJump,1.121926403618857 +TactileTextureRecognition,2.005757775262857 +Tiselac,1.357585809737937 +UCDHE-Rowing-MC,1.512952259612872 +UCIActivity,0.11104400509082067 +UIPRMD-DS-C,0.5731393849864266 +USCActivity,1.776896473816155 +UWaveGestureLibrary,0.8681988024148264 +WISDM,2.09257477768673 diff --git a/results/multiverse/TimesNet/TimesNet_sensitivity.csv b/results/multiverse/TimesNet/TimesNet_sensitivity.csv new file mode 100644 index 0000000..824b7f9 --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_sensitivity.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.37209302325581395 +AppliancesEnergy_disc,0.0 +ArticularyWordRecognition,0.9466666666666667 +AsphaltObstaclesCoordinates,0.7544757033248082 +AsphaltRegularityCoordinates,0.9513513513513514 +AtrialFibrillation,0.26666666666666666 +AustraliaRainfall_disc,0.7796634845365114 +AutomotiveRoadTrials,0.21052631578947367 +BIDMC32HR_disc,0.5839933305543976 +BIDMC32SpO2_disc,0.1478770131771596 +BeijingPM10Quality_disc,0.6721536351165981 +BeijingPM25Quality_disc,0.7981651376146789 +BenzeneConcentration_disc,0.7971285892634207 +Blink,0.995 +BoneIntensitiesAgeGroup,0.7393258426966293 +BoneProbAgeGroup,0.5370786516853933 +CharacterTrajectories,0.9811977715877437 +CounterMovementJump,0.6983240223463687 +Cricket,0.9444444444444444 +CrowdSourced,0.864406779661017 +DuckDuckGeese,0.42 +ERing,0.7592592592592593 +EigenWorms,0.4732824427480916 +Epilepsy,0.8985507246376812 +EthanolConcentration,0.2965779467680608 +EyesOpenShut,1.0 +FaceDetection,0.7105561861520999 +FordChallenge,0.8320526893523601 +HandMovementDirection,0.5675675675675675 +Handwriting,0.17411764705882352 +Heartbeat,0.543859649122807 +HouseholdPowerConsumption1_disc,0.9256559766763849 +HouseholdPowerConsumption2_disc,0.6895043731778425 +IEEEPPG_disc,0.3825301204819277 +IRDS-SFL,0.875 +JapaneseVowels,0.9783783783783784 +KERAAL-RTK,1.0 +KIMORE-PR-C,1.0 +KINECAL-QSEO,0.0 +LSST,0.5762368207623683 +Libras,0.75 +Locust2022,0.03754266211604096 +LowCost,0.6866666666666666 +MindReading,0.39203675344563554 +MotionSenseHAR,0.8981132075471698 +MotorImagery,0.62 +NATOPS,0.8333333333333334 +PEMS-SF,0.7283236994219653 +PenDigits,0.9745568896512292 +PhonemeSpectra,0.1258574410975246 +PhotoStimulation,0.3888888888888889 +RacketSports,0.8355263157894737 +STEW,0.6652637485970819 +SelfRegulationSCP1,0.958904109589041 +Skoda,0.9267775026529891 +SpokenArabicDigits,0.9813551614370168 +StandWalkJump,0.26666666666666666 +TactileTextureRecognition,0.8311306901615272 +Tiselac,0.7634984833164813 +UCDHE-Rowing-MC,0.7204545454545455 +UCIActivity,0.9779258642232403 +UIPRMD-DS-C,0.4444444444444444 +USCActivity,0.6661603888213852 +UWaveGestureLibrary,0.790625 +WISDM,0.8516515356384006 diff --git a/results/multiverse/TimesNet/TimesNet_specificity.csv b/results/multiverse/TimesNet/TimesNet_specificity.csv new file mode 100644 index 0000000..c5032a7 --- /dev/null +++ b/results/multiverse/TimesNet/TimesNet_specificity.csv @@ -0,0 +1,66 @@ +Resamples:,0 +Alzheimers,0.37209302325581395 +AppliancesEnergy_disc,0.9705882352941176 +ArticularyWordRecognition,0.9466666666666667 +AsphaltObstaclesCoordinates,0.7544757033248082 +AsphaltRegularityCoordinates,0.937007874015748 +AtrialFibrillation,0.26666666666666666 +AustraliaRainfall_disc,0.7796634845365114 +AutomotiveRoadTrials,0.9655172413793104 +BIDMC32HR_disc,0.5839933305543976 +BIDMC32SpO2_disc,0.9283216783216783 +BeijingPM10Quality_disc,0.8774373259052924 +BeijingPM25Quality_disc,0.8932424758659853 +BenzeneConcentration_disc,0.9772535804549284 +Blink,0.524 +BoneIntensitiesAgeGroup,0.7393258426966293 +BoneProbAgeGroup,0.5370786516853933 +CharacterTrajectories,0.9811977715877437 +CounterMovementJump,0.6983240223463687 +Cricket,0.9444444444444444 +CrowdSourced,0.5589273112208892 +DuckDuckGeese,0.42 +ERing,0.7592592592592593 +EigenWorms,0.4732824427480916 +Epilepsy,0.8985507246376812 +EthanolConcentration,0.2965779467680608 +EyesOpenShut,0.19047619047619047 +FaceDetection,0.5964812712826334 +FordChallenge,0.9245742092457421 +HandMovementDirection,0.5675675675675675 +Handwriting,0.17411764705882352 +Heartbeat,0.7364864864864865 +HouseholdPowerConsumption1_disc,0.9256559766763849 +HouseholdPowerConsumption2_disc,0.6895043731778425 +IEEEPPG_disc,0.3825301204819277 +IRDS-SFL,0.75 +JapaneseVowels,0.9783783783783784 +KERAAL-RTK,0.125 +KIMORE-PR-C,0.16666666666666666 +KINECAL-QSEO,1.0 +LSST,0.5762368207623683 +Libras,0.75 +Locust2022,0.998003992015968 +LowCost,0.58 +MindReading,0.39203675344563554 +MotionSenseHAR,0.8981132075471698 +MotorImagery,0.42 +NATOPS,0.8333333333333334 +PEMS-SF,0.7283236994219653 +PenDigits,0.9745568896512292 +PhonemeSpectra,0.1258574410975246 +PhotoStimulation,0.3888888888888889 +RacketSports,0.8355263157894737 +STEW,0.8078002244668911 +SelfRegulationSCP1,0.6802721088435374 +Skoda,0.9267775026529891 +SpokenArabicDigits,0.9813551614370168 +StandWalkJump,0.26666666666666666 +TactileTextureRecognition,0.8311306901615272 +Tiselac,0.7634984833164813 +UCDHE-Rowing-MC,0.7204545454545455 +UCIActivity,0.9779258642232403 +UIPRMD-DS-C,1.0 +USCActivity,0.6661603888213852 +UWaveGestureLibrary,0.790625 +WISDM,0.8516515356384006 diff --git a/results/multiverse/datasets.html b/results/multiverse/datasets.html new file mode 100644 index 0000000..5302b7b --- /dev/null +++ b/results/multiverse/datasets.html @@ -0,0 +1,119 @@ +Multiverse-core datasets: accuracy

Multiverse-core datasets: accuracy

66 datasets · accuracy · best of up to 24 estimators against the Dummy baseline · built 2026-09-03

DatasetDummyMedianBestBest estimatorGain over dummySpreadEstimators
AtrialFibrillation0.33330.23330.3333CIF0.00000.266724
KINECAL-QSEO0.94120.94120.9412Arsenal0.00000.117624
BIDMC32SpO2_disc0.71530.67030.7203ROCKET0.00500.171322
Locust20220.91120.90850.9206MRHydra0.00940.045224
Heartbeat0.72200.74390.7854CIF0.06340.126824
HouseholdPowerConsumption2_disc0.72160.76750.7872HC20.06560.141424
AutomotiveRoadTrials0.75320.79870.8442CIF0.09090.233824
AustraliaRainfall_disc0.68600.77460.7808LITETime-MV0.09480.088116
EyesOpenShut0.50000.50000.5952STSF0.09520.190524
MotorImagery0.50000.51000.6000FreshPRINCE0.10000.140024
BeijingPM10Quality_disc0.71120.82420.8417FreshPRINCE0.13050.107024
Alzheimers0.41860.37210.5581MRHydra0.13950.302323
AppliancesEnergy_disc0.80950.83330.9524FreshPRINCE0.14290.452424
EmoPain0.78310.84080.92681NN-DTW0.14370.242320
PhotoStimulation0.41670.38890.5833ROCKET0.16670.388923
FaceDetection0.50000.63250.6850H-InceptionTime0.18500.158623
BeijingPM25Quality_disc0.69770.87510.8879ConvTran0.19020.128824
HouseholdPowerConsumption1_disc0.77840.91470.9825FreshPRINCE0.20410.218724
LowCost0.50000.63750.7300TSF0.23000.248324
BoneProbAgeGroup0.47640.65170.7124H-InceptionTime0.23600.193324
StandWalkJump0.33330.43330.6000MRHydra0.26670.400024
CrowdSourced0.50020.71270.7734LITETime-MV0.27320.176824
BenzeneConcentration_disc0.68970.82010.9768STSF0.28700.577024
FordChallenge0.62320.88960.9360QUANT0.31280.312823
BIDMC32HR_disc0.65070.80830.9637RIST0.31300.635322
STEW0.50000.74020.8385Arsenal0.33850.210022
BoneIntensitiesAgeGroup0.47640.79890.8202HC20.34380.296624
PhonemeSpectra0.02560.27890.3746H-InceptionTime0.34890.291124
KERAAL-RTK0.57140.78570.9286HC20.35710.571424
LSST0.31510.63080.7040FreshPRINCE0.38890.480924
HandMovementDirection0.20270.41890.6081TSF0.40540.418924
DuckDuckGeese0.20000.46000.6400H-InceptionTime0.44000.480024
SelfRegulationSCP10.50170.85320.9454MRHydra0.44370.208224
IEEEPPG_disc0.26050.43980.7078ConvTran0.44730.438324
AsphaltRegularityCoordinates0.50730.97870.9947H-InceptionTime0.48740.291624
MindReading0.23120.52830.7243LITETime-MV0.49310.332324
EthanolConcentration0.25100.43350.7490STC0.49810.532324
UIPRMD-DS-C0.50000.83331.0000Catch220.50000.388924
WISDM0.36640.86580.8965MRHydra0.53000.131024
Blink0.44440.99221.0000Arsenal0.55560.415624
EigenWorms0.41980.89310.9771MRHydra0.55730.557323
KIMORE-PR-C0.14290.42860.7143LITETime-MV0.57140.571424
AsphaltObstaclesCoordinates0.28390.82100.8670MRHydra0.58310.289024
CounterMovementJump0.33520.74860.9274Arsenal0.59220.458124
Handwriting0.03760.37290.6529H-InceptionTime0.61530.478824
USCActivity0.11380.69240.7354LITETime-MV0.62160.137920
RacketSports0.28290.88160.9079RDST0.62500.125024
UCDHE-Rowing-MC0.20450.73410.8295PatchMTSC0.62500.338624
IRDS-SFL0.20690.78880.8621RDST0.65520.448324
Skoda0.23560.94730.9646H-InceptionTime0.72900.119923
Epilepsy0.26810.98551.0000HC20.73190.101424
Tiselac0.06280.81860.8373STSF0.77450.204419
MotionSenseHAR0.20380.98681.0000DrCIF0.79620.101924
NATOPS0.16670.88610.9667LITETime-MV0.80000.155624
UCIActivity0.19160.97520.9983LITETime-MV0.80670.174924
UWaveGestureLibrary0.12500.90620.9406Arsenal0.81560.553124
ERing0.16670.93330.9963MRHydra0.82960.237024
PEMS-SF0.11560.93061.0000CIF0.88440.317924
PenDigits0.10380.97780.9911H-InceptionTime0.88740.234422
SpokenArabicDigits0.10000.97820.9941RDST0.89400.128724
Libras0.06670.88890.9722RIST0.90560.338924
JapaneseVowels0.08380.96080.9946LiteTIME0.91080.208124
Cricket0.08330.97921.00001NN-DTW0.91670.069424
CharacterTrajectories0.06480.98890.9958H-InceptionTime0.93110.044624
TactileTextureRecognition0.05140.99851.0000H-InceptionTime0.94860.168924
ArticularyWordRecognition0.04000.98170.9933Arsenal0.95330.050024

One row per dataset. Dummy is the no-skill floor. Median, best and spread are over the other estimators, so the baseline cannot flatter them. Gain over dummy is best minus dummy, how much skill was found at all; spread is best minus worst, how much the choice of estimator mattered. The two answer different questions, and a single range would conflate them.

4 of 66 datasets gained 0.05 or less over the baseline (shaded amber) and 15 have a best of 0.99 or more (shaded green). Both separate estimators poorly, for opposite reasons. Best is a maximum over many estimators, so it is optimistic by construction: read it as what the archive can currently do on a problem, not as what any one method delivers.

\ No newline at end of file diff --git a/results/multiverse/leaderboard.html b/results/multiverse/leaderboard.html index fdccb82..229d26c 100644 --- a/results/multiverse/leaderboard.html +++ b/results/multiverse/leaderboard.html @@ -60,12 +60,12 @@ details { margin-top: .6rem; } summary { cursor: pointer; color: var(--accent); } code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .9em; } -

Multiverse-core leaderboard

24 estimators on 52 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-01

#EstimatorAccuracyBalanced accuracyAUROCF1Log loss ↓SensitivitySpecificity
ScoreRankScoreRankScoreRankScoreRankScoreRankScoreRankScoreRank
1HC20.79097.860.75188.490.89906.060.72738.260.53836.130.74598.660.79437.46
2MRHydra0.78378.310.75647.970.810516.720.73167.857.797419.740.76428.200.77579.12
3RDST0.77349.160.73339.720.791217.370.699110.038.166719.970.710910.680.78748.73
4RIST0.77209.760.739710.580.87488.030.714710.380.62188.310.740810.740.765510.80
5DrCIF0.774710.070.742910.520.88138.380.717310.860.64849.170.739711.400.770810.90
6FreshPRINCE0.774310.100.748710.480.87457.940.721110.380.60076.210.741410.870.777010.97
7CIF0.778110.130.747110.380.89088.250.721210.390.64309.210.744111.200.775310.28
8QUANT0.772010.550.746210.480.88317.550.718910.570.61757.370.752110.680.758111.56
9Arsenal0.768010.700.732110.570.845713.060.702410.693.863116.800.725710.780.773210.53
10ROCKET0.769010.820.732610.610.792518.120.701911.098.324920.750.720011.660.776411.02
11LITETime-MV0.750611.140.72999.790.85139.930.68209.871.320612.150.71329.700.763711.10
12STSF0.772411.550.747711.180.88049.830.708011.750.64328.380.734512.100.782612.19
13H-InceptionTime0.740811.630.719010.830.849610.380.683810.511.322712.900.722310.330.737812.43
14LiteTIME0.734112.340.710411.470.839411.520.668011.741.477613.150.711310.570.733612.21
15PatchMTSC0.742813.120.689713.930.826112.590.653313.490.76559.440.685213.150.735212.98
16ConvTran0.746213.140.710213.340.859211.020.676712.840.81909.520.715912.740.734513.56
17Catch220.747513.150.718113.610.869710.740.692213.620.714710.790.724013.540.737413.62
18STC0.754513.950.717214.130.874411.310.694013.960.63919.920.718514.110.753713.89
19TSF0.751514.000.723613.560.874011.580.688313.910.725210.500.709314.600.760613.87
20TDE0.726214.630.681314.790.837412.450.638214.250.886911.500.671413.820.734413.07
21Summary0.685816.660.657416.400.826815.290.623016.400.912313.250.657416.350.684416.85
22TimesURL0.695816.950.653317.050.790617.710.596717.361.005515.480.625717.320.697316.01
231NN-DTW0.671218.450.645417.630.719721.240.613617.6311.850622.930.652116.330.663618.53
24Dummy0.364521.810.302922.500.500022.950.150722.171.406716.400.285520.480.381618.33

Average score and average rank over the 52 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.

Missing results

Scoring uses the 52 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.

Reproducing this page

from aeon.datasets.tsc_datasets import multiverse_core
+

Multiverse-core leaderboard

25 estimators on 52 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-03

#EstimatorAccuracyBalanced accuracyAUROCF1Log loss ↓SensitivitySpecificity
ScoreRankScoreRankScoreRankScoreRankScoreRankScoreRankScoreRank
1HC20.79098.060.75188.710.89906.210.72738.460.53836.270.74598.910.79437.66
2MRHydra0.78378.510.75648.170.810517.380.73168.057.797420.640.76428.430.77579.36
3RDST0.77349.350.73339.940.791218.040.699110.238.166720.840.710910.960.78748.90
4RIST0.77209.970.739710.790.87488.220.714710.560.62188.520.740811.040.765511.08
5DrCIF0.774710.250.742910.760.88138.560.717311.060.64849.400.739711.690.770811.14
6FreshPRINCE0.774310.300.748710.660.87458.090.721110.580.60076.370.741411.150.777011.21
7CIF0.778110.370.747110.610.89088.370.721210.610.64309.460.744111.510.775310.53
8QUANT0.772010.810.746210.720.88317.700.718910.790.61757.580.752110.990.758111.85
9Arsenal0.768010.960.732110.810.845713.510.702410.933.863117.590.725711.060.773210.76
10ROCKET0.769011.090.732610.880.792518.810.701911.358.324921.630.720011.960.776411.29
11LITETime-MV0.750611.500.729910.080.851310.180.682010.121.320612.600.71329.970.763711.42
12STSF0.772411.790.747711.380.880410.020.708011.930.64328.630.734512.400.782612.44
13H-InceptionTime0.740811.930.719011.080.849610.600.683810.771.322713.400.722310.550.737812.75
14LiteTIME0.734112.680.710411.760.839411.800.668012.031.477613.630.711310.810.733612.55
15PatchMTSC0.742813.380.689714.230.826112.880.653313.780.76559.600.685213.440.735213.17
16ConvTran0.746213.390.710213.590.859211.250.676713.070.81909.710.715913.030.734513.81
17Catch220.747513.420.718113.920.869710.930.692213.900.714711.040.724013.880.737413.95
18STC0.754514.250.717214.450.874411.580.694014.300.639110.190.718514.420.753714.20
19TSF0.751514.260.723613.840.874011.850.688314.240.725210.750.709314.970.760614.12
20TDE0.726215.070.681315.170.837412.840.638214.670.886911.900.671414.260.734413.41
21Summary0.685817.180.657416.880.826815.810.623016.880.912313.650.657416.940.684417.38
22TimesNet0.701317.260.665917.340.827516.060.628117.581.161314.690.672616.600.689817.34
23TimesURL0.695817.460.653317.580.790618.330.596717.861.005515.940.625717.840.697316.51
241NN-DTW0.671219.080.645418.220.719722.120.613618.1711.850623.880.652116.880.663619.12
25Dummy0.364522.680.302923.430.500023.880.150723.101.406717.080.285521.290.381619.05

Average score and average rank over the 52 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.

Missing results

  • HC2 — AustraliaRainfall_disc (Time limit); STEW, Tiselac, USCActivity (cancelled before completion)
  • MRHydra — AustraliaRainfall_disc (OOM at 128GB); PenDigits (ValueError: n_timepoints must be >= 9, but found 8); Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • RDST — AustraliaRainfall_disc, Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738)); USCActivity (OOM at 64GB)
  • FreshPRINCE — FaceDetection, FordChallenge, Skoda, Tiselac (OOM at 128GB)
  • ROCKET — AustraliaRainfall_disc (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • STSF — AustraliaRainfall_disc, PenDigits (not recorded)
  • LiteTIME — BIDMC32HR_disc, BIDMC32SpO2_disc, USCActivity (not recorded)
  • PatchMTSC — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • ConvTran — Alzheimers, EigenWorms, PhotoStimulation (CUDA out of memory); EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • TSF — AustraliaRainfall_disc (not recorded)
  • TDE — AustraliaRainfall_disc, Tiselac, USCActivity (Time limit); STEW (cancelled before completion)
  • Summary — AustraliaRainfall_disc (not recorded)
  • TimesNet — EmoPain (reason not recorded)
  • TimesURL — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • 1NN-DTW — BIDMC32HR_disc (Time limit); BIDMC32SpO2_disc (not recorded)

Scoring uses the 52 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.

Reproducing this page

from aeon.datasets.tsc_datasets import multiverse_core
 from multiverse.experiments.tables import leaderboard
 
 leaderboard(
     datasets=sorted(multiverse_core),
-    estimators=["HC2", "MRHydra", "RDST", "RIST", "DrCIF", "FreshPRINCE", "CIF", "QUANT", "Arsenal", "ROCKET", "LITETime-MV", "STSF", "H-InceptionTime", "LiteTIME", "PatchMTSC", "ConvTran", "Catch22", "STC", "TSF", "TDE", "Summary", "TimesURL", "1NN-DTW", "Dummy"],
+    estimators=["HC2", "MRHydra", "RDST", "RIST", "DrCIF", "FreshPRINCE", "CIF", "QUANT", "Arsenal", "ROCKET", "LITETime-MV", "STSF", "H-InceptionTime", "LiteTIME", "PatchMTSC", "ConvTran", "Catch22", "STC", "TSF", "TDE", "Summary", "TimesNet", "TimesURL", "1NN-DTW", "Dummy"],
     metrics=["accuracy", "balacc", "auroc", "f1", "logloss", "sensitivity", "specificity"],
     sort_by="accuracy",
 )

Or python -m multiverse.experiments.tables to rebuild it with the defaults.