Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 37 additions & 24 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,30 +40,31 @@ The current paper version describes:
<!-- LEADERBOARD:START -->
| # | Estimator | Accuracy rank | Accuracy | Balanced accuracy | AUROC | F1 | Log loss &darr; | Sensitivity | Specificity |
|---|---|---|---|---|---|---|---|---|---|
| 1 | HC2 | **7.86** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** |
| 2 | MRHydra | 8.31 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 |
| 3 | RDST | 9.16 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 |
| 4 | RIST | 9.76 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 |
| 5 | DrCIF | 10.07 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 |
| 6 | FreshPRINCE | 10.10 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 |
| 7 | CIF | 10.13 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 |
| 8 | QUANT | 10.55 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 |
| 9 | Arsenal | 10.70 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 |
| 10 | ROCKET | 10.82 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 |
| 11 | LITETime-MV | 11.14 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 |
| 12 | STSF | 11.55 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 |
| 13 | H-InceptionTime | 11.63 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 |
| 14 | LiteTIME | 12.34 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 |
| 15 | PatchMTSC | 13.12 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 |
| 16 | ConvTran | 13.14 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 |
| 17 | Catch22 | 13.15 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 |
| 18 | STC | 13.95 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 |
| 19 | TSF | 14.00 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 |
| 20 | TDE | 14.63 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 |
| 21 | Summary | 16.66 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 |
| 22 | TimesURL | 16.95 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 |
| 23 | 1NN-DTW | 18.45 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 |
| 24 | Dummy | 21.81 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 |
| 1 | HC2 | **8.06** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** |
| 2 | MRHydra | 8.51 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 |
| 3 | RDST | 9.35 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 |
| 4 | RIST | 9.97 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 |
| 5 | DrCIF | 10.25 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 |
| 6 | FreshPRINCE | 10.30 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 |
| 7 | CIF | 10.37 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 |
| 8 | QUANT | 10.81 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 |
| 9 | Arsenal | 10.96 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 |
| 10 | ROCKET | 11.09 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 |
| 11 | LITETime-MV | 11.50 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 |
| 12 | STSF | 11.79 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 |
| 13 | H-InceptionTime | 11.93 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 |
| 14 | LiteTIME | 12.68 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 |
| 15 | PatchMTSC | 13.38 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 |
| 16 | ConvTran | 13.39 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 |
| 17 | Catch22 | 13.42 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 |
| 18 | STC | 14.25 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 |
| 19 | TSF | 14.26 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 |
| 20 | TDE | 15.07 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 |
| 21 | Summary | 17.18 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 |
| 22 | TimesNet | 17.26 | 0.7013 | 0.6659 | 0.8275 | 0.6281 | 1.1613 | 0.6726 | 0.6898 |
| 23 | TimesURL | 17.46 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 |
| 24 | 1NN-DTW | 19.08 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 |
| 25 | Dummy | 22.68 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 |

Average over the 52 Multiverse-core datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold.
<!-- LEADERBOARD:END -->
Expand All @@ -74,6 +75,14 @@ version with per-metric ranks to
([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/leaderboard.html),
since GitHub shows HTML as source). Missing results, and why, are listed on that page.

The same command writes a per-dataset view to
[`results/multiverse/datasets.html`](results/multiverse/datasets.html)
([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/datasets.html)),
which turns the question around: for each dataset it gives the Dummy floor, the median
and best over the other estimators, which estimator was best, how much the best gained
over Dummy, and how far apart the estimators were. It is sorted by that gain, so the
problems where nothing yet beats the baseline come first.

This repository aims to make it easier to:

- load Multiverse datasets through `aeon`
Expand All @@ -91,6 +100,10 @@ This repository aims to make it easier to:
·
<a href="docs/leaderboard.md">Leaderboard</a>
·
<a href="docs/runtime.md">Runtime</a>
·
<a href="docs/memory.md">Memory</a>
·
<a href="docs/evaluation.md">Evaluation</a>
·
<a href="docs/classifiers.md">Classifiers</a>
Expand Down
48 changes: 42 additions & 6 deletions docs/classifiers.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,14 +114,50 @@ Mathematics, 9(23), 2021.
This is the only Keras port here, following the authors, so it needs `tensorflow`
rather than `torch`. Both are in the `deep-learning` extra.

The authors tune `window_size` per dataset. Their results table carries a `Win_pct`
column spread over a five point grid: 20, 40 and 60 on five datasets each, 80 on
thirteen, and 100 on two. The default here is **0.8**, the value they use most often.
The 0.2 in their `config.yml` is the worked example for BasicMotions, not a default.
### Tuning, and what we report

**The XCM results in this repository follow the authors' protocol.** Section 4.3 sets
`window_size` and `batch_size` per dataset "by grid search based on the best average
accuracy following a stratified 5-fold cross-validation on the training set", over
windows {0.2, 0.4, 0.6, 0.8, 1.0} and batches {1, 8, 32}. Selection never touches the
test data, so the published figures are tuned but not leaked, and neither are ours.

The reported run searches the window on that grid and holds batch size at 32. That is
the one departure, and it is a cost decision rather than a modelling one: batch 1 takes
roughly 32 times the gradient steps, which would turn a day of GPU time into about 900
hours, for a value the published table selects on 4 of 30 datasets.

Both parameters accept a sequence, which triggers the search; a scalar fits once. The
class default is a single fit at **0.8**, the modal published window, because a default
should be cheap, but `XCM` in `tsml-eval` supplies the grid, and `XCM-Fixed` is the
single-fit variant kept for comparison:

```python
from multiverse.classification import XCMClassifier
from multiverse.classification._xcm import PAPER_WINDOW_SIZES, PAPER_BATCH_SIZES

XCMClassifier(window_size=PAPER_WINDOW_SIZES) # window only
XCMClassifier(window_size=PAPER_WINDOW_SIZES, batch_size=PAPER_BATCH_SIZES) # full grid
```

The selected values are on `window_fraction_` and `batch_size_`, and every grid point's
mean and per-fold accuracy on `cv_results_`.

Cost is the reason the class default is a single fit. A fixed-window pass over
Multiverse-core took 1.2 GPU-hours in total; searching the window is five candidates
over five folds, about 20 times that.

That earlier fixed-window pass is what motivated the change. It averaged 0.699 across
the 23 datasets shared with the paper's table against their 0.761, and the paper itself
reports a mean relative accuracy drop of 7.0% +/- 1.3% from using a suboptimal window,
which is the size of the gap observed. Reporting a fixed window would have measured a
configuration the authors never used.

Because `window_size` is a fraction, the kernel grows with the series, and 0.8 of
EigenWorms' 17984 points would be a 14387 point kernel. `max_window` bounds the kernel
at 100 points, and it is floored at 1 for very short series.
EigenWorms' 17984 points is a 14387 point kernel. `max_window` bounds it at 100 points
and floors it at 1 for very short series. That bound is ours, not the authors': they run
kernels of this order, 40% of EigenWorms being 7193 points. Set `max_window=None` to
reproduce them, and expect the memory cost to follow.

## Notes on the ports

Expand Down
92 changes: 92 additions & 0 deletions docs/memory.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
# Memory

**Coming soon.** This page will hold memory comparisons across the Multiverse
estimators.

**We have not yet structured an experiment to compare memory.** As with
[runtime](runtime.md), every run behind the results in this repository was set up to
measure predictive performance. A job was given whatever memory ceiling got it to
finish, on whatever node was free, with whatever core count came with the partition,
and those were never held constant across estimators because nothing depended on it.
The peak figures that came out are a by-product of that, not a measurement anyone
designed.

So nothing is published here yet, deliberately. The rest of this page records what a
memory comparison would have to fix, and why the figures we already hold cannot stand
in for one.

## What is recorded

`tsml-eval` writes one `memory_usage` value per classifier, dataset and resample: the
peak memory observed during `fit`. It is in the raw prediction files, and this
repository's ingest brings across only the accuracy-style measures, so it has not been
carried over.

One number, host side, fit only.

## Why the figures we hold cannot stand in

Each of these is a condition a memory experiment would have to fix, and that these runs
left free.

**Peak process memory is not the model's memory.** It includes the interpreter, every
imported library, the loaded dataset and any transient copies made along the way.
Importing TensorFlow or PyTorch alone accounts for a large fixed cost before a single
weight is allocated, so a small model in a heavy framework can report more than a large
model in a light one. Without subtracting a per-framework baseline the figure mostly
ranks frameworks.

**The dataset can dominate the model.** Multiverse contains problems like EigenWorms at
17,984 timepoints and FaceDetection at 5,890 training cases. For an estimator with a
small parameter count, most of the peak is the data and its copies, which says
something about the problem rather than the method.

**Host and device memory are different quantities, and only one is recorded.** A model
doing its work on a GPU can look inexpensive by peak host memory while occupying tens of
gigabytes of device memory that nothing here measures. The two are not interchangeable
and cannot be added.

**Framework allocators do not report what the model needs.** PyTorch's caching allocator
and TensorFlow's default of reserving most of the visible GPU both hold memory they are
not using, so a naive reading measures the allocator's policy rather than the model's
requirement. Getting a meaningful device figure means enabling TensorFlow's memory
growth and reading PyTorch's allocated rather than reserved totals, neither of which
these runs did.

**Failures censor the measurement.** Where an estimator ran out of memory we do not have
a peak, we have a lower bound and a ceiling. Both kinds appear in
`results/multiverse/missing_results.csv`: ConvTran hit CUDA out of memory on Alzheimers,
EigenWorms and PhotoStimulation, a device-side limit; FreshPRINCE hit OOM at 128 GB on
FaceDetection, FordChallenge and Skoda after eight attempts, a host-side one. Those are
the cases where memory mattered most, and they are exactly the cases with no number. A
table built only from successful runs is a survivorship-biased view of memory use.

**Memory scales with the resources granted.** The classical ensembles allocate per
thread, so their peak moves with the cores allocated, which varied by partition. Our
controllers also escalate a job's request from 64 GB to 128 GB after a failure, so
different runs of the same estimator saw different ceilings.

**Peak is timing-dependent.** Python's peak resident memory depends on when garbage
collection happens to run and on whether an allocator returned pages to the operating
system. Repeats of an identical run differ for reasons that have nothing to do with the
method.

## What a fair comparison would need

- All compared estimators on the same node, with cores and the memory ceiling fixed and
recorded.
- A per-framework baseline measured and subtracted, so the figure is the model's cost
rather than the cost of importing its library.
- Host and device peaks reported separately, never summed, with TensorFlow memory growth
enabled and PyTorch read via its allocated totals.
- `predict` measured as well as `fit`. Deployment cost is a separate question from
training cost, and the ranking is not the same on both.
- Failures reported alongside successes, as censored observations with the ceiling that
was in force, rather than dropped.
- Repeated runs, since peak varies between identical repeats.
- Asymptotic space complexity in the number of cases, series length and channels stated
beside the measured peaks, so a reader can tell whether a figure will hold at a
different scale.

Until most of that is in place, this page stays empty. For the same reason the
[leaderboard](leaderboard.md) carries no memory column.
87 changes: 87 additions & 0 deletions docs/runtime.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
# Runtime

**Coming soon.** This page will hold runtime comparisons across the Multiverse
estimators. Memory is a separate page, [memory](memory.md), for the same reasons and a
few of its own.

**We have not yet structured an experiment to compare runtime.** Every run behind the
results in this repository was set up to measure predictive performance. Which partition
a classifier was queued on, how many cores it was given, how many epochs it trained for,
whether a job was retried at a higher memory ceiling: all of those were chosen to get
accurate results out at a reasonable cost, and none were held constant across estimators
because nothing depended on it. The timings that came out are a by-product of that, not
a measurement anyone designed.

So nothing is published here yet, deliberately. Comparing runtime needs its own
experiment, with the conditions below fixed in advance, and we have not run one. The
rest of this page records what those conditions are, and why the figures we already hold
cannot stand in for them.

## The measurements exist

Every run already records timings. `tsml-eval` writes, per classifier, dataset and
resample:

- `fit_time` and `predict_time`, in seconds;
- `benchmark_time`, the time that machine took to sort 1,000 seeded random arrays of
20,000 elements;
- `memory_usage`, the peak memory during `fit`, which [memory](memory.md) covers.

They are in the raw prediction files. What this repository ingests under `results/` is
only the accuracy-style measures, one file per metric, so the timings have not been
brought across yet. That is a small piece of work; the reason it has not been done is
below, not the effort.

## Why the figures we hold cannot stand in

Each of these is a condition a timing experiment would have to fix, and that these runs
left free.

**The runs are spread across different hardware.** Multiverse results have been produced
on GPU partitions with H200 and A100 cards and on CPU-only nodes, with different core
counts. GPU jobs in our configurations are allocated two CPUs each. A fit time from one
partition and a fit time from another are two different measurements that happen to share
a unit.

**GPU and CPU methods are not on one axis.** For the deep learners nearly all the work is
on the accelerator and the host CPU mostly feeds batches; for the classical ensembles
there is no accelerator at all and the time scales with the cores allocated. Comparing
them measures the hardware at least as much as the algorithm, and the ratio moves when
either side changes. A statement like "X is 40 times faster than Y" is, in this setting,
a statement about a purchasing decision.

**Wall-clock contains things that are not the algorithm.** Queueing, data loading, and
retries: our controllers escalate a job's memory request from 64 GB to 128 GB after a
failure, so an elapsed time can include a dead run at the lower ceiling.

**Training time is a hyperparameter, not a property of a method.** A deep learner's fit
time is close to linear in the number of epochs, and the epoch count is a choice. Two
faithful ports of the same paper can differ several-fold on time because the authors
picked 500 epochs and the toolkit's default is 2000. Early stopping and best-epoch
selection move it again. None of that is a fact about the architecture.

**Which device a run actually used is not reliably recorded.** In the version of the
experiment tooling used for these runs, the device description inspects TensorFlow only,
so a PyTorch estimator reports CPU whether or not it ran on a GPU. Any timing table built
from those records has to have its device column reconstructed from the job
configuration rather than trusted as written.

## What a fair comparison would need

- The compared estimators run on the same hardware, or CPU timings normalised by
`benchmark_time`, which exists for exactly this purpose. There is no equivalent
normaliser for GPU work.
- Thread and core counts fixed and recorded, since the classical methods scale with them.
- `fit` and `predict` reported separately. They answer different questions: fit time is
the cost of research, predict time is the cost of deployment, and the ranking is not
the same on both.
- Repeated runs. Timings vary far more between repeats than accuracy does, especially on
shared nodes.
- Asymptotic complexity in the number of cases, series length and channels reported
beside the measured times, so a reader can tell whether a result will hold at a
different scale.
- The device stated per run, from the job configuration.

Until most of that is in place, this page stays empty. For the same reason the
[leaderboard](leaderboard.md) carries no fit time or predict time columns, and
[memory](memory.md) is empty too.
2 changes: 2 additions & 0 deletions multiverse/classification/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@

__all__ = [
"ConvTranClassifier",
"DisjointCNNClassifier",
"PatchMTSCClassifier",
"TimesNetClassifier",
"TS2VecClassifier",
Expand All @@ -13,6 +14,7 @@
]

from multiverse.classification._convtran import ConvTranClassifier
from multiverse.classification._disjoint_cnn import DisjointCNNClassifier
from multiverse.classification._patchmtsc import PatchMTSCClassifier
from multiverse.classification._timesnet import TimesNetClassifier
from multiverse.classification._ts2vec import TS2VecClassifier
Expand Down
Loading
Loading