Skip to content
Merged

Docs #26

Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 25 additions & 10 deletions docs/classifiers.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,13 +114,11 @@ Mathematics, 9(23), 2021.
This is the only Keras port here, following the authors, so it needs `tensorflow`
rather than `torch`. Both are in the `deep-learning` extra.

### Tuning, and what we report

**The XCM results in this repository follow the authors' protocol.** Section 4.3 sets
The XCM results in this repository follow the authors' tuning protocol. Section 4.3 sets
`window_size` and `batch_size` per dataset "by grid search based on the best average
accuracy following a stratified 5-fold cross-validation on the training set", over
windows {0.2, 0.4, 0.6, 0.8, 1.0} and batches {1, 8, 32}. Selection never touches the
test data, so the published figures are tuned but not leaked, and neither are ours.
test data.

The reported run searches the window on that grid and holds batch size at 32. That is
the one departure, and it is a cost decision rather than a modelling one: batch 1 takes
Expand Down Expand Up @@ -161,12 +159,29 @@ reproduce them, and expect the memory cost to follow.

## Notes on the ports

All three wrappers take aeon's ``numpy3D`` collections, shape
All classifiers in this package take aeon's ``numpy3D`` collections, shape
``(n_cases, n_channels, n_timepoints)``, and transpose internally where the original
network expects a different layout. Each holds out ``validation_size`` of the training
data inside ``fit`` and restores the best epoch, so no external validation split is
required. The networks require ``torch``, which is not a hard dependency of this
package: install it with ``pip install aeon-multiverse[deep-learning]``.
network expects a different layout. Validation and epoch selection differ by port:

* ConvTran and PatchMTSC split off ``validation_size`` inside ``fit`` and restore the
epoch with the lowest validation loss. With ``validation_size=0`` they select on
training loss instead.
* TimesNet splits internally and restores the epoch with the best validation accuracy
when a nonzero split is feasible. With ``validation_size=0`` (or fewer than two
cases) it selects on training loss instead.
* DisjointCNN passes a sampled ``validation_size`` set to Keras, but samples with
replacement from the training collection and leaves those cases in training. It
restores the best monitored weights, but this is not a held-out validation split.
* XCM has no validation split or best-epoch restoration: it trains for a fixed number
of epochs. Its optional cross-validation selects hyperparameters, not an epoch.
* TimesURL and TS2Vec pretrain on the full training collection and fit their probe on
the resulting training representations; neither has an internal validation split or
best-epoch restoration.

Thus no external validation split is required by these wrappers, but the same
validation procedure does not apply to every classifier. The networks require
``torch``, which is not a hard dependency of this package: install it with
``pip install aeon-multiverse[deep-learning]``.

### What was ported, and from where

Expand Down Expand Up @@ -216,7 +231,7 @@ faithful transcription.
| ``probability=True`` on the SVM probe | TS2Vec | The authors' grid sets it False, which leaves an ``SVC`` unable to produce probability estimates. aeon classifiers must implement ``predict_proba``, so it is enabled, adding Platt scaling fitted by internal cross-validation on the training data |
| Layer imports taken from ``tensorflow.keras.layers`` | XCM | The original imports ``Conv1D`` and ``Conv2D`` from ``keras.layers.convolutional``, a path removed in Keras 3. The layers and their arguments are unchanged |
| Kernel length floored at one point | XCM | The original computes ``int(window_size * n)``, which is zero for series shorter than five points and builds an invalid layer |
| Validation split moved inside ``fit`` | all three | The originals split train/validation outside the model, which risks leakage between train and test. TSLib is explicit about it: ``exp_classification.py`` sets ``vali_data = self._get_data(flag='TEST')``, so it selects the retained epoch on the test set. See the note at the top of this page |
| Validation split moved inside ``fit`` | ConvTran, PatchMTSC, TimesNet | The originals split train/validation outside the model, which risks leakage between train and test. TSLib is explicit about it: ``exp_classification.py`` sets ``vali_data = self._get_data(flag='TEST')``, so it selects the retained epoch on the test set. See the note at the top of this page |
| Test data scaled with training statistics | TimesNet | TSLib fits its normaliser separately per split, so its test set is scaled by its own statistics. The port fits on train and applies to test |

Note that PatchMTSC's two graph blocks are, in the original, distinguished only by a
Expand Down
29 changes: 27 additions & 2 deletions docs/datasets.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ Unequal length datasets are stored in a list of 2D numpy arrays. You can control
whether to load the equal length version with the parameter ``load_equal_length``.

```python
X,y = load_classification("JapaneseVowels", load_equal_length = False) # Unequal length example
X,y = load_classification("JapaneseVowels", load_equal_length = True) # Unequal loaded by default from 1.5
```
Imputed missing value versions can be loaded with the argument ``load_no_missing``.
You can download whole archives from zenodo or in code
Expand All @@ -69,9 +69,34 @@ download_archive(archive="UEA", extract_path="C:\\Temp\\")

```
Currently should be one of "EEG","UCR","UEA","Imbalanced","TSR", "Unequal". See
``aeon`` documentation for more details. There are lists of datasets in aeon and a
``aeon`` documentation for more details. There are lists of datasets in aeon and a
dictionary of all zenodo keys.

```python
from aeon.datasets.tsc_datasets import multiverse_core, multiverse2026, eeg2026, tsc_zenodo
```

## Dataset collections used here

The project uses the archive collections exposed by aeon:

```python
from aeon.datasets.tsc_datasets import multiverse_core, multiverse2026, eeg2026

print(len(multiverse_core)) # 66
print(len(multiverse2026)) # 133
print(len(eeg2026)) # 28
```

`multiverse2026` is the full Multiverse collection. `multiverse_core` is the smaller
benchmark subset: it is more balanced across applications, removes overly similar,
very simple and zero-information datasets, and has a useful spread of dataset sizes and
series lengths. The current benchmark results use this core list unless stated
otherwise.

The EEG collection is a separate classification archive used for EEG-specific
experiments. It is based on [aeon-neuro](https://github.com/aeon-toolkit/aeon-neuro).

These lists are Python collections of dataset names, so they can be passed directly to
experiment or result-loading code. The underlying dataset files are still downloaded
and cached using `load_classification`, as described above.
207 changes: 201 additions & 6 deletions docs/evaluation.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,206 @@
# Experimental Protocols
# Experimental protocols

There are many variations on how people structure experiments and a range of metrics
used in comparison.
This repository uses [`tsml-eval`](https://github.com/time-series-machine-learning/tsml-eval)
to run classification experiments and store predictions. `tsml-eval` is the evaluation
toolkit used by the time-series machine-learning projects around aeon. It loads the
train and test files for a dataset, fits an aeon-compatible classifier, measures the
run, writes predictions and probabilities, and provides utilities for comparing
classifiers over datasets and resamples.

The results we present are for the moment the simplest: we do all
training/validation on the default train split, and evaluate once on the test set.
The basic protocol here is deliberately simple: fit on the supplied training split,
then evaluate once on the supplied test split. Resample `0` means the archive's default
train/test split. Additional resample IDs can be run when repeated evaluation is
required.

## Installation

There are alternatives: we could perform stratified resamples or cross validate.
The experiment drivers are an optional dependency because running experiments is not
needed to import the classifiers or read the result tables:

```bash
pip install -e ".[experiments]"
```

At present, the released `tsml-eval` package pins an older aeon range than this project
uses. If pip reports an aeon dependency conflict, install a checkout of the `main`
branch of `tsml-eval` until a compatible release is available:

```bash
git clone https://github.com/time-series-machine-learning/tsml-eval.git
pip install -e ./tsml-eval
```

Deep-learning classifiers also need the project's deep-learning extra:

```bash
pip install -e ".[deep-learning]"
```

## Run one experiment

The smallest complete example is
[`run_single_dataset.py`](../multiverse/experiments/run_single_dataset.py). Set the
dataset path and output path, then choose an archive dataset and a classifier name:

```python
from tsml_eval.experiments import (
get_classifier_by_name,
load_and_run_classification_experiment,
)

classifier_name = "ROCKET"
classifier = get_classifier_by_name(classifier_name, random_state=0)

load_and_run_classification_experiment(
problem_path="C:/Data/Multiverse",
results_path="./results-raw",
dataset="BasicMotions",
classifier=classifier,
classifier_name=classifier_name,
resample_id=0,
overwrite=False,
)
```

`problem_path` must contain the standard archive layout:

```text
<problem_path>/<dataset>/<dataset>_TRAIN.ts
<problem_path>/<dataset>/<dataset>_TEST.ts
```

The function loads those files, fits only on `_TRAIN.ts`, predicts `_TEST.ts`, and
writes:

```text
<results_path>/<classifier>/Predictions/<dataset>/testResample<id>.csv
```

With `overwrite=False`, an existing result is retained, which makes interrupted batch
runs safe to restart. The same call can be put inside loops over classifiers, datasets,
and resample IDs; see `run_benchmark.py` for that pattern.

## Result-file format

The output is a tsml-format classification CSV, rather than a normal rectangular table.
It is intended to be read with `load_classifier_results`, not parsed with
`pandas.read_csv`.

The file contains:

1. A metadata line containing the dataset, classifier, split (`TEST`), resample ID,
time unit, and a description.
2. A parameter-information line containing the estimator configuration.
3. A summary line containing accuracy, fit time, predict time, benchmark time, memory
usage, number of classes, and optional train-error-estimation fields.
4. One line per test case containing the true class, predicted class, one probability
for each class, prediction time, and an optional case description.

For example, a result for `ROCKET` on `BasicMotions` at resample `0` is found at:

```text
results-raw/ROCKET/Predictions/BasicMotions/testResample0.csv
```

Load and score one file as follows:

```python
from tsml_eval.evaluation.storage import load_classifier_results

result = load_classifier_results(
"results-raw/ROCKET/Predictions/BasicMotions/testResample0.csv"
)
result.calculate_statistics()
print(result.accuracy)
print(result.balanced_accuracy)
print(result.fit_time, result.predict_time)
```

The stored probabilities allow metrics such as AUROC and log loss
to be calculated later.

## Collate and compare results with tsml-eval

Once one result file exists for every requested classifier, dataset, and resample, use
`evaluate_classifiers_by_problem` to collate and compare them:

```python
from tsml_eval.evaluation import evaluate_classifiers_by_problem

evaluate_classifiers_by_problem(
load_path="./results-raw",
classifier_names=["ROCKET", "DrCIF", "ConvTran"],
dataset_names=["BasicMotions", "ItalyPowerDemand", "Trace"],
save_path="./evaluations",
resamples=1, # evaluates resample 0
eval_name="example",
continue_on_missing=False,
)
```

The evaluator finds files using the standard directory layout and writes an evaluation
directory containing per-metric CSV files, summary CSV files with mean scores and mean
ranks, and comparison figures. Set `resamples=30` to evaluate IDs `0` through `29`, or
pass an explicit list such as `[0, 1, 2]`. By default a missing file is an error; use
`continue_on_missing=True` when deliberately allowing incomplete comparisons. The
default behaviour removes incomplete datasets from summary comparisons, which keeps
all classifiers on the same set of completed problems.

## Collate into this repository's result tables

The repository's `ingest.py` converts the raw tsml-eval files into the smaller tables
used by the leaderboard. It loads each file with `load_classifier_results`, calculates
the standard metrics, and writes one file per classifier and metric:

```text
results/multiverse/<classifier>/<classifier>_<metric>.csv
```

For example:

```python
from multiverse.experiments.ingest import ingest

ingest(
classifier="ROCKET",
predictions_path="./results-raw",
datasets=["BasicMotions", "ItalyPowerDemand", "Trace"],
resample=0,
)
```

The generated metric files have one row per dataset and one column per resample. Their
index is labelled `Resamples:`; when there is one resample, the single column is usually
`0`. Missing prediction files are left out rather than filled with a score, so failures
remain visible and can be reported separately.

To ingest the default Multiverse-core dataset list and the configured classifiers, edit
the settings at the top of `multiverse/experiments/ingest.py` and run:

```bash
python -m multiverse.experiments.ingest
```

Finally, build the HTML leaderboard from the collated tables:

```bash
python -m multiverse.experiments.tables
```

The lower-level table API is also available for inspection:

```python
from multiverse.experiments.tables import load_metric, leaderboard

accuracy = load_metric("ROCKET", "accuracy")
leaderboard(
datasets=["BasicMotions", "ItalyPowerDemand", "Trace"],
estimators=["ROCKET", "DrCIF", "ConvTran"],
metrics=["accuracy", "balacc", "logloss"],
output_path="./results/multiverse/leaderboard.html",
)
```

The ingested tables are a convenient
summary for this repository's leaderboards and should be regenerated if raw results are
changed.
46 changes: 3 additions & 43 deletions docs/leaderboard.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# Leaderboards

The leaderboards can be interactively generated on the WEBSITE. These are some
The leaderboards can be interactively generated on the WEBSITE COMING SOON. These are some
illustrative static leaderboards ranked on classification accuracy. We will embed
the interactive version and update this dynamic in time.
the interactive version when its ready. In the interim, we present some generative tools.

## Generating a leaderboard

Expand All @@ -29,9 +29,7 @@ are used, so each column describes the same problems; anything left out is liste
the page with the reason taken from `results/multiverse/missing_results.csv`.

A critical difference diagram can be added with `critical_difference=True`. It is off
by default because on the current results the omnibus Friedman test does not reject
over the leading estimators, so the diagram is a single clique and shows nothing the
table does not.
by default because we generate the front page table for all estimators.

Building the same page from the command line:

Expand All @@ -57,41 +55,3 @@ Every column in the generated table is sortable: click a heading to sort by it,
click again to reverse. The first click puts the best value on top, so ascending for
ranks and for log loss, descending for the rest.

**A warning on `max_cd_estimators`.** Truncating the critical difference diagram to the
best `n` estimators changes the statistics rather than just hiding rows. Ranks, the
omnibus test and the corrected alpha are all computed over the subset shown. The
diagram starts with an omnibus Friedman test, and dropping the weakest estimators
compresses the spread of average ranks, which can take that test from rejecting to not
rejecting. When Friedman does not reject, aeon places every estimator in a single clique
and runs no pairwise tests at all, so no differences appear.

On the current Multiverse-core results this is not hypothetical:

| Estimators in the diagram | Friedman p | Outcome |
|---|---|---|
| top 6 | 0.44 | one clique, no pairwise tests |
| top 8 | 0.14 | one clique, no pairwise tests |
| all 10 | 0.0003 | 7 significant pairs at alpha/(k-1) = 0.011 |

Treat a truncated diagram as a statement about that subset only.

`available_estimators()` lists the estimators that have results, and `load_metric()`
returns one estimator's scores for one metric as a `pandas.Series` if you would rather
build your own table.

## Multiverse

## Multiverse-core


## EEG archive

The EEG archive is a collection of EEG classification problems, described in [1]. On
release, it contains 30 datasets. Two of these are univariate and two are not
available on zenodo. The resulting list is contained in the multiverse


## UEA archive

People will still use the UEA archive, so it is worth maintaining a list for sanity
checks. The archive contains 30 datasets, but
7 changes: 2 additions & 5 deletions docs/memory.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,11 +18,8 @@ in for one.
## What is recorded

`tsml-eval` writes one `memory_usage` value per classifier, dataset and resample: the
peak memory observed during `fit`. It is in the raw prediction files, and this
repository's ingest brings across only the accuracy-style measures, so it has not been
carried over.

One number, host side, fit only.
peak memory observed during `fit`. It is in the raw prediction files, but it is only a
proxy for the memory footprint of a classifier.

## Why the figures we hold cannot stand in

Expand Down
Loading
Loading