From 110c776d21e8e529a45daa6a2df9b5c67e91efb5 Mon Sep 17 00:00:00 2001 From: Tony Bagnall Date: Tue, 1 Sep 2026 18:23:50 +0100 Subject: [PATCH 1/8] Port Disjoint-CNN, following the authors' training procedure Set up so our port and aeon's DisjointCNNClassifier can be run against each other under tsml-eval, to attribute the gap in aeon issue #3775, where aeon scored 20.3 accuracy points below the published numbers on all 23 shared UEA datasets. Comparing the two graphs found the likely cause, and it is not the training procedure. aeon applies a Permute after the final block, so the tensor entering the pooling is (time, filters, 1) rather than (time, 1, filters). GlobalAveragePooling2D reduces the two leading axes, so aeon's classifier head receives one scalar per case where it should receive 64 features. Its Dense(128) has 256 parameters, 1*128 + 128, which confirms the input width is 1. That explains the shape of the failure in the issue: accuracy correlates -0.425 with class count and collapses toward chance on the many-class problems, because a single scalar cannot separate 26 handwriting classes. Head to head on ERing, 200 epochs, default split: aeon 0.641, this port 0.952, published 0.964. The port is the authors' Keras network transcribed, with their training procedure: class-weighted loss, batch size min(n_cases // 10, 8), 500 epochs, and learning rate schedule and retained epoch both monitoring validation loss. aeon does none of those four. They are documented on the class so that a run of the two can separate the head bug from the training differences. Two things recorded rather than reproduced. The authors' DCNN_4L applies the third block's spatial convolution to conv2 rather than to its own temporal convolution, so a layer is computed and discarded; this port builds the clean stack, matching aeon, so the architecture is not a second moving part. And the data pipeline is not a difference at all: Main.py passes normalise=False, so the authors train on raw series, as aeon does. Registered in tsml-eval as DisjointCNN-MV, beside aeon's DisjointCNN rather than replacing it. Co-Authored-By: Claude Opus 5 --- multiverse/classification/__init__.py | 2 + multiverse/classification/_disjoint_cnn.py | 389 ++++++++++++++++++ .../tests/test_ported_classifiers.py | 9 + 3 files changed, 400 insertions(+) create mode 100644 multiverse/classification/_disjoint_cnn.py diff --git a/multiverse/classification/__init__.py b/multiverse/classification/__init__.py index cb5dd87..c97a596 100644 --- a/multiverse/classification/__init__.py +++ b/multiverse/classification/__init__.py @@ -5,6 +5,7 @@ __all__ = [ "ConvTranClassifier", + "DisjointCNNClassifier", "PatchMTSCClassifier", "TimesNetClassifier", "TS2VecClassifier", @@ -13,6 +14,7 @@ ] from multiverse.classification._convtran import ConvTranClassifier +from multiverse.classification._disjoint_cnn import DisjointCNNClassifier from multiverse.classification._patchmtsc import PatchMTSCClassifier from multiverse.classification._timesnet import TimesNetClassifier from multiverse.classification._ts2vec import TS2VecClassifier diff --git a/multiverse/classification/_disjoint_cnn.py b/multiverse/classification/_disjoint_cnn.py new file mode 100644 index 0000000..eb4af7e --- /dev/null +++ b/multiverse/classification/_disjoint_cnn.py @@ -0,0 +1,389 @@ +"""Disjoint-CNN classifier for aeon. + +Adapted from the authors' Disjoint-CNN implementation: +https://github.com/Navidfoumani/Disjoint-CNN + +Disjoint-CNN factorises a multivariate convolution into two disjoint steps: a +temporal convolution applied within each channel, then a spatial convolution +across channels, with a non-linearity between them. Stacking these 1+1D blocks +is the paper's alternative to a single joint convolution over both axes. + +This exists alongside ``aeon.classification.deep_learning.DisjointCNNClassifier`` +deliberately. That implementation scores far below the authors' published +numbers, 20.3 accuracy points below on the 23 shared UEA datasets in our runs +(aeon issue #3775). This port follows the authors' own training procedure rather +than aeon's, so the two can be run against each other and the gap attributed. +The training differences are recorded under "Deviations from aeon" below; the +data pipeline is not one of them, since the authors pass ``normalise=False`` in +``Main.py`` and so train on raw series, as aeon does. + +This wrapper is designed for aeon and therefore assumes input X is a 3D NumPy +array with shape (n_cases, n_channels, n_timepoints). The original expects +(n_cases, n_timepoints, n_channels, 1), so X is transposed and expanded +internally. + +The original source is distributed under the MIT License. + +MIT License + +Copyright (c) 2021 Navid Mohammadi Foumani + +Permission is hereby granted, free of charge, to any person obtaining a copy +of this software and associated documentation files (the "Software"), to deal +in the Software without restriction, including without limitation the rights +to use, copy, modify, merge, publish, distribute, sublicense, and/or sell +copies of the Software, and to permit persons to whom the Software is +furnished to do so, subject to the following conditions: + +The above copyright notice and this permission notice shall be included in all +copies or substantial portions of the Software. + +THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR +IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, +FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE +AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER +LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, +OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE +SOFTWARE. +""" + +from __future__ import annotations + +__maintainer__ = ["TonyBagnall"] +__all__ = ["DisjointCNNClassifier"] + +import math + +import numpy as np +from aeon.classification import BaseClassifier +from sklearn.utils import check_random_state + +# Temporal kernel length per block, from the authors' DCNN_2L, DCNN_3L and +# DCNN_4L. The kernel shortens as the stack deepens. +KERNEL_SIZES = {2: (8, 5), 3: (8, 5, 3), 4: (8, 5, 5, 3)} +# Pool size after the last block, which also differs per variant. +POOL_SIZES = {2: 3, 3: 3, 4: 5} + + +def _build_disjoint_cnn(n_timepoints, n_channels, n_classes, n_layers, n_filters): + """Build the authors' Disjoint-CNN network. + + A transcription of ``classifiers/DCNN_{2,3,4}L.py``. Each block is a + temporal convolution over time within a channel, then a spatial convolution + that spans every channel at once, then a permute that puts the filter axis + back where the next block's spatial convolution expects it. + + The authors' arrays are (case, time, channel), from + ``utils/data_loader.py::process_ts_data``, so time is axis 1 and the channel + axis is the one the spatial convolution spans entirely. + """ + from tensorflow.keras.layers import ( + BatchNormalization, + Conv2D, + Dense, + ELU, + GlobalAveragePooling2D, + Input, + MaxPooling2D, + Permute, + ) + from tensorflow.keras.models import Model + + kernels = KERNEL_SIZES[n_layers] + input_layer = Input((n_timepoints, n_channels, 1)) + + x = input_layer + for block, kernel in enumerate(kernels): + # Temporal: (kernel, 1) slides along time inside each channel. + x = Conv2D( + n_filters, + (kernel, 1), + strides=1, + padding="same", + kernel_initializer="he_uniform", + )(x) + x = BatchNormalization()(x) + x = ELU(alpha=1.0)(x) + + # Spatial: (1, width) spans the whole channel axis in one filter, so + # padding is valid and the axis collapses to length 1. + width = int(x.shape[2]) + x = Conv2D( + n_filters, + (1, width), + strides=1, + padding="valid", + kernel_initializer="he_uniform", + )(x) + x = BatchNormalization()(x) + x = ELU(alpha=1.0)(x) + + # The authors permute after every block except the last. + if block < len(kernels) - 1: + x = Permute((1, 3, 2))(x) + + x = MaxPooling2D(pool_size=(POOL_SIZES[n_layers], 1), strides=None, padding="valid")(x) + x = GlobalAveragePooling2D()(x) + output_layer = Dense(n_classes, activation="softmax")(x) + return Model(inputs=input_layer, outputs=output_layer) + + +def _class_weights(one_hot, mu=2.0): + """Reproduce ``utils/classifier_tools.py::create_class_weight``. + + Weight for a class is ``log(mu * total / count)``, floored at 1.0, so rare + classes are upweighted and common ones are never downweighted below 1. + """ + counts = one_hot.sum(axis=0) + total = counts.sum() + weights = {} + for index, count in enumerate(counts): + score = math.log(mu * total / float(count)) if count > 0 else 1.0 + weights[index] = score if score > 1.0 else 1.0 + return weights + + +class DisjointCNNClassifier(BaseClassifier): + """Disjoint-CNN, following the authors' training procedure. + + A stack of 1+1D blocks: a temporal convolution within each channel, then a + spatial convolution across all channels, with ELU and batch normalisation + between. The stack is max pooled, globally average pooled and classified. + + Parameters + ---------- + n_layers : int, default=4 + Number of disjoint blocks, 2, 3 or 4, selecting the authors' DCNN_2L, + DCNN_3L or DCNN_4L. The temporal kernel lengths and the final pool size + follow the variant. + n_filters : int, default=64 + Filters in every convolution. + n_epochs : int, default=500 + Training epochs. The authors' ``Main.py`` sets 500. + batch_size : int, default=8 + Upper bound on the batch size. The batch actually used is + ``min(n_cases // 10, batch_size)``, the authors' rule, so small + collections train with very small batches. + validation_size : float, default=0.1 + Fraction of the training collection sampled to monitor validation loss. + The authors sample this from the training set *with replacement* and do + not hold it out, so it overlaps the training data. Reproduced because + the learning rate schedule and the retained epoch both depend on it. + Set to 0 to monitor training loss instead, which is what aeon does. + use_class_weights : bool, default=True + Whether to weight the loss by class, as the authors do. aeon does not. + verbose : bool, default=False + Whether Keras prints training progress. + random_state : int, RandomState instance or None, default=None + Seed controlling weight initialisation, batch shuffling and the + validation sample. + + Attributes + ---------- + model_ : keras.Model + The fitted network, with the retained epoch's weights. + history_ : dict + Keras training history, one entry per metric. + batch_size_ : int + Batch size actually used, after the authors' rule. + class_weights_ : dict or None + Weights applied to the loss, or None when disabled. + n_channels_ : int + Number of channels seen in ``fit``. + n_timepoints_ : int + Series length seen in ``fit``. + classes_ : np.ndarray + Class labels, from ``BaseClassifier``. + n_classes_ : int + Number of classes, from ``BaseClassifier``. + + Notes + ----- + Deviations from ``aeon.classification.deep_learning.DisjointCNNClassifier``, + all of them places where aeon departs from the authors' code, and the + candidates for the published-versus-obtained gap in aeon issue #3775: + + - **Class weighting.** The authors weight the loss per class; aeon does not. + - **Batch size.** The authors use ``min(n_cases // 10, 8)``; aeon defaults + to a flat 16, since ``use_mini_batch_size`` is False. + - **Epochs.** The authors train for 500; aeon defaults to 2000. + - **What is monitored.** The authors monitor validation loss for both the + learning rate schedule and the retained epoch; aeon monitors training + loss, having no validation split. + + The architecture here is the clean stack, matching aeon, so that a + difference between the two is attributable to the training procedure alone. + It is worth recording that the authors' own ``DCNN_4L.py`` does not build + that stack: in the third block the spatial convolution is applied to + ``conv2``, the second block's output, rather than to the third block's + temporal convolution, which is therefore computed and discarded. It reads + like a copy-paste slip, but it is the graph their published numbers came + from, so replicating those exactly would need it. ``DCNN_2L`` and + ``DCNN_3L`` are unaffected. + + References + ---------- + .. [1] Foumani, S. N. M., Tan, C. W. and Salehi, M. "Disjoint-CNN for + Multivariate Time Series Classification." ICDMW, 2021. + + Examples + -------- + >>> from aeon.testing.data_generation import make_example_3d_numpy + >>> from multiverse.classification import DisjointCNNClassifier + >>> X, y = make_example_3d_numpy(n_cases=8, n_channels=2, n_timepoints=20) + >>> clf = DisjointCNNClassifier(n_epochs=2) # doctest: +SKIP + >>> clf.fit(X, y) # doctest: +SKIP + """ + + _tags = { + "X_inner_type": "numpy3D", + "capability:multivariate": True, + "capability:unequal_length": False, + "algorithm_type": "deeplearning", + "non_deterministic": True, + "python_dependencies": "tensorflow", + } + + def __init__( + self, + n_layers: int = 4, + n_filters: int = 64, + n_epochs: int = 500, + batch_size: int = 8, + validation_size: float = 0.1, + use_class_weights: bool = True, + verbose: bool = False, + random_state=None, + ): + self.n_layers = n_layers + self.n_filters = n_filters + self.n_epochs = n_epochs + self.batch_size = batch_size + self.validation_size = validation_size + self.use_class_weights = use_class_weights + self.verbose = verbose + self.random_state = random_state + super().__init__() + + def _validate_parameters(self) -> None: + """Check constructor parameters before any work is done.""" + if self.n_layers not in KERNEL_SIZES: + raise ValueError( + f"n_layers must be one of {sorted(KERNEL_SIZES)}, got {self.n_layers}" + ) + for name in ["n_filters", "n_epochs", "batch_size"]: + value = getattr(self, name) + if not isinstance(value, int) or value <= 0: + raise ValueError(f"{name} must be a positive integer") + if not 0 <= self.validation_size < 1: + raise ValueError("validation_size must be in [0, 1)") + + @staticmethod + def _to_original_layout(X: np.ndarray) -> np.ndarray: + """Convert an aeon collection to the authors' (case, time, channel, 1).""" + series = np.transpose(np.asarray(X, dtype=np.float32), (0, 2, 1)) + return series[..., np.newaxis] + + def _fit(self, X: np.ndarray, y): + self._validate_parameters() + + import tensorflow as tf + + rng = check_random_state(self.random_state) + seed = int(rng.randint(np.iinfo(np.int32).max)) + tf.keras.utils.set_random_seed(seed) + + self.n_channels_, self.n_timepoints_ = X.shape[1], X.shape[2] + + encoded_y = np.asarray( + [self._class_dictionary[label] for label in y], dtype=np.int64 + ) + one_hot = np.eye(self.n_classes_, dtype=np.float32)[encoded_y] + + # The authors' rule, which can floor at zero on tiny collections. + self.batch_size_ = max(1, min(X.shape[0] // 10, self.batch_size)) + + self.model_ = _build_disjoint_cnn( + self.n_timepoints_, + self.n_channels_, + self.n_classes_, + self.n_layers, + self.n_filters, + ) + self.model_.compile( + loss="categorical_crossentropy", + optimizer=tf.keras.optimizers.Adam(), + metrics=["accuracy"], + ) + + data = self._to_original_layout(X) + validation_data = None + monitor = "loss" + if self.validation_size: + # Sampled from the training set with replacement and left in it, + # exactly as the authors do. + size = max(1, int(X.shape[0] * self.validation_size)) + index = rng.randint(0, X.shape[0], size) + validation_data = (data[index], one_hot[index]) + monitor = "val_loss" + + self.class_weights_ = ( + _class_weights(one_hot) if self.use_class_weights else None + ) + + callbacks = [ + tf.keras.callbacks.ReduceLROnPlateau( + monitor=monitor, factor=0.5, patience=50, min_lr=0.0001 + ), + # The authors checkpoint to disk and reload; restoring the best + # weights in memory is the same selection without the file. + tf.keras.callbacks.EarlyStopping( + monitor=monitor, + patience=self.n_epochs, + restore_best_weights=True, + ), + ] + + history = self.model_.fit( + data, + one_hot, + validation_data=validation_data, + class_weight=self.class_weights_, + epochs=self.n_epochs, + batch_size=self.batch_size_, + verbose=1 if self.verbose else 0, + callbacks=callbacks, + ) + self.history_ = history.history + return self + + def _check_shape(self, X: np.ndarray) -> None: + if X.shape[1] != self.n_channels_: + raise ValueError( + f"X has {X.shape[1]} channels, but the classifier was fitted " + f"with {self.n_channels_}." + ) + if X.shape[2] != self.n_timepoints_: + raise ValueError( + f"X has length {X.shape[2]}, but the classifier was fitted with " + f"length {self.n_timepoints_}." + ) + + def _predict_proba(self, X: np.ndarray) -> np.ndarray: + self._check_shape(X) + return self.model_.predict(self._to_original_layout(X), verbose=0) + + def _predict(self, X: np.ndarray): + return self.classes_[np.argmax(self._predict_proba(X), axis=1)] + + @classmethod + def _get_test_params(cls, parameter_set: str = "default") -> dict: + """Return a small parameter set for aeon estimator checks.""" + return { + "n_layers": 2, + "n_filters": 4, + "n_epochs": 2, + "batch_size": 4, + "validation_size": 0.0, + "random_state": 0, + } diff --git a/multiverse/classification/tests/test_ported_classifiers.py b/multiverse/classification/tests/test_ported_classifiers.py index e421f7a..b821c56 100644 --- a/multiverse/classification/tests/test_ported_classifiers.py +++ b/multiverse/classification/tests/test_ported_classifiers.py @@ -14,6 +14,7 @@ from multiverse.classification import ( ConvTranClassifier, + DisjointCNNClassifier, PatchMTSCClassifier, TimesNetClassifier, TimesURLClassifier, @@ -67,6 +68,14 @@ "device": "cpu", "random_state": 0, }, + DisjointCNNClassifier: { + "n_layers": 2, + "n_filters": 4, + "n_epochs": 2, + "batch_size": 4, + "validation_size": 0.0, + "random_state": 0, + }, XCMClassifier: { "window_size": 0.2, "n_filters": 4, From 3e4c9759fb990fe98be569a17c58896455a72baa Mon Sep 17 00:00:00 2001 From: Tony Bagnall Date: Wed, 2 Sep 2026 12:40:50 +0100 Subject: [PATCH 2/8] Add a per-dataset accuracy page The leaderboard answers "which estimator is best". This answers "what is this dataset worth", which is the question an archive has to keep asking of its own problems. Per dataset: the Dummy floor, the median, best and worst over the other estimators, which estimator was best, the gain over Dummy, the spread, and how many estimators contributed. Sorted by gain, so the problems where nothing yet beats the baseline are at the top. Two departures from the columns as first sketched. A single range would have been close to redundant, since Dummy is almost always the weakest entry, so a best-minus-worst range mostly restates best-minus-Dummy; they are split into gain, meaning how much skill was found at all, and spread, meaning how much the choice of estimator mattered, because those answer different questions. And the median is reported beside the best, because best is a maximum over 23 estimators and so optimistic by construction. Unlike the leaderboard this keeps every dataset rather than only those every estimator finished, since a partly covered dataset is still informative; the estimator count carries the caveat. Rows are shaded where a dataset separates estimators poorly: amber for a gain of 0.05 or less, green for a best of 0.99 or more. On the current results that is 4 and 15 of 66. AtrialFibrillation and KINECAL-QSEO have a gain of exactly zero, so no estimator has yet beaten the majority class on them. Co-Authored-By: Claude Opus 5 --- README.md | 8 + multiverse/experiments/tables.py | 304 ++++++++++++++++++++ multiverse/experiments/tests/test_tables.py | 71 +++++ results/multiverse/datasets.html | 119 ++++++++ results/multiverse/leaderboard.html | 2 +- 5 files changed, 503 insertions(+), 1 deletion(-) create mode 100644 results/multiverse/datasets.html diff --git a/README.md b/README.md index 0bad635..e8b6ad3 100644 --- a/README.md +++ b/README.md @@ -74,6 +74,14 @@ version with per-metric ranks to ([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/leaderboard.html), since GitHub shows HTML as source). Missing results, and why, are listed on that page. +The same command writes a per-dataset view to +[`results/multiverse/datasets.html`](results/multiverse/datasets.html) +([preview](https://raw.githack.com/aeon-toolkit/multiverse/main/results/multiverse/datasets.html)), +which turns the question around: for each dataset it gives the Dummy floor, the median +and best over the other estimators, which estimator was best, how much the best gained +over Dummy, and how far apart the estimators were. It is sorted by that gain, so the +problems where nothing yet beats the baseline come first. + This repository aims to make it easier to: - load Multiverse datasets through `aeon` diff --git a/multiverse/experiments/tables.py b/multiverse/experiments/tables.py index bcef1e6..ec14367 100644 --- a/multiverse/experiments/tables.py +++ b/multiverse/experiments/tables.py @@ -24,6 +24,9 @@ __maintainer__ = ["TonyBagnall"] __all__ = [ "leaderboard", + "dataset_summary", + "dataset_page", + "dataset_markdown", "leaderboard_markdown", "write_markdown_table", "available_estimators", @@ -42,6 +45,12 @@ DEFAULT_RESULTS_DIR = Path(__file__).resolve().parents[2] / "results" / "multiverse" +#: A dataset whose best estimator gains no more than this over the baseline +#: shows little signal; one whose best reaches SATURATED_BEST is solved. +#: Both separate estimators poorly, for opposite reasons. +NO_SIGNAL_GAIN = 0.05 +SATURATED_BEST = 0.99 + #: Metrics where a smaller value is a better result. Anything not listed here is #: treated as higher-is-better. LOWER_IS_BETTER = {"logloss"} @@ -846,6 +855,293 @@ def write_markdown_table( return True +def dataset_summary( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", +) -> pd.DataFrame: + """Summarise one metric per dataset rather than per estimator. + + The leaderboard answers "which estimator is best"; this answers "what is + this dataset worth", which is the question an archive has to keep asking of + its own problems. + + Unlike :func:`leaderboard` this does not restrict to the datasets every + estimator has results for. A dataset only some estimators finished is still + informative, so every dataset is kept and ``estimators`` records how many + contributed. + + Parameters + ---------- + datasets : list of str + Datasets to include, in any order. + estimators : list of str + Estimator names, matching their results directories. + metric : str + Metric to summarise. + results_dir : Path or str + Directory holding one sub-directory per estimator. + baseline : str + Estimator treated as the no-skill floor, excluded from best, worst, + median and spread. Pass None to keep it in. + + Returns + ------- + pd.DataFrame + One row per dataset, indexed by dataset name, with the baseline score, + the median, best and worst over the remaining estimators, the estimator + achieving the best, the gain over the baseline, the spread, and the + number of estimators contributing. + """ + frames = _load_all(list(estimators), [metric], results_dir) + if metric not in frames: + raise ValueError(f"no results for metric {metric!r}") + scores = frames[metric].reindex(sorted(dict.fromkeys(datasets))) + + others = scores.drop(columns=[baseline], errors="ignore") + lower_better = metric in LOWER_IS_BETTER + best = others.min(axis=1) if lower_better else others.max(axis=1) + worst = others.max(axis=1) if lower_better else others.min(axis=1) + + # idxmin/idxmax raise on an all-NaN row, so only ask where there is a value. + winner = pd.Series(pd.NA, index=others.index, dtype=object) + present = others.notna().any(axis=1) + if present.any(): + rows = others[present] + winner[present] = rows.idxmin(axis=1) if lower_better else rows.idxmax(axis=1) + + summary = pd.DataFrame( + { + "baseline": scores[baseline] if baseline in scores else np.nan, + "median": others.median(axis=1), + "best": best, + "worst": worst, + "best_estimator": winner, + "estimators": others.notna().sum(axis=1), + } + ) + # The results files carry a "Resamples:" header, which pandas takes as the + # index name and would otherwise appear as the first column heading. + summary.index.name = "dataset" + # Gain is how much skill the best estimator found beyond the baseline; + # spread is how much the choice of estimator mattered. Reporting a single + # range would conflate the two, and since the baseline is almost always the + # weakest entry that range would just restate the gain. + summary["gain"] = ( + summary["baseline"] - summary["best"] + if lower_better + else summary["best"] - summary["baseline"] + ) + summary["spread"] = ( + summary["worst"] - summary["best"] + if lower_better + else summary["best"] - summary["worst"] + ) + return summary + + +def _dataset_table_html(summary, decimals) -> str: + """Render the per-dataset table, one row per dataset.""" + head = ( + '' + "Dataset" + 'Dummy' + 'Median' + 'Best' + '' + "Best estimator" + 'Gain over dummy' + 'Spread' + 'Estimators' + ) + + def number(value): + return "—" if pd.isna(value) else f"{value:.{decimals}f}" + + rows = [] + for dataset, row in summary.iterrows(): + # Two flags worth seeing at a glance: nothing beat the baseline by much, + # and everything solves it. Both make a dataset weak at separating + # estimators, for opposite reasons. + classes = [] + if pd.notna(row["gain"]) and row["gain"] <= NO_SIGNAL_GAIN: + classes.append("nosignal") + if pd.notna(row["best"]) and row["best"] >= SATURATED_BEST: + classes.append("saturated") + attribute = f' class="{" ".join(classes)}"' if classes else "" + winner = ( + "—" + if pd.isna(row["best_estimator"]) + else escape(str(row["best_estimator"])) + ) + rows.append( + f"{escape(str(dataset))}" + f'{number(row["baseline"])}' + f'{number(row["median"])}' + f'{number(row["best"])}' + f"{winner}" + f'{number(row["gain"])}' + f'{number(row["spread"])}' + f'{int(row["estimators"])}' + ) + + return ( + '
' + f"{head}" + f"{''.join(rows)}
" + ) + + +def dataset_page( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", + sort_by: str = "gain", + output_path: Path | str | None = None, + title: str | None = None, + decimals: int = 4, +) -> Path: + """Write a self-contained page summarising one metric per dataset. + + Parameters + ---------- + datasets : list of str + Datasets to include. + estimators : list of str + Estimator names, matching their results directories. + metric : str + Metric to summarise. + results_dir : Path or str + Directory holding one sub-directory per estimator. + baseline : str + Estimator treated as the no-skill floor. + sort_by : str + Column to order rows by, one of the columns of + :func:`dataset_summary`. The default puts the datasets where the best + estimator gained least over the baseline at the top, because those are + the ones worth looking at. + output_path : Path or str, optional + Where to write. Defaults to ``datasets.html`` beside the results. + title : str, optional + Page heading. + decimals : int + Decimal places for scores. + + Returns + ------- + Path + The file written. + """ + summary = dataset_summary(datasets, estimators, metric, results_dir, baseline) + if sort_by not in summary.columns: + raise ValueError(f"sort_by={sort_by!r} is not a column") + ascending = sort_by not in {"median", "best", "worst", "spread", "estimators"} + summary = summary.sort_values(sort_by, ascending=ascending, na_position="last") + + label = METRIC_LABELS.get(metric, metric) + title = title or f"Multiverse datasets: {label.lower()}" + scored = int((summary["estimators"] > 0).sum()) + no_signal = int(summary["gain"].le(NO_SIGNAL_GAIN).sum()) + saturated = int(summary["best"].ge(SATURATED_BEST).sum()) + + parts = [ + f"

{escape(title)}

", + f'

{scored} datasets · {escape(label.lower())}' + f" · best of up to {int(summary['estimators'].max())} estimators" + f" against the {escape(baseline)} baseline · built " + f"{date.today().isoformat()}

", + _dataset_table_html(summary, decimals), + '

One row per dataset. Dummy is the ' + "no-skill floor. Median, best and " + "spread are over the other estimators, so the baseline " + "cannot flatter them. Gain over dummy is best minus " + "dummy, how much skill was found at all; spread is " + "best minus worst, how much the choice of estimator mattered. The two " + "answer different questions, and a single range would conflate them.

", + f'

{no_signal} of {scored} datasets gained ' + f"{NO_SIGNAL_GAIN:.2f} or less over the baseline (shaded amber) and " + f"{saturated} have a best of {SATURATED_BEST:.2f} or more (shaded " + "green). Both separate estimators poorly, for opposite reasons. Best is " + "a maximum over many estimators, so it is optimistic by construction: " + "read it as what the archive can currently do on a problem, not as what " + "any one method delivers.

", + ] + + page = ( + '' + '' + f"{escape(title)}" + f"
{''.join(parts)}
" + f"" + ) + + output_path = ( + Path(results_dir) / "datasets.html" + if output_path is None + else Path(output_path) + ) + output_path.parent.mkdir(parents=True, exist_ok=True) + output_path.write_text(page, encoding="utf-8") + return output_path + + +def dataset_markdown( + datasets, + estimators, + metric: str = "accuracy", + results_dir: Path | str = DEFAULT_RESULTS_DIR, + baseline: str = "Dummy", + sort_by: str = "gain", + decimals: int = 4, +) -> str: + """Return the per-dataset summary as a Markdown table.""" + summary = dataset_summary(datasets, estimators, metric, results_dir, baseline) + ascending = sort_by not in {"median", "best", "worst", "spread", "estimators"} + summary = summary.sort_values(sort_by, ascending=ascending, na_position="last") + + header = [ + "Dataset", "Dummy", "Median", "Best", "Best estimator", + "Gain over dummy", "Spread", "Estimators", + ] + rows = ["| " + " | ".join(header) + " |", "|" + "---|" * len(header)] + + def number(value): + return "—" if pd.isna(value) else f"{value:.{decimals}f}" + + for dataset, row in summary.iterrows(): + winner = ( + "—" + if pd.isna(row["best_estimator"]) + else str(row["best_estimator"]) + ) + cells = [ + str(dataset), + number(row["baseline"]), + number(row["median"]), + f'**{number(row["best"])}**', + winner, + number(row["gain"]), + number(row["spread"]), + str(int(row["estimators"])), + ] + rows.append("| " + " | ".join(cells) + " |") + + rows.append("") + rows.append( + f"{METRIC_LABELS.get(metric, metric)} per dataset. Median, best and " + f"spread are over the estimators other than {baseline}. Gain over dummy " + "is best minus dummy; spread is best minus worst." + ) + return "\n".join(rows) + + def main() -> None: """Build the Multiverse-core leaderboard. @@ -872,6 +1168,14 @@ def main() -> None: ) print(f"wrote {path}") + datasets_path = dataset_page( + datasets, + estimators, + metric="accuracy", + title="Multiverse-core datasets: accuracy", + ) + print(f"wrote {datasets_path}") + table = leaderboard_markdown(datasets, estimators, sort_by="accuracy") readme = Path(__file__).resolve().parents[2] / "README.md" if write_markdown_table(readme, table): diff --git a/multiverse/experiments/tests/test_tables.py b/multiverse/experiments/tests/test_tables.py index e6f6e0b..b580ec2 100644 --- a/multiverse/experiments/tests/test_tables.py +++ b/multiverse/experiments/tests/test_tables.py @@ -12,6 +12,9 @@ LOWER_IS_BETTER, METRIC_LABELS, available_estimators, + dataset_markdown, + dataset_page, + dataset_summary, leaderboard, leaderboard_markdown, load_metric, @@ -222,3 +225,71 @@ def test_write_markdown_table_without_markers(tmp_path): path.write_text("no markers here\n", encoding="utf-8") assert not write_markdown_table(path, "new", marker="T") assert path.read_text(encoding="utf-8") == "no markers here\n" + + +def test_dataset_summary_excludes_the_baseline(results_dir): + """Best, median and spread ignore the baseline; gain is measured against it.""" + summary = dataset_summary( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], baseline="Carol", + results_dir=results_dir, + ) + # d1: Alice 0.9, Bob 0.6, baseline Carol 0.3 + assert summary.at["d1", "best"] == pytest.approx(0.9) + assert summary.at["d1", "best_estimator"] == "Alice" + assert summary.at["d1", "worst"] == pytest.approx(0.6) + assert summary.at["d1", "baseline"] == pytest.approx(0.3) + assert summary.at["d1", "gain"] == pytest.approx(0.6) + assert summary.at["d1", "spread"] == pytest.approx(0.3) + assert summary.at["d1", "estimators"] == 2 + + +def test_dataset_summary_keeps_partly_covered_datasets(results_dir): + """A dataset only some estimators finished is kept, with a lower count. + + This is where it differs from the leaderboard, which has to drop d3 to keep + the averages comparable. + """ + summary = dataset_summary( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], results_dir=results_dir + ) + assert list(summary.index) == ["d1", "d2", "d3"] + # Carol has no d3, and is not the baseline here, so only two contribute + assert summary.at["d3", "estimators"] == 2 + assert summary.at["d1", "estimators"] == 3 + + +def test_dataset_summary_lower_is_better_flips_best_and_gain(results_dir): + """For log loss the best score is the smallest, and gain stays positive.""" + summary = dataset_summary( + ["d1"], ["Alice", "Bob", "Carol"], metric="logloss", baseline="Carol", + results_dir=results_dir, + ) + # stored as 1 - score, so Alice 0.1, Bob 0.4, Carol 0.7 + assert summary.at["d1", "best"] == pytest.approx(0.1) + assert summary.at["d1", "best_estimator"] == "Alice" + assert summary.at["d1", "gain"] == pytest.approx(0.6) + assert summary.at["d1", "spread"] == pytest.approx(0.3) + + +def test_dataset_page_is_written_and_sorted_by_gain(results_dir, tmp_path): + """The page lists every dataset, worst gain first.""" + path = dataset_page( + ["d1", "d2", "d3"], ["Alice", "Bob", "Carol"], baseline="Carol", + results_dir=results_dir, output_path=tmp_path / "datasets.html", + ) + html = path.read_text(encoding="utf-8") + for dataset in ["d1", "d2", "d3"]: + assert f"{dataset}" in html + # d3 has no baseline score, so its gain is NaN and it sorts last + assert html.index("d2") < html.index("d3") + assert "Gain over dummy" in html + + +def test_dataset_markdown_marks_the_best(results_dir): + """The Markdown table bolds the best score and names the estimator.""" + table = dataset_markdown( + ["d1"], ["Alice", "Bob", "Carol"], baseline="Carol", results_dir=results_dir + ) + assert "| d1 |" in table + assert "**0.9000**" in table + assert "Alice" in table diff --git a/results/multiverse/datasets.html b/results/multiverse/datasets.html new file mode 100644 index 0000000..2729334 --- /dev/null +++ b/results/multiverse/datasets.html @@ -0,0 +1,119 @@ +Multiverse-core datasets: accuracy

Multiverse-core datasets: accuracy

66 datasets · accuracy · best of up to 23 estimators against the Dummy baseline · built 2026-09-02

DatasetDummyMedianBestBest estimatorGain over dummySpreadEstimators
AtrialFibrillation0.33330.20000.3333CIF0.00000.266723
KINECAL-QSEO0.94120.94120.9412Arsenal0.00000.117623
BIDMC32SpO2_disc0.71530.66280.7203ROCKET0.00500.171321
Locust20220.91120.90820.9206MRHydra0.00940.045223
Heartbeat0.72200.74630.7854CIF0.06340.126823
HouseholdPowerConsumption2_disc0.72160.76820.7872HC20.06560.141423
AutomotiveRoadTrials0.75320.80520.8442CIF0.09090.233823
AustraliaRainfall_disc0.68600.77310.7808LITETime-MV0.09480.088115
EyesOpenShut0.50000.50000.5952STSF0.09520.190523
MotorImagery0.50000.51000.6000FreshPRINCE0.10000.140023
BeijingPM10Quality_disc0.71120.82470.8417FreshPRINCE0.13050.107023
Alzheimers0.41860.37210.5581MRHydra0.13950.302322
AppliancesEnergy_disc0.80950.83330.9524FreshPRINCE0.14290.452423
EmoPain0.78310.84080.92681NN-DTW0.14370.242320
PhotoStimulation0.41670.37500.5833ROCKET0.16670.388922
FaceDetection0.50000.63040.6850H-InceptionTime0.18500.158622
BeijingPM25Quality_disc0.69770.87560.8879ConvTran0.19020.128823
HouseholdPowerConsumption1_disc0.77840.91250.9825FreshPRINCE0.20410.218723
LowCost0.50000.64170.7300TSF0.23000.248323
BoneProbAgeGroup0.47640.65390.7124H-InceptionTime0.23600.193323
StandWalkJump0.33330.46670.6000MRHydra0.26670.400023
CrowdSourced0.50020.71370.7734LITETime-MV0.27320.176823
BenzeneConcentration_disc0.68970.81950.9768STSF0.28700.577023
FordChallenge0.62320.88390.9360QUANT0.31280.312822
BIDMC32HR_disc0.65070.82580.9637RIST0.31300.635321
STEW0.50000.74380.8385Arsenal0.33850.210021
BoneIntensitiesAgeGroup0.47640.80450.8202HC20.34380.296623
PhonemeSpectra0.02560.27950.3746H-InceptionTime0.34890.291123
KERAAL-RTK0.57140.78570.9286HC20.35710.571423
LSST0.31510.63220.7040FreshPRINCE0.38890.480923
HandMovementDirection0.20270.41890.6081TSF0.40540.418923
DuckDuckGeese0.20000.46000.6400H-InceptionTime0.44000.480023
SelfRegulationSCP10.50170.85320.9454MRHydra0.44370.208223
IEEEPPG_disc0.26050.44730.7078ConvTran0.44730.438323
AsphaltRegularityCoordinates0.50730.98000.9947H-InceptionTime0.48740.291623
MindReading0.23120.53140.7243LITETime-MV0.49310.313923
EthanolConcentration0.25100.43350.7490STC0.49810.532323
UIPRMD-DS-C0.50000.83331.0000Catch220.50000.388923
WISDM0.36640.86730.8965MRHydra0.53000.131023
Blink0.44440.99331.0000Arsenal0.55560.415623
EigenWorms0.41980.89310.9771MRHydra0.55730.557322
KIMORE-PR-C0.14290.42860.7143LITETime-MV0.57140.571423
AsphaltObstaclesCoordinates0.28390.82100.8670MRHydra0.58310.289023
CounterMovementJump0.33520.75420.9274Arsenal0.59220.458123
Handwriting0.03760.37760.6529H-InceptionTime0.61530.467123
USCActivity0.11380.69360.7354LITETime-MV0.62160.137919
RacketSports0.28290.88160.9079RDST0.62500.125023
UCDHE-Rowing-MC0.20450.73860.8295PatchMTSC0.62500.338623
IRDS-SFL0.20690.79310.8621RDST0.65520.448323
Skoda0.23560.94800.9646H-InceptionTime0.72900.119922
Epilepsy0.26810.98551.0000HC20.73190.058023
Tiselac0.06280.81920.8373STSF0.77450.204418
MotionSenseHAR0.20380.98871.0000DrCIF0.79620.052823
NATOPS0.16670.88890.9667LITETime-MV0.80000.155623
UCIActivity0.19160.97460.9983LITETime-MV0.80670.174923
UWaveGestureLibrary0.12500.90940.9406Arsenal0.81560.553123
ERing0.16670.93330.9963MRHydra0.82960.233323
PEMS-SF0.11560.96531.0000CIF0.88440.317923
PenDigits0.10380.97860.9911H-InceptionTime0.88740.234421
SpokenArabicDigits0.10000.97730.9941RDST0.89400.128723
Libras0.06670.88890.9722RIST0.90560.338923
JapaneseVowels0.08380.95950.9946LiteTIME0.91080.208123
Cricket0.08330.98611.00001NN-DTW0.91670.069423
CharacterTrajectories0.06480.98960.9958H-InceptionTime0.93110.044623
TactileTextureRecognition0.05140.99851.0000H-InceptionTime0.94860.016223
ArticularyWordRecognition0.04000.98330.9933Arsenal0.95330.050023

One row per dataset. Dummy is the no-skill floor. Median, best and spread are over the other estimators, so the baseline cannot flatter them. Gain over dummy is best minus dummy, how much skill was found at all; spread is best minus worst, how much the choice of estimator mattered. The two answer different questions, and a single range would conflate them.

4 of 66 datasets gained 0.05 or less over the baseline (shaded amber) and 15 have a best of 0.99 or more (shaded green). Both separate estimators poorly, for opposite reasons. Best is a maximum over many estimators, so it is optimistic by construction: read it as what the archive can currently do on a problem, not as what any one method delivers.

\ No newline at end of file diff --git a/results/multiverse/leaderboard.html b/results/multiverse/leaderboard.html index fdccb82..032371d 100644 --- a/results/multiverse/leaderboard.html +++ b/results/multiverse/leaderboard.html @@ -60,7 +60,7 @@ details { margin-top: .6rem; } summary { cursor: pointer; color: var(--accent); } code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .9em; } -

Multiverse-core leaderboard

24 estimators on 52 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-01

#EstimatorAccuracyBalanced accuracyAUROCF1Log loss ↓SensitivitySpecificity
ScoreRankScoreRankScoreRankScoreRankScoreRankScoreRankScoreRank
1HC20.79097.860.75188.490.89906.060.72738.260.53836.130.74598.660.79437.46
2MRHydra0.78378.310.75647.970.810516.720.73167.857.797419.740.76428.200.77579.12
3RDST0.77349.160.73339.720.791217.370.699110.038.166719.970.710910.680.78748.73
4RIST0.77209.760.739710.580.87488.030.714710.380.62188.310.740810.740.765510.80
5DrCIF0.774710.070.742910.520.88138.380.717310.860.64849.170.739711.400.770810.90
6FreshPRINCE0.774310.100.748710.480.87457.940.721110.380.60076.210.741410.870.777010.97
7CIF0.778110.130.747110.380.89088.250.721210.390.64309.210.744111.200.775310.28
8QUANT0.772010.550.746210.480.88317.550.718910.570.61757.370.752110.680.758111.56
9Arsenal0.768010.700.732110.570.845713.060.702410.693.863116.800.725710.780.773210.53
10ROCKET0.769010.820.732610.610.792518.120.701911.098.324920.750.720011.660.776411.02
11LITETime-MV0.750611.140.72999.790.85139.930.68209.871.320612.150.71329.700.763711.10
12STSF0.772411.550.747711.180.88049.830.708011.750.64328.380.734512.100.782612.19
13H-InceptionTime0.740811.630.719010.830.849610.380.683810.511.322712.900.722310.330.737812.43
14LiteTIME0.734112.340.710411.470.839411.520.668011.741.477613.150.711310.570.733612.21
15PatchMTSC0.742813.120.689713.930.826112.590.653313.490.76559.440.685213.150.735212.98
16ConvTran0.746213.140.710213.340.859211.020.676712.840.81909.520.715912.740.734513.56
17Catch220.747513.150.718113.610.869710.740.692213.620.714710.790.724013.540.737413.62
18STC0.754513.950.717214.130.874411.310.694013.960.63919.920.718514.110.753713.89
19TSF0.751514.000.723613.560.874011.580.688313.910.725210.500.709314.600.760613.87
20TDE0.726214.630.681314.790.837412.450.638214.250.886911.500.671413.820.734413.07
21Summary0.685816.660.657416.400.826815.290.623016.400.912313.250.657416.350.684416.85
22TimesURL0.695816.950.653317.050.790617.710.596717.361.005515.480.625717.320.697316.01
231NN-DTW0.671218.450.645417.630.719721.240.613617.6311.850622.930.652116.330.663618.53
24Dummy0.364521.810.302922.500.500022.950.150722.171.406716.400.285520.480.381618.33

Average score and average rank over the 52 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.

Missing results

  • HC2 — AustraliaRainfall_disc (Time limit); STEW, Tiselac, USCActivity (cancelled before completion)
  • MRHydra — AustraliaRainfall_disc (OOM at 128GB); PenDigits (ValueError: n_timepoints must be >= 9, but found 8); Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • RDST — AustraliaRainfall_disc, Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738)); USCActivity (OOM at 64GB)
  • FreshPRINCE — FaceDetection, FordChallenge, Skoda, Tiselac (OOM at 128GB)
  • ROCKET — AustraliaRainfall_disc (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • STSF — AustraliaRainfall_disc, PenDigits (not recorded)
  • LiteTIME — BIDMC32HR_disc, BIDMC32SpO2_disc, USCActivity (not recorded)
  • PatchMTSC — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • ConvTran — Alzheimers, EigenWorms, PhotoStimulation (CUDA out of memory); EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • TSF — AustraliaRainfall_disc (not recorded)
  • TDE — AustraliaRainfall_disc, Tiselac, USCActivity (Time limit); STEW (cancelled before completion)
  • Summary — AustraliaRainfall_disc (not recorded)
  • TimesURL — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • 1NN-DTW — BIDMC32HR_disc (Time limit); BIDMC32SpO2_disc (not recorded)

Scoring uses the 52 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.

Reproducing this page

from aeon.datasets.tsc_datasets import multiverse_core
+

Multiverse-core leaderboard

24 estimators on 52 datasets · 7 metrics · ordered by average accuracy rank · built 2026-09-02

#EstimatorAccuracyBalanced accuracyAUROCF1Log loss ↓SensitivitySpecificity
ScoreRankScoreRankScoreRankScoreRankScoreRankScoreRankScoreRank
1HC20.79097.860.75188.490.89906.060.72738.260.53836.130.74598.660.79437.46
2MRHydra0.78378.310.75647.970.810516.720.73167.857.797419.740.76428.200.77579.12
3RDST0.77349.160.73339.720.791217.370.699110.038.166719.970.710910.680.78748.73
4RIST0.77209.760.739710.580.87488.030.714710.380.62188.310.740810.740.765510.80
5DrCIF0.774710.070.742910.520.88138.380.717310.860.64849.170.739711.400.770810.90
6FreshPRINCE0.774310.100.748710.480.87457.940.721110.380.60076.210.741410.870.777010.97
7CIF0.778110.130.747110.380.89088.250.721210.390.64309.210.744111.200.775310.28
8QUANT0.772010.550.746210.480.88317.550.718910.570.61757.370.752110.680.758111.56
9Arsenal0.768010.700.732110.570.845713.060.702410.693.863116.800.725710.780.773210.53
10ROCKET0.769010.820.732610.610.792518.120.701911.098.324920.750.720011.660.776411.02
11LITETime-MV0.750611.140.72999.790.85139.930.68209.871.320612.150.71329.700.763711.10
12STSF0.772411.550.747711.180.88049.830.708011.750.64328.380.734512.100.782612.19
13H-InceptionTime0.740811.630.719010.830.849610.380.683810.511.322712.900.722310.330.737812.43
14LiteTIME0.734112.340.710411.470.839411.520.668011.741.477613.150.711310.570.733612.21
15PatchMTSC0.742813.120.689713.930.826112.590.653313.490.76559.440.685213.150.735212.98
16ConvTran0.746213.140.710213.340.859211.020.676712.840.81909.520.715912.740.734513.56
17Catch220.747513.150.718113.610.869710.740.692213.620.714710.790.724013.540.737413.62
18STC0.754513.950.717214.130.874411.310.694013.960.63919.920.718514.110.753713.89
19TSF0.751514.000.723613.560.874011.580.688313.910.725210.500.709314.600.760613.87
20TDE0.726214.630.681314.790.837412.450.638214.250.886911.500.671413.820.734413.07
21Summary0.685816.660.657416.400.826815.290.623016.400.912313.250.657416.350.684416.85
22TimesURL0.695816.950.653317.050.790617.710.596717.361.005515.480.625717.320.697316.01
231NN-DTW0.671218.450.645417.630.719721.240.613617.6311.850622.930.652116.330.663618.53
24Dummy0.364521.810.302922.500.500022.950.150722.171.406716.400.285520.480.381618.33

Average score and average rank over the 52 datasets with results for every estimator on every metric. Best in each column is highlighted. Metrics marked ↓ are better when lower.

Missing results

  • HC2 — AustraliaRainfall_disc (Time limit); STEW, Tiselac, USCActivity (cancelled before completion)
  • MRHydra — AustraliaRainfall_disc (OOM at 128GB); PenDigits (ValueError: n_timepoints must be >= 9, but found 8); Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • RDST — AustraliaRainfall_disc, Tiselac (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738)); USCActivity (OOM at 64GB)
  • FreshPRINCE — FaceDetection, FordChallenge, Skoda, Tiselac (OOM at 128GB)
  • ROCKET — AustraliaRainfall_disc (LAPACK integer overflow in the RidgeClassifierCV SVD (aeon issue 3738))
  • STSF — AustraliaRainfall_disc, PenDigits (not recorded)
  • LiteTIME — BIDMC32HR_disc, BIDMC32SpO2_disc, USCActivity (not recorded)
  • PatchMTSC — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • ConvTran — Alzheimers, EigenWorms, PhotoStimulation (CUDA out of memory); EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • TSF — AustraliaRainfall_disc (not recorded)
  • TDE — AustraliaRainfall_disc, Tiselac, USCActivity (Time limit); STEW (cancelled before completion)
  • Summary — AustraliaRainfall_disc (not recorded)
  • TimesURL — EmoPain (ValueError: input collection has too little variation (std <= 1e-07))
  • 1NN-DTW — BIDMC32HR_disc (Time limit); BIDMC32SpO2_disc (not recorded)

Scoring uses the 52 datasets every estimator completed, so a dataset any one of them is missing is left out for all. Reasons are from the job logs of these runs.

Reproducing this page

from aeon.datasets.tsc_datasets import multiverse_core
 from multiverse.experiments.tables import leaderboard
 
 leaderboard(

From b4b1392a38b17cc670d61be31f17d3a6c91cc918 Mon Sep 17 00:00:00 2001
From: Tony Bagnall 
Date: Wed, 2 Sep 2026 13:10:47 +0100
Subject: [PATCH 3/8] Add a runtime holding page

A placeholder for runtime and memory comparisons, with nothing published yet and
the reasons why written down.

The measurements already exist: tsml-eval records fit_time, predict_time,
memory_usage and benchmark_time per run, and the ingest simply does not bring
them across. The reason to hold off is that the numbers we hold are not
comparable with each other. Results span H200, A100 and CPU-only nodes; GPU and
CPU methods are not on one axis, so a ratio between them is partly a statement
about hardware; wall-clock includes queueing and the controller's memory
escalation retries; a deep learner's fit time is close to linear in an epoch
count that is a choice rather than a property; memory_usage is peak host memory,
so a model doing its work on a GPU looks cheap while occupying device memory
nobody measured; and the recorded device is unreliable, since the tooling used
for these runs inspects TensorFlow only and reports CPU for PyTorch estimators
either way.

The page also sets out what a fair comparison would need, including normalising
CPU times by benchmark_time, which exists for that purpose, and notes there is
no GPU equivalent.

This is the same reasoning that keeps time and memory columns off the
leaderboard, so the two now cross-reference.

Co-Authored-By: Claude Opus 5 
---
 README.md       |  2 ++
 docs/runtime.md | 80 +++++++++++++++++++++++++++++++++++++++++++++++++
 2 files changed, 82 insertions(+)
 create mode 100644 docs/runtime.md

diff --git a/README.md b/README.md
index e8b6ad3..db6f3ca 100644
--- a/README.md
+++ b/README.md
@@ -99,6 +99,8 @@ This repository aims to make it easier to:
   ·
   Leaderboard
   ·
+  Runtime
+  ·
   Evaluation
   ·
   Classifiers
diff --git a/docs/runtime.md b/docs/runtime.md
new file mode 100644
index 0000000..65edd12
--- /dev/null
+++ b/docs/runtime.md
@@ -0,0 +1,80 @@
+# Runtime
+
+**Coming soon.** This page will hold runtime and memory comparisons across the
+Multiverse estimators. Nothing is published here yet, deliberately: the numbers we
+currently hold are not comparable with each other, and publishing them would invite
+exactly the conclusions they cannot support.
+
+The rest of this page records why, so that when the page is populated it is clear what
+the figures do and do not mean.
+
+## The measurements exist
+
+Every run already records timings. `tsml-eval` writes, per classifier, dataset and
+resample:
+
+- `fit_time` and `predict_time`, in seconds;
+- `memory_usage`, the peak memory during `fit`;
+- `benchmark_time`, the time that machine took to sort 1,000 seeded random arrays of
+  20,000 elements.
+
+They are in the raw prediction files. What this repository ingests under `results/` is
+only the accuracy-style measures, one file per metric, so the timings have not been
+brought across yet. That is a small piece of work; the reason it has not been done is
+below, not the effort.
+
+## Why the numbers are not comparable
+
+**The runs are spread across different hardware.** Multiverse results have been produced
+on GPU partitions with H200 and A100 cards and on CPU-only nodes, with different core
+counts. GPU jobs in our configurations are allocated two CPUs each. A fit time from one
+partition and a fit time from another are two different measurements that happen to share
+a unit.
+
+**GPU and CPU methods are not on one axis.** For the deep learners nearly all the work is
+on the accelerator and the host CPU mostly feeds batches; for the classical ensembles
+there is no accelerator at all and the time scales with the cores allocated. Comparing
+them measures the hardware at least as much as the algorithm, and the ratio moves when
+either side changes. A statement like "X is 40 times faster than Y" is, in this setting,
+a statement about a purchasing decision.
+
+**Wall-clock contains things that are not the algorithm.** Queueing, data loading, and
+retries: our controllers escalate a job's memory request from 64 GB to 128 GB after a
+failure, so an elapsed time can include a dead run at the lower ceiling.
+
+**Training time is a hyperparameter, not a property of a method.** A deep learner's fit
+time is close to linear in the number of epochs, and the epoch count is a choice. Two
+faithful ports of the same paper can differ several-fold on time because the authors
+picked 500 epochs and the toolkit's default is 2000. Early stopping and best-epoch
+selection move it again. None of that is a fact about the architecture.
+
+**Memory is measured in one place and spent in another.** `memory_usage` is peak host
+memory during `fit`. A model doing all its work on a GPU can look inexpensive by that
+measure while occupying tens of gigabytes of device memory, which the figure never sees.
+Peak host memory and peak device memory are different quantities and only one is
+recorded.
+
+**Which device a run actually used is not reliably recorded.** In the version of the
+experiment tooling used for these runs, the device description inspects TensorFlow only,
+so a PyTorch estimator reports CPU whether or not it ran on a GPU. Any timing table built
+from those records has to have its device column reconstructed from the job
+configuration rather than trusted as written.
+
+## What a fair comparison would need
+
+- The compared estimators run on the same hardware, or CPU timings normalised by
+  `benchmark_time`, which exists for exactly this purpose. There is no equivalent
+  normaliser for GPU work.
+- Thread and core counts fixed and recorded, since the classical methods scale with them.
+- `fit` and `predict` reported separately. They answer different questions: fit time is
+  the cost of research, predict time is the cost of deployment, and the ranking is not
+  the same on both.
+- Repeated runs. Timings vary far more between repeats than accuracy does, especially on
+  shared nodes.
+- Asymptotic complexity in the number of cases, series length and channels reported
+  beside the measured times, so a reader can tell whether a result will hold at a
+  different scale.
+- The device stated per run, from the job configuration.
+
+Until most of that is in place, this page stays empty. For the same reason the
+[leaderboard](leaderboard.md) carries no fit time, predict time or memory columns.

From 10ecf362e89a62ef59f7fff6993f87de5df4d6df Mon Sep 17 00:00:00 2001
From: Tony Bagnall 
Date: Wed, 2 Sep 2026 17:02:00 +0100
Subject: [PATCH 4/8] Lead the runtime page with the experiment not being
 designed for it

The page opened on the numbers not being comparable, which reads as a data
problem to be cleaned up. The more basic point is that no timing experiment has
been designed or run: every result here came from runs set up to measure
predictive performance, where partition, cores, epochs and memory retries were
chosen to get accurate results out cheaply and left free across estimators
because nothing depended on them. The timings are a by-product.

That now leads, and the reasons below it are framed as the conditions such an
experiment would have to fix rather than as defects in the current figures.

Co-Authored-By: Claude Opus 5 
---
 docs/runtime.md | 23 +++++++++++++++++------
 1 file changed, 17 insertions(+), 6 deletions(-)

diff --git a/docs/runtime.md b/docs/runtime.md
index 65edd12..02f314c 100644
--- a/docs/runtime.md
+++ b/docs/runtime.md
@@ -1,12 +1,20 @@
 # Runtime
 
 **Coming soon.** This page will hold runtime and memory comparisons across the
-Multiverse estimators. Nothing is published here yet, deliberately: the numbers we
-currently hold are not comparable with each other, and publishing them would invite
-exactly the conclusions they cannot support.
+Multiverse estimators.
 
-The rest of this page records why, so that when the page is populated it is clear what
-the figures do and do not mean.
+**We have not yet structured an experiment to compare runtime.** Every run behind the
+results in this repository was set up to measure predictive performance. Which partition
+a classifier was queued on, how many cores it was given, how many epochs it trained for,
+whether a job was retried at a higher memory ceiling: all of those were chosen to get
+accurate results out at a reasonable cost, and none were held constant across estimators
+because nothing depended on it. The timings that came out are a by-product of that, not
+a measurement anyone designed.
+
+So nothing is published here yet, deliberately. Comparing runtime needs its own
+experiment, with the conditions below fixed in advance, and we have not run one. The
+rest of this page records what those conditions are, and why the figures we already hold
+cannot stand in for them.
 
 ## The measurements exist
 
@@ -23,7 +31,10 @@ only the accuracy-style measures, one file per metric, so the timings have not b
 brought across yet. That is a small piece of work; the reason it has not been done is
 below, not the effort.
 
-## Why the numbers are not comparable
+## Why the figures we hold cannot stand in
+
+Each of these is a condition a timing experiment would have to fix, and that these runs
+left free.
 
 **The runs are spread across different hardware.** Multiverse results have been produced
 on GPU partitions with H200 and A100 cards and on CPU-only nodes, with different core

From 7c5d467195eac3dbdae2e1e757f1d168043f8638 Mon Sep 17 00:00:00 2001
From: Tony Bagnall 
Date: Wed, 2 Sep 2026 17:03:40 +0100
Subject: [PATCH 5/8] Add a memory holding page

Same shape as the runtime page and leading on the same point: no experiment has
been structured to compare memory. A job was given whatever ceiling got it to
finish, on whatever node was free, with whatever core count came with the
partition, none of it held constant across estimators.

Memory has enough of its own difficulties to warrant a separate page rather than
a section. What is recorded is one number, host side, fit only. Peak process
memory includes the interpreter, the imported framework and the dataset, so
without a per-framework baseline it largely ranks frameworks; on problems like
EigenWorms the data dominates the model. Host and device memory are different
quantities and only the host one is recorded. Framework allocators hold memory
they are not using, so a naive device reading measures allocator policy.

The point that decided the separate page is censoring. Where memory mattered
most we have no number at all, only a bound: missing_results.csv records
ConvTran hitting CUDA out of memory on Alzheimers, EigenWorms and
PhotoStimulation, and FreshPRINCE hitting OOM at 128 GB on FaceDetection,
FordChallenge and Skoda after eight attempts. A table built from successful runs
alone would be survivorship-biased in a way a runtime table is not.

The runtime page is narrowed to runtime and the two cross-reference.

Co-Authored-By: Claude Opus 5 
---
 README.md       |  2 ++
 docs/memory.md  | 92 +++++++++++++++++++++++++++++++++++++++++++++++++
 docs/runtime.md | 18 ++++------
 3 files changed, 101 insertions(+), 11 deletions(-)
 create mode 100644 docs/memory.md

diff --git a/README.md b/README.md
index db6f3ca..6b0a432 100644
--- a/README.md
+++ b/README.md
@@ -101,6 +101,8 @@ This repository aims to make it easier to:
   ·
   Runtime
   ·
+  Memory
+  ·
   Evaluation
   ·
   Classifiers
diff --git a/docs/memory.md b/docs/memory.md
new file mode 100644
index 0000000..00e2070
--- /dev/null
+++ b/docs/memory.md
@@ -0,0 +1,92 @@
+# Memory
+
+**Coming soon.** This page will hold memory comparisons across the Multiverse
+estimators.
+
+**We have not yet structured an experiment to compare memory.** As with
+[runtime](runtime.md), every run behind the results in this repository was set up to
+measure predictive performance. A job was given whatever memory ceiling got it to
+finish, on whatever node was free, with whatever core count came with the partition,
+and those were never held constant across estimators because nothing depended on it.
+The peak figures that came out are a by-product of that, not a measurement anyone
+designed.
+
+So nothing is published here yet, deliberately. The rest of this page records what a
+memory comparison would have to fix, and why the figures we already hold cannot stand
+in for one.
+
+## What is recorded
+
+`tsml-eval` writes one `memory_usage` value per classifier, dataset and resample: the
+peak memory observed during `fit`. It is in the raw prediction files, and this
+repository's ingest brings across only the accuracy-style measures, so it has not been
+carried over.
+
+One number, host side, fit only.
+
+## Why the figures we hold cannot stand in
+
+Each of these is a condition a memory experiment would have to fix, and that these runs
+left free.
+
+**Peak process memory is not the model's memory.** It includes the interpreter, every
+imported library, the loaded dataset and any transient copies made along the way.
+Importing TensorFlow or PyTorch alone accounts for a large fixed cost before a single
+weight is allocated, so a small model in a heavy framework can report more than a large
+model in a light one. Without subtracting a per-framework baseline the figure mostly
+ranks frameworks.
+
+**The dataset can dominate the model.** Multiverse contains problems like EigenWorms at
+17,984 timepoints and FaceDetection at 5,890 training cases. For an estimator with a
+small parameter count, most of the peak is the data and its copies, which says
+something about the problem rather than the method.
+
+**Host and device memory are different quantities, and only one is recorded.** A model
+doing its work on a GPU can look inexpensive by peak host memory while occupying tens of
+gigabytes of device memory that nothing here measures. The two are not interchangeable
+and cannot be added.
+
+**Framework allocators do not report what the model needs.** PyTorch's caching allocator
+and TensorFlow's default of reserving most of the visible GPU both hold memory they are
+not using, so a naive reading measures the allocator's policy rather than the model's
+requirement. Getting a meaningful device figure means enabling TensorFlow's memory
+growth and reading PyTorch's allocated rather than reserved totals, neither of which
+these runs did.
+
+**Failures censor the measurement.** Where an estimator ran out of memory we do not have
+a peak, we have a lower bound and a ceiling. Both kinds appear in
+`results/multiverse/missing_results.csv`: ConvTran hit CUDA out of memory on Alzheimers,
+EigenWorms and PhotoStimulation, a device-side limit; FreshPRINCE hit OOM at 128 GB on
+FaceDetection, FordChallenge and Skoda after eight attempts, a host-side one. Those are
+the cases where memory mattered most, and they are exactly the cases with no number. A
+table built only from successful runs is a survivorship-biased view of memory use.
+
+**Memory scales with the resources granted.** The classical ensembles allocate per
+thread, so their peak moves with the cores allocated, which varied by partition. Our
+controllers also escalate a job's request from 64 GB to 128 GB after a failure, so
+different runs of the same estimator saw different ceilings.
+
+**Peak is timing-dependent.** Python's peak resident memory depends on when garbage
+collection happens to run and on whether an allocator returned pages to the operating
+system. Repeats of an identical run differ for reasons that have nothing to do with the
+method.
+
+## What a fair comparison would need
+
+- All compared estimators on the same node, with cores and the memory ceiling fixed and
+  recorded.
+- A per-framework baseline measured and subtracted, so the figure is the model's cost
+  rather than the cost of importing its library.
+- Host and device peaks reported separately, never summed, with TensorFlow memory growth
+  enabled and PyTorch read via its allocated totals.
+- `predict` measured as well as `fit`. Deployment cost is a separate question from
+  training cost, and the ranking is not the same on both.
+- Failures reported alongside successes, as censored observations with the ceiling that
+  was in force, rather than dropped.
+- Repeated runs, since peak varies between identical repeats.
+- Asymptotic space complexity in the number of cases, series length and channels stated
+  beside the measured peaks, so a reader can tell whether a figure will hold at a
+  different scale.
+
+Until most of that is in place, this page stays empty. For the same reason the
+[leaderboard](leaderboard.md) carries no memory column.
diff --git a/docs/runtime.md b/docs/runtime.md
index 02f314c..c259acd 100644
--- a/docs/runtime.md
+++ b/docs/runtime.md
@@ -1,7 +1,8 @@
 # Runtime
 
-**Coming soon.** This page will hold runtime and memory comparisons across the
-Multiverse estimators.
+**Coming soon.** This page will hold runtime comparisons across the Multiverse
+estimators. Memory is a separate page, [memory](memory.md), for the same reasons and a
+few of its own.
 
 **We have not yet structured an experiment to compare runtime.** Every run behind the
 results in this repository was set up to measure predictive performance. Which partition
@@ -22,9 +23,9 @@ Every run already records timings. `tsml-eval` writes, per classifier, dataset a
 resample:
 
 - `fit_time` and `predict_time`, in seconds;
-- `memory_usage`, the peak memory during `fit`;
 - `benchmark_time`, the time that machine took to sort 1,000 seeded random arrays of
-  20,000 elements.
+  20,000 elements;
+- `memory_usage`, the peak memory during `fit`, which [memory](memory.md) covers.
 
 They are in the raw prediction files. What this repository ingests under `results/` is
 only the accuracy-style measures, one file per metric, so the timings have not been
@@ -59,12 +60,6 @@ faithful ports of the same paper can differ several-fold on time because the aut
 picked 500 epochs and the toolkit's default is 2000. Early stopping and best-epoch
 selection move it again. None of that is a fact about the architecture.
 
-**Memory is measured in one place and spent in another.** `memory_usage` is peak host
-memory during `fit`. A model doing all its work on a GPU can look inexpensive by that
-measure while occupying tens of gigabytes of device memory, which the figure never sees.
-Peak host memory and peak device memory are different quantities and only one is
-recorded.
-
 **Which device a run actually used is not reliably recorded.** In the version of the
 experiment tooling used for these runs, the device description inspects TensorFlow only,
 so a PyTorch estimator reports CPU whether or not it ran on a GPU. Any timing table built
@@ -88,4 +83,5 @@ configuration rather than trusted as written.
 - The device stated per run, from the job configuration.
 
 Until most of that is in place, this page stays empty. For the same reason the
-[leaderboard](leaderboard.md) carries no fit time, predict time or memory columns.
+[leaderboard](leaderboard.md) carries no fit time or predict time columns, and
+[memory](memory.md) is empty too.

From 0fcf08dd8f5a04c3a3f7bab0dae7113e822fc146 Mon Sep 17 00:00:00 2001
From: Tony Bagnall 
Date: Thu, 3 Sep 2026 10:59:17 +0100
Subject: [PATCH 6/8] Ingest XCM

65 of 66 Multiverse-core datasets, seven metrics, from the i7_h200 run at the
port's defaults: window_size 0.8 capped at 100 points, batch size 32, 100
epochs. EmoPain is the one gap and is the archive-wide one, aeon's input
validation rejecting it before fit for every classifier, now recorded in
missing_results.csv beside the other four.

XCM lands 21st of 25 on average accuracy rank, mean accuracy 0.6690.

The ingest logged "y_prob values do not sum to one" on many datasets. The
probabilities are fine: the largest deviation across all 65 is 2.8e-07, ordinary
float32 rounding from the Keras softmax, well inside anything that would affect
log loss or AUROC.

Against the paper's own table, on the 23 shared datasets, we average 0.699 where
Fauvel et al. report 0.761, a gap of 6.2 points. That is the expected size: they
grid-search window_size per dataset over {20,40,60,80,100}% by cross-validation
on the training set, and report a 7.0% mean relative accuracy drop from using a
suboptimal window, where this run uses the modal 0.8 everywhere. Unlike
DisjointCNN's 20.3 points, this gap is explained by a documented protocol
difference rather than an implementation fault.

Co-Authored-By: Claude Opus 5 
---
 README.md                                  | 49 ++++++++--------
 results/multiverse/XCM/XCM_accuracy.csv    | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_auroc.csv       | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_balacc.csv      | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_f1.csv          | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_logloss.csv     | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_sensitivity.csv | 66 ++++++++++++++++++++++
 results/multiverse/XCM/XCM_specificity.csv | 66 ++++++++++++++++++++++
 results/multiverse/datasets.html           |  2 +-
 results/multiverse/leaderboard.html        |  4 +-
 results/multiverse/missing_results.csv     |  1 +
 11 files changed, 491 insertions(+), 27 deletions(-)
 create mode 100644 results/multiverse/XCM/XCM_accuracy.csv
 create mode 100644 results/multiverse/XCM/XCM_auroc.csv
 create mode 100644 results/multiverse/XCM/XCM_balacc.csv
 create mode 100644 results/multiverse/XCM/XCM_f1.csv
 create mode 100644 results/multiverse/XCM/XCM_logloss.csv
 create mode 100644 results/multiverse/XCM/XCM_sensitivity.csv
 create mode 100644 results/multiverse/XCM/XCM_specificity.csv

diff --git a/README.md b/README.md
index 6b0a432..7d075d6 100644
--- a/README.md
+++ b/README.md
@@ -40,30 +40,31 @@ The current paper version describes:
 
 | # | Estimator | Accuracy rank | Accuracy | Balanced accuracy | AUROC | F1 | Log loss ↓ | Sensitivity | Specificity |
 |---|---|---|---|---|---|---|---|---|---|
-| 1 | HC2 | **7.86** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** |
-| 2 | MRHydra | 8.31 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 |
-| 3 | RDST | 9.16 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 |
-| 4 | RIST | 9.76 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 |
-| 5 | DrCIF | 10.07 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 |
-| 6 | FreshPRINCE | 10.10 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 |
-| 7 | CIF | 10.13 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 |
-| 8 | QUANT | 10.55 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 |
-| 9 | Arsenal | 10.70 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 |
-| 10 | ROCKET | 10.82 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 |
-| 11 | LITETime-MV | 11.14 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 |
-| 12 | STSF | 11.55 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 |
-| 13 | H-InceptionTime | 11.63 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 |
-| 14 | LiteTIME | 12.34 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 |
-| 15 | PatchMTSC | 13.12 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 |
-| 16 | ConvTran | 13.14 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 |
-| 17 | Catch22 | 13.15 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 |
-| 18 | STC | 13.95 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 |
-| 19 | TSF | 14.00 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 |
-| 20 | TDE | 14.63 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 |
-| 21 | Summary | 16.66 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 |
-| 22 | TimesURL | 16.95 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 |
-| 23 | 1NN-DTW | 18.45 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 |
-| 24 | Dummy | 21.81 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 |
+| 1 | HC2 | **8.07** | **0.7909** | 0.7518 | **0.8990** | 0.7273 | **0.5383** | 0.7459 | **0.7943** |
+| 2 | MRHydra | 8.49 | 0.7837 | **0.7564** | 0.8105 | **0.7316** | 7.7974 | **0.7642** | 0.7757 |
+| 3 | RDST | 9.34 | 0.7734 | 0.7333 | 0.7912 | 0.6991 | 8.1667 | 0.7109 | 0.7874 |
+| 4 | RIST | 10.01 | 0.7720 | 0.7397 | 0.8748 | 0.7147 | 0.6218 | 0.7408 | 0.7655 |
+| 5 | DrCIF | 10.35 | 0.7747 | 0.7429 | 0.8813 | 0.7173 | 0.6484 | 0.7397 | 0.7708 |
+| 6 | FreshPRINCE | 10.35 | 0.7743 | 0.7487 | 0.8745 | 0.7211 | 0.6007 | 0.7414 | 0.7770 |
+| 7 | CIF | 10.41 | 0.7781 | 0.7471 | 0.8908 | 0.7212 | 0.6430 | 0.7441 | 0.7753 |
+| 8 | QUANT | 10.82 | 0.7720 | 0.7462 | 0.8831 | 0.7189 | 0.6175 | 0.7521 | 0.7581 |
+| 9 | Arsenal | 11.00 | 0.7680 | 0.7321 | 0.8457 | 0.7024 | 3.8631 | 0.7257 | 0.7732 |
+| 10 | ROCKET | 11.08 | 0.7690 | 0.7326 | 0.7925 | 0.7019 | 8.3249 | 0.7200 | 0.7764 |
+| 11 | LITETime-MV | 11.41 | 0.7506 | 0.7299 | 0.8513 | 0.6820 | 1.3206 | 0.7132 | 0.7637 |
+| 12 | STSF | 11.86 | 0.7724 | 0.7477 | 0.8804 | 0.7080 | 0.6432 | 0.7345 | 0.7826 |
+| 13 | H-InceptionTime | 11.88 | 0.7408 | 0.7190 | 0.8496 | 0.6838 | 1.3227 | 0.7223 | 0.7378 |
+| 14 | LiteTIME | 12.63 | 0.7341 | 0.7104 | 0.8394 | 0.6680 | 1.4776 | 0.7113 | 0.7336 |
+| 15 | PatchMTSC | 13.49 | 0.7428 | 0.6897 | 0.8261 | 0.6533 | 0.7655 | 0.6852 | 0.7352 |
+| 16 | ConvTran | 13.51 | 0.7462 | 0.7102 | 0.8592 | 0.6767 | 0.8190 | 0.7159 | 0.7345 |
+| 17 | Catch22 | 13.55 | 0.7475 | 0.7181 | 0.8697 | 0.6922 | 0.7147 | 0.7240 | 0.7374 |
+| 18 | STC | 14.30 | 0.7545 | 0.7172 | 0.8744 | 0.6940 | 0.6391 | 0.7185 | 0.7537 |
+| 19 | TSF | 14.35 | 0.7515 | 0.7236 | 0.8740 | 0.6883 | 0.7252 | 0.7093 | 0.7606 |
+| 20 | TDE | 15.02 | 0.7262 | 0.6813 | 0.8374 | 0.6382 | 0.8869 | 0.6714 | 0.7344 |
+| 21 | XCM | 16.68 | 0.6743 | 0.6412 | 0.7982 | 0.5802 | 1.8918 | 0.6219 | 0.6851 |
+| 22 | Summary | 17.22 | 0.6858 | 0.6574 | 0.8268 | 0.6230 | 0.9123 | 0.6574 | 0.6844 |
+| 23 | TimesURL | 17.53 | 0.6958 | 0.6533 | 0.7906 | 0.5967 | 1.0055 | 0.6257 | 0.6973 |
+| 24 | 1NN-DTW | 19.03 | 0.6712 | 0.6454 | 0.7197 | 0.6136 | 11.8506 | 0.6521 | 0.6636 |
+| 25 | Dummy | 22.64 | 0.3645 | 0.3029 | 0.5000 | 0.1507 | 1.4067 | 0.2855 | 0.3816 |
 
 Average over the 52 Multiverse-core datasets with results for every estimator on every metric, ordered by average accuracy rank. Best in each column in bold.
 
diff --git a/results/multiverse/XCM/XCM_accuracy.csv b/results/multiverse/XCM/XCM_accuracy.csv
new file mode 100644
index 0000000..fcddf01
--- /dev/null
+++ b/results/multiverse/XCM/XCM_accuracy.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.3023255813953488
+AppliancesEnergy_disc,0.8095238095238095
+ArticularyWordRecognition,0.9466666666666667
+AsphaltObstaclesCoordinates,0.7468030690537084
+AsphaltRegularityCoordinates,0.9733688415446072
+AtrialFibrillation,0.2
+AustraliaRainfall_disc,0.7725504877186414
+AutomotiveRoadTrials,0.7142857142857143
+BIDMC32HR_disc,0.6348478532721967
+BIDMC32SpO2_disc,0.5410587744893706
+BeijingPM10Quality_disc,0.715729001584786
+BeijingPM25Quality_disc,0.8332012678288431
+BenzeneConcentration_disc,0.7640906449738524
+Blink,0.9244444444444444
+BoneIntensitiesAgeGroup,0.7528089887640449
+BoneProbAgeGroup,0.6764044943820224
+CharacterTrajectories,0.9937325905292479
+CounterMovementJump,0.8212290502793296
+Cricket,0.9861111111111112
+CrowdSourced,0.7518531591951995
+DuckDuckGeese,0.54
+ERing,0.3037037037037037
+EigenWorms,0.5038167938931297
+Epilepsy,0.9130434782608695
+EthanolConcentration,0.2889733840304182
+EyesOpenShut,0.5
+FaceDetection,0.6174801362088536
+FordChallenge,0.6232423490488007
+HandMovementDirection,0.28378378378378377
+Handwriting,0.3611764705882353
+Heartbeat,0.7560975609756098
+HouseholdPowerConsumption1_disc,0.11807580174927114
+HouseholdPowerConsumption2_disc,0.6967930029154519
+IEEEPPG_disc,0.6453313253012049
+IRDS-SFL,0.896551724137931
+JapaneseVowels,0.9783783783783784
+KERAAL-RTK,0.5714285714285714
+KIMORE-PR-C,0.42857142857142855
+KINECAL-QSEO,0.29411764705882354
+LSST,0.45174371451743717
+Libras,0.8277777777777777
+Locust2022,0.8681418611700515
+LowCost,0.52
+MindReading,0.4119448698315467
+MotionSenseHAR,0.969811320754717
+MotorImagery,0.51
+NATOPS,0.9
+PEMS-SF,0.8323699421965318
+PenDigits,0.949685534591195
+PhonemeSpectra,0.1258574410975246
+PhotoStimulation,0.4166666666666667
+RacketSports,0.8618421052631579
+STEW,0.6557239057239057
+SelfRegulationSCP1,0.552901023890785
+Skoda,0.9405730456314114
+SpokenArabicDigits,0.9845384265575261
+StandWalkJump,0.2
+TactileTextureRecognition,0.7650513950073421
+Tiselac,0.7832153690596562
+UCDHE-Rowing-MC,0.6772727272727272
+UCIActivity,0.9837567680133278
+UIPRMD-DS-C,0.7222222222222222
+USCActivity,0.6830953827460511
+UWaveGestureLibrary,0.8375
+WISDM,0.873092524628163
diff --git a/results/multiverse/XCM/XCM_auroc.csv b/results/multiverse/XCM/XCM_auroc.csv
new file mode 100644
index 0000000..e1ed21e
--- /dev/null
+++ b/results/multiverse/XCM/XCM_auroc.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.40807638331996793
+AppliancesEnergy_disc,0.7022058823529411
+ArticularyWordRecognition,0.9983680555555554
+AsphaltObstaclesCoordinates,0.9282288995088815
+AsphaltRegularityCoordinates,0.9947222813364545
+AtrialFibrillation,0.36000000000000004
+AustraliaRainfall_disc,0.8531741621798045
+AutomotiveRoadTrials,0.6814882032667877
+BIDMC32HR_disc,0.8257225160774645
+BIDMC32SpO2_disc,0.4466540048531264
+BeijingPM10Quality_disc,0.8228868866803458
+BeijingPM25Quality_disc,0.9209492402371762
+BenzeneConcentration_disc,0.7906529888748304
+Blink,0.94976
+BoneIntensitiesAgeGroup,0.8763879280693906
+BoneProbAgeGroup,0.8482297247796746
+CharacterTrajectories,0.999966482814142
+CounterMovementJump,0.962704411373488
+Cricket,0.9983164983164983
+CrowdSourced,0.7769353372486634
+DuckDuckGeese,0.8835
+ERing,0.9387489711934157
+EigenWorms,0.8730167577171467
+Epilepsy,0.9757065171400083
+EthanolConcentration,0.5558911771202598
+EyesOpenShut,0.2380952380952381
+FaceDetection,0.6725854880623994
+FordChallenge,0.5
+HandMovementDirection,0.5678910034842238
+Handwriting,0.8556383071775312
+Heartbeat,0.7316263632053106
+HouseholdPowerConsumption1_disc,0.6215145245466448
+HouseholdPowerConsumption2_disc,0.6832528326022078
+IEEEPPG_disc,0.8553718178469635
+IRDS-SFL,0.9637681159420289
+JapaneseVowels,0.9995787852093061
+KERAAL-RTK,0.5
+KIMORE-PR-C,0.5
+KINECAL-QSEO,0.25
+LSST,0.8052505045504954
+Libras,0.9869378306878308
+Locust2022,0.7439370405945788
+LowCost,0.5296222222222222
+MindReading,0.7188556235726715
+MotionSenseHAR,0.9927211095965354
+MotorImagery,0.5826
+NATOPS,0.9912222222222221
+PEMS-SF,0.9661758891337312
+PenDigits,0.9975886205572294
+PhonemeSpectra,0.8201195011847978
+PhotoStimulation,0.38638380138380135
+RacketSports,0.9724577717009012
+STEW,0.8272616975969951
+SelfRegulationSCP1,0.6061876805516726
+Skoda,0.9959321424656511
+SpokenArabicDigits,0.9998641964066783
+StandWalkJump,0.4
+TactileTextureRecognition,0.9863492753156763
+Tiselac,0.9526128233026999
+UCDHE-Rowing-MC,0.9200922262025203
+UCIActivity,0.9995001270676686
+UIPRMD-DS-C,0.8148148148148149
+USCActivity,0.9533753889936101
+UWaveGestureLibrary,0.9718861607142857
+WISDM,0.9644757073546718
diff --git a/results/multiverse/XCM/XCM_balacc.csv b/results/multiverse/XCM/XCM_balacc.csv
new file mode 100644
index 0000000..9211a6b
--- /dev/null
+++ b/results/multiverse/XCM/XCM_balacc.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.367965367965368
+AppliancesEnergy_disc,0.5
+ArticularyWordRecognition,0.9466666666666665
+AsphaltObstaclesCoordinates,0.750158283785592
+AsphaltRegularityCoordinates,0.9735191884798184
+AtrialFibrillation,0.20000000000000004
+AustraliaRainfall_disc,0.38160086861129444
+AutomotiveRoadTrials,0.6333938294010889
+BIDMC32HR_disc,0.43818415153971185
+BIDMC32SpO2_disc,0.4103754347165767
+BeijingPM10Quality_disc,0.5304937507403205
+BeijingPM25Quality_disc,0.8490841689347542
+BenzeneConcentration_disc,0.6232841845755148
+Blink,0.9315
+BoneIntensitiesAgeGroup,0.791771432890703
+BoneProbAgeGroup,0.673252210372801
+CharacterTrajectories,0.9933931956260587
+CounterMovementJump,0.8216572504708098
+Cricket,0.9861111111111112
+CrowdSourced,0.7519220303099171
+DuckDuckGeese,0.54
+ERing,0.30370370370370375
+EigenWorms,0.36
+Epilepsy,0.9141494435612083
+EthanolConcentration,0.28805361305361304
+EyesOpenShut,0.5
+FaceDetection,0.6174801362088536
+FordChallenge,0.5
+HandMovementDirection,0.30238095238095236
+Handwriting,0.35229571326491305
+Heartbeat,0.599158368895211
+HouseholdPowerConsumption1_disc,0.3253012048192771
+HouseholdPowerConsumption2_disc,0.3894931045615977
+IEEEPPG_disc,0.604697162742921
+IRDS-SFL,0.9347826086956521
+JapaneseVowels,0.9757319741190709
+KERAAL-RTK,0.5
+KIMORE-PR-C,0.6666666666666666
+KINECAL-QSEO,0.15625
+LSST,0.22954875180846224
+Libras,0.8277777777777777
+Locust2022,0.5672494601241204
+LowCost,0.52
+MindReading,0.39793855343461887
+MotionSenseHAR,0.9599330551208486
+MotorImagery,0.51
+NATOPS,0.8999999999999999
+PEMS-SF,0.8300086387042909
+PenDigits,0.9503633306991516
+PhonemeSpectra,0.1258725314812866
+PhotoStimulation,0.3414141414141414
+RacketSports,0.8704734219269104
+STEW,0.6557239057239057
+SelfRegulationSCP1,0.554398471717454
+Skoda,0.9295284795279287
+SpokenArabicDigits,0.9845433789954339
+StandWalkJump,0.20000000000000004
+TactileTextureRecognition,0.7676131153558414
+Tiselac,0.5813305834575884
+UCDHE-Rowing-MC,0.697642857142857
+UCIActivity,0.9842652561389942
+UIPRMD-DS-C,0.7222222222222222
+USCActivity,0.6872296744351977
+UWaveGestureLibrary,0.8374999999999999
+WISDM,0.547020066421741
diff --git a/results/multiverse/XCM/XCM_f1.csv b/results/multiverse/XCM/XCM_f1.csv
new file mode 100644
index 0000000..2576c12
--- /dev/null
+++ b/results/multiverse/XCM/XCM_f1.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.21816168327796231
+AppliancesEnergy_disc,0.0
+ArticularyWordRecognition,0.9462743362917276
+AsphaltObstaclesCoordinates,0.7491477045617893
+AsphaltRegularityCoordinates,0.9732620320855615
+AtrialFibrillation,0.16190476190476188
+AustraliaRainfall_disc,0.75719188582855
+AutomotiveRoadTrials,0.45
+BIDMC32HR_disc,0.6518311866365318
+BIDMC32SpO2_disc,0.11708099438652766
+BeijingPM10Quality_disc,0.15736934820904286
+BeijingPM25Quality_disc,0.7632170978627671
+BenzeneConcentration_disc,0.39881539980256664
+Blink,0.9212962962962963
+BoneIntensitiesAgeGroup,0.7478544452020218
+BoneProbAgeGroup,0.6779403896390017
+CharacterTrajectories,0.9937101179249307
+CounterMovementJump,0.8167930322758735
+Cricket,0.9860139860139859
+CrowdSourced,0.7923190546528803
+DuckDuckGeese,0.5323809523809524
+ERing,0.1517735755778025
+EigenWorms,0.40159309658148024
+Epilepsy,0.9112318840579711
+EthanolConcentration,0.24591210249853798
+EyesOpenShut,0.6666666666666666
+FaceDetection,0.6516795865633075
+FordChallenge,0.0
+HandMovementDirection,0.2781612959451537
+Handwriting,0.3377079846984321
+Heartbeat,0.358974358974359
+HouseholdPowerConsumption1_disc,0.025756350972902773
+HouseholdPowerConsumption2_disc,0.6557375281190868
+IEEEPPG_disc,0.6225822295456521
+IRDS-SFL,0.8
+JapaneseVowels,0.9781367493415885
+KERAAL-RTK,0.0
+KIMORE-PR-C,0.3333333333333333
+KINECAL-QSEO,0.0
+LSST,0.3791312736547366
+Libras,0.819467577984202
+Locust2022,0.21338155515370705
+LowCost,0.6317135549872123
+MindReading,0.40190722407655577
+MotionSenseHAR,0.9728201579348978
+MotorImagery,0.07547169811320754
+NATOPS,0.9003106182590119
+PEMS-SF,0.8278898581282377
+PenDigits,0.9496829926785575
+PhonemeSpectra,0.09434448848932879
+PhotoStimulation,0.2851037851037851
+RacketSports,0.8588118410519711
+STEW,0.7288397790055249
+SelfRegulationSCP1,0.6888361045130641
+Skoda,0.9407162240986361
+SpokenArabicDigits,0.9845196910744067
+StandWalkJump,0.1574074074074074
+TactileTextureRecognition,0.7638645349615292
+Tiselac,0.7722507644925846
+UCDHE-Rowing-MC,0.6681854984389035
+UCIActivity,0.9837608130719198
+UIPRMD-DS-C,0.6153846153846154
+USCActivity,0.6796853915511324
+UWaveGestureLibrary,0.8385274008139714
+WISDM,0.8608499413848607
diff --git a/results/multiverse/XCM/XCM_logloss.csv b/results/multiverse/XCM/XCM_logloss.csv
new file mode 100644
index 0000000..2a5c73e
--- /dev/null
+++ b/results/multiverse/XCM/XCM_logloss.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,1.281995397783109
+AppliancesEnergy_disc,1.220613288803126
+ArticularyWordRecognition,0.20913287122187446
+AsphaltObstaclesCoordinates,0.9458313154759524
+AsphaltRegularityCoordinates,0.1266260204259907
+AtrialFibrillation,1.1886000723607923
+AustraliaRainfall_disc,0.5192454015724045
+AutomotiveRoadTrials,1.531994513704828
+BIDMC32HR_disc,2.6153268183918863
+BIDMC32SpO2_disc,3.2864579433764276
+BeijingPM10Quality_disc,1.192763567851187
+BeijingPM25Quality_disc,0.5700050251191676
+BenzeneConcentration_disc,3.346307447421879
+Blink,0.6552198963655975
+BoneIntensitiesAgeGroup,0.7577387737965597
+BoneProbAgeGroup,1.3781207487249014
+CharacterTrajectories,0.025087426830150132
+CounterMovementJump,0.4458332415980918
+Cricket,0.23899100869077017
+CrowdSourced,5.681748011476673
+DuckDuckGeese,1.4476585961906254
+ERing,1.5511330924068825
+EigenWorms,2.7894721126602673
+Epilepsy,0.3715004381698749
+EthanolConcentration,1.529273939296933
+EyesOpenShut,4.248528533555482
+FaceDetection,4.314418813211354
+FordChallenge,13.571439293456832
+HandMovementDirection,1.5513074783505933
+Handwriting,2.480634334028785
+Heartbeat,0.8112859994036347
+HouseholdPowerConsumption1_disc,16.637074689936
+HouseholdPowerConsumption2_disc,1.9238153847023434
+IEEEPPG_disc,1.4806723442802172
+IRDS-SFL,0.5406823350671576
+JapaneseVowels,0.07008161110194283
+KERAAL-RTK,5.5033371337501675
+KIMORE-PR-C,8.192123345424639
+KINECAL-QSEO,1.509511533283483
+LSST,2.6029082409706916
+Libras,0.5640605291689049
+Locust2022,0.33809055946445105
+LowCost,1.1380963668204618
+MindReading,2.8786347532333054
+MotionSenseHAR,0.19610327002131758
+MotorImagery,2.977766816238155
+NATOPS,0.2578673378413515
+PEMS-SF,0.7632564329767687
+PenDigits,0.2884284956829766
+PhonemeSpectra,4.032035001666196
+PhotoStimulation,1.1993444304140304
+RacketSports,0.39196803556227966
+STEW,3.4722974016081136
+SelfRegulationSCP1,2.4810720946580234
+Skoda,0.24444171624728298
+SpokenArabicDigits,0.06498082123670214
+StandWalkJump,1.838150270556138
+TactileTextureRecognition,1.3949109714966077
+Tiselac,0.9638490690786189
+UCDHE-Rowing-MC,1.7944030264704292
+UCIActivity,0.06438419313786149
+UIPRMD-DS-C,2.579849189540059
+USCActivity,1.7029310258175514
+UWaveGestureLibrary,0.6536157560205249
+WISDM,1.999502230004768
diff --git a/results/multiverse/XCM/XCM_sensitivity.csv b/results/multiverse/XCM/XCM_sensitivity.csv
new file mode 100644
index 0000000..e9242f0
--- /dev/null
+++ b/results/multiverse/XCM/XCM_sensitivity.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.3023255813953488
+AppliancesEnergy_disc,0.0
+ArticularyWordRecognition,0.9466666666666667
+AsphaltObstaclesCoordinates,0.7468030690537084
+AsphaltRegularityCoordinates,0.9837837837837838
+AtrialFibrillation,0.2
+AustraliaRainfall_disc,0.7725504877186414
+AutomotiveRoadTrials,0.47368421052631576
+BIDMC32HR_disc,0.6348478532721967
+BIDMC32SpO2_disc,0.10688140556368961
+BeijingPM10Quality_disc,0.09190672153635117
+BeijingPM25Quality_disc,0.8892529488859764
+BenzeneConcentration_disc,0.25218476903870163
+Blink,0.995
+BoneIntensitiesAgeGroup,0.7528089887640449
+BoneProbAgeGroup,0.6764044943820224
+CharacterTrajectories,0.9937325905292479
+CounterMovementJump,0.8212290502793296
+Cricket,0.9861111111111112
+CrowdSourced,0.9470338983050848
+DuckDuckGeese,0.54
+ERing,0.3037037037037037
+EigenWorms,0.5038167938931297
+Epilepsy,0.9130434782608695
+EthanolConcentration,0.2889733840304182
+EyesOpenShut,1.0
+FaceDetection,0.7156640181611805
+FordChallenge,0.0
+HandMovementDirection,0.28378378378378377
+Handwriting,0.3611764705882353
+Heartbeat,0.24561403508771928
+HouseholdPowerConsumption1_disc,0.11807580174927114
+HouseholdPowerConsumption2_disc,0.6967930029154519
+IEEEPPG_disc,0.6453313253012049
+IRDS-SFL,1.0
+JapaneseVowels,0.9783783783783784
+KERAAL-RTK,0.0
+KIMORE-PR-C,1.0
+KINECAL-QSEO,0.0
+LSST,0.45174371451743717
+Libras,0.8277777777777777
+Locust2022,0.20136518771331058
+LowCost,0.8233333333333334
+MindReading,0.4119448698315467
+MotionSenseHAR,0.969811320754717
+MotorImagery,0.04
+NATOPS,0.9
+PEMS-SF,0.8323699421965318
+PenDigits,0.949685534591195
+PhonemeSpectra,0.1258574410975246
+PhotoStimulation,0.4166666666666667
+RacketSports,0.8618421052631579
+STEW,0.925364758698092
+SelfRegulationSCP1,0.9931506849315068
+Skoda,0.9405730456314114
+SpokenArabicDigits,0.9845384265575261
+StandWalkJump,0.2
+TactileTextureRecognition,0.7650513950073421
+Tiselac,0.7832153690596562
+UCDHE-Rowing-MC,0.6772727272727272
+UCIActivity,0.9837567680133278
+UIPRMD-DS-C,0.4444444444444444
+USCActivity,0.6830953827460511
+UWaveGestureLibrary,0.8375
+WISDM,0.873092524628163
diff --git a/results/multiverse/XCM/XCM_specificity.csv b/results/multiverse/XCM/XCM_specificity.csv
new file mode 100644
index 0000000..bc472f0
--- /dev/null
+++ b/results/multiverse/XCM/XCM_specificity.csv
@@ -0,0 +1,66 @@
+Resamples:,0
+Alzheimers,0.3023255813953488
+AppliancesEnergy_disc,1.0
+ArticularyWordRecognition,0.9466666666666667
+AsphaltObstaclesCoordinates,0.7468030690537084
+AsphaltRegularityCoordinates,0.963254593175853
+AtrialFibrillation,0.2
+AustraliaRainfall_disc,0.7725504877186414
+AutomotiveRoadTrials,0.7931034482758621
+BIDMC32HR_disc,0.6348478532721967
+BIDMC32SpO2_disc,0.7138694638694638
+BeijingPM10Quality_disc,0.9690807799442896
+BeijingPM25Quality_disc,0.8089153889835321
+BenzeneConcentration_disc,0.994383600112328
+Blink,0.868
+BoneIntensitiesAgeGroup,0.7528089887640449
+BoneProbAgeGroup,0.6764044943820224
+CharacterTrajectories,0.9937325905292479
+CounterMovementJump,0.8212290502793296
+Cricket,0.9861111111111112
+CrowdSourced,0.5568101623147494
+DuckDuckGeese,0.54
+ERing,0.3037037037037037
+EigenWorms,0.5038167938931297
+Epilepsy,0.9130434782608695
+EthanolConcentration,0.2889733840304182
+EyesOpenShut,0.0
+FaceDetection,0.5192962542565267
+FordChallenge,1.0
+HandMovementDirection,0.28378378378378377
+Handwriting,0.3611764705882353
+Heartbeat,0.9527027027027027
+HouseholdPowerConsumption1_disc,0.11807580174927114
+HouseholdPowerConsumption2_disc,0.6967930029154519
+IEEEPPG_disc,0.6453313253012049
+IRDS-SFL,0.8695652173913043
+JapaneseVowels,0.9783783783783784
+KERAAL-RTK,1.0
+KIMORE-PR-C,0.3333333333333333
+KINECAL-QSEO,0.3125
+LSST,0.45174371451743717
+Libras,0.8277777777777777
+Locust2022,0.9331337325349301
+LowCost,0.21666666666666667
+MindReading,0.4119448698315467
+MotionSenseHAR,0.969811320754717
+MotorImagery,0.98
+NATOPS,0.9
+PEMS-SF,0.8323699421965318
+PenDigits,0.949685534591195
+PhonemeSpectra,0.1258574410975246
+PhotoStimulation,0.4166666666666667
+RacketSports,0.8618421052631579
+STEW,0.38608305274971944
+SelfRegulationSCP1,0.11564625850340136
+Skoda,0.9405730456314114
+SpokenArabicDigits,0.9845384265575261
+StandWalkJump,0.2
+TactileTextureRecognition,0.7650513950073421
+Tiselac,0.7832153690596562
+UCDHE-Rowing-MC,0.6772727272727272
+UCIActivity,0.9837567680133278
+UIPRMD-DS-C,1.0
+USCActivity,0.6830953827460511
+UWaveGestureLibrary,0.8375
+WISDM,0.873092524628163
diff --git a/results/multiverse/datasets.html b/results/multiverse/datasets.html
index 2729334..6cab76c 100644
--- a/results/multiverse/datasets.html
+++ b/results/multiverse/datasets.html
@@ -60,7 +60,7 @@
 details { margin-top: .6rem; }
 summary { cursor: pointer; color: var(--accent); }
 code { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: .9em; }
-tr.nosignal td { background: rgba(214, 158, 46, .16); }tr.saturated td { background: rgba(56, 161, 105, .14); }

Multiverse-core datasets: accuracy

66 datasets · accuracy · best of up to 23 estimators against the Dummy baseline · built 2026-09-02

DatasetDummyMedianBestBest estimatorGain over dummySpreadEstimators
AtrialFibrillation0.33330.20000.3333CIF0.00000.266723
KINECAL-QSEO0.94120.94120.9412Arsenal0.00000.117623
BIDMC32SpO2_disc0.71530.66280.7203ROCKET0.00500.171321
Locust20220.91120.90820.9206MRHydra0.00940.045223
Heartbeat0.72200.74630.7854CIF0.06340.126823
HouseholdPowerConsumption2_disc0.72160.76820.7872HC20.06560.141423
AutomotiveRoadTrials0.75320.80520.8442CIF0.09090.233823
AustraliaRainfall_disc0.68600.77310.7808LITETime-MV0.09480.088115
EyesOpenShut0.50000.50000.5952STSF0.09520.190523
MotorImagery0.50000.51000.6000FreshPRINCE0.10000.140023
BeijingPM10Quality_disc0.71120.82470.8417FreshPRINCE0.13050.107023
Alzheimers0.41860.37210.5581MRHydra0.13950.302322
AppliancesEnergy_disc0.80950.83330.9524FreshPRINCE0.14290.452423
EmoPain0.78310.84080.92681NN-DTW0.14370.242320
PhotoStimulation0.41670.37500.5833ROCKET0.16670.388922
FaceDetection0.50000.63040.6850H-InceptionTime0.18500.158622
BeijingPM25Quality_disc0.69770.87560.8879ConvTran0.19020.128823
HouseholdPowerConsumption1_disc0.77840.91250.9825FreshPRINCE0.20410.218723
LowCost0.50000.64170.7300TSF0.23000.248323
BoneProbAgeGroup0.47640.65390.7124H-InceptionTime0.23600.193323
StandWalkJump0.33330.46670.6000MRHydra0.26670.400023
CrowdSourced0.50020.71370.7734LITETime-MV0.27320.176823
BenzeneConcentration_disc0.68970.81950.9768STSF0.28700.577023
FordChallenge0.62320.88390.9360QUANT0.31280.312822
BIDMC32HR_disc0.65070.82580.9637RIST0.31300.635321
STEW0.50000.74380.8385Arsenal0.33850.210021
BoneIntensitiesAgeGroup0.47640.80450.8202HC20.34380.296623
PhonemeSpectra0.02560.27950.3746H-InceptionTime0.34890.291123
KERAAL-RTK0.57140.78570.9286HC20.35710.571423
LSST0.31510.63220.7040FreshPRINCE0.38890.480923
HandMovementDirection0.20270.41890.6081TSF0.40540.418923
DuckDuckGeese0.20000.46000.6400H-InceptionTime0.44000.480023
SelfRegulationSCP10.50170.85320.9454MRHydra0.44370.208223
IEEEPPG_disc0.26050.44730.7078ConvTran0.44730.438323
AsphaltRegularityCoordinates0.50730.98000.9947H-InceptionTime0.48740.291623
MindReading0.23120.53140.7243LITETime-MV0.49310.313923
EthanolConcentration0.25100.43350.7490STC0.49810.532323
UIPRMD-DS-C0.50000.83331.0000Catch220.50000.388923
WISDM0.36640.86730.8965MRHydra0.53000.131023
Blink0.44440.99331.0000Arsenal0.55560.415623
EigenWorms0.41980.89310.9771MRHydra0.55730.557322
KIMORE-PR-C0.14290.42860.7143LITETime-MV0.57140.571423
AsphaltObstaclesCoordinates0.28390.82100.8670MRHydra0.58310.289023
CounterMovementJump0.33520.75420.9274Arsenal0.59220.458123
Handwriting0.03760.37760.6529H-InceptionTime0.61530.467123
USCActivity0.11380.69360.7354LITETime-MV0.62160.137919
RacketSports0.28290.88160.9079RDST0.62500.125023
UCDHE-Rowing-MC0.20450.73860.8295PatchMTSC0.62500.338623
IRDS-SFL0.20690.79310.8621RDST0.65520.448323
Skoda0.23560.94800.9646H-InceptionTime0.72900.119922
Epilepsy0.26810.98551.0000HC20.73190.058023
Tiselac0.06280.81920.8373STSF0.77450.204418
MotionSenseHAR0.20380.98871.0000DrCIF0.79620.052823
NATOPS0.16670.88890.9667LITETime-MV0.80000.155623
UCIActivity0.19160.97460.9983LITETime-MV0.80670.174923
UWaveGestureLibrary0.12500.90940.9406Arsenal0.81560.553123
ERing0.16670.93330.9963MRHydra0.82960.233323
PEMS-SF0.11560.96531.0000CIF0.88440.317923
PenDigits0.10380.97860.9911H-InceptionTime0.88740.234421
SpokenArabicDigits0.10000.97730.9941RDST0.89400.128723
Libras0.06670.88890.9722RIST0.90560.338923
JapaneseVowels0.08380.95950.9946LiteTIME0.91080.208123
Cricket0.08330.98611.00001NN-DTW0.91670.069423
CharacterTrajectories0.06480.98960.9958H-InceptionTime0.93110.044623
TactileTextureRecognition0.05140.99851.0000H-InceptionTime0.94860.016223
ArticularyWordRecognition0.04000.98330.9933Arsenal0.95330.050023

One row per dataset. Dummy is the no-skill floor. Median, best and spread are over the other estimators, so the baseline cannot flatter them. Gain over dummy is best minus dummy, how much skill was found at all; spread is best minus worst, how much the choice of estimator mattered. The two answer different questions, and a single range would conflate them.

4 of 66 datasets gained 0.05 or less over the baseline (shaded amber) and 15 have a best of 0.99 or more (shaded green). Both separate estimators poorly, for opposite reasons. Best is a maximum over many estimators, so it is optimistic by construction: read it as what the archive can currently do on a problem, not as what any one method delivers.