Skip to content

[ENH] Add global scalers for time series pre-treatment - #3794

Open
melpiro wants to merge 10 commits into
mainfrom
mp/global-scaler
Open

melpiro wants to merge 10 commits into
mainfrom
mp/global-scaler

Conversation

@melpiro

@melpiro melpiro commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

Reference Issues/PRs

#3792

What does this implement/fix? Explain your changes.

Implement global versions of Centerer, MinMaxScaler, and Normalizer :

  • GlobalCenterer: Transform a collection of time series so that the mean of the collection is equal to zero
  • GlobalMinMaxScaler : Transform a collection of time series so that the minimal value and the maximal value of the collection are respectively zero and one.
  • GlobalNormalizer: Transform a collection of time series so that the mean and the standard deviation of the collection are respectively zero and one.

Does your contribution introduce a new dependency? If yes, which one?

No

Any other comments?

PR checklist

For all contributions
  • I've added myself to the list of contributors. Alternatively, you can use the @all-contributors bot to do this for you after the PR has been merged.
  • The PR title starts with either [ENH], [MNT], [DOC], [BUG], [REF], [DEP] or [GOV] indicating whether the PR topic is related to enhancement, maintenance, documentation, bugs, refactoring, deprecation or governance.
For new estimators and functions
  • I've added the estimator/function to the online API documentation.
  • (OPTIONAL) I've added myself as a __maintainer__ at the top of relevant files and want to be contacted regarding its maintenance. Unmaintained files may be removed. This is for the full file, and you should not add yourself if you are just making minor changes or do not want to help maintain its contents.
For developers with write access
  • (OPTIONAL) I've updated aeon's CODEOWNERS to receive notifications about future changes to these files.

@aeon-actions-bot aeon-actions-bot Bot added enhancement New feature, improvement request or other non-bug code enhancement transformations Transformations package labels Sep 4, 2026
@aeon-actions-bot

Copy link
Copy Markdown
Contributor

Thank you for contributing to aeon

I have added the following labels to this PR based on the title: [ enhancement ].
I have added the following labels to this PR based on the changes made: [ transformations ]. Feel free to change these if they do not properly represent the PR.

The Checks tab will show the status of our automated tests. You can click on individual test runs in the tab or "Details" in the panel below to see more information if there is a failure.

If our pre-commit code quality check fails, please run pre-commit locally and push the fixes to your PR branch.

Don't hesitate to ask questions on the aeon Discord channel if you have any.

PR CI actions

These checkboxes will add labels to enable or disable CI functionality for this PR. This may not take effect immediately, and a new commit may be required to run the new configuration.

  • Run pre-commit checks for all files
  • Run mypy typecheck tests
  • Run all pytest tests and configurations
  • Run all notebook example tests
  • Run numba-disabled codecov tests
  • Disable numba cache loading
  • Regenerate expected results for testing
  • Push an empty commit to re-run CI checks

@melpiro melpiro assigned melpiro and unassigned hadifawaz1999 Sep 7, 2026

@TonyBagnall TonyBagnall left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this, I agree its a useful thing to add in. I have one issue with the unequal-length normalisation: this calculates the standard deviation of the individual series' standard deviations, rather than the pooled standard deviation of all observations.

For example, if X = [np.array([[0, 2]]), np.array([[10, 12]])], both series have std 1, so the current calculation returns zero, whereas the pooled std is approximately 5.099.

Im not sure how this should work, but I think, given its global, a pooled mean and variance per channel would be right? Happy to discuss

@melpiro

melpiro commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Hello, thank you for taking the time to review my code.
Yes, you are right; the previous implementation for unequal-length time series was wrong and has now been fixed.

I took the opportunity to improve the generalisation of global scalers by adding an axis parameter, allowing to use the scaler on any numpy.array like :

  • [sample, channel, time] axis = 2 | default with 3D
  • [sample, time, channel] axis = 1
  • list[(channels, time)] unequal lenght collection
  • [sample, time] default case with 2D (univariate collection)
  • [channels, time] axis = 1 (multivariate serie) useful for forecasting
  • [time, channels] axis = 0 (multivariate serie) useful for forecasting

To go further, the code is in reality, able to run on any NumPy array dimensionality. Even if it makes poor sense for time series, I found interesting to test my scalers with 1D numpy array and 4D NumPy array to check that they are able to generalise correctly. In theory, it could even be used with any higher dimensionality.

From there, very weird scenarios with the following input shape could be imagined :

  • 1D array of shape [time] (univariate serie)
  • Collections of videos: [sample, x, y, time] axis=3 | default with 4D
  • Even with two time axes?? [sample, channels, time1, time2] axis=(2, 3)

In short, it turns out to be useless but shows robustness and was fun to see that it works on unusual inputs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature, improvement request or other non-bug code enhancement transformations Transformations package

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants