Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

MotionMaestro

MotionMaestro: Masked Tokenization for Unified Motion Generation

Yun Chen     Munchurl Kim†     Jeonghyeok Do†
Korea Advanced Institute of Science and Technology (KAIST), South Korea
†Co-corresponding authors

arXiv GitHub Repo stars


This repository is the official implementation of "MotionMaestro: Masked Tokenization for Unified Motion Generation".

MotionMaestro demo: one model, nine tasks on a RoMo motion

One model, nine tasks. Orange: condition, blue: generated. Demo settings: keyframes every 32nd frame; the lower body is given for partial completion and zero-shot editing.

▶️ More videos on the project page

📧 News

  • Sep 2026: This repository is created. The code will be released soon.

📖 Abstract

Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.

📊 Results

One Model, Nine Tasks

MotionMaestro supports nine motion generation tasks with a single model

Figure 1. Nine tasks with a single MotionMaestro model. Orange: provided motion conditions; blue: generated motion.

Quantitative Comparison

MotionMaestro achieves the best results on all ten tasks on RoMo and on eight of ten on MotionMillion, with one generator checkpoint per dataset.

Table 1: Unified motion generation on RoMo

Table 1. Unified motion generation on RoMo. Each method is evaluated on its supported tasks. FI and partial completion report MPJPE (mm) over unobserved frames and joints, respectively; the remaining tasks report FID. Lower is better. Bold and underlined values mark the best and second-best results per task. Results for MotionMaestro-5B, scaled up from our default 1.3B model, are provided in the Appendix.


Table 2: Unified motion generation and computational cost on MotionMillion

Table 2. Unified motion generation and computational cost on MotionMillion. Evaluation settings, metrics, and notation follow Table 1. We additionally report computational cost.

More results and videos: project page.

🖼️ Method Overview

Overview of MotionMaestro

Every task is a masking pattern over the motion sequence.

  • Masked motion tokenizer. Task-aligned masked reconstruction learns one latent space for complete motions and partial observations.
  • Decoder refinement. With the encoder frozen, the decoder is fine-tuned on clean motions to improve reconstruction.
  • Conditional flow matching. A flow-matching generator in the frozen latent space takes an observation map; an observation loss encourages outputs to match the given conditions.

🚀 Code Release Plan

The code and pretrained models will be released soon.

  • Inference code
  • Pretrained models
  • Motion tokenizer
  • Training scripts
  • Evaluation scripts

📑 Citation

If you find MotionMaestro useful, please consider citing:

@article{chen2026motionmaestro,
  title={MotionMaestro: Masked Tokenization for Unified Motion Generation},
  author={Chen, Yun and Kim, Munchurl and Do, Jeonghyeok},
  journal={arXiv preprint arXiv:2609.37495},
  year={2026}
}

About

MotionMaestro: Masked Tokenization for Unified Motion Generation

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Contributors