Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

InternW0-Δ

A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

English · 中文
🚀 Quick start · 🔥 Training · 🦾 Evaluation & deployment · 📦 Models

🌐 Project page · 📄 Paper / arXiv · 🤗 Hugging Face (coming soon)

InternW0-delta overview: open data, predictive dynamics, 4D distillation, and robot manipulation.

✨ Overview

InternW0-delta (InternW0-Δ) is a unified world–action model that learns from robot and human demonstrations. It combines pretrained video dynamics, a frozen vision-language model, and training-only 4D supervision to generate robot actions.

  • 🧩 World–Action MoT. A pretrained video expert and an action expert exchange information through masked self-attention, with visual memory and task-conditioned scene semantics.
  • 🔮 Causal Imprint. Compact queries learn action-relevant scene changes from future supervision, making predictive features available to the action expert.
  • 🌌 4D distillation. A frozen Track4World teacher transfers geometry and motion knowledge during training. The teacher and distillation branch are absent from action inference.

InternW0-delta architecture: sparse visual memory, a video expert and an action expert coupled by directed MoT, with Causal Imprint queries and VLM conditioning.

The repository covers data preparation, pretraining, 4D distillation, post-training, evaluation, and real robot deployment.

🚀 Quick start

1. Install

Use Linux, Python 3.11, a CUDA GPU, and FFmpeg. DeepSpeed extensions require a C++ compiler and a matching CUDA toolkit.

git clone https://github.com/InternRobotics/InternW0-Delta.git
cd InternW0-Delta
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -e '.[train]'

2. Prepare models and data

Download the backbones and policy weights following the model guide. For post-training, place the initialization weights at checkpoints/pretrain.pt.

Prepare the selected dataset using the data guide. Run the commands below from the repository root.

3. Post-train a policy

For LIBERO:

export LIBERO_DATA_ROOT=data/libero
python tools/text_cache.py task=libero
NPROC_PER_NODE=8 bash run.sh libero

Continue to LIBERO Plus evaluation, or choose another task below.

🔥 Training

Pretraining

Choose datasets through source_files in dataset.yaml. Training loads each selected source's normalization statistics automatically.

Prepare the action expert initialization, then run:

python tools/path_index.py
python tools/pretrain_text_cache.py
NPROC_PER_NODE=8 bash run.sh pretrain

To continue from pretrained weights, add resume=checkpoints/pretrain.pt model.skip_dit_load_from_pretrain=true; action expert initialization is then unnecessary. For custom data, see normalization.

4D distillation

Select data, cache Track4World teacher features, and continue training with the 4D distillation guide.

Task configurations

Use the task name with python tools/text_cache.py task=<task> and bash run.sh <task>. See post-training data for the required exports.

Task Configuration Evaluation / deployment
LIBERO libero LIBERO Plus
RoboTwin robotwin RoboTwin
RoboDojo robodojo RoboDojo
Ebench ebench, ebench_joint_ee —
Real robots / RTC rtc Deployment

🦾 Evaluation and deployment

Download the task weights, then follow the corresponding environment setup and run instructions:

Guide Use
LIBERO Plus Evaluate LIBERO policies under task perturbations
RoboTwin Post-train on clean demonstrations and evaluate clean → random
RoboDojo Run the Isaac Sim benchmark
Real robot deployment RTC training, robot adapters, and synchronous / asynchronous execution

⚡ Infra

Building on training optimization work from the LiteGen, we provide optional acceleration for InternW0-Δ, including sparse video decoding, VLM execution optimizations, VAE/VLM caching, and MoT compilation.

bash run.sh libero +infra=online
bash run.sh robotwin +infra=compile

See the infra guide for profiles, cache generation, and an A800 performance reference.

🗂️ Data and tools

Guide Use
Data preparation Dataset selection, path indices, text caches, and normalization
Visualization Video, URDF state/action playback, and synchronized curves

The default paths are configurable through environment variables or YAML:

Variable Default
WAM_DATA_ROOT data
WAM_CACHE_ROOT .cache/internw0

📐 80D state and action representation

Robot data shares a masked 80D layout across embodiments.

Dimension map and conventions
Dimensions (zero-based, end-exclusive) State Action
[0,7) / [40,47) Left / right arm joints Joint targets
[7,10) / [47,50) Left / right end-effector position Translation delta
[10,16) / [50,56) End-effector rotation, 6D Rotation-vector delta in the first 3 dimensions; last 3 masked
[16,17) / [56,57) Left / right gripper Gripper command
[17,29) / [57,69) Left / right hand, 12 control channels Hand command
[29,34) Torso joints, up to 5 Absolute joint targets
[34,39) / [69,74) Reserved, masked Reserved, masked
[39,40) Independent lift height Absolute height target
[74,77) Head joints, up to 3 Absolute joint targets
[77,80) Base (x, y, yaw) when available Local (forward, left, yaw) displacement, m / m / rad per source row

Missing channels are padded and masked. Hand channels represent independent finger controls; torso/head ordering and native joint, gripper, and lift units are specified by the source adapter. Base actions are finite displacements, not velocities; unavailable base states are masked. End-effector reference frames and delta conventions follow each adapter. The layout is defined in robot.yaml.

📚 Citation

@misc{miao2026internw0deltaworldactionmodel,
      title={InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data},
      author={Xingyu Miao and Zizun Li and Baole Fang and Kaiwen Song and Tenghui Wang and Hanxue Zhang and Yating Wang and Xudong Li and Yuping He and Xueyuan Wei and Chao Gao and Xijie Yang and Yingxiang Xu and Kerui Ren and Wenqi Guo and Jianjun Zhou and Xinzhe Wang and Weiguang Zhao and Ni Yang and Zetao Cai and Yufei Xue and Hengjie Li and Zeyu He and Yuanzhen Zhou and Rong Fu and Jianyang Zhang and Siwei Cui and Fuxian Huang and Yunsong Zhou and Xing Gao and Yifei Yao and Qiaojun Yu and Kailin Li and Ming Zhou and Mu Huang and Xinyue Li and Wenze Cui and Bingqi Jiang and Xueyue Zhu and Junting Dong and Haoyu Guo and Tao Lu and Mulin Yu and Bowen Zhou and Bin Zhao and Tianfan Xue and Weinan Zhang and Chunhua Shen},
      year={2026},
      eprint={2609.31394},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.31394},
}

📜 License

The code is released under the MIT License. Third-party components retain their respective licenses; see NOTICE.

About

InternW0-Δ : A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

Resources

Stars

51 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages