Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

23 Commits

Folders and files

Repository files navigation

Duke Humanoid V2: simulation and training

The reproduction package for the visible-reachable workspace: the workspace study, the two-target reach-and-grasp benchmark, the robot assets, and the checkpoints behind the numbers.

Paper · Project entry point · Hardware · Onboard control stack · See it run · Reproduce the paper

License Python GPU

Visible-reachable volumes of six humanoid platforms

What this repository computes: the visible-reachable workspace of six humanoid platforms. Magenta to orange is visible-reachable, blue is reachable but blind.

VRW asks where a robot can both reach and see. Take the reachable workspace, then keep only the targets that can also be observed from a configuration that reaches them. The joints spent realizing the reach do not count as gaze actuation, which is why a wrist camera on the reaching arm is not an independent view. Running that measure over eight platforms is what the code here does, and it is what chose this robot's camera count, mounting, and articulation.

The robot, its hardware specifications, and the onboard control stack live at duke_humanoid_v2, which is the entry point for the project. This repository is one of its two submodules.

This is a generated export of the research repository; the directory layout matches it, so import paths in the paper's scripts work unchanged.

Contents

Install

Needs an NVIDIA GPU and Python 3.12. From nothing on the machine:

# 1. micromamba, if you have no Python 3.12 yet. Single static binary, no root.
"${SHELL}" <(curl -L micro.mamba.pm/install.sh)
micromamba create -n vrw python=3.12 -y
micromamba activate vrw

# 2. uv, the installer used below
micromamba install -c conda-forge uv -y

# 3. dependencies, into the environment activated above
uv pip install -r requirements.txt

# 4. check the GPU is actually visible; see below if this prints False
python -c "import torch; print(torch.cuda.is_available())"

uv pip resolves the whole set at once instead of one package at a time, which is what keeps a torch + warp + mjlab install from taking a coffee break. It targets whatever environment is active, micromamba or venv alike; pass --python "$(which python)" if you want to be explicit. Plain pip install -r requirements.txt works too and installs the same versions.

If step 4 prints False, the default torch wheel was built against a newer CUDA than your driver. Check yours with nvidia-smi, then reinstall torch from the matching index, for example on a CUDA 12.8 driver:

uv pip install --reinstall torch --index-url https://download.pytorch.org/whl/cu128

--reinstall is the part that matters: without it uv sees torch as already satisfied and the index is ignored. Nothing else needs reinstalling.

The benchmark and generate_workspace_curobo.py additionally need nvidia-curobo, which is not pip-installable from here. Follow its own instructions after the step above. Everything else in this repository runs without it, including both workspace figures, which read the shipped caches.

requirements.txt is deliberately unversioned. This study tracks current mjlab and mujoco-warp, and the numbers here come from the shipped caches and checkpoints rather than from a live solve, so a pin would go stale without protecting a result. Developed against Python 3.12, mjlab 1.6, mujoco 3.11, warp-lang 1.16, torch 2.10, CUDA 12.9, on Ubuntu 22.04. Verified on a fresh clone with mjlab 1.6, mujoco 3.11, warp-lang 1.16 and torch 2.11+cu128 on Ubuntu 22.04.

Headless machines need MUJOCO_GL=egl in front of any command that renders. play is the one exception: it opens a native viewer and needs a real display, so on a headless box use play --headless-eval 300 instead, which steps the policy and prints the reward breakdown.

See it run

Three things you can watch, one command each. Every one of them reads weights or caches already in this checkout: no training, no sweep, no downloads.

1. Watch the locomotion policy. Opens a MuJoCo viewer with the whole-body policy driving the robot:

python mj_envs/run.py play --task HumanoidRmaVelEstArmFlashSacv2ybsk_yaw_s4MixedArmsCam

No checkpoint argument needed. A fresh clone has no runs/, so play falls back to the pinned weight in mj_envs/tasks/visual_manipulation/test/checkpoints/ and prints which one it picked. The other two shipped policies are ...SingleCam and G1RmaVelEstArmFlashSacStudentOnlyg1bsk2.

2. Watch the two-target task. One robot, one scenario, the full mission the benchmark scores:

python mj_envs/tasks/visual_manipulation/test/curobo_reach_verify.py     --robot v2 --scenario bimanual_mixed_close --dynamic --mpc --walk --camera --view

Needs cuRobo. This is the paper's comparison in one command: run it again with --robot v2_fixed to watch the same mission with the cameras welded instead of actuated. Scenarios are left_right_close, left_right_far, front_back_close, front_back_far, bimanual_mixed_close, bimanual_mixed_front_back_close. Drop --view to run headless and print only the verdict. What each flag changes is in mj_envs/tasks/visual_manipulation/.

Two targets left and right, benches close Two targets front and back, benches far
left_right_close. Both targets within reach. front_back_far. Benches at 0.8 m, so the robot locates the targets, then walks.

Each panel above is one camera configuration: v2_single_fixed, v2_single on the top row, v2_fixed, v2 on the bottom.

3. Regenerate the workspace figures. Seconds, from the shipped caches, no GPU sweep:

MUJOCO_GL=egl python mj_envs/asset_zoo/reachability_study/plot_workspace_curobo.py --reach-visible-compare
python mj_envs/asset_zoo/reachability_study/camera_count_ablation.py

Visible-reachable workspace, fixed versus actuated cameras

The measure those figures report, as a volume: the same arms and the same body, differing only in whether the camera joints are free. Blue is what VRW exists to expose, space the arm can reach and the cameras cannot see. Rendered by plot_workspace_curobo.py --vrw-video --robot v2.

What is in here

The four directories worth opening first each have their own README, so you can navigate by browsing rather than by grepping:

Directory What it holds
asset/ Every robot model, with the upstream source and license of each comparison platform.
asset/duke_v2/ This robot: body, head-camera gimbals, end effectors, and how to view them.
mj_envs/asset_zoo/reachability_study/ The visible-reachable workspace computation and both workspace figures.
mj_envs/tasks/visual_manipulation/ The two-target benchmark, its scenarios, and the mission flags.
mj_envs/tasks/humanoid_velocity/ The locomotion policy: what it observes, how it is trained, and why it is built that way.

The rest:

  • asset/create/ is the MJCF/URDF build tooling. The export command behind each shipped cuRobo URDF is recorded in the corresponding curobo/*_robot_cfg.py.
  • asset/toddlerbot_2xm_gripper/ is a sixth supported platform, included as a worked example though no paper figure reports it. run_eta2_platforms.py scores it as its own column, it is the smallest robot here and so the cheapest to regenerate a payload for, and it is the only robot that exercises the coupled-neck branch of gpu_visibility (_apply_coupling, -1/0.909 gear).
  • mj_envs/asset_zoo/ has robot constants, scene objects, and the reachability study.
  • mj_envs/flash_sac/ and mj_envs/ppo/ are the RL training stacks. flash_sac is a distributional SAC and is what every shipped policy was trained with.
  • mj_envs/tasks/visual_manipulation/test/checkpoints/ has the pinned policy weights, with the md5 and training command for each.

Precomputed data:

  • mj_envs/asset_zoo/reachability_study/aggregated_cache/ holds the aggregated VRW payloads that the figure scripts read, which is why Fig. 2 and Fig. 5 regenerate in seconds, not a GPU sweep.

  • mj_envs/asset_zoo/cache/ holds safe-arm-pose and arm-collision-graph payloads for the two robots that are trained and evaluated, humanoid_v21 and unitree_g1. The other platforms' safe-arm-pose caches (~1.4 GB) are left out because the figures read aggregated_cache. Three samplers regenerate them, split by how a pose is certified collision-free:

    robots script
    humanoid_v21, unitree_g1 mj_envs/asset_zoo/generate_safe_arm_poses.py
    booster_t1, fourier_gr3, pal_talos mj_envs/asset_zoo/reachability_study/generate_mjcf_safe_arm_poses.py
    toddlerbot, apptronik_apollo mj_envs/asset_zoo/reachability_study/generate_curobo_safe_arm_poses.py

    Each takes --robot <name> --samples 4096 --seed 42 and writes into mj_envs/asset_zoo/cache/. Run one before generate_workspace_curobo.py --robot <name>.

Long-form method notes are in mj_envs/asset_zoo/reachability_study/readme_reachability.md and mj_envs/tasks/visual_manipulation/readme_visual_manipulation.md.

What this export leaves out

The original CAD (*.step), the high-resolution render meshes (*_high_res.obj and humanoid_v21_high_res.xml), recorded rollout videos (media/), and the platforms that belong to separate projects or that no result uses: Argus, the ballbot, Berkeley Humanoid Lite, OpenArm. Scripts that reference those platforms still import; only those --robot values are unavailable.

Reproducing the paper

Every figure this package covers regenerates from data already in the repository. Only the benchmark needs a sweep.

regenerates in command
Fig. 2, cross-platform VRW comparison seconds, from shipped caches plot_workspace_curobo.py --reach-visible-compare
Fig. 5, camera count x articulation seconds, from shipped caches camera_count_ablation.py
Fig. 6, task keyframe grid seconds, from shipped frames make_keyframe_figure.py
Table IV, 900-trial benchmark hours, multi-GPU dyn_sweep.py
The locomotion policy days run.py train

REPRODUCE.md has the full recipe for each: exact commands, the checkpoint provenance table, the expected benchmark numbers, the regression check, and a record of what was verified on which hardware.

Citation

@misc{xia2026visiblereachableworkspaceperception,
      title={Visible-Reachable Workspace for Perception-Aware Humanoid Design}, 
      author={Boxi Xia and Zijiang Yang and Ryan Shin and Bokuan Li and Eric Wun-Hao Lu and Jacob Lee and Jiaxun Liu and Boyuan Chen},
      year={2026},
      eprint={2609.08905},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.08905}, 
}

License

Apache-2.0, see LICENSE. Third-party robot models under asset/ keep their upstream licenses alongside their meshes.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages