The reproduction package for the visible-reachable workspace: the workspace study, the two-target reach-and-grasp benchmark, the robot assets, and the checkpoints behind the numbers.
Paper · Project entry point · Hardware · Onboard control stack · See it run · Reproduce the paper
VRW asks where a robot can both reach and see. Take the reachable workspace, then keep only the targets that can also be observed from a configuration that reaches them. The joints spent realizing the reach do not count as gaze actuation, which is why a wrist camera on the reaching arm is not an independent view. Running that measure over eight platforms is what the code here does, and it is what chose this robot's camera count, mounting, and articulation.
The robot, its hardware specifications, and the onboard control stack live at duke_humanoid_v2, which is the entry point for the project. This repository is one of its two submodules.
This is a generated export of the research repository; the directory layout matches it, so import paths in the paper's scripts work unchanged.
Needs an NVIDIA GPU and Python 3.12. From nothing on the machine:
# 1. micromamba, if you have no Python 3.12 yet. Single static binary, no root.
"${SHELL}" <(curl -L micro.mamba.pm/install.sh)
micromamba create -n vrw python=3.12 -y
micromamba activate vrw
# 2. uv, the installer used below
micromamba install -c conda-forge uv -y
# 3. dependencies, into the environment activated above
uv pip install -r requirements.txt
# 4. check the GPU is actually visible; see below if this prints False
python -c "import torch; print(torch.cuda.is_available())"uv pip resolves the whole set at once instead of one package at a time, which is what keeps a
torch + warp + mjlab install from taking a coffee break. It targets whatever environment is
active, micromamba or venv alike; pass --python "$(which python)" if you want to be explicit.
Plain pip install -r requirements.txt works too and installs the same versions.
If step 4 prints False, the default torch wheel was built against a newer CUDA than your
driver. Check yours with nvidia-smi, then reinstall torch from the matching index, for example
on a CUDA 12.8 driver:
uv pip install --reinstall torch --index-url https://download.pytorch.org/whl/cu128--reinstall is the part that matters: without it uv sees torch as already satisfied and the
index is ignored. Nothing else needs reinstalling.
The benchmark and generate_workspace_curobo.py additionally need nvidia-curobo, which is not
pip-installable from here. Follow
its own instructions after the step
above. Everything else in this repository runs without it, including both workspace figures, which
read the shipped caches.
requirements.txt is deliberately unversioned. This study tracks current mjlab and
mujoco-warp, and the numbers here come from the shipped caches and checkpoints rather than from
a live solve, so a pin would go stale without protecting a result. Developed against Python 3.12,
mjlab 1.6, mujoco 3.11, warp-lang 1.16, torch 2.10, CUDA 12.9, on Ubuntu 22.04. Verified on
a fresh clone with mjlab 1.6, mujoco 3.11, warp-lang 1.16 and torch 2.11+cu128 on Ubuntu
22.04.
Headless machines need MUJOCO_GL=egl in front of any command that renders. play is the one
exception: it opens a native viewer and needs a real display, so on a headless box use
play --headless-eval 300 instead, which steps the policy and prints the reward breakdown.
Three things you can watch, one command each. Every one of them reads weights or caches already in this checkout: no training, no sweep, no downloads.
1. Watch the locomotion policy. Opens a MuJoCo viewer with the whole-body policy driving the robot:
python mj_envs/run.py play --task HumanoidRmaVelEstArmFlashSacv2ybsk_yaw_s4MixedArmsCamNo checkpoint argument needed. A fresh clone has no runs/, so play falls back to the pinned
weight in mj_envs/tasks/visual_manipulation/test/checkpoints/ and prints which one it picked.
The other two shipped policies are ...SingleCam and G1RmaVelEstArmFlashSacStudentOnlyg1bsk2.
2. Watch the two-target task. One robot, one scenario, the full mission the benchmark scores:
python mj_envs/tasks/visual_manipulation/test/curobo_reach_verify.py --robot v2 --scenario bimanual_mixed_close --dynamic --mpc --walk --camera --viewNeeds cuRobo. This is the paper's comparison in one command: run it again with --robot v2_fixed
to watch the same mission with the cameras welded instead of actuated. Scenarios are
left_right_close, left_right_far, front_back_close, front_back_far,
bimanual_mixed_close, bimanual_mixed_front_back_close. Drop --view to run headless and print
only the verdict. What each flag changes is in
mj_envs/tasks/visual_manipulation/.
![]() |
![]() |
| left_right_close. Both targets within reach. | front_back_far. Benches at 0.8 m, so the robot locates the targets, then walks. |
Each panel above is one camera configuration: v2_single_fixed, v2_single on the top row,
v2_fixed, v2 on the bottom.
3. Regenerate the workspace figures. Seconds, from the shipped caches, no GPU sweep:
MUJOCO_GL=egl python mj_envs/asset_zoo/reachability_study/plot_workspace_curobo.py --reach-visible-compare
python mj_envs/asset_zoo/reachability_study/camera_count_ablation.pyThe measure those figures report, as a volume: the same arms and the same body, differing only in
whether the camera joints are free. Blue is what VRW exists to expose, space the arm can reach and
the cameras cannot see. Rendered by plot_workspace_curobo.py --vrw-video --robot v2.
The four directories worth opening first each have their own README, so you can navigate by browsing rather than by grepping:
| Directory | What it holds |
|---|---|
asset/ |
Every robot model, with the upstream source and license of each comparison platform. |
asset/duke_v2/ |
This robot: body, head-camera gimbals, end effectors, and how to view them. |
mj_envs/asset_zoo/reachability_study/ |
The visible-reachable workspace computation and both workspace figures. |
mj_envs/tasks/visual_manipulation/ |
The two-target benchmark, its scenarios, and the mission flags. |
mj_envs/tasks/humanoid_velocity/ |
The locomotion policy: what it observes, how it is trained, and why it is built that way. |
The rest:
asset/create/is the MJCF/URDF build tooling. The export command behind each shipped cuRobo URDF is recorded in the correspondingcurobo/*_robot_cfg.py.asset/toddlerbot_2xm_gripper/is a sixth supported platform, included as a worked example though no paper figure reports it.run_eta2_platforms.pyscores it as its own column, it is the smallest robot here and so the cheapest to regenerate a payload for, and it is the only robot that exercises the coupled-neck branch ofgpu_visibility(_apply_coupling, -1/0.909 gear).mj_envs/asset_zoo/has robot constants, scene objects, and the reachability study.mj_envs/flash_sac/andmj_envs/ppo/are the RL training stacks.flash_sacis a distributional SAC and is what every shipped policy was trained with.mj_envs/tasks/visual_manipulation/test/checkpoints/has the pinned policy weights, with the md5 and training command for each.
Precomputed data:
-
mj_envs/asset_zoo/reachability_study/aggregated_cache/holds the aggregated VRW payloads that the figure scripts read, which is why Fig. 2 and Fig. 5 regenerate in seconds, not a GPU sweep. -
mj_envs/asset_zoo/cache/holds safe-arm-pose and arm-collision-graph payloads for the two robots that are trained and evaluated,humanoid_v21andunitree_g1. The other platforms' safe-arm-pose caches (~1.4 GB) are left out because the figures readaggregated_cache. Three samplers regenerate them, split by how a pose is certified collision-free:robots script humanoid_v21,unitree_g1mj_envs/asset_zoo/generate_safe_arm_poses.pybooster_t1,fourier_gr3,pal_talosmj_envs/asset_zoo/reachability_study/generate_mjcf_safe_arm_poses.pytoddlerbot,apptronik_apollomj_envs/asset_zoo/reachability_study/generate_curobo_safe_arm_poses.pyEach takes
--robot <name> --samples 4096 --seed 42and writes intomj_envs/asset_zoo/cache/. Run one beforegenerate_workspace_curobo.py --robot <name>.
Long-form method notes are in mj_envs/asset_zoo/reachability_study/readme_reachability.md and
mj_envs/tasks/visual_manipulation/readme_visual_manipulation.md.
The original CAD (*.step), the high-resolution render meshes (*_high_res.obj and
humanoid_v21_high_res.xml), recorded rollout videos (media/), and the platforms that belong to
separate projects or that no result uses: Argus, the ballbot, Berkeley Humanoid Lite, OpenArm.
Scripts that reference those platforms still import; only those --robot values are unavailable.
Every figure this package covers regenerates from data already in the repository. Only the benchmark needs a sweep.
| regenerates in | command | |
|---|---|---|
| Fig. 2, cross-platform VRW comparison | seconds, from shipped caches | plot_workspace_curobo.py --reach-visible-compare |
| Fig. 5, camera count x articulation | seconds, from shipped caches | camera_count_ablation.py |
| Fig. 6, task keyframe grid | seconds, from shipped frames | make_keyframe_figure.py |
| Table IV, 900-trial benchmark | hours, multi-GPU | dyn_sweep.py |
| The locomotion policy | days | run.py train |
REPRODUCE.md has the full recipe for each: exact commands, the checkpoint provenance table, the expected benchmark numbers, the regression check, and a record of what was verified on which hardware.
@misc{xia2026visiblereachableworkspaceperception,
title={Visible-Reachable Workspace for Perception-Aware Humanoid Design},
author={Boxi Xia and Zijiang Yang and Ryan Shin and Bokuan Li and Eric Wun-Hao Lu and Jacob Lee and Jiaxun Liu and Boyuan Chen},
year={2026},
eprint={2609.08905},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.08905},
}Apache-2.0, see LICENSE. Third-party robot models under asset/ keep their upstream
licenses alongside their meshes.



