Skip to content

Add Cosmos3-Edge variant of the LIBERO-10 action-policy SFT recipe - #278

Open
filipemartinsubrobotics wants to merge 2 commits into
NVIDIA:mainfrom
ubrobotics-ai:libero-edge-recipe
Open

filipemartinsubrobotics wants to merge 2 commits into
NVIDIA:mainfrom
ubrobotics-ai:libero-edge-recipe

Conversation

@filipemartinsubrobotics

Copy link
Copy Markdown

What

action_policy_libero_edge: the libero_10 action-policy SFT recipe on the public
nvidia/Cosmos3-Edge base. It mirrors action_policy_libero_nano and differs only in

  • the model config (EDGE_MODEL_CONFIG, dense Nemotron-2B-Dense-VL backbone), and
  • one extra trainable key, k_norm_und_for_gen, as in vision_sft_edge.

Adds the experiment, a run TOML and launcher mirroring Nano preset A (HSDP 2x8,
global batch 2048), and a docs section covering the one-node setting
(replicate 1, grad_accum 2) and running outside the Docker image.

Validation

  • 50-iteration single-node smoke run, 8x RTX PRO 6000 Blackwell 96 GB, global batch 2048:
    ~79 s/iteration, loss 15.3 -> 3.6, DCP checkpoint saved.
  • Not yet measured: a full 2000-iteration run and the closed-loop libero_10 success rate.
    Happy to run them if useful for review.

Notes for running outside the container

torchcodec needs NPP (nvidia/cu13/lib on LD_LIBRARY_PATH) and an FFmpeg with an AV1
decoder (dav1d) for LIBERO's AV1 videos; opencv-python's bundled FFmpeg lacks it.

action_policy_libero_edge is action_policy_libero_nano on the public
nvidia/Cosmos3-Edge base. It differs only in the model config
(EDGE_MODEL_CONFIG, the dense Nemotron-2B-Dense-VL backbone) and one extra
trainable key, k_norm_und_for_gen, which trains with the generation pathway
as in vision_sft_edge.

Adds the experiment, a run TOML and launcher mirroring the Nano preset A
(HSDP 2x8, global batch 2048), and a docs section with the one-node setting
(replicate 1, grad_accum 2) and notes for running outside the Docker image
(NPP and an AV1-capable FFmpeg for torchcodec).

Validated with a 50-iteration single-node smoke run on 8x RTX PRO 6000
Blackwell: about 79 s/iteration at global batch 2048, loss 15.3 -> 3.6,
DCP checkpoint saved. A full 2000-iteration run and the closed-loop libero_10
success rate have not been measured for Edge yet.

Signed-off-by: Filipe Martins <293984334+filipemartinsubrobotics@users.noreply.github.com>
Signed-off-by: Filipe Martins <293984334+filipemartinsubrobotics@users.noreply.github.com>
@filipemartinsubrobotics
filipemartinsubrobotics marked this pull request as ready for review October 2, 2026 10:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant