Add Cosmos3-Edge variant of the LIBERO-10 action-policy SFT recipe - #278
Open
filipemartinsubrobotics wants to merge 2 commits into
Open
filipemartinsubrobotics wants to merge 2 commits into
filipemartinsubrobotics wants to merge 2 commits into
Conversation
action_policy_libero_edge is action_policy_libero_nano on the public nvidia/Cosmos3-Edge base. It differs only in the model config (EDGE_MODEL_CONFIG, the dense Nemotron-2B-Dense-VL backbone) and one extra trainable key, k_norm_und_for_gen, which trains with the generation pathway as in vision_sft_edge. Adds the experiment, a run TOML and launcher mirroring the Nano preset A (HSDP 2x8, global batch 2048), and a docs section with the one-node setting (replicate 1, grad_accum 2) and notes for running outside the Docker image (NPP and an AV1-capable FFmpeg for torchcodec). Validated with a 50-iteration single-node smoke run on 8x RTX PRO 6000 Blackwell: about 79 s/iteration at global batch 2048, loss 15.3 -> 3.6, DCP checkpoint saved. A full 2000-iteration run and the closed-loop libero_10 success rate have not been measured for Edge yet. Signed-off-by: Filipe Martins <293984334+filipemartinsubrobotics@users.noreply.github.com>
Signed-off-by: Filipe Martins <293984334+filipemartinsubrobotics@users.noreply.github.com>
filipemartinsubrobotics
marked this pull request as ready for review
October 2, 2026 10:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
action_policy_libero_edge: the libero_10 action-policy SFT recipe on the publicnvidia/Cosmos3-Edgebase. It mirrorsaction_policy_libero_nanoand differs only inEDGE_MODEL_CONFIG, dense Nemotron-2B-Dense-VL backbone), andk_norm_und_for_gen, as invision_sft_edge.Adds the experiment, a run TOML and launcher mirroring Nano preset A (HSDP 2x8,
global batch 2048), and a docs section covering the one-node setting
(replicate 1, grad_accum 2) and running outside the Docker image.
Validation
~79 s/iteration, loss 15.3 -> 3.6, DCP checkpoint saved.
Happy to run them if useful for review.
Notes for running outside the container
torchcodec needs NPP (
nvidia/cu13/libonLD_LIBRARY_PATH) and an FFmpeg with an AV1decoder (dav1d) for LIBERO's AV1 videos; opencv-python's bundled FFmpeg lacks it.