Skip to content

Add reproducible Cosmos3 Nano and Edge fine-tuning - #277

Draft
ramanathan831 wants to merge 17 commits into
NVIDIA:mainfrom
ramanathan831:dev/ram/cosmos3-tao-reproducibility
Draft

ramanathan831 wants to merge 17 commits into
NVIDIA:mainfrom
ramanathan831:dev/ram/cosmos3-tao-reproducibility

Conversation

@ramanathan831

@ramanathan831 ramanathan831 commented Oct 1, 2026 •

Copy link
Copy Markdown

Summary

Add repository-owned Cosmos3 Nano and Edge fine-tuning support, including:

  • Dataset-neutral video-QA recipes, deterministic data overrides, resumable batching, and worker prefetch.
  • LoRA training, Qwen patch-embedding compatibility, token metrics, epoch checkpointing, and structured lifecycle reporting.
  • Repeated-video validation caching, batched vision attention, and opt-in gradient-spike rollback.
  • Edge attention/mRoPE compatibility, portable FP8 cache decoding, native VLM checkpoint export, and image/source provenance.
  • Non-root container support and the validated pre-commit dependency bootstrap.

Cleanup and scope

  • Replace the recipe re-export and dataset-specific monolith with direct registration of four generic Nano/Edge conversation/reasoning recipes.
  • Move reusable video processing/collation into the existing dataflow roles module and media-grouped sharding into the existing distributor module.
  • Use the Framework-native reasoning dataset and lifecycle callback directly; training has no external toolkit runtime dependency. Their integration is promoted from the dependent runtime PR.
  • Retain three generic examples because their model profiles and training schedules differ. Generalize the launcher and environment names without changing those schedules.
  • Retain the rollback guard, A100 compatibility helper, offline VLM exporter, and regression coverage: these provide independent capabilities.
  • Preserve strict provenance for explicitly identified release images while allowing ordinary builds to report unverified source identity. No CI gates, configurations, or pins changed in this cleanup.

This is structural consolidation, not a blanket file deletion: the current diff has 82 changed files, including 23 new files. Shared-module reuse and moving the native dataset into this parent add edits to existing files; the dependent runtime PR shrinks from 102 changed files to 80.

History and dependent PRs

Supersedes #164, which this account cannot reopen after its maintainer closure. The existing source branch is retained.

The initial 58-commit source history was consolidated to 14 logical commits with an identical final tree. The subsequently requested implementation cleanup adds three focused commits: native video SFT consolidation, generic examples, and default/release image provenance. The current tree therefore intentionally differs from the earlier history-only snapshot.

The dependent PRs are restacked, not replaced:

Validation at e4481fc

  • 103 targeted regression tests passed on this parent, including callbacks, TOML configuration, video recipes, resume, validation caching, loss metrics, attention/mRoPE compatibility, LoRA, export, and Dockerfile/provenance contracts.
  • The four registered recipe configurations match the previous versions. The five relocated classes have equivalent executable ASTs (apart from the distributor's generalized dataset type annotation).
  • All three generic example TOMLs load through the real Framework configuration loader.
  • Both original pre-commit configurations pass locally on all three branches.
  • The restacked dependent branches pass 118 native runtime tests, 557 skills CPU-contract tests under Python 3.11, and strict version stamps.
  • Fresh hosted fork pre-commit and CPU-contract checks passed on the restacked heads. Upstream workflows still require NVIDIA maintainer approval; these fork checks do not establish an upstream CI pass.
  • Previous A100/H100 functional validation is recorded separately. GPU flows and container builds were not rerun for this structural cleanup, and the unchanged H200-specific cases remain unverified without the upstream hardware. The existing Docker Build job is disabled by its upstream if: false condition.

This replacement remains a draft pending upstream CI and review.

ramanathan831 and others added 17 commits October 1, 2026 21:28
Make iopath available to converters, declare the supported Liger API, and
retain optional storage imports. Include the validated uv 0.12.21 lockfile.
Preserve runtime permissions, local package installation, and image
provenance. Exclude generated Python bytecode from the build context.

Signed-off-by: Ramanathan Arunachalam <rarunachalam@nvidia.com>
Restore model export declarations, honor process CPU affinity, and
limit benchmark NCCL environment defaults to CLI startup.
Normalize distributor seeds, preserve contiguous batches across epochs,
and support bounded sequences with spawned preprocessing workers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Use the shared video metadata probe and preserve Qwen timestamps and
modality metadata throughout preprocessing.
Pad the variable sequence dimension before restoring position axes.
Select supported attention implementations while honoring explicit overrides.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Ramanathan Arunachalam <rarunachalam@nvidia.com>
Decode FP8 cache entries on pre-Ada devices while retaining native
FP8 masked-load semantics on supported architectures.
Inject adapters before sharding, restore adapter-only trainability after
materialization, and report logical parameter scope and token loss statistics.
Include Qwen patch-embedding compatibility and the scalar Liger loss API.
Batch uniform vision attention chunks and cache deterministic validation
features without desynchronizing distributed ranks. Honor frozen-weight
precision without FSDP, deterministic FlashAttention, and cuDNN 9.15 support.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ramanathan Arunachalam <rarunachalam@nvidia.com>
Expose checkpoint and validation cadence through TOML, report progress
and terminal failures, and handle padded checkpoint markers. Preserve
ranked loss logs and resolve configuration types without optional Megatron.
Provide local video-QA datasets, deterministic overrides, runtime model
profiles, rank-local decoding, decoded-frame caching, and launch examples.
Retain scoped decoder mocks in the existing recipe regression coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ramanathan Arunachalam <rarunachalam@nvidia.com>
…overy

Restore parameter and optimizer snapshots after gradient spikes. Bound
baseline inflation, coalesce clustered backoff, and gate rate recovery
on demonstrated health. Keep the guard disabled by default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ramanathan Arunachalam <rarunachalam@nvidia.com>
Export native distributed reasoner checkpoints and resolve source
provenance correctly from both ordinary and linked Git worktrees.
Build rumdl 0.1.62 from its pinned release source using an isolated Rust
toolchain, then run the original full hook configuration unchanged.
Move reusable video processing and sharding into existing shared modules, replace the external dataset and status contracts with Framework-native implementations, and register the four dataset-neutral recipes directly. Preserve their configuration values and behavioral tests; add native dataset coverage.
Replace dataset-specific names and environment variables with the shared Cosmos video contract while retaining the Nano, Edge, and task-aware schedules.
Keep explicit release provenance fail-closed, allow ordinary builds without inventing a source identity, and use Framework-native manifest paths. Cover missing, partial, dirty, and complete provenance.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant