Skip to content

(fsdp2 dev. perf) Add sequence packing for EAGLE3, DFlash, and DFlash2 - #928

Draft
yushengsu-thu wants to merge 2 commits into
sgl-project:mainfrom
yushengsu-thu:codex/sequence-packing
Draft

yushengsu-thu wants to merge 2 commits into
sgl-project:mainfrom
yushengsu-thu:codex/sequence-packing

Conversation

@yushengsu-thu

@yushengsu-thu yushengsu-thu commented Oct 4, 2026 •

Copy link
Copy Markdown

Motivation

Variable-length samples waste draft-training work on padding. Add opt-in sequence packing for text EAGLE3, DFlash, and DFlash2 with offline features and online disaggregated consumers.

Modifications

  • Enable training.sequence_packing: true with training.attention_backend: flex_attention; the default remains disabled.
  • Pack each existing microbatch with document-local positions, isolated attention and boundary-safe labels/teacher shifts. Preserve sample order, logical batch size, loss normalization, accumulation and acknowledgements.
  • Preserve DFlash/DFlash2 anchor sampling, use sparse document-aware masks, skip invalid backbone proposals, and restore the original objective layout. Online target capture remains per sample.
  • Add focused configuration, collation, model-parity, lifecycle and online-integration tests, plus usage documentation.

USP, multimodal inputs, compact teacher, loss-position trimming and EAGLE3 LK are unsupported. DFlash/DFlash2 LK and D-PACE are supported.

Accuracy Test

Prior Linux/H200 validation (2026-10-02): 204 CPU tests/540 subtests; 57 GPU model/host-sync tests/45 subtests; five live-online/metadata tests/seven subtests. Two-rank FULL_SHARD loss/gradient comparisons passed for all three architectures. Full-model runs verified every trainable tensor received 256 optimizer updates, all five draft layers updated, and weights/losses stayed finite.

Current checks: repository-wide pre-commit and 27 configuration/package-architecture tests (21 subtests) pass. All 17 changed/new production files match the full-model measured-source hashes. The local macOS Torch installation is missing libtorch_cpu.dylib, so model/GPU results above are from the prior completed H200 runs.

BF16 parity is tolerance-based; convergence, final checkpoint quality and speculative-serving acceptance were not evaluated.

Benchmark & Profiling

Complete Qwen3-4B target (36 layers), five-layer drafts, two H200 GPUs, 1,024 ShareGPT conversations, 256 optimizer steps, four runs per mode. Medians:

Model Training step: padded → packed Completion: padded → packed Completion speedup
DFlash 452.79 → 433.64 ms 145.526 → 140.378 s 1.0367×
DFlash2 426.98 → 405.76 ms 137.320 → 132.905 s 1.0332×

Completion includes live target capture, transport, training, acknowledgements and checkpoints at steps 128/256. It excludes preparation, startup and a separate corpus warmup. Training-step diagnostics cover the first 250 steps (1.0442×/1.0523× faster). Gains are workload-specific; the full-vocabulary objective still uses the restored proposal layout.

Related Issues

None.

Checklist

  • Format code with repository pre-commit hooks.
  • Add unit and integration tests.
  • Update usage documentation.
  • Provide training performance and numerical-parity results with validation limits.

Copilot AI balanced review requested due to automatic review settings October 4, 2026 10:34

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@yushengsu-thu
yushengsu-thu marked this pull request as draft October 4, 2026 11:08
@yushengsu-thu yushengsu-thu changed the title [Feature] Add sequence packing for EAGLE3, DFlash, and DFlash2 (fsdp2 dev. perf) Add sequence packing for EAGLE3, DFlash, and DFlash2 Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants