Skip to content

(fsdp2 dev.) bench: FSDP1 vs FSDP2 backend benchmark harness (step-level and end-to-end) - #926

Closed
yushengsu-thu wants to merge 5 commits into
sgl-project:mainfrom
yushengsu-thu:bench/fsdp-backend-compare
Closed

yushengsu-thu wants to merge 5 commits into
sgl-project:mainfrom
yushengsu-thu:bench/fsdp-backend-compare

Conversation

@yushengsu-thu

@yushengsu-thu yushengsu-thu commented Oct 4, 2026 •

Copy link
Copy Markdown

Stacked on #915 (codex/fsdp2-backend): the diff below also shows #915's commits until it merges. Increment-only view for review: yushengsu-thu#5

Conclusion (what the harness measured on 4 × B300, torch 2.13.0+cu130)

Motivation

A repeatable way to answer "what does a backend change buy at production shapes" without a dataset or model download, at two levels: the trainer step in isolation and the full training loop. The numbers in #915's benchmark comments and in #922–#925 come from this harness.

What it does

  • benchmarks/fsdp_backend/bench_fsdp_backends.py (torchrun) builds the real composite training models at production shapes with random weights — EAGLE3 (configs/qwen3-8b-eagle3.json), DFlash2 (configs/qwen3-8b-dflash.json + DFlash2 conv/selector keys), DSpark (configs/qwen3-8b-dspark.json) — feeds synthetic TrainBatches through the real TrainerCore / strategy / create_training_backend seam (EAGLE3 through the offline reader/normalizer/collator) and records per rank: optimizer-step and micro-step time, peak allocated/reserved memory, memory after wrap, host synchronizations and NCCL collective counts for one profiled step, loss and grad norm for parity, and optionally checkpoint save timings (--measure-checkpoint). --label names a variant; --compile-blocks, --fp8-linear, --shard-frozen-tables request the option PRs' BackendOptions when the checkout defines them.
  • benchmarks/fsdp_backend/e2e_offline.py (torchrun) is the end-to-end counterpart: specforge.launch.build_offline_runtime → Trainer.fit() on synthetic feature files on disk (offline reader, FeatureDataLoader workers, TrainerController acks/logging, interval and final checkpoints), reporting steady-state throughput from the controller's own log callbacks plus total wall time — so a kernel-level win that is eaten elsewhere in the loop shows up.
  • e2e_offline.py --variable-mask zeroes a random 20–80% prefix of every synthetic sample so the valid-anchor count changes per micro-batch, as it does with real conversations; --variable-length gives every sample a random length (25–100% of --seq-len) so pad-to-longest batches change shape; --static-shapes turns on (fsdp2 dev. perf) Add training.compile_blocks: per-block torch.compile before fully_shard #922's training.static_shapes, --shape-buckets 512,1024,1536 (fsdp2 dev.) - follow-up : Add training.static_shape_buckets: a few static lengths for compiled blocks #929's buckets; --compile-dynamic requests dynamic-shape compilation when the checkout supports it. These reproduce offline what the online runs showed for per-block compilation (see "What the online runs taught" below).
  • summarize.py / summarize_e2e.py render one column per label with absolute values and deltas against fsdp2; run_matrix.sh runs the FSDP1/FSDP2 matrix on one node. The online (disaggregated) runs below used the production specforge train managed-local path with real prompts; their driver scripts are not part of this PR.

Features are random tensors in the offline runs: losses are a parity check between backends, not a training signal.

End-to-end comparison — offline (Trainer.fit)

Offline e2e = build_offline_runtime → Trainer.fit(): offline reader, FeatureDataLoader with 4 workers, TrainerController acks and logging, final checkpoint; 256 synthetic samples × 6 epochs = 48 optimizer steps, seq 2048, batch 2 × accum 4, DP=4 on 4 × B300 (torch 2.13.0+cu130); steady state = last 24 steps from the controller's log callbacks. Step-level = the real TrainerCore step on a resident batch, batch 2 × accum 8, seq 4096. Production-shaped Qwen3-8B drafts with random weights and synthetic features; deltas are per rank.

draft metric fsdp fsdp2 fsdp2 + compile fsdp2 + compile + fp8 fsdp2 + shard_frozen
DFlash2 steady step time 978 ms 995 ms 813 ms (−182 ms) 770 ms (−225 ms) 994 ms (−1 ms)
DFlash2 samples/s (4 GPUs) 32.7 32.2 39.4 41.6 32.2
DFlash2 peak allocated per rank 37650 MB 37652 MB 34400 MB (−3252 MB) 28783 MB (−8869 MB) 38247 MB (+595 MB)
DSpark steady step time 740 ms 726 ms 638 ms (−88 ms) 638 ms (−88 ms) 729 ms (+3 ms)
DSpark samples/s (4 GPUs) 43.3 44.1 50.2 50.2 43.9
DSpark peak allocated per rank 25057 MB 25058 MB 22328 MB (−2730 MB) 20796 MB (−4262 MB) 25651 MB (+593 MB)
EAGLE3 steady step time 796 ms 664 ms (−132 ms vs fsdp) 569 ms (−95 ms) 592 ms (−72 ms) 966 ms (no-op; see #924)
EAGLE3 samples/s (4 GPUs) 40.2 48.2 56.2 54.1 33.1
EAGLE3 peak allocated per rank 22950 MB 21168 MB (−1782 MB vs fsdp) 20384 MB (−784 MB) 19912 MB (−1256 MB) 21168 MB

With save_interval: 8 (#925): sync → async step time 2286 → 1614 ms (DFlash2), 1980 → 1333 ms (DSpark), 1149 → 947 ms (EAGLE3); GPU peak unchanged.

End-to-end comparison — online (disaggregated)

Online e2e = the disaggregated path exactly as specforge train runs it (managed-local): one SGLang 0.5.18 capture server with the spec-capture patch on GPU 0 (Qwen3-8B, mem_fraction_static 0.5), Mooncake store, producers feeding real ShareGPT prompts (6000 conversations, max_length 2048, qwen chat template), and a DP=3 trainer on GPUs 1–3; batch 2 × accum 4, 40 optimizer steps, num_anchors 512; steady state = the controller's perf/* counters (averaged over the ranks that logged them) over the last 24 optimizer steps. samples/s = 24 samples per optimizer step (batch 2 × accum 4 × 3 trainer GPUs) divided by the mean step time of that window; train compute and data wait are per optimizer step. GPU memory is the nvidia-smi maximum over the run for the trainer GPUs.

DFlash2

variant steady samples/s (3 trainer GPUs) train compute per step Δ compute vs fsdp2 data wait per step trainer GPU memory max Δ memory vs fsdp2
fsdp (FSDP1) 22.22 1074 ms +16 ms (+1.5%) 1 ms 55794 MB −288 MB
fsdp2 22.62 1058 ms – 1 ms 56082 MB –
fsdp2 + shard_frozen_tables 22.59 1060 ms +2 ms (+0.2%) 1 ms 57996 MB +1914 MB
fsdp2, save_interval 8 9.29 1076 ms +18 ms (+1.7%) 1 ms 55850 MB −232 MB
fsdp2, save_interval 8, checkpoint_async 13.69 1097 ms +39 ms (+3.7%) 1 ms 55666 MB −416 MB

DSpark

variant steady samples/s (3 trainer GPUs) train compute per step Δ compute vs fsdp2 data wait per step trainer GPU memory max Δ memory vs fsdp2
fsdp (FSDP1) 33.26 720 ms −0 ms (−0.0%) 1 ms 41814 MB +974 MB
fsdp2 33.26 720 ms – 1 ms 40840 MB –
fsdp2 + shard_frozen_tables not run
fsdp2, save_interval 8 not run
fsdp2, save_interval 8, checkpoint_async not run

The first online matrix of this series ran with the options silently ignored: the disaggregated launch path did not forward them (fixed in #922–#925); the tables above are from the runs after the fix. Online memory is nvidia-smi used on the trainer GPUs, i.e. what the caching allocator reserved, a coarse measure; the precise peak-allocated numbers are in the offline table.

Second devbox, with static_shapes (#922 / #923):

End-to-end comparison — online (disaggregated)

Online e2e = the disaggregated path exactly as specforge train runs it (managed-local): one SGLang 0.5.18 capture server with the spec-capture patch on GPU 0 (Qwen3-8B, mem_fraction_static 0.5), Mooncake store, producers feeding real ShareGPT prompts (6000 conversations, max_length 2048, qwen chat template), and a DP=3 trainer on GPUs 1–3; batch 2 × accum 4, 40 optimizer steps, num_anchors 512; steady state = the controller's perf/* counters (averaged over the ranks that logged them) over the last 24 optimizer steps. samples/s = 24 samples per optimizer step (batch 2 × accum 4 × 3 trainer GPUs) divided by the mean step time of that window; train compute and data wait are per optimizer step. GPU memory is the nvidia-smi maximum over the run for the trainer GPUs.

DFlash2

variant steady samples/s (3 trainer GPUs) train compute per step Δ compute vs fsdp2 data wait per step trainer GPU memory max Δ memory vs fsdp2
fsdp2 22.15 1081 ms – 1 ms 53982 MB –
fsdp2 + static_shapes 24.84 963 ms −117 ms (−10.9%) 1 ms 57392 MB +3410 MB
fsdp2 + static_shapes + compile 31.24 765 ms −315 ms (−29.2%) 1 ms 52432 MB −1550 MB
fsdp2 + static_shapes + compile + fp8 31.35 763 ms −318 ms (−29.4%) 1 ms 46542 MB −7440 MB

DSpark

variant steady samples/s (3 trainer GPUs) train compute per step Δ compute vs fsdp2 data wait per step trainer GPU memory max Δ memory vs fsdp2
fsdp2 33.30 719 ms – 1 ms 39652 MB –
fsdp2 + static_shapes 35.88 667 ms −52 ms (−7.2%) 1 ms 40950 MB +1298 MB
fsdp2 + static_shapes + compile 41.99 569 ms −150 ms (−20.9%) 1 ms 38982 MB −670 MB
fsdp2 + static_shapes + compile + fp8 40.75 587 ms −132 ms (−18.3%) 1 ms 35708 MB −3944 MB

Measured on a second 4 × B300 devbox (same type as the first) after the fixes described above; the fsdp2 baseline was re-run there. Online memory is nvidia-smi used on the trainer GPUs, i.e. what the caching allocator reserved, a coarse measure; the precise peak-allocated numbers are in the offline tables.

What the online runs taught

  • The options were not reaching the online trainer. The first online matrix came back identical to plain fsdp2 for every option. The reason was not noise: specforge/training/disaggregated.py builds the offline-disaggregated and online trainers through its own call sites, which never forwarded backend_options / checkpoint_async (only the single-process path via assembly._common_launch_kwargs did). Each option PR now forwards them at both call sites with tests; the online tables above are from runs after that fix. The harness itself was not affected (it calls build_offline_runtime directly), which is exactly why the discrepancy was visible.
  • Per-block compilation needs static shapes, and real data does not have them. With the options actually forwarded, the online compile_blocks runs crashed on torch 2.13 at the first micro-batch whose padded length differed: InductorError: CantSplit: 8*s50 + 65536 not divisible by s50 + 8192 (online batches are padded to the longest sample, so the context length that the block's flex_attention sees became symbolic after Dynamo's first recompile; 8192 = 512 anchors × 16 draft tokens, 8 = KV heads). The fixed-length offline e2e never hit it; with fixed length but variable masks the graph compiles and the step time stays at −0.1% (dynamic-shape kernels, no speed-up). (fsdp2 dev. perf) Add training.compile_blocks: per-block torch.compile before fully_shard #922 therefore adds training.static_shapes (every micro-batch padded to data.max_length, exactly num_anchors anchor slots per sample); the variable-mask table below shows what that buys: the fsdp2 baseline itself gets much faster (no per-shape rebuild of the flex block mask and the anchor tensors) and the compile gain comes back on top.
  • fp8 has two more shape constraints. torchao's float8 GEMM needs the token count of every micro-batch to be a multiple of 16 (RuntimeError: Expected self.size(1) to be divisible by 16, but got self.size(1)=3858, a block's K/V projection over the context positions), and its float8 all-gather breaks on FSDP2's padded uneven shards (setStorage … out of bounds with DP=3 for 4096 rows). (fsdp2 dev. perf) Add training.fp8_linear: torchao Float8Linear with float8 all-gather #923 warns without static_shapes, enforces data.max_length % 16 == 0 with it, and falls back to a bf16 all-gather when the DP size does not divide the weight rows.

End-to-end comparison — offline with variable loss masks (Trainer.fit)

Offline e2e with variable loss masks = the same Trainer.fit() run as above, but every synthetic sample has a random 20–80% prefix masked out of the loss, so the number of valid anchors differs between micro-batches as it does with real conversations (the harness's --variable-mask); 256 samples × 6 epochs, seq 2048, batch 2 × accum 4, DP=4 on 4 × B300, steady state = last 24 steps.

DFlash2

variant steady step time Δ vs fsdp2 samples/s (4 GPUs) peak allocated per rank Δ memory
fsdp2 1750 ms – 18.3 37652 MB –
fsdp2 + static_shapes 1012 ms -738 ms (−42.2%) 31.6 37444 MB -208 MB
fsdp2 + static_shapes + compile_blocks 838 ms -912 ms (−52.1%) 38.2 34103 MB -3549 MB
fsdp2 + static_shapes + compile_blocks + fp8_linear 795 ms -955 ms (−54.6%) 40.2 28783 MB -8868 MB

DSpark

variant steady step time Δ vs fsdp2 samples/s (4 GPUs) peak allocated per rank Δ memory
fsdp2 823 ms – 38.9 25156 MB –
fsdp2 + static_shapes 734 ms -89 ms (−10.8%) 43.6 25058 MB -99 MB
fsdp2 + static_shapes + compile_blocks 641 ms -182 ms (−22.1%) 49.9 22329 MB -2828 MB
fsdp2 + static_shapes + compile_blocks + fp8_linear 648 ms -176 ms (−21.3%) 49.4 20797 MB -4359 MB

At step level (not e2e) FSDP2 also removes the 6 cudaEventSynchronize per optimizer step of FSDP1's limit_all_gathers rate limiter on the per-block path and builds/wraps in half the time (≈25 s vs 41–54 s).

benchmarks/fsdp_backend drives the real TrainerCore / strategy / backend seam
with production-shaped EAGLE3, DFlash2 and DSpark drafts (random weights,
synthetic batches) under torchrun and records step time, peak memory, host
synchronizations and NCCL collective counts per rank; summarize.py renders
the FSDP1/FSDP2 comparison as markdown.
…summary

Runs can be labelled (--label) and request the sibling PRs' BackendOptions
(--compile-blocks, --fp8-linear, --shard-frozen-tables) when the checkout
defines them; --measure-checkpoint times the full-state gather and the
synchronous/asynchronous write; summarize.py renders one column per label
with deltas against the fsdp2 baseline.
e2e_offline.py drives build_offline_runtime -> Trainer.fit() on synthetic
feature files (offline reader, FeatureDataLoader workers, controller acks,
logging, final checkpoint) and reports steady-state throughput from the
controller's log callbacks plus total wall time; summarize_e2e.py renders it.
summarize.py now takes profiler and checkpoint fields from rank 0.
@yushengsu-thu

Copy link
Copy Markdown
Author

Closing per author request; the harness stays on branch bench/fsdp-backend-compare of yushengsu-thu/SpecForge for reproducing the numbers in #915 and #922–#925.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant