Skip to content

feat(capture): NVLink publication in the v0.5.18 capture sink - #938

Open
maocheng23 wants to merge 4 commits into
mainfrom
maocheng/nvlink-capture-sink
Open

maocheng23 wants to merge 4 commits into
mainfrom
maocheng/nvlink-capture-sink

Conversation

@maocheng23

@maocheng23 maocheng23 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Multi-node NVLink capture transport, part 1 of 2 (server side). Part 2: #939.

Why

On GB200/GB300 NVL72 the capture servers and trainers sit in one NVLink domain, but online capture still moves hidden states through host memory or the NICs. Mooncake's NVLink (MNNVL) transport can only export memory allocated with a fabric handle, and the Mooncake store cannot place objects in such memory, so the store path cannot use it.

Changes

MOONCAKE_PROTOCOL=nvlink switches the v0.5.18 sink from the store to a CaptureArena, memory the capture server owns. The arena is named by placement (where samples wait) rather than transport; currently only online NVLink capture uses it, decided in one place (arena_enabled()).

  • Arena: one fabric allocation from the TransferEngine on the writer GPU, created on the first capture (like the store connection today). Its size is the required SGLANG_SPEC_CAPTURE_ARENA_BYTES. It's a separate variable from the store's MOONCAKE_GLOBAL_SEGMENT_SIZE, whose 32 GiB default was chosen for host memory, so GPU memory is never taken silently.
  • Objects: they keep the store keys ({store_id}/{sample_id}/g{gen}/{name}), one object per tensor. Each result adds "arena": {"session", "control"}, and every feature gains an "address".
  • Lifetime: like hard-pinned store objects, they stay until a client frees them, the counterpart of the store's remove. Clients write a JSON list of keys, one per line, to the control TCP endpoint (port SGLANG_SPEC_CAPTURE_CONTROL_PORT, ephemeral by default).
  • No server thread: frees are one-way. The writer thread, the only thread that touches the arena, reads them before it allocates, so nothing competes with the scheduler for the GIL and the arena needs no locks. The first revision served frees over HTTP from a thread; under scheduler load each request took tens of milliseconds and slowed training.
  • Full arena: a batch that does not fit waits up to 30 s for frees, then fails with a hint to raise the arena size or lower the producer's resident watermark.
  • Hooks: no scheduler hook changes. NVLink always selects device publication, and SGLANG_SPEC_CAPTURE_GPU_PUT=0 is rejected.

Not included: the Kimi K3 patch and the SGLang plugin port (maocheng/sglang-capture-plugin) still need the same sink change.

Tests

  • CPU: arena allocation, reuse and coalescing; waiting for frees; frees split across reads and closed connections; the arena size being required; NVLink forcing device publication.
  • GB300 container (sglang-v0.5.18-cu130-arm64, Mooncake 0.3.13.post1): the sink tests pass, including the CUDA ones (publishing into the arena through a fake engine).
  • GB300, across two trays (c001 → c004), real TransferEngine: this sink published 64 capture-shaped samples (6.47 GiB). A trainer-side MooncakeFeatureStore from part 2 read every byte back correctly (3.5 ms per ~100 MB sample on average), then freed them, and the arena returned to fully free.

🤖 Generated with Claude Code

With MOONCAKE_PROTOCOL=nvlink the sink keeps capture objects in one
fabric-memory arena that the Mooncake TransferEngine allocates on the writer
GPU: the NVLink transport can only export such memory, and the Mooncake store
cannot place objects there. Responses carry each object's device address
plus the TransferEngine session and a control endpoint, where POST /free is
the counterpart of the store's remove. Capture hooks are unchanged; NVLink
always selects device publication.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
maocheng23 and others added 2 commits October 6, 2026 02:25
Serving frees from an HTTP thread competed with the SGLang scheduler for the
GIL: each request took tens of milliseconds under load. Clients now write
JSON key lists to a TCP control connection, and the writer thread, the only
thread that touches the arena, drains them before it allocates. The arena
needs no locks and the HTTP server is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The arena reused MOONCAKE_GLOBAL_SEGMENT_SIZE, whose 32 GiB default was
chosen for the store's host-memory segment; with NVLink it is HBM on the
writer GPU. SGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTES now sets it and is
required, so the arena never silently takes GPU memory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sport

NvlinkArena becomes CaptureArena, the result field "nvlink" becomes "arena"
and SGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTES becomes
SGLANG_SPEC_CAPTURE_ARENA_BYTES. Where a sample waits (store or arena) is
independent of how its bytes move; currently only online NVLink capture
uses the arena, decided in one place (arena_enabled), so an arena over
another transport can be added without renaming again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant