Skip to content

feat(capture): NVLink publication in the SGLang capture plugin's sink - #26

Draft
maocheng23 wants to merge 1 commit into
maocheng/sglang-capture-pluginfrom
maocheng/sglang-capture-plugin-nvlink
Draft

maocheng23 wants to merge 1 commit into
maocheng/sglang-capture-pluginfrom
maocheng/sglang-capture-plugin-nvlink

Conversation

@maocheng23

Copy link
Copy Markdown
Owner

Ports the multi-node NVLink (MNNVL) publication path from sgl-project#938 into the SGLang capture plugin (#25). With it, plugin capture servers can feed trainers over NVLink the way patched v0.5.18 servers do. The trainer side is sgl-project#939.

Base: stacked on #25, not on sgl-project#938

Changes

  • sink.py gains NvlinkArena, nvlink_enabled(), the NVLink branch of gpu_put_enabled(), a lazy _arena() and _publish_nvlink(). The code matches the patch sink line for line, with two exceptions: the plugin's constants are public, and the arena-size check lives in nvlink_arena_bytes().
    • With MOONCAKE_PROTOCOL=nvlink, put_samples never connects to the store. It copies each batch into a fabric-memory arena that the Mooncake TransferEngine allocates on the writer GPU, using its own stream.
    • Results gain nvlink: {session, control}, and every feature gains an address.
    • Frees arrive as JSON key lines on the control TCP connection. The writer thread drains them before it allocates, so there is no HTTP thread.
  • capture.py: install() rejects a missing SGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTES when the scheduler starts, like the plugin's flag checks. The patch only fails at the first capture, and then fails every batch. The startup log now says nvlink publication.
    • Device publication needed no change. gpu_put_enabled() is true under nvlink, so CaptureDeviceOutput.copy_to_host keeps device tensors and records the event the writer waits on.
  • tcp/rdma are unchanged. The only edits on the shared path are the store-connection guard and a list of feature metadata.
  • Docs: the plugin README gains an NVLink section with the new env vars, and the patch inventory gains one line.

Tests

tests/test_runtime/test_sglang_capture_plugin_sink.py mirrors sgl-project#938's sink tests:

  • the arena size is required;
  • NVLink forces device publication and rejects SGLANG_SPEC_CAPTURE_GPU_PUT=0;
  • allocation, reuse and coalescing;
  • a full arena waits for frees, then fails;
  • frees split across reads, and closed connections.

Plugin-specific tests:

  • tcp and rdma publish through the store, with no NVLink fields and no arena;
  • NVLink never connects to the store;
  • TestInstall covers the startup checks.

CUDA tests, using a stand-in TransferEngine:

  • publishing into the arena, followed by a free through the control endpoint;
  • the copy is ordered after the ready event through submit_samples;
  • replace semantics: a clashing key releases the objects already placed, and replace=True republishes.

CPU (Mac): 36 plugin tests pass, with the 3 CUDA tests skipped. Across tests/test_runtime and tests/test_config: 832 passed, with the same failure set as the base commit (test_flag_tables_match_the_pinned_sglang_parser and two collection errors). All three come from the local SGLang being main rather than the pinned 0.5.18.

GB300 (node c002 in allocation 3146). Only one node was free, so both halves ran on different GPUs of one tray, not across trays as in sgl-project#938. Image: lmsysorg/sglang:nightly-dev-cu13-20261006-f3fe7534 with the hooks overlay (spec-capture/integration 7567fcfc4e) and Mooncake 0.3.13.post1 (SUPPORT_MNNVL).

median, steps 3–16 nvlink tcp (host publication)
samples/s 16.99 13.94
fetch per sample 0.85 ms 6.07 ms
server publish per batch 1.23 ms 3.10 ms

Step-1 loss is identical (25.387348); later steps differ in the third or fourth digit, as in sgl-project#938's runs. Each arm ran once, in a fresh container. The smoke prompts are short (about 7 MB per sample), so these numbers mostly reflect fixed costs.

Not covered: cross-tray MNNVL with the plugin, since no second node was free. sgl-project#938 validated that path with the same arena code.

🤖 Generated with Claude Code

Port the NvlinkArena from the v0.5.18 capture patch (sgl-project#938
at 9ad9ff8) into the plugin's sink. With MOONCAKE_PROTOCOL=nvlink the sink
copies each batch into one fabric-memory arena that the Mooncake
TransferEngine allocates on the writer GPU, sized by the required
SGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTES. Results gain
"nvlink": {"session", "control"} and a per-feature "address". Clients free
objects by writing JSON key lists to the control TCP connection, which the
writer thread drains before it allocates. NVLink always selects device
publication, and SGLANG_SPEC_CAPTURE_GPU_PUT=0 is rejected.

The scheduler now checks the arena size at startup, like the plugin's flag
checks, instead of failing every capture batch. The tcp and rdma paths are
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant