Repository navigation
feat(capture): NVLink publication in the SGLang capture plugin's sink - #26
Draft
maocheng23 wants to merge 1 commit into
Draft
maocheng23 wants to merge 1 commit into
maocheng23 wants to merge 1 commit into
Conversation
Port the NvlinkArena from the v0.5.18 capture patch (sgl-project#938 at 9ad9ff8) into the plugin's sink. With MOONCAKE_PROTOCOL=nvlink the sink copies each batch into one fabric-memory arena that the Mooncake TransferEngine allocates on the writer GPU, sized by the required SGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTES. Results gain "nvlink": {"session", "control"} and a per-feature "address". Clients free objects by writing JSON key lists to the control TCP connection, which the writer thread drains before it allocates. NVLink always selects device publication, and SGLANG_SPEC_CAPTURE_GPU_PUT=0 is rejected. The scheduler now checks the arena size at startup, like the plugin's flag checks, instead of failing every capture batch. The tcp and rdma paths are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports the multi-node NVLink (MNNVL) publication path from sgl-project#938 into the SGLang capture plugin (#25). With it, plugin capture servers can feed trainers over NVLink the way patched v0.5.18 servers do. The trainer side is sgl-project#939.
Base: stacked on #25, not on sgl-project#938
sink.pyis a separate copy of the patch's sink, so nothing in feat(capture): NVLink publication in the v0.5.18 capture sink sgl-project/SpecForge#938's patch diff is required.9ad9ff8a. Its third commit (2026-10-06 20:07 UTC) sizes the arena with a requiredSGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTESinstead ofMOONCAKE_GLOBAL_SEGMENT_SIZE, and this port follows it.Changes
sink.pygainsNvlinkArena,nvlink_enabled(), the NVLink branch ofgpu_put_enabled(), a lazy_arena()and_publish_nvlink(). The code matches the patch sink line for line, with two exceptions: the plugin's constants are public, and the arena-size check lives innvlink_arena_bytes().MOONCAKE_PROTOCOL=nvlink,put_samplesnever connects to the store. It copies each batch into a fabric-memory arena that the Mooncake TransferEngine allocates on the writer GPU, using its own stream.nvlink: {session, control}, and every feature gains anaddress.capture.py:install()rejects a missingSGLANG_SPEC_CAPTURE_NVLINK_ARENA_BYTESwhen the scheduler starts, like the plugin's flag checks. The patch only fails at the first capture, and then fails every batch. The startup log now saysnvlink publication.gpu_put_enabled()is true under nvlink, soCaptureDeviceOutput.copy_to_hostkeeps device tensors and records the event the writer waits on.Tests
tests/test_runtime/test_sglang_capture_plugin_sink.pymirrors sgl-project#938's sink tests:SGLANG_SPEC_CAPTURE_GPU_PUT=0;Plugin-specific tests:
TestInstallcovers the startup checks.CUDA tests, using a stand-in TransferEngine:
submit_samples;replace=Truerepublishes.CPU (Mac): 36 plugin tests pass, with the 3 CUDA tests skipped. Across
tests/test_runtimeandtests/test_config: 832 passed, with the same failure set as the base commit (test_flag_tables_match_the_pinned_sglang_parserand two collection errors). All three come from the local SGLang being main rather than the pinned 0.5.18.GB300 (node c002 in allocation 3146). Only one node was free, so both halves ran on different GPUs of one tray, not across trays as in sgl-project#938. Image:
lmsysorg/sglang:nightly-dev-cu13-20261006-f3fe7534with the hooks overlay (spec-capture/integration7567fcfc4e) and Mooncake 0.3.13.post1 (SUPPORT_MNNVL).MC_FORCE_MNNVL=1). The plugin sink ran on GPU 0 throughsubmit_samples(ready_event=...); feat(mooncake): read server captures over multi-node NVLink sgl-project/SpecForge#939'sMooncakeFeatureStore+NvlinkObjectClientread on GPU 2.get()took 1.56 ms per sample on average.aecbaeb1).Step-1 loss is identical (25.387348); later steps differ in the third or fourth digit, as in sgl-project#938's runs. Each arm ran once, in a fresh container. The smoke prompts are short (about 7 MB per sample), so these numbers mostly reflect fixed costs.
Not covered: cross-tray MNNVL with the plugin, since no second node was free. sgl-project#938 validated that path with the same arena code.
🤖 Generated with Claude Code