Skip to content

[refactor] add hardware-specific installation instruction - #861

Draft
FrankLeeeee wants to merge 6 commits into
sgl-project:mainfrom
FrankLeeeee:refactor-deps
Draft

FrankLeeeee wants to merge 6 commits into
sgl-project:mainfrom
FrankLeeeee:refactor-deps

Conversation

@FrankLeeeee

@FrankLeeeee FrankLeeeee commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

SpecForge pinned a single CUDA torch in its base dependencies and kept a separate requirements-rocm.txt for AMD, so installs on other hardware needed manual workarounds and the docs had to spell out each variant by hand. SGLang 0.5.18 only ships CUDA 13 wheels, so the installation guidance also needed to say so clearly. This PR makes the install path hardware-aware in pyproject.toml and lets users pick their setup in the docs and get the right command.

Modifications

  • Move torch out of the base dependencies into cuda and rocm extras. The cuda extra also pins sglang-kernel and mooncake-transfer-engine-cuda13, and [tool.uv.sources] routes each extra to its wheel index (cu130, rocm7.2).
  • Fold openai into the base dependencies and remove the data extra; update the data preparation doc and the regenerate script hint accordingly.
  • Delete requirements-rocm.txt. ROCm and Ascend NPU install with --no-deps inside the vendor stack, as the docs already recommended.
  • Rewrite the installation guide around an interactive configurator (hardware, version, installer, extras) that generates the install command, with per-hardware notes below it. The landing-page selector is extended and reused for this.

the installation selector in doc:

Screenshot 2026-09-09 at 5 04 19 PM

Related Issues

Accuracy Test

Benchmark & Profiling

Checklist

Move torch into cuda/rocm extras routed to their wheel indexes via
[tool.uv.sources], fold openai into the base deps, drop requirements-rocm.txt,
and add a hardware/version/installer/extras configurator to the docs.
@FrankLeeeee
FrankLeeeee marked this pull request as draft September 8, 2026 09:45
@maocheng23

Copy link
Copy Markdown
Collaborator

LGTM, but have we tested this on different hardwares?

@curnane-lab

Copy link
Copy Markdown
Collaborator

LGTM, but have we tested this on different hardwares?

From the Ascend NPU side: I'll run the install flow on Ascend NPU 910C (A3 node) (driver + CANN + torch_npu, then pip install -e . --no-deps with SGLang 0.5.18 NPU build) and post the results here shortly

@FrankLeeeee

Copy link
Copy Markdown
Collaborator Author

@maocheng23 , CUDA works well. still testing on AMD and NPU.

@FrankLeeeee

FrankLeeeee commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator Author

I have talked to @Fridge003 and SGLang is deprecating the support for CUDA 12.9 as well, so SpecForge can go directly with CUDA 13. Nonetheless, I will provide a CUDA 12.9 docker image for specforge since sglang still has it.

@curnane-lab

curnane-lab commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

NPU test results from the Ascend side (following up on my comment above and @maocheng23's question on hardware testing):

Tested on a 910C (A3) node, sglang_v0.5.18 NPU container, 16 NPUs (6 capture + 10 trainer), config qwen3.5-4b-dflash-disaggregated-npu.yaml.

With the container's vendor stack (torch 2.10.0+cpu + torch_npu 2.10.0), training comes up fine: capture servers, rollout workers and the 10-rank trainer all connect, epoch 1 is running and features are going into the mooncake store.

I also tried the latest torch_npu release (26.1.1, torch_npu 2.12.0.post2, paired with PyTorch 2.12.0). On CANN 9.0.0 this fails: triton_ascend 3.2.0's Ascend backend compiles npu_utils.cpp against RT_LIMIT_TYPE_SIMT_WARP_STACK_SIZE, which CANN 9.0.0's base.h has renamed to RT_LIMIT_TYPE_SIMT_DVG_WARP_STACK_SIZE, and the compile error takes down all 6 capture servers. Note this combo officially pairs with CANN 9.1.X, so I'll upgrade CANN and re-run — will report back here.

@FrankLeeeee

Copy link
Copy Markdown
Collaborator Author

thanks @curnane-lab

@maocheng23

Copy link
Copy Markdown
Collaborator

LGTM, thanks a lot!
But this PR is still a draft, shall we merge it after combining #843 ?

Fix CUDA CI dependencies and preserve prepared vendor environments. Integrate Intel XPU dependencies and installation commands, declare accelerator conflicts for uv, and remove the separate XPU and ROCm dependency files.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants