Repository navigation
[refactor] add hardware-specific installation instruction - #861
FrankLeeeee wants to merge 6 commits into
Conversation
Move torch into cuda/rocm extras routed to their wheel indexes via [tool.uv.sources], fold openai into the base deps, drop requirements-rocm.txt, and add a hardware/version/installer/extras configurator to the docs.
fbd4b61 to
7cf51e9
Compare
|
LGTM, but have we tested this on different hardwares? |
From the Ascend NPU side: I'll run the install flow on Ascend NPU 910C (A3 node) (driver + CANN + torch_npu, then pip install -e . --no-deps with SGLang 0.5.18 NPU build) and post the results here shortly |
|
@maocheng23 , CUDA works well. still testing on AMD and NPU. |
|
I have talked to @Fridge003 and SGLang is deprecating the support for CUDA 12.9 as well, so SpecForge can go directly with CUDA 13. Nonetheless, I will provide a CUDA 12.9 docker image for specforge since sglang still has it. |
|
NPU test results from the Ascend side (following up on my comment above and @maocheng23's question on hardware testing): Tested on a 910C (A3) node, With the container's vendor stack (torch 2.10.0+cpu + torch_npu 2.10.0), training comes up fine: capture servers, rollout workers and the 10-rank trainer all connect, epoch 1 is running and features are going into the mooncake store. I also tried the latest torch_npu release (26.1.1, torch_npu 2.12.0.post2, paired with PyTorch 2.12.0). On CANN 9.0.0 this fails: triton_ascend 3.2.0's Ascend backend compiles |
|
thanks @curnane-lab |
|
LGTM, thanks a lot! |
Fix CUDA CI dependencies and preserve prepared vendor environments. Integrate Intel XPU dependencies and installation commands, declare accelerator conflicts for uv, and remove the separate XPU and ROCm dependency files.
Motivation
SpecForge pinned a single CUDA
torchin its base dependencies and kept a separaterequirements-rocm.txtfor AMD, so installs on other hardware needed manual workarounds and the docs had to spell out each variant by hand. SGLang 0.5.18 only ships CUDA 13 wheels, so the installation guidance also needed to say so clearly. This PR makes the install path hardware-aware inpyproject.tomland lets users pick their setup in the docs and get the right command.Modifications
torchout of the base dependencies intocudaandrocmextras. Thecudaextra also pinssglang-kernelandmooncake-transfer-engine-cuda13, and[tool.uv.sources]routes each extra to its wheel index (cu130, rocm7.2).openaiinto the base dependencies and remove thedataextra; update the data preparation doc and the regenerate script hint accordingly.requirements-rocm.txt. ROCm and Ascend NPU install with--no-depsinside the vendor stack, as the docs already recommended.the installation selector in doc:
Related Issues
Accuracy Test
Benchmark & Profiling
Checklist