Skip to content

Rework RHEL/manylinux tutorial for the manylinux-native build (26.08+) - #170

Merged
nv-rinig merged 1 commit into
mainfrom
feat/rhel-manylinux-26.08
Sep 21, 2026
Merged

nv-rinig merged 1 commit into
mainfrom
feat/rhel-manylinux-26.08

Conversation

@nv-rinig

@nv-rinig nv-rinig commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Since 26.08, build.py --target-platform=rhel is manylinux-native: it expects a pypa manylinux base image (its /opt/_internal CPython layout and pipx shared venv) and installs its own build tooling into it, and the released RHEL artifacts moved from manylinux_2_28 to manylinux_2_34. The previous tutorial reconstructed a pyenv-on-Rocky-8 base for 26.05 and no longer builds. This reworks it to build 26.08 and later (including current main) from public sources only.

Changes

  • Dockerfile.base.rhel starts from quay.io/pypa/manylinux_2_34_x86_64 (AlmaLinux 9) and adds the CUDA 13.4 toolkit, cuDNN and TensorRT from the public cuda-rhel9 repo, EPEL and CRB (RHEL 9's CodeReady Builder repo, the successor of PowerTools) for the -devel packages build.py installs, gcc-toolset-13 (the compiler NVIDIA's 26.08 artifacts use; EL9's RapidJSON 1.1.0 does not compile with GCC 14), a Python 3.12 pipx shared venv, and an empty libpython archive that Step 3 passes as Python_LIBRARY so the core's Python bindings link without pulling in pypa's static libpython.
  • Dockerfile.pytorch.rhel installs the public torch wheel into the same image and backfills the files the backend's --image=pytorch extraction expects from NVIDIA's internal wheel image; Dockerfile.pytorch-runtime.rhel completes the served image with torch, NCCL and cuSPARSELt.
  • The README sets the Triton ref and every version in one exported block, with a per-release table; the build command has no version flags (build.py derives versions and repo tags from the checkout); the Step 4 wheel extraction workaround is gone (26.08 stages manylinux wheels itself), and each check shows its actual output.

Testing

  • Full run of the README on an x86-64 host with an A100 (TRITON_REF=main): Steps 1 to 4, wheels tagged manylinux_2_34, CPU serving of the python and onnxruntime models and GPU serving of the pytorch model all return [11, 22, 33, 44].
  • NVIDIA's internal CI job for this tutorial, updated alongside this PR, runs the same steps against server main on every nightly: the build, the CPU and GPU serve checks, and a subset of the PyTorch backend's libtorch tests (those that don't need torchvision). All pass.

build.py's rhel target expects a pypa manylinux base on 26.08+ and main;
the base is manylinux_2_34 plus cuda-rhel9 rpms and gcc-toolset-13.
@nv-rinig nv-rinig changed the title Rework RHEL/manylinux tutorial for 26.08's manylinux-native build path Rework RHEL/manylinux tutorial for the manylinux-native build (26.08+) Sep 18, 2026
@nv-rinig
nv-rinig force-pushed the feat/rhel-manylinux-26.08 branch from a5bc5d0 to ed3deab Compare September 18, 2026 21:56
@nv-rinig
nv-rinig marked this pull request as ready for review September 18, 2026 21:59
@greptile-apps

greptile-apps Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

The PR appears safe to merge after considering two non-blocking reproducibility and validation improvements.

Findings

  1. P2 Wheel Contents Aren't Verified ▶
  2. P2 Runtime Libraries Aren't Pinned ▶

Summary

This PR rewrites the RHEL tutorial around the manylinux-native build introduced for 26.08:

  • Replaces the Rocky Linux 8 base with a pinned manylinux_2_34 image and adds CUDA 13.4, cuDNN, TensorRT, GCC 13, and the expected Python environment.
  • Updates the PyTorch build and runtime images for Torch 2.14 and system-provided NCCL/cuSPARSELt.
  • Simplifies the server build command by deriving versions and repository tags from the checkout.
  • Refreshes the documented validation and inference examples for EL9 artifacts.
  • The artifact validation should inspect wheel contents rather than only filenames, and runtime library versions should be constrained for reproducible release builds.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Pinned manylinux_2_34 image] --> B[Install CUDA, cuDNN, TensorRT, and GCC 13]
  B --> C[triton-manylinux-base]
  A --> D[Install pinned Torch wheel]
  D --> E[PyTorch extraction image]
  C --> F[build.py target-platform rhel]
  E --> F
  F --> G[Manylinux wheels and tritonserver image]
  G --> H[Install Torch, NCCL, and cuSPARSELt]
  H --> I[PyTorch-capable runtime image]
  G --> J[CPU Python and ONNX inference checks]
  I --> K[GPU PyTorch inference check]
Loading

Reviews (1) · Last reviewed commit: "Rework RHEL/manylinux tutorial for the m..."

bash -c 'auditwheel show /b/install/python/tritonserver-*.whl'
# ... is consistent with the following platform tag: "manylinux_2_27_x86_64"
# ... external versioned symbols in system libraries: libc.so.6 (GLIBC_2.2.5 ... 2.27)
find build/install -name '*.whl' # ...-cp312-cp312-manylinux_2_34_x86_64.whl

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Wheel Contents Aren't Verified

This check only validates the wheel filename. A wheel tagged manylinux_2_34 could still depend on symbols newer than glibc 2.34 and pass this step, even though the tutorial says compatibility is verified. Retain an auditwheel show check, or another inspection of the wheel's actual symbol requirements, so the validation covers the artifact contents.

RUN SP=/opt/pyenv_build/versions/3.12/lib/python3.12/site-packages/nvidia; \
printf '%s\n' "$SP/cusparselt/lib" "$SP/nccl/lib" "$SP/nvshmem/lib" \
> /etc/ld.so.conf.d/torch-cuda.conf && ldconfig
RUN dnf install -y libnccl libcusparselt0-cuda-13 && dnf clean all

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Runtime Libraries Aren't Pinned

The Torch wheel is pinned, but NCCL and cuSPARSELt are installed without versions from a live CUDA repository. Rebuilding the same documented release can therefore select different library versions than the known-good run, reducing reproducibility and potentially causing model-load compatibility problems. Pin these RPM versions or use a versioned repository snapshot for each release row.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@nv-rinig
nv-rinig merged commit f44bc0f into main Sep 21, 2026
4 checks passed
@nv-rinig
nv-rinig deleted the feat/rhel-manylinux-26.08 branch September 21, 2026 22:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants