Skip to content

fix: Build on cuda-dl-base devel and ship on inference-runtime - #8994

Merged
mc-nv merged 10 commits into
mainfrom
mchornyi/TRI-1930/inference-image
Oct 5, 2026
Merged

mc-nv merged 10 commits into
mainfrom
mchornyi/TRI-1930/inference-image

Conversation

@mc-nv

@mc-nv mc-nv commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What does the PR do?

  • build.py defaulted base to tritonserver:<ver>-py3-min, which served as
    both the build toolchain and the runtime base. Splits the roles: build on
    cuda-dl-base:*-devel, ship on cuda-dl-base:*-inference-runtime.
  • Always passes a resolved TRITON_BUILD_CONTAINER to ORT/OpenVINO; previously
    only a version went through and their CMake rebuilt a py3-min tag.
  • compose.py min, Dockerfile.sdk and docs follow suit.

Checklist

  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field
  • Added test plan and verified test passes.
  • Verified that the PR passes existing CI.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging ref.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

Check the conventional commit type
box here and add the label to the github PR.

  • build
  • chore
  • ci
  • docs
  • feat
  • fix
  • perf
  • refactor
  • revert
  • style
  • test

Related PRs:

Where should the reviewer start?

  • build.py — create_build_dockerfiles() for the two new defaults and the
    guards on the inference fallback, and onnxruntime_cmake_args() /
    openvino_cmake_args() for the build-container routing.

Test plan

./build.py --dryrun with no --image overrides, then reading the generated
Dockerfiles and cmake_build:

Case Expected
GPU default Dockerfile.buildbase / .cibase on *-devel; Dockerfile on *-inference-runtime
ORT / OpenVINO TRITON_BUILD_CONTAINER = the resolved devel image, no bare version
CPU-only backend build container stays CUDA-capable; Triton base stays ubuntu:24.04
--image=base,X / --image=inference,Y override still wins
--cuda-dl-base-version drives both images
CPU-only + pytorch CUDA-stub donor stage unchanged
vLLM runtime image still nvcr.io/nvidia/vllm:<ver>-py3

All verified. Both image tags confirmed to exist via docker manifest inspect.
pre-commit clean on every changed file.

  • CI Pipeline ID:

Caveats:

  • --upstream-container-version no longer influences the base image, since the
    cuda-dl-base tag couples the train to a CUDA version and cannot be assembled
    from it. The pair is held in DEFAULT_TRITON_VERSION_MAP as
    cuda_dl_base_version, with a --cuda-dl-base-version override.
  • gpu-base (build.py) and gpu-min (compose.py) are deliberately left
    on py3-min. They are not runtime bases — they are scratch donor stages that
    CPU-only PyTorch builds copy CUDA stubs out of, and are never shipped.
    docs/customization_guide/build.md documents that flag and is likewise
    unchanged.
  • fil is untouched. Its ops/Dockerfile expects a full tritonserver:*-py3
    image rather than a minimal base, and -py3 is not being retired, so it
    needs separate assessment.

Background

py3-min is being retired. Because it was doing double duty as toolchain and
runtime base, replacing it required separating those roles rather than swapping
one string. The docs commits also standardise container version placeholders on
YY.MM (previously a mix of xx.yy, yy.mm and YY.MM) and correct the GKE
demo bucket path, which used an underscore the deployer never used.

Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)

  • Resolves: TRI-1930

Switch Dockerfile.sdk BASE_IMAGE from the tritonserver py3-min image to
the cuda-dl-base inference-runtime image, matching the corresponding
tritonserver CI change (TRI-1930).
mc-nv added 7 commits October 2, 2026 03:46
py3-min served as both the build toolchain and the runtime base. build.py
already routed the `base` and `inference` --image keys to those two distinct
roles, but defaulted both to py3-min, so an invocation without explicit
--image overrides built and shipped on the image being retired.

Default `base` to the cuda-dl-base devel image and `inference` to the
inference-runtime image. The `inference` default is guarded on --enable-gpu,
non-RHEL and non-vllm so the CPU-only ubuntu base, the RHEL explicit-base
requirement and the existing vLLM runtime image keep their behaviour.

Routing the ONNX Runtime and OpenVINO build containers is part of the same
change rather than a follow-up, because without it the swap does not reach
them: build.py only passed TRITON_BUILD_CONTAINER when --image=base was given,
and otherwise passed a bare version that each backend's CMakeLists
interpolated back into a py3-min tag. They now always receive the resolved
image. Its default stays CUDA-capable regardless of --enable-gpu, since those
backends need the CUDA toolchain even when Triton itself is built CPU-only --
which is what the old py3-min fallback was quietly providing.

The tag is held as a train plus CUDA version pair because cuda-dl-base
publishes exactly one CUDA version per train, so the two cannot be bumped
independently. fil is deliberately untouched: its Dockerfile expects a full
tritonserver -py3 image rather than a minimal base, so it needs separate
assessment.
The `min` container is the base the composed output image is built from, so it
takes the same inference-runtime image the server runtime now defaults to.

`gpu-min` is deliberately left on py3-min: it is not a runtime base but a
scratch donor stage that CPU-only PyTorch builds copy CUDA stubs out of, and
is never shipped.

The version is read from build.py's map rather than duplicated here, since
compose.py already imports build for container_versions(). Verified the
replacement image carries CUDA_VERSION in its environment -- create_argmap
inspects the min container and fails a GPU compose without it.
build.md and compose.md still named py3-min as the image the buildbase is
built on and the one compose uses as its `min` container. Point both at
cuda-dl-base, and note in compose.md that `min` no longer tracks the branch
the way `full` does.

Two deliberate edits beyond a string swap:

- The "Building Without Docker" section asserted that no Dockerfile is
  available for the base image. That claim was about py3-min and is not
  verified for cuda-dl-base, so it is reworded rather than carried over; the
  practical instruction to install CUDA/cuDNN/TensorRT by hand is unchanged.
- The same paragraph said the non-GPU min image is ubuntu:22.04 while the code
  has used ubuntu:24.04 for some time. Corrected, since leaving a known-wrong
  statement inside a sentence being rewritten is worse than the small scope
  increase.

build.md's `--image=gpu-base,...py3-min` example is intentionally unchanged:
gpu-base is a scratch donor stage for CPU-only PyTorch builds, not a runtime
base, and still uses py3-min.
The docs used three different placeholders for a container version: xx.yy,
yy.mm and YY.MM. xx.yy is the most common but conveys nothing -- it reads as
two arbitrary numbers rather than a year and month.

Standardise on YY.MM, which is already what build.py's own --backend,
--repo-tag, --repoagent and --cache help text uses ("version YY.MM -> branch
rYY.MM"), and which matches the release-train naming of both the tritonserver
and cuda-dl-base containers on NGC. The CUDA placeholder is uppercased to X.Y
to match.

deploy/gke-marketplace-app/README.md is left alone: its xx.yy sits next to a
`gs://triton_sample_models/xx_yy` bucket path, and changing only the prose
would be inconsistent while changing the path risks misstating a real bucket
naming convention that could not be verified here.
… path

Completes the placeholder standardisation started in the previous commit;
these two were held back because the version placeholder sat next to a GCS
bucket path and it was not clear whether that path mirrored a real naming
convention.

It does not. The demo bucket is versioned with a dot, not an underscore --
`gs://triton_sample_models/26.09` in server-deployer/schema.yaml,
chart/triton/values.yaml and trt-engine/README.md -- so the README's
`gs://triton_sample_models/xx_yy` was simply wrong. Corrected to
`gs://triton_sample_models/<YY.MM>` to match what the deployer actually uses.
The previous commit left the bucket path as
`gs://triton_sample_models/YY_MM`, carrying over the underscore from the old
`xx_yy` text. The demo bucket is versioned with a dot: `26.09` in
server-deployer/schema.yaml, chart/triton/values.yaml and
trt-engine/README.md.
@mc-nv
mc-nv force-pushed the mchornyi/TRI-1930/inference-image branch from b0a91fa to 031fdb7 Compare October 2, 2026 04:17
@mc-nv mc-nv added Documentation Improvements or additions to documentation (docs: PRs) chore Maintenance work, no production code change (chore: PRs) and removed refactor Code change that neither fixes a bug nor adds a feature (refactor: PRs) labels Oct 2, 2026
The min container took its version straight from build.py's map, so it ignored
--container-version entirely: composing with --container-version 26.07 paired a
26.07 full container with a 26.09 min container, silently, with no way to
correct it short of --image min,<name>.

Add --cuda-dl-base-version, mirroring build.py's flag, and use it for the min
container. The default is unchanged, so composing without flags behaves exactly
as before.

The pairing cannot be derived automatically: cuda-dl-base couples its train to
a CUDA version (26.07 is cuda13.3, 26.09 is cuda13.4) and there is no table
mapping a Triton container version to the right one. Where a mismatch is
therefore possible -- --container-version given without --cuda-dl-base-version
-- warn instead of pairing them silently.
@mc-nv
mc-nv marked this pull request as ready for review October 2, 2026 17:47
@greptile-apps

greptile-apps Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[High risk] Changes base container image and build configuration.

The PR does not yet appear safe to merge because the SDK build-stage image concern remains outstanding.

Findings

  1. P1 SDK builds on runtime image ▶

Summary

The PR separates the default GPU build and runtime images, routes the resolved build image to ORT and OpenVINO, updates compose and SDK image defaults, and refreshes related documentation. The latest change restores the existing runtime behavior of an explicit --image=base override.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  D[cuda-dl-base devel] --> B[Triton build]
  D --> O[ORT and OpenVINO builds]
  R[cuda-dl-base inference-runtime] --> F[Final Triton image]
  B --> F
  X[Explicit base override] --> B
  X --> F
Loading

Reviews (2) · Last reviewed commit: "fix: Keep --image=base supplying the run..."

Comment thread build.py
Comment thread Dockerfile.sdk
…unset

Adding a default for the `inference` image changed what --image=base,X alone
does: the build stage still used X, but the final image silently became the
cuda-dl-base default, discarding whatever runtime dependencies X was chosen
for. Previously the base override supplied both.

Restore that: when `base` is given and `inference` is not, leave the inference
image unset so the runtime stage falls back to the override, as before. The
default split is unaffected when neither is given, and passing both still wins.

Reported by Greptile on #8994.
@mc-nv
mc-nv requested review from Vinya567, nv-rinig and whoisj October 2, 2026 18:57
@mc-nv
mc-nv merged commit ac2c218 into main Oct 5, 2026
4 checks passed
@mc-nv
mc-nv deleted the mchornyi/TRI-1930/inference-image branch October 5, 2026 19:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

chore Maintenance work, no production code change (chore: PRs) Documentation Improvements or additions to documentation (docs: PRs) fix Bug fix (fix: PRs)

Development

Successfully merging this pull request may close these issues.

2 participants