Repository navigation
fix: Build on cuda-dl-base devel and ship on inference-runtime - #8994
Merged
Merged
Conversation
Switch Dockerfile.sdk BASE_IMAGE from the tritonserver py3-min image to the cuda-dl-base inference-runtime image, matching the corresponding tritonserver CI change (TRI-1930).
py3-min served as both the build toolchain and the runtime base. build.py already routed the `base` and `inference` --image keys to those two distinct roles, but defaulted both to py3-min, so an invocation without explicit --image overrides built and shipped on the image being retired. Default `base` to the cuda-dl-base devel image and `inference` to the inference-runtime image. The `inference` default is guarded on --enable-gpu, non-RHEL and non-vllm so the CPU-only ubuntu base, the RHEL explicit-base requirement and the existing vLLM runtime image keep their behaviour. Routing the ONNX Runtime and OpenVINO build containers is part of the same change rather than a follow-up, because without it the swap does not reach them: build.py only passed TRITON_BUILD_CONTAINER when --image=base was given, and otherwise passed a bare version that each backend's CMakeLists interpolated back into a py3-min tag. They now always receive the resolved image. Its default stays CUDA-capable regardless of --enable-gpu, since those backends need the CUDA toolchain even when Triton itself is built CPU-only -- which is what the old py3-min fallback was quietly providing. The tag is held as a train plus CUDA version pair because cuda-dl-base publishes exactly one CUDA version per train, so the two cannot be bumped independently. fil is deliberately untouched: its Dockerfile expects a full tritonserver -py3 image rather than a minimal base, so it needs separate assessment.
The `min` container is the base the composed output image is built from, so it takes the same inference-runtime image the server runtime now defaults to. `gpu-min` is deliberately left on py3-min: it is not a runtime base but a scratch donor stage that CPU-only PyTorch builds copy CUDA stubs out of, and is never shipped. The version is read from build.py's map rather than duplicated here, since compose.py already imports build for container_versions(). Verified the replacement image carries CUDA_VERSION in its environment -- create_argmap inspects the min container and fails a GPU compose without it.
build.md and compose.md still named py3-min as the image the buildbase is built on and the one compose uses as its `min` container. Point both at cuda-dl-base, and note in compose.md that `min` no longer tracks the branch the way `full` does. Two deliberate edits beyond a string swap: - The "Building Without Docker" section asserted that no Dockerfile is available for the base image. That claim was about py3-min and is not verified for cuda-dl-base, so it is reworded rather than carried over; the practical instruction to install CUDA/cuDNN/TensorRT by hand is unchanged. - The same paragraph said the non-GPU min image is ubuntu:22.04 while the code has used ubuntu:24.04 for some time. Corrected, since leaving a known-wrong statement inside a sentence being rewritten is worse than the small scope increase. build.md's `--image=gpu-base,...py3-min` example is intentionally unchanged: gpu-base is a scratch donor stage for CPU-only PyTorch builds, not a runtime base, and still uses py3-min.
The docs used three different placeholders for a container version: xx.yy,
yy.mm and YY.MM. xx.yy is the most common but conveys nothing -- it reads as
two arbitrary numbers rather than a year and month.
Standardise on YY.MM, which is already what build.py's own --backend,
--repo-tag, --repoagent and --cache help text uses ("version YY.MM -> branch
rYY.MM"), and which matches the release-train naming of both the tritonserver
and cuda-dl-base containers on NGC. The CUDA placeholder is uppercased to X.Y
to match.
deploy/gke-marketplace-app/README.md is left alone: its xx.yy sits next to a
`gs://triton_sample_models/xx_yy` bucket path, and changing only the prose
would be inconsistent while changing the path risks misstating a real bucket
naming convention that could not be verified here.
… path Completes the placeholder standardisation started in the previous commit; these two were held back because the version placeholder sat next to a GCS bucket path and it was not clear whether that path mirrored a real naming convention. It does not. The demo bucket is versioned with a dot, not an underscore -- `gs://triton_sample_models/26.09` in server-deployer/schema.yaml, chart/triton/values.yaml and trt-engine/README.md -- so the README's `gs://triton_sample_models/xx_yy` was simply wrong. Corrected to `gs://triton_sample_models/<YY.MM>` to match what the deployer actually uses.
The previous commit left the bucket path as `gs://triton_sample_models/YY_MM`, carrying over the underscore from the old `xx_yy` text. The demo bucket is versioned with a dot: `26.09` in server-deployer/schema.yaml, chart/triton/values.yaml and trt-engine/README.md.
mc-nv
force-pushed
the
mchornyi/TRI-1930/inference-image
branch
from
October 2, 2026 04:17
b0a91fa to
031fdb7
Compare
The min container took its version straight from build.py's map, so it ignored --container-version entirely: composing with --container-version 26.07 paired a 26.07 full container with a 26.09 min container, silently, with no way to correct it short of --image min,<name>. Add --cuda-dl-base-version, mirroring build.py's flag, and use it for the min container. The default is unchanged, so composing without flags behaves exactly as before. The pairing cannot be derived automatically: cuda-dl-base couples its train to a CUDA version (26.07 is cuda13.3, 26.09 is cuda13.4) and there is no table mapping a Triton container version to the right one. Where a mismatch is therefore possible -- --container-version given without --cuda-dl-base-version -- warn instead of pairing them silently.
mc-nv
marked this pull request as ready for review
October 2, 2026 17:47
|
…unset Adding a default for the `inference` image changed what --image=base,X alone does: the build stage still used X, but the final image silently became the cuda-dl-base default, discarding whatever runtime dependencies X was chosen for. Previously the base override supplied both. Restore that: when `base` is given and `inference` is not, leave the inference image unset so the runtime stage falls back to the override, as before. The default split is unaffected when neither is given, and passing both still wins. Reported by Greptile on #8994.
whoisj
approved these changes
Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does the PR do?
build.pydefaultedbasetotritonserver:<ver>-py3-min, which served asboth the build toolchain and the runtime base. Splits the roles: build on
cuda-dl-base:*-devel, ship oncuda-dl-base:*-inference-runtime.TRITON_BUILD_CONTAINERto ORT/OpenVINO; previouslyonly a version went through and their CMake rebuilt a py3-min tag.
compose.pymin,Dockerfile.sdkand docs follow suit.Checklist
<commit_type>: <Title>Commit Type:
Check the conventional commit type
box here and add the label to the github PR.
Related PRs:
Where should the reviewer start?
build.py—create_build_dockerfiles()for the two new defaults and theguards on the
inferencefallback, andonnxruntime_cmake_args()/openvino_cmake_args()for the build-container routing.Test plan
./build.py --dryrunwith no--imageoverrides, then reading the generatedDockerfiles and
cmake_build:Dockerfile.buildbase/.cibaseon*-devel;Dockerfileon*-inference-runtimeTRITON_BUILD_CONTAINER= the resolved devel image, no bare versionubuntu:24.04--image=base,X/--image=inference,Y--cuda-dl-base-versionnvcr.io/nvidia/vllm:<ver>-py3All verified. Both image tags confirmed to exist via
docker manifest inspect.pre-commitclean on every changed file.Caveats:
--upstream-container-versionno longer influences the base image, since thecuda-dl-base tag couples the train to a CUDA version and cannot be assembled
from it. The pair is held in
DEFAULT_TRITON_VERSION_MAPascuda_dl_base_version, with a--cuda-dl-base-versionoverride.gpu-base(build.py) andgpu-min(compose.py) are deliberately lefton py3-min. They are not runtime bases — they are scratch donor stages that
CPU-only PyTorch builds copy CUDA stubs out of, and are never shipped.
docs/customization_guide/build.mddocuments that flag and is likewiseunchanged.
filis untouched. Itsops/Dockerfileexpects a fulltritonserver:*-py3image rather than a minimal base, and
-py3is not being retired, so itneeds separate assessment.
Background
py3-min is being retired. Because it was doing double duty as toolchain and
runtime base, replacing it required separating those roles rather than swapping
one string. The docs commits also standardise container version placeholders on
YY.MM(previously a mix ofxx.yy,yy.mmandYY.MM) and correct the GKEdemo bucket path, which used an underscore the deployer never used.
Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)