llama.cpp b11080 → b11211, CUDA 13.4, ROCm 10 (TheRock), sccache on every Windows Ninja job - #456
Merged
Merged
Conversation
First step of the b11080 -> b11209 series. 23 upstream commits, all internal to this project's surface: CUDA/Metal/SYCL/OpenCL/hexagon kernels, cpp-httplib 0.57.1, a Muse Glimmer tool-call parser fix and MiMo-V2.6 template detection. No priority-list header moves in this range, and all eight local patches apply unchanged (replayed against every tag, not just the endpoints). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
upstream #28690 lets --host take a comma-separated list and binds every address. It removed server_http_context::thread and ::listening_address in favour of join() and listening_addresses. patches/0007 still applied cleanly at b11104 -- its hunks are nowhere near the changed lines -- but two of its own + lines in llama_server_attach named the removed members, so the native build would have failed. It now logs every listening address and blocks in ctx_http.join(), as upstream's llama_server() does. NativeServer gains getHosts() (the parsed list, never empty); getHost() returns its first element instead of the raw "a,b" string. Two new NativeServerSmokeTest cases pin the split/trim and the fallback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
56 upstream commits, compatible from this project's side: backend kernels, the llama.cpp 0.5.0 version bump, router fixes, new model support (Ling 3.0 VL, Gemma 4 DSpark draft) and two OAI-layer features -- input_image as a Responses function_call_output and video_url as an alias of input_video. Only a private member of server-context.h moves; all eight patches apply at every tag. The two OAI features are recorded in the history row as Java-API candidates (ContentPart has no video part; ResponsesApiSupport flattens function_call_output to text), not changed here. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
upstream #24669 adds the extended batch API (llama_batch_ext + llama_process) next to the classic llama_batch/llama_decode, which is unchanged. Purely additive for this project: nothing in src/main/cpp calls the decode or batch helpers directly. Kept as its own step because it is the one new public llama.h surface in the b11080 -> b11209 series. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Last step of the b11080 -> b11209 series. 46 upstream commits, compatible from this project's side: backend kernels, a GGUF-driven W4A4 precision policy with no public API, shared Unicode helpers in common.h, cpp-httplib 0.58.0, cleanup after a failed state restore, a grammar token_id fix, and a revert of the --fit context-length change. All eight patches are still required. Verified at b11209 from a fresh configure: patches applied (8, stamp matches), Release build with 0 warnings, ctest 559/559, 40 Java_* exports, and mvn clean verify 1774 run / 0 failures with NativeLibraryLoadSmokeTest 4/4 confirming the pin against the linked build-info. The CHANGELOG entry covers the whole series. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
AMD now builds and releases ROCm through TheRock. Both rocm-* jobs install its Python wheels (rocm[libraries,devel]==10.0.0 from stable.repo.amd.com/rocm/whl-next) the way upstream llama.cpp's own ubuntu-rocm / windows-rocm release jobs do at b11209, replacing the repo.radeon.com 6.3.4 apt repo on Linux and the HIP SDK 26.Q1 installer on Windows. Paths are read back with rocm-sdk path. Linux now uses CMake's native HIP language with ROCm's clang as the HIP compiler (upstream's form); Windows uses the TheRock clang under lib\llvm\bin. The GPU target lists are copied from upstream: newer architectures are added, gfx900/gfx906 are dropped because llama.cpp no longer ships them and TheRock never marks them release-ready. Not validated locally: stable.repo.amd.com is unreachable from the sandbox, so the first CI run is the proof. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
CUDA 13.3 -> 13.4, matching upstream llama.cpp. Linux installs cuda-toolkit-13-4 from NVIDIA's rhel8 repo. Windows drops Jimver/cuda-toolkit, which has no 13.4 in any release or on master, and assembles the toolkit from NVIDIA's per-component redist archives with upstream's windows-setup-cuda component list. Classifiers stay cuda13-*. ROCm Linux keeps gfx900/gfx906 on top of upstream's target list: TheRock 10 still builds them. They stay only as long as they build without patches and do not hold back a newer ROCm. ggml-org/free-disk-space now runs first in the CUDA, ROCm and both SYCL Linux jobs, and replaces the hand-rolled toolchain removal in the two emulator jobs (with android, large-packages, tool-cache and swap kept). The dockcross wrappers' help text named the tags they were first generated from; it now names the pinned 20260712-79e54f9 image. Not validated locally: the NVIDIA and AMD package servers are unreachable from the sandbox, so the first CI run is the proof. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
The ROCm target lists were upstream llama.cpp's, which differed between Linux and Windows partly by accident. They are now every target TheRock builds per OS (its SUPPORTED_GPUS.md): Linux 27 targets, Windows 23. That adds gfx90c and gfx1153 on Linux and gfx900/gfx906/gfx90c on Windows. The two lists now differ only by the Instinct parts (gfx908/gfx90a/gfx942/gfx950), which ROCm supports on Linux alone. The extras upstream omits are build-passing only in TheRock; they stay as long as they build without patches and do not block a newer ROCm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Five Windows build jobs never installed sccache: arm64, arm64-OpenCL, ROCm, SYCL and OpenVINO. They now get the same USE_CACHE/SCCACHE_WEBDAV env and install step as the other Ninja jobs. The two arm64 jobs use the native aarch64-pc-windows-msvc release, which exists (the comment that said only x86_64 was available is gone). Only the two MSVC-classifier jobs stay uncached: the Visual Studio generator ignores launchers. build.bat's probe only proves sccache can wrap cl.exe, while these jobs compile with clang-cl, ROCm's clang or icx. So a configure or build that fails with sccache as the launcher is now retried once from a clean build dir without it. The retry is unconditional because cmd cannot tee output to match an error signature; a real compile error costs one extra uncached attempt, a cache incompatibility ends green. Every configure also passes -DGGML_CCACHE=OFF. Without it ggml self-enables any sccache on PATH whenever no launcher is set, i.e. in the probe-failed and retry cases, and the uncached build went through sccache after all. No goto/labels: the file is checked out with LF line endings, where cmd's label search is unreliable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Two upstream commits: #29440 (RPC waits on the RDMA completion channel instead of spinning) and #29514 (upstream CI only). Version-only from this project's side: GGML_RPC is OFF here, so the changed transport is never compiled into jllama, and upstream's release.yml is untouched, so the CUDA/ROCm/OpenVINO pins derived from it stay as they are. All eight patches apply unchanged (fresh configure; the applier's stamp lists all eight against d7fb90e8e). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
bernardladenthin
had a problem deploying
to
maven-central
September 27, 2026 11:00 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
startgate
September 27, 2026 11:00 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 27, 2026 11:00 — with
GitHub Actions
Failure
|
This was referenced Sep 27, 2026
Merged
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
b11080→b11211in six steps (b11103,b11104,b11160,b11163,b11209,b11211). The only incompatible step isb11104(upstream #28690, multi-address--host). Therepatches/0007still applied but named the removedserver_http_context::thread/listening_addressmembers, so it now usesjoin()/listening_addresses.NativeServergainsgetHosts(), andgetHost()now returns the first address. All eight patches still apply and are still required. Each range has its own row indocs/history/llama-cpp-breaking-changes.md.cuda-toolkit-13-4. On Windows,Jimver/cuda-toolkithas no 13.4 in any release or on master, so it is replaced by NVIDIA's redist archives, using upstream'swindows-setup-cudacomponent list.rocm-*classifiers need a ROCm 10 runtime.build.batretries uncached. It now retries once without sccache after a failed configure or build, and passes-DGGML_CCACHE=OFFso the uncached path really is uncached.ggml-org/free-disk-space. It now runs in the CUDA, ROCm and SYCL Linux jobs and in both emulator jobs.20260712-79e54f9image.Test plan
ctest) pass locally, and Java tests pass against the locally built lib.GGML_RPCisOFF, so it is never compiled intojllama.NativeServerSmokeTest: coversgetHosts()(comma-separated, trimmed, empty → default).cuda-toolkit-13-4package, the CUDA 13.4 redist archives and AMD'sstable.repo.amd.comwheel index: the package servers are blocked here.build.batchanges: there is no Windows machine here.build WITH sccache failed -- retrying ONCE. If SYCL does it every time, sccache should come out of that job.Related issues / PRs
Refs ggml-org/llama.cpp#28690, ggml-org/llama.cpp#24669
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.md🤖 Generated with Claude Code
https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Generated by Claude Code