Skip to content

llama.cpp b11080 → b11211, CUDA 13.4, ROCm 10 (TheRock), sccache on every Windows Ninja job - #456

Merged
bernardladenthin merged 10 commits into
mainfrom
claude/busy-archimedes-lo9uyy
Sep 27, 2026
Merged

bernardladenthin merged 10 commits into
mainfrom
claude/busy-archimedes-lo9uyy

Conversation

@bernardladenthin

Copy link
Copy Markdown
Owner

Summary

  • llama.cpp b11080 → b11211 in six steps (b11103, b11104, b11160, b11163, b11209, b11211). The only incompatible step is b11104 (upstream #28690, multi-address --host). There patches/0007 still applied but named the removed server_http_context::thread / listening_address members, so it now uses join() / listening_addresses. NativeServer gains getHosts(), and getHost() now returns the first address. All eight patches still apply and are still required. Each range has its own row in docs/history/llama-cpp-breaking-changes.md.
  • CUDA 13.3 → 13.4 to match upstream. Linux installs cuda-toolkit-13-4. On Windows, Jimver/cuda-toolkit has no 13.4 in any release or on master, so it is replaced by NVIDIA's redist archives, using upstream's windows-setup-cuda component list.
  • ROCm 6.3.4 / HIP SDK 26.Q1 → ROCm 10.0.0 from TheRock on both ROCm jobs, installed from pip wheels as upstream's release jobs do. Each OS now builds every GPU target TheRock builds for it: Linux 27, Windows 23. That includes gfx900/gfx906/gfx90c/gfx1153, which upstream omits; they stay only as long as they build without patches. BREAKING (runtime): consumers of the rocm-* classifiers need a ROCm 10 runtime.
  • CI hardening:
    • sccache in every Windows Ninja job. Added to arm64, arm64-OpenCL, ROCm, SYCL and OpenVINO; the arm64 jobs use the native aarch64 release.
    • build.bat retries uncached. It now retries once without sccache after a failed configure or build, and passes -DGGML_CCACHE=OFF so the uncached path really is uncached.
    • ggml-org/free-disk-space. It now runs in the CUDA, ROCm and SYCL Linux jobs and in both emulator jobs.
    • dockcross help text. The wrapper scripts' help text now names the pinned 20260712-79e54f9 image.

Test plan

  • Native build at b11209: full native build and C++ suite (ctest) pass locally, and Java tests pass against the locally built lib.
  • b11211 checked by configure only: a fresh configure applies all eight patches. The only source change in this step is RPC transport code, and GGML_RPC is OFF, so it is never compiled into jllama.
  • NativeServerSmokeTest: covers getHosts() (comma-separated, trimmed, empty → default).
  • CI is green on this branch. Not run yet. Some of this could not be checked from the sandbox:
    • NVIDIA's rhel8 cuda-toolkit-13-4 package, the CUDA 13.4 redist archives and AMD's stable.repo.amd.com wheel index: the package servers are blocked here.
    • All Windows build.bat changes: there is no Windows machine here.
    • Worth checking in the first run: whether any Windows job logs build WITH sccache failed -- retrying ONCE. If SYCL does it every time, sccache should come out of that job.
  • Docs / CHANGELOG updated (README classifier rows, CLAUDE.md, CHANGELOG, breaking-changes history).

Related issues / PRs

Refs ggml-org/llama.cpp#28690, ggml-org/llama.cpp#24669

Checklist

  • I have read CONTRIBUTING.md and CODE_OF_CONDUCT.md
  • My commits follow Conventional Commits (the commits use the repo's existing "Upgrade llama.cpp from …" style)
  • No security-sensitive changes

🤖 Generated with Claude Code

https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8


Generated by Claude Code

First step of the b11080 -> b11209 series. 23 upstream commits, all internal to
this project's surface: CUDA/Metal/SYCL/OpenCL/hexagon kernels, cpp-httplib
0.57.1, a Muse Glimmer tool-call parser fix and MiMo-V2.6 template detection.
No priority-list header moves in this range, and all eight local patches apply
unchanged (replayed against every tag, not just the endpoints).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
upstream #28690 lets --host take a comma-separated list and binds every address.
It removed server_http_context::thread and ::listening_address in favour of
join() and listening_addresses.

patches/0007 still applied cleanly at b11104 -- its hunks are nowhere near the
changed lines -- but two of its own + lines in llama_server_attach named the
removed members, so the native build would have failed. It now logs every
listening address and blocks in ctx_http.join(), as upstream's llama_server()
does.

NativeServer gains getHosts() (the parsed list, never empty); getHost() returns
its first element instead of the raw "a,b" string. Two new NativeServerSmokeTest
cases pin the split/trim and the fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
56 upstream commits, compatible from this project's side: backend kernels, the
llama.cpp 0.5.0 version bump, router fixes, new model support (Ling 3.0 VL,
Gemma 4 DSpark draft) and two OAI-layer features -- input_image as a Responses
function_call_output and video_url as an alias of input_video. Only a private
member of server-context.h moves; all eight patches apply at every tag.

The two OAI features are recorded in the history row as Java-API candidates
(ContentPart has no video part; ResponsesApiSupport flattens function_call_output
to text), not changed here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
upstream #24669 adds the extended batch API (llama_batch_ext + llama_process)
next to the classic llama_batch/llama_decode, which is unchanged. Purely
additive for this project: nothing in src/main/cpp calls the decode or batch
helpers directly. Kept as its own step because it is the one new public
llama.h surface in the b11080 -> b11209 series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Last step of the b11080 -> b11209 series. 46 upstream commits, compatible from
this project's side: backend kernels, a GGUF-driven W4A4 precision policy with no
public API, shared Unicode helpers in common.h, cpp-httplib 0.58.0, cleanup after
a failed state restore, a grammar token_id fix, and a revert of the --fit
context-length change. All eight patches are still required.

Verified at b11209 from a fresh configure: patches applied (8, stamp matches),
Release build with 0 warnings, ctest 559/559, 40 Java_* exports, and
mvn clean verify 1774 run / 0 failures with NativeLibraryLoadSmokeTest 4/4
confirming the pin against the linked build-info. The CHANGELOG entry covers the
whole series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
AMD now builds and releases ROCm through TheRock. Both rocm-* jobs install
its Python wheels (rocm[libraries,devel]==10.0.0 from
stable.repo.amd.com/rocm/whl-next) the way upstream llama.cpp's own
ubuntu-rocm / windows-rocm release jobs do at b11209, replacing the
repo.radeon.com 6.3.4 apt repo on Linux and the HIP SDK 26.Q1 installer
on Windows. Paths are read back with rocm-sdk path.

Linux now uses CMake's native HIP language with ROCm's clang as the HIP
compiler (upstream's form); Windows uses the TheRock clang under
lib\llvm\bin. The GPU target lists are copied from upstream: newer
architectures are added, gfx900/gfx906 are dropped because llama.cpp no
longer ships them and TheRock never marks them release-ready.

Not validated locally: stable.repo.amd.com is unreachable from the
sandbox, so the first CI run is the proof.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
CUDA 13.3 -> 13.4, matching upstream llama.cpp. Linux installs
cuda-toolkit-13-4 from NVIDIA's rhel8 repo. Windows drops
Jimver/cuda-toolkit, which has no 13.4 in any release or on master, and
assembles the toolkit from NVIDIA's per-component redist archives with
upstream's windows-setup-cuda component list. Classifiers stay cuda13-*.

ROCm Linux keeps gfx900/gfx906 on top of upstream's target list: TheRock
10 still builds them. They stay only as long as they build without
patches and do not hold back a newer ROCm.

ggml-org/free-disk-space now runs first in the CUDA, ROCm and both SYCL
Linux jobs, and replaces the hand-rolled toolchain removal in the two
emulator jobs (with android, large-packages, tool-cache and swap kept).

The dockcross wrappers' help text named the tags they were first
generated from; it now names the pinned 20260712-79e54f9 image.

Not validated locally: the NVIDIA and AMD package servers are
unreachable from the sandbox, so the first CI run is the proof.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
The ROCm target lists were upstream llama.cpp's, which differed between
Linux and Windows partly by accident. They are now every target TheRock
builds per OS (its SUPPORTED_GPUS.md): Linux 27 targets, Windows 23.
That adds gfx90c and gfx1153 on Linux and gfx900/gfx906/gfx90c on
Windows. The two lists now differ only by the Instinct parts
(gfx908/gfx90a/gfx942/gfx950), which ROCm supports on Linux alone.

The extras upstream omits are build-passing only in TheRock; they stay
as long as they build without patches and do not block a newer ROCm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Five Windows build jobs never installed sccache: arm64, arm64-OpenCL,
ROCm, SYCL and OpenVINO. They now get the same USE_CACHE/SCCACHE_WEBDAV
env and install step as the other Ninja jobs. The two arm64 jobs use the
native aarch64-pc-windows-msvc release, which exists (the comment that
said only x86_64 was available is gone). Only the two MSVC-classifier
jobs stay uncached: the Visual Studio generator ignores launchers.

build.bat's probe only proves sccache can wrap cl.exe, while these jobs
compile with clang-cl, ROCm's clang or icx. So a configure or build
that fails with sccache as the launcher is now retried once from a
clean build dir without it. The retry is unconditional because cmd
cannot tee output to match an error signature; a real compile error
costs one extra uncached attempt, a cache incompatibility ends green.

Every configure also passes -DGGML_CCACHE=OFF. Without it ggml
self-enables any sccache on PATH whenever no launcher is set, i.e. in
the probe-failed and retry cases, and the uncached build went through
sccache after all. No goto/labels: the file is checked out with LF
line endings, where cmd's label search is unreliable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Two upstream commits: #29440 (RPC waits on the RDMA completion channel
instead of spinning) and #29514 (upstream CI only). Version-only from
this project's side: GGML_RPC is OFF here, so the changed transport is
never compiled into jllama, and upstream's release.yml is untouched, so
the CUDA/ROCm/OpenVINO pins derived from it stay as they are.

All eight patches apply unchanged (fresh configure; the applier's stamp
lists all eight against d7fb90e8e).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
@sonarqubecloud

Copy link
Copy Markdown

This branch had an error being deployed

1 failed deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants