From d52e6ce024a5234ee46d76e9f40d6008ce8071c3 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 09:58:21 +0000 Subject: [PATCH 01/10] Upgrade llama.cpp from b11080 to b11103 First step of the b11080 -> b11209 series. 23 upstream commits, all internal to this project's surface: CUDA/Metal/SYCL/OpenCL/hexagon kernels, cpp-httplib 0.57.1, a Muse Glimmer tool-call parser fix and MiMo-V2.6 template detection. No priority-list header moves in this range, and all eight local patches apply unchanged (replayed against every tag, not just the endpoints). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index b655cf90c..595d758ce 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11080** +Current llama.cpp pinned version: **b11103** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11080 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11103 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11080`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11103`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1655,7 +1655,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11080`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11103`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index daa243ebd..87d2ac22c 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11080](https://img.shields.io/badge/llama.cpp-%23b11080-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11080) +[![llama.cpp b11103](https://img.shields.io/badge/llama.cpp-%23b11103-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11103) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 807d7ff64..e2da834ba 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -744,3 +744,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11062–b11069 | patches + upstream verification | **Eight patches now: `0011` dropped, the other eight apply unchanged.** #29161 deleted the very `INVALID` branch of `common_peg_until_parser` that `0011` patched: the until-parser now consumes an undecodable run in every mode (strict included), records it on the result, and `common_chat_peg_mapper` renders the node through `sanitized_text()` — one U+FFFD per run, per the Unicode "maximal subpart" rule (`\xE4\xB8` + `c` → one replacement, `\xFF\xFE` → two), with the text after the run kept. `0011` had returned only the text *before* the byte. The lenient incomplete-at-end branch is unchanged (trailing bytes withheld). So the applier failed loud — `patch failed: common/peg-parser.cpp:680` — exactly as designed, and the patch was **dropped, not refreshed**, per the `0009`/`0013` precedent; it had never been filed upstream, so nothing to close. Its runnable guard was kept and re-pointed: `ContentOnlyParseUtf8` in `src/test/cpp/test_utils.cpp` now pins upstream's replacement contract (six cases, one more than before, the extra one pinning the run boundary `\xFF\xFE` → two U+FFFD), so a future upstream revert to `FAIL` still reds `C++ Tests` everywhere. **Replayed in filename order against pristine b11069**: `0001` `0002` `0003` `0006` `0007` `0008` `0010` `0012` apply clean, `0011` is the only failure. **All standing drop-checks still say "still required"** at the pristine tag: `0001` (`common_params_parse_main` 0 occurrences in `b11069:common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1281` — i.e. [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open upstream and is **not** what this range fixed), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`), and `0003`/`0006`/`0007`/`0008` absent upstream. `.github/verify-patches-applied.sh`'s header comment no longer lists `0011`. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `68d9053af` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty, 0010 cast present), extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean, `ctest` **552/552** (551 → 552: the re-pointed guard gained one case), `nm -D` 40 `Java_*` exports and 0 mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `mvn clean` — `nativeBuildInfoMatchesPinnedVersionConstant` confirms `LlamaCppVersion.LLAMA_CPP_VERSION` (`b11069`) against the linked `build-info`. `test_utils.cpp` is clean under the CI-pinned clang-format 23.1.1; `spotless:check` clean. Model-backed Java tests were not run (HF-blocked sandbox); `NativeServerAttachIntegrationTest.completion_overHttp_served`, the test that first surfaced the `0011` failure, is the CI-side confirmation for this drop. | | b11069–b11080 | Eleven commits, 72 files, **1244 KiB** (`tools/ui` untouched, so the WebUI-excluded figure is the same). The size is almost entirely one commit: **#29197** ("hexagon: overhaul of buffer and DMA handling to support 64bit mappings") rewrites 46 files under `ggml/src/ggml-hexagon/` for +7085/−6108 on its own, and this project builds no hexagon classifier. The rest of `ggml/src` is three self-contained backend changes — `ggml-cpu` (**#23492**, ARM repack kernels for Q1_0, +762/−0, purely additive), `ggml-metal` (**#29206**, fusion-pattern op list simplification) and `ggml-sycl` (**#28918**, MKL-FA softmax load coalescing). **`ggml/include` is byte-identical**, so no public ggml header moved. Outside ggml: `.github/workflows` (**#28991**, upstream's own self-hosted CI refactor — not consumed here), `scripts/snapdragon`, `tests/test-backend-ops.cpp` (**#29204**, regex `-o` filter), `docs/backend/snapdragon`, and three tool READMEs regenerated for the new env vars. **One priority-list file changed and it is the consequential one: `common/json.h`** (**#28518**, "json: Fixed json enum handling", +4/−0) — see the patch row below. `common/arg.cpp` gains six `set_env(...)` calls (**#27380**: `LLAMA_ARG_TEMPERATURE` / `_TOP_P` / `_MIN_P` / `_REPEAT_PENALTY` / `_PRESENCE_PENALTY` / `_FREQUENCY_PENALTY`) on existing options — additive, no option added or removed, so the `ModelFlag`/`ModelOption` contract is untouched; note only that those six now read an environment default when the flag is absent, which an embedding host with those vars set would inherit. `tools/server/` moves by exactly one functional line (**#28938**: `unset_reserved_args` also unsets `LLAMA_ARG_API_KEY_FILE`, so a router no longer forwards it to spawned children) plus its README. | **No project-source change; one patch dropped.** The three mechanical server-contract greps have real input this time (`tools/server/` is in the range) and all three come back **identical**: the request-field set is 68 names, the bounded-field set 23, and the response-key set 142, unchanged between the two tags — `server-schema.cpp`, `server-task.cpp`, `server-context.cpp` and `server-common.h` are all byte-unchanged, so a contract change behind a stable signature is ruled out by construction rather than by reading. Wire-name extraction re-ran against b11080's sources and is unchanged at **138 CLI / 57 request / 15 trainer** names. `common/json.h`'s change is additive (a new constructor overload and one line in a type trait) and cannot break a caller; what it does is retire a patch. | | b11069–b11080 | patches + upstream verification | **Eight patches now: `0010` dropped, the other eight apply unchanged.** **#28518** gives `common_json_value` an `std::is_enum`-gated constructor that delegates to `std::underlying_type`, and adds `std::is_enum` to `common_json_is_value`. That fixes at its root the trap `0010` worked around at the emit site: an unscoped enum no longer binds to `common_json_value(bool)`, so upstream's own `get_res_model_info()` emits a numeric `vocab_type` with no cast. **This drop is the case the by-hand drop-check exists for, and it is worth recording precisely: `0010` still applied cleanly at b11080** — upstream never touched the emit site — so the fail-loud applier said nothing and a redundant carry would have shipped silently. The standing check is worded "did upstream cast the value themselves?"; the answer was no and the correct verdict was still *drop*, because the defect is gone. Dropped, not refreshed, per the `0009`/`0011`/`0013` precedent; it had never been filed upstream, so nothing to close. **The runnable guard was kept and re-pointed** (the `0011` precedent): the `CommonJsonEnumTrap` pair in `src/test/cpp/test_json_helpers.cpp` is now the `CommonJsonEnum` trio and pins upstream's contract — an uncast enum serialises as its numeric value, an explicit `static_cast` is equivalent, and a real `bool` is still a boolean (the new overload sits next to `common_json_value(bool)`, so that one is worth pinning too). A bump that loses the enum constructor therefore reds `C++ Tests` on every platform, and the response is to reinstate both the cast and the patch; `jllama.cpp` keeps its own two `"vocab_type"` casts, which are correct either way. `.github/verify-patches-applied.sh` lost its third check with the patch — `0010` was the only patch with no runnable guard, which is precisely what that check was for — and keeps its two generic assertions (every patch on disk is in the stamp; the patched tree is dirty); the matching `TODO.md` coverage-gap entry is resolved and removed. **Replayed in filename order against pristine b11080**: all nine of the previous set apply clean, `0010` included — which is the point. **The remaining four standing drop-checks all say "still required"** at the pristine tag: `0001` (`common_params_parse_main` 0 occurrences in `b11080:common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`, so [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero-sum guard), `0014` (`common_log_set_callback` 0 occurrences in `b11080:common/log.h`); `0003`/`0006`/`0007`/`0008` remain absent upstream. **Only two patch-target files were in the range at all** (`common/arg.cpp` for `0001`, `tools/server/server-models.cpp` for `0008`), both far from the patched hunks, and both proven by replay rather than by reading. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `1d72b05d3` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty), the fetched `server-context.cpp` confirmed **uncast** at the emit site and `common/json.h` confirmed to carry the `is_enum` overload, Release build clean (0 errors, 0 warnings in project sources), `ctest` **559/559** (558 → 559: the re-pointed guard gained one case), `nm -D` 40 `Java_*` exports, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `mvn clean` — `nativeBuildInfoMatchesPinnedVersionConstant` confirms `LlamaCppVersion.LLAMA_CPP_VERSION` (`b11080`) against the linked `build-info`. Full `mvn test` **1772 run, 0 failures, 0 errors** (272 skipped — the model-gated classes, no GGUF in this sandbox). `test_json_helpers.cpp` is clean under the CI-pinned clang-format 23.1.1; `spotless:check` clean; SpotBugs **0** findings. Model-backed Java tests were not run (HF-blocked sandbox); `NativeServerAttachIntegrationTest.models_reportNumericVocabType` — whose failure message now names #28518 instead of the retired patch — is the CI-side confirmation that the wire value stayed numeric across this drop. | +| b11080–b11103 | 23 commits, 49 files outside `tools/ui`, **301 KiB**. Backend-internal almost throughout: CUDA (#29224 sm_70 tile fix, #29155 contiguous convert, #29135 conv2d implicit GEMM, #28536 FA swizzle refactor), Metal (#29220 FA mask bounds), SYCL (#29132, #28895), OpenCL (#29055 A8 Q4_0 dp4a kernel), hexagon (#29199). `vendor/cpp-httplib` moves 0.56 → **0.57.1** (#29214, #29239). Outside ggml: `common/chat.cpp` + `common/parsers/muse-glimmer.cpp` (#29242, tool-call parser fix; #29257 MiMo-V2.6 template detection), `cmake/llama-config.cmake.in` (#29228, repeated `find_package` — not consumed here, this project uses `FetchContent`), `src/llama-context.cpp` (#26625, sched-reserve reporting), tests and the Python converter. | **No project-source change.** No header on the priority list moves (`include/`, `common/common.h`, `common/arg.h`, `tools/server/*.h` and `tools/mtmd/*.h` are byte-identical between the two tags). The server contract greps come back identical over the whole b11080–b11209 range (68 request fields, 23 bounded fields, 142 response keys), and so does the set of `common/arg.cpp` option names. | +| b11080–b11103 | patches + upstream verification | **All eight patches apply unchanged**, replayed in filename order against every tag of the range, not only the endpoints. No patch-target file is in the range. All standing drop-checks still say "still required" (checked at b11209, the end of this bump series — see the b11163–b11209 row). | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 93cc34211..ae3f20529 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11080 + GIT_TAG b11103 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index d4e43a049..bd0e98146 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11080"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11103"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11080-"} — call + * plus the resolved upstream commit, e.g. {@code "b11103-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11080"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11103"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11080"; + public static final String LLAMA_CPP_VERSION = "b11103"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 7dfb2187561f54472e5967eb381283d332b7a32d Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:00:07 +0000 Subject: [PATCH 02/10] Upgrade llama.cpp from b11103 to b11104: multi-address --host upstream #28690 lets --host take a comma-separated list and binds every address. It removed server_http_context::thread and ::listening_address in favour of join() and listening_addresses. patches/0007 still applied cleanly at b11104 -- its hunks are nowhere near the changed lines -- but two of its own + lines in llama_server_attach named the removed members, so the native build would have failed. It now logs every listening address and blocks in ctx_http.join(), as upstream's llama_server() does. NativeServer gains getHosts() (the parsed list, never empty); getHost() returns its first element instead of the raw "a,b" string. Two new NativeServerSmokeTest cases pin the split/trim and the fallback. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CLAUDE.md | 10 +++--- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../0007-server-attach-http-frontend.patch | 10 +++--- .../ladenthin/llama/server/NativeServer.java | 32 +++++++++++++++++-- .../llama/value/LlamaCppVersion.java | 8 ++--- .../llama/server/NativeServerSmokeTest.java | 17 ++++++++++ 8 files changed, 64 insertions(+), 19 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 595d758ce..1f3601b3c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11103** +Current llama.cpp pinned version: **b11104** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11103 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11104 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11103`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11104`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -751,7 +751,7 @@ Current patches: | `0001-win32-arg-parse-embed-guard.patch` | Windows JNI regression from llama.cpp **#24779** (introduced b9739): on Windows `common_params_parse` re-derived argv from the **process** command line (`GetCommandLineW`) and adopted it, so an embedded/JNI caller (`java.exe`) lost its `--model …` args → "Failed to parse model parameters". b9789 narrowed the unconditional override to a **count-guard** (`if (static_cast(utf8.buf.size()) == argc) { argv = utf8.ptrs.data(); }`), but that is exactly the variant the project already found breaks its Windows server-integration tests (when the embedded argv length coincides with `java.exe`'s). The patch carries the **complete upstream change** (so it can be submitted to llama.cpp verbatim and then dropped here): **(1)** `common_params_parse` parses **exactly the argv it is given** (no `GetCommandLineW` magic) and a new `common_params_parse_main()` wrapper holds the UTF-8 recovery for the standalone tools' `main()` (`common/arg.{cpp,h}`); **(2)** the **~34 standalone `main()` call sites** (every `common_params_parse(argc, argv, …)` across `tools/*`, `examples/*` and the `tests/*` programs) flip to `common_params_parse_main()`; **(3)** a `tests/test-arg-parser.cpp` regression case pins that `common_params_parse` honors a caller-supplied argv. The embedded caller (`jllama.cpp`) keeps calling `common_params_parse` and is never overridden. **Our subproject build compiles only the `arg.{cpp,h}` core** — `LLAMA_BUILD_TOOLS`/`LLAMA_BUILD_TESTS` are OFF for a FetchContent subproject — so the flips + test are applied-but-not-compiled here; they were validated via a one-off `-DLLAMA_BUILD_TOOLS=ON -DLLAMA_BUILD_TESTS=ON` build (the new test compiles and its asserts pass; `test-arg-parser`'s only red there is the live `ggml.ai` download check, which is sandbox-network, not the patch). Because it spans **36 files** it must be refreshed on every llama.cpp bump (the applier fails loud). **Refreshed at the b10679 bump:** upstream rewrote `tests/test-save-load-state.cpp`'s `main()` to take a `--models DIR` option, which it strips itself into a `filtered_argv` before calling `common_params_parse(fargc, filtered_argv.data(), …)`. That call site therefore stopped qualifying for the `_main()` flip — by this patch's own rule a caller that builds its own argv must use `common_params_parse` directly, so its argv is kept — and the hunk was **dropped** rather than refreshed (37 → 36 files). Caveat for whoever submits this upstream: that `main()` now filters a possibly-mojibake Windows argv *before* any UTF-8 recovery, so the fully correct upstream form there is recover-then-filter, not a one-line flip. It is out of scope for the downstream carry because `LLAMA_BUILD_TESTS` is OFF here, so the file is never compiled. **Still required at b10679, verified rather than assumed:** `common_params_parse` in pristine `b10679:common/arg.cpp` still carries the `#ifdef _WIN32` count-guarded `argv = utf8.ptrs.data()` override, and `common_params_parse_main` appears nowhere in `b10679:common/arg.h` — upstream has not adopted the fix. The upstream-facing write-up, including a standalone reproducer that makes llama.cpp's own `test-arg-parser` fail on unmodified `master`, lives in [docs/upstream-investigation-win32-argv-substitution.md](docs/upstream-investigation-win32-argv-substitution.md). **Reported upstream as [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416)** (2026-08-01, label `bug-unconfirmed`, first bad commit `508a475`); the issue asks which of the two directions the maintainers prefer before a PR is opened, so this patch stays downstream until they answer. | | `0002-server-preserve-caller-load-progress-callback.patch` | Load-progress-callback regression introduced in llama.cpp **b9789**: `server_context::load_model` (`tools/server/server-context.cpp`) now **unconditionally** installs the server's own load-progress reporter on `params_base.load_progress_callback` immediately before `common_init_from_params`, clobbering any callback the embedding caller already set. libjllama's `LoadProgressCallback` feature wires `common_params.load_progress_callback` to a JNI trampoline *before* calling `load_model`, so the bump silently killed it — `LoadProgressCallbackTest` saw zero progress updates and the abort-on-`false` path never threw. The patch guards the assignment with `if (params_base.load_progress_callback == nullptr)`, so the server installs its own reporter **only when the caller hasn't** — a caller-supplied callback survives and fires during load. Standalone `llama-server` (no caller callback, so the field is null) is unaffected. Same JNI-vs-standalone divergence class as `0001`. **The guard is `== nullptr || == load_progress_callback`, and the second disjunct must never be dropped:** `load_progress_text` is a **local** of `load_model()`, and upstream re-assigns both fields on every call so the `user_data` always points at the current frame. `load_model()` runs a **second** time when resuming from the sleeping state (`--sleep-idle-seconds`), and by then `params_base` holds *our own* callback from the first load — a bare nullptr check skips the re-assignment and leaves `user_data` pointing into a **dead stack frame**, which segfaults inside `load_progress_callback()` on the first request after an idle window. That was a latent defect in this patch from the day it was written; only a second `load_model()` can reach it, and nothing exercised sleep until `IdleSleepWakeIntegrationTest` was added. | | `0003-pr22393-server-add-slot-prompt-similarity-getter-setter.patch` | **Upstream-PR carry** of [ggml-org/llama.cpp#22393](https://github.com/ggml-org/llama.cpp/pull/22393) ("server : add slot_prompt_similarity getter/setter"). Purely additive: adds `server_context::get_slot_prompt_similarity()` / `set_slot_prompt_similarity(float)` (`tools/server/server-context.{cpp,h}`) so an embedding/JNI caller can query and tune the slot-selection threshold at runtime without reloading the model. Verbatim copy of the PR, which **upstream closed without merging** (rejected as exposing unsafe internal state — see the patch header). Carried permanently; it will not be droppable via a version bump. | -| `0007-server-attach-http-frontend.patch` | **Adds `llama_server_attach(argc, argv, server_context&)`** so the `NativeServer` *attach mode* can serve an **already-loaded `LlamaModel`** over the upstream HTTP frontend — no second model load, no `start_loop()`; the LlamaModel's worker keeps driving the shared `server_context` and the HTTP routes post tasks to its queue (the queue is the synchronization point). Mechanically: (1) extracts the **pure core route table** (`health` … `slots`) out of `llama_server()` into `static void llama_server_register_common_routes(ctx_http, routes)` (shared, so the two entry points cannot drift on the core endpoint set). **Scope note (narrowed at the b10154 bump):** the helper deliberately carries **only** the stable, state-independent route table — **not** the resumable-streaming routes (their handlers differ between router / non-router), the GCP-compat shim, or the experimental **CORS-proxy / MCP-server / built-in-tools** wiring. b10154 (upstream MCP-server support) moved the streaming routes into the middle of that block and coupled tools/CORS to a per-call `server_mcp mcp_mgr` lifecycle, so the earlier contiguous "route-table + CORS-proxy + tools" extraction is no longer possible; `llama_server()` keeps all of that inline, **byte-identical to upstream b10154** (only the route-table block is factored out). (2) adds `llama_server_attach`, which parses only the HTTP-side argv via `common_params_parse`, starts the stream-session GC + `server_http_context`, registers the common route table, the **non-router** resumable-streaming handlers (upstream b10154 paths `/v1/stream` GET/DEL + `/v1/streams/lookup` POST), the GCP-compat shim, and **403 "disabled" stubs for `/cors-proxy` + `/tools`** (attach mode does not wire the experimental CORS-proxy / MCP / built-in-tools host — those belong to a full `llama-server`, not an embedded model), marks ready immediately (model already loaded), and blocks on the HTTP thread until `llama_server_request_shutdown()` — never calling `common_init()`, backend init, `ctx_server.terminate()` or `llama_backend_free()` (the embedding caller owns those). Applies after `0001`+`0006` (same file); closes the "NativeServer — reuse an already-loaded LlamaModel" TODO. Upstream-submittable ("server: let embedding callers attach the HTTP frontend to an existing server_context"). **Refreshed at the b10519 bump:** upstream #26347 dropped the API key from the `/models` + `/v1/models` public-endpoint set and deleted the two trailing `// public endpoint (no API key check)` comments on those route registrations. Those two lines sit inside this patch's route-table removal block, so `git apply` failed ("patch does not apply", `server.cpp:258`) at **every** tag from b10519 on; the fix was to drop the now-wrong comment from all four affected lines (2 on the `-` side, 2 in the extracted helper on the `+` side), keeping the helper byte-identical to the block it replaces. **This is the invariant to re-check on every bump:** the `+` side of `llama_server_register_common_routes()` must stay a verbatim copy of the route table it factors out of `llama_server()`. | +| `0007-server-attach-http-frontend.patch` | **Adds `llama_server_attach(argc, argv, server_context&)`** so the `NativeServer` *attach mode* can serve an **already-loaded `LlamaModel`** over the upstream HTTP frontend — no second model load, no `start_loop()`; the LlamaModel's worker keeps driving the shared `server_context` and the HTTP routes post tasks to its queue (the queue is the synchronization point). Mechanically: (1) extracts the **pure core route table** (`health` … `slots`) out of `llama_server()` into `static void llama_server_register_common_routes(ctx_http, routes)` (shared, so the two entry points cannot drift on the core endpoint set). **Scope note (narrowed at the b10154 bump):** the helper deliberately carries **only** the stable, state-independent route table — **not** the resumable-streaming routes (their handlers differ between router / non-router), the GCP-compat shim, or the experimental **CORS-proxy / MCP-server / built-in-tools** wiring. b10154 (upstream MCP-server support) moved the streaming routes into the middle of that block and coupled tools/CORS to a per-call `server_mcp mcp_mgr` lifecycle, so the earlier contiguous "route-table + CORS-proxy + tools" extraction is no longer possible; `llama_server()` keeps all of that inline, **byte-identical to upstream b10154** (only the route-table block is factored out). (2) adds `llama_server_attach`, which parses only the HTTP-side argv via `common_params_parse`, starts the stream-session GC + `server_http_context`, registers the common route table, the **non-router** resumable-streaming handlers (upstream b10154 paths `/v1/stream` GET/DEL + `/v1/streams/lookup` POST), the GCP-compat shim, and **403 "disabled" stubs for `/cors-proxy` + `/tools`** (attach mode does not wire the experimental CORS-proxy / MCP / built-in-tools host — those belong to a full `llama-server`, not an embedded model), marks ready immediately (model already loaded), and blocks on the HTTP thread until `llama_server_request_shutdown()` — never calling `common_init()`, backend init, `ctx_server.terminate()` or `llama_backend_free()` (the embedding caller owns those). Applies after `0001`+`0006` (same file); closes the "NativeServer — reuse an already-loaded LlamaModel" TODO. Upstream-submittable ("server: let embedding callers attach the HTTP frontend to an existing server_context"). **Refreshed at the b10519 bump:** upstream #26347 dropped the API key from the `/models` + `/v1/models` public-endpoint set and deleted the two trailing `// public endpoint (no API key check)` comments on those route registrations. Those two lines sit inside this patch's route-table removal block, so `git apply` failed ("patch does not apply", `server.cpp:258`) at **every** tag from b10519 on; the fix was to drop the now-wrong comment from all four affected lines (2 on the `-` side, 2 in the extracted helper on the `+` side), keeping the helper byte-identical to the block it replaces. **This is the invariant to re-check on every bump:** the `+` side of `llama_server_register_common_routes()` must stay a verbatim copy of the route table it factors out of `llama_server()`. **Refreshed at the b11104 bump** (upstream #28690, multi-address `--host`): `server_http_context` lost its single `thread` and `listening_address` members in favour of `join()` and a `listening_addresses` vector, one listener thread per bound address. The patch still *applied* cleanly there — only its own `+` lines named the removed members — so the applier could not see it; `llama_server_attach` now logs every address and blocks in `ctx_http.join()`, exactly as upstream's `llama_server()` does. | | `0008-server-models-worker-cmd-override.patch` | **Makes router mode usable in-JVM.** The router (`server-models.cpp`) spawns each model worker by re-executing its own binary (`get_server_exec_path()` = `/proc/self/exe` & friends) — inside a JVM that binary is `java`, not a llama-server, so embedded router workers could never start. The patch adds env `LLAMA_SERVER_WORKER_CMD` (whitespace-split; read in `server_model_meta::update_args`) which replaces only the leading binary-path token of the rendered worker args, letting an embedding host relaunch workers through its own bootstrap — e.g. `java -cp app.jar net.ladenthin.llama.server.NativeServer` (each worker is then a fresh JVM running the classic single-model `NativeServer`). Exposed in Java as `NativeServer.setWorkerCommand(String...)` (JNI `setenv`); exercised by `RouterModeIntegrationTest` (Linux CI). Upstream-submittable (also useful for containerized/wrapped deployments). | | `0006-server-embed-native-server-jni.patch` | **Makes `server.cpp`'s `llama_server` embeddable in the JVM** so the `NativeServer` JNI bridge can run the full upstream HTTP server (WebUI included) inside `libjllama` — see "Two server modes" below. b9870 already exposes `int llama_server(int, char**)` (non-static; no `main` in the file), so the patch only adds embedded-mode support: (1) a `g_llama_server_embedded` flag + `llama_server_set_embedded()` / `llama_server_request_shutdown()` (declared in the committed `src/main/cpp/native_server_bridge.h`); (2) skips installing the process-wide SIGINT/SIGTERM handlers when embedded (they would hijack the JVM's); (3) in embedded mode parses the **forwarded** argv via `common_params_parse` instead of `common_params_parse_main` (whose `GetCommandLineW` recovery would pick up `java.exe`'s command line — the same Windows class of bug `0001` fixes). `llama_server_request_shutdown()` mirrors the SIGTERM path (invokes the installed `shutdown_handler` → `ctx_server.terminate()` unblocks `start_loop()`), giving JNI an out-of-band stop since `ctx_server` is loop-local. Applies **after `0001`** (which flips this call site to `common_params_parse_main`), so its context is the post-`0001` tree; regenerate against `0001`+source on a bump. Only touches `tools/server/server.cpp`. | | `0012-model-guard-zero-split-sum-and-name-the-device-index.patch` | **A GPU that reports zero free memory makes every model load fail with the unactionable `error loading model: vector`.** `llama_model_base::load_tensors` (`src/llama-model.cpp`) weights the per-device layer split by `ggml_backend_dev_memory()`'s `free`, then normalises: `splits[i] /= split_sum`. With a single device reporting `free == 0` that is `0/0` → **NaN** in every split point; NaN compares false against everything, so the `std::upper_bound` below returns the end iterator, `layer_gpu == n_devices()`, and `devices.at(layer_gpu)` throws `std::out_of_range` — whose libc++ `what()` is the bare string `"vector"`, which `llama.cpp`'s `catch (const std::exception &)` prints verbatim. Upstream's `free == 0 && total == 0` host-memory fallback does **not** fire, because `total` is `recommendedMaxWorkingSetSize` and is non-zero. **Reachable since b10618..b10797**: upstream `8c0b9cd04` ("metal : fix memory query under low-memory conditions", [#27701](https://github.com/ggml-org/llama.cpp/pull/27701)) changed `ggml-metal-device.m` to `*free = *total > cur ? *total - cur : 0`; before that clamp an over-committed device (`currentAllocatedSize > recommendedMaxWorkingSetSize`) *underflowed* to a huge `size_t`, which normalised fine, so the same precondition was harmless. That is why the `Java Tests macOS …` jobs went red at the b10792→b10797 step while every Linux/Windows job stayed green — **and why only a GPU build can fail this way at all**: `act_gpu_layers` is `devices.empty() ? 0 : …`, so with no GPU backend `devices` is empty, every layer returns early on `cpu_dev`, and the `.at()` line is unreachable. **Shape:** the two blocks are lifted out of `load_tensors` into free functions declared in `src/llama-model.h`, purely so they can be driven by a test — the failing state needs a real over-committed GPU and cannot be arranged through any public API. `llama_model_splits_normalize()` carries **the fix**: on `split_sum == 0` it `LLAMA_LOG_WARN`s and falls back to an even split (`splits[i] = float(i+1)/splits.size()`), the only neutral choice when no device can be preferred and exactly right for a single device. `llama_model_splits_select_device()` carries **the diagnostic**: it bounds-checks the index and throws a `std::runtime_error` naming the function, the offloaded layer, the device index, the split-point count **and the split points themselves** — with NaN splits that message prints `nan` and names the cause outright, which is precisely what was missing when this had to be diagnosed by reading source. **A second, backend-independent trigger reaches the same line**, found while writing this up and verified against the unfixed library: `--tensor-split` values are parsed with `std::stof` and never range-checked (`common/arg.cpp`), so `-ts 1,-1` cancels out, `split_sum` is 0 again, the split points become `[inf, -nan]`, and every layer maps one past the last device — on CUDA, Vulkan or ROCm just as much as on Metal, with no memory pressure involved. That is what makes this an ordinary upstream defect rather than a Metal edge case, and the warning names both causes rather than only the memory one. Also adds upstream `tests/test-model-split.cpp` (5 cases in upstream's `testing.h` style) + its `llama_build_and_test` registration. Touches `src/llama-model.{cpp,h}`, `tests/test-model-split.cpp` and `tests/CMakeLists.txt` — **none** of which any other patch touches, so it is independent of all of them. Upstream-submittable ("model: fall back to an even split when no device reports free memory"); **not yet filed upstream**. **Runnable guard: `src/test/cpp/test_model_split.cpp`** — a FetchContent subproject builds with `LLAMA_BUILD_TESTS=OFF`, so the upstream test above is applied-but-never-compiled here (same as `0001`'s test). That file drives the same two functions from `jllama_test`, which runs on **every** platform in `C++ Tests`, so a bump that drops this patch fails the build at link time everywhere instead of surfacing as one red macOS Java job. **Verification limit — read before assuming this can be dropped:** the *failing path* still cannot be reached without a GPU backend, so the guard pins the arithmetic (what actually broke), not the end-to-end load; the end-to-end proof is the macOS CI job. On a bump, re-check whether upstream added its own `split_sum == 0` guard (grep `split_sum` in `src/llama-model.cpp`) and **drop this patch rather than refreshing it** if they did — the fail-loud applier detects "does not apply", never "upstream already fixed this". | @@ -1655,7 +1655,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11103`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11104`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 87d2ac22c..4aac4aa6d 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11103](https://img.shields.io/badge/llama.cpp-%23b11103-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11103) +[![llama.cpp b11104](https://img.shields.io/badge/llama.cpp-%23b11104-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11104) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index e2da834ba..6fda5e72b 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -746,3 +746,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11069–b11080 | patches + upstream verification | **Eight patches now: `0010` dropped, the other eight apply unchanged.** **#28518** gives `common_json_value` an `std::is_enum`-gated constructor that delegates to `std::underlying_type`, and adds `std::is_enum` to `common_json_is_value`. That fixes at its root the trap `0010` worked around at the emit site: an unscoped enum no longer binds to `common_json_value(bool)`, so upstream's own `get_res_model_info()` emits a numeric `vocab_type` with no cast. **This drop is the case the by-hand drop-check exists for, and it is worth recording precisely: `0010` still applied cleanly at b11080** — upstream never touched the emit site — so the fail-loud applier said nothing and a redundant carry would have shipped silently. The standing check is worded "did upstream cast the value themselves?"; the answer was no and the correct verdict was still *drop*, because the defect is gone. Dropped, not refreshed, per the `0009`/`0011`/`0013` precedent; it had never been filed upstream, so nothing to close. **The runnable guard was kept and re-pointed** (the `0011` precedent): the `CommonJsonEnumTrap` pair in `src/test/cpp/test_json_helpers.cpp` is now the `CommonJsonEnum` trio and pins upstream's contract — an uncast enum serialises as its numeric value, an explicit `static_cast` is equivalent, and a real `bool` is still a boolean (the new overload sits next to `common_json_value(bool)`, so that one is worth pinning too). A bump that loses the enum constructor therefore reds `C++ Tests` on every platform, and the response is to reinstate both the cast and the patch; `jllama.cpp` keeps its own two `"vocab_type"` casts, which are correct either way. `.github/verify-patches-applied.sh` lost its third check with the patch — `0010` was the only patch with no runnable guard, which is precisely what that check was for — and keeps its two generic assertions (every patch on disk is in the stamp; the patched tree is dirty); the matching `TODO.md` coverage-gap entry is resolved and removed. **Replayed in filename order against pristine b11080**: all nine of the previous set apply clean, `0010` included — which is the point. **The remaining four standing drop-checks all say "still required"** at the pristine tag: `0001` (`common_params_parse_main` 0 occurrences in `b11080:common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`, so [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero-sum guard), `0014` (`common_log_set_callback` 0 occurrences in `b11080:common/log.h`); `0003`/`0006`/`0007`/`0008` remain absent upstream. **Only two patch-target files were in the range at all** (`common/arg.cpp` for `0001`, `tools/server/server-models.cpp` for `0008`), both far from the patched hunks, and both proven by replay rather than by reading. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `1d72b05d3` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty), the fetched `server-context.cpp` confirmed **uncast** at the emit site and `common/json.h` confirmed to carry the `is_enum` overload, Release build clean (0 errors, 0 warnings in project sources), `ctest` **559/559** (558 → 559: the re-pointed guard gained one case), `nm -D` 40 `Java_*` exports, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `mvn clean` — `nativeBuildInfoMatchesPinnedVersionConstant` confirms `LlamaCppVersion.LLAMA_CPP_VERSION` (`b11080`) against the linked `build-info`. Full `mvn test` **1772 run, 0 failures, 0 errors** (272 skipped — the model-gated classes, no GGUF in this sandbox). `test_json_helpers.cpp` is clean under the CI-pinned clang-format 23.1.1; `spotless:check` clean; SpotBugs **0** findings. Model-backed Java tests were not run (HF-blocked sandbox); `NativeServerAttachIntegrationTest.models_reportNumericVocabType` — whose failure message now names #28518 instead of the retired patch — is the CI-side confirmation that the wire value stayed numeric across this drop. | | b11080–b11103 | 23 commits, 49 files outside `tools/ui`, **301 KiB**. Backend-internal almost throughout: CUDA (#29224 sm_70 tile fix, #29155 contiguous convert, #29135 conv2d implicit GEMM, #28536 FA swizzle refactor), Metal (#29220 FA mask bounds), SYCL (#29132, #28895), OpenCL (#29055 A8 Q4_0 dp4a kernel), hexagon (#29199). `vendor/cpp-httplib` moves 0.56 → **0.57.1** (#29214, #29239). Outside ggml: `common/chat.cpp` + `common/parsers/muse-glimmer.cpp` (#29242, tool-call parser fix; #29257 MiMo-V2.6 template detection), `cmake/llama-config.cmake.in` (#29228, repeated `find_package` — not consumed here, this project uses `FetchContent`), `src/llama-context.cpp` (#26625, sched-reserve reporting), tests and the Python converter. | **No project-source change.** No header on the priority list moves (`include/`, `common/common.h`, `common/arg.h`, `tools/server/*.h` and `tools/mtmd/*.h` are byte-identical between the two tags). The server contract greps come back identical over the whole b11080–b11209 range (68 request fields, 23 bounded fields, 142 response keys), and so does the set of `common/arg.cpp` option names. | | b11080–b11103 | patches + upstream verification | **All eight patches apply unchanged**, replayed in filename order against every tag of the range, not only the endpoints. No patch-target file is in the range. All standing drop-checks still say "still required" (checked at b11209, the end of this bump series — see the b11163–b11209 row). | +| b11103–b11104 | One commit, 8 files, **23 KiB**: **#28690** ("server: Add support for binding to multiple addresses"). `common_params::hostname` (a `std::string`) becomes `hostnames` (a `std::vector`, default `{"127.0.0.1"}`); `--host` now takes a comma-separated list of IP addresses and/or `.sock` paths, blanks around the commas are stripped, and an empty list is rejected. `server_http_context` (`tools/server/server-http.h`) drops `std::thread thread`, `std::string hostname` and `std::string listening_address`, and gains `void join()`, `std::vector listening_addresses` and a private `init_listener()`; `llama_server()` now calls `ctx_http.join()` instead of joining the thread, and logs one "listening on" line per address. | **Two project changes.** (1) `patches/0007` — see the patch row. (2) `server.NativeServer` gains `getHosts()` (the parsed `--host` list, never empty) and `getHost()` now returns its first element, the same address upstream prints as the default; before this, a `--host a,b` would have been reported as the literal string `"a,b"`. Pinned by two new `NativeServerSmokeTest` cases (split + trim + skip empties; an all-separator value falls back to the default). No other project source names the removed members: `jllama.cpp` never touched `hostname`, and the Java `OpenAiCompatServer` binds through the JDK's `HttpServer` with its own single-host config. | +| b11103–b11104 | patches + upstream verification | **`0007` refreshed; it still applied cleanly, which is the trap.** Its hunks sit far from the lines #28690 changed, so `git apply` succeeded against b11104 — but two of its own `+` lines, inside `llama_server_attach`, named `ctx_http.listening_address` and `ctx_http.thread`, both of which no longer exist. The applier cannot see that class of break; only a compile can. Fixed to match upstream's own new shape: log every entry of `listening_addresses`, then `ctx_http.join()`. The other seven patches are untouched and replay clean against every tag to b11209. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index ae3f20529..6409f6ac0 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11103 + GIT_TAG b11104 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/patches/0007-server-attach-http-frontend.patch b/llama/patches/0007-server-attach-http-frontend.patch index ec80ef913..6906c1bd6 100644 --- a/llama/patches/0007-server-attach-http-frontend.patch +++ b/llama/patches/0007-server-attach-http-frontend.patch @@ -197,13 +197,13 @@ index 102692c74..8389fe6bf 100644 + ctx_http.stop(); + }; + -+ SRV_INF("attached to existing server context, listening on %s\n", ctx_http.listening_address.c_str()); -+ -+ // block until llama_server_request_shutdown() stops the HTTP thread -+ if (ctx_http.thread.joinable()) { -+ ctx_http.thread.join(); ++ for (const auto & address : ctx_http.listening_addresses) { ++ SRV_INF("attached to existing server context, listening on %s\n", address.c_str()); + } + ++ // block until llama_server_request_shutdown() stops the HTTP listener threads ++ ctx_http.join(); ++ + server_stream_session_manager_stop(); + return 0; +} diff --git a/llama/src/main/java/net/ladenthin/llama/server/NativeServer.java b/llama/src/main/java/net/ladenthin/llama/server/NativeServer.java index 8144ea018..d11a65ae8 100644 --- a/llama/src/main/java/net/ladenthin/llama/server/NativeServer.java +++ b/llama/src/main/java/net/ladenthin/llama/server/NativeServer.java @@ -4,6 +4,9 @@ package net.ladenthin.llama.server; +import java.util.ArrayList; +import java.util.Collections; +import java.util.List; import java.util.Objects; import java.util.concurrent.CountDownLatch; import java.util.concurrent.TimeUnit; @@ -190,17 +193,40 @@ public boolean isRunning() { /** * Returns the bind host parsed from the arguments ({@code --host}), or {@code 127.0.0.1} when * absent. Best-effort convenience for logging; the authoritative value is what the native server - * parsed. + * parsed. When {@code --host} lists several addresses, this is the first of them — see + * {@link #getHosts()}. * * @return the configured bind host */ public String getHost() { + return getHosts().get(0); + } + + /** + * Returns every bind address parsed from the arguments ({@code --host}), or a single + * {@code 127.0.0.1} when absent. Since llama.cpp b11104 {@code --host} takes a comma-separated + * list (IP addresses and/or UNIX socket paths ending in {@code .sock}) and the server listens on + * all of them; blanks around the commas are ignored, as upstream does. Best-effort convenience for + * logging, like {@link #getHost()}. + * + * @return the configured bind addresses, never empty + */ + public List getHosts() { for (int i = 0; i < args.length - 1; i++) { if ("--host".equals(args[i])) { - return args[i + 1]; + final List hosts = new ArrayList<>(); + for (final String host : args[i + 1].split(",", -1)) { + final String trimmed = host.trim(); + if (!trimmed.isEmpty()) { + hosts.add(trimmed); + } + } + if (!hosts.isEmpty()) { + return Collections.unmodifiableList(hosts); + } } } - return DEFAULT_HOST; + return Collections.singletonList(DEFAULT_HOST); } /** diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index bd0e98146..81d58a05f 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11103"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11104"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11103-"} — call + * plus the resolved upstream commit, e.g. {@code "b11104-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11103"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11104"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11103"; + public static final String LLAMA_CPP_VERSION = "b11104"; // Constants holder — not instantiable. private LlamaCppVersion() {} diff --git a/llama/src/test/java/net/ladenthin/llama/server/NativeServerSmokeTest.java b/llama/src/test/java/net/ladenthin/llama/server/NativeServerSmokeTest.java index 892da39e6..c994867eb 100644 --- a/llama/src/test/java/net/ladenthin/llama/server/NativeServerSmokeTest.java +++ b/llama/src/test/java/net/ladenthin/llama/server/NativeServerSmokeTest.java @@ -5,6 +5,7 @@ package net.ladenthin.llama.server; import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; import static org.hamcrest.Matchers.is; import static org.junit.jupiter.api.Assertions.assertThrows; @@ -27,6 +28,21 @@ public void parsesHostAndPortFromArgs() { assertThat(server.isRunning(), is(false)); } + @Test + public void commaSeparatedHostsAreSplitAndTrimmedAndTheFirstIsTheHost() { + // since llama.cpp b11104, --host takes a comma-separated list and binds all of them + NativeServer server = new NativeServer("-m", "m.gguf", "--host", "127.0.0.1, ::1 ,,/tmp/l.sock"); + assertThat(server.getHosts(), contains("127.0.0.1", "::1", "/tmp/l.sock")); + assertThat(server.getHost(), is("127.0.0.1")); + } + + @Test + public void hostListOfOnlySeparatorsFallsBackToDefault() { + NativeServer server = new NativeServer("-m", "m.gguf", "--host", " , "); + assertThat(server.getHosts(), contains("127.0.0.1")); + assertThat(server.getHost(), is("127.0.0.1")); + } + @Test public void shortPortFlagParsed() { NativeServer server = new NativeServer("-m", "m.gguf", "-p", "9099"); @@ -37,6 +53,7 @@ public void shortPortFlagParsed() { public void defaultsWhenFlagsAbsent() { NativeServer server = new NativeServer("-m", "m.gguf"); assertThat(server.getHost(), is("127.0.0.1")); + assertThat(server.getHosts(), contains("127.0.0.1")); assertThat(server.getPort(), is(8080)); } From a243ed6e909085c70ec906385065197d349884d5 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:00:32 +0000 Subject: [PATCH 03/10] Upgrade llama.cpp from b11104 to b11160 56 upstream commits, compatible from this project's side: backend kernels, the llama.cpp 0.5.0 version bump, router fixes, new model support (Ling 3.0 VL, Gemma 4 DSpark draft) and two OAI-layer features -- input_image as a Responses function_call_output and video_url as an alias of input_video. Only a private member of server-context.h moves; all eight patches apply at every tag. The two OAI features are recorded in the history row as Java-API candidates (ContentPart has no video part; ResponsesApiSupport flattens function_call_output to text), not changed here. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 1f3601b3c..dd265f045 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11104** +Current llama.cpp pinned version: **b11160** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11104 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11160 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11104`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11160`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1655,7 +1655,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11104`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11160`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 4aac4aa6d..c63f485d9 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11104](https://img.shields.io/badge/llama.cpp-%23b11104-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11104) +[![llama.cpp b11160](https://img.shields.io/badge/llama.cpp-%23b11160-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11160) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 6fda5e72b..efeed1de2 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -748,3 +748,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11080–b11103 | patches + upstream verification | **All eight patches apply unchanged**, replayed in filename order against every tag of the range, not only the endpoints. No patch-target file is in the range. All standing drop-checks still say "still required" (checked at b11209, the end of this bump series — see the b11163–b11209 row). | | b11103–b11104 | One commit, 8 files, **23 KiB**: **#28690** ("server: Add support for binding to multiple addresses"). `common_params::hostname` (a `std::string`) becomes `hostnames` (a `std::vector`, default `{"127.0.0.1"}`); `--host` now takes a comma-separated list of IP addresses and/or `.sock` paths, blanks around the commas are stripped, and an empty list is rejected. `server_http_context` (`tools/server/server-http.h`) drops `std::thread thread`, `std::string hostname` and `std::string listening_address`, and gains `void join()`, `std::vector listening_addresses` and a private `init_listener()`; `llama_server()` now calls `ctx_http.join()` instead of joining the thread, and logs one "listening on" line per address. | **Two project changes.** (1) `patches/0007` — see the patch row. (2) `server.NativeServer` gains `getHosts()` (the parsed `--host` list, never empty) and `getHost()` now returns its first element, the same address upstream prints as the default; before this, a `--host a,b` would have been reported as the literal string `"a,b"`. Pinned by two new `NativeServerSmokeTest` cases (split + trim + skip empties; an all-separator value falls back to the default). No other project source names the removed members: `jllama.cpp` never touched `hostname`, and the Java `OpenAiCompatServer` binds through the JDK's `HttpServer` with its own single-host config. | | b11103–b11104 | patches + upstream verification | **`0007` refreshed; it still applied cleanly, which is the trap.** Its hunks sit far from the lines #28690 changed, so `git apply` succeeded against b11104 — but two of its own `+` lines, inside `llama_server_attach`, named `ctx_http.listening_address` and `ctx_http.thread`, both of which no longer exist. The applier cannot see that class of break; only a compile can. Fixed to match upstream's own new shape: log every entry of `listening_addresses`, then `ctx_http.join()`. The other seven patches are untouched and replay clean against every tag to b11209. | +| b11104–b11160 | 56 commits, 93 files outside `tools/ui`, **4169 KiB** — nearly all of it backend kernels and generated shader/template sources (Vulkan IQ4_XS MMQ/MMV #28415, Intel Xe FA #24406, AMD int8 coopmat #27952; CUDA conv3d implicit GEMM #29137, top-k MoE #28432; SYCL, OpenCL, hexagon, Metal). llama.cpp's own version goes **0.4.1 → 0.5.0** (#29333) and ggml's to 0.25.1 — no ABI or link-target change for this project. Server-side, all behavioural: `server-chat.cpp` accepts `input_image` as a Responses-API `function_call_output` (**#22575**), `server-common.cpp` accepts OpenAI's `video_url` content part as an alias of `input_video` (**#27921**), the token-count endpoint no longer crashes while the server sleeps (**#29309**, which also narrows the private `handle_count_tokens` signature in `server-context.h`), and the router fixes eviction races (#29217), stops passing its log file to children (#29212), lets a preset set one (#29334) and dedups a draft HF model via `dedup-cache-models` (#27934). New model support arrives with no API: Ling 3.0 VL (#29151, adds `tools/mtmd/models/ling3vl.cpp`), Gemma 4 DSpark draft backbone (#29226), DFlash vision targets (#29339). Jinja: unary `+`/`-` parsing (#29244). | **No project-source change.** The only priority-list header that moves is `server-context.h`, and only in a private member. `ling3vl.cpp` is compiled by upstream's own `mtmd` target, which this project links, so the new source file needs no entry here. The three server-contract greps are unchanged (the new behaviour lives in the OAI translation layer, not the schema); the two OAI-layer features are candidates for the Java API — `ContentPart` has no video part yet, and `ResponsesApiSupport` flattens a `function_call_output` to text, so an `input_image` output is dropped before it reaches the native layer. Recorded, not changed in this bump. | +| b11104–b11160 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range (`0007` in its b11104 form). Patch-target files touched: `tools/server/server-context.cpp` (#29309, #29325 — `0002`/`0003`) and `tools/server/server-models.cpp` (#29212, #29217, #27934, #29339 — `0008`), in each case away from the patched hunks. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 6409f6ac0..debde451f 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11104 + GIT_TAG b11160 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 81d58a05f..8ff1ccd9f 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11104"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11160"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11104-"} — call + * plus the resolved upstream commit, e.g. {@code "b11160-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11104"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11160"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11104"; + public static final String LLAMA_CPP_VERSION = "b11160"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 2f7fbb809e614fbc4b8f905a5d5f0513bef6a976 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:00:57 +0000 Subject: [PATCH 04/10] Upgrade llama.cpp from b11160 to b11163: llama_batch_ext upstream #24669 adds the extended batch API (llama_batch_ext + llama_process) next to the classic llama_batch/llama_decode, which is unchanged. Purely additive for this project: nothing in src/main/cpp calls the decode or batch helpers directly. Kept as its own step because it is the one new public llama.h surface in the b11080 -> b11209 series. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index dd265f045..8a6261ba3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11160** +Current llama.cpp pinned version: **b11163** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11160 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11163 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11160`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11163`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1655,7 +1655,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11160`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11163`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index c63f485d9..39d795a30 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11160](https://img.shields.io/badge/llama.cpp-%23b11160-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11160) +[![llama.cpp b11163](https://img.shields.io/badge/llama.cpp-%23b11163-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11163) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index efeed1de2..d211c75ea 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -750,3 +750,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11103–b11104 | patches + upstream verification | **`0007` refreshed; it still applied cleanly, which is the trap.** Its hunks sit far from the lines #28690 changed, so `git apply` succeeded against b11104 — but two of its own `+` lines, inside `llama_server_attach`, named `ctx_http.listening_address` and `ctx_http.thread`, both of which no longer exist. The applier cannot see that class of break; only a compile can. Fixed to match upstream's own new shape: log every entry of `listening_addresses`, then `ctx_http.join()`. The other seven patches are untouched and replay clean against every tag to b11209. | | b11104–b11160 | 56 commits, 93 files outside `tools/ui`, **4169 KiB** — nearly all of it backend kernels and generated shader/template sources (Vulkan IQ4_XS MMQ/MMV #28415, Intel Xe FA #24406, AMD int8 coopmat #27952; CUDA conv3d implicit GEMM #29137, top-k MoE #28432; SYCL, OpenCL, hexagon, Metal). llama.cpp's own version goes **0.4.1 → 0.5.0** (#29333) and ggml's to 0.25.1 — no ABI or link-target change for this project. Server-side, all behavioural: `server-chat.cpp` accepts `input_image` as a Responses-API `function_call_output` (**#22575**), `server-common.cpp` accepts OpenAI's `video_url` content part as an alias of `input_video` (**#27921**), the token-count endpoint no longer crashes while the server sleeps (**#29309**, which also narrows the private `handle_count_tokens` signature in `server-context.h`), and the router fixes eviction races (#29217), stops passing its log file to children (#29212), lets a preset set one (#29334) and dedups a draft HF model via `dedup-cache-models` (#27934). New model support arrives with no API: Ling 3.0 VL (#29151, adds `tools/mtmd/models/ling3vl.cpp`), Gemma 4 DSpark draft backbone (#29226), DFlash vision targets (#29339). Jinja: unary `+`/`-` parsing (#29244). | **No project-source change.** The only priority-list header that moves is `server-context.h`, and only in a private member. `ling3vl.cpp` is compiled by upstream's own `mtmd` target, which this project links, so the new source file needs no entry here. The three server-contract greps are unchanged (the new behaviour lives in the OAI translation layer, not the schema); the two OAI-layer features are candidates for the Java API — `ContentPart` has no video part yet, and `ResponsesApiSupport` flattens a `function_call_output` to text, so an `input_image` output is dropped before it reaches the native layer. Recorded, not changed in this bump. | | b11104–b11160 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range (`0007` in its b11104 form). Patch-target files touched: `tools/server/server-context.cpp` (#29309, #29325 — `0002`/`0003`) and `tools/server/server-models.cpp` (#29212, #29217, #27934, #29339 — `0008`), in each case away from the patched hunks. | +| b11160–b11163 | Three commits, 12 files, **78 KiB**; one matters: **#24669** ("llama: add llama_batch_ext"). `include/llama.h` gains an opaque `llama_batch_ext` with `_init`/`_free`/`_clear`, `_add`/`_add_token`/`_add_embd`/`_add_seq`, setters for token/state embeddings, output flags and (multi-)positions, a `llama_embd {data, n_rows, n_embd}` view, and `llama_process(ctx, LLAMA_PROCESS_TYPE_{ENCODE,DECODE}, batch)`, which returns what `llama_decode` returns. `llama-cpp.h` adds `llama_batch_ext_ptr`; `common.h` adds `common_batch_ext_get_one()` and moves `common_prompt_batch_decode` onto the new API (its parameter type is now spelled `llama_tokens`, the same `std::vector`). The classic `llama_batch` / `llama_decode` API is **unchanged**. The other two commits are upstream CI pytest-worker settings. | **No project-source change — the addition is purely additive.** Nothing in `src/main/cpp` calls `llama_decode`, `llama_batch_get_one`, `common_batch_add` or `common_prompt_batch_decode` (inference goes through the server library, TTS through `mtmd_helper::gen_audio`), so there is nothing to migrate. Not a Java-API candidate either: the batch API is a lower-level decode path than anything this library exposes, and the server library already chooses its own. Worth watching on later bumps — if upstream moves the server onto `llama_process`, that shows up in `server-context.cpp`, which `0002`/`0003` patch. | +| b11160–b11163 | patches + upstream verification | **All eight patches apply unchanged.** No patch-target file is in the range. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index debde451f..81dca2a1e 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11160 + GIT_TAG b11163 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 8ff1ccd9f..5cd681c0b 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11160"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11163"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11160-"} — call + * plus the resolved upstream commit, e.g. {@code "b11163-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11160"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11163"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11160"; + public static final String LLAMA_CPP_VERSION = "b11163"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 49932e3f769174edd4eb04049d23e6bc713dc85a Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:06:54 +0000 Subject: [PATCH 05/10] Upgrade llama.cpp from b11163 to b11209 Last step of the b11080 -> b11209 series. 46 upstream commits, compatible from this project's side: backend kernels, a GGUF-driven W4A4 precision policy with no public API, shared Unicode helpers in common.h, cpp-httplib 0.58.0, cleanup after a failed state restore, a grammar token_id fix, and a revert of the --fit context-length change. All eight patches are still required. Verified at b11209 from a fresh configure: patches applied (8, stamp matches), Release build with 0 warnings, ctest 559/559, 40 Java_* exports, and mvn clean verify 1774 run / 0 failures with NativeLibraryLoadSmokeTest 4/4 confirming the pin against the linked build-info. The CHANGELOG entry covers the whole series. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CHANGELOG.md | 13 +++++++++++++ CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 6 files changed, 25 insertions(+), 10 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7920ca88a..9aae31fe9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -101,6 +101,19 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by where the backend cannot provide it, `OFF` disables it. ### Changed +- **llama.cpp `b11080` → `b11209`, in five steps sized by what they change here.** 129 upstream + commits. Four steps are version-only from this project's side (`b11103`, `b11160`, `b11163`, + `b11209`); the one incompatible change got a step of its own: **`b11104`** (llama.cpp #28690) + lets `--host` take a comma-separated list of addresses and removed `server_http_context::thread` + and `::listening_address`. `patches/0007` still applied cleanly there but named both members, so + it was refreshed to upstream's new `join()` / `listening_addresses` shape; and + `NativeServer` gains **`getHosts()`**, with `getHost()` now returning the first address instead + of the raw comma-separated string. New upstream surface that needs no project change: the + extended batch API (`llama_batch_ext` + `llama_process`, #24669, additive), `input_image` accepted + as a Responses-API `function_call_output` and OpenAI's `video_url` as an alias of `input_video` + on the native server (#22575, #27921), cleanup of K/V and recurrent state after a failed state + restore (#27530), and new model support (Ling 3.0 VL, Gemma 4 DSpark draft, MiMo-V2.6). + cpp-httplib moves to 0.58.0. All eight local patches are still required. - **llama.cpp `b11069` → `b11080`, and local patch `0010` dropped — upstream fixed the enum-to-JSON-boolean trap at its root.** Eleven upstream commits, 1244 KiB, no project-source change. The size is one commit that does not concern this project (llama.cpp #29197 rewrites 46 diff --git a/CLAUDE.md b/CLAUDE.md index 8a6261ba3..1d685dc1f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11163** +Current llama.cpp pinned version: **b11209** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11163 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11209 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11163`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11209`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1655,7 +1655,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11163`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11209`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 39d795a30..b9500f5cb 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11163](https://img.shields.io/badge/llama.cpp-%23b11163-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11163) +[![llama.cpp b11209](https://img.shields.io/badge/llama.cpp-%23b11209-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11209) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index d211c75ea..9861fd7f3 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -752,3 +752,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11104–b11160 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range (`0007` in its b11104 form). Patch-target files touched: `tools/server/server-context.cpp` (#29309, #29325 — `0002`/`0003`) and `tools/server/server-models.cpp` (#29212, #29217, #27934, #29339 — `0008`), in each case away from the patched hunks. | | b11160–b11163 | Three commits, 12 files, **78 KiB**; one matters: **#24669** ("llama: add llama_batch_ext"). `include/llama.h` gains an opaque `llama_batch_ext` with `_init`/`_free`/`_clear`, `_add`/`_add_token`/`_add_embd`/`_add_seq`, setters for token/state embeddings, output flags and (multi-)positions, a `llama_embd {data, n_rows, n_embd}` view, and `llama_process(ctx, LLAMA_PROCESS_TYPE_{ENCODE,DECODE}, batch)`, which returns what `llama_decode` returns. `llama-cpp.h` adds `llama_batch_ext_ptr`; `common.h` adds `common_batch_ext_get_one()` and moves `common_prompt_batch_decode` onto the new API (its parameter type is now spelled `llama_tokens`, the same `std::vector`). The classic `llama_batch` / `llama_decode` API is **unchanged**. The other two commits are upstream CI pytest-worker settings. | **No project-source change — the addition is purely additive.** Nothing in `src/main/cpp` calls `llama_decode`, `llama_batch_get_one`, `common_batch_add` or `common_prompt_batch_decode` (inference goes through the server library, TTS through `mtmd_helper::gen_audio`), so there is nothing to migrate. Not a Java-API candidate either: the batch API is a lower-level decode path than anything this library exposes, and the server library already chooses its own. Worth watching on later bumps — if upstream moves the server onto `llama_process`, that shows up in `server-context.cpp`, which `0002`/`0003` patch. | | b11160–b11163 | patches + upstream verification | **All eight patches apply unchanged.** No patch-target file is in the range. | +| b11163–b11209 | 46 commits, 157 files outside `tools/ui`, **5339 KiB** — again dominated by backend kernels (ggml-cpu tiled mul_mat for k-quants #27851, Metal per-dtype FA libraries #29329, CUDA RMS_NORM+SCALE fusion, SYCL sparse FA, OpenCL, hexagon backend sampler). **#24364** adds a model-driven W4A4 precision policy (`llama_prec_policy`), selected from GGUF metadata inside `src/` — no public `llama.h` surface. `common/common.h` gains shared Unicode helpers (**#29415**: `fs_path_to_utf8`, and `utf8_to_wstring`/`wstring_to_utf8` on Windows) and pulls in ``; `common/common.cpp` simplifies `fs_create_directory_with_parents` (#29432). Behavioural fixes worth knowing: K/V and recurrent state are cleaned up after a **failed state restore** (**#27530**, the path behind session load), grammar `token_id` parsing no longer truncates large ids (#29382), fused-QKV tensor split with uneven K/V heads (#29294), and #29437 **reverts** #28849's change to the max context length chosen by `--fit` with a unified KV cache. `vendor/cpp-httplib` moves to **0.58.0** (#29407). Jinja gains `sameas` and argument-taking test statements (#29448, #29443). | **No project-source change.** `common.h`'s additions are new free functions and an include; nothing in `src/main/cpp` names the renamed or new helpers. `tools/mtmd/mtmd-audio.h` and `tools/server/server-common.h` (a `#ifndef _WIN32` around `wake_fd`, #29479) change in ways no project file reaches. The three server-contract greps and the `arg.cpp` option set are unchanged over the whole series. | +| b11163–b11209 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range. Patch-target file touched: `src/llama-model.cpp` (#29294, and #29151 earlier in the series — `0012`), away from the split-normalisation block. **Standing drop-checks at pristine b11209, all "still required":** `0001` (`common_params_parse_main` 0 occurrences in `common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`, so [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1582`), `0014` (`common_log_set_callback` 0 occurrences in `common/log.h`); `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `187664b53` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty), Release build clean with **0 warnings** — which is what proves the `0007` refresh of the b11103–b11104 step, since only a compile can — `ctest` **559/559**, `nm -D` 40 `Java_*` exports, and `mvn clean verify` **1774 run, 0 failures, 0 errors** (272 skipped, the model-gated classes; no GGUF in this sandbox) with `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (`nativeBuildInfoMatchesPinnedVersionConstant` confirms `b11209` against the linked `build-info`), `LlamaLoggerTest` 6/6 against the real library, spotless and SpotBugs clean. Model-backed Java tests were not run (HF-blocked sandbox). The four intermediate pins were checked by patch replay and header diff, not by a build of their own. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 81dca2a1e..1b6d7a00a 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11163 + GIT_TAG b11209 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 5cd681c0b..8adab355b 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11163"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11209"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11163-"} — call + * plus the resolved upstream commit, e.g. {@code "b11209-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11163"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11209"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11163"; + public static final String LLAMA_CPP_VERSION = "b11209"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 9907dad0239c6b4d7b9a0d62a91149c0a62dfcb5 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:24:25 +0000 Subject: [PATCH 06/10] Build the ROCm classifiers against ROCm 10 from TheRock AMD now builds and releases ROCm through TheRock. Both rocm-* jobs install its Python wheels (rocm[libraries,devel]==10.0.0 from stable.repo.amd.com/rocm/whl-next) the way upstream llama.cpp's own ubuntu-rocm / windows-rocm release jobs do at b11209, replacing the repo.radeon.com 6.3.4 apt repo on Linux and the HIP SDK 26.Q1 installer on Windows. Paths are read back with rocm-sdk path. Linux now uses CMake's native HIP language with ROCm's clang as the HIP compiler (upstream's form); Windows uses the TheRock clang under lib\llvm\bin. The GPU target lists are copied from upstream: newer architectures are added, gfx900/gfx906 are dropped because llama.cpp no longer ships them and TheRock never marks them release-ready. Not validated locally: stable.repo.amd.com is unreachable from the sandbox, so the first CI run is the proof. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- .github/workflows/publish.yml | 83 +++++++++++++++++++++++++---------- CHANGELOG.md | 7 +++ CLAUDE.md | 14 +++++- README.md | 4 +- 4 files changed, 81 insertions(+), 27 deletions(-) diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 26b620fb4..70f117d54 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -2020,6 +2020,12 @@ jobs: SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: + - name: Free up disk space + # The TheRock wheels unpack to several GB; upstream's ubuntu-rocm job frees the runner + # the same way. Runs first, before setup-java puts the JDK into the tool cache. + uses: ggml-org/free-disk-space@v1.3.1 + with: + tool-cache: true - uses: actions/checkout@v7 - name: Download shared WebUI assets uses: actions/download-artifact@v8 @@ -2030,21 +2036,36 @@ jobs: with: distribution: 'temurin' java-version: ${{ env.JAVA_VERSION }} - - name: Install ROCm/HIP (AMD apt repo) + - name: Install ROCm/HIP (TheRock wheels) + # KEEP IN SYNC WITH UPSTREAM: ROCM_VERSION and the GPU target list track the ubuntu-rocm + # job in llama.cpp's .github/workflows/release.yml at the pinned GIT_TAG. Since ROCm 7.14 + # AMD builds and releases ROCm through TheRock (https://github.com/ROCm/TheRock), replacing + # the monolithic releases behind the repo.radeon.com apt repo this job used before (it was + # pinned at 6.3.4). The wheels carry the HIP runtime + CMake configs ("libraries") and the + # compilers/headers ("devel"); rocm-sdk reports where they landed. + env: + ROCM_VERSION: "10.0.0" run: | - sudo mkdir --parents --mode=0755 /etc/apt/keyrings - wget -qO- https://repo.radeon.com/rocm/rocm.gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/rocm.gpg > /dev/null - echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/rocm.gpg] https://repo.radeon.com/rocm/apt/6.3.4 noble main" | sudo tee /etc/apt/sources.list.d/rocm.list - printf 'Package: *\nPin: release o=repo.radeon.com\nPin-Priority: 600\n' | sudo tee /etc/apt/preferences.d/rocm-pin-600 - sudo apt-get update - sudo apt-get install -y rocm-hip-sdk rocblas-dev hipblas-dev - echo "/opt/rocm/bin" >> "$GITHUB_PATH" - echo "ROCM_PATH=/opt/rocm" >> "$GITHUB_ENV" + python3 -m venv "$RUNNER_TEMP/rocm-venv" + source "$RUNNER_TEMP/rocm-venv/bin/activate" + python -m pip install --upgrade pip + python -m pip install --index-url https://stable.repo.amd.com/rocm/whl-next/ "rocm[libraries,devel]==${ROCM_VERSION}" + ROCM_PATH=$(rocm-sdk path --root) + echo "ROCM_PATH=$ROCM_PATH" >> "$GITHUB_ENV" + echo "HIP_PATH=$ROCM_PATH" >> "$GITHUB_ENV" + echo "CMAKE_PREFIX_PATH=$(rocm-sdk path --cmake)" >> "$GITHUB_ENV" + echo "LD_LIBRARY_PATH=$ROCM_PATH/lib:${LD_LIBRARY_PATH:-}" >> "$GITHUB_ENV" + echo "$(rocm-sdk path --bin)" >> "$GITHUB_PATH" + echo "$RUNNER_TEMP/rocm-venv/bin" >> "$GITHUB_PATH" - name: Build libraries shell: bash + # Native CMake HIP language (upstream's form): the HIP compiler is ROCm's clang, the C/C++ + # TUs stay on the runner's gcc. The target list is upstream's; architectures llama.cpp no + # longer ships (gfx900/gfx906 — "build passing" only in TheRock, never release-ready) are + # not carried here. run: | mvn --no-transfer-progress -f llama/pom.xml compile - .github/build.sh "-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx900;gfx906;gfx908;gfx90a;gfx1030;gfx1100;gfx1101;gfx1102 -DCMAKE_C_COMPILER=/opt/rocm/llvm/bin/clang -DCMAKE_CXX_COMPILER=/opt/rocm/llvm/bin/clang++ -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" + .github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx908;gfx90a;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" - name: Upload artifacts uses: actions/upload-artifact@v7 with: @@ -2059,7 +2080,7 @@ jobs: # HIP clang headers (__clang_hip_cmath.h) cannot overload the __host__ __device__ # isgreater/isless/... that the very new MSVC declares via _CLANG_BUILTIN2, so the # device-code compile fails. Upstream llama.cpp builds win-hip on windows-2022 for the same - # reason (it drives the HIP SDK's own clang and relies on the older MSVC STL). + # reason (it drives ROCm's own clang and relies on the older MSVC STL). runs-on: windows-2022 steps: - uses: actions/checkout@v7 @@ -2072,23 +2093,39 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: x64 - - name: Install AMD HIP SDK for Windows + - name: Install ROCm/HIP (TheRock wheels) shell: pwsh - # Mirrors upstream llama.cpp's windows-hip release job: HIP SDK 26.Q1, then - # resolve HIP_PATH from the installed ROCm dir and point the compilers + - # CMAKE_PREFIX_PATH at it so ggml-hip's find_package(hip) resolves. + # KEEP IN SYNC WITH UPSTREAM: mirrors llama.cpp's windows-rocm release job + # (.github/actions/windows-setup-rocm at the pinned GIT_TAG) — the same TheRock wheels as + # the Linux job, in place of the former AMD-Software-PRO-Edition HIP SDK installer. The + # venv lives under C:\TheRock\build, the path upstream uses. + env: + ROCM_VERSION: "10.0.0" run: | - $url = "https://download.amd.com/developer/eula/rocm-hub/AMD-Software-PRO-Edition-26.Q1-Win11-For-HIP.exe" - Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\rocm-install.exe" - $proc = Start-Process "$env:RUNNER_TEMP\rocm-install.exe" -ArgumentList '-install' -NoNewWindow -PassThru -Wait - if ($proc.ExitCode -ne 0) { Write-Error "HIP SDK install failed with exit code $($proc.ExitCode)"; exit 1 } - $hip = $(Resolve-Path 'C:\Program Files\AMD\ROCm\*\bin\clang.exe' | Split-Path | Split-Path) - "HIP_PATH=$hip" | Out-File -FilePath $env:GITHUB_ENV -Append - "$hip\bin" | Out-File -FilePath $env:GITHUB_PATH -Append + $ErrorActionPreference = "Stop" + New-Item -Path "C:\TheRock\build" -ItemType Directory -Force | Out-Null + python -m venv C:\TheRock\build\.venv + & C:\TheRock\build\.venv\Scripts\Activate.ps1 + python -m pip install --upgrade pip + python -m pip install --index-url https://stable.repo.amd.com/rocm/whl-next/ "rocm[libraries,devel]==$env:ROCM_VERSION" + if ($LASTEXITCODE -ne 0) { throw "ROCm wheel install failed with exit code $LASTEXITCODE" } + rocm-sdk init + if ($LASTEXITCODE -ne 0) { throw "rocm-sdk init failed with exit code $LASTEXITCODE" } + $rocm = (rocm-sdk path --root) + if (-not $rocm) { throw "rocm-sdk path --root returned empty - devel package may not be installed" } + $rocm = $rocm.Trim() + "HIP_PATH=$rocm" | Out-File -FilePath $env:GITHUB_ENV -Append + "HIP_DEVICE_LIB_PATH=$rocm\lib\llvm\amdgcn\bitcode" | Out-File -FilePath $env:GITHUB_ENV -Append + "HIP_PLATFORM=amd" | Out-File -FilePath $env:GITHUB_ENV -Append + "LLVM_PATH=$rocm\lib\llvm" | Out-File -FilePath $env:GITHUB_ENV -Append + (rocm-sdk path --bin).Trim() | Out-File -FilePath $env:GITHUB_PATH -Append + "C:\TheRock\build\.venv\Scripts" | Out-File -FilePath $env:GITHUB_PATH -Append - name: Build libraries shell: cmd + # Upstream's compiler wiring for TheRock (clang under lib\llvm\bin, not bin\) and its + # GPU target list; -Wno-error=incompatible-pointer-types is upstream's too. run: | - .github\build.bat -G "Ninja Multi-Config" -DGGML_HIP=ON -DGPU_TARGETS=gfx1030;gfx1100;gfx1101;gfx1102 -DCMAKE_PREFIX_PATH="%HIP_PATH%" -DCMAKE_C_COMPILER="%HIP_PATH%\bin\clang.exe" -DCMAKE_CXX_COMPILER="%HIP_PATH%\bin\clang++.exe" -DOS_NAME=Windows -DOS_ARCH=x86_64 + .github\build.bat -G "Ninja Multi-Config" -DGGML_HIP=ON -DGPU_TARGETS=gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DCMAKE_PREFIX_PATH="%HIP_PATH%" -DHIP_PATH="%HIP_PATH%" -DCMAKE_C_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_CXX_COMPILER="%HIP_PATH%\lib\llvm\bin\clang++.exe" -DCMAKE_HIP_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_C_FLAGS="-Wno-error=incompatible-pointer-types" -DOS_NAME=Windows -DOS_ARCH=x86_64 - name: Upload artifacts uses: actions/upload-artifact@v7 with: diff --git a/CHANGELOG.md b/CHANGELOG.md index 9aae31fe9..657dbf537 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -114,6 +114,13 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by on the native server (#22575, #27921), cleanup of K/V and recurrent state after a failed state restore (#27530), and new model support (Ling 3.0 VL, Gemma 4 DSpark draft, MiMo-V2.6). cpp-httplib moves to 0.58.0. All eight local patches are still required. +- **BREAKING (runtime): the ROCm classifiers are built against ROCm 10.0 (TheRock), not 6.3.** + AMD now releases ROCm through [TheRock](https://github.com/ROCm/TheRock); both `rocm-*` jobs install + its wheels the way upstream llama.cpp's own release jobs do, replacing the `repo.radeon.com` 6.3.4 apt + repo (Linux) and the HIP SDK 26.Q1 installer (Windows). The GPU target lists now follow upstream too: + newer architectures are added (Linux: gfx942/gfx950, RDNA1–RDNA4 incl. gfx1150–1152 and + gfx1200/1201; Windows: RDNA1–RDNA4), and gfx900/gfx906 are dropped — llama.cpp no longer ships them + and TheRock never marks them release-ready. Consumers need a ROCm 10 runtime. - **llama.cpp `b11069` → `b11080`, and local patch `0010` dropped — upstream fixed the enum-to-JSON-boolean trap at its root.** Eleven upstream commits, 1244 KiB, no project-source change. The size is one commit that does not concern this project (llama.cpp #29197 rewrites 46 diff --git a/CLAUDE.md b/CLAUDE.md index 1d685dc1f..b18636dbb 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -322,8 +322,8 @@ matching GPU) and bundle **no** vendor runtime. | Classifier | GGML flag(s) | Job runner / toolchain | Tree | |---|---|---|---| -| `rocm-linux-x86-64` | `GGML_HIP=ON -DAMDGPU_TARGETS=…` | `ubuntu-latest` + ROCm apt repo (`/opt/rocm/llvm/bin/clang`) | `resources_linux_rocm` | -| `rocm-windows-x86-64` | `GGML_HIP=ON` | `windows-2025-vs2026` + AMD HIP SDK | `resources_windows_rocm` | +| `rocm-linux-x86-64` | `GGML_HIP=ON -DCMAKE_HIP_COMPILER=… -DGPU_TARGETS=…` | `ubuntu-latest` + ROCm 10 TheRock wheels (pip, `rocm-sdk path`) | `resources_linux_rocm` | +| `rocm-windows-x86-64` | `GGML_HIP=ON` | `windows-2022` + ROCm 10 TheRock wheels (pip) | `resources_windows_rocm` | | `sycl-fp16-linux-x86-64` | `GGML_SYCL=ON -DGGML_SYCL_F16=ON` (`icx`/`icpx`) | `ubuntu-latest` + Intel oneAPI apt | `resources_linux_sycl_fp16` | | `sycl-fp32-linux-x86-64` | `GGML_SYCL=ON` (`icx`/`icpx`) | `ubuntu-latest` + Intel oneAPI apt | `resources_linux_sycl_fp32` | | `sycl-windows-x86-64` | `GGML_SYCL=ON` (`icx`) | `windows-2025-vs2026` + oneAPI installer | `resources_windows_sycl` | @@ -331,6 +331,16 @@ matching GPU) and bundle **no** vendor runtime. | `openvino-linux-x86-64` | `GGML_OPENVINO=ON` | `ubuntu-latest` + OpenVINO apt | `resources_linux_openvino` | | `openvino-windows-x86-64` | `GGML_OPENVINO=ON` | `windows-2025-vs2026` + OpenVINO archive | `resources_windows_openvino` | +**ROCm comes from TheRock, and the version and GPU targets follow upstream.** Since ROCm 7.14 AMD +builds and releases ROCm through [TheRock](https://github.com/ROCm/TheRock); both ROCm jobs install +its Python wheels (`rocm[libraries,devel]` from `stable.repo.amd.com/rocm/whl-next/`) exactly as +llama.cpp's own `ubuntu-rocm` / `windows-rocm` release jobs do, and read the paths back with +`rocm-sdk path`. The ROCm version and the `GPU_TARGETS` list are copied from upstream's +`release.yml` at the pinned `GIT_TAG` — **re-check both on every llama.cpp bump**. Architectures +upstream no longer builds are deliberately **not** carried along (gfx900/gfx906 were dropped at the +switch from the old 6.3.4 apt repo: TheRock marks them "build passing" only, never release-ready): +supporting hardware llama.cpp itself does not ship for is not worth holding back a newer toolchain. + Two routing notes mirror existing precedent: **Linux SYCL** ships two precision variants at the *same* arch, so `CMakeLists.txt` routes them to two *distinct* trees by `GGML_SYCL_F16` (fp16 vs fp32). **Windows OpenCL** now holds both `x86_64` (desktop ICD) and `aarch64` (Snapdragon/Adreno) in the one diff --git a/README.md b/README.md index b9500f5cb..905902b14 100644 --- a/README.md +++ b/README.md @@ -191,8 +191,8 @@ exclusive — and optionally a CPU Windows build. | `vulkan-linux-x86-64` | Vulkan | Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) | A Vulkan runtime (`libvulkan.so.1`), which current GPU drivers install. No Vulkan SDK is needed at runtime. The most portable Linux GPU option (vendor-independent, no CUDA toolkit). Built natively on `ubuntu-latest`, so it shares the aarch64 build's higher glibc floor (≈ 2.39). | | `vulkan-linux-aarch64` | Vulkan | Linux aarch64 with a Vulkan 1.2+ GPU | A Vulkan runtime (`libvulkan.so.1`) from the device/driver. glibc ≥ 2.39 (built on `ubuntu-24.04-arm`). | | `opencl-android-aarch64` | OpenCL (Adreno) | Android aarch64 with Qualcomm Adreno GPU | A device-supplied OpenCL ICD (`libOpenCL.so`). Devices without an ICD (e.g. most non-Snapdragon Android hardware) must use the default CPU JAR. | -| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host. Not bundled; native load fails without it. No CPU fallback. | -| `rocm-windows-x86-64` | ROCm / HIP | Windows x86-64 with AMD GPU | The AMD HIP SDK runtime DLLs (`amdhip64.dll`, `rocblas.dll`, `hipblas.dll`) on `PATH`. Not bundled. No CPU fallback. | +| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm **10** runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host — built against ROCm 10.0 (TheRock), the same version upstream llama.cpp ships; GPUs older than gfx908 (e.g. gfx900/gfx906) are not targeted. Not bundled; native load fails without it. No CPU fallback. | +| `rocm-windows-x86-64` | ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm **10** runtime DLLs (`amdhip64.dll`, `rocblas.dll`, `hipblas.dll`) on `PATH` — built against ROCm 10.0 (TheRock); RDNA1 (gfx1010) and newer. Not bundled. No CPU fallback. | | `sycl-fp16-linux-x86-64` | SYCL (Intel oneAPI, fp16) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp16 accumulation (faster, slightly lower precision). Not bundled. | | `sycl-fp32-linux-x86-64` | SYCL (Intel oneAPI, fp32) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp32 accumulation (higher precision). Not bundled. | | `sycl-windows-x86-64` | SYCL (Intel oneAPI) | Windows x86-64 with Intel GPU (Arc / iGPU) | The Intel oneAPI / Level-Zero runtime DLLs on `PATH`. Not bundled. | From d9c21cd57fe33f4c295e5e4f3c67e7112b34ade0 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:26:57 +0000 Subject: [PATCH 07/10] CUDA 13.4, gfx900/gfx906 kept, free-disk-space on heavy Linux jobs CUDA 13.3 -> 13.4, matching upstream llama.cpp. Linux installs cuda-toolkit-13-4 from NVIDIA's rhel8 repo. Windows drops Jimver/cuda-toolkit, which has no 13.4 in any release or on master, and assembles the toolkit from NVIDIA's per-component redist archives with upstream's windows-setup-cuda component list. Classifiers stay cuda13-*. ROCm Linux keeps gfx900/gfx906 on top of upstream's target list: TheRock 10 still builds them. They stay only as long as they build without patches and do not hold back a newer ROCm. ggml-org/free-disk-space now runs first in the CUDA, ROCm and both SYCL Linux jobs, and replaces the hand-rolled toolchain removal in the two emulator jobs (with android, large-packages, tool-cache and swap kept). The dockcross wrappers' help text named the tags they were first generated from; it now names the pinned 20260712-79e54f9 image. Not validated locally: the NVIDIA and AMD package servers are unreachable from the sandbox, so the first CI run is the proof. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- .github/build_cuda_linux.sh | 6 +- .github/dockcross/dockcross-android-arm | 6 +- .github/dockcross/dockcross-android-arm64 | 6 +- .github/dockcross/dockcross-android-x86_64 | 6 +- .github/dockcross/dockcross-linux-arm64-lts | 6 +- .github/dockcross/dockcross-manylinux2014-x64 | 6 +- .../dockcross/dockcross-manylinux_2_28-x64 | 6 +- .github/workflows/publish.yml | 102 ++++++++++++++---- CHANGELOG.md | 8 +- CLAUDE.md | 56 +++++----- README.md | 2 +- 11 files changed, 141 insertions(+), 69 deletions(-) diff --git a/.github/build_cuda_linux.sh b/.github/build_cuda_linux.sh index e2c39c03e..66aad22fb 100755 --- a/.github/build_cuda_linux.sh +++ b/.github/build_cuda_linux.sh @@ -5,7 +5,7 @@ # # SPDX-License-Identifier: MIT -# A Cuda 13.3 install script for RHEL8/Rocky8/Manylinux_2.28 +# A Cuda 13.4 install script for RHEL8/Rocky8/Manylinux_2.28 # Available versions can be found at: # https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/ @@ -13,7 +13,7 @@ sudo dnf install -y kernel-devel kernel-headers sudo dnf install -y https://dl.fedoraproject.org/pub/epel/epel-release-latest-8.noarch.rpm sudo dnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-rhel8.repo -sudo dnf install -y cuda-toolkit-13-3 +sudo dnf install -y cuda-toolkit-13-4 # CUDA target architectures — LOCAL-dev build-speed knob. # @@ -38,4 +38,4 @@ case "${CUDA_FAST_BUILD:-}" in ;; esac -exec .github/build.sh $@ -DGGML_CUDA=1 -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.3/bin/nvcc $CUDA_ARCH_ARGS +exec .github/build.sh $@ -DGGML_CUDA=1 -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.4/bin/nvcc $CUDA_ARCH_ARGS diff --git a/.github/dockcross/dockcross-android-arm b/.github/dockcross/dockcross-android-arm index 314ac6bef..0520414be 100755 --- a/.github/dockcross/dockcross-android-arm +++ b/.github/dockcross/dockcross-android-arm @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/android-arm:20240418-88c04a4 image, run: +# dockcross/android-arm:20260712-79e54f9 image, run: # -# docker run --rm dockcross/android-arm:20240418-88c04a4 > dockcross-android-arm-20240418-88c04a4 -# chmod +x dockcross-android-arm-20240418-88c04a4 +# docker run --rm dockcross/android-arm:20260712-79e54f9 > dockcross-android-arm-20260712-79e54f9 +# chmod +x dockcross-android-arm-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/dockcross/dockcross-android-arm64 b/.github/dockcross/dockcross-android-arm64 index 2972e6b8a..2fe099480 100755 --- a/.github/dockcross/dockcross-android-arm64 +++ b/.github/dockcross/dockcross-android-arm64 @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/android-arm64:20240418-88c04a4 image, run: +# dockcross/android-arm64:20260712-79e54f9 image, run: # -# docker run --rm dockcross/android-arm64:20240418-88c04a4 > dockcross-android-arm64-20240418-88c04a4 -# chmod +x dockcross-android-arm64-20240418-88c04a4 +# docker run --rm dockcross/android-arm64:20260712-79e54f9 > dockcross-android-arm64-20260712-79e54f9 +# chmod +x dockcross-android-arm64-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/dockcross/dockcross-android-x86_64 b/.github/dockcross/dockcross-android-x86_64 index 7f15c9921..dca0778c2 100755 --- a/.github/dockcross/dockcross-android-x86_64 +++ b/.github/dockcross/dockcross-android-x86_64 @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/android-x86_64:20240418-88c04a4 image, run: +# dockcross/android-x86_64:20260712-79e54f9 image, run: # -# docker run --rm dockcross/android-x86_64:20240418-88c04a4 > dockcross-android-x86_64-20240418-88c04a4 -# chmod +x dockcross-android-x86_64-20240418-88c04a4 +# docker run --rm dockcross/android-x86_64:20260712-79e54f9 > dockcross-android-x86_64-20260712-79e54f9 +# chmod +x dockcross-android-x86_64-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/dockcross/dockcross-linux-arm64-lts b/.github/dockcross/dockcross-linux-arm64-lts index 4f868e88b..4f0c3a084 100755 --- a/.github/dockcross/dockcross-linux-arm64-lts +++ b/.github/dockcross/dockcross-linux-arm64-lts @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/linux-arm64-lts:20230601-c2f5366 image, run: +# dockcross/linux-arm64-lts:20260712-79e54f9 image, run: # -# docker run --rm dockcross/linux-arm64-lts:20230601-c2f5366 > dockcross-linux-arm64-lts-20230601-c2f5366 -# chmod +x dockcross-linux-arm64-lts-20230601-c2f5366 +# docker run --rm dockcross/linux-arm64-lts:20260712-79e54f9 > dockcross-linux-arm64-lts-20260712-79e54f9 +# chmod +x dockcross-linux-arm64-lts-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/dockcross/dockcross-manylinux2014-x64 b/.github/dockcross/dockcross-manylinux2014-x64 index ec9bb0578..2f1950187 100755 --- a/.github/dockcross/dockcross-manylinux2014-x64 +++ b/.github/dockcross/dockcross-manylinux2014-x64 @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/manylinux2014-x64:20230601-c2f5366 image, run: +# dockcross/manylinux2014-x64:20260712-79e54f9 image, run: # -# docker run --rm dockcross/manylinux2014-x64:20230601-c2f5366 > dockcross-manylinux2014-x64-20230601-c2f5366 -# chmod +x dockcross-manylinux2014-x64-20230601-c2f5366 +# docker run --rm dockcross/manylinux2014-x64:20260712-79e54f9 > dockcross-manylinux2014-x64-20260712-79e54f9 +# chmod +x dockcross-manylinux2014-x64-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/dockcross/dockcross-manylinux_2_28-x64 b/.github/dockcross/dockcross-manylinux_2_28-x64 index 8280b8a4c..163f7bfef 100755 --- a/.github/dockcross/dockcross-manylinux_2_28-x64 +++ b/.github/dockcross/dockcross-manylinux_2_28-x64 @@ -268,10 +268,10 @@ exit $run_exit_code # This image is not intended to be run manually. # # To create a dockcross helper script for the -# dockcross/manylinux_2_28-x64:20240812-60fa1b0 image, run: +# dockcross/manylinux_2_28-x64:20260712-79e54f9 image, run: # -# docker run --rm dockcross/manylinux_2_28-x64:20240812-60fa1b0 > dockcross-manylinux_2_28-x64-20240812-60fa1b0 -# chmod +x dockcross-manylinux_2_28-x64-20240812-60fa1b0 +# docker run --rm dockcross/manylinux_2_28-x64:20260712-79e54f9 > dockcross-manylinux_2_28-x64-20260712-79e54f9 +# chmod +x dockcross-manylinux_2_28-x64-20260712-79e54f9 # # You may then wish to move the dockcross script to your PATH. # diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 70f117d54..566b02c9d 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -743,6 +743,11 @@ jobs: SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} DOCKCROSS_ARGS: "-e SCCACHE_WEBDAV_ENDPOINT -e SCCACHE_WEBDAV_TOKEN -e USE_CACHE" steps: + - name: Free up disk space + # The GPU toolkit this job installs runs to several GB on top of the llama.cpp build tree; + # same guard upstream llama.cpp's CUDA/ROCm jobs use. Linux-only action; the tool cache is + # kept (default) so nothing a later setup-* step relies on is removed. + uses: ggml-org/free-disk-space@v1.3.1 - uses: actions/checkout@v7 - name: Download shared WebUI assets uses: actions/download-artifact@v8 @@ -1252,6 +1257,16 @@ jobs: needs: [crosscompile-android-aarch64, crosscompile-android-x86_64, verify-model-cache] runs-on: ubuntu-latest steps: + - name: Free up disk space + # The AVD userdata partition needs ~7.4 GB. Android SDK and the apt "large packages" + # (it removes libgl1-mesa-dri among others) are kept for the emulator, and so is the swap + # file it may lean on; the tool cache is kept because setup-java/setup-gradle install into it. + uses: ggml-org/free-disk-space@v1.3.1 + with: + android: false + large-packages: false + tool-cache: false + swap-storage: false - uses: actions/checkout@v7 - uses: actions/setup-java@v6 with: @@ -1300,12 +1315,12 @@ jobs: - name: Free disk space for the emulator (only the draft model is needed on-device) # Same guard as test-android-llmservice: the AVD userdata partition needs ~7.4 GB; drop the # rest of the restored ~10 GB GGUF cache (this job adb-pushes only the tiny draft model) plus - # large preinstalled toolchains so the emulator can create userdata and boot. + # the few preinstalled toolchains the free-disk-space step leaves, so the emulator can create + # userdata and boot. run: | echo "Disk before cleanup:"; df -h / | tail -1 find models -type f ! -name "${DRAFT_MODEL_NAME}" -delete 2>/dev/null || true - sudo rm -rf /usr/share/dotnet /opt/ghc /usr/local/share/powershell /opt/hostedtoolcache/CodeQL 2>/dev/null || true - docker image prune -af 2>/dev/null || true + sudo rm -rf /usr/local/share/powershell /opt/hostedtoolcache/CodeQL 2>/dev/null || true echo "Disk after cleanup:"; df -h / | tail -1 - name: Run on-emulator instrumentation (connectedDebugAndroidTest) uses: reactivecircus/android-emulator-runner@v2 @@ -1432,6 +1447,16 @@ jobs: needs: [crosscompile-android-aarch64, crosscompile-android-x86_64, verify-model-cache] runs-on: ubuntu-latest steps: + - name: Free up disk space + # The AVD userdata partition needs ~7.4 GB. Android SDK and the apt "large packages" + # (it removes libgl1-mesa-dri among others) are kept for the emulator, and so is the swap + # file it may lean on; the tool cache is kept because setup-java/setup-gradle install into it. + uses: ggml-org/free-disk-space@v1.3.1 + with: + android: false + large-packages: false + tool-cache: false + swap-storage: false - uses: actions/checkout@v7 - uses: actions/setup-java@v6 with: @@ -1477,12 +1502,11 @@ jobs: # The AVD userdata partition needs ~7.4 GB; the full ~10 GB GGUF cache restore leaves too # little free, so the emulator FATALs ("Not enough space to create userdata partition") and # the boot poll loops until timeout. This job only adb-pushes the tiny draft model, so drop - # the rest of the restored cache plus large preinstalled toolchains it never uses. + # the rest of the restored cache plus what the free-disk-space step leaves (PowerShell, CodeQL). run: | echo "Disk before cleanup:"; df -h / | tail -1 find models -type f ! -name "${DRAFT_MODEL_NAME}" -delete 2>/dev/null || true - sudo rm -rf /usr/share/dotnet /opt/ghc /usr/local/share/powershell /opt/hostedtoolcache/CodeQL 2>/dev/null || true - docker image prune -af 2>/dev/null || true + sudo rm -rf /usr/local/share/powershell /opt/hostedtoolcache/CodeQL 2>/dev/null || true echo "Disk after cleanup:"; df -h / | tail -1 - name: Run LLM Service UI test on emulator (connectedDebugAndroidTest) uses: reactivecircus/android-emulator-runner@v2 @@ -1871,15 +1895,44 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: x64 - - name: Install CUDA Toolkit - # Full toolkit install (default method: local, no sub-packages restriction). - # A reduced network sub-package set ("nvcc","cudart","cublas",…) omitted the - # nvcc crt headers (crt/host_config.h), so cmake's CUDA compiler detection - # failed at configure. The full installer ships every header reliably. - uses: Jimver/cuda-toolkit@v0.2.36 - id: cuda-toolkit - with: - cuda: '13.3.1' + - name: Install CUDA Toolkit 13.4 (NVIDIA redist archives) + # KEEP IN SYNC WITH UPSTREAM: the component set and versions mirror llama.cpp's + # .github/actions/windows-setup-cuda ("Install Cuda Toolkit 13.4 for x64") at the pinned + # GIT_TAG. Jimver/cuda-toolkit (used here up to CUDA 13.3.1) has no 13.4 in any release or + # on master, so the toolkit is assembled from NVIDIA's per-component redist zips instead — + # which is also how upstream builds its own Windows CUDA release. cuda_crt is listed + # explicitly: since 13.x the nvcc crt headers (crt/host_config.h) ship in their own + # archive, and without them CMake's CUDA compiler detection fails at configure. + shell: pwsh + run: | + $ErrorActionPreference = "Stop" + $cuda = "C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.4" + $base = "https://developer.download.nvidia.com/compute/cuda/redist" + $parts = @( + "cuda_crt/cuda_crt-windows-x86_64-13.4.59", + "cuda_cudart/cuda_cudart-windows-x86_64-13.4.49", + "cuda_nvcc/cuda_nvcc-windows-x86_64-13.4.59", + "cuda_nvrtc/cuda_nvrtc-windows-x86_64-13.4.59", + "libcublas/libcublas-windows-x86_64-13.7.0.27", + "libnvvm/libnvvm-windows-x86_64-13.4.59", + "cuda_nvtx/cuda_nvtx-windows-x86_64-13.4.49", + "cuda_profiler_api/cuda_profiler_api-windows-x86_64-13.4.49", + "visual_studio_integration/visual_studio_integration-windows-x86_64-13.4.49", + "cccl/cccl-windows-x86_64-13.3.4.2.1" + ) + New-Item -ItemType Directory -Force -Path $cuda | Out-Null + foreach ($p in $parts) { + $dir, $name = $p.Split("/") + $url = "$base/$dir/windows-x86_64/$name-archive.zip" + $zip = "$env:RUNNER_TEMP\$name.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile $zip + Expand-Archive -Path $zip -DestinationPath "$env:RUNNER_TEMP\cuda-redist" -Force + Copy-Item -Path "$env:RUNNER_TEMP\cuda-redist\$name-archive\*" -Destination $cuda -Recurse -Force + } + "$cuda\bin" | Out-File -FilePath $env:GITHUB_PATH -Append -Encoding utf8 + "CUDA_PATH=$cuda" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8 + "CUDA_PATH_V13_4=$cuda" | Out-File -FilePath $env:GITHUB_ENV -Append -Encoding utf8 - name: Install sccache (shared compiler cache) if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' continue-on-error: true @@ -2060,12 +2113,13 @@ jobs: - name: Build libraries shell: bash # Native CMake HIP language (upstream's form): the HIP compiler is ROCm's clang, the C/C++ - # TUs stay on the runner's gcc. The target list is upstream's; architectures llama.cpp no - # longer ships (gfx900/gfx906 — "build passing" only in TheRock, never release-ready) are - # not carried here. + # TUs stay on the runner's gcc. The target list is upstream's plus gfx900/gfx906, which this + # classifier shipped before and which TheRock 10 still builds ("build passing", not + # release-ready) although upstream dropped them from its own list. Keep them only while + # they cost nothing: drop them the moment they need a patch or block a newer ROCm. run: | mvn --no-transfer-progress -f llama/pom.xml compile - .github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx908;gfx90a;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" + .github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx900;gfx906;gfx908;gfx90a;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" - name: Upload artifacts uses: actions/upload-artifact@v7 with: @@ -2142,6 +2196,11 @@ jobs: SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: + - name: Free up disk space + # The GPU toolkit this job installs runs to several GB on top of the llama.cpp build tree; + # same guard upstream llama.cpp's CUDA/ROCm jobs use. Linux-only action; the tool cache is + # kept (default) so nothing a later setup-* step relies on is removed. + uses: ggml-org/free-disk-space@v1.3.1 - uses: actions/checkout@v7 - name: Download shared WebUI assets uses: actions/download-artifact@v8 @@ -2180,6 +2239,11 @@ jobs: SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: + - name: Free up disk space + # The GPU toolkit this job installs runs to several GB on top of the llama.cpp build tree; + # same guard upstream llama.cpp's CUDA/ROCm jobs use. Linux-only action; the tool cache is + # kept (default) so nothing a later setup-* step relies on is removed. + uses: ggml-org/free-disk-space@v1.3.1 - uses: actions/checkout@v7 - name: Download shared WebUI assets uses: actions/download-artifact@v8 diff --git a/CHANGELOG.md b/CHANGELOG.md index 657dbf537..53374288d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -119,8 +119,12 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by its wheels the way upstream llama.cpp's own release jobs do, replacing the `repo.radeon.com` 6.3.4 apt repo (Linux) and the HIP SDK 26.Q1 installer (Windows). The GPU target lists now follow upstream too: newer architectures are added (Linux: gfx942/gfx950, RDNA1–RDNA4 incl. gfx1150–1152 and - gfx1200/1201; Windows: RDNA1–RDNA4), and gfx900/gfx906 are dropped — llama.cpp no longer ships them - and TheRock never marks them release-ready. Consumers need a ROCm 10 runtime. + gfx1200/1201; Windows: RDNA1–RDNA4). gfx900/gfx906 stay in the Linux build on a best-effort basis: + llama.cpp no longer ships them and TheRock marks them build-passing only, so they are kept exactly + as long as they build without patches. Consumers need a ROCm 10 runtime. +- **CUDA 13.3 → 13.4** for both `cuda13-*` classifiers, matching upstream llama.cpp. Linux installs + `cuda-toolkit-13-4`; Windows assembles the toolkit from NVIDIA's redist archives (upstream's + component list) because `Jimver/cuda-toolkit` never shipped 13.4. Classifier names are unchanged. - **llama.cpp `b11069` → `b11080`, and local patch `0010` dropped — upstream fixed the enum-to-JSON-boolean trap at its root.** Eleven upstream commits, 1244 KiB, no project-source change. The size is one commit that does not concern this project (llama.cpp #29197 rewrites 46 diff --git a/CLAUDE.md b/CLAUDE.md index b18636dbb..1164d01cb 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -10,34 +10,30 @@ Current llama.cpp pinned version: **b11209** ## Upgrading CUDA Version -Current CUDA version: **13.3** - -To change the CUDA version, update the following **three** places: - -1. **`.github/build_cuda_linux.sh`** — Line 16: `sudo dnf install -y cuda-toolkit-13-3` -2. **`.github/build_cuda_linux.sh`** — Line 41: `-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.3/bin/nvcc` -3. **`llama/pom.xml`** — The `` tag in the `cuda` jar execution: `cuda13-linux-x86-64` - -Also update the header comment in `build_cuda_linux.sh` and the job name in `.github/workflows/release.yaml` for clarity. +Current CUDA version: **13.4** (Linux `cuda-toolkit-13-4` from NVIDIA's rhel8 repo; Windows 13.4 redist archives) + +To change the CUDA version, update the following places: + +1. **`.github/build_cuda_linux.sh`** — the `sudo dnf install -y cuda-toolkit-13-4` line and the + `-DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.4/bin/nvcc` line (plus the header comment). +2. **`.github/workflows/publish.yml`** — the `build-windows-x86_64-cuda` job's + "Install CUDA Toolkit … (NVIDIA redist archives)" step: the `v13.x` directory, `CUDA_PATH_V13_x`, + and the per-component archive list. **Copy that list from upstream's + `.github/actions/windows-setup-cuda/action.yml` at the pinned `GIT_TAG`** — the component versions + differ per component (cuBLAS and CCCL have their own numbering) and cannot be derived from the CUDA + version. (This replaced `Jimver/cuda-toolkit`, which never shipped 13.4.) +3. **`llama/pom.xml`** — the ``s `cuda13-linux-x86-64` / `cuda13-windows-x86-64` + (major version only — no change for a minor bump). +4. **`CLAUDE.md`** — the "Current CUDA version" line above. Available CUDA versions for RHEL8/Manylinux_2_28 can be browsed at: ``` https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/ ``` +and the Windows redist components at `https://developer.download.nvidia.com/compute/cuda/redist/`. **Note:** Each CUDA version supports only certain GCC versions. If the dockcross container uses a newer GCC than CUDA supports, the build will fail with `unsupported GNU version`. Check NVIDIA's compatibility table before downgrading CUDA. -Example: To upgrade from 13.3 to a hypothetical 13.4: -```bash -# Edit .github/build_cuda_linux.sh: -# line 10: cuda-toolkit-13-3 -> cuda-toolkit-13-4 -# line 12: /usr/local/cuda-13.3/bin/nvcc -> /usr/local/cuda-13.4/bin/nvcc -# Edit llama/pom.xml classifier: cuda13-linux-x86-64 (major version only, no need to change for minor bumps) -# Edit CLAUDE.md line: Current CUDA version: **13.3** -> **13.4** -git add .github/build_cuda_linux.sh llama/pom.xml CLAUDE.md -git commit -m "Upgrade CUDA from 13.3 to 13.4" -``` - ### Fast local CUDA builds (`CUDA_FAST_BUILD`) — single-arch speed knob The CUDA artifact must ship kernels for **every supported GPU generation**, so the default @@ -217,7 +213,8 @@ Wiring (mirrors the CUDA-Linux / OpenCL-Android classifier pattern): - `build-windows-x86_64` / `build-windows-x86` — **Ninja CPU**, artifacts `Windows-{arch}-libraries` → picked up by the `package` job's `pattern: "*-libraries"` into the **default** tree. - `build-windows-x86_64-msvc` / `build-windows-x86-msvc` — **MSVC CPU**, artifacts `Windows-{arch}-msvc`. - - `build-windows-x86_64-cuda` — `Jimver/cuda-toolkit@v0.2.36` (CUDA `13.3.1`) + `-DGGML_CUDA=ON`, + - `build-windows-x86_64-cuda` — CUDA `13.4` assembled from NVIDIA's redist archives (upstream's + `windows-setup-cuda` component list; `Jimver/cuda-toolkit` stops at 13.3.1) + `-DGGML_CUDA=ON`, artifact `Windows-x86_64-cuda`. - `build-windows-x86_64-vulkan` — `jakoch/install-vulkan-sdk-action` + `-DGGML_VULKAN=ON`, artifact `Windows-x86_64-vulkan`. @@ -336,10 +333,12 @@ builds and releases ROCm through [TheRock](https://github.com/ROCm/TheRock); bot its Python wheels (`rocm[libraries,devel]` from `stable.repo.amd.com/rocm/whl-next/`) exactly as llama.cpp's own `ubuntu-rocm` / `windows-rocm` release jobs do, and read the paths back with `rocm-sdk path`. The ROCm version and the `GPU_TARGETS` list are copied from upstream's -`release.yml` at the pinned `GIT_TAG` — **re-check both on every llama.cpp bump**. Architectures -upstream no longer builds are deliberately **not** carried along (gfx900/gfx906 were dropped at the -switch from the old 6.3.4 apt repo: TheRock marks them "build passing" only, never release-ready): -supporting hardware llama.cpp itself does not ship for is not worth holding back a newer toolchain. +`release.yml` at the pinned `GIT_TAG` — **re-check both on every llama.cpp bump**. The one deliberate +deviation: the Linux job additionally keeps **gfx900/gfx906**, which this classifier shipped before +and which TheRock still builds ("build passing", never release-ready) although upstream dropped them. +The rule for such extras: carry them only while they build without problems and without local +patches; the moment one needs a patch or holds back a newer ROCm/llama.cpp, drop it — hardware +llama.cpp itself does not ship for is never worth blocking something newer. Two routing notes mirror existing precedent: **Linux SYCL** ships two precision variants at the *same* arch, so `CMakeLists.txt` routes them to two *distinct* trees by `GGML_SYCL_F16` (fp16 vs fp32). @@ -347,6 +346,10 @@ arch, so `CMakeLists.txt` routes them to two *distinct* trees by `GGML_SYCL_F16` `resources_windows_opencl` tree, split by the `opencl-windows` / `opencl-windows-aarch64` profiles' arch-scoped `` — exactly like the `vulkan-linux` / `vulkan-linux-aarch64` split. +The Linux jobs that install a multi-GB vendor toolchain (CUDA, ROCm, both SYCL) start with +`ggml-org/free-disk-space` — the same guard upstream llama.cpp's CUDA/ROCm jobs use (ROCm also clears the +tool cache, as upstream does, which is safe only because the step runs before `setup-java`). + The vendor toolchain install steps in `publish.yml` are **first-pass** (apt repos / vendor installers pinned to a specific version): if a URL/version 404s in CI, the job fails loud and the step is adjusted — the failure is intentional signal, not a regression to hide behind `continue-on-error`. @@ -2675,7 +2678,8 @@ native jobs; the test job additionally `needs: verify-model-cache`), so the inst Both emulator jobs (`test-android-llmservice` + `test-android-emulator`) prepend a **free-disk step** before the emulator (delete every restored model except `DRAFT_MODEL_NAME` + large unused preinstalled -toolchains) because the AVD userdata partition needs ~7.4 GB and the full ~10 GB GGUF cache restore +toolchains — the latter via `ggml-org/free-disk-space` with `android`, `large-packages`, `tool-cache` +and `swap-storage` switched **off**, since the emulator needs the SDK, mesa and the JDK) because the AVD userdata partition needs ~7.4 GB and the full ~10 GB GGUF cache restore otherwise FATALs the emulator ("Not enough space to create userdata partition"). Neither llmservice job is **yet a publish gate** (not in the `publish-snapshot`/`publish-release` `needs:` graphs) so a Compose/AGP version-pin hiccup can't block a library release. diff --git a/README.md b/README.md index 905902b14..4ca1356d4 100644 --- a/README.md +++ b/README.md @@ -191,7 +191,7 @@ exclusive — and optionally a CPU Windows build. | `vulkan-linux-x86-64` | Vulkan | Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) | A Vulkan runtime (`libvulkan.so.1`), which current GPU drivers install. No Vulkan SDK is needed at runtime. The most portable Linux GPU option (vendor-independent, no CUDA toolkit). Built natively on `ubuntu-latest`, so it shares the aarch64 build's higher glibc floor (≈ 2.39). | | `vulkan-linux-aarch64` | Vulkan | Linux aarch64 with a Vulkan 1.2+ GPU | A Vulkan runtime (`libvulkan.so.1`) from the device/driver. glibc ≥ 2.39 (built on `ubuntu-24.04-arm`). | | `opencl-android-aarch64` | OpenCL (Adreno) | Android aarch64 with Qualcomm Adreno GPU | A device-supplied OpenCL ICD (`libOpenCL.so`). Devices without an ICD (e.g. most non-Snapdragon Android hardware) must use the default CPU JAR. | -| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm **10** runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host — built against ROCm 10.0 (TheRock), the same version upstream llama.cpp ships; GPUs older than gfx908 (e.g. gfx900/gfx906) are not targeted. Not bundled; native load fails without it. No CPU fallback. | +| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm **10** runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host — built against ROCm 10.0 (TheRock), the same version upstream llama.cpp ships; targets gfx900 and newer (gfx900/gfx906 best effort — built, but not release-ready in ROCm 10). Not bundled; native load fails without it. No CPU fallback. | | `rocm-windows-x86-64` | ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm **10** runtime DLLs (`amdhip64.dll`, `rocblas.dll`, `hipblas.dll`) on `PATH` — built against ROCm 10.0 (TheRock); RDNA1 (gfx1010) and newer. Not bundled. No CPU fallback. | | `sycl-fp16-linux-x86-64` | SYCL (Intel oneAPI, fp16) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp16 accumulation (faster, slightly lower precision). Not bundled. | | `sycl-fp32-linux-x86-64` | SYCL (Intel oneAPI, fp32) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp32 accumulation (higher precision). Not bundled. | From ea29df7de2a191dce1dc3619d71a7d948ee777b9 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:30:34 +0000 Subject: [PATCH 08/10] ROCm: build every GPU target TheRock supports per OS The ROCm target lists were upstream llama.cpp's, which differed between Linux and Windows partly by accident. They are now every target TheRock builds per OS (its SUPPORTED_GPUS.md): Linux 27 targets, Windows 23. That adds gfx90c and gfx1153 on Linux and gfx900/gfx906/gfx90c on Windows. The two lists now differ only by the Instinct parts (gfx908/gfx90a/gfx942/gfx950), which ROCm supports on Linux alone. The extras upstream omits are build-passing only in TheRock; they stay as long as they build without patches and do not block a newer ROCm. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- .github/workflows/publish.yml | 24 ++++++++++++++---------- CHANGELOG.md | 11 ++++++----- CLAUDE.md | 15 ++++++++------- README.md | 4 ++-- 4 files changed, 30 insertions(+), 24 deletions(-) diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 566b02c9d..4050c5d85 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -2090,8 +2090,8 @@ jobs: distribution: 'temurin' java-version: ${{ env.JAVA_VERSION }} - name: Install ROCm/HIP (TheRock wheels) - # KEEP IN SYNC WITH UPSTREAM: ROCM_VERSION and the GPU target list track the ubuntu-rocm - # job in llama.cpp's .github/workflows/release.yml at the pinned GIT_TAG. Since ROCm 7.14 + # KEEP IN SYNC WITH UPSTREAM: ROCM_VERSION tracks the ubuntu-rocm job in llama.cpp's + # .github/workflows/release.yml at the pinned GIT_TAG (GPU targets: see the build step). Since ROCm 7.14 # AMD builds and releases ROCm through TheRock (https://github.com/ROCm/TheRock), replacing # the monolithic releases behind the repo.radeon.com apt repo this job used before (it was # pinned at 6.3.4). The wheels carry the HIP runtime + CMake configs ("libraries") and the @@ -2113,13 +2113,15 @@ jobs: - name: Build libraries shell: bash # Native CMake HIP language (upstream's form): the HIP compiler is ROCm's clang, the C/C++ - # TUs stay on the runner's gcc. The target list is upstream's plus gfx900/gfx906, which this - # classifier shipped before and which TheRock 10 still builds ("build passing", not - # release-ready) although upstream dropped them from its own list. Keep them only while - # they cost nothing: drop them the moment they need a patch or block a newer ROCm. + # TUs stay on the runner's gcc. The target list is every Linux target TheRock builds + # (its SUPPORTED_GPUS.md) — a superset of upstream llama.cpp's list, which omits + # gfx900/gfx906/gfx90c/gfx1153 ("build passing" only, not release-ready). Better too many + # than too few, but only while they cost nothing: drop an extra the moment it needs a + # patch or blocks a newer ROCm. Linux alone has the Instinct parts (gfx908/90a/942/950): + # ROCm does not support them on Windows. run: | mvn --no-transfer-progress -f llama/pom.xml compile - .github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx900;gfx906;gfx908;gfx90a;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" + .github/build.sh "-DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$(hipconfig -l)/clang -DGPU_TARGETS=gfx900;gfx906;gfx908;gfx90a;gfx90c;gfx942;gfx950;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DGGML_NATIVE=OFF -DOS_NAME=Linux -DOS_ARCH=x86_64" - name: Upload artifacts uses: actions/upload-artifact@v7 with: @@ -2176,10 +2178,12 @@ jobs: "C:\TheRock\build\.venv\Scripts" | Out-File -FilePath $env:GITHUB_PATH -Append - name: Build libraries shell: cmd - # Upstream's compiler wiring for TheRock (clang under lib\llvm\bin, not bin\) and its - # GPU target list; -Wno-error=incompatible-pointer-types is upstream's too. + # Upstream's compiler wiring for TheRock (clang under lib\llvm\bin, not bin\); + # -Wno-error=incompatible-pointer-types is upstream's too. Targets: every Windows target + # TheRock builds — the Linux list minus the Instinct parts, which ROCm has no Windows + # support for. Same rule for the extras upstream omits as on Linux. run: | - .github\build.bat -G "Ninja Multi-Config" -DGGML_HIP=ON -DGPU_TARGETS=gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DCMAKE_PREFIX_PATH="%HIP_PATH%" -DHIP_PATH="%HIP_PATH%" -DCMAKE_C_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_CXX_COMPILER="%HIP_PATH%\lib\llvm\bin\clang++.exe" -DCMAKE_HIP_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_C_FLAGS="-Wno-error=incompatible-pointer-types" -DOS_NAME=Windows -DOS_ARCH=x86_64 + .github\build.bat -G "Ninja Multi-Config" -DGGML_HIP=ON -DGPU_TARGETS=gfx900;gfx906;gfx90c;gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201 -DCMAKE_PREFIX_PATH="%HIP_PATH%" -DHIP_PATH="%HIP_PATH%" -DCMAKE_C_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_CXX_COMPILER="%HIP_PATH%\lib\llvm\bin\clang++.exe" -DCMAKE_HIP_COMPILER="%HIP_PATH%\lib\llvm\bin\clang.exe" -DCMAKE_C_FLAGS="-Wno-error=incompatible-pointer-types" -DOS_NAME=Windows -DOS_ARCH=x86_64 - name: Upload artifacts uses: actions/upload-artifact@v7 with: diff --git a/CHANGELOG.md b/CHANGELOG.md index 53374288d..b9b2f8004 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -117,11 +117,12 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by - **BREAKING (runtime): the ROCm classifiers are built against ROCm 10.0 (TheRock), not 6.3.** AMD now releases ROCm through [TheRock](https://github.com/ROCm/TheRock); both `rocm-*` jobs install its wheels the way upstream llama.cpp's own release jobs do, replacing the `repo.radeon.com` 6.3.4 apt - repo (Linux) and the HIP SDK 26.Q1 installer (Windows). The GPU target lists now follow upstream too: - newer architectures are added (Linux: gfx942/gfx950, RDNA1–RDNA4 incl. gfx1150–1152 and - gfx1200/1201; Windows: RDNA1–RDNA4). gfx900/gfx906 stay in the Linux build on a best-effort basis: - llama.cpp no longer ships them and TheRock marks them build-passing only, so they are kept exactly - as long as they build without patches. Consumers need a ROCm 10 runtime. + repo (Linux) and the HIP SDK 26.Q1 installer (Windows). The GPU target lists are now every target + TheRock builds per OS — a superset of upstream llama.cpp's: Linux adds gfx90c, gfx942/gfx950, RDNA1, + the rest of RDNA2/RDNA3, RDNA3.5 (gfx1150–1153) and RDNA4 (gfx1200/1201) to the previous eight; + Windows goes from four RDNA2/3 targets to 23, gfx900 through RDNA4. The extras upstream omits + (gfx900/gfx906/gfx90c/gfx1153, build-passing only in TheRock) are kept as long as they build + without patches. Consumers need a ROCm 10 runtime. - **CUDA 13.3 → 13.4** for both `cuda13-*` classifiers, matching upstream llama.cpp. Linux installs `cuda-toolkit-13-4`; Windows assembles the toolkit from NVIDIA's redist archives (upstream's component list) because `Jimver/cuda-toolkit` never shipped 13.4. Classifier names are unchanged. diff --git a/CLAUDE.md b/CLAUDE.md index 1164d01cb..499d2f21d 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -332,13 +332,14 @@ matching GPU) and bundle **no** vendor runtime. builds and releases ROCm through [TheRock](https://github.com/ROCm/TheRock); both ROCm jobs install its Python wheels (`rocm[libraries,devel]` from `stable.repo.amd.com/rocm/whl-next/`) exactly as llama.cpp's own `ubuntu-rocm` / `windows-rocm` release jobs do, and read the paths back with -`rocm-sdk path`. The ROCm version and the `GPU_TARGETS` list are copied from upstream's -`release.yml` at the pinned `GIT_TAG` — **re-check both on every llama.cpp bump**. The one deliberate -deviation: the Linux job additionally keeps **gfx900/gfx906**, which this classifier shipped before -and which TheRock still builds ("build passing", never release-ready) although upstream dropped them. -The rule for such extras: carry them only while they build without problems and without local -patches; the moment one needs a patch or holds back a newer ROCm/llama.cpp, drop it — hardware -llama.cpp itself does not ship for is never worth blocking something newer. +`rocm-sdk path`. The ROCm version follows upstream's `release.yml` at the pinned `GIT_TAG` — +**re-check it on every llama.cpp bump**. The `GPU_TARGETS` lists deliberately go **further than +upstream's**: they are every target TheRock builds for that OS (its `SUPPORTED_GPUS.md`), which adds +gfx900/gfx906/gfx90c/gfx1153 — "build passing" there, not release-ready, and omitted by llama.cpp. +Supporting more rather than fewer is the policy, with one limit: an extra stays only while it builds +without problems and without local patches; the moment one needs a patch or holds back a newer +ROCm/llama.cpp, drop it. The two lists differ **only** by the Instinct parts +(gfx908/gfx90a/gfx942/gfx950), which ROCm supports on Linux alone. Two routing notes mirror existing precedent: **Linux SYCL** ships two precision variants at the *same* arch, so `CMakeLists.txt` routes them to two *distinct* trees by `GGML_SYCL_F16` (fp16 vs fp32). diff --git a/README.md b/README.md index 4ca1356d4..3f53ec0cf 100644 --- a/README.md +++ b/README.md @@ -191,8 +191,8 @@ exclusive — and optionally a CPU Windows build. | `vulkan-linux-x86-64` | Vulkan | Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) | A Vulkan runtime (`libvulkan.so.1`), which current GPU drivers install. No Vulkan SDK is needed at runtime. The most portable Linux GPU option (vendor-independent, no CUDA toolkit). Built natively on `ubuntu-latest`, so it shares the aarch64 build's higher glibc floor (≈ 2.39). | | `vulkan-linux-aarch64` | Vulkan | Linux aarch64 with a Vulkan 1.2+ GPU | A Vulkan runtime (`libvulkan.so.1`) from the device/driver. glibc ≥ 2.39 (built on `ubuntu-24.04-arm`). | | `opencl-android-aarch64` | OpenCL (Adreno) | Android aarch64 with Qualcomm Adreno GPU | A device-supplied OpenCL ICD (`libOpenCL.so`). Devices without an ICD (e.g. most non-Snapdragon Android hardware) must use the default CPU JAR. | -| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm **10** runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host — built against ROCm 10.0 (TheRock), the same version upstream llama.cpp ships; targets gfx900 and newer (gfx900/gfx906 best effort — built, but not release-ready in ROCm 10). Not bundled; native load fails without it. No CPU fallback. | -| `rocm-windows-x86-64` | ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm **10** runtime DLLs (`amdhip64.dll`, `rocblas.dll`, `hipblas.dll`) on `PATH` — built against ROCm 10.0 (TheRock); RDNA1 (gfx1010) and newer. Not bundled. No CPU fallback. | +| `rocm-linux-x86-64` | ROCm / HIP | Linux x86-64 with AMD GPU | An installed AMD ROCm **10** runtime (`libamdhip64.so`, `librocblas.so`, `libhipblas.so`) on the host — built against ROCm 10.0 (TheRock), the same version upstream llama.cpp ships; targets every GPU TheRock builds for Linux, Instinct included (gfx900/gfx906/gfx90c/gfx1153 best effort — built, but not release-ready in ROCm 10). Not bundled; native load fails without it. No CPU fallback. | +| `rocm-windows-x86-64` | ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm **10** runtime DLLs (`amdhip64.dll`, `rocblas.dll`, `hipblas.dll`) on `PATH` — built against ROCm 10.0 (TheRock); every Radeon target TheRock builds for Windows, gfx900 through RDNA4 (gfx900/gfx906/gfx90c/gfx1153 best effort). Not bundled. No CPU fallback. | | `sycl-fp16-linux-x86-64` | SYCL (Intel oneAPI, fp16) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp16 accumulation (faster, slightly lower precision). Not bundled. | | `sycl-fp32-linux-x86-64` | SYCL (Intel oneAPI, fp32) | Linux x86-64 with Intel GPU (Arc / iGPU) | An installed Intel oneAPI / Level-Zero runtime. fp32 accumulation (higher precision). Not bundled. | | `sycl-windows-x86-64` | SYCL (Intel oneAPI) | Windows x86-64 with Intel GPU (Arc / iGPU) | The Intel oneAPI / Level-Zero runtime DLLs on `PATH`. Not bundled. | From 607b423e76106e4f352b864d816c7f5dd185d648 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:50:33 +0000 Subject: [PATCH 09/10] sccache on every Windows Ninja job, with an uncached retry in build.bat Five Windows build jobs never installed sccache: arm64, arm64-OpenCL, ROCm, SYCL and OpenVINO. They now get the same USE_CACHE/SCCACHE_WEBDAV env and install step as the other Ninja jobs. The two arm64 jobs use the native aarch64-pc-windows-msvc release, which exists (the comment that said only x86_64 was available is gone). Only the two MSVC-classifier jobs stay uncached: the Visual Studio generator ignores launchers. build.bat's probe only proves sccache can wrap cl.exe, while these jobs compile with clang-cl, ROCm's clang or icx. So a configure or build that fails with sccache as the launcher is now retried once from a clean build dir without it. The retry is unconditional because cmd cannot tee output to match an error signature; a real compile error costs one extra uncached attempt, a cache incompatibility ends green. Every configure also passes -DGGML_CCACHE=OFF. Without it ggml self-enables any sccache on PATH whenever no launcher is set, i.e. in the probe-failed and retry cases, and the uncached build went through sccache after all. No goto/labels: the file is checked out with LF line endings, where cmd's label search is unreliable. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- .github/build.bat | 37 +++++++++++-- .github/workflows/publish.yml | 100 +++++++++++++++++++++++++++++++++- CLAUDE.md | 18 +++++- 3 files changed, 146 insertions(+), 9 deletions(-) diff --git a/.github/build.bat b/.github/build.bat index d2c755b3e..7976f75f9 100755 --- a/.github/build.bat +++ b/.github/build.bat @@ -49,11 +49,40 @@ REM nvcc command line (it dies with `sccache: error: Could not parse shell line` REM fails every .cu compile). So CUDA device code is built by nvcc directly (uncached) REM here; the cl.exe C/C++ TUs still cache via the C/CXX launcher set above. +REM The probe above only proves sccache can wrap cl.exe. Several jobs compile with a +REM different compiler (clang-cl on arm64, ROCm's clang, Intel icx), which sccache may +REM refuse or mishandle -- so a configure or build that fails WITH the launcher is +REM retried ONCE from a clean build dir WITHOUT it (the build.sh retry, but unconditional: +REM cmd cannot tee the output to match an sccache error signature). A genuine compile +REM error therefore still fails, just after one uncached attempt; a cache/compiler +REM incompatibility ends green and uncached, never red. +REM -DGGML_CCACHE=OFF on every configure: ggml otherwise self-enables any ccache/sccache it +REM finds on PATH when no launcher is set -- i.e. exactly in the probe-failed and retry +REM cases -- via the global RULE_LAUNCH_COMPILE, silently re-caching the uncached build. +REM No goto/labels on purpose: this file is checked out with LF line endings +REM (.gitattributes eol=lf), and cmd's label search is unreliable in LF-only scripts. +set "RETRY=" mkdir build -cmake -Bbuild %LAUNCH% %* -if errorlevel 1 exit /b 1 -cmake --build build --config Release -set "BUILD_RC=!ERRORLEVEL!" +cmake -Bbuild %LAUNCH% -DGGML_CCACHE=OFF %* +if errorlevel 1 ( + if not defined LAUNCH exit /b 1 + set "RETRY=1" +) else ( + cmake --build build --config Release + set "BUILD_RC=!ERRORLEVEL!" + if not "!BUILD_RC!"=="0" if defined LAUNCH set "RETRY=1" +) +if defined RETRY ( + echo build.bat: build WITH sccache failed -- retrying ONCE from scratch WITHOUT the cache. + sccache --show-stats + set "LAUNCH=" + rmdir /s /q build + mkdir build + cmake -Bbuild -DGGML_CCACHE=OFF %* + if errorlevel 1 exit /b 1 + cmake --build build --config Release + set "BUILD_RC=!ERRORLEVEL!" +) REM Print cache stats (best-effort) regardless of build outcome -- only when sccache REM was wired in as the launcher. diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 4050c5d85..42b46f623 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -1816,9 +1816,8 @@ jobs: # Native arm64 build on GitHub's free windows-11-arm runner. Goes into the DEFAULT JAR (no # classifier): OSInfo maps a Windows-on-ARM JVM (os.arch=aarch64) to Windows/aarch64, the same # path CMake emits here, and the `*-libraries` glob in the package/publish jobs merges it into - # src/main/resources. sccache is intentionally omitted (the existing install step pulls the - # x86_64 sccache zip; an arm64 build would need the aarch64 release — not worth the extra path - # for one CPU job, so build.bat just builds uncached when sccache is absent). + # src/main/resources. sccache: the native aarch64-pc-windows-msvc release, wrapping clang-cl + # (guarded by build.bat's probe + uncached retry like every other Windows job). # # Compiler: clang-cl, NOT MSVC cl.exe. ggml's ggml-cpu/CMakeLists.txt aborts with "MSVC is not # supported for ARM, use clang" via `if (MSVC AND NOT CMAKE_C_COMPILER_ID STREQUAL "Clang")`. @@ -1834,6 +1833,10 @@ jobs: # off makes ggml use its own std::thread threadpool, so the arm64 jllama.dll (and the test exe) are # self-contained with no libomp dependency to ship. The x86_64/x86 jobs keep OpenMP (MSVC vcomp). runs-on: windows-11-arm + env: + USE_CACHE: ${{ github.event_name != 'workflow_dispatch' || inputs.use_cache }} + SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev + SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: - uses: actions/checkout@v7 - name: Download shared WebUI assets @@ -1845,6 +1848,21 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: arm64 + - name: Install sccache (shared compiler cache) + if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' + continue-on-error: true + shell: pwsh + # sccache ships a native Windows-on-ARM build; it wraps clang-cl here. + # build.bat probes sccache before trusting it and, should a build fail with it as the + # launcher, retries once uncached -- so this can speed the job up but never red it. + run: | + $ver = "0.18.0" + $rel = "sccache-v$ver-aarch64-pc-windows-msvc" + $url = "https://github.com/mozilla/sccache/releases/download/v$ver/$rel.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\sccache.zip" + Expand-Archive -Path "$env:RUNNER_TEMP\sccache.zip" -DestinationPath "$env:RUNNER_TEMP\sccache" -Force + Add-Content -Path $env:GITHUB_PATH -Value "$env:RUNNER_TEMP\sccache\$rel" - name: Build libraries shell: cmd # No mvn compile needed: the JNI header (jllama.h) is committed and the native build @@ -2138,6 +2156,10 @@ jobs: # device-code compile fails. Upstream llama.cpp builds win-hip on windows-2022 for the same # reason (it drives ROCm's own clang and relies on the older MSVC STL). runs-on: windows-2022 + env: + USE_CACHE: ${{ github.event_name != 'workflow_dispatch' || inputs.use_cache }} + SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev + SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: - uses: actions/checkout@v7 - name: Download shared WebUI assets @@ -2176,6 +2198,21 @@ jobs: "LLVM_PATH=$rocm\lib\llvm" | Out-File -FilePath $env:GITHUB_ENV -Append (rocm-sdk path --bin).Trim() | Out-File -FilePath $env:GITHUB_PATH -Append "C:\TheRock\build\.venv\Scripts" | Out-File -FilePath $env:GITHUB_PATH -Append + - name: Install sccache (shared compiler cache) + if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' + continue-on-error: true + shell: pwsh + # Wraps ROCm's clang (HIP is compiled as C++ on Windows, so the C/CXX launcher covers it). + # build.bat probes sccache before trusting it and, should a build fail with it as the + # launcher, retries once uncached -- so this can speed the job up but never red it. + run: | + $ver = "0.18.0" + $rel = "sccache-v$ver-x86_64-pc-windows-msvc" + $url = "https://github.com/mozilla/sccache/releases/download/v$ver/$rel.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\sccache.zip" + Expand-Archive -Path "$env:RUNNER_TEMP\sccache.zip" -DestinationPath "$env:RUNNER_TEMP\sccache" -Force + Add-Content -Path $env:GITHUB_PATH -Value "$env:RUNNER_TEMP\sccache\$rel" - name: Build libraries shell: cmd # Upstream's compiler wiring for TheRock (clang under lib\llvm\bin, not bin\); @@ -2281,6 +2318,10 @@ jobs: name: Build Windows 2025 x86_64 SYCL (Intel oneAPI) needs: [startgate, build-webui] runs-on: windows-2025-vs2026 + env: + USE_CACHE: ${{ github.event_name != 'workflow_dispatch' || inputs.use_cache }} + SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev + SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: - uses: actions/checkout@v7 - name: Download shared WebUI assets @@ -2300,6 +2341,21 @@ jobs: curl -fSL -o "%RUNNER_TEMP%\oneapi.exe" "https://registrationcenter-download.intel.com/akdlm/IRC_NAS/b60765d1-2b85-4e85-86b6-cb0e9563a699/intel-deep-learning-essentials-2025.3.3.18_offline.exe" "%RUNNER_TEMP%\oneapi.exe" -s -x -f "%RUNNER_TEMP%\oneapi_extracted" --log "%RUNNER_TEMP%\extract.log" "%RUNNER_TEMP%\oneapi_extracted\bootstrapper.exe" -s --action install --components=intel.oneapi.win.cpp-dpcpp-common:intel.oneapi.win.mkl.devel:intel.oneapi.win.dnnl:intel.oneapi.win.tbb.devel --eula=accept -p=NEED_VS2022_INTEGRATION=0 --log-dir="%RUNNER_TEMP%" + - name: Install sccache (shared compiler cache) + if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' + continue-on-error: true + shell: pwsh + # Wraps cl.exe (C) and icx (C++); if sccache cannot handle icx the retry builds uncached. + # build.bat probes sccache before trusting it and, should a build fail with it as the + # launcher, retries once uncached -- so this can speed the job up but never red it. + run: | + $ver = "0.18.0" + $rel = "sccache-v$ver-x86_64-pc-windows-msvc" + $url = "https://github.com/mozilla/sccache/releases/download/v$ver/$rel.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\sccache.zip" + Expand-Archive -Path "$env:RUNNER_TEMP\sccache.zip" -DestinationPath "$env:RUNNER_TEMP\sccache" -Force + Add-Content -Path $env:GITHUB_PATH -Value "$env:RUNNER_TEMP\sccache\$rel" - name: Build libraries shell: cmd run: | @@ -2321,6 +2377,10 @@ jobs: # Maven profile packages only that subtree. build_opencl_windows.bat stages the # OpenCL headers + ICD loader before delegating to build.bat. runs-on: windows-11-arm + env: + USE_CACHE: ${{ github.event_name != 'workflow_dispatch' || inputs.use_cache }} + SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev + SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: - uses: actions/checkout@v7 - name: Download shared WebUI assets @@ -2332,6 +2392,21 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: arm64 + - name: Install sccache (shared compiler cache) + if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' + continue-on-error: true + shell: pwsh + # sccache ships a native Windows-on-ARM build; it wraps clang-cl here. + # build.bat probes sccache before trusting it and, should a build fail with it as the + # launcher, retries once uncached -- so this can speed the job up but never red it. + run: | + $ver = "0.18.0" + $rel = "sccache-v$ver-aarch64-pc-windows-msvc" + $url = "https://github.com/mozilla/sccache/releases/download/v$ver/$rel.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\sccache.zip" + Expand-Archive -Path "$env:RUNNER_TEMP\sccache.zip" -DestinationPath "$env:RUNNER_TEMP\sccache" -Force + Add-Content -Path $env:GITHUB_PATH -Value "$env:RUNNER_TEMP\sccache\$rel" - name: Build libraries shell: cmd run: | @@ -2398,6 +2473,10 @@ jobs: name: Build Windows 2025 x86_64 OpenVINO (Intel) needs: [startgate, build-webui] runs-on: windows-2025-vs2026 + env: + USE_CACHE: ${{ github.event_name != 'workflow_dispatch' || inputs.use_cache }} + SCCACHE_WEBDAV_ENDPOINT: https://cache.depot.dev + SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }} steps: - uses: actions/checkout@v7 - name: Download shared WebUI assets @@ -2424,6 +2503,21 @@ jobs: # The archive extracts into a nested versioned folder; point OpenVINO_DIR at its runtime/cmake. $root = (Get-ChildItem "C:\openvino" -Directory | Select-Object -First 1).FullName "OpenVINO_DIR=$root\runtime\cmake" | Out-File -FilePath $env:GITHUB_ENV -Append + - name: Install sccache (shared compiler cache) + if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != '' + continue-on-error: true + shell: pwsh + # Plain cl.exe build, same as the Vulkan/OpenCL jobs. + # build.bat probes sccache before trusting it and, should a build fail with it as the + # launcher, retries once uncached -- so this can speed the job up but never red it. + run: | + $ver = "0.18.0" + $rel = "sccache-v$ver-x86_64-pc-windows-msvc" + $url = "https://github.com/mozilla/sccache/releases/download/v$ver/$rel.zip" + Write-Host "Downloading $url" + Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\sccache.zip" + Expand-Archive -Path "$env:RUNNER_TEMP\sccache.zip" -DestinationPath "$env:RUNNER_TEMP\sccache" -Force + Add-Content -Path $env:GITHUB_PATH -Value "$env:RUNNER_TEMP\sccache\$rel" - name: Build libraries shell: cmd # vcpkg toolchain file wires in the OpenCL (incl. cl2.hpp) that ggml-openvino needs. diff --git a/CLAUDE.md b/CLAUDE.md index 499d2f21d..d24a116b8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -240,6 +240,20 @@ FindVulkan-compatible). Because all five Windows build jobs are in the `package` GPU-toolchain failure blocks packaging — the same release-gating policy the Linux-CUDA / Android-OpenCL jobs already follow. +**sccache on every Windows Ninja job.** All ten Ninja build jobs install sccache (x86_64 or the +native `aarch64` release) with the same `USE_CACHE` / `SCCACHE_WEBDAV_*` env; only the two MSVC-classifier +jobs cannot, because the Visual Studio generator ignores compiler launchers. It cannot red a build, by +three guards in `build.bat`: the install step is `continue-on-error`; a probe compiles through sccache +before it is trusted; and — because that probe only proves `cl.exe`, while arm64 builds with `clang-cl`, +ROCm with its own `clang` and SYCL with `icx` — **a configure or build that fails with sccache as the +launcher is retried once from a clean build dir without it**. The retry is unconditional (cmd cannot +tee the output to match an error signature the way `build.sh` does), so a genuine compile error costs +one extra uncached attempt before it fails. Every configure also passes `-DGGML_CCACHE=OFF`: without +it ggml self-enables any sccache it finds on `PATH` whenever no launcher is set — exactly the +probe-failed and retry cases — and the "uncached" build silently goes through sccache after all (the +same trap `build.sh`'s retry hit with nvcc). `build.bat` uses no `goto`/labels on purpose: it is +checked out with LF line endings, where cmd's label search is unreliable. + **Local sanity builds** (need MSVC + Ninja on PATH; sccache optional; GPU builds also need the matching SDK): ```bat mvn -q compile @@ -291,8 +305,8 @@ build + `ctest`). It emits to the **canonical** `resources/.../Windows/aarch64/` tree — so it ships in the **default** JAR alongside Windows x86-64 / x86 (like those, it is not a classifier). No Java change was needed: `OSInfo` already maps a Windows-on-ARM JVM (`os.arch=aarch64`) to `Windows/aarch64` (it isn't in `archMapping`, so it falls through `translateArchNameToFolderName`). -sccache is intentionally omitted (the shared install step pulls the x86_64 sccache zip; not worth an -arm64 path for one CPU job — `build.bat` just builds uncached). **Compiler: `clang-cl`, not MSVC +sccache runs here too, from its native `aarch64-pc-windows-msvc` release (it wraps `clang-cl`; see +"sccache on every Windows Ninja job" above). **Compiler: `clang-cl`, not MSVC `cl.exe`.** ggml's `ggml-cpu/CMakeLists.txt` aborts with *"MSVC is not supported for ARM, use clang"* via `if (MSVC AND NOT CMAKE_C_COMPILER_ID STREQUAL "Clang")`; `clang-cl` (LLVM's MSVC-compatible driver) satisfies that guard (compiler id `"Clang"`) while keeping CMake's `MSVC=TRUE`, so the static `/MT` CRT From 1a969ca07791a4910a1f6232d756d14213a785ad Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 10:59:35 +0000 Subject: [PATCH 10/10] Upgrade llama.cpp from b11209 to b11211 Two upstream commits: #29440 (RPC waits on the RDMA completion channel instead of spinning) and #29514 (upstream CI only). Version-only from this project's side: GGML_RPC is OFF here, so the changed transport is never compiled into jllama, and upstream's release.yml is untouched, so the CUDA/ROCm/OpenVINO pins derived from it stay as they are. All eight patches apply unchanged (fresh configure; the applier's stamp lists all eight against d7fb90e8e). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CHANGELOG.md | 6 +++--- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 6 files changed, 15 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b9b2f8004..975b57b94 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -101,9 +101,9 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by where the backend cannot provide it, `OFF` disables it. ### Changed -- **llama.cpp `b11080` → `b11209`, in five steps sized by what they change here.** 129 upstream - commits. Four steps are version-only from this project's side (`b11103`, `b11160`, `b11163`, - `b11209`); the one incompatible change got a step of its own: **`b11104`** (llama.cpp #28690) +- **llama.cpp `b11080` → `b11211`, in six steps sized by what they change here.** 131 upstream + commits. Five steps are version-only from this project's side (`b11103`, `b11160`, `b11163`, + `b11209`, `b11211`); the one incompatible change got a step of its own: **`b11104`** (llama.cpp #28690) lets `--host` take a comma-separated list of addresses and removed `server_http_context::thread` and `::listening_address`. `patches/0007` still applied cleanly there but named both members, so it was refreshed to upstream's new `join()` / `listening_addresses` shape; and diff --git a/CLAUDE.md b/CLAUDE.md index d24a116b8..30ea41fb5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11209** +Current llama.cpp pinned version: **b11211** ## Upgrading CUDA Version @@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11209 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11211 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11209`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11211`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1683,7 +1683,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11209`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11211`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 3f53ec0cf..d443db084 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11209](https://img.shields.io/badge/llama.cpp-%23b11209-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11209) +[![llama.cpp b11211](https://img.shields.io/badge/llama.cpp-%23b11211-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11211) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 9861fd7f3..9e57689ef 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -754,3 +754,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11160–b11163 | patches + upstream verification | **All eight patches apply unchanged.** No patch-target file is in the range. | | b11163–b11209 | 46 commits, 157 files outside `tools/ui`, **5339 KiB** — again dominated by backend kernels (ggml-cpu tiled mul_mat for k-quants #27851, Metal per-dtype FA libraries #29329, CUDA RMS_NORM+SCALE fusion, SYCL sparse FA, OpenCL, hexagon backend sampler). **#24364** adds a model-driven W4A4 precision policy (`llama_prec_policy`), selected from GGUF metadata inside `src/` — no public `llama.h` surface. `common/common.h` gains shared Unicode helpers (**#29415**: `fs_path_to_utf8`, and `utf8_to_wstring`/`wstring_to_utf8` on Windows) and pulls in ``; `common/common.cpp` simplifies `fs_create_directory_with_parents` (#29432). Behavioural fixes worth knowing: K/V and recurrent state are cleaned up after a **failed state restore** (**#27530**, the path behind session load), grammar `token_id` parsing no longer truncates large ids (#29382), fused-QKV tensor split with uneven K/V heads (#29294), and #29437 **reverts** #28849's change to the max context length chosen by `--fit` with a unified KV cache. `vendor/cpp-httplib` moves to **0.58.0** (#29407). Jinja gains `sameas` and argument-taking test statements (#29448, #29443). | **No project-source change.** `common.h`'s additions are new free functions and an include; nothing in `src/main/cpp` names the renamed or new helpers. `tools/mtmd/mtmd-audio.h` and `tools/server/server-common.h` (a `#ifndef _WIN32` around `wake_fd`, #29479) change in ways no project file reaches. The three server-contract greps and the `arg.cpp` option set are unchanged over the whole series. | | b11163–b11209 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range. Patch-target file touched: `src/llama-model.cpp` (#29294, and #29151 earlier in the series — `0012`), away from the split-normalisation block. **Standing drop-checks at pristine b11209, all "still required":** `0001` (`common_params_parse_main` 0 occurrences in `common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`, so [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1582`), `0014` (`common_log_set_callback` 0 occurrences in `common/log.h`); `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `187664b53` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty), Release build clean with **0 warnings** — which is what proves the `0007` refresh of the b11103–b11104 step, since only a compile can — `ctest` **559/559**, `nm -D` 40 `Java_*` exports, and `mvn clean verify` **1774 run, 0 failures, 0 errors** (272 skipped, the model-gated classes; no GGUF in this sandbox) with `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (`nativeBuildInfoMatchesPinnedVersionConstant` confirms `b11209` against the linked `build-info`), `LlamaLoggerTest` 6/6 against the real library, spotless and SpotBugs clean. Model-backed Java tests were not run (HF-blocked sandbox). The four intermediate pins were checked by patch replay and header diff, not by a build of their own. | +| b11209–b11211 | Two commits, 5 files, 69 lines: **#29440** (RPC: wait on the RDMA completion channel instead of spinning, `ggml/src/ggml-rpc/transport{,-apple}.cpp`) and **#29514** (upstream CI only: `GGML_SCHED_DEBUG_REALLOC=1` for the ctest workflows). Version-only from this project's side: `GGML_RPC` is `OFF` here, so the one source change is never compiled into `jllama`. The upstream `release.yml` is untouched, so the CUDA (13.4) / ROCm (10.0.0) / OpenVINO pins derived from it are unchanged. | +| b11209–b11211 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir; the applier's stamp lists all eight against `d7fb90e8e`). No patch-target file is in the range. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 1b6d7a00a..0917fa7c4 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11209 + GIT_TAG b11211 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 8adab355b..8abe339e7 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11209"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11211"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11209-"} — call + * plus the resolved upstream commit, e.g. {@code "b11211-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11209"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11211"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11209"; + public static final String LLAMA_CPP_VERSION = "b11211"; // Constants holder — not instantiable. private LlamaCppVersion() {}