From 80c08dca2715db477c9e160a7674691f087b8869 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 12:54:58 +0000 Subject: [PATCH] Upgrade llama.cpp from b11211 to b11214 Three upstream commits (HIP fattn-mma on CDNA, Vulkan argsort for Adreno, and #29516: an llguidance grammar without llguidance now throws instead of GGML_ABORT, so it no longer kills the JVM). No API or workflow change; all eight patches apply unchanged. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CHANGELOG.md | 6 ++++++ CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 6 files changed, 18 insertions(+), 10 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 317782d6f..2474f4b28 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -70,6 +70,12 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by without a native library or model. ### Changed +- **llama.cpp `b11211` → `b11214`.** Three upstream commits, version-only from this project's side: + a HIP flash-attention kernel choice for CDNA (#28907), a Vulkan argsort fix for Adreno (#29469), and + **#29516**, which makes `common_sampler_init` *throw* `failed to parse grammar: llguidance is not + enabled` instead of calling `GGML_ABORT` when a `%llguidance` grammar reaches a build without + llguidance. That one matters in a JVM: these builds do not enable llguidance, so such a grammar used + to abort the whole process; it now surfaces as an ordinary request error. All patches apply unchanged. - **BREAKING (runtime): the shipped SLF4J binding is now `slf4j-simple`, not `logback-classic`.** Two independent reasons, and the first is a hard failure rather than a preference: diff --git a/CLAUDE.md b/CLAUDE.md index 0ad7a342d..b0d981ddd 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11211** +Current llama.cpp pinned version: **b11214** ## Upgrading CUDA Version @@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11211 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11214 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11211`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11214`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1683,7 +1683,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11211`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11214`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 08150937e..9d39dbcbf 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11211](https://img.shields.io/badge/llama.cpp-%23b11211-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11211) +[![llama.cpp b11214](https://img.shields.io/badge/llama.cpp-%23b11214-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11214) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 9e57689ef..45af228ed 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -756,3 +756,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11163–b11209 | patches + upstream verification | **All eight patches apply unchanged** at every tag of the range. Patch-target file touched: `src/llama-model.cpp` (#29294, and #29151 earlier in the series — `0012`), away from the split-normalisation block. **Standing drop-checks at pristine b11209, all "still required":** `0001` (`common_params_parse_main` 0 occurrences in `common/arg.h`; the count-guarded `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`, so [ggml-org/llama.cpp#26416](https://github.com/ggml-org/llama.cpp/issues/26416) remains open), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1582`), `0014` (`common_log_set_callback` 0 occurrences in `common/log.h`); `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure** (`rm -rf build && cmake -B build -DBUILD_TESTING=ON`, the real `FetchContent` path): stamp head `187664b53` with **eight** SHA-256 lines, `verify-patches-applied.sh` green (8 applied, tree dirty), Release build clean with **0 warnings** — which is what proves the `0007` refresh of the b11103–b11104 step, since only a compile can — `ctest` **559/559**, `nm -D` 40 `Java_*` exports, and `mvn clean verify` **1774 run, 0 failures, 0 errors** (272 skipped, the model-gated classes; no GGUF in this sandbox) with `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (`nativeBuildInfoMatchesPinnedVersionConstant` confirms `b11209` against the linked `build-info`), `LlamaLoggerTest` 6/6 against the real library, spotless and SpotBugs clean. Model-backed Java tests were not run (HF-blocked sandbox). The four intermediate pins were checked by patch replay and header diff, not by a build of their own. | | b11209–b11211 | Two commits, 5 files, 69 lines: **#29440** (RPC: wait on the RDMA completion channel instead of spinning, `ggml/src/ggml-rpc/transport{,-apple}.cpp`) and **#29514** (upstream CI only: `GGML_SCHED_DEBUG_REALLOC=1` for the ctest workflows). Version-only from this project's side: `GGML_RPC` is `OFF` here, so the one source change is never compiled into `jllama`. The upstream `release.yml` is untouched, so the CUDA (13.4) / ROCm (10.0.0) / OpenVINO pins derived from it are unchanged. | | b11209–b11211 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir; the applier's stamp lists all eight against `d7fb90e8e`). No patch-target file is in the range. | +| b11211–b11214 | Three commits, 5 files, 29 lines: **#28907** (HIP: fattn-mma on CDNA for `dkq > 256` at large batch, `ggml-cuda/fattn*.cu[h]` + `scripts/hip/gcn-cdna-vgpr-check.py`), **#29469** (Vulkan argsort kernel selection for Adreno) and **#29516** (`common/sampling.cpp`: an llguidance grammar without `LLAMA_LLGUIDANCE` now throws `std::runtime_error` instead of `GGML_ABORT` — a process abort turned into a catchable error, which for a JNI host is the difference between a failed request and a dead JVM). Version-only from this project's side: no API change, no workflow/ROCm/CUDA-component change upstream in the range. | +| b11211–b11214 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir). No patch-target file is in the range. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 0917fa7c4..56c67d23e 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11211 + GIT_TAG b11214 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 8abe339e7..6ace75120 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11211"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11214"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11211-"} — call + * plus the resolved upstream commit, e.g. {@code "b11214-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11211"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11214"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11211"; + public static final String LLAMA_CPP_VERSION = "b11214"; // Constants holder — not instantiable. private LlamaCppVersion() {}