From 9e68122fb0507ec2d9ec8fd45fe0e7dd6ffc1878 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 17 Sep 2026 06:32:30 +0000 Subject: [PATCH 1/5] feat: upgrade llama.cpp from b10988 to b11002 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit First chunk toward b11012. 14 commits, 75 KiB. Two files on the priority review list, both implementation-only behind unchanged signatures — no compile or link consequence, but one of them changes behaviour. common/fit.cpp (#28849, "Change max context length for auto-fitting with unified KV") is the one worth knowing. common_params_fit_impl now sizes n_ctx_max by n_seq_max rather than n_streams, so an auto-sized context (-c 0) with --parallel > 1 and a UNIFIED KV cache gets n_ctx_train * n_seq_max instead of n_ctx_train; kv_unified stays n_streams == 1, so the non-unified path is unchanged. The file is compiled into llama-common and linked into jllama, and nothing in this project calls it directly — it is reached through common_init_from_params when n_ctx == 0. So the effect here is "an auto-sized parallel server may now ask for more context", which is upstream's intended fix, not a regression to work around. common/parsers/qwen3-coder.cpp (#28869) adds "\n" ahead of "" in thinking_end_tags so the newline lands inside the forced message. Parser internals; no API surface. Everything else is backend or upstream tooling: CUDA/HIP im2col access patterns, HIP MoE tile heuristic and AllReduce, Vulkan MUL_MAT_ID tail, Metal mul_mm_id NaN, hexagon copy/DMA paths, spacemit int16 transpose, an RPC compute-graph cache invalidation, qwen4exp hc ops (src/models/), and llama-bench --version (never compiled here — LLAMA_BUILD_TOOLS is OFF). No tools/server/ file moved, so the three mechanical server-contract greps have no input. No patch-target file is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 93acbd5f2..51188afdc 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b10988** +Current llama.cpp pinned version: **b11002** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b10988 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11002 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10988`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11002`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10988`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11002`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index fd8f8410d..d35f54082 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b10988](https://img.shields.io/badge/llama.cpp-%23b10988-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10988) +[![llama.cpp b11002](https://img.shields.io/badge/llama.cpp-%23b11002-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11002) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 181ae78ea..2c3dc8d34 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b10988 + GIT_TAG b11002 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index ea38ab654..b1d52cdc4 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10988"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11002"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b10988-"} — call + * plus the resolved upstream commit, e.g. {@code "b11002-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10988"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11002"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b10988"; + public static final String LLAMA_CPP_VERSION = "b11002"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 132a5d012d9649fc19fb7be6bd70702ff63714da Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 17 Sep 2026 06:32:52 +0000 Subject: [PATCH 2/5] feat: upgrade llama.cpp from b11002 to b11005 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Second chunk toward b11012. 3 commits, 33 KiB, dominated by #27625 which adds a new architecture, HrmTextForCausalLM (DFM Mimir 1B). This is the chunk that touches patches/0012's targets, src/llama-model.{cpp,h}, so the collision risk was checked rather than assumed. The three edits to llama-model.cpp are new-arch dispatch entries — an LLM_ARCH_HRM_TEXT case in llama_model_mapping, a MIRRORED meta-split branch for its aliased cache slots, and a rope-type case — at lines ~316, ~477 and ~3030. 0012's hunks are the load_tensors split arithmetic at 1493-1511. They are a thousand lines apart and cannot interact; llama-model.h likewise gains only an additive hrm_z_l_init tensor member. 0012's standing drop-check was re-run against pristine b11005 for the same reason: splits[i] /= split_sum is still bare, with no zero-sum guard, so the patch is still required and is not droppable here. The rest of the arch addition is upstream-internal and never reaches this project's surface: llama-arch.{cpp,h} (enum + tensor names), llama-hparams.h, llama-context.cpp, llama-model-saver.cpp, src/models/hrm-text.cpp and models.h. None is a header this project includes. Also here: #28989 lets Nemotron-H models define only layer_norm_epsilon, and #28995 has hexagon accept the zeroed rope probe in supports_op. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 51188afdc..0e2a52600 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11002** +Current llama.cpp pinned version: **b11005** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11002 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11005 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11002`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11005`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11002`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11005`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index d35f54082..97992f9c7 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11002](https://img.shields.io/badge/llama.cpp-%23b11002-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11002) +[![llama.cpp b11005](https://img.shields.io/badge/llama.cpp-%23b11005-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11005) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 2c3dc8d34..cb34d4698 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11002 + GIT_TAG b11005 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index b1d52cdc4..19ec45893 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11002"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11005"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11002-"} — call + * plus the resolved upstream commit, e.g. {@code "b11005-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11002"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11005"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11002"; + public static final String LLAMA_CPP_VERSION = "b11005"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 4c12869b2fd074486594dd24b0e1f3aaa22afda3 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 17 Sep 2026 06:33:21 +0000 Subject: [PATCH 3/5] feat: upgrade llama.cpp from b11005 to b11011 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Third chunk toward b11012. 6 commits, 88 KiB. Nothing under common/, include/, tools/server/ or tools/mtmd/, so every row of the API-compatibility table is vacuously satisfied and the server-contract greps have no input. The three src/ files are upstream-internal. #28549 (CUDA graph for MTP draft) splits llama_context::gf_res_prev into a two-element array so batches with and without outputs get distinct CUDA graph cache keys, and adds a private get_gf_res_prev(); that is a private member of llama_context declared in src/llama-context.h, which is an INTERNAL header this project does not include. The one internal upstream header it does include is src/llama-model.h — from test_model_split.cpp, via the include dir patches/0012 adds — and that file does not move in this chunk. #28965 fixes tensor-parallel split state and granularity for fused QKV on gemma4/qwen35 (llama-model.cpp). The rest is backend-side: #28981 fixes the ggml_backend_sycl_split_buffer_type signature (it gains a leading int main_device), #28975 works around an NVIDIA bug in argsort_large.comp, #28994 adds hexagon Q4_K/Q6_K support, and #28959 is upstream CI. The SYCL signature change is a public ggml-sycl.h symbol but a backend-internal one — nothing in this project calls any ggml backend split-buffer API, and the three sycl-* classifier jobs compile upstream's own consistent tree. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 0e2a52600..c6e4289fd 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11005** +Current llama.cpp pinned version: **b11011** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11005 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11011 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11005`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11011`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11005`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11011`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 97992f9c7..a4102c1fb 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11005](https://img.shields.io/badge/llama.cpp-%23b11005-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11005) +[![llama.cpp b11011](https://img.shields.io/badge/llama.cpp-%23b11011-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11011) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index cb34d4698..265e8818f 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11005 + GIT_TAG b11011 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 19ec45893..729ebee86 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11005"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11011"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11005-"} — call + * plus the resolved upstream commit, e.g. {@code "b11011-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11005"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11011"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11005"; + public static final String LLAMA_CPP_VERSION = "b11011"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 92515d2d6c61443c3ee64f4aa36faa84ce94a4cc Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 17 Sep 2026 06:33:21 +0000 Subject: [PATCH 4/5] feat: upgrade llama.cpp from b11011 to b11012 Fourth and final chunk, reaching the target release. One commit, 13 KiB: #28988 adds Vulkan support for the qwen4exp hc ops that #28901 introduced earlier in this range (chunk 1), so the two Vulkan classifiers can run that architecture instead of falling back. Entirely ggml/src/ggml-vulkan/** plus its shader generator. Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, and no patch target. Kept as its own chunk rather than folded into b11005 -> b11011: that would have made one 100.6 KiB step, over the runbook's threshold by a hair. "Barely over" is exactly the rationalisation the threshold exists to prevent, and the cost of honouring it here is one extra commit, not an extra build. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c6e4289fd..d19da9722 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11011** +Current llama.cpp pinned version: **b11012** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11011 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11012 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11011`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11012`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11011`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11012`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index a4102c1fb..d6714f996 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11011](https://img.shields.io/badge/llama.cpp-%23b11011-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11011) +[![llama.cpp b11012](https://img.shields.io/badge/llama.cpp-%23b11012-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11012) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 265e8818f..a99a20b18 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11011 + GIT_TAG b11012 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 729ebee86..c4038e1e6 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11011"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11012"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11011-"} — call + * plus the resolved upstream commit, e.g. {@code "b11012-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11011"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11012"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11011"; + public static final String LLAMA_CPP_VERSION = "b11012"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 845d7eea5a96e2230af976e356fece5013f60c6a Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 17 Sep 2026 06:40:08 +0000 Subject: [PATCH 5/5] docs: record the b10988-b11012 range MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two rows per the runbook's step 4. The first covers what moved: the four-chunk walk and why the 13 KiB tail was kept separate, the two common/ files on the review surface (with #28849's auto-context behaviour change spelled out), the new HRM-TEXT architecture, and the two ggml/include headers — one purely additive, one a real but backend-internal signature break. The second is the patch row, and it matters more than usual: this is the first range in a while where a patch target actually moved. It records the collision check against patches/0012 by line, the note that 0012's anchor shifted 1493 -> 1518 under the new arch code (which is why drop-checks are by content, not by line), the six standing drop-checks against the pristine tag, and the full verification numbers. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP --- docs/history/llama-cpp-breaking-changes.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 00b76ce1e..0cb929df1 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -734,3 +734,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10969–b10976 | patches + upstream verification | End of the two-chunk walk, and **the patch set is unchanged at nine** — nothing dropped, nothing refreshed. Verified against the pristine target rather than inferred: `git apply -p1` of all nine, in filename order, into a clean `b10976` worktree, every one clean. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10976:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10976:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10976:src/llama-model.cpp:1493`, no zero-sum guard), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`/`llama_server_attach` and `LLAMA_SERVER_WORKER_CMD` all absent upstream). The whole b10948→b10976 range leaves **every** patch target except `src/llama-model.cpp` and `tests/CMakeLists.txt` byte-identical, and both of those move only at a distance from the patched regions. Verified end-to-end for real: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `987498f4592a76897863cf53711dce38380c082b` (= `b10976`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors; `ctest` **537/537**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** — the check that cross-validates the bumped `LLAMA_CPP_VERSION` constant against the linked `build-info`, needing the `clean` because the constant is inlined into the already-compiled test class. | | b10976–b10988 | 12 commits, **323 KiB** — and the size is the story, so the chunking decision is recorded rather than taken quietly. The range was walked in **three** commits: `b10976→b10980` (54 KiB), `b10980→b10981` (191 KiB) and `b10981→b10988` (78 KiB). The middle step is the one that breaks the runbook's 100 KiB rule, and it is **irreducible**: it is a *single* upstream commit, **#28638** ("OpenVINO: optimize stateful decode and GPU MoE inference"), and there is no intermediate `b` tag inside one commit, so no smaller step exists to take. 35 of its 37 files are `ggml/src/ggml-openvino/**`; the other two are `ci/run.sh` and `docs/backend/OPENVINO.md`. Across the **whole** range, `ggml/src` accounts for 293 KiB of the 323 — the remainder is `src/models` 4.6 KiB, `tests/` 5.5 KiB (never compiled here), `.github` 14 KiB of upstream's own CI, `docs`/`ci`/`CONTRIBUTING.md` 5.3 KiB, and `ggml/include` **0.3 KiB**. Backend work by vendor: **#28599** Metal FA kernels for HSK=96/HSV=64 (MiniCPM3), **#28881** a generic OpenCL `ssm_scan`, **#27637** OpenCL MoE expert-matmul selection by batch size, **#28105** Vulkan sparse flash attention (5 shaders + `ggml-vulkan.cpp`), **#26308** CUDA row-contiguous `SUM_ROWS`, **#28789** RPC weight-only hash caching. The one non-backend change is **#28934**, pure code motion: `build_arch_graph()` moves below the `graph()` template specializations in `src/models/{dflash,eagle3,t5}.cpp`. 67 files, 2828 insertions, 892 deletions. | **No project source change, and zero files on the review surface — literally zero.** Nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/` moved at all, so every row of the priority API-compatibility table is vacuously satisfied and the three mechanical `tools/server` contract greps have **no input to compare**: the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have moved. The three `src/models/*.cpp` files are internal upstream TUs and #28934 moves no signature. **One "safe to skip" header does move and was checked rather than waved past**: `ggml/include/ggml-rpc.h` bumps `RPC_PROTO_MAJOR_VERSION` 6 → 7. It is inert here — `GGML_RPC` appears nowhere in `llama/CMakeLists.txt`, `publish.yml`, `build.sh` or `build.bat`, so `ggml-rpc` is never built and the wire protocol it versions is never spoken. The practical risk of the 191 KiB OpenVINO step is correspondingly narrow: `llama/CMakeLists.txt` routes `GGML_OPENVINO` to the `resources_linux_openvino` / `resources_windows_openvino` **classifier** trees only — it is not in the default JAR — and both `openvino-*` jobs are build-only on GPU-less runners, so "must still compile" is the whole of it, and CI checks that directly. | | b10976–b10988 | patches + upstream verification | **The patch set is unchanged at nine — nothing dropped, nothing refreshed, and nothing even had to be re-examined.** Not one patch-target file is touched anywhere in the range: `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.{cpp,h}`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are all byte-unchanged across b10976→b10988, verified by diffing those paths explicitly rather than inferred from the aggregate. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10988:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10988:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10988:src/llama-model.cpp:1493`, no zero-sum guard). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `9f31776c3773cf03f98535c19b7e6d394af374b4` (= `b10988`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (the `clean` is load-bearing — the `LLAMA_CPP_VERSION` constant is inlined into the already-compiled test class, so without it the cross-check against the linked `build-info` compares the old value); full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **One defect was found by this bump rather than by the range**: `.github/verify-patches-applied.sh` counted the stamp's patch lines as "total lines minus one", which silently went stale when the applier's content oracle added a second metadata line (`tree `) in the next commit of the session that introduced the script. The guard therefore failed on every correct tree and would have redded `C++ Tests` on the next pipeline run; it is fixed in this branch by counting patch lines by their own shape instead of by subtraction. | +| b10988–b11012 | 24 commits, **206 KiB**, 65 files — over the runbook's threshold, walked in **four** chunks: `b10988→b11002` (75 KiB / 14), `b11002→b11005` (33 KiB / 3), `b11005→b11011` (88 KiB / 6) and `b11011→b11012` (13 KiB / 1). The last was kept separate on purpose: folding it into the third makes one **100.6 KiB** step, over by a hair, and "barely over" is the rationalisation the threshold exists to prevent — the cost of honouring it is one extra commit, not an extra build. **Two files on the priority review list**, both `common/` implementation behind unchanged signatures. **#28849** (`common/fit.cpp`) is the one with teeth: `common_params_fit_impl` now sizes `n_ctx_max` by `n_seq_max` rather than `n_streams`, so an auto-sized context (`-c 0`) with `--parallel > 1` and a **unified** KV cache gets `n_ctx_train * n_seq_max` instead of `n_ctx_train`; `kv_unified` keeps `n_streams == 1`, so the non-unified path is unchanged. **#28869** (`common/parsers/qwen3-coder.cpp`) puts `"\n"` ahead of `""` in `thinking_end_tags` so the newline lands inside the forced message. **#27625** adds a whole architecture, `HrmTextForCausalLM` (DFM Mimir 1B), across `src/llama-arch.{cpp,h}`, `llama-hparams.h`, `llama-context.cpp`, `llama-model-saver.cpp`, `src/models/hrm-text.cpp` and — the part that matters here — `src/llama-model.{cpp,h}`. **#28549** splits `llama_context::gf_res_prev` into a two-element array so batches with and without outputs get distinct CUDA-graph cache keys. The remaining ~33 `ggml/src` files are CUDA/HIP im2col and MoE heuristics, HIP AllReduce, Vulkan MUL_MAT_ID / argsort / qwen4exp hc ops, Metal `mul_mm_id` NaN, hexagon copy/DMA and K-quants, spacemit int16 transpose, and an RPC graph-cache invalidation. | **No project source change.** Neither `common/` change is a compile or link consequence — both are implementation-only in TUs upstream compiles into `llama-common`, and nothing here calls `common_params_fit_impl` directly (it is reached through `common_init_from_params` when `n_ctx == 0`), so #28849 surfaces as "an auto-sized parallel server may now ask for more context", upstream's intended fix rather than a regression to absorb. **Zero `tools/server/` files moved**, so the three mechanical server-contract greps have no input and the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have changed. **Two `ggml/include` headers do move and both were opened rather than waved past**: `ggml.h` is purely additive (a new `ggml_dsv4_hc_pre_gated` plus one comment line — no existing signature moves, and this project calls no `ggml_dsv4_*`), and `ggml-sycl.h`'s `ggml_backend_sycl_split_buffer_type` gains a leading `int main_device` — a real signature break, but of a backend-internal symbol no project code calls; the three `sycl-*` classifier jobs compile upstream's own self-consistent tree. `src/llama-context.h` is an **internal** header this project does not include (the only internal one it does is `src/llama-model.h`, from `test_model_split.cpp`), and #28549's change there is a private member. | +| b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. |