From a581a9acaf27169bb8c975853f09dc0b49d1d3ab Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:55:57 +0000 Subject: [PATCH 01/12] feat: upgrade llama.cpp from b11018 to b11020 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit First chunk toward b11062. 2 commits, 10 KiB. #28993 ("gguf : align the data section relative to the GGUF start, not the file") is the substantive one: gguf.cpp now computes the data-section padding from the GGUF's own start offset rather than the file offset, so a GGUF embedded at a non-zero offset in a larger file stays internally consistent. llama-model-loader.cpp gains the matching guard — loading through a FILE* with mmap now throws when the data section is not aligned to the CPU tensor alignment — and include/llama.h grows one new entry point, llama_adapter_lora_init_from_file_ptr, implemented in llama-adapter.cpp. Both are additive: the project loads models and LoRA adapters by path, never by FILE*, so nothing here is on a path jllama reaches. #29008 adds message_delimiters to the DeepSeek V3.2/V4 parser so the server can locate user turns for context checkpoints. Parser data only, no API surface. No file under tools/server/, so the three mechanical server-contract greps have no input, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 1e1b6f612..99b556149 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11018** +Current llama.cpp pinned version: **b11020** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11018 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11020 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11018`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11020`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11018`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11020`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 5054455d9..aa84e391e 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11018](https://img.shields.io/badge/llama.cpp-%23b11018-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11018) +[![llama.cpp b11020](https://img.shields.io/badge/llama.cpp-%23b11020-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11020) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 028b45081..e5a7acc6c 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11018 + GIT_TAG b11020 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 135013783..e76e3acc2 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11018"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11020"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11018-"} — call + * plus the resolved upstream commit, e.g. {@code "b11020-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11018"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11020"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11018"; + public static final String LLAMA_CPP_VERSION = "b11020"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 57328202b60206867641ddffc08e64edf23aae8c Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:56:32 +0000 Subject: [PATCH 02/12] feat: upgrade llama.cpp from b11020 to b11022 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Second chunk. 2 commits, 579 KiB — well over the runbook's 100 KiB threshold, and irreducibly so: b11021 does not exist as a tag, so #28732 cannot be split into a smaller step. The size is also almost entirely bookkeeping. #28732 ("vulkan: split buffers and debug code into separate files, add shared headers") moves ~5.4k lines out of ggml-vulkan.cpp into ggml-vulkan-buffers.cpp, ggml-vulkan-debug.cpp and three new headers (types, push-constants, common). Diffed as a rename-aware move it is a pure reorganisation: ggml-vulkan/CMakeLists.txt gains exactly the five new files in the ggml_add_backend_library() call and nothing else changes about how the backend is built or linked, so the two Vulkan classifiers (vulkan-linux-*, vulkan-windows-x86-64) are unaffected beyond compiling more, smaller translation units. #28947 adds an API/ABI check to upstream's own make-release workflow: a GitHub workflow plus three scripts/ files. Never executed by this project's CI. Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, so every row of the API-compatibility table is vacuously satisfied and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 99b556149..d4cb49d00 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11020** +Current llama.cpp pinned version: **b11022** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11020 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11022 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11020`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11022`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11020`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11022`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index aa84e391e..668aaaa45 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11020](https://img.shields.io/badge/llama.cpp-%23b11020-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11020) +[![llama.cpp b11022](https://img.shields.io/badge/llama.cpp-%23b11022-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11022) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index e5a7acc6c..8434fc8cf 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11020 + GIT_TAG b11022 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index e76e3acc2..0ba437da7 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11020"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11022"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11020-"} — call + * plus the resolved upstream commit, e.g. {@code "b11022-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11020"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11022"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11020"; + public static final String LLAMA_CPP_VERSION = "b11022"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 74b03858e0a70e18eb650b4304a0760b0b8b464b Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:57:30 +0000 Subject: [PATCH 03/12] feat: upgrade llama.cpp from b11022 to b11024 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Third chunk. 2 commits, 224 KiB — over the threshold, and again irreducibly: b11023 does not exist as a tag, so #29009 cannot be split. #29009 ("openvino : Update OpenVINO to 2026.4; fix clangd, MSVC warnings") is the whole of the size. It rewrites large parts of ggml-openvino's quant and utils translation units, but every symbol it touches is inside ggml/src/ggml-openvino/**: no ggml public header, no llama header, nothing this project compiles against. The remaining files are upstream's own release/ self-hosted workflows and docs/backend/OPENVINO.md. #27985 fixes the reasoning menu in the WebUI's single-model desktop view. The WebUI auto-follows GIT_TAG — build-webui re-reads the tag and rebuilds the matching Svelte UI — so it needs no action here. Watch item, recorded rather than acted on: upstream's own OpenVINO jobs move their SDK pin 2026.3.1 -> 2026.4, while this project's two OpenVINO classifier jobs still install the 2026.2.1 archive (publish.yml). That gap already existed one release back and the jobs built, and nothing in this diff uses an API newer than what 2026.2.1 exposes (the ov:: surface it adds is Core/CompiledModel/ InferRequest/Tensor, all long-standing). Per the classifier policy these vendor install steps are first-pass and fail loud, so if 2026.2.1 stops compiling ggml-openvino the job reds the pipeline and the pin gets moved then — bumping it speculatively here cannot be validated on a GPU-less runner. Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index d4cb49d00..57201eee3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11022** +Current llama.cpp pinned version: **b11024** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11022 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11024 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11022`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11024`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11022`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11024`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 668aaaa45..ec60a5a76 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11022](https://img.shields.io/badge/llama.cpp-%23b11022-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11022) +[![llama.cpp b11024](https://img.shields.io/badge/llama.cpp-%23b11024-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11024) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 8434fc8cf..e92ba8e1a 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11022 + GIT_TAG b11024 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 0ba437da7..d98f005e6 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11022"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11024"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11022-"} — call + * plus the resolved upstream commit, e.g. {@code "b11024-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11022"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11024"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11022"; + public static final String LLAMA_CPP_VERSION = "b11024"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 6f8dc58f98d12107681f7c0e7ea8fa788eb0d98f Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:58:21 +0000 Subject: [PATCH 04/12] feat: upgrade llama.cpp from b11024 to b11042 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fourth chunk, and the first back under the threshold: 18 commits, 87 KiB. Two themes on the review surface, both additive: Allocation-failure hardening (#28149, #26070, #28978). ggml now checks allocation results instead of assuming success, and graph-buffer reservation failure is handled rather than fatal; tools/mtmd/clip.cpp picks up the matching call-site change — clip_encode() returns false when ggml_backend_sched_alloc_graph() fails instead of proceeding with an unallocated graph. No ggml public header moved anywhere in this range (ggml/include is byte-identical b11018..b11062), and no project C++ calls the scheduler or allocator directly, so this reaches jllama only as better-behaved-on-OOM inside upstream translation units. Model plumbing (#29042, #29014, #29018, #29033). llama_model_base gains load_swa_pattern() — a helper that reads the SWA pattern either as one flag per layer or as a period to expand — and 20 src/models/*.cpp architectures are rewritten onto it; create_tensor_gate_up_exps() learns to honour TENSOR_SKIP; llama-model-saver writes the SWA pattern so 15 more architectures round-trip; and llama-vocab adds LLAMA_VOCAB_PRE_TYPE_UFAKZEKA = 59. All of it is inside src/, compiled into the static llama library, none of it visible in include/llama.h. llama-model.{cpp,h} are patch 0012's target files, so this chunk is the first that could have disturbed a patch. It does not: load_swa_pattern() lands ~1700 lines above load_tensors()'s split arithmetic, and 0012 still applies clean. The rest is backends and upstream CI: an OpenCL bin kernel, F16 FWHT on CPU, Vulkan IQ3_S MMQ kernels and a raised mul_mat_id expert limit, an RPC ACCEL skip, a GGML_CPU=OFF/GGML_CUDA=ON cmake fix, a gguf-py Q8_1 block-size fix, and five workflow tweaks. No file under tools/server/, so the three mechanical server-contract greps have no input. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 57201eee3..259fba0af 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11024** +Current llama.cpp pinned version: **b11042** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11024 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11042 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11024`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11042`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11024`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11042`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index ec60a5a76..db8f75e54 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11024](https://img.shields.io/badge/llama.cpp-%23b11024-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11024) +[![llama.cpp b11042](https://img.shields.io/badge/llama.cpp-%23b11042-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11042) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index e92ba8e1a..366b68a04 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11024 + GIT_TAG b11042 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index d98f005e6..49045fb27 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11024"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11042"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11024-"} — call + * plus the resolved upstream commit, e.g. {@code "b11042-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11024"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11042"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11024"; + public static final String LLAMA_CPP_VERSION = "b11042"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 670dd09b044ce712e280612d4ff94d756295ccc0 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:58:38 +0000 Subject: [PATCH 05/12] feat: upgrade llama.cpp from b11042 to b11045 Fifth chunk. 3 commits, 98 KiB, all three in the Hexagon backend (ggml/src/ggml-hexagon/**): ROLL op support (#29105), an im2col update (#29103), and HMX flash-attention head_dim padding for DK=DV=72 (#26539). Nothing under common/, include/, src/, tools/ or the project CMakeLists, so every row of the API-compatibility table is vacuously satisfied, the three mechanical server-contract greps have no input, and no patch target is in this chunk. This project ships no Hexagon classifier, so the code is not even compiled here. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 259fba0af..ee634f2ec 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11042** +Current llama.cpp pinned version: **b11045** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11042 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11045 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11042`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11045`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11042`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11045`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index db8f75e54..e1ecedfec 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11042](https://img.shields.io/badge/llama.cpp-%23b11042-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11042) +[![llama.cpp b11045](https://img.shields.io/badge/llama.cpp-%23b11045-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11045) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 366b68a04..4f9a4c906 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11042 + GIT_TAG b11045 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 49045fb27..f7e454842 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11042"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11045"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11042-"} — call + * plus the resolved upstream commit, e.g. {@code "b11045-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11042"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11045"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11042"; + public static final String LLAMA_CPP_VERSION = "b11045"; // Constants holder — not instantiable. private LlamaCppVersion() {} From d5464f03c95297a58d5005f5cdca6d368f348dfb Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:59:05 +0000 Subject: [PATCH 06/12] feat: upgrade llama.cpp from b11045 to b11050 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sixth chunk. 5 commits, 57 KiB. One line of public API moves: #29084 adds LLAMA_VOCAB_TYPE_TEST = 7 to include/llama.h's llama_vocab_type enum, with the tokenizer itself in llama-vocab.cpp (a rolling hash over fixed-size chunks, for generating a dummy vocab in test-llama-archs). Purely additive, and the value is a tail append, so no existing constant renumbers. This project surfaces vocab_type as a raw int in two places — jllama.cpp's two "vocab_type" emit sites, both already static_cast-ed, and ModelMeta.getVocabType() which reads it back as an int — so a new enumerator needs no Java-side constant and cannot go stale. (The b10585 common_json enum trap is unrelated and already guarded; nothing here changes it.) llama-model-saver.cpp continues the #29042 round-trip work from chunk 4. The remaining four are backends: Metal FA support checks (#29122) and qwen4exp hc ops (#29000), a CUDA CUB argsort in-place-keys corruption fix (#28389 — real bug fix, reaches the cuda13 classifiers), and an OpenCL flash_attn bin kernel (#29046). No file under tools/server/, so the three mechanical server-contract greps have no input, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index ee634f2ec..fbd10dd1a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11045** +Current llama.cpp pinned version: **b11050** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11045 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11050 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11045`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11050`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11045`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11050`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index e1ecedfec..591154934 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11045](https://img.shields.io/badge/llama.cpp-%23b11045-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11045) +[![llama.cpp b11050](https://img.shields.io/badge/llama.cpp-%23b11050-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11050) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 4f9a4c906..3cb6bc340 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11045 + GIT_TAG b11050 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index f7e454842..0d758af1f 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11045"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11050"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11045-"} — call + * plus the resolved upstream commit, e.g. {@code "b11050-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11045"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11050"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11045"; + public static final String LLAMA_CPP_VERSION = "b11050"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 53c9d3bc1055fd334839ad4c11f3f32f84d4e20c Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 09:59:40 +0000 Subject: [PATCH 07/12] feat: upgrade llama.cpp from b11050 to b11052 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Seventh chunk. 2 commits, 115 KiB — over the threshold, irreducibly again (b11051 does not exist as a tag). #29127 is the one line on the review surface and it is a behaviour fix worth naming: common/json-schema-to-grammar.cpp's gbnf_escape_length() now accepts '-' as an escapable character, matching parse_char() in llama-grammar.cpp. A JSON-schema "pattern" containing \- previously produced a grammar the parser then rejected, so a structured-output request using such a pattern failed. This project reaches it through the grammar routing in eval_llama_cmpl_schema(), so it is a real (if narrow) fix for callers passing json_schema/response_format. #28948 is the rest of the bytes: Metal MoE and SSM_CONV fusion optimizations plus a new argsort.metal kernel, with test-backend-ops and the MTL fusion CSV updated alongside. That reaches the default macOS arm64 dylib (Metal is in the default jar, not a classifier), so it is exercised by the three macOS Java jobs and the smoke-fatjar-macos gate. No file under tools/server/, so the three mechanical server-contract greps have no input, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index fbd10dd1a..10a38f94e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11050** +Current llama.cpp pinned version: **b11052** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11050 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11052 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11050`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11052`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11050`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11052`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 591154934..2dcc096fe 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11050](https://img.shields.io/badge/llama.cpp-%23b11050-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11050) +[![llama.cpp b11052](https://img.shields.io/badge/llama.cpp-%23b11052-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11052) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 3cb6bc340..885c313aa 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11050 + GIT_TAG b11052 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 0d758af1f..7a68cd654 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11050"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11052"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11050-"} — call + * plus the resolved upstream commit, e.g. {@code "b11052-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11050"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11052"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11050"; + public static final String LLAMA_CPP_VERSION = "b11052"; // Constants holder — not instantiable. private LlamaCppVersion() {} From c5168fd955d835a49cc3e9d89c7097e3c11bc6c5 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 10:02:51 +0000 Subject: [PATCH 08/12] feat: upgrade llama.cpp from b11052 to b11055, refresh patches 0001 and 0006 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Eighth chunk. 3 commits, 97 KiB. Two are Hexagon ops (GEGLU_QUICK #29114, TOP_K #29113, not compiled here); the third is the one that matters. #29125 ("server : improve startup log messages") is the first commit in this whole range to touch a patch target. It rewrites llama_server()'s startup logging — an "initializing ..." line before the argv parse, the CORS and enabled-features warnings collapsed from multi-line banners to one line each, and the :8080 port notice likewise — plus a TODO comment above common_params_parse() in common/arg.h and SRV_INF -> SRV_TRC / a source-tagged "Available models" listing in server-models.cpp. Two patches went stale on it, both purely on context, and both are refreshed here with no change to what they do: 0001 — its common/arg.h hunk anchored on the two comment lines above the common_params_parse() declaration; upstream inserted two more (the TODO), so the hunk now anchors on the new last comment line. Its tools/server/server.cpp hunk anchored on the three lines above the parse call, one of which is now the new SRV_INF("initializing ..."). 0006 — same server.cpp call site, same cause (its hunk 2 replaces the line 0001 just flipped). Hunks 1 and 3 were untouched; only line offsets moved. The refresh is context-only: `git diff` of the two patch files shows the added/ removed lines byte-identical, with only @@ line numbers, three context lines and the index blob hashes changing. Verified by replaying the whole stack in filename order against pristine b11055 AND pristine b11062 — all nine apply clean at both tags. 0007's standing invariant is intact for the same reason it is checkable at all: its `-` side is a verbatim copy of the route table it factors out of llama_server(), so a clean apply proves upstream did not touch that block. It did not — #29125's edits sit above it (the CORS warning) and below it (the warn_names loop), never inside. tools/server/ is touched, so the three mechanical contract greps are in scope this time. They have no input regardless: server-schema.cpp, server-task.cpp and server-context.cpp are byte-identical between b11018 and b11062 (verified by blob hash, not by diff reading), so the request-field set, the field bounds and the response-key set cannot have moved anywhere in this bump. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../patches/0001-win32-arg-parse-embed-guard.patch | 14 +++++++------- .../0006-server-embed-native-server-jni.patch | 10 +++++----- .../net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 6 files changed, 22 insertions(+), 22 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 10a38f94e..8620066a1 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11052** +Current llama.cpp pinned version: **b11055** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11052 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11055 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11052`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11055`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11052`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11055`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 2dcc096fe..d644f5928 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11052](https://img.shields.io/badge/llama.cpp-%23b11052-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11052) +[![llama.cpp b11055](https://img.shields.io/badge/llama.cpp-%23b11055-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11055) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 885c313aa..6ce9bd0ca 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11052 + GIT_TAG b11055 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/patches/0001-win32-arg-parse-embed-guard.patch b/llama/patches/0001-win32-arg-parse-embed-guard.patch index 8527089ad..56fd186c8 100644 --- a/llama/patches/0001-win32-arg-parse-embed-guard.patch +++ b/llama/patches/0001-win32-arg-parse-embed-guard.patch @@ -44,11 +44,11 @@ index 79480e06f..ed7793b4d 100644 std::vector supported_tmpl; int32_t res = llama_chat_builtin_templates(nullptr, 0); diff --git a/common/arg.h b/common/arg.h -index 8f609e356..62c615d29 100644 +index 203d1b4e1..a70f8882c 100644 --- a/common/arg.h +++ b/common/arg.h -@@ -123,6 +123,11 @@ struct common_params_context { - // if one argument has invalid value, it will automatically display usage of the specific argument (and not the full usage message) +@@ -126,6 +126,11 @@ struct common_params_context { + // this is a side-effect that should be avoided bool common_params_parse(int argc, char ** argv, common_params & params, llama_example ex, void(*print_usage)(int, char **) = nullptr); +// Like common_params_parse(), but first recovers the process command line as UTF-8 argv on @@ -483,12 +483,12 @@ index f2179ed27..6d958a861 100644 } if (params.out_file.empty()) { diff --git a/tools/server/server.cpp b/tools/server/server.cpp -index a3b2a8b0f..80d6a3ff6 100644 +index 1167c0aea..28f18c1bb 100644 --- a/tools/server/server.cpp +++ b/tools/server/server.cpp -@@ -102,7 +102,7 @@ int llama_server(int argc, char ** argv) { - // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free - server_stream_session_manager_start(); +@@ -104,7 +104,7 @@ int llama_server(int argc, char ** argv) { + + SRV_INF("%s", "initializing ...\n"); - if (!common_params_parse(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { diff --git a/llama/patches/0006-server-embed-native-server-jni.patch b/llama/patches/0006-server-embed-native-server-jni.patch index c9e2f4d61..23d73c014 100644 --- a/llama/patches/0006-server-embed-native-server-jni.patch +++ b/llama/patches/0006-server-embed-native-server-jni.patch @@ -1,5 +1,5 @@ diff --git a/tools/server/server.cpp b/tools/server/server.cpp -index 7f9ca414..9c0caf18 100644 +index 28f18c1bb..8aeed4b99 100644 --- a/tools/server/server.cpp +++ b/tools/server/server.cpp @@ -25,6 +25,28 @@ @@ -31,9 +31,9 @@ index 7f9ca414..9c0caf18 100644 static inline void signal_handler(int signal) { if (is_terminating.test_and_set()) { // in case it hangs, we can force terminate the server by hitting Ctrl+C twice -@@ -97,7 +119,13 @@ int llama_server(int argc, char ** argv) { - // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free - server_stream_session_manager_start(); +@@ -104,7 +126,13 @@ int llama_server(int argc, char ** argv) { + + SRV_INF("%s", "initializing ...\n"); - if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + // [jllama] embedded (JNI) callers forward a clean UTF-8 argv, so honor it exactly via @@ -46,7 +46,7 @@ index 7f9ca414..9c0caf18 100644 return 1; } -@@ -433,7 +461,10 @@ int llama_server(common_params & params, int argc, char ** argv) { +@@ -493,7 +521,10 @@ int llama_server(common_params & params, int argc, char ** argv) { } // register signal handler if not running by CLI diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 7a68cd654..951f01aee 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11052"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11055"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11052-"} — call + * plus the resolved upstream commit, e.g. {@code "b11055-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11052"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11055"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11052"; + public static final String LLAMA_CPP_VERSION = "b11055"; // Constants holder — not instantiable. private LlamaCppVersion() {} From a90405a210dbb208682d2555e59102edab14a862 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 10:03:23 +0000 Subject: [PATCH 09/12] feat: upgrade llama.cpp from b11055 to b11062 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Ninth and final chunk, reaching the target release. 7 commits, 84 KiB. Chat parsing is where the substance is, and both items reach this project — common/chat.cpp is compiled into llama-common and the server path runs every completion through common_chat_parse(): #28682 adds a dedicated Ling 3.0 / Bailing V3 parser (common/parsers/ling3.cpp, 194 lines, declared in parsers.h, registered in sources.cmake) and a detection arm in common_chat_try_specialized_template() keyed on "ASSISTANT" + "". Additive: a template that did not match any specialized arm before still does not. #29115 fixes the gemma4 required-tool grammar — with tool_choice == COMMON_CHAT_TOOL_CHOICE_REQUIRED the grammar now ends at the tool call instead of continuing into the content scan, so a request that demands a tool call can no longer come back as prose. A real fix for callers setting tool_choice=required against a gemma4 template. #28832 makes the mamba time-step projection input contiguous before the matmul (ggml_cont on the non-norm branch) — a correctness fix in src/models/. The rest is backend and UI: CUDA sparse FA for qwen4 (#28770, with the matching src/models/qwen4exp.cpp tweak), Metal F16 FWHT (#29094), Hexagon I32 GET_ROWS (#29116), and a WebUI mobile-breakpoint/overflow fix (#29108, auto-followed). No file under tools/server/ in this chunk, and no patch target. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 8 ++++---- README.md | 2 +- llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 4 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 8620066a1..0df5670b6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11055** +Current llama.cpp pinned version: **b11062** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11055 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11062 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11055`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11062`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11055`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11062`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index d644f5928..e6c809805 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11055](https://img.shields.io/badge/llama.cpp-%23b11055-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11055) +[![llama.cpp b11062](https://img.shields.io/badge/llama.cpp-%23b11062-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11062) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 6ce9bd0ca..9486e622f 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11055 + GIT_TAG b11062 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 951f01aee..7936ed6b7 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11055"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11062"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11055-"} — call + * plus the resolved upstream commit, e.g. {@code "b11062-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11055"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11062"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11055"; + public static final String LLAMA_CPP_VERSION = "b11062"; // Constants holder — not instantiable. private LlamaCppVersion() {} From 13ddd5947fe6ceced5889e9e938d8f60634763aa Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 10:14:00 +0000 Subject: [PATCH 10/12] docs: record the b11018-b11062 range MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two rows per the runbook's step 4. The first records what moved and, unusually, the chunking itself: nine steps, three of which break the 100 KiB rule and all three irreducibly — b11021, b11023 and b11051 do not exist as tags, so each of those steps is one upstream commit with no smaller step available. Worth having written down, because "over the threshold" has so far always meant "should have been split" and here three times it does not. The second is the patch/verification row. It carries the first patch refresh in several ranges (0001 and 0006, context-only, both caused by upstream #29125), the evidence that 0007's route-table invariant survived it, and the reason the three tools/server contract greps have no input despite tools/server being touched: server-schema.cpp, server-task.cpp and server-context.cpp are byte-identical across the whole range by blob hash. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- docs/history/llama-cpp-breaking-changes.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 670c4afd9..4771da1d0 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -738,3 +738,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. | | b11012–b11018 | 6 commits, **28 KiB**, 9 files — comfortably under the chunking threshold, so a single step. **Eight of the nine files are `ggml/src` backend internals and the ninth is `CODEOWNERS`.** SYCL: **#28953** fixes a B70 allocation failure above 19.3 GB, **#28929** fuses the SiLU epilogue into the `ssm_conv` kernel (new `ssm_conv.{cpp,hpp}` + `fusion.cpp`). Vulkan: **#25483** skips unneeded MoE work in the `mul_mm` coopmat1 path, **#28996** fixes `buffer_reference` alignment in `im2col.comp` / `im2col_3d.comp`. OpenCL: **#28984** clears various warnings. Docs: **#29003** removes a code owner for `test-llama-archs`. | **No project source change, and zero files on the review surface** — nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical server-contract greps have no input. Every change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. | | b11012–b11018 | patches + upstream verification | **Nine patches, none touched and none droppable.** No patch-target file appears anywhere in the range, so `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.cpp`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are byte-unchanged. **All six standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11018:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `src/llama-model.cpp:1518` — unmoved from b11012), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11018:tools/server/`). Verified from a fresh configure: stamp at head `c9a5eeeb3` with nine SHA-256 lines, `verify-patches-applied.sh` green, extraction unchanged at **138 CLI / 57 request / 15 trainer** names, Release build clean with zero errors and zero warnings, `ctest` **551/551**, `nm -D` **40** `Java_*` exports and **0** mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`, full `mvn test` **1763/0**, SpotBugs **0**, spotless clean. **Context worth recording: the previous range's PR run (#958, the b11012 PR) was the first full-matrix execution since b10948** — 66 jobs, **58 success / 2 failure / 6 skipped**, the two failures being the `Verify GPG signing key` pair that `publish.yml` documents as an expected red on a `pull_request` event (the `maven-central` environment withholds secrets there). That run is what first exercised the trainer-model wiring, `verify-test-counts.sh` and both aarch64 fat-jar smoke jobs added earlier in the same session; all passed. | +| b11018–b11062 | 36 commits, **1 347 KiB**, and the chunking is the first thing worth recording: the range was walked in **nine** steps — `b11018→b11020` (10 KiB / 2 commits), `→b11022` (579 / 2), `→b11024` (224 / 2), `→b11042` (87 / 18), `→b11045` (98 / 3), `→b11050` (57 / 5), `→b11052` (115 / 2), `→b11055` (97 / 3) and `→b11062` (84 / 7). **Three steps break the 100 KiB rule and all three are irreducible**: b11021, b11023 and b11051 do not exist as tags, so each of those steps is a *single* upstream commit with no smaller step available — **#28732** (Vulkan: split `ggml-vulkan.cpp` into buffers/debug translation units plus three shared headers, ~5.4k lines moved, `ggml-vulkan/CMakeLists.txt` gains exactly the five new files), **#29009** (OpenVINO update to 2026.4, entirely inside `ggml/src/ggml-openvino/**`), and **#28948** (Metal MoE + SSM_CONV fusion, new `argsort.metal`). **The review surface is 43 files, all additive or implementation-only.** `include/llama.h` gains two things and loses nothing: `LLAMA_VOCAB_TYPE_TEST = 7` (a tail append — no existing enumerator renumbers, and this project reads `vocab_type` as a raw int in `ModelMeta.getVocabType()` and emits it `static_cast`-ed in `jllama.cpp`, so no Java-side constant can go stale) and `llama_adapter_lora_init_from_file_ptr` (#28993, additive; adapters are loaded by path here, never by `FILE*`). `common/chat.cpp` picks up a Ling 3.0 / Bailing V3 detection arm and `common/parsers/ling3.cpp` (#28682), and `common/parsers/gemma4.cpp` **fixes a real bug on a path this project serves**: with `tool_choice == required` the grammar now terminates at the tool call instead of falling through to the content scan (#29115). `common/json-schema-to-grammar.cpp` fixes a second one — `gbnf_escape_length()` now accepts `\-`, so a JSON-schema `pattern` containing an escaped hyphen no longer produces a grammar the parser rejects (#29127). `src/llama-model.{cpp,h}` gain `load_swa_pattern()` with 20 `src/models/*.cpp` architectures rewritten onto it and `TENSOR_SKIP` honoured in `create_tensor_gate_up_exps()` (#29042, #29014); `tools/mtmd/clip.cpp` returns false instead of proceeding when `ggml_backend_sched_alloc_graph()` fails (#28149 / #26070). **`ggml/include` is byte-identical across the whole range**, so no ggml public API moved at all. | +| b11018–b11062 | patches + upstream verification | **Nine patches still, none dropped — but two needed a refresh, the first in several ranges.** One upstream commit is responsible: **#29125** ("server : improve startup log messages", first tagged b11053) adds an `SRV_INF("initializing ...")` line immediately above `llama_server()`'s argv parse and a two-line `TODO` comment above `common_params_parse()` in `common/arg.h`. `0001` anchors hunks on both spots and `0006` replaces the very line `0001` flips, so both went stale **on context only** — the refresh changes `@@` line numbers, three context lines and the index blob hashes, and not one added or removed line. Replayed in filename order against pristine **b11055 and b11062**: all nine apply clean at both. **`0007`'s standing invariant is intact and provably so** — its `-` side is a verbatim copy of the route table it factors out of `llama_server()`, so a clean apply *is* the proof upstream did not touch that block; #29125's edits sit above it (the CORS warning) and below it (the `warn_names` loop), never inside. **The three mechanical `tools/server/` contract greps have no input** despite `tools/server/` being touched: `server-schema.cpp`, `server-task.cpp` and `server-context.cpp` are byte-identical b11018→b11062, verified by blob hash rather than by reading a diff, so the request-field set, the field bounds and the response-key set cannot have moved. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11062:common/arg.h`, WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0011` (the `common_peg_until_parser` `INVALID` branch still returns `FAIL` unconditionally, ignoring `ctx.is_lenient()`, while the `INCOMPLETE` branch right above it honours it), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero guard), and `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure**: stamp head `3cf03257f` with nine SHA-256 lines, `verify-patches-applied.sh` green (9 applied, 0010 cast present), extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean (0 errors, 0 warnings), `ctest` 551/551, `nm -D` 40 `Java_*` exports and 0 mangled, `NativeLibraryLoadSmokeTest` 4/4 with 0 skipped after a `mvn clean`, `mvn test` 1763/0 (269 model-gated skips in a HF-blocked sandbox), `verify-test-counts.sh` 1763 across 119 classes, SpotBugs 0, spotless clean. **One watch item, recorded not acted on**: upstream moved its own OpenVINO SDK pin 2026.3.1→2026.4 in #29009 while this project's two OpenVINO classifier jobs still install the 2026.2.1 archive. The gap predates this range and those jobs built; nothing in the diff uses an API newer than 2026.2.1 exposes. Per the classifier policy these vendor steps are first-pass and fail loud, so the pin moves when a job reds — bumping it speculatively cannot be validated on a GPU-less runner. | From c04451cd02f4944900287d9b40652aef10e9462b Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 10:45:57 +0000 Subject: [PATCH 11/12] docs: correct the stale C++ test counts in CLAUDE.md Two numbers in the "C++ unit tests" section had drifted from the suite. The per-file table said test_jni_helpers.cpp carries 63 tests (it carries 70) and the footer said 544 in total (551). Both are now what `ctest` reports on the b11062 pin, and every other row in the table was checked against a per-file count rather than assumed. The jni_helpers row's prose needed a second correction that the numbers alone would have hidden: it said "The last 7 pin jni_guard_impl", which was true when those tests were the tail of the file but stopped being true once the four ReleaseJllamaContext / JllamaContextGuard tests were appended after them. Positional wording goes stale silently, so it now says "Seven of them". Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- CLAUDE.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 0df5670b6..1bd0dd218 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1591,14 +1591,14 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" | `src/test/cpp/test_server.cpp` | 206 | Upstream result types: `server_slot_stats` (the `timings` JSON payload; replaced `result_timings` in b10408), `task_params::to_json()` (incl. `dry_sequence_breakers`, `preserved_tokens`, `timings_per_token`), `completion_token_output`, `server_task_result_cmpl_partial` (non-oaicompat + `to_json_oaicompat` + logprobs + `to_json_oaicompat_chat` + `to_json_anthropic` + dispatcher), `server_task_result_cmpl_final` (non-oaicompat + `to_json_oaicompat` + `to_json_oaicompat_chat` + `to_json_oaicompat_chat_stream` + `to_json_anthropic` + `to_json_anthropic_stream` + tool_calls + dispatcher), `server_task_result_embd`, `server_task_result_rerank`, `server_task_result_metrics` (`to_metrics()` = the `/metrics` Prometheus exposition text; its `to_json()` has been unused since b10519 and returns `json{}` = JSON null), `server_task_result_slots` (`to_json()` = the `/slots` array, fed by the b10519 `SERVER_TASK_TYPE_SLOT_GET` task), `server_task_result_slot_save_load`, `server_task_result_slot_erase`, `server_task_result_apply_lora`, `server_task_result_get_lora`, `server_task_result_error`, `format_error_response`, `server_task::need_sampling()`, `server_task::n_tokens()`, `server_schema::eval_llama_cmpl_schema()` (parsing pipeline + grammar routing + error paths + per-request `dry_*` and `sse_ping_interval` field round-trips incl. hard-limit + server-default inheritance), `response_fields` projection | | `src/test/cpp/test_json_helpers.cpp` | 63 | All functions in `json_helpers.hpp`: `get_result_error_message`, `results_to_json`, `rerank_results_to_json` (incl. missing/out-of-range `index` rejection), `parse_encoding_format`, `extract_embedding_prompt`, `is_infill_request`, `parse_slot_prompt_similarity`, `parse_positive_int_config`, `wrap_stream_chunk`, `server_metrics_to_json` | | `src/test/cpp/test_log_helpers.cpp` | 13 | All functions in `log_helpers.hpp`: `log_level_name`, `format_log_as_json` | -| `src/test/cpp/test_jni_helpers.cpp` | 63 | All functions in `jni_helpers.hpp` using a zero-filled `JNINativeInterface_` mock (incl. the `utf8_to_jstring_impl` byte-array string path: emoji byte-preservation, truncated-UTF-8 replace-not-throw). The last 7 pin `jni_guard_impl` — the JNI exception boundary every `Java_*` entry point runs inside — including the `catch (...)` arm that is the only backstop for a non-`std::exception` type, and its two refusals (never `ThrowNew` over a pending Java exception, never with a null class). | +| `src/test/cpp/test_jni_helpers.cpp` | 70 | All functions in `jni_helpers.hpp` using a zero-filled `JNINativeInterface_` mock (incl. the `utf8_to_jstring_impl` byte-array string path: emoji byte-preservation, truncated-UTF-8 replace-not-throw). Seven of them pin `jni_guard_impl` — the JNI exception boundary every `Java_*` entry point runs inside — including the `catch (...)` arm that is the only backstop for a non-`std::exception` type, and its two refusals (never `ThrowNew` over a pending Java exception, never with a null class). | | `src/test/cpp/test_tts_wav.cpp` | 2 | The in-memory WAV writer `pcm_to_wav16_bytes` in `tts_wav.hpp` (WAV header/payload + little-endian clamping) — our own code, not upstream. The Qwen3-TTS pipeline it pairs with (`mtmd_helper::gen_audio`) is entirely upstream-owned (no project-side DSP to unit-test here). The load path is additionally covered by `test_tts_params.cpp` (3 tests over `tts_params.hpp`'s `build_tts_params`, plus 2 pinning the upstream `-1` default it depends on), which pins the CPU-thread resolution whose absence used to crash the JVM on every platform — see the `TODO.md` entry for the mechanism. End-to-end coverage is `TtsIntegrationTest`, which is model-gated. | | `src/test/cpp/test_tts_params.cpp` | 13 | The **three** builders every hand-assembled `common_params` goes through: `build_tts_params` (`tts_params.hpp`), `build_train_params` (`train_params.hpp`) and the shared `jllama::resolve_cpu_params` (`cpu_params.hpp`). Each builder is guarded separately on purpose — testing the resolver alone does **not** cover its call sites, because `train_engine.cpp` is compiled into `jllama` only, never into `jllama_test`, and `LlamaTrainerIntegrationTest` is gated on `net.ladenthin.llama.train.model`, which no CI job sets. Without these the JVM-abort bug could regress in the trainer on every platform, unseen. | | `src/test/cpp/test_model_split.cpp` | 7 | The two `load_tensors()` split helpers that `patches/0012` extracts out of llama.cpp's `src/llama-model.cpp` — `llama_model_splits_normalize` (proportional split, single device, and the zero-sum case that used to produce NaN, **and the cancelling `--tensor-split` case** — `-ts 1,-1` reaches the identical line on any backend with no GPU memory pressure at all) and `llama_model_splits_select_device` (every layer maps to a real device index; malformed split points throw a message that names the function, the layer, the index and the split values instead of libc++'s bare `"vector"`). **This is the runnable guard for `0012`**: the patch also ships an upstream `tests/test-model-split.cpp`, but a FetchContent subproject builds with `LLAMA_BUILD_TESTS=OFF`, so that one is applied-but-never-compiled here. This file is the only place the two functions are linked in CI, on every platform — so a bump that drops the patch fails the `C++ Tests` build outright rather than resurfacing as one red macOS Java job. It is the one test file that includes an **internal** upstream header (`llama-model.h`, via the `${llama.cpp_SOURCE_DIR}/src` include dir added for it), which is deliberate: a signature drift should fail loudly at compile time. | | `src/test/cpp/test_model_flags.cpp` | 4 | **The contract between the Java CLI-flag registries and llama.cpp's server argument parser.** CMake reads `ModelFlag.java` + `ModelOption.java` (`cmake/extract-java-wire-names.cmake` → a generated header of `{name, contract}` pairs), and this file asserts every `SERVER_PARSER` name is in `common_params_parser_init(params, LLAMA_EXAMPLE_SERVER).options`. It exists because **no Java test can catch this class**: `ModelFlagTest`/`ModelParametersExtendedTest` pin the *string mapping* (`hasKey("--mlock")`), never that llama.cpp still accepts the string, so they stay green forever while the flag is dead — and `common_params_parse` treats an unregistered option as a hard error, so the affected builder method makes the model **unloadable**, not merely ineffective. **A grep over `arg.cpp` is not a substitute**: `--grp-attn-n`/`-w` are present there at every pinned tag but `set_examples()`-scoped to `LLAMA_EXAMPLE_COMPLETION`/`PASSKEY`, so the server parser rejects them exactly like a deleted flag — only the real option table sees that. `--vocab-only` is the one exemption, and it declares itself `CliContract.PROJECT_PSEUDO` on its own constant rather than appearing in a list inside this file; the test asserts such a name is **still unknown** to the parser (an exemption upstream later registers would be hiding a real check) and that the exempt set is non-empty. | | `src/test/cpp/test_wire_contracts.cpp` | 6 | **The same contract for the two quieter surfaces.** `RequestField` against `server_schema::make_llama_cmpl_schema(...)` (5 tests) and `TrainingField` against `jllama_train::config_keys()` (1 test). Both receivers *silently ignore* an unknown key — the schema skips it, `train_engine.cpp` reads with `j.value(key, default)` and falls back — so a dead field produces no error anywhere and every string-mapping test keeps passing. `OAI_LAYER`-declared keys (consumed by `oaicompat_*_params_parse` before the schema) are exempt from the schema check, and are checked **both** ways: still unknown to the schema (the inverted check), and read by at least one upstream reader-shaped site (the configure-time sweep — this is what caught `chat_template`, a key a public builder wrote and nothing read). See [`docs/history/parameter-wire-surface.md`](docs/history/parameter-wire-surface.md). | -**Current total: 544 tests (all passing).** +**Current total: 551 tests (all passing).** #### Upstream source location (in CMake build tree) From fc5d9548c8c9ba4c9421d9e198ac7f1ed0260c17 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 20 Sep 2026 10:47:54 +0000 Subject: [PATCH 12/12] ci: track upstream's OpenVINO SDK pin, 2026.2.1 -> 2026.4 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both OpenVINO classifier jobs installed the 2026.2.1 archive while upstream llama.cpp moved its own OPENVINO_VERSION_MAJOR / OPENVINO_VERSION_FULL to 2026.3.1 and then, in #29009 (inside this bump's b11022->b11024 chunk), to 2026.4. ggml-openvino is developed against whichever pair upstream pins, so the drift is what eventually breaks the compile — and it does not break on the bump that introduces it, it breaks on some later one, in a job whose runner has no Intel GPU to reproduce on. Tracking upstream is the cheaper end of that trade. Both jobs now use major 2026.4 / full 2026.4.0.22959.99c81491cc3, from the same URL template upstream's linux-setup-openvino and windows-setup-openvino actions use (verified against those two action.yml files at the pinned tag, not guessed from the old string). Each job gains a keep-in-sync note naming upstream's two variables as the source of truth, because the coupling is otherwise invisible: nothing in this repo points at release.yml, and the previous drift happened by simply not looking. Verification limit, stated rather than glossed: the two archive URLs could NOT be reached from the bump sandbox — storage.openvinotoolkit.org is blocked by the network policy (the proxy answers 403 to CONNECT), so no HEAD check was possible. What stands behind them is that upstream's own release jobs download exactly these two URLs at b11062. Per the classifier policy the step is fail-loud, so a wrong URL reds the job rather than shipping a backend-less jar. The b11018-b11062 history row is updated accordingly — it recorded this as a watch item deliberately not acted on, which is no longer what happened. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1 --- .github/workflows/publish.yml | 21 +++++++++++++++------ docs/history/llama-cpp-breaking-changes.md | 2 +- 2 files changed, 16 insertions(+), 7 deletions(-) diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 6edb00294..2ce6f8df7 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -2162,15 +2162,22 @@ jobs: with: distribution: 'temurin' java-version: ${{ env.JAVA_VERSION }} - - name: Install OpenCL dev + Intel OpenVINO 2026.2.1 (archive) + - name: Install OpenCL dev + Intel OpenVINO 2026.4 (archive) run: | # Intel's OpenVINO APT repo only publishes up to ~2025 (the /openvino/2026 path 404s), and # 2025.x has the older ov::Allocator API that breaks ggml-openvino's template compile. So use - # the ARCHIVE for 2026.2.1 — exactly what upstream llama.cpp's linux-setup-openvino action does. + # the ARCHIVE — exactly what upstream llama.cpp's linux-setup-openvino action does, from the + # same URL template. + # + # KEEP IN SYNC WITH UPSTREAM. The version tracks llama.cpp's own OPENVINO_VERSION_MAJOR / + # OPENVINO_VERSION_FULL (.github/workflows/release.yml at the pinned GIT_TAG); ggml-openvino + # is developed against that pair, so lagging it is what eventually breaks the compile. Both + # OpenVINO jobs here (Linux + Windows) use the same two values — bump them together: + # major = 2026.4 full = 2026.4.0.22959.99c81491cc3 # OpenCL headers (incl. the C++ CL/cl2.hpp via opencl-clhpp-headers) come from Ubuntu's own repos. sudo apt-get update sudo apt-get install -y ocl-icd-opencl-dev opencl-headers opencl-clhpp-headers intel-opencl-icd - url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.2.1/linux/openvino_toolkit_ubuntu24_2026.2.1.21919.ede283a88e3_x86_64.tgz" + url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/linux/openvino_toolkit_ubuntu24_2026.4.0.22959.99c81491cc3_x86_64.tgz" sudo mkdir -p /opt/intel/openvino curl -fSL "$url" | sudo tar -xz --strip-components=1 -C /opt/intel/openvino echo "OpenVINO_DIR=/opt/intel/openvino/runtime/cmake" >> "$GITHUB_ENV" @@ -2202,14 +2209,16 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: x64 - - name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.2.1 + - name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.4 shell: pwsh # vcpkg's opencl port ships the full C++ headers incl. CL/cl2.hpp that OpenVINO's # ocl_wrapper.hpp needs (the Khronos OpenCL-Headers dropped cl2.hpp) — same as upstream - # llama.cpp's windows-openvino job. OpenVINO 2026.2.1 matches ggml-openvino's target API. + # llama.cpp's windows-openvino job. OpenVINO 2026.4 matches ggml-openvino's target API. + # Keep the version in sync with the Linux OpenVINO job above (and with upstream's + # OPENVINO_VERSION_MAJOR / OPENVINO_VERSION_FULL) — see the note there. run: | C:\vcpkg\vcpkg install opencl:x64-windows - $url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.2.1/windows/openvino_toolkit_windows_2026.2.1.21919.ede283a88e3_x86_64.zip" + $url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/windows/openvino_toolkit_windows_2026.4.0.22959.99c81491cc3_x86_64.zip" Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\openvino.zip" Expand-Archive -Path "$env:RUNNER_TEMP\openvino.zip" -DestinationPath "C:\openvino" -Force # The archive extracts into a nested versioned folder; point OpenVINO_DIR at its runtime/cmake. diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 4771da1d0..7a218e25b 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -739,4 +739,4 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11012–b11018 | 6 commits, **28 KiB**, 9 files — comfortably under the chunking threshold, so a single step. **Eight of the nine files are `ggml/src` backend internals and the ninth is `CODEOWNERS`.** SYCL: **#28953** fixes a B70 allocation failure above 19.3 GB, **#28929** fuses the SiLU epilogue into the `ssm_conv` kernel (new `ssm_conv.{cpp,hpp}` + `fusion.cpp`). Vulkan: **#25483** skips unneeded MoE work in the `mul_mm` coopmat1 path, **#28996** fixes `buffer_reference` alignment in `im2col.comp` / `im2col_3d.comp`. OpenCL: **#28984** clears various warnings. Docs: **#29003** removes a code owner for `test-llama-archs`. | **No project source change, and zero files on the review surface** — nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical server-contract greps have no input. Every change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. | | b11012–b11018 | patches + upstream verification | **Nine patches, none touched and none droppable.** No patch-target file appears anywhere in the range, so `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.cpp`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are byte-unchanged. **All six standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11018:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `src/llama-model.cpp:1518` — unmoved from b11012), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11018:tools/server/`). Verified from a fresh configure: stamp at head `c9a5eeeb3` with nine SHA-256 lines, `verify-patches-applied.sh` green, extraction unchanged at **138 CLI / 57 request / 15 trainer** names, Release build clean with zero errors and zero warnings, `ctest` **551/551**, `nm -D` **40** `Java_*` exports and **0** mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`, full `mvn test` **1763/0**, SpotBugs **0**, spotless clean. **Context worth recording: the previous range's PR run (#958, the b11012 PR) was the first full-matrix execution since b10948** — 66 jobs, **58 success / 2 failure / 6 skipped**, the two failures being the `Verify GPG signing key` pair that `publish.yml` documents as an expected red on a `pull_request` event (the `maven-central` environment withholds secrets there). That run is what first exercised the trainer-model wiring, `verify-test-counts.sh` and both aarch64 fat-jar smoke jobs added earlier in the same session; all passed. | | b11018–b11062 | 36 commits, **1 347 KiB**, and the chunking is the first thing worth recording: the range was walked in **nine** steps — `b11018→b11020` (10 KiB / 2 commits), `→b11022` (579 / 2), `→b11024` (224 / 2), `→b11042` (87 / 18), `→b11045` (98 / 3), `→b11050` (57 / 5), `→b11052` (115 / 2), `→b11055` (97 / 3) and `→b11062` (84 / 7). **Three steps break the 100 KiB rule and all three are irreducible**: b11021, b11023 and b11051 do not exist as tags, so each of those steps is a *single* upstream commit with no smaller step available — **#28732** (Vulkan: split `ggml-vulkan.cpp` into buffers/debug translation units plus three shared headers, ~5.4k lines moved, `ggml-vulkan/CMakeLists.txt` gains exactly the five new files), **#29009** (OpenVINO update to 2026.4, entirely inside `ggml/src/ggml-openvino/**`), and **#28948** (Metal MoE + SSM_CONV fusion, new `argsort.metal`). **The review surface is 43 files, all additive or implementation-only.** `include/llama.h` gains two things and loses nothing: `LLAMA_VOCAB_TYPE_TEST = 7` (a tail append — no existing enumerator renumbers, and this project reads `vocab_type` as a raw int in `ModelMeta.getVocabType()` and emits it `static_cast`-ed in `jllama.cpp`, so no Java-side constant can go stale) and `llama_adapter_lora_init_from_file_ptr` (#28993, additive; adapters are loaded by path here, never by `FILE*`). `common/chat.cpp` picks up a Ling 3.0 / Bailing V3 detection arm and `common/parsers/ling3.cpp` (#28682), and `common/parsers/gemma4.cpp` **fixes a real bug on a path this project serves**: with `tool_choice == required` the grammar now terminates at the tool call instead of falling through to the content scan (#29115). `common/json-schema-to-grammar.cpp` fixes a second one — `gbnf_escape_length()` now accepts `\-`, so a JSON-schema `pattern` containing an escaped hyphen no longer produces a grammar the parser rejects (#29127). `src/llama-model.{cpp,h}` gain `load_swa_pattern()` with 20 `src/models/*.cpp` architectures rewritten onto it and `TENSOR_SKIP` honoured in `create_tensor_gate_up_exps()` (#29042, #29014); `tools/mtmd/clip.cpp` returns false instead of proceeding when `ggml_backend_sched_alloc_graph()` fails (#28149 / #26070). **`ggml/include` is byte-identical across the whole range**, so no ggml public API moved at all. | -| b11018–b11062 | patches + upstream verification | **Nine patches still, none dropped — but two needed a refresh, the first in several ranges.** One upstream commit is responsible: **#29125** ("server : improve startup log messages", first tagged b11053) adds an `SRV_INF("initializing ...")` line immediately above `llama_server()`'s argv parse and a two-line `TODO` comment above `common_params_parse()` in `common/arg.h`. `0001` anchors hunks on both spots and `0006` replaces the very line `0001` flips, so both went stale **on context only** — the refresh changes `@@` line numbers, three context lines and the index blob hashes, and not one added or removed line. Replayed in filename order against pristine **b11055 and b11062**: all nine apply clean at both. **`0007`'s standing invariant is intact and provably so** — its `-` side is a verbatim copy of the route table it factors out of `llama_server()`, so a clean apply *is* the proof upstream did not touch that block; #29125's edits sit above it (the CORS warning) and below it (the `warn_names` loop), never inside. **The three mechanical `tools/server/` contract greps have no input** despite `tools/server/` being touched: `server-schema.cpp`, `server-task.cpp` and `server-context.cpp` are byte-identical b11018→b11062, verified by blob hash rather than by reading a diff, so the request-field set, the field bounds and the response-key set cannot have moved. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11062:common/arg.h`, WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0011` (the `common_peg_until_parser` `INVALID` branch still returns `FAIL` unconditionally, ignoring `ctx.is_lenient()`, while the `INCOMPLETE` branch right above it honours it), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero guard), and `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure**: stamp head `3cf03257f` with nine SHA-256 lines, `verify-patches-applied.sh` green (9 applied, 0010 cast present), extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean (0 errors, 0 warnings), `ctest` 551/551, `nm -D` 40 `Java_*` exports and 0 mangled, `NativeLibraryLoadSmokeTest` 4/4 with 0 skipped after a `mvn clean`, `mvn test` 1763/0 (269 model-gated skips in a HF-blocked sandbox), `verify-test-counts.sh` 1763 across 119 classes, SpotBugs 0, spotless clean. **One watch item, recorded not acted on**: upstream moved its own OpenVINO SDK pin 2026.3.1→2026.4 in #29009 while this project's two OpenVINO classifier jobs still install the 2026.2.1 archive. The gap predates this range and those jobs built; nothing in the diff uses an API newer than 2026.2.1 exposes. Per the classifier policy these vendor steps are first-pass and fail loud, so the pin moves when a job reds — bumping it speculatively cannot be validated on a GPU-less runner. | +| b11018–b11062 | patches + upstream verification | **Nine patches still, none dropped — but two needed a refresh, the first in several ranges.** One upstream commit is responsible: **#29125** ("server : improve startup log messages", first tagged b11053) adds an `SRV_INF("initializing ...")` line immediately above `llama_server()`'s argv parse and a two-line `TODO` comment above `common_params_parse()` in `common/arg.h`. `0001` anchors hunks on both spots and `0006` replaces the very line `0001` flips, so both went stale **on context only** — the refresh changes `@@` line numbers, three context lines and the index blob hashes, and not one added or removed line. Replayed in filename order against pristine **b11055 and b11062**: all nine apply clean at both. **`0007`'s standing invariant is intact and provably so** — its `-` side is a verbatim copy of the route table it factors out of `llama_server()`, so a clean apply *is* the proof upstream did not touch that block; #29125's edits sit above it (the CORS warning) and below it (the `warn_names` loop), never inside. **The three mechanical `tools/server/` contract greps have no input** despite `tools/server/` being touched: `server-schema.cpp`, `server-task.cpp` and `server-context.cpp` are byte-identical b11018→b11062, verified by blob hash rather than by reading a diff, so the request-field set, the field bounds and the response-key set cannot have moved. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11062:common/arg.h`, WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0011` (the `common_peg_until_parser` `INVALID` branch still returns `FAIL` unconditionally, ignoring `ctx.is_lenient()`, while the `INCOMPLETE` branch right above it honours it), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero guard), and `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure**: stamp head `3cf03257f` with nine SHA-256 lines, `verify-patches-applied.sh` green (9 applied, 0010 cast present), extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean (0 errors, 0 warnings), `ctest` 551/551, `nm -D` 40 `Java_*` exports and 0 mangled, `NativeLibraryLoadSmokeTest` 4/4 with 0 skipped after a `mvn clean`, `mvn test` 1763/0 (269 model-gated skips in a HF-blocked sandbox), `verify-test-counts.sh` 1763 across 119 classes, SpotBugs 0, spotless clean. **The OpenVINO SDK pin moved with it**: #29009 takes upstream's own `OPENVINO_VERSION_MAJOR`/`OPENVINO_VERSION_FULL` to 2026.4, and this project's two OpenVINO classifier jobs — which had drifted two releases behind at 2026.2.1 — now install `2026.4` / `2026.4.0.22959.99c81491cc3` from the same URL template upstream's `{linux,windows}-setup-openvino` actions use. ggml-openvino is developed against whatever pair upstream pins, so tracking it is the cheaper end of the trade: a lagging pin does not fail on the bump that introduces the drift, it fails on some later one, in a job whose runner has no Intel GPU to reproduce on. **Not verifiable from the bump sandbox** — `storage.openvinotoolkit.org` is blocked by the network policy, so neither archive URL could be HEAD-checked here; the evidence they resolve is that upstream's own release jobs download exactly these two URLs at b11062. Per the classifier policy the step is fail-loud, so a wrong URL reds the job rather than shipping a backend-less jar. Both jobs now carry a keep-in-sync note naming upstream's two variables as the source of truth, so the next bump has somewhere to look instead of rediscovering the coupling. |