From b8a1d36e5471fe0f929e87a321963cedadb10fa0 Mon Sep 17 00:00:00 2001 From: Claude Date: Mon, 28 Sep 2026 21:09:46 +0000 Subject: [PATCH] Upgrade llama.cpp from b11236 to b11237 One upstream commit (#29603): the OpenVINO backend marks unaligned batch-stride views unsupported. Backend-internal; no API change, no patch touches the file, no workflow change upstream. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8 --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 2 +- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 12 insertions(+), 10 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 8e79fb97..994e691b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11236** +Current llama.cpp pinned version: **b11237** ## Upgrading CUDA Version @@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11236 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11237 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11236`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11237`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1795,7 +1795,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11236`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11237`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 7dcbfe41..caaeca52 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11236](https://img.shields.io/badge/llama.cpp-%23b11236-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11236) +[![llama.cpp b11237](https://img.shields.io/badge/llama.cpp-%23b11237-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11237) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 65a1b4f1..b573eea4 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -762,3 +762,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11214–b11222 | patches + upstream verification | **`0001` and `0006` refreshed, six apply unchanged.** #29537 moved `SRV_INF("initializing ...")` from above to below the `common_params_parse` call in `llama_server()`, which is the context of both patches' parse-call hunk; the hunks now anchor on the preceding `server_stream_session_manager_start()` lines instead. Content unchanged, all eight apply in order on pristine b11222. **Drop-check `0001`: still required** — `common_params_parse_main` has 0 occurrences in `b11222:common/arg.h`, and `common/arg.cpp` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override. | | b11222–b11236 | Fourteen commits, 51 files, ~2.2k lines. **#29385** migrates speculative decoding, `mtmd` and the server to `batch_ext`: `common/common.h` gains the `common_batch` wrapper and `common_batch_get_one`/`common_batch_from_llama_batch`, drops `string_from(ctx, llama_batch)` and `common_batch_ext_get_one`; `common/speculative.h` adds a `common_batch` overload of `common_speculative_process`; `tools/mtmd/mtmd-helper.h` changes `mtmd_helper_post_decode_callback` to take a `mtmd_helper_embd_batch`. **None of these symbols is used by project code** (`jllama.cpp`, `tts_engine.cpp`, the test sources; `gen_audio` is unchanged). `include/llama.h` only adds `llama_get_causal_attn`. **#28876** lets RANK pooling split batches for causal-LLM rerankers (Qwen3 / Qwen3-VL) inside `server-context.cpp`, no interface change. No `tools/server/*.h`, `server-schema.cpp` or `server-task.cpp` change, so the request/response contract checks have nothing to compare. **#28362** enables a Windows ARM64 build with MSVC `cl.exe` in `ggml-cpu/CMakeLists.txt`; the arm64 job keeps `clang-cl`. No upstream `.github` change, so the ROCm and Windows-CUDA component pins stay. Version-only from this project's side. | | b11222–b11236 | patches + upstream verification | **`0001` refreshed, eight apply unchanged.** #29426 rewrote `tests/test-recurrent-state-rollback.cpp` so its `main()` strips `--models DIR` into its own `filtered_argv` before `common_params_parse(fargc, filtered_argv.data(), …)` — the same shape `test-save-load-state.cpp` took at b10679 — so by `0001`'s own rule that call site keeps `common_params_parse` and the hunk was **dropped** (37 → 36 file diffs). Verified by applying all nine patches in order to a clean b11236 worktree. **Drop checks:** `0002` still needed (`server-context.cpp` still assigns `params_base.load_progress_callback` unconditionally); `common/arg.cpp`, `src/llama-model.cpp`, `common/log.{h,cpp}` and `ggml/src/ggml-rpc/` are untouched in the range, so `0001`, `0012`, `0014` and `0015` still carry live fixes. | +| b11236–b11237 | One commit, 1 file, 22 lines. **#29603** (`ggml/src/ggml-openvino/ggml-openvino.cpp`): the OpenVINO backend reports views with an unaligned batch stride as unsupported, so the scheduler falls back instead of computing them wrong. Backend-internal, no API change, no workflow change upstream. Version-only from this project's side. | +| b11236–b11237 | patches + upstream verification | **All nine patches apply unchanged.** No patch-target file is in the range. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 756c17cc..d8ac14c1 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -188,7 +188,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11236 + GIT_TAG b11237 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 5d6605f8..a462661d 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11236"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11237"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11236-"} — call + * plus the resolved upstream commit, e.g. {@code "b11237-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11236"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11237"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11236"; + public static final String LLAMA_CPP_VERSION = "b11237"; // Constants holder — not instantiable. private LlamaCppVersion() {}