diff --git a/CLAUDE.md b/CLAUDE.md index 93acbd5f2..d19da9722 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b10988** +Current llama.cpp pinned version: **b11012** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b10988 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11012 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10988`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11012`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10988`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11012`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index fd8f8410d..d6714f996 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b10988](https://img.shields.io/badge/llama.cpp-%23b10988-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10988) +[![llama.cpp b11012](https://img.shields.io/badge/llama.cpp-%23b11012-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11012) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 00b76ce1e..0cb929df1 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -734,3 +734,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10969–b10976 | patches + upstream verification | End of the two-chunk walk, and **the patch set is unchanged at nine** — nothing dropped, nothing refreshed. Verified against the pristine target rather than inferred: `git apply -p1` of all nine, in filename order, into a clean `b10976` worktree, every one clean. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10976:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10976:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10976:src/llama-model.cpp:1493`, no zero-sum guard), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`/`llama_server_attach` and `LLAMA_SERVER_WORKER_CMD` all absent upstream). The whole b10948→b10976 range leaves **every** patch target except `src/llama-model.cpp` and `tests/CMakeLists.txt` byte-identical, and both of those move only at a distance from the patched regions. Verified end-to-end for real: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `987498f4592a76897863cf53711dce38380c082b` (= `b10976`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors; `ctest` **537/537**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** — the check that cross-validates the bumped `LLAMA_CPP_VERSION` constant against the linked `build-info`, needing the `clean` because the constant is inlined into the already-compiled test class. | | b10976–b10988 | 12 commits, **323 KiB** — and the size is the story, so the chunking decision is recorded rather than taken quietly. The range was walked in **three** commits: `b10976→b10980` (54 KiB), `b10980→b10981` (191 KiB) and `b10981→b10988` (78 KiB). The middle step is the one that breaks the runbook's 100 KiB rule, and it is **irreducible**: it is a *single* upstream commit, **#28638** ("OpenVINO: optimize stateful decode and GPU MoE inference"), and there is no intermediate `b` tag inside one commit, so no smaller step exists to take. 35 of its 37 files are `ggml/src/ggml-openvino/**`; the other two are `ci/run.sh` and `docs/backend/OPENVINO.md`. Across the **whole** range, `ggml/src` accounts for 293 KiB of the 323 — the remainder is `src/models` 4.6 KiB, `tests/` 5.5 KiB (never compiled here), `.github` 14 KiB of upstream's own CI, `docs`/`ci`/`CONTRIBUTING.md` 5.3 KiB, and `ggml/include` **0.3 KiB**. Backend work by vendor: **#28599** Metal FA kernels for HSK=96/HSV=64 (MiniCPM3), **#28881** a generic OpenCL `ssm_scan`, **#27637** OpenCL MoE expert-matmul selection by batch size, **#28105** Vulkan sparse flash attention (5 shaders + `ggml-vulkan.cpp`), **#26308** CUDA row-contiguous `SUM_ROWS`, **#28789** RPC weight-only hash caching. The one non-backend change is **#28934**, pure code motion: `build_arch_graph()` moves below the `graph()` template specializations in `src/models/{dflash,eagle3,t5}.cpp`. 67 files, 2828 insertions, 892 deletions. | **No project source change, and zero files on the review surface — literally zero.** Nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/` moved at all, so every row of the priority API-compatibility table is vacuously satisfied and the three mechanical `tools/server` contract greps have **no input to compare**: the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have moved. The three `src/models/*.cpp` files are internal upstream TUs and #28934 moves no signature. **One "safe to skip" header does move and was checked rather than waved past**: `ggml/include/ggml-rpc.h` bumps `RPC_PROTO_MAJOR_VERSION` 6 → 7. It is inert here — `GGML_RPC` appears nowhere in `llama/CMakeLists.txt`, `publish.yml`, `build.sh` or `build.bat`, so `ggml-rpc` is never built and the wire protocol it versions is never spoken. The practical risk of the 191 KiB OpenVINO step is correspondingly narrow: `llama/CMakeLists.txt` routes `GGML_OPENVINO` to the `resources_linux_openvino` / `resources_windows_openvino` **classifier** trees only — it is not in the default JAR — and both `openvino-*` jobs are build-only on GPU-less runners, so "must still compile" is the whole of it, and CI checks that directly. | | b10976–b10988 | patches + upstream verification | **The patch set is unchanged at nine — nothing dropped, nothing refreshed, and nothing even had to be re-examined.** Not one patch-target file is touched anywhere in the range: `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.{cpp,h}`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are all byte-unchanged across b10976→b10988, verified by diffing those paths explicitly rather than inferred from the aggregate. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10988:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10988:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10988:src/llama-model.cpp:1493`, no zero-sum guard). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `9f31776c3773cf03f98535c19b7e6d394af374b4` (= `b10988`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (the `clean` is load-bearing — the `LLAMA_CPP_VERSION` constant is inlined into the already-compiled test class, so without it the cross-check against the linked `build-info` compares the old value); full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **One defect was found by this bump rather than by the range**: `.github/verify-patches-applied.sh` counted the stamp's patch lines as "total lines minus one", which silently went stale when the applier's content oracle added a second metadata line (`tree `) in the next commit of the session that introduced the script. The guard therefore failed on every correct tree and would have redded `C++ Tests` on the next pipeline run; it is fixed in this branch by counting patch lines by their own shape instead of by subtraction. | +| b10988–b11012 | 24 commits, **206 KiB**, 65 files — over the runbook's threshold, walked in **four** chunks: `b10988→b11002` (75 KiB / 14), `b11002→b11005` (33 KiB / 3), `b11005→b11011` (88 KiB / 6) and `b11011→b11012` (13 KiB / 1). The last was kept separate on purpose: folding it into the third makes one **100.6 KiB** step, over by a hair, and "barely over" is the rationalisation the threshold exists to prevent — the cost of honouring it is one extra commit, not an extra build. **Two files on the priority review list**, both `common/` implementation behind unchanged signatures. **#28849** (`common/fit.cpp`) is the one with teeth: `common_params_fit_impl` now sizes `n_ctx_max` by `n_seq_max` rather than `n_streams`, so an auto-sized context (`-c 0`) with `--parallel > 1` and a **unified** KV cache gets `n_ctx_train * n_seq_max` instead of `n_ctx_train`; `kv_unified` keeps `n_streams == 1`, so the non-unified path is unchanged. **#28869** (`common/parsers/qwen3-coder.cpp`) puts `"\n"` ahead of `""` in `thinking_end_tags` so the newline lands inside the forced message. **#27625** adds a whole architecture, `HrmTextForCausalLM` (DFM Mimir 1B), across `src/llama-arch.{cpp,h}`, `llama-hparams.h`, `llama-context.cpp`, `llama-model-saver.cpp`, `src/models/hrm-text.cpp` and — the part that matters here — `src/llama-model.{cpp,h}`. **#28549** splits `llama_context::gf_res_prev` into a two-element array so batches with and without outputs get distinct CUDA-graph cache keys. The remaining ~33 `ggml/src` files are CUDA/HIP im2col and MoE heuristics, HIP AllReduce, Vulkan MUL_MAT_ID / argsort / qwen4exp hc ops, Metal `mul_mm_id` NaN, hexagon copy/DMA and K-quants, spacemit int16 transpose, and an RPC graph-cache invalidation. | **No project source change.** Neither `common/` change is a compile or link consequence — both are implementation-only in TUs upstream compiles into `llama-common`, and nothing here calls `common_params_fit_impl` directly (it is reached through `common_init_from_params` when `n_ctx == 0`), so #28849 surfaces as "an auto-sized parallel server may now ask for more context", upstream's intended fix rather than a regression to absorb. **Zero `tools/server/` files moved**, so the three mechanical server-contract greps have no input and the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have changed. **Two `ggml/include` headers do move and both were opened rather than waved past**: `ggml.h` is purely additive (a new `ggml_dsv4_hc_pre_gated` plus one comment line — no existing signature moves, and this project calls no `ggml_dsv4_*`), and `ggml-sycl.h`'s `ggml_backend_sycl_split_buffer_type` gains a leading `int main_device` — a real signature break, but of a backend-internal symbol no project code calls; the three `sycl-*` classifier jobs compile upstream's own self-consistent tree. `src/llama-context.h` is an **internal** header this project does not include (the only internal one it does is `src/llama-model.h`, from `test_model_split.cpp`), and #28549's change there is a private member. | +| b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 181ae78ea..a99a20b18 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b10988 + GIT_TAG b11012 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index ea38ab654..c4038e1e6 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10988"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11012"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b10988-"} — call + * plus the resolved upstream commit, e.g. {@code "b11012-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10988"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11012"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b10988"; + public static final String LLAMA_CPP_VERSION = "b11012"; // Constants holder — not instantiable. private LlamaCppVersion() {}