diff --git a/CLAUDE.md b/CLAUDE.md index d19da9722..1e1b6f612 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11012** +Current llama.cpp pinned version: **b11018** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11012 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11018 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11012`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11018`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11012`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11018`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index d6714f996..5054455d9 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11012](https://img.shields.io/badge/llama.cpp-%23b11012-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11012) +[![llama.cpp b11018](https://img.shields.io/badge/llama.cpp-%23b11018-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11018) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 0cb929df1..670c4afd9 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -736,3 +736,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10976–b10988 | patches + upstream verification | **The patch set is unchanged at nine — nothing dropped, nothing refreshed, and nothing even had to be re-examined.** Not one patch-target file is touched anywhere in the range: `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.{cpp,h}`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are all byte-unchanged across b10976→b10988, verified by diffing those paths explicitly rather than inferred from the aggregate. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10988:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10988:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10988:src/llama-model.cpp:1493`, no zero-sum guard). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `9f31776c3773cf03f98535c19b7e6d394af374b4` (= `b10988`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (the `clean` is load-bearing — the `LLAMA_CPP_VERSION` constant is inlined into the already-compiled test class, so without it the cross-check against the linked `build-info` compares the old value); full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **One defect was found by this bump rather than by the range**: `.github/verify-patches-applied.sh` counted the stamp's patch lines as "total lines minus one", which silently went stale when the applier's content oracle added a second metadata line (`tree `) in the next commit of the session that introduced the script. The guard therefore failed on every correct tree and would have redded `C++ Tests` on the next pipeline run; it is fixed in this branch by counting patch lines by their own shape instead of by subtraction. | | b10988–b11012 | 24 commits, **206 KiB**, 65 files — over the runbook's threshold, walked in **four** chunks: `b10988→b11002` (75 KiB / 14), `b11002→b11005` (33 KiB / 3), `b11005→b11011` (88 KiB / 6) and `b11011→b11012` (13 KiB / 1). The last was kept separate on purpose: folding it into the third makes one **100.6 KiB** step, over by a hair, and "barely over" is the rationalisation the threshold exists to prevent — the cost of honouring it is one extra commit, not an extra build. **Two files on the priority review list**, both `common/` implementation behind unchanged signatures. **#28849** (`common/fit.cpp`) is the one with teeth: `common_params_fit_impl` now sizes `n_ctx_max` by `n_seq_max` rather than `n_streams`, so an auto-sized context (`-c 0`) with `--parallel > 1` and a **unified** KV cache gets `n_ctx_train * n_seq_max` instead of `n_ctx_train`; `kv_unified` keeps `n_streams == 1`, so the non-unified path is unchanged. **#28869** (`common/parsers/qwen3-coder.cpp`) puts `"\n"` ahead of `""` in `thinking_end_tags` so the newline lands inside the forced message. **#27625** adds a whole architecture, `HrmTextForCausalLM` (DFM Mimir 1B), across `src/llama-arch.{cpp,h}`, `llama-hparams.h`, `llama-context.cpp`, `llama-model-saver.cpp`, `src/models/hrm-text.cpp` and — the part that matters here — `src/llama-model.{cpp,h}`. **#28549** splits `llama_context::gf_res_prev` into a two-element array so batches with and without outputs get distinct CUDA-graph cache keys. The remaining ~33 `ggml/src` files are CUDA/HIP im2col and MoE heuristics, HIP AllReduce, Vulkan MUL_MAT_ID / argsort / qwen4exp hc ops, Metal `mul_mm_id` NaN, hexagon copy/DMA and K-quants, spacemit int16 transpose, and an RPC graph-cache invalidation. | **No project source change.** Neither `common/` change is a compile or link consequence — both are implementation-only in TUs upstream compiles into `llama-common`, and nothing here calls `common_params_fit_impl` directly (it is reached through `common_init_from_params` when `n_ctx == 0`), so #28849 surfaces as "an auto-sized parallel server may now ask for more context", upstream's intended fix rather than a regression to absorb. **Zero `tools/server/` files moved**, so the three mechanical server-contract greps have no input and the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have changed. **Two `ggml/include` headers do move and both were opened rather than waved past**: `ggml.h` is purely additive (a new `ggml_dsv4_hc_pre_gated` plus one comment line — no existing signature moves, and this project calls no `ggml_dsv4_*`), and `ggml-sycl.h`'s `ggml_backend_sycl_split_buffer_type` gains a leading `int main_device` — a real signature break, but of a backend-internal symbol no project code calls; the three `sycl-*` classifier jobs compile upstream's own self-consistent tree. `src/llama-context.h` is an **internal** header this project does not include (the only internal one it does is `src/llama-model.h`, from `test_model_split.cpp`), and #28549's change there is a private member. | | b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. | +| b11012–b11018 | 6 commits, **28 KiB**, 9 files — comfortably under the chunking threshold, so a single step. **Eight of the nine files are `ggml/src` backend internals and the ninth is `CODEOWNERS`.** SYCL: **#28953** fixes a B70 allocation failure above 19.3 GB, **#28929** fuses the SiLU epilogue into the `ssm_conv` kernel (new `ssm_conv.{cpp,hpp}` + `fusion.cpp`). Vulkan: **#25483** skips unneeded MoE work in the `mul_mm` coopmat1 path, **#28996** fixes `buffer_reference` alignment in `im2col.comp` / `im2col_3d.comp`. OpenCL: **#28984** clears various warnings. Docs: **#29003** removes a code owner for `test-llama-archs`. | **No project source change, and zero files on the review surface** — nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical server-contract greps have no input. Every change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. | +| b11012–b11018 | patches + upstream verification | **Nine patches, none touched and none droppable.** No patch-target file appears anywhere in the range, so `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.cpp`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are byte-unchanged. **All six standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11018:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `src/llama-model.cpp:1518` — unmoved from b11012), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11018:tools/server/`). Verified from a fresh configure: stamp at head `c9a5eeeb3` with nine SHA-256 lines, `verify-patches-applied.sh` green, extraction unchanged at **138 CLI / 57 request / 15 trainer** names, Release build clean with zero errors and zero warnings, `ctest` **551/551**, `nm -D` **40** `Java_*` exports and **0** mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`, full `mvn test` **1763/0**, SpotBugs **0**, spotless clean. **Context worth recording: the previous range's PR run (#958, the b11012 PR) was the first full-matrix execution since b10948** — 66 jobs, **58 success / 2 failure / 6 skipped**, the two failures being the `Verify GPG signing key` pair that `publish.yml` documents as an expected red on a `pull_request` event (the `maven-central` environment withholds secrets there). That run is what first exercised the trainer-model wiring, `verify-test-counts.sh` and both aarch64 fat-jar smoke jobs added earlier in the same session; all passed. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index a99a20b18..028b45081 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11012 + GIT_TAG b11018 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index c4038e1e6..135013783 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11012"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11018"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11012-"} — call + * plus the resolved upstream commit, e.g. {@code "b11018-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11012"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11018"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11012"; + public static final String LLAMA_CPP_VERSION = "b11018"; // Constants holder — not instantiable. private LlamaCppVersion() {}