Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI.

Current llama.cpp pinned version: **b11012**
Current llama.cpp pinned version: **b11018**

## Upgrading CUDA Version

Expand Down Expand Up @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi
ships no UI):
```bash
# needs node/npm + network for the asset build; the embed step is plain cmake -P
git clone --depth 1 --branch b11012 https://github.com/ggml-org/llama.cpp /tmp/lc
git clone --depth 1 --branch b11018 https://github.com/ggml-org/llama.cpp /tmp/lc
( cd /tmp/lc/tools/ui && npm ci && npm run build )
mkdir -p webui-generated /tmp/ui-gen
cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \
Expand Down Expand Up @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend:
- `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored
as the repo secret **`DEPOT_TOKEN`**.

Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11012`), the
Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11018`), the
~280 upstream object files are byte-identical every run, so a warm cache recompiles only the
*changed* files. Depot's cache is **shared across all branches** (unlike GitHub's
per-branch `actions/cache`), so every branch builds incrementally; a `b<nnnn>` version bump
Expand Down Expand Up @@ -1602,7 +1602,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson"

#### Upstream source location (in CMake build tree)

llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11012`.
llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11018`.

**GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely
by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
**Build:**
![Java 8+](https://img.shields.io/badge/Java-8%2B-informational)
![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey)
[![llama.cpp b11012](https://img.shields.io/badge/llama.cpp-%23b11012-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11012)
[![llama.cpp b11018](https://img.shields.io/badge/llama.cpp-%23b11018-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11018)
[![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/)
![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162)
[![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev)
Expand Down
2 changes: 2 additions & 0 deletions docs/history/llama-cpp-breaking-changes.md
Original file line number Diff line number Diff line change
Expand Up @@ -736,3 +736,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r
| b10976–b10988 | patches + upstream verification | **The patch set is unchanged at nine — nothing dropped, nothing refreshed, and nothing even had to be re-examined.** Not one patch-target file is touched anywhere in the range: `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.{cpp,h}`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are all byte-unchanged across b10976→b10988, verified by diffing those paths explicitly rather than inferred from the aggregate. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b10988:common/arg.h`; the WIN32 `argv = utf8.ptrs.data()` override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `b10988:tools/server/server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b10988:src/llama-model.cpp:1493`, no zero-sum guard). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `9f31776c3773cf03f98535c19b7e6d394af374b4` (= `b10988`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; full `cmake --build --config Release` clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped** (the `clean` is load-bearing — the `LLAMA_CPP_VERSION` constant is inlined into the already-compiled test class, so without it the cross-check against the linked `build-info` compares the old value); full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **One defect was found by this bump rather than by the range**: `.github/verify-patches-applied.sh` counted the stamp's patch lines as "total lines minus one", which silently went stale when the applier's content oracle added a second metadata line (`tree <fingerprint>`) in the next commit of the session that introduced the script. The guard therefore failed on every correct tree and would have redded `C++ Tests` on the next pipeline run; it is fixed in this branch by counting patch lines by their own shape instead of by subtraction. |
| b10988–b11012 | 24 commits, **206 KiB**, 65 files — over the runbook's threshold, walked in **four** chunks: `b10988→b11002` (75 KiB / 14), `b11002→b11005` (33 KiB / 3), `b11005→b11011` (88 KiB / 6) and `b11011→b11012` (13 KiB / 1). The last was kept separate on purpose: folding it into the third makes one **100.6 KiB** step, over by a hair, and "barely over" is the rationalisation the threshold exists to prevent — the cost of honouring it is one extra commit, not an extra build. **Two files on the priority review list**, both `common/` implementation behind unchanged signatures. **#28849** (`common/fit.cpp`) is the one with teeth: `common_params_fit_impl` now sizes `n_ctx_max` by `n_seq_max` rather than `n_streams`, so an auto-sized context (`-c 0`) with `--parallel > 1` and a **unified** KV cache gets `n_ctx_train * n_seq_max` instead of `n_ctx_train`; `kv_unified` keeps `n_streams == 1`, so the non-unified path is unchanged. **#28869** (`common/parsers/qwen3-coder.cpp`) puts `"\n</think>"` ahead of `"</think>"` in `thinking_end_tags` so the newline lands inside the forced message. **#27625** adds a whole architecture, `HrmTextForCausalLM` (DFM Mimir 1B), across `src/llama-arch.{cpp,h}`, `llama-hparams.h`, `llama-context.cpp`, `llama-model-saver.cpp`, `src/models/hrm-text.cpp` and — the part that matters here — `src/llama-model.{cpp,h}`. **#28549** splits `llama_context::gf_res_prev` into a two-element array so batches with and without outputs get distinct CUDA-graph cache keys. The remaining ~33 `ggml/src` files are CUDA/HIP im2col and MoE heuristics, HIP AllReduce, Vulkan MUL_MAT_ID / argsort / qwen4exp hc ops, Metal `mul_mm_id` NaN, hexagon copy/DMA and K-quants, spacemit int16 transpose, and an RPC graph-cache invalidation. | **No project source change.** Neither `common/` change is a compile or link consequence — both are implementation-only in TUs upstream compiles into `llama-common`, and nothing here calls `common_params_fit_impl` directly (it is reached through `common_init_from_params` when `n_ctx == 0`), so #28849 surfaces as "an auto-sized parallel server may now ask for more context", upstream's intended fix rather than a regression to absorb. **Zero `tools/server/` files moved**, so the three mechanical server-contract greps have no input and the request-field set, its `set_hard_limits` bounds and the emitted response keys cannot have changed. **Two `ggml/include` headers do move and both were opened rather than waved past**: `ggml.h` is purely additive (a new `ggml_dsv4_hc_pre_gated` plus one comment line — no existing signature moves, and this project calls no `ggml_dsv4_*`), and `ggml-sycl.h`'s `ggml_backend_sycl_split_buffer_type` gains a leading `int main_device` — a real signature break, but of a backend-internal symbol no project code calls; the three `sycl-*` classifier jobs compile upstream's own self-consistent tree. `src/llama-context.h` is an **internal** header this project does not include (the only internal one it does is `src/llama-model.h`, from `test_model_split.cpp`), and #28549's change there is a private member. |
| b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. |
| b11012–b11018 | 6 commits, **28 KiB**, 9 files — comfortably under the chunking threshold, so a single step. **Eight of the nine files are `ggml/src` backend internals and the ninth is `CODEOWNERS`.** SYCL: **#28953** fixes a B70 allocation failure above 19.3 GB, **#28929** fuses the SiLU epilogue into the `ssm_conv` kernel (new `ssm_conv.{cpp,hpp}` + `fusion.cpp`). Vulkan: **#25483** skips unneeded MoE work in the `mul_mm` coopmat1 path, **#28996** fixes `buffer_reference` alignment in `im2col.comp` / `im2col_3d.comp`. OpenCL: **#28984** clears various warnings. Docs: **#29003** removes a code owner for `test-llama-archs`. | **No project source change, and zero files on the review surface** — nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical server-contract greps have no input. Every change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. |
| b11012–b11018 | patches + upstream verification | **Nine patches, none touched and none droppable.** No patch-target file appears anywhere in the range, so `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.cpp`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are byte-unchanged. **All six standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11018:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `src/llama-model.cpp:1518` — unmoved from b11012), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11018:tools/server/`). Verified from a fresh configure: stamp at head `c9a5eeeb3` with nine SHA-256 lines, `verify-patches-applied.sh` green, extraction unchanged at **138 CLI / 57 request / 15 trainer** names, Release build clean with zero errors and zero warnings, `ctest` **551/551**, `nm -D` **40** `Java_*` exports and **0** mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`, full `mvn test` **1763/0**, SpotBugs **0**, spotless clean. **Context worth recording: the previous range's PR run (#958, the b11012 PR) was the first full-matrix execution since b10948** — 66 jobs, **58 success / 2 failure / 6 skipped**, the two failures being the `Verify GPG signing key` pair that `publish.yml` documents as an expected red on a `pull_request` event (the `maven-central` environment withholds secrets there). That run is what first exercised the trainer-model wiring, `verify-test-counts.sh` and both aarch64 fat-jar smoke jobs added earlier in the same session; all passed. |
2 changes: 1 addition & 1 deletion llama/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE)
FetchContent_Declare(
llama.cpp
GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git
GIT_TAG b11012
GIT_TAG b11018
PATCH_COMMAND ${CMAKE_COMMAND}
-DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches
-DLLAMA_SRC=<SOURCE_DIR>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,28 +10,28 @@
* library was compiled against, exposed as a compile-time constant so callers can render a badge or
* emit a startup log line without loading the native library.
*
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11012"}) that mirrors the
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11018"}) that mirrors the
* {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is
* absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a
* lightweight version badge in Android or other UIs.</p>
*
* <p>For the <em>authoritative</em> value that is baked into the native binary — the build number
* plus the resolved upstream commit, e.g. {@code "b11012-<commit>"} — call
* plus the resolved upstream commit, e.g. {@code "b11018-<commit>"} — call
* {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own
* {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires
* the native library to be loaded).</p>
*/
public final class LlamaCppVersion {

/**
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11012"}.
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11018"}.
*
* <p>Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the
* "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the
* compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the
* value actually linked into the native binary.</p>
*/
public static final String LLAMA_CPP_VERSION = "b11012";
public static final String LLAMA_CPP_VERSION = "b11018";

// Constants holder — not instantiable.
private LlamaCppVersion() {}
Expand Down
Loading