diff --git a/CHANGELOG.md b/CHANGELOG.md index 2474f4b28..837535706 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -70,6 +70,14 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by without a native library or model. ### Changed +- **llama.cpp `b11214` → `b11222`.** Eight upstream commits. Two touch argument parsing: + **#29518** makes `string_split` throw `invalid value: "…"` for a list element that does not parse + (only the benchmark options `-npp`/`-ntg`/`-npl` use a numeric split, so nothing a server or `jllama` + argument reaches changes); and **#29537** registers `--rpc` + in every build and rejects it at parse time with `RPC not supported in this build` (this project builds + with `GGML_RPC=OFF`), where before the option did not exist at all. The rest is CUDA/SYCL/OpenCL kernel + work, a Jinja `dict` builtin and conversion scripts. `patches/0001` and `0006` were refreshed for a + moved log line in `tools/server/server.cpp`; their content is unchanged. - **llama.cpp `b11211` → `b11214`.** Three upstream commits, version-only from this project's side: a HIP flash-attention kernel choice for CDNA (#28907), a Vulkan argsort fix for Adreno (#29469), and **#29516**, which makes `common_sampler_init` *throw* `failed to parse grammar: llguidance is not diff --git a/CLAUDE.md b/CLAUDE.md index 3c86b53fe..e934ca4f2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11214** +Current llama.cpp pinned version: **b11222** ## Upgrading CUDA Version @@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11214 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11222 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11214`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11222`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1683,7 +1683,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11214`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11222`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 9d39dbcbf..4065c43bd 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11214](https://img.shields.io/badge/llama.cpp-%23b11214-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11214) +[![llama.cpp b11222](https://img.shields.io/badge/llama.cpp-%23b11222-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11222) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 45af228ed..e6934165e 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -758,3 +758,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b11209–b11211 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir; the applier's stamp lists all eight against `d7fb90e8e`). No patch-target file is in the range. | | b11211–b11214 | Three commits, 5 files, 29 lines: **#28907** (HIP: fattn-mma on CDNA for `dkq > 256` at large batch, `ggml-cuda/fattn*.cu[h]` + `scripts/hip/gcn-cdna-vgpr-check.py`), **#29469** (Vulkan argsort kernel selection for Adreno) and **#29516** (`common/sampling.cpp`: an llguidance grammar without `LLAMA_LLGUIDANCE` now throws `std::runtime_error` instead of `GGML_ABORT` — a process abort turned into a catchable error, which for a JNI host is the difference between a failed request and a dead JVM). Version-only from this project's side: no API change, no workflow/ROCm/CUDA-component change upstream in the range. | | b11211–b11214 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir). No patch-target file is in the range. | +| b11214–b11222 | Eight commits, 11 files, 276 lines. **#29518** (`common/common.h`: `string_split` throws `std::invalid_argument` on a token that does not parse, instead of pushing an uninitialised/zero value) and **#29537** (`common/arg.cpp`: `--rpc` is registered unconditionally and throws `RPC not supported in this build` when `llama_supports_rpc()` is false; `tools/server/server.cpp`: the `initializing ...` log moved below `common_params_parse`). Neither needs a project source change: `jllama` calls neither `string_split` nor `--rpc`, `--rpc` is not in `ModelOption`, and the only numeric `string_split` callers are the benchmark options `-npp`/`-ntg`/`-npl` (every server-reachable split is `string_split`, which cannot fail). The rest: #26289 (CUDA fp16 tile FA configs), #29243 (SYCL FWHT > 512), #29503 (OpenCL bin kernel loading), #29477 (Jinja `dict` builtin), #29528 (PLaMo-3 YaRN conversion), #29529 (upstream CI). No `release.yml` change, so the CUDA/ROCm/OpenVINO pins stay. | +| b11214–b11222 | patches + upstream verification | **`0001` and `0006` refreshed, six apply unchanged.** #29537 moved `SRV_INF("initializing ...")` from above to below the `common_params_parse` call in `llama_server()`, which is the context of both patches' parse-call hunk; the hunks now anchor on the preceding `server_stream_session_manager_start()` lines instead. Content unchanged, all eight apply in order on pristine b11222. **Drop-check `0001`: still required** — `common_params_parse_main` has 0 occurrences in `b11222:common/arg.h`, and `common/arg.cpp` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 56c67d23e..9c6e87470 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11214 + GIT_TAG b11222 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/patches/0001-win32-arg-parse-embed-guard.patch b/llama/patches/0001-win32-arg-parse-embed-guard.patch index 56fd186c8..900929d38 100644 --- a/llama/patches/0001-win32-arg-parse-embed-guard.patch +++ b/llama/patches/0001-win32-arg-parse-embed-guard.patch @@ -486,9 +486,9 @@ diff --git a/tools/server/server.cpp b/tools/server/server.cpp index 1167c0aea..28f18c1bb 100644 --- a/tools/server/server.cpp +++ b/tools/server/server.cpp -@@ -104,7 +104,7 @@ int llama_server(int argc, char ** argv) { - - SRV_INF("%s", "initializing ...\n"); +@@ -102,7 +102,7 @@ int llama_server(int argc, char ** argv) { + // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free + server_stream_session_manager_start(); - if (!common_params_parse(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { diff --git a/llama/patches/0006-server-embed-native-server-jni.patch b/llama/patches/0006-server-embed-native-server-jni.patch index 23d73c014..f90f275ba 100644 --- a/llama/patches/0006-server-embed-native-server-jni.patch +++ b/llama/patches/0006-server-embed-native-server-jni.patch @@ -31,9 +31,9 @@ index 28f18c1bb..8aeed4b99 100644 static inline void signal_handler(int signal) { if (is_terminating.test_and_set()) { // in case it hangs, we can force terminate the server by hitting Ctrl+C twice -@@ -104,7 +126,13 @@ int llama_server(int argc, char ** argv) { - - SRV_INF("%s", "initializing ...\n"); +@@ -102,7 +124,13 @@ int llama_server(int argc, char ** argv) { + // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free + server_stream_session_manager_start(); - if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + // [jllama] embedded (JNI) callers forward a clean UTF-8 argv, so honor it exactly via diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 6ace75120..e520cea1e 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11214"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11222"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11214-"} — call + * plus the resolved upstream commit, e.g. {@code "b11222-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11214"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11222"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11214"; + public static final String LLAMA_CPP_VERSION = "b11222"; // Constants holder — not instantiable. private LlamaCppVersion() {}