Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,14 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
without a native library or model.

### Changed
- **llama.cpp `b11214` → `b11222`.** Eight upstream commits. Two touch argument parsing:
**#29518** makes `string_split<T>` throw `invalid value: "…"` for a list element that does not parse
(only the benchmark options `-npp`/`-ntg`/`-npl` use a numeric split, so nothing a server or `jllama`
argument reaches changes); and **#29537** registers `--rpc`
in every build and rejects it at parse time with `RPC not supported in this build` (this project builds
with `GGML_RPC=OFF`), where before the option did not exist at all. The rest is CUDA/SYCL/OpenCL kernel
work, a Jinja `dict` builtin and conversion scripts. `patches/0001` and `0006` were refreshed for a
moved log line in `tools/server/server.cpp`; their content is unchanged.
- **llama.cpp `b11211` → `b11214`.** Three upstream commits, version-only from this project's side:
a HIP flash-attention kernel choice for CDNA (#28907), a Vulkan argsort fix for Adreno (#29469), and
**#29516**, which makes `common_sampler_init` *throw* `failed to parse grammar: llguidance is not
Expand Down
8 changes: 4 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI.

Current llama.cpp pinned version: **b11214**
Current llama.cpp pinned version: **b11222**

## Upgrading CUDA Version

Expand Down Expand Up @@ -538,7 +538,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi
ships no UI):
```bash
# needs node/npm + network for the asset build; the embed step is plain cmake -P
git clone --depth 1 --branch b11214 https://github.com/ggml-org/llama.cpp /tmp/lc
git clone --depth 1 --branch b11222 https://github.com/ggml-org/llama.cpp /tmp/lc
( cd /tmp/lc/tools/ui && npm ci && npm run build )
mkdir -p webui-generated /tmp/ui-gen
cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \
Expand Down Expand Up @@ -578,7 +578,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend:
- `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored
as the repo secret **`DEPOT_TOKEN`**.

Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11214`), the
Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11222`), the
~280 upstream object files are byte-identical every run, so a warm cache recompiles only the
*changed* files. Depot's cache is **shared across all branches** (unlike GitHub's
per-branch `actions/cache`), so every branch builds incrementally; a `b<nnnn>` version bump
Expand Down Expand Up @@ -1683,7 +1683,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson"

#### Upstream source location (in CMake build tree)

llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11214`.
llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11222`.

**GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely
by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
**Build:**
![Java 8+](https://img.shields.io/badge/Java-8%2B-informational)
![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey)
[![llama.cpp b11214](https://img.shields.io/badge/llama.cpp-%23b11214-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11214)
[![llama.cpp b11222](https://img.shields.io/badge/llama.cpp-%23b11222-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11222)
[![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/)
![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162)
[![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev)
Expand Down
2 changes: 2 additions & 0 deletions docs/history/llama-cpp-breaking-changes.md
Original file line number Diff line number Diff line change
Expand Up @@ -758,3 +758,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r
| b11209–b11211 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir; the applier's stamp lists all eight against `d7fb90e8e`). No patch-target file is in the range. |
| b11211–b11214 | Three commits, 5 files, 29 lines: **#28907** (HIP: fattn-mma on CDNA for `dkq > 256` at large batch, `ggml-cuda/fattn*.cu[h]` + `scripts/hip/gcn-cdna-vgpr-check.py`), **#29469** (Vulkan argsort kernel selection for Adreno) and **#29516** (`common/sampling.cpp`: an llguidance grammar without `LLAMA_LLGUIDANCE` now throws `std::runtime_error` instead of `GGML_ABORT` — a process abort turned into a catchable error, which for a JNI host is the difference between a failed request and a dead JVM). Version-only from this project's side: no API change, no workflow/ROCm/CUDA-component change upstream in the range. |
| b11211–b11214 | patches + upstream verification | **All eight patches apply unchanged** (fresh configure into an empty build dir). No patch-target file is in the range. |
| b11214–b11222 | Eight commits, 11 files, 276 lines. **#29518** (`common/common.h`: `string_split<T>` throws `std::invalid_argument` on a token that does not parse, instead of pushing an uninitialised/zero value) and **#29537** (`common/arg.cpp`: `--rpc` is registered unconditionally and throws `RPC not supported in this build` when `llama_supports_rpc()` is false; `tools/server/server.cpp`: the `initializing ...` log moved below `common_params_parse`). Neither needs a project source change: `jllama` calls neither `string_split` nor `--rpc`, `--rpc` is not in `ModelOption`, and the only numeric `string_split<int>` callers are the benchmark options `-npp`/`-ntg`/`-npl` (every server-reachable split is `string_split<std::string>`, which cannot fail). The rest: #26289 (CUDA fp16 tile FA configs), #29243 (SYCL FWHT > 512), #29503 (OpenCL bin kernel loading), #29477 (Jinja `dict` builtin), #29528 (PLaMo-3 YaRN conversion), #29529 (upstream CI). No `release.yml` change, so the CUDA/ROCm/OpenVINO pins stay. |
| b11214–b11222 | patches + upstream verification | **`0001` and `0006` refreshed, six apply unchanged.** #29537 moved `SRV_INF("initializing ...")` from above to below the `common_params_parse` call in `llama_server()`, which is the context of both patches' parse-call hunk; the hunks now anchor on the preceding `server_stream_session_manager_start()` lines instead. Content unchanged, all eight apply in order on pristine b11222. **Drop-check `0001`: still required** — `common_params_parse_main` has 0 occurrences in `b11222:common/arg.h`, and `common/arg.cpp` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override. |
2 changes: 1 addition & 1 deletion llama/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE)
FetchContent_Declare(
llama.cpp
GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git
GIT_TAG b11214
GIT_TAG b11222
PATCH_COMMAND ${CMAKE_COMMAND}
-DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches
-DLLAMA_SRC=<SOURCE_DIR>
Expand Down
6 changes: 3 additions & 3 deletions llama/patches/0001-win32-arg-parse-embed-guard.patch
Original file line number Diff line number Diff line change
Expand Up @@ -486,9 +486,9 @@ diff --git a/tools/server/server.cpp b/tools/server/server.cpp
index 1167c0aea..28f18c1bb 100644
--- a/tools/server/server.cpp
+++ b/tools/server/server.cpp
@@ -104,7 +104,7 @@ int llama_server(int argc, char ** argv) {

SRV_INF("%s", "initializing ...\n");
@@ -102,7 +102,7 @@ int llama_server(int argc, char ** argv) {
// touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free
server_stream_session_manager_start();

- if (!common_params_parse(argc, argv, params, LLAMA_EXAMPLE_SERVER)) {
+ if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) {
Expand Down
6 changes: 3 additions & 3 deletions llama/patches/0006-server-embed-native-server-jni.patch
Original file line number Diff line number Diff line change
Expand Up @@ -31,9 +31,9 @@ index 28f18c1bb..8aeed4b99 100644
static inline void signal_handler(int signal) {
if (is_terminating.test_and_set()) {
// in case it hangs, we can force terminate the server by hitting Ctrl+C twice
@@ -104,7 +126,13 @@ int llama_server(int argc, char ** argv) {

SRV_INF("%s", "initializing ...\n");
@@ -102,7 +124,13 @@ int llama_server(int argc, char ** argv) {
// touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free
server_stream_session_manager_start();

- if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) {
+ // [jllama] embedded (JNI) callers forward a clean UTF-8 argv, so honor it exactly via
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,28 +10,28 @@
* library was compiled against, exposed as a compile-time constant so callers can render a badge or
* emit a startup log line without loading the native library.
*
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11214"}) that mirrors the
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11222"}) that mirrors the
* {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is
* absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a
* lightweight version badge in Android or other UIs.</p>
*
* <p>For the <em>authoritative</em> value that is baked into the native binary — the build number
* plus the resolved upstream commit, e.g. {@code "b11214-<commit>"} — call
* plus the resolved upstream commit, e.g. {@code "b11222-<commit>"} — call
* {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own
* {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires
* the native library to be loaded).</p>
*/
public final class LlamaCppVersion {

/**
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11214"}.
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11222"}.
*
* <p>Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the
* "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the
* compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the
* value actually linked into the native binary.</p>
*/
public static final String LLAMA_CPP_VERSION = "b11214";
public static final String LLAMA_CPP_VERSION = "b11222";

// Constants holder — not instantiable.
private LlamaCppVersion() {}
Expand Down
Loading