diff --git a/.github/workflows/publish.yml b/.github/workflows/publish.yml index 6edb00294..2ce6f8df7 100644 --- a/.github/workflows/publish.yml +++ b/.github/workflows/publish.yml @@ -2162,15 +2162,22 @@ jobs: with: distribution: 'temurin' java-version: ${{ env.JAVA_VERSION }} - - name: Install OpenCL dev + Intel OpenVINO 2026.2.1 (archive) + - name: Install OpenCL dev + Intel OpenVINO 2026.4 (archive) run: | # Intel's OpenVINO APT repo only publishes up to ~2025 (the /openvino/2026 path 404s), and # 2025.x has the older ov::Allocator API that breaks ggml-openvino's template compile. So use - # the ARCHIVE for 2026.2.1 — exactly what upstream llama.cpp's linux-setup-openvino action does. + # the ARCHIVE — exactly what upstream llama.cpp's linux-setup-openvino action does, from the + # same URL template. + # + # KEEP IN SYNC WITH UPSTREAM. The version tracks llama.cpp's own OPENVINO_VERSION_MAJOR / + # OPENVINO_VERSION_FULL (.github/workflows/release.yml at the pinned GIT_TAG); ggml-openvino + # is developed against that pair, so lagging it is what eventually breaks the compile. Both + # OpenVINO jobs here (Linux + Windows) use the same two values — bump them together: + # major = 2026.4 full = 2026.4.0.22959.99c81491cc3 # OpenCL headers (incl. the C++ CL/cl2.hpp via opencl-clhpp-headers) come from Ubuntu's own repos. sudo apt-get update sudo apt-get install -y ocl-icd-opencl-dev opencl-headers opencl-clhpp-headers intel-opencl-icd - url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.2.1/linux/openvino_toolkit_ubuntu24_2026.2.1.21919.ede283a88e3_x86_64.tgz" + url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/linux/openvino_toolkit_ubuntu24_2026.4.0.22959.99c81491cc3_x86_64.tgz" sudo mkdir -p /opt/intel/openvino curl -fSL "$url" | sudo tar -xz --strip-components=1 -C /opt/intel/openvino echo "OpenVINO_DIR=/opt/intel/openvino/runtime/cmake" >> "$GITHUB_ENV" @@ -2202,14 +2209,16 @@ jobs: uses: ilammy/msvc-dev-cmd@v1 with: arch: x64 - - name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.2.1 + - name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.4 shell: pwsh # vcpkg's opencl port ships the full C++ headers incl. CL/cl2.hpp that OpenVINO's # ocl_wrapper.hpp needs (the Khronos OpenCL-Headers dropped cl2.hpp) — same as upstream - # llama.cpp's windows-openvino job. OpenVINO 2026.2.1 matches ggml-openvino's target API. + # llama.cpp's windows-openvino job. OpenVINO 2026.4 matches ggml-openvino's target API. + # Keep the version in sync with the Linux OpenVINO job above (and with upstream's + # OPENVINO_VERSION_MAJOR / OPENVINO_VERSION_FULL) — see the note there. run: | C:\vcpkg\vcpkg install opencl:x64-windows - $url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.2.1/windows/openvino_toolkit_windows_2026.2.1.21919.ede283a88e3_x86_64.zip" + $url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/windows/openvino_toolkit_windows_2026.4.0.22959.99c81491cc3_x86_64.zip" Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\openvino.zip" Expand-Archive -Path "$env:RUNNER_TEMP\openvino.zip" -DestinationPath "C:\openvino" -Force # The archive extracts into a nested versioned folder; point OpenVINO_DIR at its runtime/cmake. diff --git a/CLAUDE.md b/CLAUDE.md index 1e1b6f612..1bd0dd218 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b11018** +Current llama.cpp pinned version: **b11062** ## Upgrading CUDA Version @@ -510,7 +510,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b11018 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b11062 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -550,7 +550,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11018`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11062`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1591,18 +1591,18 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" | `src/test/cpp/test_server.cpp` | 206 | Upstream result types: `server_slot_stats` (the `timings` JSON payload; replaced `result_timings` in b10408), `task_params::to_json()` (incl. `dry_sequence_breakers`, `preserved_tokens`, `timings_per_token`), `completion_token_output`, `server_task_result_cmpl_partial` (non-oaicompat + `to_json_oaicompat` + logprobs + `to_json_oaicompat_chat` + `to_json_anthropic` + dispatcher), `server_task_result_cmpl_final` (non-oaicompat + `to_json_oaicompat` + `to_json_oaicompat_chat` + `to_json_oaicompat_chat_stream` + `to_json_anthropic` + `to_json_anthropic_stream` + tool_calls + dispatcher), `server_task_result_embd`, `server_task_result_rerank`, `server_task_result_metrics` (`to_metrics()` = the `/metrics` Prometheus exposition text; its `to_json()` has been unused since b10519 and returns `json{}` = JSON null), `server_task_result_slots` (`to_json()` = the `/slots` array, fed by the b10519 `SERVER_TASK_TYPE_SLOT_GET` task), `server_task_result_slot_save_load`, `server_task_result_slot_erase`, `server_task_result_apply_lora`, `server_task_result_get_lora`, `server_task_result_error`, `format_error_response`, `server_task::need_sampling()`, `server_task::n_tokens()`, `server_schema::eval_llama_cmpl_schema()` (parsing pipeline + grammar routing + error paths + per-request `dry_*` and `sse_ping_interval` field round-trips incl. hard-limit + server-default inheritance), `response_fields` projection | | `src/test/cpp/test_json_helpers.cpp` | 63 | All functions in `json_helpers.hpp`: `get_result_error_message`, `results_to_json`, `rerank_results_to_json` (incl. missing/out-of-range `index` rejection), `parse_encoding_format`, `extract_embedding_prompt`, `is_infill_request`, `parse_slot_prompt_similarity`, `parse_positive_int_config`, `wrap_stream_chunk`, `server_metrics_to_json` | | `src/test/cpp/test_log_helpers.cpp` | 13 | All functions in `log_helpers.hpp`: `log_level_name`, `format_log_as_json` | -| `src/test/cpp/test_jni_helpers.cpp` | 63 | All functions in `jni_helpers.hpp` using a zero-filled `JNINativeInterface_` mock (incl. the `utf8_to_jstring_impl` byte-array string path: emoji byte-preservation, truncated-UTF-8 replace-not-throw). The last 7 pin `jni_guard_impl` — the JNI exception boundary every `Java_*` entry point runs inside — including the `catch (...)` arm that is the only backstop for a non-`std::exception` type, and its two refusals (never `ThrowNew` over a pending Java exception, never with a null class). | +| `src/test/cpp/test_jni_helpers.cpp` | 70 | All functions in `jni_helpers.hpp` using a zero-filled `JNINativeInterface_` mock (incl. the `utf8_to_jstring_impl` byte-array string path: emoji byte-preservation, truncated-UTF-8 replace-not-throw). Seven of them pin `jni_guard_impl` — the JNI exception boundary every `Java_*` entry point runs inside — including the `catch (...)` arm that is the only backstop for a non-`std::exception` type, and its two refusals (never `ThrowNew` over a pending Java exception, never with a null class). | | `src/test/cpp/test_tts_wav.cpp` | 2 | The in-memory WAV writer `pcm_to_wav16_bytes` in `tts_wav.hpp` (WAV header/payload + little-endian clamping) — our own code, not upstream. The Qwen3-TTS pipeline it pairs with (`mtmd_helper::gen_audio`) is entirely upstream-owned (no project-side DSP to unit-test here). The load path is additionally covered by `test_tts_params.cpp` (3 tests over `tts_params.hpp`'s `build_tts_params`, plus 2 pinning the upstream `-1` default it depends on), which pins the CPU-thread resolution whose absence used to crash the JVM on every platform — see the `TODO.md` entry for the mechanism. End-to-end coverage is `TtsIntegrationTest`, which is model-gated. | | `src/test/cpp/test_tts_params.cpp` | 13 | The **three** builders every hand-assembled `common_params` goes through: `build_tts_params` (`tts_params.hpp`), `build_train_params` (`train_params.hpp`) and the shared `jllama::resolve_cpu_params` (`cpu_params.hpp`). Each builder is guarded separately on purpose — testing the resolver alone does **not** cover its call sites, because `train_engine.cpp` is compiled into `jllama` only, never into `jllama_test`, and `LlamaTrainerIntegrationTest` is gated on `net.ladenthin.llama.train.model`, which no CI job sets. Without these the JVM-abort bug could regress in the trainer on every platform, unseen. | | `src/test/cpp/test_model_split.cpp` | 7 | The two `load_tensors()` split helpers that `patches/0012` extracts out of llama.cpp's `src/llama-model.cpp` — `llama_model_splits_normalize` (proportional split, single device, and the zero-sum case that used to produce NaN, **and the cancelling `--tensor-split` case** — `-ts 1,-1` reaches the identical line on any backend with no GPU memory pressure at all) and `llama_model_splits_select_device` (every layer maps to a real device index; malformed split points throw a message that names the function, the layer, the index and the split values instead of libc++'s bare `"vector"`). **This is the runnable guard for `0012`**: the patch also ships an upstream `tests/test-model-split.cpp`, but a FetchContent subproject builds with `LLAMA_BUILD_TESTS=OFF`, so that one is applied-but-never-compiled here. This file is the only place the two functions are linked in CI, on every platform — so a bump that drops the patch fails the `C++ Tests` build outright rather than resurfacing as one red macOS Java job. It is the one test file that includes an **internal** upstream header (`llama-model.h`, via the `${llama.cpp_SOURCE_DIR}/src` include dir added for it), which is deliberate: a signature drift should fail loudly at compile time. | | `src/test/cpp/test_model_flags.cpp` | 4 | **The contract between the Java CLI-flag registries and llama.cpp's server argument parser.** CMake reads `ModelFlag.java` + `ModelOption.java` (`cmake/extract-java-wire-names.cmake` → a generated header of `{name, contract}` pairs), and this file asserts every `SERVER_PARSER` name is in `common_params_parser_init(params, LLAMA_EXAMPLE_SERVER).options`. It exists because **no Java test can catch this class**: `ModelFlagTest`/`ModelParametersExtendedTest` pin the *string mapping* (`hasKey("--mlock")`), never that llama.cpp still accepts the string, so they stay green forever while the flag is dead — and `common_params_parse` treats an unregistered option as a hard error, so the affected builder method makes the model **unloadable**, not merely ineffective. **A grep over `arg.cpp` is not a substitute**: `--grp-attn-n`/`-w` are present there at every pinned tag but `set_examples()`-scoped to `LLAMA_EXAMPLE_COMPLETION`/`PASSKEY`, so the server parser rejects them exactly like a deleted flag — only the real option table sees that. `--vocab-only` is the one exemption, and it declares itself `CliContract.PROJECT_PSEUDO` on its own constant rather than appearing in a list inside this file; the test asserts such a name is **still unknown** to the parser (an exemption upstream later registers would be hiding a real check) and that the exempt set is non-empty. | | `src/test/cpp/test_wire_contracts.cpp` | 6 | **The same contract for the two quieter surfaces.** `RequestField` against `server_schema::make_llama_cmpl_schema(...)` (5 tests) and `TrainingField` against `jllama_train::config_keys()` (1 test). Both receivers *silently ignore* an unknown key — the schema skips it, `train_engine.cpp` reads with `j.value(key, default)` and falls back — so a dead field produces no error anywhere and every string-mapping test keeps passing. `OAI_LAYER`-declared keys (consumed by `oaicompat_*_params_parse` before the schema) are exempt from the schema check, and are checked **both** ways: still unknown to the schema (the inverted check), and read by at least one upstream reader-shaped site (the configure-time sweep — this is what caught `chat_template`, a key a public builder wrote and nothing read). See [`docs/history/parameter-wire-surface.md`](docs/history/parameter-wire-surface.md). | -**Current total: 544 tests (all passing).** +**Current total: 551 tests (all passing).** #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11018`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11062`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index 5054455d9..e6c809805 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b11018](https://img.shields.io/badge/llama.cpp-%23b11018-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11018) +[![llama.cpp b11062](https://img.shields.io/badge/llama.cpp-%23b11062-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11062) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 670c4afd9..7a218e25b 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -738,3 +738,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10988–b11012 | patches + upstream verification | **The patch set is unchanged at nine, and this is the first range in a while where a patch target actually moved** — so the collision was checked, not assumed. #27625 edits `src/llama-model.cpp` in three places (an `LLM_ARCH_HRM_TEXT` case in `llama_model_mapping`, a `MIRRORED` meta-split branch for that arch's aliased cache slots, and a rope-type case) at roughly lines 316, 477 and 3030, while `patches/0012`'s hunks are the `load_tensors` split arithmetic at **1493–1518**. A thousand lines apart; `src/llama-model.h` likewise gains only an additive `hrm_z_l_init` member. Every other patch target — `common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/*.cpp`, `tests/CMakeLists.txt` — is byte-unchanged across the range. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11012:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `b11012:src/llama-model.cpp:1518`, no zero-sum guard — note it moved 1493 → 1518 under the new arch code, which is exactly why this is re-checked by content rather than by line), `0002` (`params_base.load_progress_callback = load_progress_callback;` still unguarded at `server-context.cpp:1095`), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11012:tools/server/`). Verified end-to-end from a fresh configure: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `35822afe58475e0506cd51e6573903e46d4c67c9` (= `b11012`) with **nine** SHA-256 lines; extraction unchanged at **138 CLI / 57 request / 15 trainer** names; Release build clean with **zero** errors and zero warnings; `ctest` **551/551**; `nm -D` **40** `Java_*` exports, **0** mangled; `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`; full `mvn test` **1763/0**; SpotBugs **0**; spotless and `javadoc:jar` clean. **This is also the first bump to exercise the `verify-patches-applied.sh` fix from the previous range** — the guard that had been failing on every correct tree since its own content-oracle sibling landed now reports "9 applied, tree dirty, patches/0010 cast present" on a real build, as it always should have. | | b11012–b11018 | 6 commits, **28 KiB**, 9 files — comfortably under the chunking threshold, so a single step. **Eight of the nine files are `ggml/src` backend internals and the ninth is `CODEOWNERS`.** SYCL: **#28953** fixes a B70 allocation failure above 19.3 GB, **#28929** fuses the SiLU epilogue into the `ssm_conv` kernel (new `ssm_conv.{cpp,hpp}` + `fusion.cpp`). Vulkan: **#25483** skips unneeded MoE work in the `mul_mm` coopmat1 path, **#28996** fixes `buffer_reference` alignment in `im2col.comp` / `im2col_3d.comp`. OpenCL: **#28984** clears various warnings. Docs: **#29003** removes a code owner for `test-llama-archs`. | **No project source change, and zero files on the review surface** — nothing under `common/`, `include/`, `tools/server/` or `tools/mtmd/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical server-contract greps have no input. Every change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. | | b11012–b11018 | patches + upstream verification | **Nine patches, none touched and none droppable.** No patch-target file appears anywhere in the range, so `common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*.cpp`, `src/llama-model.{cpp,h}` and `tests/CMakeLists.txt` are byte-unchanged. **All six standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11018:common/arg.h`; WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0012` (bare `splits[i] /= split_sum;` still at `src/llama-model.cpp:1518` — unmoved from b11012), and `0003`/`0006`/`0008` (`get_slot_prompt_similarity`, `llama_server_set_embedded`, `LLAMA_SERVER_WORKER_CMD` all absent from `b11018:tools/server/`). Verified from a fresh configure: stamp at head `c9a5eeeb3` with nine SHA-256 lines, `verify-patches-applied.sh` green, extraction unchanged at **138 CLI / 57 request / 15 trainer** names, Release build clean with zero errors and zero warnings, `ctest` **551/551**, `nm -D` **40** `Java_*` exports and **0** mangled, `NativeLibraryLoadSmokeTest` **4/4, 0 skipped** after a `clean`, full `mvn test` **1763/0**, SpotBugs **0**, spotless clean. **Context worth recording: the previous range's PR run (#958, the b11012 PR) was the first full-matrix execution since b10948** — 66 jobs, **58 success / 2 failure / 6 skipped**, the two failures being the `Verify GPG signing key` pair that `publish.yml` documents as an expected red on a `pull_request` event (the `maven-central` environment withholds secrets there). That run is what first exercised the trainer-model wiring, `verify-test-counts.sh` and both aarch64 fat-jar smoke jobs added earlier in the same session; all passed. | +| b11018–b11062 | 36 commits, **1 347 KiB**, and the chunking is the first thing worth recording: the range was walked in **nine** steps — `b11018→b11020` (10 KiB / 2 commits), `→b11022` (579 / 2), `→b11024` (224 / 2), `→b11042` (87 / 18), `→b11045` (98 / 3), `→b11050` (57 / 5), `→b11052` (115 / 2), `→b11055` (97 / 3) and `→b11062` (84 / 7). **Three steps break the 100 KiB rule and all three are irreducible**: b11021, b11023 and b11051 do not exist as tags, so each of those steps is a *single* upstream commit with no smaller step available — **#28732** (Vulkan: split `ggml-vulkan.cpp` into buffers/debug translation units plus three shared headers, ~5.4k lines moved, `ggml-vulkan/CMakeLists.txt` gains exactly the five new files), **#29009** (OpenVINO update to 2026.4, entirely inside `ggml/src/ggml-openvino/**`), and **#28948** (Metal MoE + SSM_CONV fusion, new `argsort.metal`). **The review surface is 43 files, all additive or implementation-only.** `include/llama.h` gains two things and loses nothing: `LLAMA_VOCAB_TYPE_TEST = 7` (a tail append — no existing enumerator renumbers, and this project reads `vocab_type` as a raw int in `ModelMeta.getVocabType()` and emits it `static_cast`-ed in `jllama.cpp`, so no Java-side constant can go stale) and `llama_adapter_lora_init_from_file_ptr` (#28993, additive; adapters are loaded by path here, never by `FILE*`). `common/chat.cpp` picks up a Ling 3.0 / Bailing V3 detection arm and `common/parsers/ling3.cpp` (#28682), and `common/parsers/gemma4.cpp` **fixes a real bug on a path this project serves**: with `tool_choice == required` the grammar now terminates at the tool call instead of falling through to the content scan (#29115). `common/json-schema-to-grammar.cpp` fixes a second one — `gbnf_escape_length()` now accepts `\-`, so a JSON-schema `pattern` containing an escaped hyphen no longer produces a grammar the parser rejects (#29127). `src/llama-model.{cpp,h}` gain `load_swa_pattern()` with 20 `src/models/*.cpp` architectures rewritten onto it and `TENSOR_SKIP` honoured in `create_tensor_gate_up_exps()` (#29042, #29014); `tools/mtmd/clip.cpp` returns false instead of proceeding when `ggml_backend_sched_alloc_graph()` fails (#28149 / #26070). **`ggml/include` is byte-identical across the whole range**, so no ggml public API moved at all. | +| b11018–b11062 | patches + upstream verification | **Nine patches still, none dropped — but two needed a refresh, the first in several ranges.** One upstream commit is responsible: **#29125** ("server : improve startup log messages", first tagged b11053) adds an `SRV_INF("initializing ...")` line immediately above `llama_server()`'s argv parse and a two-line `TODO` comment above `common_params_parse()` in `common/arg.h`. `0001` anchors hunks on both spots and `0006` replaces the very line `0001` flips, so both went stale **on context only** — the refresh changes `@@` line numbers, three context lines and the index blob hashes, and not one added or removed line. Replayed in filename order against pristine **b11055 and b11062**: all nine apply clean at both. **`0007`'s standing invariant is intact and provably so** — its `-` side is a verbatim copy of the route table it factors out of `llama_server()`, so a clean apply *is* the proof upstream did not touch that block; #29125's edits sit above it (the CORS warning) and below it (the `warn_names` loop), never inside. **The three mechanical `tools/server/` contract greps have no input** despite `tools/server/` being touched: `server-schema.cpp`, `server-task.cpp` and `server-context.cpp` are byte-identical b11018→b11062, verified by blob hash rather than by reading a diff, so the request-field set, the field bounds and the response-key set cannot have moved. **All standing drop-checks still say "still required"**, run against the pristine tag because the fail-loud applier detects "does not apply" but never "upstream already fixed this": `0001` (`common_params_parse_main` 0 occurrences in `b11062:common/arg.h`, WIN32 override still at `common/arg.cpp:1282`), `0002` (`params_base.load_progress_callback = load_progress_callback` still unguarded at `server-context.cpp:1095`), `0010` (`{"vocab_type", meta.model_vocab_type}` still uncast at `server-context.cpp:4554`), `0011` (the `common_peg_until_parser` `INVALID` branch still returns `FAIL` unconditionally, ignoring `ctx.is_lenient()`, while the `INCOMPLETE` branch right above it honours it), `0012` (bare `splits[i] /= split_sum` at `llama-model.cpp:1518`, no zero guard), and `0003`/`0006`/`0007`/`0008` absent upstream. **Verified at the target from a fresh configure**: stamp head `3cf03257f` with nine SHA-256 lines, `verify-patches-applied.sh` green (9 applied, 0010 cast present), extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean (0 errors, 0 warnings), `ctest` 551/551, `nm -D` 40 `Java_*` exports and 0 mangled, `NativeLibraryLoadSmokeTest` 4/4 with 0 skipped after a `mvn clean`, `mvn test` 1763/0 (269 model-gated skips in a HF-blocked sandbox), `verify-test-counts.sh` 1763 across 119 classes, SpotBugs 0, spotless clean. **The OpenVINO SDK pin moved with it**: #29009 takes upstream's own `OPENVINO_VERSION_MAJOR`/`OPENVINO_VERSION_FULL` to 2026.4, and this project's two OpenVINO classifier jobs — which had drifted two releases behind at 2026.2.1 — now install `2026.4` / `2026.4.0.22959.99c81491cc3` from the same URL template upstream's `{linux,windows}-setup-openvino` actions use. ggml-openvino is developed against whatever pair upstream pins, so tracking it is the cheaper end of the trade: a lagging pin does not fail on the bump that introduces the drift, it fails on some later one, in a job whose runner has no Intel GPU to reproduce on. **Not verifiable from the bump sandbox** — `storage.openvinotoolkit.org` is blocked by the network policy, so neither archive URL could be HEAD-checked here; the evidence they resolve is that upstream's own release jobs download exactly these two URLs at b11062. Per the classifier policy the step is fail-loud, so a wrong URL reds the job rather than shipping a backend-less jar. Both jobs now carry a keep-in-sync note naming upstream's two variables as the source of truth, so the next bump has somewhere to look instead of rediscovering the coupling. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index 028b45081..9486e622f 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b11018 + GIT_TAG b11062 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/patches/0001-win32-arg-parse-embed-guard.patch b/llama/patches/0001-win32-arg-parse-embed-guard.patch index 8527089ad..56fd186c8 100644 --- a/llama/patches/0001-win32-arg-parse-embed-guard.patch +++ b/llama/patches/0001-win32-arg-parse-embed-guard.patch @@ -44,11 +44,11 @@ index 79480e06f..ed7793b4d 100644 std::vector supported_tmpl; int32_t res = llama_chat_builtin_templates(nullptr, 0); diff --git a/common/arg.h b/common/arg.h -index 8f609e356..62c615d29 100644 +index 203d1b4e1..a70f8882c 100644 --- a/common/arg.h +++ b/common/arg.h -@@ -123,6 +123,11 @@ struct common_params_context { - // if one argument has invalid value, it will automatically display usage of the specific argument (and not the full usage message) +@@ -126,6 +126,11 @@ struct common_params_context { + // this is a side-effect that should be avoided bool common_params_parse(int argc, char ** argv, common_params & params, llama_example ex, void(*print_usage)(int, char **) = nullptr); +// Like common_params_parse(), but first recovers the process command line as UTF-8 argv on @@ -483,12 +483,12 @@ index f2179ed27..6d958a861 100644 } if (params.out_file.empty()) { diff --git a/tools/server/server.cpp b/tools/server/server.cpp -index a3b2a8b0f..80d6a3ff6 100644 +index 1167c0aea..28f18c1bb 100644 --- a/tools/server/server.cpp +++ b/tools/server/server.cpp -@@ -102,7 +102,7 @@ int llama_server(int argc, char ** argv) { - // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free - server_stream_session_manager_start(); +@@ -104,7 +104,7 @@ int llama_server(int argc, char ** argv) { + + SRV_INF("%s", "initializing ...\n"); - if (!common_params_parse(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { diff --git a/llama/patches/0006-server-embed-native-server-jni.patch b/llama/patches/0006-server-embed-native-server-jni.patch index c9e2f4d61..23d73c014 100644 --- a/llama/patches/0006-server-embed-native-server-jni.patch +++ b/llama/patches/0006-server-embed-native-server-jni.patch @@ -1,5 +1,5 @@ diff --git a/tools/server/server.cpp b/tools/server/server.cpp -index 7f9ca414..9c0caf18 100644 +index 28f18c1bb..8aeed4b99 100644 --- a/tools/server/server.cpp +++ b/tools/server/server.cpp @@ -25,6 +25,28 @@ @@ -31,9 +31,9 @@ index 7f9ca414..9c0caf18 100644 static inline void signal_handler(int signal) { if (is_terminating.test_and_set()) { // in case it hangs, we can force terminate the server by hitting Ctrl+C twice -@@ -97,7 +119,13 @@ int llama_server(int argc, char ** argv) { - // touch it. lifecycle is symmetric, stop_gc() runs in clean_up() before backend free - server_stream_session_manager_start(); +@@ -104,7 +126,13 @@ int llama_server(int argc, char ** argv) { + + SRV_INF("%s", "initializing ...\n"); - if (!common_params_parse_main(argc, argv, params, LLAMA_EXAMPLE_SERVER)) { + // [jllama] embedded (JNI) callers forward a clean UTF-8 argv, so honor it exactly via @@ -46,7 +46,7 @@ index 7f9ca414..9c0caf18 100644 return 1; } -@@ -433,7 +461,10 @@ int llama_server(common_params & params, int argc, char ** argv) { +@@ -493,7 +521,10 @@ int llama_server(common_params & params, int argc, char ** argv) { } // register signal handler if not running by CLI diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 135013783..7936ed6b7 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11018"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11062"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b11018-"} — call + * plus the resolved upstream commit, e.g. {@code "b11062-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11018"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b11062"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b11018"; + public static final String LLAMA_CPP_VERSION = "b11062"; // Constants holder — not instantiable. private LlamaCppVersion() {}