Upgrade llama.cpp from b11018 to b11062 - #441
Merged
Merged
Conversation
First chunk toward b11062. 2 commits, 10 KiB.
#28993 ("gguf : align the data section relative to the GGUF start, not the
file") is the substantive one: gguf.cpp now computes the data-section padding
from the GGUF's own start offset rather than the file offset, so a GGUF
embedded at a non-zero offset in a larger file stays internally consistent.
llama-model-loader.cpp gains the matching guard — loading through a FILE* with
mmap now throws when the data section is not aligned to the CPU tensor
alignment — and include/llama.h grows one new entry point,
llama_adapter_lora_init_from_file_ptr, implemented in llama-adapter.cpp. Both
are additive: the project loads models and LoRA adapters by path, never by
FILE*, so nothing here is on a path jllama reaches.
#29008 adds message_delimiters to the DeepSeek V3.2/V4 parser so the server can
locate user turns for context checkpoints. Parser data only, no API surface.
No file under tools/server/, so the three mechanical server-contract greps have
no input, and no patch target is in this chunk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Second chunk. 2 commits, 579 KiB — well over the runbook's 100 KiB threshold,
and irreducibly so: b11021 does not exist as a tag, so #28732 cannot be split
into a smaller step. The size is also almost entirely bookkeeping.
#28732 ("vulkan: split buffers and debug code into separate files, add shared
headers") moves ~5.4k lines out of ggml-vulkan.cpp into ggml-vulkan-buffers.cpp,
ggml-vulkan-debug.cpp and three new headers (types, push-constants, common).
Diffed as a rename-aware move it is a pure reorganisation: ggml-vulkan/CMakeLists.txt
gains exactly the five new files in the ggml_add_backend_library() call and nothing
else changes about how the backend is built or linked, so the two Vulkan classifiers
(vulkan-linux-*, vulkan-windows-x86-64) are unaffected beyond compiling more,
smaller translation units.
#28947 adds an API/ABI check to upstream's own make-release workflow: a GitHub
workflow plus three scripts/ files. Never executed by this project's CI.
Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, so every row
of the API-compatibility table is vacuously satisfied and no patch target is in
this chunk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Third chunk. 2 commits, 224 KiB — over the threshold, and again irreducibly:
b11023 does not exist as a tag, so #29009 cannot be split.
#29009 ("openvino : Update OpenVINO to 2026.4; fix clangd, MSVC warnings") is
the whole of the size. It rewrites large parts of ggml-openvino's quant and
utils translation units, but every symbol it touches is inside
ggml/src/ggml-openvino/**: no ggml public header, no llama header, nothing this
project compiles against. The remaining files are upstream's own release/
self-hosted workflows and docs/backend/OPENVINO.md.
#27985 fixes the reasoning menu in the WebUI's single-model desktop view. The
WebUI auto-follows GIT_TAG — build-webui re-reads the tag and rebuilds the
matching Svelte UI — so it needs no action here.
Watch item, recorded rather than acted on: upstream's own OpenVINO jobs move
their SDK pin 2026.3.1 -> 2026.4, while this project's two OpenVINO classifier
jobs still install the 2026.2.1 archive (publish.yml). That gap already existed
one release back and the jobs built, and nothing in this diff uses an API newer
than what 2026.2.1 exposes (the ov:: surface it adds is Core/CompiledModel/
InferRequest/Tensor, all long-standing). Per the classifier policy these vendor
install steps are first-pass and fail loud, so if 2026.2.1 stops compiling
ggml-openvino the job reds the pipeline and the pin gets moved then — bumping
it speculatively here cannot be validated on a GPU-less runner.
Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, and no
patch target is in this chunk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Fourth chunk, and the first back under the threshold: 18 commits, 87 KiB.
Two themes on the review surface, both additive:
Allocation-failure hardening (#28149, #26070, #28978). ggml now checks
allocation results instead of assuming success, and graph-buffer reservation
failure is handled rather than fatal; tools/mtmd/clip.cpp picks up the matching
call-site change — clip_encode() returns false when
ggml_backend_sched_alloc_graph() fails instead of proceeding with an
unallocated graph. No ggml public header moved anywhere in this range
(ggml/include is byte-identical b11018..b11062), and no project C++ calls the
scheduler or allocator directly, so this reaches jllama only as
better-behaved-on-OOM inside upstream translation units.
Model plumbing (#29042, #29014, #29018, #29033). llama_model_base gains
load_swa_pattern() — a helper that reads the SWA pattern either as one flag per
layer or as a period to expand — and 20 src/models/*.cpp architectures are
rewritten onto it; create_tensor_gate_up_exps() learns to honour TENSOR_SKIP;
llama-model-saver writes the SWA pattern so 15 more architectures round-trip;
and llama-vocab adds LLAMA_VOCAB_PRE_TYPE_UFAKZEKA = 59. All of it is inside
src/, compiled into the static llama library, none of it visible in
include/llama.h.
llama-model.{cpp,h} are patch 0012's target files, so this chunk is the first
that could have disturbed a patch. It does not: load_swa_pattern() lands ~1700
lines above load_tensors()'s split arithmetic, and 0012 still applies clean.
The rest is backends and upstream CI: an OpenCL bin kernel, F16 FWHT on CPU,
Vulkan IQ3_S MMQ kernels and a raised mul_mat_id expert limit, an RPC ACCEL
skip, a GGML_CPU=OFF/GGML_CUDA=ON cmake fix, a gguf-py Q8_1 block-size fix, and
five workflow tweaks.
No file under tools/server/, so the three mechanical server-contract greps have
no input.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Fifth chunk. 3 commits, 98 KiB, all three in the Hexagon backend (ggml/src/ggml-hexagon/**): ROLL op support (#29105), an im2col update (#29103), and HMX flash-attention head_dim padding for DK=DV=72 (#26539). Nothing under common/, include/, src/, tools/ or the project CMakeLists, so every row of the API-compatibility table is vacuously satisfied, the three mechanical server-contract greps have no input, and no patch target is in this chunk. This project ships no Hexagon classifier, so the code is not even compiled here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Sixth chunk. 5 commits, 57 KiB. One line of public API moves: #29084 adds LLAMA_VOCAB_TYPE_TEST = 7 to include/llama.h's llama_vocab_type enum, with the tokenizer itself in llama-vocab.cpp (a rolling hash over fixed-size chunks, for generating a dummy vocab in test-llama-archs). Purely additive, and the value is a tail append, so no existing constant renumbers. This project surfaces vocab_type as a raw int in two places — jllama.cpp's two "vocab_type" emit sites, both already static_cast<int>-ed, and ModelMeta.getVocabType() which reads it back as an int — so a new enumerator needs no Java-side constant and cannot go stale. (The b10585 common_json enum trap is unrelated and already guarded; nothing here changes it.) llama-model-saver.cpp continues the #29042 round-trip work from chunk 4. The remaining four are backends: Metal FA support checks (#29122) and qwen4exp hc ops (#29000), a CUDA CUB argsort in-place-keys corruption fix (#28389 — real bug fix, reaches the cuda13 classifiers), and an OpenCL flash_attn bin kernel (#29046). No file under tools/server/, so the three mechanical server-contract greps have no input, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Seventh chunk. 2 commits, 115 KiB — over the threshold, irreducibly again (b11051 does not exist as a tag). #29127 is the one line on the review surface and it is a behaviour fix worth naming: common/json-schema-to-grammar.cpp's gbnf_escape_length() now accepts '-' as an escapable character, matching parse_char() in llama-grammar.cpp. A JSON-schema "pattern" containing \- previously produced a grammar the parser then rejected, so a structured-output request using such a pattern failed. This project reaches it through the grammar routing in eval_llama_cmpl_schema(), so it is a real (if narrow) fix for callers passing json_schema/response_format. #28948 is the rest of the bytes: Metal MoE and SSM_CONV fusion optimizations plus a new argsort.metal kernel, with test-backend-ops and the MTL fusion CSV updated alongside. That reaches the default macOS arm64 dylib (Metal is in the default jar, not a classifier), so it is exercised by the three macOS Java jobs and the smoke-fatjar-macos gate. No file under tools/server/, so the three mechanical server-contract greps have no input, and no patch target is in this chunk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
…nd 0006
Eighth chunk. 3 commits, 97 KiB. Two are Hexagon ops (GEGLU_QUICK #29114,
TOP_K #29113, not compiled here); the third is the one that matters.
#29125 ("server : improve startup log messages") is the first commit in this
whole range to touch a patch target. It rewrites llama_server()'s startup
logging — an "initializing ..." line before the argv parse, the CORS and
enabled-features warnings collapsed from multi-line banners to one line each,
and the :8080 port notice likewise — plus a TODO comment above
common_params_parse() in common/arg.h and SRV_INF -> SRV_TRC / a source-tagged
"Available models" listing in server-models.cpp.
Two patches went stale on it, both purely on context, and both are refreshed
here with no change to what they do:
0001 — its common/arg.h hunk anchored on the two comment lines above the
common_params_parse() declaration; upstream inserted two more (the TODO), so
the hunk now anchors on the new last comment line. Its tools/server/server.cpp
hunk anchored on the three lines above the parse call, one of which is now the
new SRV_INF("initializing ...").
0006 — same server.cpp call site, same cause (its hunk 2 replaces the line
0001 just flipped). Hunks 1 and 3 were untouched; only line offsets moved.
The refresh is context-only: `git diff` of the two patch files shows the added/
removed lines byte-identical, with only @@ line numbers, three context lines and
the index blob hashes changing. Verified by replaying the whole stack in
filename order against pristine b11055 AND pristine b11062 — all nine apply
clean at both tags.
0007's standing invariant is intact for the same reason it is checkable at all:
its `-` side is a verbatim copy of the route table it factors out of
llama_server(), so a clean apply proves upstream did not touch that block. It
did not — #29125's edits sit above it (the CORS warning) and below it (the
warn_names loop), never inside.
tools/server/ is touched, so the three mechanical contract greps are in scope
this time. They have no input regardless: server-schema.cpp, server-task.cpp
and server-context.cpp are byte-identical between b11018 and b11062 (verified
by blob hash, not by diff reading), so the request-field set, the field bounds
and the response-key set cannot have moved anywhere in this bump.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Ninth and final chunk, reaching the target release. 7 commits, 84 KiB. Chat parsing is where the substance is, and both items reach this project — common/chat.cpp is compiled into llama-common and the server path runs every completion through common_chat_parse(): #28682 adds a dedicated Ling 3.0 / Bailing V3 parser (common/parsers/ling3.cpp, 194 lines, declared in parsers.h, registered in sources.cmake) and a detection arm in common_chat_try_specialized_template() keyed on "<role>ASSISTANT</role>" + "<arg_key>". Additive: a template that did not match any specialized arm before still does not. #29115 fixes the gemma4 required-tool grammar — with tool_choice == COMMON_CHAT_TOOL_CHOICE_REQUIRED the grammar now ends at the tool call instead of continuing into the content scan, so a request that demands a tool call can no longer come back as prose. A real fix for callers setting tool_choice=required against a gemma4 template. #28832 makes the mamba time-step projection input contiguous before the matmul (ggml_cont on the non-norm branch) — a correctness fix in src/models/. The rest is backend and UI: CUDA sparse FA for qwen4 (#28770, with the matching src/models/qwen4exp.cpp tweak), Metal F16 FWHT (#29094), Hexagon I32 GET_ROWS (#29116), and a WebUI mobile-breakpoint/overflow fix (#29108, auto-followed). No file under tools/server/ in this chunk, and no patch target. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Two rows per the runbook's step 4. The first records what moved and, unusually, the chunking itself: nine steps, three of which break the 100 KiB rule and all three irreducibly — b11021, b11023 and b11051 do not exist as tags, so each of those steps is one upstream commit with no smaller step available. Worth having written down, because "over the threshold" has so far always meant "should have been split" and here three times it does not. The second is the patch/verification row. It carries the first patch refresh in several ranges (0001 and 0006, context-only, both caused by upstream #29125), the evidence that 0007's route-table invariant survived it, and the reason the three tools/server contract greps have no input despite tools/server being touched: server-schema.cpp, server-task.cpp and server-context.cpp are byte-identical across the whole range by blob hash. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
bernardladenthin
had a problem deploying
to
startgate
September 20, 2026 10:44 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 20, 2026 10:44 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
September 20, 2026 10:44 — with
GitHub Actions
Failure
Two numbers in the "C++ unit tests" section had drifted from the suite. The per-file table said test_jni_helpers.cpp carries 63 tests (it carries 70) and the footer said 544 in total (551). Both are now what `ctest` reports on the b11062 pin, and every other row in the table was checked against a per-file count rather than assumed. The jni_helpers row's prose needed a second correction that the numbers alone would have hidden: it said "The last 7 pin jni_guard_impl", which was true when those tests were the tail of the file but stopped being true once the four ReleaseJllamaContext / JllamaContextGuard tests were appended after them. Positional wording goes stale silently, so it now says "Seven of them". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
Both OpenVINO classifier jobs installed the 2026.2.1 archive while upstream llama.cpp moved its own OPENVINO_VERSION_MAJOR / OPENVINO_VERSION_FULL to 2026.3.1 and then, in #29009 (inside this bump's b11022->b11024 chunk), to 2026.4. ggml-openvino is developed against whichever pair upstream pins, so the drift is what eventually breaks the compile — and it does not break on the bump that introduces it, it breaks on some later one, in a job whose runner has no Intel GPU to reproduce on. Tracking upstream is the cheaper end of that trade. Both jobs now use major 2026.4 / full 2026.4.0.22959.99c81491cc3, from the same URL template upstream's linux-setup-openvino and windows-setup-openvino actions use (verified against those two action.yml files at the pinned tag, not guessed from the old string). Each job gains a keep-in-sync note naming upstream's two variables as the source of truth, because the coupling is otherwise invisible: nothing in this repo points at release.yml, and the previous drift happened by simply not looking. Verification limit, stated rather than glossed: the two archive URLs could NOT be reached from the bump sandbox — storage.openvinotoolkit.org is blocked by the network policy (the proxy answers 403 to CONNECT), so no HEAD check was possible. What stands behind them is that upstream's own release jobs download exactly these two URLs at b11062. Per the classifier policy the step is fail-loud, so a wrong URL reds the job rather than shipping a backend-less jar. The b11018-b11062 history row is updated accordingly — it recorded this as a watch item deliberately not acted on, which is no longer what happened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1
bernardladenthin
had a problem deploying
to
startgate
September 20, 2026 10:48 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 20, 2026 10:48 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
September 20, 2026 10:48 — with
GitHub Actions
Failure
|
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
b11018tob11062across all configuration files and documentation0006-server-embed-native-server-jni.patchto account for upstream line number shiftsTest plan
Related issues / PRs
Routine upstream version bump as part of ongoing llama.cpp integration maintenance.
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.mdhttps://claude.ai/code/session_01PLMJFZvnNPM9hRuL8LGPC1