Skip to content

feat: upgrade llama.cpp to b11012 - #439

Merged
bernardladenthin merged 5 commits into
mainfrom
claude/serene-goodall-cq2rgc
Sep 17, 2026
Merged

bernardladenthin merged 5 commits into
mainfrom
claude/serene-goodall-cq2rgc

Conversation

@bernardladenthin

Copy link
Copy Markdown
Owner

Summary

  • Upgrades llama.cpp b10988 → b11012 (24 commits, 206 KiB, 65 files), walked in four chunks. The tail chunk is 13 KiB and was kept separate on purpose: folding it into the third makes one 100.6 KiB step — over the runbook's threshold by a hair, and "barely over" is the rationalisation the threshold exists to prevent.
  • Two files on the priority review list, both common/ implementation behind unchanged signatures. #28849 (common/fit.cpp) is the one with teeth — an auto-sized context (-c 0) with --parallel > 1 and a unified KV cache now gets n_ctx_train * n_seq_max instead of n_ctx_train.
  • The first range in a while where a patch target actually moved. #27625 adds a whole architecture (HrmTextForCausalLM) and edits src/llama-model.{cpp,h} — patches/0012's targets. The collision was checked, not assumed, and 0012's anchor shifted 1493 → 1518 underneath it.
The four chunks
Step Size Contents
b10988 → b11002 75 KiB, 14 commits #28849 fit.cpp auto-context sizing, #28869 qwen3-coder \n</think>, #28901 qwen4exp hc ops, CUDA/HIP/Vulkan/Metal/hexagon/RPC backend work
b11002 → b11005 33 KiB, 3 commits #27625 HrmTextForCausalLM (DFM Mimir 1B), #28989 Nemotron-H, #28995 hexagon rope probe
b11005 → b11011 88 KiB, 6 commits #28549 CUDA graph for MTP draft, #28981 SYCL signature fix, #28975 NV argsort workaround, #28965 TP split state, #28994 hexagon K-quants
b11011 → b11012 13 KiB, 1 commit #28988 Vulkan support for the qwen4exp hc ops added in chunk 1

Neither common/ change is a compile or link consequence — both are implementation-only in TUs upstream compiles into llama-common. Nothing here calls common_params_fit_impl directly (it's reached via common_init_from_params when n_ctx == 0), so #28849 surfaces as "an auto-sized parallel server may now ask for more context" — upstream's intended fix, not a regression to absorb.

Zero tools/server/ files moved, so the three mechanical server-contract greps have no input: the request-field set, its set_hard_limits bounds and the emitted response keys cannot have changed.

Two ggml/include headers move and both were opened rather than waved past. ggml.h is purely additive (a new ggml_dsv4_hc_pre_gated plus a comment — no existing signature moves, and this project calls no ggml_dsv4_*). ggml-sycl.h's ggml_backend_sycl_split_buffer_type gains a leading int main_device — a real signature break, but of a backend-internal symbol no project code calls; the three sycl-* classifier jobs compile upstream's own self-consistent tree. src/llama-context.h is an internal header this project does not include (the only internal one it does is src/llama-model.h, from test_model_split.cpp), and #28549's change there is a private member.

The patch-collision check (patches/0012)

#27625 edits src/llama-model.cpp in three places — an LLM_ARCH_HRM_TEXT case in llama_model_mapping, a MIRRORED meta-split branch for that arch's aliased cache slots, and a rope-type case — at roughly lines 316, 477 and 3030. patches/0012's hunks are the load_tensors split arithmetic at 1493–1518. A thousand lines apart, and src/llama-model.h gains only an additive hrm_z_l_init member.

Worth noting for the next bump: 0012's anchor moved 1493 → 1518 under the new arch code. That is exactly why the drop-checks are written to match content rather than a line number.

Test plan

  • Affected unit / integration tests pass locally
  • CI is green on this branch — the full matrix cannot run automatically here; see the note below
  • Docs / CHANGELOG updated where applicable — two rows appended to docs/history/llama-cpp-breaking-changes.md

Verified at the target tag from a fresh configure (rm -rf build, real FetchContent path):

Check Result
Configure clean; stamp head 35822afe58475e0506cd51e6573903e46d4c67c9 (= b11012) with nine SHA-256 lines
verify-patches-applied.sh 9 applied, tree dirty, patches/0010 cast present
Wire-name extraction 138 CLI / 57 request / 15 trainer — unchanged
Release build clean, 0 errors, 0 warnings
ctest 551/551
nm -D 40 Java_* exports, 0 C++-mangled
NativeLibraryLoadSmokeTest 4/4, 0 skipped after a clean — cross-validates LLAMA_CPP_VERSION against the linked build-info
mvn test 1763 tests, 0 failures
SpotBugs / spotless / javadoc:jar 0 bugs, clean, clean

All six standing drop-checks run against the pristine tag — all still required, none droppable:

0001 (common_params_parse_main 0 occurrences in b11012:common/arg.h, WIN32 override still at common/arg.cpp:1282) · 0002 (load_progress_callback still unguarded at server-context.cpp:1095) · 0003/0006/0008 (get_slot_prompt_similarity, llama_server_set_embedded, LLAMA_SERVER_WORKER_CMD all absent from b11012:tools/server/) · 0010 (vocab_type still uncast at server-context.cpp:4554) · 0012 (bare splits[i] /= split_sum; still at src/llama-model.cpp:1518)

This is also the first bump to exercise the verify-patches-applied.sh fix from #438 — the guard that had been failing on every correct tree now passes on a real build, as it always should have.

Note

The full Publish matrix will not run on this PR automatically — it is dispatch-only in practice here (recent push-to-main and pull_request runs are cancelled in the startgate abort window). Validating the branch end-to-end needs a workflow_dispatch with publish_to_central left at its default false.

The same four checks will be red and none is from this diff — all inherited, triaged in detail on #437: claude-review (account-side credentials), both Verify GPG signing key jobs (environment: maven-central secrets are not delivered on a pull_request event, which publish.yml documents as an expected red), and analyze/CodeQL (KotlinVersionTooRecentError on Kotlin 2.4.20, red on main).

Related issues / PRs

Follows #438. No upstream issue — nothing in this range required a report.

Checklist

  • I have read CONTRIBUTING.md and CODE_OF_CONDUCT.md
  • My commits follow Conventional Commits
  • No security-sensitive changes

🤖 Generated with Claude Code

https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP


Generated by Claude Code

First chunk toward b11012. 14 commits, 75 KiB. Two files on the priority
review list, both implementation-only behind unchanged signatures — no
compile or link consequence, but one of them changes behaviour.

common/fit.cpp (#28849, "Change max context length for auto-fitting with
unified KV") is the one worth knowing. common_params_fit_impl now sizes
n_ctx_max by n_seq_max rather than n_streams, so an auto-sized context
(-c 0) with --parallel > 1 and a UNIFIED KV cache gets n_ctx_train *
n_seq_max instead of n_ctx_train; kv_unified stays n_streams == 1, so the
non-unified path is unchanged. The file is compiled into llama-common and
linked into jllama, and nothing in this project calls it directly — it is
reached through common_init_from_params when n_ctx == 0. So the effect here
is "an auto-sized parallel server may now ask for more context", which is
upstream's intended fix, not a regression to work around.

common/parsers/qwen3-coder.cpp (#28869) adds "\n</think>" ahead of "</think>"
in thinking_end_tags so the newline lands inside the forced message. Parser
internals; no API surface.

Everything else is backend or upstream tooling: CUDA/HIP im2col access
patterns, HIP MoE tile heuristic and AllReduce, Vulkan MUL_MAT_ID tail,
Metal mul_mm_id NaN, hexagon copy/DMA paths, spacemit int16 transpose, an RPC
compute-graph cache invalidation, qwen4exp hc ops (src/models/), and
llama-bench --version (never compiled here — LLAMA_BUILD_TOOLS is OFF).

No tools/server/ file moved, so the three mechanical server-contract greps
have no input. No patch-target file is in this chunk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Second chunk toward b11012. 3 commits, 33 KiB, dominated by #27625 which
adds a new architecture, HrmTextForCausalLM (DFM Mimir 1B).

This is the chunk that touches patches/0012's targets, src/llama-model.{cpp,h},
so the collision risk was checked rather than assumed. The three edits to
llama-model.cpp are new-arch dispatch entries — an LLM_ARCH_HRM_TEXT case in
llama_model_mapping, a MIRRORED meta-split branch for its aliased cache slots,
and a rope-type case — at lines ~316, ~477 and ~3030. 0012's hunks are the
load_tensors split arithmetic at 1493-1511. They are a thousand lines apart
and cannot interact; llama-model.h likewise gains only an additive
hrm_z_l_init tensor member.

0012's standing drop-check was re-run against pristine b11005 for the same
reason: splits[i] /= split_sum is still bare, with no zero-sum guard, so the
patch is still required and is not droppable here.

The rest of the arch addition is upstream-internal and never reaches this
project's surface: llama-arch.{cpp,h} (enum + tensor names), llama-hparams.h,
llama-context.cpp, llama-model-saver.cpp, src/models/hrm-text.cpp and
models.h. None is a header this project includes. Also here: #28989 lets
Nemotron-H models define only layer_norm_epsilon, and #28995 has hexagon
accept the zeroed rope probe in supports_op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Third chunk toward b11012. 6 commits, 88 KiB. Nothing under common/,
include/, tools/server/ or tools/mtmd/, so every row of the API-compatibility
table is vacuously satisfied and the server-contract greps have no input.

The three src/ files are upstream-internal. #28549 (CUDA graph for MTP draft)
splits llama_context::gf_res_prev into a two-element array so batches with and
without outputs get distinct CUDA graph cache keys, and adds a private
get_gf_res_prev(); that is a private member of llama_context declared in
src/llama-context.h, which is an INTERNAL header this project does not
include. The one internal upstream header it does include is src/llama-model.h
— from test_model_split.cpp, via the include dir patches/0012 adds — and that
file does not move in this chunk. #28965 fixes tensor-parallel split state and
granularity for fused QKV on gemma4/qwen35 (llama-model.cpp).

The rest is backend-side: #28981 fixes the ggml_backend_sycl_split_buffer_type
signature (it gains a leading int main_device), #28975 works around an NVIDIA
bug in argsort_large.comp, #28994 adds hexagon Q4_K/Q6_K support, and #28959 is
upstream CI. The SYCL signature change is a public ggml-sycl.h symbol but a
backend-internal one — nothing in this project calls any ggml backend
split-buffer API, and the three sycl-* classifier jobs compile upstream's own
consistent tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Fourth and final chunk, reaching the target release. One commit, 13 KiB:
#28988 adds Vulkan support for the qwen4exp hc ops that #28901 introduced
earlier in this range (chunk 1), so the two Vulkan classifiers can run that
architecture instead of falling back.

Entirely ggml/src/ggml-vulkan/** plus its shader generator. Nothing under
common/, include/, src/, tools/server/ or tools/mtmd/, and no patch target.

Kept as its own chunk rather than folded into b11005 -> b11011: that would
have made one 100.6 KiB step, over the runbook's threshold by a hair. "Barely
over" is exactly the rationalisation the threshold exists to prevent, and the
cost of honouring it here is one extra commit, not an extra build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Two rows per the runbook's step 4.

The first covers what moved: the four-chunk walk and why the 13 KiB tail was
kept separate, the two common/ files on the review surface (with #28849's
auto-context behaviour change spelled out), the new HRM-TEXT architecture, and
the two ggml/include headers — one purely additive, one a real but
backend-internal signature break.

The second is the patch row, and it matters more than usual: this is the first
range in a while where a patch target actually moved. It records the collision
check against patches/0012 by line, the note that 0012's anchor shifted
1493 -> 1518 under the new arch code (which is why drop-checks are by content,
not by line), the six standing drop-checks against the pristine tag, and the
full verification numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
@sonarqubecloud

Copy link
Copy Markdown

@bernardladenthin
bernardladenthin merged commit f5a1219 into main Sep 17, 2026
61 of 66 checks passed
@bernardladenthin
bernardladenthin deleted the claude/serene-goodall-cq2rgc branch September 17, 2026 07:28

This branch had an error being deployed

1 failed and 1 active deployments
startgate — 845d7eea Deployed Sep 17, 2026 by bernardladenthin via Start gate (abort window) #958
maven-central — 845d7eea Deployed Sep 17, 2026 by bernardladenthin via Verify GPG signing key (no secrets printed) #958
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants