feat: upgrade llama.cpp to b11012 - #439
Merged
Merged
Conversation
First chunk toward b11012. 14 commits, 75 KiB. Two files on the priority review list, both implementation-only behind unchanged signatures — no compile or link consequence, but one of them changes behaviour. common/fit.cpp (#28849, "Change max context length for auto-fitting with unified KV") is the one worth knowing. common_params_fit_impl now sizes n_ctx_max by n_seq_max rather than n_streams, so an auto-sized context (-c 0) with --parallel > 1 and a UNIFIED KV cache gets n_ctx_train * n_seq_max instead of n_ctx_train; kv_unified stays n_streams == 1, so the non-unified path is unchanged. The file is compiled into llama-common and linked into jllama, and nothing in this project calls it directly — it is reached through common_init_from_params when n_ctx == 0. So the effect here is "an auto-sized parallel server may now ask for more context", which is upstream's intended fix, not a regression to work around. common/parsers/qwen3-coder.cpp (#28869) adds "\n</think>" ahead of "</think>" in thinking_end_tags so the newline lands inside the forced message. Parser internals; no API surface. Everything else is backend or upstream tooling: CUDA/HIP im2col access patterns, HIP MoE tile heuristic and AllReduce, Vulkan MUL_MAT_ID tail, Metal mul_mm_id NaN, hexagon copy/DMA paths, spacemit int16 transpose, an RPC compute-graph cache invalidation, qwen4exp hc ops (src/models/), and llama-bench --version (never compiled here — LLAMA_BUILD_TOOLS is OFF). No tools/server/ file moved, so the three mechanical server-contract greps have no input. No patch-target file is in this chunk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Second chunk toward b11012. 3 commits, 33 KiB, dominated by #27625 which
adds a new architecture, HrmTextForCausalLM (DFM Mimir 1B).
This is the chunk that touches patches/0012's targets, src/llama-model.{cpp,h},
so the collision risk was checked rather than assumed. The three edits to
llama-model.cpp are new-arch dispatch entries — an LLM_ARCH_HRM_TEXT case in
llama_model_mapping, a MIRRORED meta-split branch for its aliased cache slots,
and a rope-type case — at lines ~316, ~477 and ~3030. 0012's hunks are the
load_tensors split arithmetic at 1493-1511. They are a thousand lines apart
and cannot interact; llama-model.h likewise gains only an additive
hrm_z_l_init tensor member.
0012's standing drop-check was re-run against pristine b11005 for the same
reason: splits[i] /= split_sum is still bare, with no zero-sum guard, so the
patch is still required and is not droppable here.
The rest of the arch addition is upstream-internal and never reaches this
project's surface: llama-arch.{cpp,h} (enum + tensor names), llama-hparams.h,
llama-context.cpp, llama-model-saver.cpp, src/models/hrm-text.cpp and
models.h. None is a header this project includes. Also here: #28989 lets
Nemotron-H models define only layer_norm_epsilon, and #28995 has hexagon
accept the zeroed rope probe in supports_op.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Third chunk toward b11012. 6 commits, 88 KiB. Nothing under common/, include/, tools/server/ or tools/mtmd/, so every row of the API-compatibility table is vacuously satisfied and the server-contract greps have no input. The three src/ files are upstream-internal. #28549 (CUDA graph for MTP draft) splits llama_context::gf_res_prev into a two-element array so batches with and without outputs get distinct CUDA graph cache keys, and adds a private get_gf_res_prev(); that is a private member of llama_context declared in src/llama-context.h, which is an INTERNAL header this project does not include. The one internal upstream header it does include is src/llama-model.h — from test_model_split.cpp, via the include dir patches/0012 adds — and that file does not move in this chunk. #28965 fixes tensor-parallel split state and granularity for fused QKV on gemma4/qwen35 (llama-model.cpp). The rest is backend-side: #28981 fixes the ggml_backend_sycl_split_buffer_type signature (it gains a leading int main_device), #28975 works around an NVIDIA bug in argsort_large.comp, #28994 adds hexagon Q4_K/Q6_K support, and #28959 is upstream CI. The SYCL signature change is a public ggml-sycl.h symbol but a backend-internal one — nothing in this project calls any ggml backend split-buffer API, and the three sycl-* classifier jobs compile upstream's own consistent tree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Fourth and final chunk, reaching the target release. One commit, 13 KiB: #28988 adds Vulkan support for the qwen4exp hc ops that #28901 introduced earlier in this range (chunk 1), so the two Vulkan classifiers can run that architecture instead of falling back. Entirely ggml/src/ggml-vulkan/** plus its shader generator. Nothing under common/, include/, src/, tools/server/ or tools/mtmd/, and no patch target. Kept as its own chunk rather than folded into b11005 -> b11011: that would have made one 100.6 KiB step, over the runbook's threshold by a hair. "Barely over" is exactly the rationalisation the threshold exists to prevent, and the cost of honouring it here is one extra commit, not an extra build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Two rows per the runbook's step 4. The first covers what moved: the four-chunk walk and why the 13 KiB tail was kept separate, the two common/ files on the review surface (with #28849's auto-context behaviour change spelled out), the new HRM-TEXT architecture, and the two ggml/include headers — one purely additive, one a real but backend-internal signature break. The second is the patch row, and it matters more than usual: this is the first range in a while where a patch target actually moved. It records the collision check against patches/0012 by line, the note that 0012's anchor shifted 1493 -> 1518 under the new arch code (which is why drop-checks are by content, not by line), the six standing drop-checks against the pristine tag, and the full verification numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
bernardladenthin
had a problem deploying
to
maven-central
September 17, 2026 06:40 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
September 17, 2026 06:40 — with
GitHub Actions
Failure
|
6 tasks
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
b10988→b11012(24 commits, 206 KiB, 65 files), walked in four chunks. The tail chunk is 13 KiB and was kept separate on purpose: folding it into the third makes one 100.6 KiB step — over the runbook's threshold by a hair, and "barely over" is the rationalisation the threshold exists to prevent.common/implementation behind unchanged signatures. #28849 (common/fit.cpp) is the one with teeth — an auto-sized context (-c 0) with--parallel > 1and a unified KV cache now getsn_ctx_train * n_seq_maxinstead ofn_ctx_train.HrmTextForCausalLM) and editssrc/llama-model.{cpp,h}—patches/0012's targets. The collision was checked, not assumed, and0012's anchor shifted 1493 → 1518 underneath it.The four chunks
b10988 → b11002fit.cppauto-context sizing, #28869 qwen3-coder\n</think>, #28901 qwen4exp hc ops, CUDA/HIP/Vulkan/Metal/hexagon/RPC backend workb11002 → b11005HrmTextForCausalLM(DFM Mimir 1B), #28989 Nemotron-H, #28995 hexagon rope probeb11005 → b11011b11011 → b11012Neither
common/change is a compile or link consequence — both are implementation-only in TUs upstream compiles intollama-common. Nothing here callscommon_params_fit_impldirectly (it's reached viacommon_init_from_paramswhenn_ctx == 0), so #28849 surfaces as "an auto-sized parallel server may now ask for more context" — upstream's intended fix, not a regression to absorb.Zero
tools/server/files moved, so the three mechanical server-contract greps have no input: the request-field set, itsset_hard_limitsbounds and the emitted response keys cannot have changed.Two
ggml/includeheaders move and both were opened rather than waved past.ggml.his purely additive (a newggml_dsv4_hc_pre_gatedplus a comment — no existing signature moves, and this project calls noggml_dsv4_*).ggml-sycl.h'sggml_backend_sycl_split_buffer_typegains a leadingint main_device— a real signature break, but of a backend-internal symbol no project code calls; the threesycl-*classifier jobs compile upstream's own self-consistent tree.src/llama-context.his an internal header this project does not include (the only internal one it does issrc/llama-model.h, fromtest_model_split.cpp), and #28549's change there is a private member.The patch-collision check (patches/0012)
#27625 edits
src/llama-model.cppin three places — anLLM_ARCH_HRM_TEXTcase inllama_model_mapping, aMIRROREDmeta-split branch for that arch's aliased cache slots, and a rope-type case — at roughly lines 316, 477 and 3030.patches/0012's hunks are theload_tensorssplit arithmetic at 1493–1518. A thousand lines apart, andsrc/llama-model.hgains only an additivehrm_z_l_initmember.Worth noting for the next bump:
0012's anchor moved 1493 → 1518 under the new arch code. That is exactly why the drop-checks are written to match content rather than a line number.Test plan
docs/history/llama-cpp-breaking-changes.mdVerified at the target tag from a fresh configure (
rm -rf build, realFetchContentpath):35822afe58475e0506cd51e6573903e46d4c67c9(= b11012) with nine SHA-256 linesverify-patches-applied.sh9 applied, tree dirty, patches/0010 cast presentctestnm -DJava_*exports, 0 C++-mangledNativeLibraryLoadSmokeTestclean— cross-validatesLLAMA_CPP_VERSIONagainst the linkedbuild-infomvn testjavadoc:jarAll six standing drop-checks run against the pristine tag — all still required, none droppable:
0001(common_params_parse_main0 occurrences inb11012:common/arg.h, WIN32 override still atcommon/arg.cpp:1282) ·0002(load_progress_callbackstill unguarded atserver-context.cpp:1095) ·0003/0006/0008(get_slot_prompt_similarity,llama_server_set_embedded,LLAMA_SERVER_WORKER_CMDall absent fromb11012:tools/server/) ·0010(vocab_typestill uncast atserver-context.cpp:4554) ·0012(baresplits[i] /= split_sum;still atsrc/llama-model.cpp:1518)This is also the first bump to exercise the
verify-patches-applied.shfix from #438 — the guard that had been failing on every correct tree now passes on a real build, as it always should have.Note
The full
Publishmatrix will not run on this PR automatically — it is dispatch-only in practice here (recent push-to-mainandpull_requestruns are cancelled in thestartgateabort window). Validating the branch end-to-end needs aworkflow_dispatchwithpublish_to_centralleft at its defaultfalse.The same four checks will be red and none is from this diff — all inherited, triaged in detail on #437:
claude-review(account-side credentials), bothVerify GPG signing keyjobs (environment: maven-centralsecrets are not delivered on apull_requestevent, whichpublish.ymldocuments as an expected red), andanalyze/CodeQL (KotlinVersionTooRecentErroron Kotlin 2.4.20, red onmain).Related issues / PRs
Follows #438. No upstream issue — nothing in this range required a report.
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.md🤖 Generated with Claude Code
https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Generated by Claude Code