feat: upgrade llama.cpp to b11018 - #440
Merged
Merged
Conversation
Small, clean range: 6 commits, 28 KiB, 9 files — under the runbook's 100 KiB threshold, so a single step with no chunking. Eight of the nine files are ggml backend internals and the ninth is CODEOWNERS: #28953 fixes a SYCL allocation failure above 19.3 GB on B70, #28929 fuses the SiLU epilogue into the SYCL ssm_conv kernel, #25483 skips unneeded MoE work in the Vulkan mul_mm coopmat1 path, #28996 fixes buffer_reference alignment in the Vulkan im2col shaders, #28984 clears OpenCL warnings, and #29003 drops a code owner for test-llama-archs. Zero files under common/, include/, tools/server/ or tools/mtmd/, so every row of the API-compatibility table is vacuously satisfied and the three mechanical tools/server contract greps have no input. No patch-target file is in the range, so all nine patches are untouched. Verified at the target tag from a fresh configure: stamp head c9a5eeeb3 with nine SHA-256 lines, verify-patches-applied.sh green, extraction unchanged at 138 CLI / 57 request / 15 trainer names, Release build clean (0 errors, 0 warnings), ctest 551/551, nm -D 40 Java_* exports and 0 mangled, NativeLibraryLoadSmokeTest 4/4 with 0 skipped after a clean, mvn test 1763/0, SpotBugs 0, spotless clean. All six standing drop-checks still report "still required" against the pristine tag: 0001 (common_params_parse_main absent from arg.h, WIN32 override at arg.cpp:1282), 0002 (load_progress_callback unguarded at server-context.cpp:1095), 0010 (vocab_type uncast at server-context.cpp:4554), 0012 (bare splits[i] /= split_sum at llama-model.cpp:1518), and 0003/0006/0008 absent upstream. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Two rows per the runbook's step 4: what moved (all ggml backend internals plus CODEOWNERS; zero review-surface files) and the patch/verification row. The second row also records something the project had not been able to state before: run #958 on the b11012 PR was the first full-matrix CI execution since b10948, 58/66 green with only the two documented-expected GPG jobs red, and it is what finally exercised the trainer-model wiring, verify-test-counts.sh and both aarch64 fat-jar smoke jobs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
bernardladenthin
had a problem deploying
to
maven-central
September 17, 2026 10:42 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
startgate
September 17, 2026 10:42 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
September 17, 2026 10:42 — with
GitHub Actions
Failure
|
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
b11012→b11018— 6 commits, 28 KiB, 9 files. Under the runbook's 100 KiB threshold, so a single step with no chunking.ggml/srcbackend internals; the ninth isCODEOWNERS. Zero files undercommon/,include/,tools/server/ortools/mtmd/, and no patch target touched.What changed
ssm_convkernel (newssm_conv.{cpp,hpp},fusion.cpp)mul_mmcoopmat1 pathbuffer_referencealignment inim2col.comp/im2col_3d.comptest-llama-archsEvery change is confined to a backend the classifier jobs build but whose internals this project never calls; the default JAR's CPU path is untouched. Every row of the API-compatibility table is vacuously satisfied and the three mechanical
tools/servercontract greps have no input.Test plan
docs/history/llama-cpp-breaking-changes.mdVerified at the target tag from a fresh configure (
rm -rf build, realFetchContentpath):c9a5eeeb3with nine SHA-256 linesverify-patches-applied.sh9 applied, tree dirty, patches/0010 cast presentctestnm -DJava_*exports, 0 C++-mangledNativeLibraryLoadSmokeTestclean— cross-validatesLLAMA_CPP_VERSIONagainst the linkedbuild-infomvn testAll six standing drop-checks against the pristine tag — all still required:
0001(common_params_parse_main0 occurrences inb11018:common/arg.h, WIN32 override still atcommon/arg.cpp:1282) ·0002(load_progress_callbackstill unguarded atserver-context.cpp:1095) ·0003/0006/0008(all three symbols absent fromb11018:tools/server/) ·0010(vocab_typestill uncast atserver-context.cpp:4554) ·0012(baresplits[i] /= split_sum;still atsrc/llama-model.cpp:1518, unmoved from b11012)Note
The previous PR (#439) got a real full-matrix run — #958, 2h20m, 66 jobs: 58 success / 2 failure / 6 skipped. The two failures are the
Verify GPG signing keypair, whichpublish.ymldocuments as the expected red on apull_requestevent (themaven-centralenvironment withholds secrets there).That run was the first full-matrix execution since b10948, and it is what finally exercised the CI wiring added earlier in this series: the trainer-model download,
verify-test-counts.shon all six Java jobs, and both aarch64 fat-jar smoke jobs — all green.This PR's own run is subject to the usual start-gate abort window, so it may or may not execute; the same two GPG jobs will be red either way.
Related issues / PRs
Follows #439. No upstream issue — nothing in this range required a report.
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.md🤖 Generated with Claude Code
https://claude.ai/code/session_01Cft7guQngfyKycdJfBEfJP
Generated by Claude Code