Skip to content

feat: llama.cpp RPC — --rpc client and an in-JVM RpcServer - #461

Merged
bernardladenthin merged 6 commits into
mainfrom
claude/busy-archimedes-lo9uyy
Sep 28, 2026
Merged

bernardladenthin merged 6 commits into
mainfrom
claude/busy-archimedes-lo9uyy

Conversation

@bernardladenthin

Copy link
Copy Markdown
Owner

Summary

  • RPC client. A model can offload its layers to llama.cpp RPC servers on other machines.
    • Java: ModelParameters.setRpcServers(RpcEndpoint...).
    • CLI: --rpc host:port[,…], on both HTTP servers.
    • value.RpcEndpoint validates endpoints: IPv4 or host name. The transport has no IPv6, and upstream's parser would cut an IPv6 literal at its first colon.
  • RPC server. RpcServer serves this machine's devices (every GPU found, else the CPU) and is the in-JVM counterpart of rpc-server.
    • Binds to loopback only; startOnNetwork is the explicit opt-in and logs a warning.
    • One server per process.
    • Also runnable as java -cp <jar> net.ladenthin.llama.RpcServer.
  • Still one library, no new runtime dependency.
    • ggml-rpc is a static backend linked into jllama on every platform and in every classifier.
    • The transport is plain TCP. jllama.dll already imported WS2_32.dll through cpp-httplib.
    • GGML_RPC_RDMA is forced off. Upstream enables it whenever the build host has libibverbs/librdma, which would make the library unloadable on any machine without rdma-core.
  • patches/0015 makes ggml-rpc safe to embed. Upstream GGML_ABORTs on every client-side connection problem, which in a JVM kills the application.
    • Registration fails softly instead: an unreachable, malformed or non-RPC endpoint becomes a LlamaException naming it.
    • A cached registration is re-checked.
    • A gone server's device reports 0/0 memory instead of aborting.
    • The server can be stopped (ggml_backend_rpc_stop_server) and reports whether it is listening (ggml_backend_rpc_server_listening).
    • Transport fixes: file-descriptor leaks and SIGPIPE.
  • rpc_support.hpp runs before every parse (LlamaModel and NativeServer).
    • It registers the --rpc servers up front.
    • ggml never forgets a registered RPC device and puts RPC devices first by default. So when an earlier load in the same JVM registered one, a load that did not ask for it gets an explicit --device list without it. That list is chosen by llama.cpp's own default rules.
    • Without this, a load after a stopped server aborts the JVM. That was confirmed by a falsification run with the filter disabled.
  • .github/verify-native-deps.py (in the package job) checks every shipped native library against the dependencies it is allowed to have. It reads ELF, PE (including Windows arm64) and Mach-O directly.
    • It found a pre-existing defect on its first run: the macOS dylib links Homebrew's openssl@3, so it does not load on a Mac without that formula. This is true for published 5.1.0 and for the current snapshot.
    • The two paths are allowlisted as a known defect so the check only flags new dependencies. The fix is in TODO.md because it needs a macOS CI run.
  • Side effect. llama_supports_gpu_offload() is now true on CPU-only builds: the -ngl "no usable GPU" warning disappears and the load log shows "offloading" lines. Upstream's own release binaries (all built with RPC) behave the same, and nothing else reads it.

Test plan

  • C++: test_rpc.cpp, 22 tests. ctest 581/581 on a fresh build directory, where the applier applied all nine patches, 0015 included. Covered:
    • the selection rules;
    • a real loopback server/client (mul_mat over RPC equals the local CPU result);
    • stop, including a still-connected client;
    • the unreachable, malformed and stale-server paths;
    • memory 0/0 for a gone server.
  • Java unit tests: RpcEndpointTest, RpcServerOptionsTest, plus additions to ModelParametersExtendedTest, OpenAiServerCliTest and WireNameRegistryTest. PIT on value.*: 388/388.
  • Java native, model-free: RpcServerTest (8 tests): lifecycle, restart on the same port, single instance, port taken, unreachable --rpc load. The JVM survives.
  • Java model-backed: RpcIntegrationTest. It checks that the layers are on the RPC server (the load log shows model buffer size for the endpoint). It then loads a model without --rpc after the server stopped, at -lv 4, which queries every registered device's memory.
    • Run locally against a tiny generated GGUF, because HuggingFace is blocked in this sandbox.
    • With the device filter disabled, the second load aborts the JVM.
  • Two-JVM fat-jar smoke .github/smoke-rpc-fatjar.sh, run locally on the default fat jar. It is wired into smoke-fatjar-linux and requires:
    • a chat completion over RPC;
    • an RPC model buffer in the log;
    • the server accepting a client;
    • a clean non-SIGABRT exit naming the endpoint for a server nobody runs.
  • Full llama Java suite: 1817 tests, 0 failures. spotbugs:check, spotless, javadoc, clang-format 23.1.1 and REUSE all clean.
  • CI is green on this branch. This is the first run on Windows (x86, x64, arm64), macOS, Android, s390x and the GPU classifiers.
  • Docs / CHANGELOG updated where applicable: README "Distributed inference over RPC", CLAUDE.md (RPC section, patch row, test table, JNI entry-point count), CHANGELOG, TODO.

Related issues / PRs

Follows #460. The deferred follow-ups are in TODO.md "RPC backend — follow-ups":

  • a server lost mid-inference still aborts the process (an upstream limitation);
  • filing 0015 upstream;
  • Android on a device;
  • TLS;
  • RDMA.

Checklist

  • I have read CONTRIBUTING.md and CODE_OF_CONDUCT.md
  • My commits follow Conventional Commits
  • No security-sensitive changes. This adds a network server. It is loopback-only by default, remote binding is an explicit opt-in with a warning, and the README documents that the protocol has no authentication or encryption.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8


Generated by Claude Code

Offload a model's layers to llama.cpp RPC servers on other machines
(ModelParameters.setRpcServers, --rpc on both HTTP servers), and serve this
machine's devices with RpcServer, the in-JVM rpc-server (loopback-only unless
startOnNetwork; also runnable as a main class). ggml-rpc is a static backend,
so every artifact stays one jllama library; GGML_RPC_RDMA is forced off so no
runtime dependency is added.

patches/0015 makes ggml-rpc safe to embed. Registering an unreachable,
malformed or non-RPC server is a failure instead of a GGML_ABORT, and a cached
registration is re-checked. A gone server's device reports 0/0 memory. The
server can be stopped (ggml_backend_rpc_stop_server) and reports whether it
is listening. The transport gets fd-leak and SIGPIPE fixes. A server lost
mid-inference still aborts; that is on TODO.

rpc_support.hpp runs before every parse (LlamaModel, NativeServer). It
registers --rpc servers up front with a clear error, and since ggml never
forgets a registered RPC device, a load that did not ask for one gets an
explicit --device list without it.

Tests: test_rpc.cpp (22, incl. a real loopback server/client and the
stop/unreachable/stale paths), RpcEndpointTest, RpcServerOptionsTest,
RpcServerTest (native, model-free), RpcIntegrationTest (draft model,
layers proven on the RPC server, then a stale-server load at -lv 4), and a
two-JVM fat-jar smoke in smoke-fatjar-linux. verify-native-deps.py holds
every shipped library to its known dependencies; on its first run it found
that the macOS dylib links Homebrew OpenSSL (known defect, on TODO).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
value.* is gated at 100%. The digit check's mutants survived because
Integer.parseInt also rejects a non-digit with a NumberFormatException (an
IllegalArgumentException), so the test now asserts the port message and a port
containing 0 and 9. The second-colon check is written without a boundary
(lastIndexOf != colon): ">= 0" vs "> 0" was an equivalent mutant.
PIT 388/388.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8

Copy link
Copy Markdown
Owner Author

Two checks are red that are not caused by this PR:

  • Verify GPG signing key (no secrets printed)
  • Verify GPG signing key — Gradle/BouncyCastle path (no secrets printed)

Both jobs run in the maven-central environment. Its secrets are not available on a PR branch, so the jobs end within seconds without writing a log. They fail the same way on every PR run of this branch, including #459 and #460.

This PR changes neither job. No change on this branch can fix them, and a re-run would fail the same way. The checks that matter here are the build, test and smoke jobs, which run after the start gate.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

claude-review failed too, but not because of this diff. The job ended 245 ms after the model started, with total_cost_usd: 0 and an empty modelUsage, and reported Claude result reported subtype success with is_error:true. The review never read any code, so the error came from the review action or its API access before any review began.

Nothing on this branch can fix it. A re-run of the job, or a later push, will start a fresh review.


Generated by Claude Code

…cle without them

SonarCloud's analysis build has no libjllama, so RpcServer counted as
uncovered: its static initialiser loaded the natives and every lifecycle
branch needed JNI (61.3% coverage on new code, gate 80%). The four native
methods move to the package-private RpcServerNative (JNI symbols renamed in
rpc_bridge.cpp). RpcServer uses them through RpcServer.Backend, so the class
loads without natives. RpcServerLifecycleTest covers start/stop,
idempotent close, bind failure, a throwing serve, start timeout, interrupt,
cache directory creation and failure, and the single-instance slot being
released on every failure path, all over a fake backend.

RpcServerTest, RpcIntegrationTest and the two-JVM fat-jar smoke pass against
the rebuilt library.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
The macOS 15 Metal job aborted in RpcLoopback.ServerComputesTheSameResultAsTheLocalCpu:
the loopback server served the runner's paravirtual Metal GPU, which has no MUL_MAT,
and ggml-rpc's client answers every supports_op with true (upstream TODO), so the
scheduler never falls back and the server hits GGML_ABORT("unsupported op").
RpcIntegrationTest would have taken the JVM down the same way.

- rpc_support.hpp: server_devices(names) resolves devices by name
  (case-insensitive, deduplicated); an unknown name lists the available ones,
  a registered RPC device is refused.
- RpcServer: startLocal/startOnNetwork overloads with a device list, --device on
  the command line (upstream rpc-server -d), getDevices(); names are resolved
  before the server thread starts.
- Every test that computes over RPC serves CPU; the fat-jar smoke uses
  --device CPU and asserts it.
- Docs: README, CLAUDE.md, CHANGELOG; the upstream supports_op gap in TODO.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
…roj device

CI (Java Tests Ubuntu): RpcIntegrationTest stopped its server, then
TtsIntegrationTest in the same fork aborted the JVM in get_dispatcher() while
common_fit_params built a context over the gone server. TTS and the trainer
build common_params themselves, so prepare_argv never saw them.

- rpc_support.hpp: exclude_stale_devices(params.devices) applies the same device
  selection to a hand-built, null-terminated device vector; tts_engine and
  train_engine call it before common_init_from_params.
- The multimodal projector does not use the device list: clip takes the first
  registered GPU-type device, and RPC devices are GPU-type and registered last,
  so on a GPU-less host it would pick the stale server. prepare_argv now pins
  --mmproj-device (or --no-mmproj-offload) in that case unless the caller named
  it; TTS applies the same answer to its mtmd params.
- Tests: 3 new C++ cases (587 total) plus params-level checks in the stale-server
  test; RpcIntegrationTest ends with a TextToSpeech load over the stale registry
  (verified to abort the JVM with the new call removed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
CI (Smoke test all-backends fat jar, Linux): RpcServer never reported listening
within 60 s. RpcServer is the first entry point that loads the library before
LlamaModel. Loading runs JNI_OnLoad, whose GetFieldID on LlamaModel initializes
that class; its static block called LlamaLoader.initialize() again on the loading
thread, and the reentrant lock let a second complete load run: it cleared the
temp files and, with a multi-backend jar, extracted and probed every GPU backend
again (CUDA, ROCm, SYCL are hundreds of MB) over the library being loaded.

- LlamaLoader.runOnceOnThisThread: the nested call returns at once; later calls
  and other threads still run the load (BackendManifestLoadTest relies on that).
  Three unit tests pin nesting, re-running and a failed load staying retryable.
- smoke-rpc-fatjar.sh waits up to 300 s like the NativeServer smoke and fails when
  the backend is selected more than once. Reproduced locally with a hand-built
  multi-backend jar: two selections before, one after; the smoke passes on it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
@sonarqubecloud

Copy link
Copy Markdown

@bernardladenthin
bernardladenthin merged commit 80ddeea into main Sep 28, 2026
80 of 84 checks passed
@bernardladenthin
bernardladenthin deleted the claude/busy-archimedes-lo9uyy branch September 28, 2026 19:54

This branch had an error being deployed

1 failed and 1 active deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants