feat: llama.cpp RPC — --rpc client and an in-JVM RpcServer - #461
Conversation
Offload a model's layers to llama.cpp RPC servers on other machines (ModelParameters.setRpcServers, --rpc on both HTTP servers), and serve this machine's devices with RpcServer, the in-JVM rpc-server (loopback-only unless startOnNetwork; also runnable as a main class). ggml-rpc is a static backend, so every artifact stays one jllama library; GGML_RPC_RDMA is forced off so no runtime dependency is added. patches/0015 makes ggml-rpc safe to embed. Registering an unreachable, malformed or non-RPC server is a failure instead of a GGML_ABORT, and a cached registration is re-checked. A gone server's device reports 0/0 memory. The server can be stopped (ggml_backend_rpc_stop_server) and reports whether it is listening. The transport gets fd-leak and SIGPIPE fixes. A server lost mid-inference still aborts; that is on TODO. rpc_support.hpp runs before every parse (LlamaModel, NativeServer). It registers --rpc servers up front with a clear error, and since ggml never forgets a registered RPC device, a load that did not ask for one gets an explicit --device list without it. Tests: test_rpc.cpp (22, incl. a real loopback server/client and the stop/unreachable/stale paths), RpcEndpointTest, RpcServerOptionsTest, RpcServerTest (native, model-free), RpcIntegrationTest (draft model, layers proven on the RPC server, then a stale-server load at -lv 4), and a two-JVM fat-jar smoke in smoke-fatjar-linux. verify-native-deps.py holds every shipped library to its known dependencies; on its first run it found that the macOS dylib links Homebrew OpenSSL (known defect, on TODO). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
value.* is gated at 100%. The digit check's mutants survived because Integer.parseInt also rejects a non-digit with a NumberFormatException (an IllegalArgumentException), so the test now asserts the port message and a port containing 0 and 9. The second-colon check is written without a boundary (lastIndexOf != colon): ">= 0" vs "> 0" was an equivalent mutant. PIT 388/388. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
|
Two checks are red that are not caused by this PR:
Both jobs run in the This PR changes neither job. No change on this branch can fix them, and a re-run would fail the same way. The checks that matter here are the build, test and smoke jobs, which run after the start gate. Generated by Claude Code |
|
Nothing on this branch can fix it. A re-run of the job, or a later push, will start a fresh review. Generated by Claude Code |
…cle without them SonarCloud's analysis build has no libjllama, so RpcServer counted as uncovered: its static initialiser loaded the natives and every lifecycle branch needed JNI (61.3% coverage on new code, gate 80%). The four native methods move to the package-private RpcServerNative (JNI symbols renamed in rpc_bridge.cpp). RpcServer uses them through RpcServer.Backend, so the class loads without natives. RpcServerLifecycleTest covers start/stop, idempotent close, bind failure, a throwing serve, start timeout, interrupt, cache directory creation and failure, and the single-instance slot being released on every failure path, all over a fake backend. RpcServerTest, RpcIntegrationTest and the two-JVM fat-jar smoke pass against the rebuilt library. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
The macOS 15 Metal job aborted in RpcLoopback.ServerComputesTheSameResultAsTheLocalCpu:
the loopback server served the runner's paravirtual Metal GPU, which has no MUL_MAT,
and ggml-rpc's client answers every supports_op with true (upstream TODO), so the
scheduler never falls back and the server hits GGML_ABORT("unsupported op").
RpcIntegrationTest would have taken the JVM down the same way.
- rpc_support.hpp: server_devices(names) resolves devices by name
(case-insensitive, deduplicated); an unknown name lists the available ones,
a registered RPC device is refused.
- RpcServer: startLocal/startOnNetwork overloads with a device list, --device on
the command line (upstream rpc-server -d), getDevices(); names are resolved
before the server thread starts.
- Every test that computes over RPC serves CPU; the fat-jar smoke uses
--device CPU and asserts it.
- Docs: README, CLAUDE.md, CHANGELOG; the upstream supports_op gap in TODO.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
…roj device CI (Java Tests Ubuntu): RpcIntegrationTest stopped its server, then TtsIntegrationTest in the same fork aborted the JVM in get_dispatcher() while common_fit_params built a context over the gone server. TTS and the trainer build common_params themselves, so prepare_argv never saw them. - rpc_support.hpp: exclude_stale_devices(params.devices) applies the same device selection to a hand-built, null-terminated device vector; tts_engine and train_engine call it before common_init_from_params. - The multimodal projector does not use the device list: clip takes the first registered GPU-type device, and RPC devices are GPU-type and registered last, so on a GPU-less host it would pick the stale server. prepare_argv now pins --mmproj-device (or --no-mmproj-offload) in that case unless the caller named it; TTS applies the same answer to its mtmd params. - Tests: 3 new C++ cases (587 total) plus params-level checks in the stale-server test; RpcIntegrationTest ends with a TextToSpeech load over the stale registry (verified to abort the JVM with the new call removed). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
CI (Smoke test all-backends fat jar, Linux): RpcServer never reported listening within 60 s. RpcServer is the first entry point that loads the library before LlamaModel. Loading runs JNI_OnLoad, whose GetFieldID on LlamaModel initializes that class; its static block called LlamaLoader.initialize() again on the loading thread, and the reentrant lock let a second complete load run: it cleared the temp files and, with a multi-backend jar, extracted and probed every GPU backend again (CUDA, ROCm, SYCL are hundreds of MB) over the library being loaded. - LlamaLoader.runOnceOnThisThread: the nested call returns at once; later calls and other threads still run the load (BackendManifestLoadTest relies on that). Three unit tests pin nesting, re-running and a failed load staying retryable. - smoke-rpc-fatjar.sh waits up to 300 s like the NativeServer smoke and fails when the backend is selected more than once. Reproduced locally with a hand-built multi-backend jar: two selections before, one after; the smoke passes on it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
|



Summary
ModelParameters.setRpcServers(RpcEndpoint...).--rpc host:port[,…], on both HTTP servers.value.RpcEndpointvalidates endpoints: IPv4 or host name. The transport has no IPv6, and upstream's parser would cut an IPv6 literal at its first colon.RpcServerserves this machine's devices (every GPU found, else the CPU) and is the in-JVM counterpart ofrpc-server.startOnNetworkis the explicit opt-in and logs a warning.java -cp <jar> net.ladenthin.llama.RpcServer.ggml-rpcis a static backend linked intojllamaon every platform and in every classifier.jllama.dllalready importedWS2_32.dllthrough cpp-httplib.GGML_RPC_RDMAis forced off. Upstream enables it whenever the build host has libibverbs/librdma, which would make the library unloadable on any machine without rdma-core.patches/0015makes ggml-rpc safe to embed. UpstreamGGML_ABORTs on every client-side connection problem, which in a JVM kills the application.LlamaExceptionnaming it.ggml_backend_rpc_stop_server) and reports whether it is listening (ggml_backend_rpc_server_listening).rpc_support.hppruns before every parse (LlamaModel and NativeServer).--rpcservers up front.--devicelist without it. That list is chosen by llama.cpp's own default rules..github/verify-native-deps.py(in thepackagejob) checks every shipped native library against the dependencies it is allowed to have. It reads ELF, PE (including Windows arm64) and Mach-O directly.openssl@3, so it does not load on a Mac without that formula. This is true for published 5.1.0 and for the current snapshot.TODO.mdbecause it needs a macOS CI run.llama_supports_gpu_offload()is now true on CPU-only builds: the-ngl"no usable GPU" warning disappears and the load log shows "offloading" lines. Upstream's own release binaries (all built with RPC) behave the same, and nothing else reads it.Test plan
test_rpc.cpp, 22 tests.ctest581/581 on a fresh build directory, where the applier applied all nine patches,0015included. Covered:mul_matover RPC equals the local CPU result);RpcEndpointTest,RpcServerOptionsTest, plus additions toModelParametersExtendedTest,OpenAiServerCliTestandWireNameRegistryTest. PIT onvalue.*: 388/388.RpcServerTest(8 tests): lifecycle, restart on the same port, single instance, port taken, unreachable--rpcload. The JVM survives.RpcIntegrationTest. It checks that the layers are on the RPC server (the load log showsmodel buffer sizefor the endpoint). It then loads a model without--rpcafter the server stopped, at-lv 4, which queries every registered device's memory..github/smoke-rpc-fatjar.sh, run locally on the default fat jar. It is wired intosmoke-fatjar-linuxand requires:llamaJava suite: 1817 tests, 0 failures. spotbugs:check, spotless, javadoc, clang-format 23.1.1 and REUSE all clean.Related issues / PRs
Follows #460. The deferred follow-ups are in
TODO.md"RPC backend — follow-ups":0015upstream;Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.md🤖 Generated with Claude Code
https://claude.ai/code/session_01QSXAwbXqx9u5xsMArMvgT8
Generated by Claude Code