Conversation
RPC, Electrum, and Esplora bind only after the store opens, catch-up finishes, and the scripthash index materializes. On a synced mainnet node those took 42 minutes (a schema backfill), 9 hours (IBD), and 33 minutes. A liveness probe on any of those listeners would restart the process in the middle of that work. --health-listen [ADDR] (conf health_listen, bare 127.0.0.1:9332) binds before run_node opens the store and answers GET /healthz with 200 in every phase. Other methods are 405 and other paths 404. Concurrency, body, and timeout limits are always on, as on Esplora. A bind failure stops the node before it touches the datadir. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
/readyz is 503 "not ready: <reason>" until the node can serve, then 200. The first failing check names the reason: - the bring-up phase (opening, starting, catch-up, indexing, stopping), written by run_p2p at each transition; - a configured RPC, Electrum, or Esplora listener that did not bind (their start only warns, so the node would otherwise follow the tip while a probe routes clients to a dead port); - initial block download, read from the same ChainHub::in_ibd that getblockchaininfo.initialblockdownload reports; - the tip more than 6 blocks behind the best header; - with --sh-index, the scripthash index more than 6 blocks behind the tip. The chain reads can touch the store, so /readyz takes them on the blocking pool. /healthz stays 200 in every phase: liveness must not fail during a migration or IBD. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--metrics (conf metrics, refused without --health-listen) adds GET
/metrics on the health listener in the Prometheus text format. It does
not grow a second metrics dialect: each gauge is a value the node already
publishes, named after its source.
- rbitcoin_blocks, _headers, _tip_time_seconds, _initial_block_download:
getblockchaininfo blocks, headers, time, initialblockdownload
- rbitcoin_connections{direction}: getnetworkinfo connections_in/_out,
now counted by one rbitcoin_net::connection_counts both call
- rbitcoin_mempool_transactions, _bytes: getmempoolinfo size, bytes
- rbitcoin_scripthash_lag_blocks: tip: accept sh_lag= (with --sh-index)
- process_resident_memory_bytes: ibd: sizes rss=; process_start_time_seconds
- rbitcoin_build_info, rbitcoin_phase, rbitcoin_ready: /readyz itself
The cross-surface journey scrapes once during IBD and once after and
checks each gauge against RPC at that moment. A scrape takes one peer
snapshot and folds the mempool once (the getmempoolinfo fold) on the
blocking pool.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Esplora, Electrum, historical block serve, and mempool accept meters were sample-and-reset: the 5 s perf tick swapped each atomic to zero, so no reader could see a running count and Prometheus had nothing to scrape. PerfCounter keeps the running total and a window mark. The perf tick takes the change since its previous sample (the DEBUG line prints the same numbers); /metrics reads the total: - rbitcoin_esplora_requests_total, rbitcoin_esplora_request_seconds_total - rbitcoin_electrum_requests_total, rbitcoin_electrum_request_seconds_total - rbitcoin_block_serve_total, rbitcoin_block_serve_bytes_total - rbitcoin_mempool_accepts_total, rbitcoin_mempool_rejects_total A window maximum cannot come from totals, so PerfMax stays a reset and is not exported. Esplora and Electrum share one RequestMeter in place of two copies of the same statics and compare-exchange loop. No extra work per request; the window mark is touched only by the tick. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
operations.md gets the two flags and a "Health probes and metrics" section: the routes, the /readyz reasons in check order, a k8s probe snippet that keeps liveness on /healthz, the metric table with the RPC field or log token each one equals, and a scrape config. SECURITY.md scopes the listener. peer-clients.md rank 4 is landed. A changelog.d fragment covers the new flags and the meters becoming running totals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
in_ibd latches off after the first exit and headers - blocks cannot grow when the node has no peers, so a node that lost every peer after IBD answered /readyz 200 with a stale tip while a load balancer kept routing wallets to it. Add the non-latching half of the IBD staleness check to readiness: tip_header age over ChainHub::max_tip_age_secs (--max-tip-age, default 24h) answers 503 "tip stale (last block Ns ago)". rbitcoin_ready inherits the gate through the same readiness(). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
With --metrics the warning listed /healthz and /readyz only, but /metrics is the endpoint that exposes peer counts and mempool size. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
bkeroack
force-pushed
the
node/health-endpoints
branch
from
September 28, 2026 21:37
177edcd to
45ed58a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements
docs/peer-clients.md"Ranked if we spend time later" rank 4:"
/healthz(and maybe/readyz) on the node listen; Prometheus later as aflag". That table says a row is implemented only when you name it, so treat
this as a proposal. The calls I made are listed under Open questions;
any of them is easy to change.
Why a separate listener
RPC, Electrum and Esplora bind only after the store opens, catch-up finishes
and the scripthash index materializes. On a synced mainnet node:
run_p2psteprun_nodestore open (schema 24→26input.*backfill)run_ibd_or_skip(full IBD, full script checks)enter_tip_mode(scripthash materialize)A liveness probe on any of those listeners fails for the whole window, so k8s
(or any restart-on-unhealthy supervisor) would restart the node in the middle
of the migration or the IBD, and it would start over. The health listener binds
at the top of
run_p2p, beforerun_node.What
--health-listen [ADDR](confhealth_listen; bare →127.0.0.1:9332)GET /healthz200 okin every phaseGET /readyz200 ok, or503 not ready: <reason>for the first failing checkGET /metrics--metricsonly; otherwise 404Other paths 404, other methods 405, no body, no auth. The same always-on
concurrency / body / timeout layers as Esplora. A non-loopback bind logs a
WARN (a k8s
httpGetprobe connects to the pod IP). A bind failure stops thenode before it touches the datadir.
/readyzchecks, in order:opening,starting,catch-up,indexing,stopping), written byrun_p2pat each transition.(
rpc not listening). Their start only warns, so without this the nodefollows the tip while a load balancer routes wallets to a dead port.
initial block download:ChainHub::in_ibd, the same callgetblockchaininfo.initialblockdownloadmakes.tip stale (last block Ns ago): the tip is older than--max-tip-age(default 24 h).
in_ibdlatches off after the first exit, so without thisa node that later loses every peer would stay ready on a stale tip.
tip N blocks behind headers(over 6).--sh-index:scripthash index N blocks behind tip(over 6).--metrics(confmetrics; refused without--health-listen) addsPrometheus text on the same listener. To avoid a second metrics dialect, each
gauge is a value the node already publishes, named after its source, and the
journey checks it against RPC at the same moment:
rbitcoin_blocks,_headers,_tip_time_seconds,_initial_block_download=getblockchaininforbitcoin_connections{direction="in"|"out"}=getnetworkinfoconnections_in/_out(both now count through onerbitcoin_net::connection_counts)rbitcoin_mempool_transactions,_bytes=getmempoolinfosize,bytesrbitcoin_scripthash_lag_blocks=tip: accept sh_lag=process_resident_memory_bytes=ibd: sizes rss=;process_start_time_secondsrbitcoin_build_info,rbitcoin_phase,rbitcoin_ready=/readyzCounters are the
tip: perfmeters: Esplora and Electrum requests andseconds, historical block serves and bytes, mempool accepts and rejects.
Those meters were sample-and-reset (
swap(0)), which a counter cannotread. They are now running totals (
PerfCounter), and the 5 s perf tick takesthe change since its previous sample, so the DEBUG line prints the same
numbers. The window max stays a reset (
PerfMax); it is not exported.Esplora and Electrum share one
RequestMeterin place of two copies of thesame three statics and CAS loop.
Cost (principle 9)
/healthz: no state read./readyz: a phase load; infollowingalsoin_ibd(an atomic oncelatched),
best_header_height(read lock over the header tips), tip heightand
sh_lag_heights. It runs on the blocking pool because the chain readscan touch the store.
/metrics: the above plus onePeerHub::snapshotand one fold over themempool for
bytes(the same fold asgetmempoolinfo), on the blockingpool. At a 15 s scrape interval that is noise next to the tip path.
atomic that only the 5 s tick touches.
ibd: perftimer.
Tests
Journeys (
rbitcoin-test --test cross_surface):esplora_broadcast_visible_in_rpc_and_electrumadds--health-listenand--metricsto the existing session. It checks/healthz200 / 405 / 404;/readyzis503 not ready: initial block downloadwhile RPC reportsinitialblockdownload: true, then200 okafter the mined block. Itscrapes
/metricsonce in IBD (mempool non-empty) and once after, andcompares every gauge above with RPC at that moment. The counters are
present, ≥ 1 where the journey used the surface, and do not go down.
readyz_names_a_listener_that_did_not_bind: RPC on a taken port →503 not ready: rpc not listeningwhile the node follows the tip, and/metricsis 404 without--metrics.health_listen_bind_failure_stops_before_store_open: a taken health port→
run_p2perrors andstore/does not exist. This pins that the listenerbinds before the store opens.
The binary smoke journey (
scenarios::node_cli_and_surface_smoke) refuses--metricswithout--health-listen(exit 1) and--health-listen bad(exit 2). The consolidated kebab-flag test checks the bare
--health-listendefault and
--metrics.Units only where a session cannot reach the case: the
readinesstable (everyphase, the lag boundaries 6 and 7, the tip-age boundary, SH off), the listener
answering
opening(a real store opens in milliseconds), andPerfCounter/PerfMax/RequestMeterwindows versus totals.Commits follow the steps:
/healthz,/readyz, gauges, counters, docs,then the stale-tip gate and the
/metricsmention in the non-loopback WARN.Docs:
docs/operator/operations.md(flags, a "Health probes and metrics"section with the reason strings, a k8s probe snippet, the metric table and a
scrape config), a
SECURITY.mdscope bullet,peer-clients.mdrank 4marked landed, and
changelog.d/health-endpoints.md.Open questions
--health-listensocket, for the reasons above,rather than routes on RPC or Esplora.
--health-listen/--metrics; bare port 9332.gate reuses
--max-tip-age. Should SH lag gate readiness at all?peer-clients.mdrow landed but did not add aquality.mdrow; say if you want one./healthzis "the process answers HTTP". Failing itwhen the main loop stops ticking would catch a wedged loop, but it would
have to stay off during
opening.health/metricsoptions yet; I can't runnixos-module-evallocally.Local results
Rebased onto
60e03731(after #800 and #802; the CLI and validate checksmoved into #802's smoke journey).
cargo fmt --check,cargo clippy --workspace --all-targets -D warnings,./scripts/ast-grep.shand
cargo deny checkpass.cargo test --workspace --no-fail-fastran on ahost at load average ~165, where only wall-clock tests failed:
integration_multinode(20 s budgets) andtip_accept::tests::drop_join_does_not_park_worker. With--test-threads=1,integration_multinodepasses 15 of 16, and the 16th(
p2p_compact_hb_getblocktxn_and_orphan, a wall timeout) passes alone;drop_join_does_not_park_workerpasses alone.cross_surface,scenariosand the node lib pass. Before the rebase the required CI jobs were green.
Side note, not changed here: on a heavily loaded host,
tip_accept::tests::drop_join_does_not_park_workercan miss its 2 s"job never started" deadline. It then panics without releasing the job it
queued. When that job runs later it blocks the shared tip-accept lane, and
every later lane test in that process hangs.
🤖 Generated with Claude Code