Where: client layer + core/connection; root cause not fully isolated.
What happens: Running 4 IpmiClient.getChassisStatus() + getSensors() workers concurrently in one JVM against the same BMC (each worker creates its own IpmiConnector on an ephemeral port) gives 1 success and 3 TimeoutException (60 s). Two simultaneous workers: 0 of 2. Chassis status (one in-session message) succeeds for all workers; the losses start on the second in-session message of each SDR walk, at the same millisecond for both sessions.
What is BMC-side and what is ours:
- Four simultaneous
ipmiutil sessions (health, and FRU reads) all succeed in < 2 s – two of them visibly retried once, so this BMC does drop some packets under concurrent sessions.
- Three walks in three separate JVMs: all succeed, with 1–2 "Message timed out" retries each.
- Four walks in one JVM with a 3 s message timeout: 2 succeed, 2 exhaust the 3 retries.
So part of the problem is the 300 s timeout and the broken sync retry (separate issues), but in-JVM concurrency is measurably worse than separate processes, which points at shared state or the races listed in the data-race issue (UdpMessenger.send() is synchronized with a 1 ms sleep per packet, static ConnectionManager.reservedTags, etc.). A packet capture comparing the two setups would settle it. MetricsHub polls many hosts from one JVM, so this matters.
Where: client layer +
core/connection; root cause not fully isolated.What happens: Running 4
IpmiClient.getChassisStatus()+getSensors()workers concurrently in one JVM against the same BMC (each worker creates its ownIpmiConnectoron an ephemeral port) gives 1 success and 3TimeoutException(60 s). Two simultaneous workers: 0 of 2. Chassis status (one in-session message) succeeds for all workers; the losses start on the second in-session message of each SDR walk, at the same millisecond for both sessions.What is BMC-side and what is ours:
ipmiutilsessions (health, and FRU reads) all succeed in < 2 s – two of them visibly retried once, so this BMC does drop some packets under concurrent sessions.So part of the problem is the 300 s timeout and the broken sync retry (separate issues), but in-JVM concurrency is measurably worse than separate processes, which points at shared state or the races listed in the data-race issue (
UdpMessenger.send()issynchronizedwith a 1 ms sleep per packet, staticConnectionManager.reservedTags, etc.). A packet capture comparing the two setups would settle it. MetricsHub polls many hosts from one JVM, so this matters.