Skip to content

Several IpmiClient instances in one JVM against the same BMC: only one of four succeeds #97

Description

@bertysentry

Where: client layer + core/connection; root cause not fully isolated.

What happens: Running 4 IpmiClient.getChassisStatus() + getSensors() workers concurrently in one JVM against the same BMC (each worker creates its own IpmiConnector on an ephemeral port) gives 1 success and 3 TimeoutException (60 s). Two simultaneous workers: 0 of 2. Chassis status (one in-session message) succeeds for all workers; the losses start on the second in-session message of each SDR walk, at the same millisecond for both sessions.

What is BMC-side and what is ours:

  • Four simultaneous ipmiutil sessions (health, and FRU reads) all succeed in < 2 s – two of them visibly retried once, so this BMC does drop some packets under concurrent sessions.
  • Three walks in three separate JVMs: all succeed, with 1–2 "Message timed out" retries each.
  • Four walks in one JVM with a 3 s message timeout: 2 succeed, 2 exhaust the 3 retries.

So part of the problem is the 300 s timeout and the broken sync retry (separate issues), but in-JVM concurrency is measurably worse than separate processes, which points at shared state or the races listed in the data-race issue (UdpMessenger.send() is synchronized with a 1 ms sleep per packet, static ConnectionManager.reservedTags, etc.). A packet capture comparing the two setups would settle it. MetricsHub polls many hosts from one JVM, so this matters.

Activity

  1. bertysentry commented on Oct 9, 2026

    @bertysentry
    ContributorAuthor

    Re-measured on 2026-10-09 with an in-JVM harness (4 threads, each its own IpmiConnector, getChassisStatus() + getSensors() against the same BMC), on main (1abf019, which carries the transport fixes of #140 and #142) and on the lifecycle/thread-safety branches (#148 and the follow-up PR for #96/#89).

    GIGABYTE ME62 (the BMC of this report), 4 workers in one JVM:

    Build Run 1 Run 2
    main 4/4 OK (2.6 s each) 4/4 OK (3.9 s, one worker 7.5 s after a lost reply)
    lifecycle + thread-safety 4/4 OK (2.5 s each) 4/4 OK (2.5 s each)

    The 1-of-4 result of the report came from the 300 s message timeout and the broken sync retry: a dropped reply stalled the worker for the whole call. #140 (5 s per-message timeout, retry waits for the resent request) fixed that, so the in-JVM scenario now passes on main.

    Supermicro SYS (ipmi-hci-stg-01), where 2 of 4 workers still fail, in ~300 ms, with IPMIException: Inactive session ID.: the failure is the BMC's, not the library's.

    • 4 workers in one JVM: 2/4 OK (five runs, always exactly 2);
    • 2 workers in one JVM: 2/2 OK; 3 workers: 2/3 OK;
    • 4 separate JVMs started at the same time: 2/4 and 3/4 OK;
    • 4 simultaneous ipmiutil health -F lan2 -J 3 sessions: 2/4 then 3/4 OK, the others fail with Unable to establish IPMI v2 / RMCP+ session; 2 simultaneous: 2/2 OK.

    That BMC accepts about two RMCP+ sessions being opened at the same time; the extra ones are evicted. The Thread safety section of the documentation already says that parallel sessions against the same BMC are not reliable.

    No packet capture was needed: the same client in separate processes and an independent client reproduce the Supermicro limit, and the GIGABYTE scenario passes.

  2. bertysentry commented on Oct 9, 2026

    @bertysentry
    ContributorAuthor

    Closing: the reported scenario passes since #140, and the remaining Supermicro failure is a BMC-side session limit (see the measurements above).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions