Skip to content
rsasaki0109Public

About

Check whether the sensor calibration on your robot is still right: one command, honest per-axis verdicts for LiDAR, IMU, camera, GNSS, and vehicle.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

27 stars

Watchers

0 watching

Forks

Calibrex

Estimate it, check it, watch it drift.
The sensor calibration of your robot, from ROS 2 and ROS 1 bags, with honest per-axis verdicts for LiDAR, IMU, camera, GNSS, and vehicle.

CI GitHub Release Python License Status

calibrex check catches a bad calibration and confirms the good one on KITTI: with a 3 degree yaw error in the deployed LiDAR transform the LiDAR points miss the bollards and lidar-vehicle fails; restoring the vendor calibration file puts them back on the bollards and the re-check passes

Real calibrex check runs on KITTI development drives: a deployed LiDAR transform with a 3° yaw error fails; restoring the calibration file and re-checking passes (pitch and yaw judged; roll is not observable from driving).

Two Ouster LiDARs on NTU VIRAL: the second LiDAR's scans move from a perturbed extrinsic to the calibrex estimate and snap onto the first LiDAR's map; the median point-to-plane residual drops from 19.1 cm to 4.4 cm A hand-held Livox MID360 sweep at 180 deg/s: the raw 0.1 s sweep is bent by the rotation, and gyro deskewing with the calibrex IMU-LiDAR estimate straightens the ground and the poles
LiDAR ↔ LiDAR snaps into focus (NTU VIRAL tnp_01). From a labelled perturbed start to the calibrex lidar-lidar estimate; median point-to-plane residual 19.1 → 4.4 cm (the design value gives 4.6 cm). Gyro deskew (RTK-SLAM, hand-held MID360). The fastest sweep of the run, 180 deg/s: raw vs deskewed with Calibrex's own IMU-LiDAR estimate; local surface thickness 3.42 → 2.19 cm.
Hilti 2022 exp21: gyro rates rotated into the camera frame and camera rates from tracked features start misaligned and lock together as rotation, time offset and gyro bias move to the calibrex estimate; the final frame compares with Kalibr The KITTI rig frame tree from /tf_static: the velo_link axes rotate with the injected yaw (drawn eight times larger) and turn red as calibrex check fails lidar-vehicle
Camera ↔ IMU curves lock together (Hilti 2022 exp21, development data). Rate RMSE 15.8 → 6.2 deg/s (the floor is camera-rate noise); the estimate is 0.45° and −0.21 ms from Kalibr's target-based calibration. The rig under an injected yaw error (KITTI /tf_static). Rotation drawn ×8 for visibility; colours follow the real calibrex check verdict at each step.

calibrex check --pairs camera-lidar on KITTI drive 0005: the LiDAR points projected through the vendor camera-LiDAR calibration land on the cyclist and the bollards and the pair passes; turning the camera tf 3 degrees about its y axis shifts them off the objects and the pair fails on pitch (3.00 degrees against a 0.50 degree tolerance); x y z are never judged

Camera ↔ LiDAR, rotation only (KITTI development drive 0005). Two real calibrex check --pairs camera-lidar runs: the vendor calibration passes (roll and pitch judged), the same calibration turned +3° about the camera y axis fails (pitch 3.00° against 0.50°). The frames in between only slide the overlay; they are not check results.

Every frame is rendered from a real recording by Calibrex's own code; generators, input digests and licenses are in readme-data-gifs.json.

The workflow

Terminal recording: calibrex estimate on a bag with no calibration (roll, pitch and yaw observed, the lever arm marked NOT MEASURED), calibrex check of the exported frames.yaml on another recording (pass on the three judged axes, x y z unchecked), and calibrex drift over three recordings flagging the one whose IMU was remounted

Real runs on the Koide hard-localization recordings (abridged output): estimate a calibration from one bag, check it on another recording, then ask whether it changed across recordings. The third recording is a copy with its IMU rotated 2° about z. The full walkthrough.

python -m pip install "https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"

calibrex estimate my_bag/ --output est/ --html est.html    # 1. no calibration yet: measure it, export frames.yaml / static transforms / URDF
calibrex check other_bag/ --tf est/frames.yaml --html check.html   # 2. is the deployed calibration still right on another recording?
calibrex drift day1/ day2/ day3/ --output drift/ --html drift.html # 3. did it change over time, and in which bag?
Step Command You get
1 calibrex estimate <bag> per pair each axis with its std and whether the data observed it; frames.yaml (loads with --tf), static_transform_publisher commands, URDF joints (and a Kalibr camchain); axes the data did not observe are marked NOT MEASURED, never written as measured; --html report
2 calibrex check <bag2> --tf frames.yaml pass / warn / fail / inconclusive per axis, the axes it could not judge, the error the data could have detected, next steps, progress while it runs; --html report. Without --tf it reads the bag's /tf_static
3 calibrex drift <bag1> <bag2> ... per axis stable / drift / inconclusive across recordings of one rig; the deviating bag and the size of the change

What it covers

Pair check judges estimate / drift
IMU ↔ LiDAR rotation, clock offset, lever arm when observed yes
camera ↔ IMU rotation, clock offset (needs intrinsics) yes
camera ↔ LiDAR rotation only, targetless edge alignment; translation is never judged needs a rough --tf prior; rotation only
LiDAR ↔ LiDAR rotation and translation needs a rough --tf prior
GNSS ↔ LiDAR, GNSS ↔ IMU antenna lever arm yes
LiDAR / IMU / INS / wheels ↔ vehicle the rotation axes the motion excites (roll is unobservable on planar driving); opt-in with --vehicle-frame yes
camera focal length deployed focal lengths check only

Every row reports per axis whether the data could judge it. camera-LiDAR is development evidence, not a claim: on KITTI and Hilti recordings the unmodified reference never warned or failed and 23 of 27 perturbed candidates were flagged (results).

Your bag format

check, estimate and drift read the bag as it is, with no conversion and no ROS installation; the format is detected from the file:

Your bag Notes
rosbag2 directory or .db3 (sqlite3) per-message zstd / lz4 needs calibrex[rosbag2-compression]
ROS 2 MCAP (.mcap) chunks none, zstd or lz4
ROS 1 .bag chunks none, bz2 or lz4 (calibrex[rosbag1-lz4] for lz4)

calibrex check my_bag/ --plan prints which sensor pairs the bag can be checked for, and why the others are skipped, without running an estimator. Vehicle pairs are opt-in with --vehicle-frame base_link. Details: supported input formats.

The ROS 1 and MCAP-lz4 readers, camera-LiDAR, the Autoware Velodyne decoder and the faster first runs are on main and arrive with the release after v0.5.1. To use them now, install from the repository: python -m pip install "git+https://github.com/rsasaki0109/Calibrex.git".

Autoware bags. Autoware's all-sensors sample records raw velodyne_msgs/VelodyneScan packets and, as a point cloud, only one concatenated cloud in base_link. Calibrex decodes the packets (VLP-16 and VLP-32C, single return) with tools/velodyne_scan_to_pointcloud2.py (in the repository), which writes a derived bag with one PointCloud2 per sensor, and reads autoware_auto_vehicle_msgs/VelocityReport as wheel speed. On that bag (36.5 s, a straight drive) lidar-lidar gives a real pass for the deployed /tf_static and fail for +1 / +3 deg and +5 cm perturbations; the IMU and GNSS pairs stay inconclusive because a straight drive excites no rotation, and the output says how much recording and turning would. See the Autoware section.

In your browser, no install: the bag check page runs calibrex check and calibrex estimate on your own rosbag2 (plan which pairs can be checked, then run them, with progress, verdicts and frames.yaml downloads). Multi-GB bags are read lazily; it runs locally under Pyodide, so no data is uploaded. It was compared with the command line on real bags (KITTI, Koide, NTU VIRAL, RTK-SLAM, Hilti; 30-70 s slices, bags up to 22 GB): same verdicts, numbers equal to 1e-6 or better, except the camera pairs, whose estimates differ within one to two sigma because Pyodide ships an older OpenCV. The page also takes --topic-kind and --camera. The browser calibration page calibrates an IMU against a sensor trajectory (TUM): rotation, clock offset, gyro bias, and lever arm, with held-out evidence.

When an axis is not observable, it says why and what would fix it. For a lever-arm (translation) axis calibrex check and calibrex estimate name the rotation the recording lacks, or how much more of the same motion would reach the bound; see what makes a lever arm observable. calibrex --help groups the commands: Start here (check, estimate, drift, doctor, demo, ...), per-pair calibration, evidence and CI, and the rest.

🔍 Three commands
estimate, check, drift: reads the bag, finds the sensors, and covers up to ten pairs (GNSS, IMU, LiDAR, camera, vehicle).
⚖️ Honest verdicts
Per-axis pass / warn / fail / inconclusive, with partial coverage and detection power reported.
🧪 Pre-registered SOTA
Thresholds are frozen before held-out data is scored; refuted audits stay on the leaderboard.

Docs · Workflow · Estimate · Check · Drift · SOTA audits · Coverage · Gallery · Quickstart · Methods

Estimate a calibration when there is none

calibrex check judges a candidate calibration. With a bag but none, every pair is skipped (no_candidate_calibration); calibrex estimate is the one-command path to one, and check then verifies it on a different recording:

calibrex estimate my_bag/ --output est/                     # add --tf rough.yaml for a prior
calibrex check other_bag/ --tf est/frames.yaml              # verify on a different recording
calibrex estimate my_bag/ --output est/ --html est.html     # also a self-contained HTML report

--html (or calibrex render est/bag_estimate.json --format html later) writes a single offline page: per pair each axis as value, 1 sigma and observability (axes the data did not observe are greyed and marked "NOT MEASURED"), the exported and omitted frames, the exported files with their digests, and copy buttons for the calibrex check --tf and static_transform_publisher commands.

The calibrex estimate HTML report for Koide indoor_easy_01: imu-lidar partial, roll pitch and yaw observed with standard deviation bars, x y z greyed as unobservable and from prior, NOT MEASURED, and the exported frame tree

It runs the same native estimators and keeps the estimates (est/bag_estimate.json, slac.bag_estimate/v0.1): per pair the transform, each axis with its standard deviation, and whether the data observed it (observed, unobservable, control_not_detected, not_estimated), with the bag digest and settings as provenance. From the observed frames it writes frames.yaml (loads directly with --tf), ROS 2 static_transform_publisher commands and a launch file, and URDF joints (plus a Kalibr camchain-imucam.yaml when --tf gave a camchain). A frame with an axis the data did not observe is never written as if measured: it is written only when a rough --tf prior supplies that axis (the YAML marks it NOT MEASURED), otherwise it is omitted and listed under next steps. Rotation-only estimators (camera-IMU, the vehicle pairs) therefore need a rough lever arm from --tf; lidar-lidar registration and camera-lidar edge alignment (rotation only) also start from a --tf prior, and camera-IMU and camera-LiDAR need the camera intrinsics (a CameraInfo topic or a Kalibr camchain).

Real-data results, with the splits, exact commands and the failures (Hilti camera-IMU, Koide imu-lidar, KITTI vehicle pairs, RTK-SLAM GNSS pairs): calibrex estimate on real data.

Check a deployed calibration

calibrex check my_bag/ --output check.json --html check.html     # reads /tf_static; add --tf rig.urdf to override

calibrex check asks whether the extrinsics deployed on a robot agree with what a recording says (step 2 of the workflow; no calibration yet? estimate one first). It reads the candidate transforms from the bag's /tf_static (or --tf: the frames.yaml of calibrex estimate, URDF, Kalibr, RTK-SLAM or Hilti calibration files), detects the sensor topics, works out which of the ten sensor pairs the bag can check, runs the native estimator of each, and judges the deployed transform per axis against the estimate and its uncertainty: pass, warn, fail or inconclusive. It also checks the estimates against each other (rig closure) and reports which axes it could not judge. Vehicle pairs (LiDAR, IMU or wheel odometry against the vehicle frame) are opt-in with --vehicle-frame base_link.

calibrex check on pooled KITTI development drives: the vendor tf is inconclusive, a +1 degree yaw error injected into velo_link makes lidar-vehicle and lidar-wheel_odometry fail (yaw 1.31 deg against a 0.5 deg tolerance)

A yaw error swept from -3 to +3 degrees into the deployed velo_link transform with the real calibrex check verdict at each step

A sweep of injected errors. Not a calibration running: a yaw error from −3° to +3° is injected into the deployed velo_link transform and each step is one real calibrex check run. Errors of about 1° or more turn lidar-vehicle to fail. Even at 0° the vendor transform is about 0.3° from the estimate, and the overall verdict stays inconclusive because imu-vehicle cannot judge on KITTI.

Demo: the KITTI raw development drives (0005, 0009, 0014, 0015, 0022) pooled into one bag, with velo_link of the deployed tf turned about its parent's z axis (script, first run about 6 minutes, the other variants seconds):

deployed tf overall lidar-vehicle lidar-wheel_odometry ins-lidar imu-vehicle
vendor (KITTI calibration) inconclusive pass, partial pass, partial pass, partial inconclusive
velo_link yaw +1 deg fail fail (yaw 1.31 deg vs 0.5) fail (yaw 1.31 deg vs 0.5) pass, partial inconclusive
velo_link yaw +3 deg fail fail (yaw 3.31 deg vs 0.5) fail (yaw 3.31 deg vs 0.5) pass, partial inconclusive
pair                  verdict                           |delta|/tolerance                                    unchecked
lidar-vehicle         fail (partial: pitch, yaw only)   pitch 0.505/0.5 deg [warn]; yaw 1.31/0.5 deg [fail]  roll
imu-vehicle           inconclusive                      -                                                    roll, pitch, yaw
ins-lidar             pass (partial: roll, pitch only)  roll 0.114/0.5 deg; pitch 0.0567/0.5 deg             yaw, x, y, z
lidar-wheel_odometry  fail (partial: pitch, yaw only)   pitch 0.505/0.5 deg [warn]; yaw 1.31/0.5 deg [fail]  roll
overall verdict: fail                                    (velo_link yaw +1 deg)

Read it honestly: every vehicle pair has partial coverage (roll is not observable from planar driving, and ins-lidar cannot see yaw), so a pass covers only the judged axes. The vendor run is inconclusive, not pass, because the INS velocities are too noisy for imu-vehicle to judge anything on KITTI; and the vendor pitch (0.496 deg) sits right at the 0.5 deg tolerance floor. The two LiDAR-motion pairs read the same odometry, so they agree by construction. Only the 1 and 3 degree yaw errors are injected here; the numbers are the development drives, not a held-out claim.

Tutorial · demo summary · HTML report of the +1 deg run (download and open; GitHub shows HTML as source).

Did the rig change between recordings? calibrex drift

Two or more bags of the same rig, in recording order:

calibrex drift day1/ day2/ day3/ --pairs imu-lidar --output drift/ --html drift.html

Each bag is estimated (the estimator cache makes a re-run cheap; drift/<bag>/bag_estimate.json) and every axis the data observed in at least two bags is tested for consistency: pairwise differences against max(3 sigma, floor) (the minimum detectable change of the axis type) and a chi-square homogeneity test. Per pair the verdict is stable, drift (with the deviating bag named when three or more bags allow it, and the size of the change) or inconclusive; exit status 1 on drift (--fail-on). Result: drift/calibration_drift.json, slac.calibration_drift/v0.1. A stable verdict means no change larger than the reported minimum detectable change was found. Per-pair clock offsets are compared too (see the drift page).

Pre-registered SOTA audits

Calibrex may call a method state of the art only for a claim that a frozen audit marks supported. The claim, scoring, held-out data, and thresholds are committed before the held-out data is scored, each piece of evidence is a digest-pinned artifact, and refuted audits stay on the SOTA leaderboard. Standings as of 2026-10-02:

Pair Standing Scope of the audited claim Details
imu-lidar supported 2, refuted 1 Livox MID360 against LI-Init (GPL, run in a container): rotation and clock offset (4/4 gates), and the full extrinsic in a second round (4/4). The first full-extrinsic round was refuted (0/8) and stays listed. MID360 IMU-LiDAR
lidar-vehicle supported 1 Motion-only LiDAR-to-vehicle rotation against KITTI's calib_imu_to_velo on eight unseen drives (3/3 gates). Pitch and yaw only; roll stays unobservable. KITTI LiDAR-vehicle
lidar-lidar refuted NTU VIRAL (two Ouster OS1-16): refuted 2/4. Accuracy gates pass, but the method does not beat scan-to-scan and cross-recording consistency fails (1.143, bound 1.0). NTU VIRAL LiDAR-LiDAR
camera-imu refuted Hilti 2022 forward cameras (cam0, cam1) against Kalibr on four unseen recordings: refuted 6/7. Rotation (max 0.70 deg), clock offset (0.33 ms), consistency (0.43 deg) and the cam0-cam1 relative rotation (0.41 deg) pass, but the shortest recording (exp04) leaves an axis unconstrained on both cameras. Hilti camera-IMU
the other six target pairs no_claim Native methods exist, but no audited claim yet. Leaderboard

A supported standing is scoped to one dataset family and one comparison; it is not a general accuracy claim. Held-out data is reserved per pair, and a spent split cannot be reused for a second claim. The machine-readable standings are in sota_leaderboard.json.

calibrex sota audit protocol.yaml --output audit.yaml   # exit 0 only for supported

Calibration coverage

All ten target sensor pairs have a native method. "Audit standing" is the standing on the leaderboard.

Pair Native command Public data used Audit standing
camera (focal lengths) calibrex camera-imu focal Hilti 2022 no_claim
camera ↔ IMU calibrex camera-imu rotation Hilti 2022 (dev exp21, exp07; audit exp01-exp04) refuted
camera ↔ LiDAR calibrex check / estimate / drift (targetless edge alignment, rotation only; results), calibrex camera-lidar ..., calibrex demo kitti-lidar-camera-evidence KITTI dev drives and Hilti 2022 (check), KITTI-shaped fixture, A2D2, ACFR no_claim
GNSS ↔ IMU calibrex gnss-imu compose (from calibrex gnss-lidar rtk-slam and IMU-LiDAR) RTK-SLAM no_claim
IMU ↔ LiDAR calibrex imu-lidar livox, livox-translation, trajectory RTK-SLAM, Zenodo MID360 driving supported (2) / refuted (1)
IMU ↔ vehicle calibrex imu-vehicle kitti KITTI raw no_claim
INS ↔ LiDAR calibrex ins-lidar kitti KITTI raw no_claim
LiDAR ↔ LiDAR calibrex lidar-lidar ros2 NTU VIRAL refuted
LiDAR ↔ vehicle calibrex lidar-vehicle kitti KITTI raw supported (1)
LiDAR ↔ wheel odometry calibrex lidar-wheel trajectory, kitti KITTI raw (OXTS stands in for wheels) no_claim

Also native or adapter-backed: hand-eye AX=XB and robot-world AX=YB (13 native methods, benchmarked below against OpenCV), radar extrinsics, RGB-D joint SLAC, and solid-state LiDAR clock-offset profiling, mostly through calibrex calibrate <config>. Open3D, Kalibr, Koide, ROS, and Autoware stay behind adapters. See calibration methods for solver-level detail, and the per-pair pages: IMU-LiDAR, LiDAR-vehicle, LiDAR-lidar, camera-IMU, INS-LiDAR, LiDAR-wheel, GNSS-LiDAR, GNSS-IMU.

Calibrex LiDAR calibration coverage map

Public-data gallery

A2D2 real camera and LiDAR projection overlay
Real A2D2 camera × LiDAR: front-left camera frames with real camera-view LiDAR returns. A2D2 distributes these points pre-registered into the camera view, so this is visual evidence—not independent extrinsic accuracy.
Livox public solid-state LiDAR calibration evidence A2D2 public front multi-LiDAR calibration evidence A2D2 public front-rear LiDAR calibration evidence
Livox Horizon ↔ Horizon
Known-bad controls and holdout point-to-plane evidence.
A2D2 front pair
Fixed-rig metadata and support accounting.
A2D2 front ↔ rear
A longer baseline with different overlap behavior.
Calibrex simultaneous localization and calibration on TIERS Indoor02 real moving-platform data TIERS LidarsCali real online solid-state LiDAR calibration Livox real point cloud calibration before and after refinement
TIERS moving platform
A Velodyne VLP-16 motion map supports online Ouster OS1 calibration, with 106 of 108 batches accepted by holdout gates.
TIERS LidarsCali online
Real Livox Horizon ↔ Avia batches replayed through the online gate.
Livox before → after
Real Horizon PCD returns through the native registration refinement replay; the public pair has no transform ground truth.
TIERS real solid-state LiDAR time offset sweep TIERS time-offset sweep
Real VLP-16 ↔ Livox Horizon candidate probes. The plot reports the algorithmic +40 ms estimate and re-optimized train/holdout RMSE; the public sequence has no independent clock ground truth.

Every gallery asset is generated from public raw data. Dataset source, protocol, parameters, and digests are recorded in readme-gif-gallery.json.

Five-minute quickstart

Run the complete evidence path without ROS or a dataset download. The command evaluates the bundled deterministic KITTI-shaped LiDAR-camera fixture, writes a schema-valid result and review report, and verifies a digest-bound provenance bundle:

python -m pip install \
  "https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"

calibrex demo kitti-lidar-camera-evidence \
  --output-dir outputs/kitti-lidar-camera-evidence \
  --strict-assessment
calibrex validate outputs/kitti-lidar-camera-evidence/result.yaml
calibrex verify outputs/kitti-lidar-camera-evidence/bundle.json

Open outputs/kitti-lidar-camera-evidence/report.html to inspect the projection evidence and known-bad perturbation probes. The checked fixture detects 16 of 24 mandatory perturbation cases and passes all six falsification-policy gates in under one minute in the clean Windows wheel smoke test. This verifies the pipeline and evidence contracts; it is not a real-sensor accuracy claim or a standalone camera-LiDAR calibration algorithm. Supply an officially downloaded KITTI raw sequence with --dataset-path when evaluating real data.

To render and validate a committed result from a source checkout:

git clone https://github.com/rsasaki0109/Calibrex.git
cd Calibrex
python -m pip install .

calibrex validate examples/precomputed/result.yaml --json
calibrex render examples/precomputed/result.yaml \
  --format evidence-card \
  --output outputs/quickstart/evidence-card.svg
calibrex render examples/precomputed/result.yaml \
  --output-dir outputs/quickstart

Before configuring a solve, diagnose a recording and save the result for review or CI:

calibrex doctor recording.mcap \
  --output outputs/doctor.json \
  --json
calibrex validate outputs/doctor.json

doctor infers supported dataset types, reports missing optional dependencies, data coverage and degeneracy warnings, and suggests compatible evidence workflows. Running it without a path retains the lightweight environment check.

Run the same evidence gates on every calibration change:

- uses: actions/checkout@v4
- uses: rsasaki0109/Calibrex@v0.5.1
  with:
    candidate: calibration/candidate.yaml
    baseline: calibration/baseline.yaml

The action writes a GitHub Step Summary, fails on FAIL or INCONCLUSIVE by default, and exposes schema-valid evidence, comparison, SVG, and calibration-ci.json artifacts. See Calibration CI.

Tried it on your rig? Share a sanitized result or a useful failure case in Discussions. If the evidence-first workflow earns a place in your calibration stack, consider starring Calibrex so other robotics teams can find it.

Public real-data benchmarks

Calibrex evaluates all 13 native hand-eye methods and seven OpenCV 4 variants on the ETHZ ASL real robot-arm dataset. Every method sees the same five digest-locked absolute-pose splits. AX=XB and AX=YB are separate equation families and are never ranked together.

Hand-eye AX=XB

Method Holdout rotation RMSE deg ↓ Holdout translation RMSE mm ↓ Failure rate Runtime s
Calibrex Park-Martin 0.867227 [0.759962, 0.958155] 13.8759 [11.7678, 15.8266] 0.0% 0.153451
Calibrex Tsai-Lenz 0.875291 [0.77004, 0.967317] 13.8652 [11.7326, 15.8293] 0.0% 0.126921
Calibrex Daniilidis 0.868969 [0.761845, 0.960775] 13.8212 [11.9121, 15.6887] 0.0% 0.322403
Calibrex Andreff 0.867199 [0.759958, 0.958223] 13.879 [11.7674, 15.8309] 0.0% 0.334728
Calibrex Shiu-Ahmad 1.6191 [0.787831, 3.10277] 31.9219 [12.7334, 52.761] 0.0% 14.1537
Calibrex Chou-Kamel 0.867205 [0.759953, 0.958241] 13.8795 [11.7684, 15.8315] 0.0% 0.129661
Calibrex Horaud-Dornaika 0.890804 [0.782116, 0.996268] 14.1434 [11.8641, 16.1088] 0.0% 0.139859
Calibrex H-D nonlinear 0.885789 [0.777979, 0.991045] 14.0634 [11.8408, 16.0027] 0.0% 19.4512
OpenCV Tsai 0.871872 [0.769051, 0.961725] 13.8681 [11.8255, 15.6992] 0.0% 0.0368567
OpenCV Park 0.867225 [0.759969, 0.958144] 13.8754 [11.7667, 15.8261] 0.0% 0.0280751
OpenCV Horaud 0.867203 [0.759961, 0.958229] 13.879 [11.7674, 15.831] 0.0% 0.0248639
OpenCV Andreff 0.868623 [0.763653, 0.959819] 15.4855 [13.6068, 17.338] 0.0% 0.0327896
OpenCV Daniilidis 0.871005 [0.764385, 0.963618] 13.9265 [11.9635, 15.6961] 0.0% 0.0284915

Robot-world hand-eye AX=YB

Method Holdout rotation RMSE deg ↓ Holdout translation RMSE mm ↓ Failure rate Runtime s
Calibrex Shah 0.624739 [0.559308, 0.688861] 10.7793 [9.53397, 11.9952] 0.0% 0.0206115
Calibrex Li-Wang-Wu 0.627131 [0.561467, 0.692796] 19.651 [15.2668, 23.3423] 0.0% 0.0321106
Calibrex Dornaika-Horaud 0.624742 [0.559311, 0.688865] 10.7793 [9.53399, 11.9952] 0.0% 0.0286939
Calibrex Zhuang-Roth-Sudhakar 0.651861 [0.569557, 0.733996] 10.961 [9.65144, 12.1956] 0.0% 0.0213057
Calibrex D-H nonlinear 0.6265 [0.560987, 0.692012] 10.7752 [9.52577, 12.0161] 0.0% 0.217955
OpenCV Shah 0.624739 [0.559308, 0.688861] 10.7793 [9.53397, 11.9952] 0.0% 0.00343018
OpenCV Li 0.627131 [0.561467, 0.692796] 19.651 [15.2668, 23.3423] 0.0% 0.00570322

These are scoped consistency results, not blanket accuracy claims: the archive does not provide an accepted ground-truth extrinsic. The tables report closure RMSE on untouched holdout poses, retain failures in the denominator, and show 95% bootstrap intervals across splits. See the full protocol, citations, limitations, and reproducible provenance.

Solid-state LiDAR quickstart

Start with the small Livox sample for a fast, no-ROS check. It writes a calibrated result.yaml, report.html, evidence sidecars, and a verified provenance bundle:

python -m pip install .
calibrex demo livox-evidence \
  --output-dir outputs/solid-state-livox-demo \
  --json
calibrex validate outputs/solid-state-livox-demo/result.yaml
calibrex verify outputs/solid-state-livox-demo/bundle.json
  • Real timing and clock offset. The public TIERS VLP-16 ↔ Livox Horizon bag (about 7.18 GB, never downloaded automatically) is profiled with calibrex continuous-time-lidar-pair; see the public-dataset tutorial. On the checked run the selected offset was +40 ms and train RMSE changed from 0.0957 m to 0.0240 m, with 0.0242 m holdout RMSE. This is an algorithmic estimate under the declared holdout: the public sequence has no independent clock ground truth.
  • Synthetic truth gate. A deterministic benchmark recovers a known extrinsic and +30 ms clock offset and rejects a fixed-clock known-bad control (tools/run_solid_state_synthetic_benchmark.py). It verifies solver mechanics, not real-sensor accuracy. See the YAML artifact and report.
  • Public cross-dataset benchmark. Paired solver variants are compared on the same capture windows, temporal holdouts, and sampling seeds. In the v0.3 matrix, adaptive wins 23/27 replicates with mean improvement 47.70% (bootstrap 95% CI [34.43, 59.99]%); AgRob Modular-e stays a visible counterexample (5/9, mean −1.34%). The public-only v0.4 candidate scores 27/27 replicates: adaptive wins 25/27, mean improvement 57.20% (95% CI [45.69, 67.43]%), and 9 of 54 variant artifacts hit the declared max_iterations category. See the v0.3 report, v0.4 report, and train-only selection note. These are ground-truth-free temporal-holdout results, not an absolute accuracy or SOTA claim.

Full commands for every step are in the solid-state public benchmark runbook. The independent physical-metrology packet remains an optional future path and is checked in as a planned template, not a result: YAML and report.

What is implemented in the current alpha?
  • typed config, result, comparison, protocol, policy, and evidence schemas
  • native solvers for camera-IMU, IMU-LiDAR, GNSS-LiDAR/IMU, LiDAR-LiDAR, LiDAR/IMU-vehicle, LiDAR-wheel odometry, and INS-LiDAR, each with held-out windows, a block jackknife, and known-bad controls
  • pair-agnostic pre-registered SOTA audits and per-pair standings (calibrex sota)
  • offline and online/streaming LiDAR calibration for rosbag1, rosbag2, and MCAP
  • motion compensation, per-point deskew, and trajectory evidence
  • targetless camera-LiDAR mutual information and online monitoring
  • native point-to-point, point-to-plane, hand-eye, and robot-world baselines
  • backend-neutral joint SLAC with typed Schur pose elimination
  • radar velocity, LiDAR-IMU rotation, and temporal-offset evidence
  • typed external-run artifacts, a Kalibr camchain importer, and Koide/Open3D, ROS, and Autoware adapter boundaries
  • a frozen, SHA-bound full-scale KITTI reference-vs-known-bad falsification runner

See the calibration methods and changelog for solver-level detail and limitations.

Real data, honest verdicts

Calibrex does not turn every run green. Weak excitation, failed controls, and refuted audits are reported as evidence, not hidden as demo noise. A sample of current results from the benchmark pages:

Pair / data Result Verdict
lidar-lidar, NTU VIRAL held-out audit Accuracy gates pass (0.489 deg, 0.077 m to design), but no gain over scan-to-scan (paired CI low −0.0048) and cross-recording consistency 1.143 vs bound 1.0; 2/4 gates ❌ Refuted
imu-lidar, first full-extrinsic audit vs LI-Init Lever-arm x unobservable on both recordings, so the claim was contradicted; kept on the leaderboard beside the later supported round ❌ Refuted
camera-imu, Hilti 2022 dev (exp21, exp07), 10 camera runs 5 pass, 1 warn, 4 inconclusive; side and down cameras on exp07 still do not constrain rotation (jackknife 0.35-0.7 deg) ⚠️ Inconclusive
ins-lidar, KITTI drives 0005 + 0009 Roll, pitch, and clock offset estimated; yaw and translations unobservable ⚠️ Inconclusive
lidar-wheel, KITTI (OXTS stand-in) Rotation matches lidar-vehicle within 0.004 deg; a 1 % speed-scale control is not detected ⚠️ Warn
imu-lidar, MID360 driving Roll estimated; yaw unobservable because a vehicle rotates almost only about the vertical axis ⚠️ Inconclusive
imu-lidar, MID360 hand-held (RTK-SLAM seq2) Every rotation axis estimated; held-out windows detect every known-bad shift ✅ Pass

The failure is part of the product: gates refuse to certify what the available data cannot falsify.

Why evidence?

Most calibration tools stop after producing a transform. Calibrex asks the next question: what evidence would falsify this transform?

Calibrex is an open-source, ROS-independent Python toolkit that turns candidate extrinsics, time offsets, and trajectories into evidence backed by holdout metrics, known-bad controls, observability checks, and reproducible provenance. It ships native solvers for every sensor pair listed below, plus adapters for Kalibr, Open3D, and Autoware, and it reports pass, warn, fail, or inconclusive instead of optimizer convergence or a single training residual.

A matrix gives you Calibrex adds
One estimated transform Candidate, reference, and selected estimates kept separate
One training residual Train/holdout metrics and temporal stability
Optimizer convergence Known-bad perturbation challenges
A covariance matrix Rank, weak directions, and degeneracy warnings
A screenshot Schema-valid artifacts with input and producer provenance
A result file A digest-locked evidence bundle that can be verified later
flowchart LR
    A[Sensor data] --> B[Candidate calibration]
    B --> C{Evidence gates}
    C -->|Holdout| D[Generalization]
    C -->|Known-bad| E[Falsification]
    C -->|Observability| F[Weak directions]
    D --> G[PASS / WARN / FAIL]
    E --> G
    F --> G
    G --> H[HTML + schema-valid sidecars]
Loading

See the evidence at a glance

The card below is not a hand-authored mockup. It is rendered from a schema-valid result.yaml, bound to the source file by SHA-256, and committed with machine-readable provenance inside the SVG.

Calibrex evidence card showing holdout metrics, observability, schema status, and source provenance

Generate the same artifact from any Calibrex result:

calibrex render result.yaml \
  --format evidence-card \
  --output evidence-card.svg

Compare candidates without hiding incompatibilities

This generated table compares the validated quickstart result with a declared 5° yaw / 0.15 m known-bad control. The amber protocol state is intentional: these compact example results do not declare a shared evidence protocol, so Calibrex shows the metric deltas but refuses to call the comparison compatible.

Calibrex comparison table contrasting an accepted result and a known-bad candidate, with protocol compatibility and source digests

calibrex compare \
  examples/precomputed/result.yaml \
  examples/precomputed/known_bad_result.yaml \
  --format evidence-table \
  --left-label accepted \
  --right-label "known bad" \
  --output comparison-table.svg

Outputs you can inspect and verify

result.yaml                  calibrated values + run provenance
├── report.html              portable human-readable report
├── metrics.json             train / holdout measurements
├── evidence.json            protocols, summaries, and known-bad cases
├── assessment.json          PASS / WARN / FAIL policy decision
├── observability.json       rank and weak directions
├── degeneracy.json          failure modes and limitations
├── protocol.json            declared evaluation contract
├── policy.json              falsification thresholds
├── transforms.json          typed transform provenance
├── bundle.json              artifact SHA-256 manifest
└── verification.json        materialized bundle verification
calibrex evidence result.yaml --output evidence.json
calibrex assess evidence.json
calibrex verify bundle.json
calibrex compare reference.yaml candidate.yaml --enforce-compatible

render never claims to recompute metrics. Cached inputs remain marked as cached, raw-input verification remains visible, and incompatible protocols are not silently ranked together.

Install

The supported no-source install is the versioned wheel attached to the v0.5.1 GitHub Release:

python -m pip install \
  "https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"

Development checkouts and optional backends remain explicit:

python -m pip install -e ".[dev]"
python -m pip install -e ".[open3d]"

The core package stays ROS-independent. ROS bags are read through typed data adapters, and GPL or ecosystem-specific tools (for example LI-Init, used only as an audit baseline in a container) stay behind optional adapter or subprocess boundaries.

Reproduce the README visuals

The evidence card is deterministic and bound to its source result:

calibrex render examples/precomputed/result.yaml \
  --format evidence-card \
  --output docs/assets/readme-evidence-card.svg

calibrex compare \
  examples/precomputed/result.yaml \
  examples/precomputed/known_bad_result.yaml \
  --format evidence-table \
  --left-label accepted \
  --right-label "known bad" \
  --output docs/assets/readme-comparison-table.svg

The public-data GIF gallery has its own schema-valid manifest:

python tools/generate_calibration_evidence_gif.py --readme-gallery

Tests fail if the committed evidence card drifts from its validated source.

Design principles

  • Keep the sensor-agnostic core ROS-independent.
  • Treat dataset calibration as reference evidence, not absolute truth.
  • Require evaluation for calibration behavior changes.
  • Record provenance for every generated result.
  • Keep GPL and ecosystem tools behind adapters or subprocess boundaries.
  • Preserve schema stability and evaluation quality over short-term convenience.

Documentation

License

Apache-2.0. If you use Calibrex in research, see CITATION.cff.

About

Check whether the sensor calibration on your robot is still right: one command, honest per-axis verdicts for LiDAR, IMU, camera, GNSS, and vehicle.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

27 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages