Estimate it, check it, watch it drift.
The sensor calibration of your robot, from ROS 2 and ROS 1 bags, with honest per-axis verdicts for
LiDAR, IMU, camera, GNSS, and vehicle.
Real calibrex check runs on KITTI development drives: a deployed
LiDAR transform with a 3° yaw error fails; restoring the calibration file and
re-checking passes (pitch and yaw judged; roll is not observable from driving).
Camera ↔ LiDAR, rotation only (KITTI development drive 0005). Two real
calibrex check --pairs camera-lidar runs: the vendor calibration passes
(roll and pitch judged), the same calibration turned +3° about the camera y axis fails
(pitch 3.00° against 0.50°). The frames in between only slide the overlay; they are not check
results.
Every frame is rendered from a real recording by Calibrex's own code; generators, input digests and licenses are in readme-data-gifs.json.
Real runs on the Koide hard-localization recordings (abridged output): estimate a calibration from one bag, check it on another recording, then ask whether it changed across recordings. The third recording is a copy with its IMU rotated 2° about z. The full walkthrough.
python -m pip install "https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"
calibrex estimate my_bag/ --output est/ --html est.html # 1. no calibration yet: measure it, export frames.yaml / static transforms / URDF
calibrex check other_bag/ --tf est/frames.yaml --html check.html # 2. is the deployed calibration still right on another recording?
calibrex drift day1/ day2/ day3/ --output drift/ --html drift.html # 3. did it change over time, and in which bag?| Step | Command | You get |
|---|---|---|
| 1 | calibrex estimate <bag> |
per pair each axis with its std and whether the data observed it; frames.yaml (loads with --tf), static_transform_publisher commands, URDF joints (and a Kalibr camchain); axes the data did not observe are marked NOT MEASURED, never written as measured; --html report |
| 2 | calibrex check <bag2> --tf frames.yaml |
pass / warn / fail / inconclusive per axis, the axes it could not judge, the error the data could have detected, next steps, progress while it runs; --html report. Without --tf it reads the bag's /tf_static |
| 3 | calibrex drift <bag1> <bag2> ... |
per axis stable / drift / inconclusive across recordings of one rig; the deviating bag and the size of the change |
| Pair | check judges |
estimate / drift |
|---|---|---|
| IMU ↔ LiDAR | rotation, clock offset, lever arm when observed | yes |
| camera ↔ IMU | rotation, clock offset (needs intrinsics) | yes |
| camera ↔ LiDAR | rotation only, targetless edge alignment; translation is never judged | needs a rough --tf prior; rotation only |
| LiDAR ↔ LiDAR | rotation and translation | needs a rough --tf prior |
| GNSS ↔ LiDAR, GNSS ↔ IMU | antenna lever arm | yes |
| LiDAR / IMU / INS / wheels ↔ vehicle | the rotation axes the motion excites (roll is unobservable on planar driving); opt-in with --vehicle-frame |
yes |
| camera focal length | deployed focal lengths | check only |
Every row reports per axis whether the data could judge it. camera-LiDAR is development evidence, not a claim: on KITTI and Hilti recordings the unmodified reference never warned or failed and 23 of 27 perturbed candidates were flagged (results).
check, estimate and drift read the bag as it is, with no conversion and no ROS installation;
the format is detected from the file:
| Your bag | Notes |
|---|---|
rosbag2 directory or .db3 (sqlite3) |
per-message zstd / lz4 needs calibrex[rosbag2-compression] |
ROS 2 MCAP (.mcap) |
chunks none, zstd or lz4 |
ROS 1 .bag |
chunks none, bz2 or lz4 (calibrex[rosbag1-lz4] for lz4) |
calibrex check my_bag/ --plan prints which sensor pairs the bag can be checked for, and why the
others are skipped, without running an estimator. Vehicle pairs are opt-in with
--vehicle-frame base_link. Details: supported input formats.
The ROS 1 and MCAP-lz4 readers, camera-LiDAR, the Autoware Velodyne decoder and the faster first runs are on
mainand arrive with the release after v0.5.1. To use them now, install from the repository:python -m pip install "git+https://github.com/rsasaki0109/Calibrex.git".
Autoware bags. Autoware's all-sensors sample records raw velodyne_msgs/VelodyneScan packets
and, as a point cloud, only one concatenated cloud in base_link. Calibrex decodes the packets
(VLP-16 and VLP-32C, single return) with tools/velodyne_scan_to_pointcloud2.py (in the repository), which writes a
derived bag with one PointCloud2 per sensor, and reads autoware_auto_vehicle_msgs/VelocityReport
as wheel speed. On that bag (36.5 s, a straight drive) lidar-lidar gives a real pass for the
deployed /tf_static and fail for +1 / +3 deg and +5 cm perturbations; the IMU and GNSS pairs
stay inconclusive because a straight drive excites no rotation, and the output says how much
recording and turning would. See the Autoware section.
In your browser, no install: the
bag check page runs calibrex check and
calibrex estimate on your own rosbag2 (plan which pairs can be checked, then run them, with
progress, verdicts and frames.yaml downloads). Multi-GB bags are read lazily; it runs locally under
Pyodide, so no data is uploaded. It was compared with the command line on real bags (KITTI, Koide,
NTU VIRAL, RTK-SLAM, Hilti; 30-70 s slices, bags up to 22 GB): same verdicts, numbers equal to 1e-6
or better, except the camera pairs, whose estimates differ within one to two sigma because Pyodide
ships an older OpenCV. The page also takes --topic-kind and --camera. The
browser calibration page calibrates an IMU against a
sensor trajectory (TUM): rotation, clock offset, gyro bias, and lever arm, with held-out evidence.
When an axis is not observable, it says why and what would fix it. For a lever-arm (translation)
axis calibrex check and calibrex estimate name the rotation the recording lacks, or how much more
of the same motion would reach the bound; see
what makes a lever arm observable. calibrex --help
groups the commands: Start here (check, estimate, drift, doctor, demo, ...), per-pair
calibration, evidence and CI, and the rest.
🔍 Three commandsestimate, check, drift: reads the bag, finds the sensors, and covers up to ten pairs (GNSS, IMU, LiDAR, camera, vehicle).
|
⚖️ Honest verdicts Per-axis pass / warn / fail / inconclusive, with partial coverage and detection power reported.
|
🧪 Pre-registered SOTA Thresholds are frozen before held-out data is scored; refuted audits stay on the leaderboard. |
Docs · Workflow · Estimate · Check · Drift · SOTA audits · Coverage · Gallery · Quickstart · Methods
calibrex check judges a candidate calibration. With a bag but none, every pair is
skipped (no_candidate_calibration); calibrex estimate is the one-command path to one, and
check then verifies it on a different recording:
calibrex estimate my_bag/ --output est/ # add --tf rough.yaml for a prior
calibrex check other_bag/ --tf est/frames.yaml # verify on a different recording
calibrex estimate my_bag/ --output est/ --html est.html # also a self-contained HTML report--html (or calibrex render est/bag_estimate.json --format html later) writes a single
offline page: per pair each axis as value, 1 sigma and observability (axes the data did not
observe are greyed and marked "NOT MEASURED"), the exported and omitted frames, the exported
files with their digests, and copy buttons for the calibrex check --tf and
static_transform_publisher commands.
It runs the same native estimators and keeps the estimates (est/bag_estimate.json,
slac.bag_estimate/v0.1): per pair the transform, each axis with its standard deviation,
and whether the data observed it (observed, unobservable, control_not_detected,
not_estimated), with the bag digest and settings as provenance. From the observed
frames it writes frames.yaml (loads directly with --tf), ROS 2
static_transform_publisher commands and a launch file, and URDF joints (plus a Kalibr
camchain-imucam.yaml when --tf gave a camchain). A frame with an axis the data did not
observe is never written as if measured: it is written only when a rough --tf prior
supplies that axis (the YAML marks it NOT MEASURED), otherwise it is omitted and listed
under next steps. Rotation-only estimators (camera-IMU, the vehicle pairs) therefore need a
rough lever arm from --tf; lidar-lidar registration and camera-lidar edge alignment (rotation only) also start from a --tf prior, and
camera-IMU and camera-LiDAR need the camera intrinsics (a CameraInfo topic or a Kalibr camchain).
Real-data results, with the splits, exact commands and the failures (Hilti camera-IMU, Koide
imu-lidar, KITTI vehicle pairs, RTK-SLAM GNSS pairs):
calibrex estimate on real data.
calibrex check my_bag/ --output check.json --html check.html # reads /tf_static; add --tf rig.urdf to overridecalibrex check asks whether the extrinsics deployed on a robot agree with what
a recording says (step 2 of the workflow; no calibration yet? estimate one first). It reads the candidate transforms from the bag's /tf_static
(or --tf: the frames.yaml of calibrex estimate, URDF, Kalibr, RTK-SLAM or Hilti calibration files), detects the
sensor topics, works out which of the ten sensor pairs the bag can check, runs
the native estimator of each, and judges the deployed transform per axis against
the estimate and its uncertainty: pass, warn, fail or inconclusive. It
also checks the estimates against each other (rig closure) and reports which
axes it could not judge. Vehicle pairs (LiDAR, IMU or wheel odometry against the
vehicle frame) are opt-in with --vehicle-frame base_link.
A sweep of injected errors. Not a calibration running: a yaw
error from −3° to +3° is injected into the deployed velo_link transform and
each step is one real calibrex check run. Errors of about 1° or more turn
lidar-vehicle to fail. Even at 0° the vendor transform is about 0.3° from the
estimate, and the overall verdict stays inconclusive because imu-vehicle
cannot judge on KITTI.
Demo: the KITTI raw development drives (0005, 0009, 0014, 0015, 0022) pooled into
one bag, with velo_link of the deployed tf turned about its parent's z axis
(script, first run about 6 minutes, the other
variants seconds):
| deployed tf | overall | lidar-vehicle | lidar-wheel_odometry | ins-lidar | imu-vehicle |
|---|---|---|---|---|---|
| vendor (KITTI calibration) | inconclusive |
pass, partial | pass, partial | pass, partial | inconclusive |
velo_link yaw +1 deg |
fail |
fail (yaw 1.31 deg vs 0.5) | fail (yaw 1.31 deg vs 0.5) | pass, partial | inconclusive |
velo_link yaw +3 deg |
fail |
fail (yaw 3.31 deg vs 0.5) | fail (yaw 3.31 deg vs 0.5) | pass, partial | inconclusive |
pair verdict |delta|/tolerance unchecked
lidar-vehicle fail (partial: pitch, yaw only) pitch 0.505/0.5 deg [warn]; yaw 1.31/0.5 deg [fail] roll
imu-vehicle inconclusive - roll, pitch, yaw
ins-lidar pass (partial: roll, pitch only) roll 0.114/0.5 deg; pitch 0.0567/0.5 deg yaw, x, y, z
lidar-wheel_odometry fail (partial: pitch, yaw only) pitch 0.505/0.5 deg [warn]; yaw 1.31/0.5 deg [fail] roll
overall verdict: fail (velo_link yaw +1 deg)
Read it honestly: every vehicle pair has partial coverage (roll is not observable
from planar driving, and ins-lidar cannot see yaw), so a pass covers only the
judged axes. The vendor run is inconclusive, not pass, because the INS
velocities are too noisy for imu-vehicle to judge anything on KITTI; and the vendor
pitch (0.496 deg) sits right at the 0.5 deg tolerance floor. The two LiDAR-motion pairs read the same
odometry, so they agree by construction. Only the 1 and 3 degree yaw errors are
injected here; the numbers are the development drives, not a held-out claim.
Tutorial · demo summary · HTML report of the +1 deg run (download and open; GitHub shows HTML as source).
Two or more bags of the same rig, in recording order:
calibrex drift day1/ day2/ day3/ --pairs imu-lidar --output drift/ --html drift.htmlEach bag is estimated (the estimator cache makes a re-run cheap; drift/<bag>/bag_estimate.json)
and every axis the data observed in at least two bags is tested for consistency: pairwise
differences against max(3 sigma, floor) (the minimum detectable change of the axis type) and a
chi-square homogeneity test. Per pair the verdict is stable, drift (with the deviating bag
named when three or more bags allow it, and the size of the change) or inconclusive; exit
status 1 on drift (--fail-on). Result: drift/calibration_drift.json,
slac.calibration_drift/v0.1. A stable verdict means no change larger than the reported
minimum detectable change was found. Per-pair clock offsets are compared too (see
the drift page).
Calibrex may call a method state of the art only for a claim that a frozen
audit marks supported. The claim, scoring, held-out data, and thresholds are
committed before the held-out data is scored, each piece of evidence is a
digest-pinned artifact, and refuted audits stay on the
SOTA leaderboard. Standings as of
2026-10-02:
| Pair | Standing | Scope of the audited claim | Details |
|---|---|---|---|
imu-lidar |
supported 2, refuted 1 | Livox MID360 against LI-Init (GPL, run in a container): rotation and clock offset (4/4 gates), and the full extrinsic in a second round (4/4). The first full-extrinsic round was refuted (0/8) and stays listed. | MID360 IMU-LiDAR |
lidar-vehicle |
supported 1 | Motion-only LiDAR-to-vehicle rotation against KITTI's calib_imu_to_velo on eight unseen drives (3/3 gates). Pitch and yaw only; roll stays unobservable. |
KITTI LiDAR-vehicle |
lidar-lidar |
refuted | NTU VIRAL (two Ouster OS1-16): refuted 2/4. Accuracy gates pass, but the method does not beat scan-to-scan and cross-recording consistency fails (1.143, bound 1.0). | NTU VIRAL LiDAR-LiDAR |
camera-imu |
refuted | Hilti 2022 forward cameras (cam0, cam1) against Kalibr on four unseen recordings: refuted 6/7. Rotation (max 0.70 deg), clock offset (0.33 ms), consistency (0.43 deg) and the cam0-cam1 relative rotation (0.41 deg) pass, but the shortest recording (exp04) leaves an axis unconstrained on both cameras. | Hilti camera-IMU |
| the other six target pairs | no_claim | Native methods exist, but no audited claim yet. | Leaderboard |
A supported standing is scoped to one dataset family and one comparison; it
is not a general accuracy claim. Held-out data is reserved per pair, and a
spent split cannot be reused for a second claim. The machine-readable
standings are in
sota_leaderboard.json.
calibrex sota audit protocol.yaml --output audit.yaml # exit 0 only for supportedAll ten target sensor pairs have a native method. "Audit standing" is the standing on the leaderboard.
| Pair | Native command | Public data used | Audit standing |
|---|---|---|---|
| camera (focal lengths) | calibrex camera-imu focal |
Hilti 2022 | no_claim |
| camera ↔ IMU | calibrex camera-imu rotation |
Hilti 2022 (dev exp21, exp07; audit exp01-exp04) | refuted |
| camera ↔ LiDAR | calibrex check / estimate / drift (targetless edge alignment, rotation only; results), calibrex camera-lidar ..., calibrex demo kitti-lidar-camera-evidence |
KITTI dev drives and Hilti 2022 (check), KITTI-shaped fixture, A2D2, ACFR | no_claim |
| GNSS ↔ IMU | calibrex gnss-imu compose (from calibrex gnss-lidar rtk-slam and IMU-LiDAR) |
RTK-SLAM | no_claim |
| IMU ↔ LiDAR | calibrex imu-lidar livox, livox-translation, trajectory |
RTK-SLAM, Zenodo MID360 driving | supported (2) / refuted (1) |
| IMU ↔ vehicle | calibrex imu-vehicle kitti |
KITTI raw | no_claim |
| INS ↔ LiDAR | calibrex ins-lidar kitti |
KITTI raw | no_claim |
| LiDAR ↔ LiDAR | calibrex lidar-lidar ros2 |
NTU VIRAL | refuted |
| LiDAR ↔ vehicle | calibrex lidar-vehicle kitti |
KITTI raw | supported (1) |
| LiDAR ↔ wheel odometry | calibrex lidar-wheel trajectory, kitti |
KITTI raw (OXTS stands in for wheels) | no_claim |
Also native or adapter-backed: hand-eye AX=XB and robot-world AX=YB
(13 native methods, benchmarked below against OpenCV), radar extrinsics,
RGB-D joint SLAC, and solid-state LiDAR clock-offset profiling, mostly through
calibrex calibrate <config>. Open3D, Kalibr, Koide, ROS, and Autoware stay
behind adapters. See calibration methods
for solver-level detail, and the per-pair pages:
IMU-LiDAR,
LiDAR-vehicle,
LiDAR-lidar,
camera-IMU,
INS-LiDAR,
LiDAR-wheel,
GNSS-LiDAR,
GNSS-IMU.
Every gallery asset is generated from public raw data. Dataset source,
protocol, parameters, and digests are recorded in
readme-gif-gallery.json.
Run the complete evidence path without ROS or a dataset download. The command evaluates the bundled deterministic KITTI-shaped LiDAR-camera fixture, writes a schema-valid result and review report, and verifies a digest-bound provenance bundle:
python -m pip install \
"https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"
calibrex demo kitti-lidar-camera-evidence \
--output-dir outputs/kitti-lidar-camera-evidence \
--strict-assessment
calibrex validate outputs/kitti-lidar-camera-evidence/result.yaml
calibrex verify outputs/kitti-lidar-camera-evidence/bundle.jsonOpen outputs/kitti-lidar-camera-evidence/report.html to inspect the projection
evidence and known-bad perturbation probes. The checked fixture detects 16 of
24 mandatory perturbation cases and passes all six falsification-policy gates
in under one minute in the clean Windows wheel smoke test. This verifies the
pipeline and evidence contracts; it is not a real-sensor accuracy claim or a
standalone camera-LiDAR calibration algorithm. Supply an officially downloaded
KITTI raw sequence with --dataset-path when evaluating real data.
To render and validate a committed result from a source checkout:
git clone https://github.com/rsasaki0109/Calibrex.git
cd Calibrex
python -m pip install .
calibrex validate examples/precomputed/result.yaml --json
calibrex render examples/precomputed/result.yaml \
--format evidence-card \
--output outputs/quickstart/evidence-card.svg
calibrex render examples/precomputed/result.yaml \
--output-dir outputs/quickstartBefore configuring a solve, diagnose a recording and save the result for review or CI:
calibrex doctor recording.mcap \
--output outputs/doctor.json \
--json
calibrex validate outputs/doctor.jsondoctor infers supported dataset types, reports missing optional dependencies,
data coverage and degeneracy warnings, and suggests compatible evidence
workflows. Running it without a path retains the lightweight environment check.
Run the same evidence gates on every calibration change:
- uses: actions/checkout@v4
- uses: rsasaki0109/Calibrex@v0.5.1
with:
candidate: calibration/candidate.yaml
baseline: calibration/baseline.yamlThe action writes a GitHub Step Summary, fails on FAIL or INCONCLUSIVE by
default, and exposes schema-valid evidence, comparison, SVG, and
calibration-ci.json artifacts. See Calibration CI.
Tried it on your rig? Share a sanitized result or a useful failure case in Discussions. If the evidence-first workflow earns a place in your calibration stack, consider starring Calibrex so other robotics teams can find it.
Calibrex evaluates all 13 native hand-eye methods and seven OpenCV 4 variants
on the ETHZ ASL real robot-arm dataset.
Every method sees the same five digest-locked absolute-pose splits. AX=XB and
AX=YB are separate equation families and are never ranked together.
| Method | Holdout rotation RMSE deg ↓ | Holdout translation RMSE mm ↓ | Failure rate | Runtime s |
|---|---|---|---|---|
| Calibrex Park-Martin | 0.867227 [0.759962, 0.958155] | 13.8759 [11.7678, 15.8266] | 0.0% | 0.153451 |
| Calibrex Tsai-Lenz | 0.875291 [0.77004, 0.967317] | 13.8652 [11.7326, 15.8293] | 0.0% | 0.126921 |
| Calibrex Daniilidis | 0.868969 [0.761845, 0.960775] | 13.8212 [11.9121, 15.6887] | 0.0% | 0.322403 |
| Calibrex Andreff | 0.867199 [0.759958, 0.958223] | 13.879 [11.7674, 15.8309] | 0.0% | 0.334728 |
| Calibrex Shiu-Ahmad | 1.6191 [0.787831, 3.10277] | 31.9219 [12.7334, 52.761] | 0.0% | 14.1537 |
| Calibrex Chou-Kamel | 0.867205 [0.759953, 0.958241] | 13.8795 [11.7684, 15.8315] | 0.0% | 0.129661 |
| Calibrex Horaud-Dornaika | 0.890804 [0.782116, 0.996268] | 14.1434 [11.8641, 16.1088] | 0.0% | 0.139859 |
| Calibrex H-D nonlinear | 0.885789 [0.777979, 0.991045] | 14.0634 [11.8408, 16.0027] | 0.0% | 19.4512 |
| OpenCV Tsai | 0.871872 [0.769051, 0.961725] | 13.8681 [11.8255, 15.6992] | 0.0% | 0.0368567 |
| OpenCV Park | 0.867225 [0.759969, 0.958144] | 13.8754 [11.7667, 15.8261] | 0.0% | 0.0280751 |
| OpenCV Horaud | 0.867203 [0.759961, 0.958229] | 13.879 [11.7674, 15.831] | 0.0% | 0.0248639 |
| OpenCV Andreff | 0.868623 [0.763653, 0.959819] | 15.4855 [13.6068, 17.338] | 0.0% | 0.0327896 |
| OpenCV Daniilidis | 0.871005 [0.764385, 0.963618] | 13.9265 [11.9635, 15.6961] | 0.0% | 0.0284915 |
| Method | Holdout rotation RMSE deg ↓ | Holdout translation RMSE mm ↓ | Failure rate | Runtime s |
|---|---|---|---|---|
| Calibrex Shah | 0.624739 [0.559308, 0.688861] | 10.7793 [9.53397, 11.9952] | 0.0% | 0.0206115 |
| Calibrex Li-Wang-Wu | 0.627131 [0.561467, 0.692796] | 19.651 [15.2668, 23.3423] | 0.0% | 0.0321106 |
| Calibrex Dornaika-Horaud | 0.624742 [0.559311, 0.688865] | 10.7793 [9.53399, 11.9952] | 0.0% | 0.0286939 |
| Calibrex Zhuang-Roth-Sudhakar | 0.651861 [0.569557, 0.733996] | 10.961 [9.65144, 12.1956] | 0.0% | 0.0213057 |
| Calibrex D-H nonlinear | 0.6265 [0.560987, 0.692012] | 10.7752 [9.52577, 12.0161] | 0.0% | 0.217955 |
| OpenCV Shah | 0.624739 [0.559308, 0.688861] | 10.7793 [9.53397, 11.9952] | 0.0% | 0.00343018 |
| OpenCV Li | 0.627131 [0.561467, 0.692796] | 19.651 [15.2668, 23.3423] | 0.0% | 0.00570322 |
These are scoped consistency results, not blanket accuracy claims: the archive does not provide an accepted ground-truth extrinsic. The tables report closure RMSE on untouched holdout poses, retain failures in the denominator, and show 95% bootstrap intervals across splits. See the full protocol, citations, limitations, and reproducible provenance.
Start with the small Livox sample for a fast, no-ROS check. It writes a
calibrated result.yaml, report.html, evidence sidecars, and a verified
provenance bundle:
python -m pip install .
calibrex demo livox-evidence \
--output-dir outputs/solid-state-livox-demo \
--json
calibrex validate outputs/solid-state-livox-demo/result.yaml
calibrex verify outputs/solid-state-livox-demo/bundle.json- Real timing and clock offset. The public TIERS VLP-16 ↔ Livox Horizon
bag (about 7.18 GB, never downloaded automatically) is profiled with
calibrex continuous-time-lidar-pair; see the public-dataset tutorial. On the checked run the selected offset was +40 ms and train RMSE changed from 0.0957 m to 0.0240 m, with 0.0242 m holdout RMSE. This is an algorithmic estimate under the declared holdout: the public sequence has no independent clock ground truth. - Synthetic truth gate. A deterministic benchmark recovers a known
extrinsic and +30 ms clock offset and rejects a fixed-clock known-bad
control (
tools/run_solid_state_synthetic_benchmark.py). It verifies solver mechanics, not real-sensor accuracy. See the YAML artifact and report. - Public cross-dataset benchmark. Paired solver variants are compared on
the same capture windows, temporal holdouts, and sampling seeds. In the v0.3
matrix, adaptive wins 23/27 replicates with mean improvement 47.70%
(bootstrap 95% CI [34.43, 59.99]%); AgRob Modular-e stays a visible
counterexample (5/9, mean −1.34%). The public-only v0.4 candidate scores
27/27 replicates: adaptive wins 25/27, mean improvement 57.20%
(95% CI [45.69, 67.43]%), and 9 of 54 variant artifacts hit the declared
max_iterationscategory. See the v0.3 report, v0.4 report, and train-only selection note. These are ground-truth-free temporal-holdout results, not an absolute accuracy or SOTA claim.
Full commands for every step are in the solid-state public benchmark runbook. The independent physical-metrology packet remains an optional future path and is checked in as a planned template, not a result: YAML and report.
What is implemented in the current alpha?
- typed config, result, comparison, protocol, policy, and evidence schemas
- native solvers for camera-IMU, IMU-LiDAR, GNSS-LiDAR/IMU, LiDAR-LiDAR, LiDAR/IMU-vehicle, LiDAR-wheel odometry, and INS-LiDAR, each with held-out windows, a block jackknife, and known-bad controls
- pair-agnostic pre-registered SOTA audits and per-pair standings (
calibrex sota) - offline and online/streaming LiDAR calibration for rosbag1, rosbag2, and MCAP
- motion compensation, per-point deskew, and trajectory evidence
- targetless camera-LiDAR mutual information and online monitoring
- native point-to-point, point-to-plane, hand-eye, and robot-world baselines
- backend-neutral joint SLAC with typed Schur pose elimination
- radar velocity, LiDAR-IMU rotation, and temporal-offset evidence
- typed external-run artifacts, a Kalibr camchain importer, and Koide/Open3D, ROS, and Autoware adapter boundaries
- a frozen, SHA-bound full-scale KITTI reference-vs-known-bad falsification runner
See the calibration methods and changelog for solver-level detail and limitations.
Calibrex does not turn every run green. Weak excitation, failed controls, and refuted audits are reported as evidence, not hidden as demo noise. A sample of current results from the benchmark pages:
| Pair / data | Result | Verdict |
|---|---|---|
lidar-lidar, NTU VIRAL held-out audit |
Accuracy gates pass (0.489 deg, 0.077 m to design), but no gain over scan-to-scan (paired CI low −0.0048) and cross-recording consistency 1.143 vs bound 1.0; 2/4 gates | ❌ Refuted |
imu-lidar, first full-extrinsic audit vs LI-Init |
Lever-arm x unobservable on both recordings, so the claim was contradicted; kept on the leaderboard beside the later supported round | ❌ Refuted |
camera-imu, Hilti 2022 dev (exp21, exp07), 10 camera runs |
5 pass, 1 warn, 4 inconclusive; side and down cameras on exp07 still do not constrain rotation (jackknife 0.35-0.7 deg) | |
ins-lidar, KITTI drives 0005 + 0009 |
Roll, pitch, and clock offset estimated; yaw and translations unobservable | |
lidar-wheel, KITTI (OXTS stand-in) |
Rotation matches lidar-vehicle within 0.004 deg; a 1 % speed-scale control is not detected | |
imu-lidar, MID360 driving |
Roll estimated; yaw unobservable because a vehicle rotates almost only about the vertical axis | |
imu-lidar, MID360 hand-held (RTK-SLAM seq2) |
Every rotation axis estimated; held-out windows detect every known-bad shift | ✅ Pass |
The failure is part of the product: gates refuse to certify what the available data cannot falsify.
Most calibration tools stop after producing a transform. Calibrex asks the next question: what evidence would falsify this transform?
Calibrex is an open-source, ROS-independent Python toolkit that turns candidate
extrinsics, time offsets, and trajectories into evidence backed by holdout
metrics, known-bad controls, observability checks, and reproducible provenance.
It ships native solvers for every sensor pair listed below, plus adapters for
Kalibr, Open3D, and Autoware, and it reports pass, warn, fail, or
inconclusive instead of optimizer convergence or a single training residual.
| A matrix gives you | Calibrex adds |
|---|---|
| One estimated transform | Candidate, reference, and selected estimates kept separate |
| One training residual | Train/holdout metrics and temporal stability |
| Optimizer convergence | Known-bad perturbation challenges |
| A covariance matrix | Rank, weak directions, and degeneracy warnings |
| A screenshot | Schema-valid artifacts with input and producer provenance |
| A result file | A digest-locked evidence bundle that can be verified later |
flowchart LR
A[Sensor data] --> B[Candidate calibration]
B --> C{Evidence gates}
C -->|Holdout| D[Generalization]
C -->|Known-bad| E[Falsification]
C -->|Observability| F[Weak directions]
D --> G[PASS / WARN / FAIL]
E --> G
F --> G
G --> H[HTML + schema-valid sidecars]
The card below is not a hand-authored mockup. It is rendered from a
schema-valid result.yaml, bound to the
source file by SHA-256, and committed with machine-readable provenance inside
the SVG.
Generate the same artifact from any Calibrex result:
calibrex render result.yaml \
--format evidence-card \
--output evidence-card.svgThis generated table compares the validated quickstart result with a declared 5° yaw / 0.15 m known-bad control. The amber protocol state is intentional: these compact example results do not declare a shared evidence protocol, so Calibrex shows the metric deltas but refuses to call the comparison compatible.
calibrex compare \
examples/precomputed/result.yaml \
examples/precomputed/known_bad_result.yaml \
--format evidence-table \
--left-label accepted \
--right-label "known bad" \
--output comparison-table.svgresult.yaml calibrated values + run provenance
├── report.html portable human-readable report
├── metrics.json train / holdout measurements
├── evidence.json protocols, summaries, and known-bad cases
├── assessment.json PASS / WARN / FAIL policy decision
├── observability.json rank and weak directions
├── degeneracy.json failure modes and limitations
├── protocol.json declared evaluation contract
├── policy.json falsification thresholds
├── transforms.json typed transform provenance
├── bundle.json artifact SHA-256 manifest
└── verification.json materialized bundle verification
calibrex evidence result.yaml --output evidence.json
calibrex assess evidence.json
calibrex verify bundle.json
calibrex compare reference.yaml candidate.yaml --enforce-compatiblerender never claims to recompute metrics. Cached inputs remain marked as
cached, raw-input verification remains visible, and incompatible protocols are
not silently ranked together.
The supported no-source install is the versioned wheel attached to the v0.5.1 GitHub Release:
python -m pip install \
"https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"Development checkouts and optional backends remain explicit:
python -m pip install -e ".[dev]"
python -m pip install -e ".[open3d]"The core package stays ROS-independent. ROS bags are read through typed data adapters, and GPL or ecosystem-specific tools (for example LI-Init, used only as an audit baseline in a container) stay behind optional adapter or subprocess boundaries.
The evidence card is deterministic and bound to its source result:
calibrex render examples/precomputed/result.yaml \
--format evidence-card \
--output docs/assets/readme-evidence-card.svg
calibrex compare \
examples/precomputed/result.yaml \
examples/precomputed/known_bad_result.yaml \
--format evidence-table \
--left-label accepted \
--right-label "known bad" \
--output docs/assets/readme-comparison-table.svgThe public-data GIF gallery has its own schema-valid manifest:
python tools/generate_calibration_evidence_gif.py --readme-galleryTests fail if the committed evidence card drifts from its validated source.
- Keep the sensor-agnostic core ROS-independent.
- Treat dataset calibration as reference evidence, not absolute truth.
- Require evaluation for calibration behavior changes.
- Record provenance for every generated result.
- Keep GPL and ecosystem tools behind adapters or subprocess boundaries.
- Preserve schema stability and evaluation quality over short-term convenience.
- Documentation site
- The workflow: estimate, check, drift
- SOTA leaderboard
- Benchmark pages: MID360 IMU-LiDAR · KITTI LiDAR-vehicle · NTU VIRAL LiDAR-LiDAR · Hilti camera-IMU · KITTI INS-LiDAR · LiDAR-wheel odometry · RTK-SLAM GNSS-LiDAR · RTK-SLAM GNSS-IMU · ETHZ hand-eye · ETHZ robot-world
- SLAC concept
- Calibration methods
- Solid-state LiDAR calibration
- Frame conventions
- Public datasets
- Your own data · Browser calibration
- Open3D adapter
- LiDAR-camera adapter
- License boundaries
- Support · Contributing · Security
Apache-2.0. If you use Calibrex in research, see
CITATION.cff.
















