The workflow: estimate, check, drift¶
Calibrex answers three questions about the sensor calibration of a robot, one command
each. They chain, and every step reads your bag as it is (rosbag2 .db3, ROS 2 .mcap, ROS 1
.bag; see your bag format) with no ROS installation.
| You have | Ask | Command | You get |
|---|---|---|---|
| a bag and no calibration yet | what do the data say the extrinsics are? | calibrex estimate |
frames.yaml, static transforms, URDF joints (and a Kalibr camchain), per-axis std and observability |
| a calibration and another recording | is the deployed calibration still right? | calibrex check |
per-axis pass / warn / fail / inconclusive, what it could not judge, next steps, progress |
| several recordings of one rig over time | did the calibration change, and in which bag? | calibrex drift |
per-axis stable / drift / inconclusive, the deviating bag, the size of the change |
Real runs on the Koide hard-localization recordings
indoor_easy_01 and indoor_easy_02 (first 130 s each; calibrex 0.5.1). The
output lines are verbatim and abridged; the transcript and the digest of each full output are in
transcript.json.
The third recording is indoor_easy_01 with its IMU rotated 2 deg about z
(known-bad control).
The commands and outputs below are the ones in the recording. A KITTI drive set shows the same chain for a ground vehicle further down.
0. Install¶
python -m pip install "https://github.com/rsasaki0109/Calibrex/releases/download/v0.5.1/calibrex-0.5.1-py3-none-any.whl"
calibrex --help # the commands are grouped: Start here, Per-pair calibration, Evidence and CI, ...
calibrex --help lists Start here first: check, estimate, drift, doctor, demo,
inspect, init, calibrate and render. The per-pair commands, the evidence and CI tools and
the benchmark helpers follow in their own groups.
1. Estimate: no calibration yet¶
With a bag but no calibration, calibrex check has no candidate to judge (every pair reads
skipped (no_candidate_calibration)). calibrex estimate runs the same native estimators and
keeps what the data observed.
calibrex estimate indoor_easy_01/ --max-duration-s 130 --output est/ --html est.html
imu-lidar (imu_link / depth_camera_link): partial T_depth_camera_link_imu_link
roll -64.8441 deg +- 0.251 observed
pitch +65.3548 deg +- 0.294 observed
yaw +72.3625 deg +- 0.247 observed
x - unobservable (not constrained by the data (reported std 0.02803 m))
y - unobservable (not constrained by the data (reported std 0.03273 m))
z - unobservable (not constrained by the data (reported std 0.02292 m))
time offset +0.62 +- 0.18 ms (estimated)
axes that are not observed hold no measurement and are not exported
...
exported frames (root depth_camera_link):
depth_camera_link -> imu_link [imu-lidar] NOT MEASURED x, y, z (from --tf prior)
next steps:
- verify on a different recording: calibrex check <other bag> --tf est/frames.yaml
- deploy: est/static_transforms.sh (ROS 2 static_transform_publisher), est/static_transforms.launch.yaml or est/joints.urdf.xml (URDF joints)
- imu-lidar lever arm z: the motion excites it (9 deg RMS rotation about the other axes) but too briefly: std 1.95 cm against the 1 cm bound. About 6 min of the same motion would reach it (this recording: 2 min). ...
What this tells you, and what it will not pretend:
- Observed vs not measured. The rotation (roll, pitch, yaw) was observed with a standard
deviation of about 0.25 deg. The lever arm (x, y, z) was not constrained by this 2-minute
recording. An axis the data did not observe is never written as if it were measured: it is
written only when a rough
--tfprior supplies it (the YAML marks itNOT MEASURED), otherwise the frame is omitted and listed undernext steps. Here the bag's own/tf_staticwas the prior. - Why, and how much motion fixes it. The last
next stepsline is the translation observability hint: the recording rotates enough about the other axes to excite the lever arm, but too briefly, and about 6 minutes of the same motion would reach the 1 cm bound. For a recording that lacks a rotation axis entirely (a level vehicle) the hint names the missing rotation and says to measure it and pass it with--tfinstead. See what makes a lever arm observable, which also compares these predictions with what happened when more data were recorded. - Files.
est/frames.yaml(loads directly with--tf),est/static_transforms.sh,est/static_transforms.launch.yaml,est/joints.urdf.xml, andest/bag_estimate.json(slac.bag_estimate/v0.1, with the bag digest as provenance).--html est.htmlwrites one offline page:
The top of est.html. Observed axes are green bars; the greyed rows
are not measurements. The page also lists the exported files with digests and copy buttons for the
calibrex check --tf and static_transform_publisher commands.
calibrex render est/bag_estimate.json --format html regenerates the page later. Rotation-only
estimators (camera-IMU, the vehicle pairs) need a rough lever arm from --tf; lidar-lidar starts
from a --tf prior; camera-IMU needs the intrinsics (a CameraInfo topic or a Kalibr camchain).
Real-data results, with the failures: calibrex estimate on real
data.
2. Check on another recording¶
Deploy the estimate (est/static_transforms.sh, the launch file or the URDF joints), then verify it
on a different recording than the one it was estimated from:
calibrex check indoor_easy_02/ --tf est/frames.yaml --max-duration-s 130 --pairs imu-lidar --html check.html
candidate sources:
frames_yaml: est/frames.yaml (1 frames) sha256=35eb35a1d7b7
...
pair sensors verdict |delta|/tolerance unchecked detectable
imu-lidar imu_link / depth_camera_link pass (partial: roll, pitch, yaw only) roll 0.255/1.5 deg; pitch 0.982/1.5 deg; yaw 0.484/1.5 deg x, y, z roll 1.75 deg; pitch 2.48 deg; yaw 1.98 deg
...
overall verdict: pass
WARNING: 1 pair(s) have partial coverage: unchecked axes were not judged, so a pass covers only the judged axes
Read it as: the rotation estimated on recording 1 agrees with recording 2 to within a quarter to
two thirds of the 1.5 deg tolerance (rigid-scan floor), and the data could detect a rotation error of
about 2 deg (detectable). The lever arm is unchecked, not passed. check also prints
the same next steps lever-arm hint, shows progress per pair while the estimators run (the
check: lines on stderr), and with --html check.html writes the per-axis verdict report. The
exit status is non-zero on fail (see --fail-on), so it gates a pipeline; see
Calibration CI.
check also works without estimate: calibrex check my_bag/ reads the calibration already in the
bag's /tf_static, and --tf accepts a URDF, a Kalibr camchain or an RTK-SLAM calibration. The full
reference is Check a deployed calibration; the 3 degree yaw-error demo that
fails and the vendor calibration that passes is the GIF at the top of the
README.
3. Drift over time¶
Over weeks a rig gets serviced, bumped or remounted. Keep one bag per occasion, in recording order, and ask whether the calibration is the same in all of them. No reference calibration is needed.
calibrex drift indoor_easy_01/ indoor_easy_02/ remounted/ --max-duration-s 130 --pairs imu-lidar \
--output drift/ --html drift.html
overall: drift (a change is flagged beyond max(3 sigma, floor) with chi-square p < 0.01; rotation in deg, translation in cm)
imu-lidar: drift T_depth_camera_link_imu_link
axis indoor_easy_01 indoor_easy_02 remounted max|diff| min detectable status
roll -64.844 +- 0.251 -64.821 +- 0.210 -65.717 +- 0.251 0.896 1.500 stable
pitch +65.355 +- 0.294 +64.647 +- 0.258 +63.962 +- 0.294 1.392 1.500 stable
yaw +72.363 +- 0.247 +71.421 +- 0.199 +70.824 +- 0.247 1.539 1.500 drift
x - - - - - inconclusive
y - - - - - inconclusive
z - - - - - inconclusive
remounted vs the others: roll -0.886 deg, pitch -0.992 deg, yaw -0.968 deg; rotation 1.42 deg
deviating bag(s): remounted
next steps:
- imu-lidar: remounted disagrees with the other recording(s); recalibrate it (calibrex estimate remounted --output DIR) or check the mounting, and confirm with calibrex check remounted --tf <the calibration>
remounted/ is a copy of indoor_easy_01 whose IMU vectors were rotated 2 deg about z
(tools/rotate_imu_in_bag.py), which is what a physically remounted IMU measures. drift flags
it, names the bag, and reports the change it found: a 1.42 deg rotation against the injected 2 deg
(the rigid-scan estimates carry about 0.25 deg of noise per axis, and the other two recordings agree
with each other to better than 1 deg). Two recordings alone, indoor_easy_01 and indoor_easy_02, read stable
with the same table. Read stable as "no change larger than min detectable was found",
and an axis the data did not observe in two bags is inconclusive, not stable. With only two bags a
difference cannot be attributed to one of them, and the output says so. How small a rotation is
detected, and the false-alarm behaviour on unmodified recordings:
Drift on real data. Reference: Detect calibration
drift.
The same chain on a ground vehicle (KITTI)¶
For a ground vehicle the vehicle pairs are opt-in (--vehicle-frame base_link). These runs use the
KITTI raw development drives converted with calibrex convert kitti-raw (the pooled bag is
drives 0005, 0014 and 0022; the held-out drives 0027-0059 were not used) and the OXTS twist as the
wheel proxy.
K="--vehicle-frame base_link --topic-kind /oxts/twist=wheel"
calibrex estimate kitti_0005_0014_0022/ $K --output est_kitti/ --html est_kitti.html
calibrex check kitti_0009/ --tf est_kitti/frames.yaml $K --html check_0009.html
calibrex drift kitti_0005_0014_0022/ kitti_0009/ kitti_0015/ $K --output drift_kitti/
lidar-vehicle (velo_link / base_link): partial T_base_link_velo_link
roll - unobservable (not constrained by the data (reported std 0.5074 deg))
pitch +0.5254 deg +- 0.0515 observed
yaw -0.3127 deg +- 0.0939 observed
exported frames (root base_link):
base_link -> velo_link [lidar-vehicle] NOT MEASURED roll, x, y, z (from --tf prior)
velo_link -> imu_link [ins-lidar] NOT MEASURED yaw, x, y, z (from --tf prior)
imu-vehicle (imu_link / base_link): failed
no_judgeable_axes: no axis was constrained by the data: roll unobservable (reported std 0.574 deg); ...
Abridged and re-spaced:
pair verdict |delta|/tolerance unchecked
lidar-vehicle pass (partial: yaw only) yaw 0.0917/0.5 deg roll, pitch
imu-vehicle inconclusive - roll, pitch, yaw
ins-lidar inconclusive - roll, pitch, yaw, x, y, z
lidar-wheel_odometry pass (partial: yaw only) yaw 0.095/0.5 deg roll, pitch
overall verdict: inconclusive
lidar-wheel_odometry: stable T_base_link_velo_link
axis kpool3 k0009 k0015 max|diff| min detectable status
roll - - - - - inconclusive
pitch +0.525 +- 0.052 - - - - inconclusive
yaw -0.304 +- 0.091 -0.215 +- 0.040 -0.258 +- 0.049 0.089 0.500 stable
overall: inconclusive
A short drive observes one or two axes of a vehicle pair, so most of the table is
inconclusive or "yaw only" and the overall verdict says so instead of passing. That is the
intended behaviour; more or longer drives, with turns, observe more axes. These are development-drive
tool runs, not a held-out claim; the pre-registered KITTI result is on the KITTI LiDAR-vehicle
page.
In the browser¶
Nothing to install: the bag check page runs calibrex check and
calibrex estimate on your own rosbag2 under Pyodide, with the same code as the
command line. Drop the bag's metadata.yaml and storage file (multi-GB bags are read in place, not
copied or uploaded), pick the pairs and a duration cap, and either plan (which pairs can be
checked and why others are skipped), check (verdicts, the same report as --html) or
estimate (download frames.yaml and the exports). Each result lists the equivalent CLI
command. Browser runs are about twice as slow as native and read only the first part of the bag;
the details and limits are in Plan and run a bag in the
browser. The browser calibration
page is the separate, smaller tool that calibrates an IMU against a
trajectory you already have.
Your bag format¶
The three commands detect the format from the file and read it directly, with no conversion:
| Your bag | Notes |
|---|---|
rosbag2 directory or .db3 (sqlite3) |
per-message zstd / lz4 needs calibrex[rosbag2-compression] |
ROS 2 MCAP (.mcap) |
chunks none, zstd or lz4 |
ROS 1 .bag |
chunks none, bz2 or lz4 (calibrex[rosbag1-lz4] for lz4) |
calibrex check my_bag/ --plan lists which sensor pairs the bag can be checked for and why the
others are skipped, without running an estimator. The details are in supported input
formats. The ROS 1 and MCAP-lz4 readers are on main
and ship with the release after v0.5.1; the 0.5.1 wheel used above reads rosbag2 .db3 and .mcap
(lz4 MCAP chunks and ROS 1 bags need main).
Which pairs. Besides the IMU-LiDAR run above, the same commands cover camera-IMU, LiDAR-LiDAR,
the GNSS pairs, the ground-vehicle pairs and, newly, camera-LiDAR. camera-LiDAR is
targetless edge alignment and judges rotation only: translation is never judged (its estimate
was up to 11 cm off with a 1-2 cm std). estimate and drift take it with a rough --tf prior;
the evidence is in the camera-LiDAR results.
Autoware bags¶
Autoware's all-sensors sample bag stores raw velodyne_msgs/VelodyneScan packets and one
concatenated cloud in base_link, so there is no per-sensor cloud to register. Calibrex decodes
the packets (VLP-16 and VLP-32C, single return) and
tools/velodyne_scan_to_pointcloud2.py (in the repository) writes a derived bag with one PointCloud2 per sensor;
autoware_auto_vehicle_msgs/VelocityReport is read as wheel speed and yaw rate. On that bag (36.5 s,
a straight drive) lidar-lidar passes for the deployed /tf_static and fails for +1 deg, +3 deg
and +5 cm perturbations. The IMU and GNSS pairs stay inconclusive: a straight drive excites no
rotation, and the output states the recording length and turning that would. The decoder is checked
against the bag's own concatenated cloud (median 0.03 cm on the VLP-16, 1.8 to 3.1 cm on the
VLP-32C); dual-return mode and other models are untested. See
the Autoware section.
Where to go next¶
| Question | Page |
|---|---|
Every option of check, the verdicts, the cache, real-data results |
Check a deployed calibration |
| Why is this axis unobservable, how much motion fixes it | Translation observability |
Detection limits and false alarms of drift |
Drift on real data |
| How good are the estimates against Kalibr, KITTI and other references | estimate on real data, SOTA leaderboard |
| Gate a pipeline on the verdict | Calibration CI |
| Your own bag, step by step | Calibrate your own data |
What the runs cost¶
The recording above ran on a laptop CPU with a cold estimator cache: estimate of the 130 s Koide
bag took 5 min 19 s, check of the second bag 5 min 2 s (it estimates that bag, then judges), and
drift over three bags 4 min 46 s (the first two came from the cache; only remounted was new).
--cache-dir is shared by all three commands, so a re-run, or a check of a bag that drift
already estimated, costs seconds. The KITTI estimate over the 1268-sweep pooled bag took 5 min 27 s
and the check of drive 0009 1 min 52 s. These times were measured before the odometry speed-up
on main, which makes uncached first runs 1.25 to 2 times faster with bit-identical results
(compared artifact by artifact on KITTI, RTK-SLAM and Koide); the cached re-run is unchanged.