{"id":"4c4f9b43-7eb4-4da9-ab2e-7a8a3214dc6b","arxiv_id":"2504.18500","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using a 7.1 kg robot sensor payload and seven real-world environments, the study quantifies how time offsets, extrinsic calibration errors, IMU grade, and camera and LiDAR choice affect odometry accuracy.","lead":"This paper presents Boxi, a tightly integrated sensor payload with two LiDARs, ten cameras, seven IMUs, and RTK GNSS, and it tests how payload design choices affect robot state estimation accuracy. It also provides a cookbook of sensor selection, synchronization, calibration, and mechanical design guidelines for building future robot perception systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GNSS modality comparison is partially self-referential: IE-TC is fused into the Holistic Fusion ground truth (App. B), so Tables IV/V likely understate GNSS ATE/RTE.","rationale":"The reader's weakest_assumption exactly identifies the self-referential GNSS baseline. I agree that this is load-bearing because the modality comparison supports one of the paper's three central claims (sensor modality has a crucial impact on state estimation). However, the concern does not invalidate the core DLIO time-synchronization and extrinsic-calibration findings (Section IV-G3), which are differential ablations evaluated against a TPS-anchored reference and are less sensitive to IE-TC inclusion. The paper even provides the tool needed to test the concern: HF-Optimized w/o IE is already computed for one dataset and is identical to the full version on that dataset, suggesting the circularity may be modest where TPS coverage is good. The check should be extended to all missions and to the GNSS row specifically. Given this, the reader's CONDITIONAL verdict remains appropriate: the quantitative modality claims need stronger controls, but the qualitative engineering guidelines are credible. No verdict change is required, but the concern should be explicitly addressed in revision.","tokens_in":35082,"tokens_out":8946,"duration_ms":90520,"concrete_test":"Recompute Tables IV and V for all seven missions using the HF-Optimized-w/o-IE ground truth (already defined in Appendix B and computed for Excavation Site in Table VIII) instead of the full HF-Optimized reference. Additionally, for the GNSS row, compute ATE/RTE restricted to time intervals where TPS measurements are available and separately for dropout intervals. If the GNSS ATE increases by more than 20% when evaluated without IE in the reference, the current GNSS comparison is substantially inflated and must be reported with the circularity caveat or removed from the modality ranking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the partially circular reference used for the modality comparison. Appendix B constructs the HF-Optimized ground truth by fusing TPS, the HG4930 IMU, and IE-TC GNSS poses. Tables IV/V then report ATE/RTE for the modality rows, including the row labeled 'GNSS [2]' (Inertial Explorer IE-TC). Because IE-TC is an input to the very ground truth it is scored against, the GNSS accuracy is inflated by construction, especially during TPS dropouts (e.g., Excavation Site, 84–121 s, Fig. 18) where HF-Optimized bridges with IE-TC. The validation in Table VIII covers only Excavation Site, only against TPS, and explicitly excludes dropout periods ('do not reflect the accuracy of position during the dropout periods'). Therefore the GNSS ATE/RTE averages (Table V: 0.06 m; Table IV: 0.008 m) are not independently established. A secondary but related effect: DLIO/OKVIS2 also use the HG4930 IMU, which is part of the same ground-truth pipeline, so during TPS outages their absolute errors may also be softened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Boxi, a compact multi-modal sensor payload for legged robots, and uses it to study how payload design decisions affect downstream state estimation. It reports modality-level comparisons (kinematic, LiDAR, camera, GNSS), LiDAR and camera choice ablations, IMU-grade comparisons, and synthetic time-synchronization and extrinsic-calibration perturbations across seven real-world environments. The paper also presents a hardware/software 'cookbook', a custom calibration toolchain, time-synchronization verification tools, and open-source code. The central claim is that time synchronization, calibration, and sensor modality have a crucial impact on state estimation, with specific quantitative findings such as DLIO being robust to small time delays and large translation offsets but sensitive to rotational extrinsic errors.","tokens_in":35340,"tokens_out":5612,"duration_ms":58153,"significance":"If the quantitative claims are supported, this is a valuable systems contribution: the open hardware/software release, the calibration chain with sub-millimeter LiDAR-camera consistency, the time-synchronization verification tool, and the multi-environment dataset are all concrete assets for the community. The paper also makes falsifiable statements about which design factors matter most, such as the priority of rotational calibration over sub-millisecond time synchronization for tightly coupled LiDAR-inertial odometry. However, several methodological gaps in the evaluation limit the strength of the quantitative conclusions, and at least one comparison is partially self-referential.","major_comments":[{"comment":"The GNSS modality comparison is partially self-referential. The HF-Optimized ground truth described in Appendix B fuses TPS measurements, the HG4930 IMU, and the Inertial Explorer IE-TC GNSS poses, and Table IV/V then score the row labeled 'GNSS [2]' (IE-TC) against that ground truth. Because IE-TC is an input to the reference trajectory, the GNSS ATE/RTE values in Tables IV and V are artificially low by construction, particularly during TPS dropouts such as the Excavation Site interval at 84-121 s shown in Figure 18, where HF-Optimized bridges with IE-TC. The validation in Table VIII covers only the Excavation Site, compares against TPS directly, and explicitly excludes dropout periods, so it does not independently establish the accuracy of the GNSS row over the full trajectories. The authors should either remove the GNSS row from the modality comparison, or re-evaluate IE-TC against a TPS-only reference on segments with continuous line-of-sight and clearly state that the reported GNSS numbers exclude dropout periods.","section":"IV-C, Tables IV-V, Appendix B"},{"comment":"The camera comparison confounds shutter type with camera interface, field-of-view overlap, timestamping quality, and sensor characteristics. In Section IV-E, the CoreResearch global-shutter cameras are compared with the HDR and ZED2i rolling-shutter cameras, but the ZED2i is connected via USB without accurate timestamping, the HDR cameras have limited overlapping FoV, and the underlying sensors differ beyond shutter type. Section IV-I itself acknowledges that effects are not measured in isolation. Despite this, Section IV-G3 concludes that 'naively replacing images recorded from a global shutter camera with rolling shutter images results in a significant performance drop.' This conclusion overreaches the evidence; the observed performance gap could be driven by FoV overlap, interface latency, or timestamping rather than shutter type alone. The authors should either add a controlled comparison in which only the shutter mechanism changes, or rephrase the conclusion to state the specific confounded configuration tested.","section":"IV-E, IV-G3, IV-I"},{"comment":"The modality comparison in Tables IV and V is not tuning-fair. Section IV-I states that DLIO parameters were tuned to achieve competitive performance across all datasets, while OKVIS2 was not fine-tuned for each camera setting. The claim in Section IV-C that 'DLIO consistently outperforms OKVIS2' therefore conflates algorithmic modality with tuning effort. The authors should either tune OKVIS2 with comparable effort, report sensitivity to parameter settings for both algorithms, or soften the comparative claim to reflect that DLIO was tuned while OKVIS2 was not.","section":"IV-C, IV-I"},{"comment":"The LiDAR and camera comparisons are based on very limited environment coverage. The LiDAR comparison in Section IV-D uses only the Hike and Warehouse datasets, and the camera comparison in Section IV-E uses only the Mountain Ascent dataset. The text nevertheless makes general claims such as 'the CoreResearch unit, utilizing global shutter cameras, consistently delivers the best accuracy across all environments' and that the Hesai outperforms the Livox. These claims are not supported by the presented data. The authors should either extend the evaluations to more of the seven environments or explicitly restrict the conclusions to the tested conditions.","section":"IV-D, IV-E, Figure 4"}],"minor_comments":[{"comment":"The text in Section IV-C says the results are 'compared against the total station (TPS) measurements', but Appendix B defines the reference as the Holistic Fusion (HF-Optimized) trajectory that fuses TPS, IMU, and IE-TC. Please align the wording to avoid ambiguity.","section":"IV-C vs Appendix B"},{"comment":"The term 'tactile-grade IMU' should be 'tactical-grade IMU' in Sections IV-F and IV-G3; Table II already uses the correct abbreviation 'Tact.'.","section":"IV-F, IV-G3, Table II"},{"comment":"The label 'Confided' on the LiDAR comparison subplot should be 'Confined'.","section":"Figure 4"},{"comment":"The table caption says 'HS indicates the output of Holistic-Fusion' but the text uses the abbreviation 'HF'; please standardize.","section":"Table VIII"},{"comment":"The sentence 'However, it lacks IEEE 1588v2 support (see Section VII)' appears to reference the wrong section; the discussion of the switch's lack of PTP support is in Section V-D2.","section":"Appendix I"}],"recommendation":"major_revision","confidential_remarks":"The circular GNSS evaluation and the confounded camera comparison are the two issues I would most want addressed before publication. Neither seems fatal to the hardware/cookbook contribution, but the quantitative modality claims in the abstract and Section IV need to be revised or supported by additional experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful engineering paper, and the GNSS row in Tables IV/V is partially self-referential, so treat those numbers as optimistic. The rest of the ablation story holds up well enough.\n\nWhat's new: Boxi is a curated payload with 2 LiDARs, 10 cameras, 7 IMUs, RTK GNSS, and a total-station prism, and the authors use it to run systematic ablations on one platform: time-sync offsets, extrinsic translation/rotation errors, IMU grades, camera shutter types, LiDAR types. Prior payload papers (VersaVIS, EverySync, PIRVS, FusionPortable) emphasized calibration and sync but didn't provide this kind of downstream state-estimation ablation. The cookbook is practical and explicit. They also ship open-source code and a time-sync verification tool. That's a real reproducible contribution.\n\nThe strong claims: DLIO is robust to time offsets under 2 ms and large translation offsets, but rotation calibration errors of a few degrees cause divergence, so prioritize rotation extrinsic accuracy. That's a useful, actionable guideline for payload designers.\n\nSoft spots, in proportion: The biggest is the GNSS comparison. The ground truth in Appendix B fuses TPS, HG4930, and IE-TC poses. Tables IV/V score IE-TC against that fused reference, so the GNSS row is partly marking its own homework, especially during TPS dropouts where HF relies on IE-TC. Table VIII validates HF against TPS only on Excavation Site and explicitly excludes dropout periods, so the reported GNSS ATE/RTE are not independently established. This is a genuine flaw in that row; it doesn't sink the paper but should be fixed (e.g., report GNSS errors only when TPS is available, or build a hold-out GNSS reference).\n\nSecondary: the camera comparison confounds shutter type with interface (GMSL2 vs USB) and FoV overlap; DLIO was tuned while OKVIS2 was not; LiDAR and camera comparisons each use one or two environments. All of these are stated in the limitations section, which is to their credit, but they do weaken the quantitative precision. The qualitative conclusions still look safe.\n\nWho it's for: any lab designing a mobile sensor suite or evaluating state estimation. It deserves a serious referee. I'd send it out with requests for the GNSS re-analysis and, if feasible, extra tuning of OKVIS2 or at least a sensitivity statement.","headline":"A useful, openly documented payload paper with real ablation data, held back partly by a partially circular GNSS baseline and some algorithm-tuning confounds, but worth sending to review.","tokens_in":35880,"tokens_out":2495,"would_cite":true,"duration_ms":25388,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that rotational extrinsic calibration, not time synchronization or translation accuracy, is the limiting factor for tightly coupled LiDAR-inertial odometry, on the basis of controlled ablations on a real 7.1 kg…","keywords":["sensor payload design","time synchronization","extrinsic calibration","LiDAR-inertial odometry","state estimation","multimodal sensing","legged robots","ablation study"],"falsifier":"Mount a tightly coupled LiDAR-inertial odometry system with verified extrinsics on a robot, then apply a controlled 2-degree rotation error between the IMU and LiDAR frames in a feature-rich environment; the paper predicts the trajectory will diverge. If the estimate stays bounded over the run, the claimed criticality of rotational extrinsic calibration at that threshold is not universal.","tokens_in":34912,"feed_emoji":"🤖","tokens_out":8025,"duration_ms":74334,"temperature":0.7,"pith_summary":"Boxi is a 7.1 kg, tightly integrated sensor payload for quadruped robots, and this paper uses it to turn payload-design folklore into measured thresholds. Across seven real-world environments, the authors compare LiDAR, camera, inertial, and GNSS state estimation, then deliberately corrupt time synchronization and extrinsic calibration to find when performance breaks. The headline result is that a tightly coupled LiDAR-inertial odometry pipeline tolerates time offsets up to roughly 5 ms and translation offsets below 5 mm, but rotation offsets of about 2 degrees between LiDAR and IMU make it diverge. Sensor choice is also quantified: global-shutter cameras beat rolling-shutter cameras in visual-inertial odometry unless rolling-shutter effects are modeled, a long-range precision LiDAR outperforms a short-range low-cost one in open terrain, and IMU grade barely matters in tightly coupled LiDAR-inertial odometry while mattering during dead-reckoning gaps. The same evidence is distilled into a payload 'cookbook' of design rules for future systems.","feed_headline":"Rotation calibration errors break LiDAR-inertial odometry at 2 degrees","feed_subtitle":"Real-world runs show sub-5 ms time errors are safe, while a 2-degree rotation offset makes state estimation diverge.","key_machinery":"The load-bearing object is Boxi itself: a rigid, machined payload whose sensor mounts are positioned to sub-millimeter mechanical tolerance, whose extrinsics are verified by dedicated calibration procedures, and whose time synchronization is validated by a gyroscope-correlation tool. The argument is carried by controlled perturbation sweeps that inject known time delays, translation offsets, and rotation offsets into an otherwise well-calibrated tightly coupled LiDAR-inertial odometry pipeline, isolating how each design quantity degrades trajectory accuracy.","core_discovery":"On Boxi's seven test environments, the paper establishes that, for a tightly coupled LiDAR-inertial odometry system, the extrinsic rotation between IMU and LiDAR is the most fragile calibration quantity: errors as small as 2 degrees cause the estimator to diverge, whereas time-delay errors up to approximately 5 ms and translation offsets below 5 mm leave performance nearly unchanged. The paper attributes the time tolerance to the 100 Hz IMU's 10 ms measurement period, and the rotation sensitivity to the long-range LiDAR, where far points amplify angular mismatch. It further reports that swapping a long-range precision LiDAR for a lower-cost short-range model measurably increases error in open spaces, that rolling-shutter cameras without shutter compensation degrade visual-inertial odometry, and that IMU grade makes little difference in tightly coupled LiDAR-inertial odometry but a large difference during dead-reckoning gaps.","pith_inferences":["Inference: The same sweep, repeated on a pipeline with online extrinsic calibration, would likely shift the rotation threshold upward; the 2-degree figure is a property of a calibration-fixed pipeline.","Inference: The time-offset tolerance likely scales with IMU rate; a 500 Hz IMU may demand synchronization below 2 ms, so the paper's 5 ms budget should not be read as universal.","Inference: A cost model for payloads can be derived from these thresholds, treating calibration accuracy as a hard constraint and sensor grade as a soft one.","Inference: The cookbook's 'keep it simple' advice has a testable consequence: a reduced autonomy version with fewer IMUs and one compute board should achieve similar LiDAR-inertial accuracy, a claim the authors state as future work."],"forward_implications":["A payload team should invest in rotation calibration hardware and procedure before pursuing sub-millisecond time synchronization or millimeter-level translation accuracy.","For tightly coupled LiDAR-inertial odometry in feature-rich environments, a low-cost consumer IMU can replace a tactical-grade IMU without meaningful trajectory error; the expensive IMU pays off only when other references drop out.","Rolling-shutter cameras need per-line exposure modeling or hardware triggering; naively substituting them for global-shutter cameras in visual-inertial odometry degrades accuracy significantly.","In open or featureless terrain, a longer-range, higher-accuracy LiDAR is worth its cost; in confined spaces the gap shrinks.","These thresholds give concrete design budgets: time synchronization below 5 ms and translation extrinsics below 5 mm are sufficient for the tested LiDAR-inertial pipeline."],"supporting_citations":[{"why":"Supplies the tightly coupled LiDAR-inertial odometry pipeline on which all time and extrinsic perturbation ablations run.","marker":"[14]"},{"why":"Supplies the multi-camera visual-inertial odometry system used to compare camera types and configurations.","marker":"[41]"},{"why":"Supplies the kinematic-inertial estimator used as the proprioceptive baseline in the modality comparison.","marker":"[5]"},{"why":"Supplies the post-processed GNSS-inertial solution that is both scored as a modality and used inside the ground-truth fusion.","marker":"[2]"},{"why":"Supplies the factor-graph fusion that combines total-station, IMU, and GNSS into the reference trajectory.","marker":"[51]"},{"why":"Supplies the intensity-based LiDAR-to-camera extrinsic calibration that gives Boxi its sub-millimeter LiDAR-camera extrinsics.","marker":"[25]"},{"why":"Supplies the Allan-variance analysis used to characterize each IMU's noise for camera-IMU calibration.","marker":"[8]"},{"why":"Supplies the gyroscope-correlation method extended into the paper's IMU time-synchronization verification tool.","marker":"[23]"}],"fun_headline_variants":["2° rotation error makes LiDAR-inertial odometry diverge","Time sync errors up to 5 ms harmless; 2° rotation is fatal","Why 2° rotation breaks LiDAR-inertial odometry but 5 ms sync is fine","Long-range LiDAR substitution degrades odometry in open spaces","Rolling shutter without compensation hurts visual-inertial odometry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ablation thresholds are measured against a reference trajectory assembled by fusing total-station, IMU, and post-processed GNSS, and the GNSS solution being scored is also an input to that reference, so the comparisons assume the reference is accurate enough to expose differences of a few centimeters.","fun_headline_variants_meta":{"raw":{"variants":["2° rotation error makes LiDAR-inertial odometry diverge","Time sync errors up to 5 ms harmless; 2° rotation is fatal","Why 2° rotation breaks LiDAR-inertial odometry but 5 ms sync is fine","Long-range LiDAR substitution degrades odometry in open spaces","Rolling shutter without compensation hurts visual-inertial odometry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000814,"raw_usage":{"total_tokens":3608,"prompt_tokens":1026,"completion_tokens":2582,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":2482}},"tokens_in":642,"tokens_out":2582,"duration_ms":18753,"temperature":1.0,"reasoning_tokens":2482,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:14:46.260434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Mount a tightly coupled LiDAR-inertial odometry system with verified extrinsics on a robot, then apply a controlled 2-degree rotation error between the IMU and LiDAR frames in a feature-rich environment; the paper predicts the trajectory will diverge. If the estimate stays bounded over the run, the claimed criticality of rotational extrinsic calibration at that threshold is not universal.","supporting_citations":[{"cited_title":"Direct lidar-inertial odometry: Lightweight lio with continuous- time motion correction","cited_arxiv_id":null,"evidence_quote":"Supplies the tightly coupled LiDAR-inertial odometry pipeline on which all time and extrinsic perturbation ablations run."},{"cited_title":"Twist-n-sync: Software clock synchronization with microseconds accuracy using mems-gyroscopes","cited_arxiv_id":null,"evidence_quote":"Supplies the gyroscope-correlation method extended into the paper's IMU time-synchronization verification tool."}],"review_version":1}