{"id":"545e21de-5f69-4ee1-b3ff-9ec682592191","arxiv_id":"2507.12920","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MoCap2GT jointly optimizes MoCap and IMU data using cubic B-splines on SE(3) with a variable time offset to generate SLAM ground truth trajectories that meet targets of ATE/ARE below 2 mm/0.2 deg and RTE/RRE below 0.4 mm/0.02 deg.","lead":"The paper presents MoCap2GT, an estimator that fuses motion capture and IMU data to produce higher precision ground truth trajectories for SLAM benchmarking. The method adds SE(3) B-spline interpolation with a variable time offset and a robust initializer, and the authors report meeting stricter rotation and inter-frame accuracy targets than existing tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-world validation is a single short-range trajectory against an unverified monocular-BA reference; the Section I-A accuracy claim is therefore not yet established.","rationale":"I read the paper as a system paper whose contribution is the joint MoCap/IMU estimator and its validation. The optimization formulation is a standard MLE framework with a correctly implemented cumulative cubic B-spline on SE(3): the basis matrix in Eq. 20 is the transpose of the cumulative coefficient matrix and yields the expected basis functions. The linear initializer is carefully constructed, and the simulated experiments are useful for relative comparison. The strongest claim, however, is an empirical 'first to meet targets' statement, and that statement is only as strong as the real-world reference used to support it. The reader's weakest-assumption analysis correctly identifies the soft spot: the theoretical conversion from 0.1 px reprojection error to 0.05 mm / 0.003 deg is not a substitute for an independent reference, and the physical range of the validation is extremely small. I would not reject the paper on this basis, because the method is plausible and the source code is promised, but the verdict should remain conditional. The concrete check proposed here, an independent reference and repeated trials over a longer, better-excited trajectory, would settle whether the accuracy targets are genuinely met with the claimed margins. Since my concern aligns with the reader's, the verdict is unchanged.","tokens_in":14160,"tokens_out":13345,"duration_ms":145707,"concrete_test":"Re-run the Section V-B real-world evaluation on a trajectory of at least several meters with rotations above the 10 deg / 5 s degeneracy threshold, and compare MoCap2GT against an independent reference, such as a second MoCap system, a laser tracker, or an encoder-instrumented motion stage, instead of the monocular-BA reference. Repeat the experiment at least five times, report mean and standard deviation for ATE, ARE, RTE, and RRE, and check whether the upper 95% confidence bounds remain below the Section I-A targets. If they do, the concern is resolved; if not, the claim should be restricted to short-range motions or weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that MoCap2GT is the first method to meet ATE/ARE < 2 mm/0.2 deg and RTE/RRE < 0.4 mm/0.02 deg with a consumer-grade IMU and MoCap system—rests on Table III, which compares against a reference trajectory from monocular global bundle adjustment (Section V-B). That reference is assumed accurate from a 0.1 px reprojection error via a theoretical conversion to 0.05 mm and 0.003 deg at 1 m, but it is not independently verified; a global BA solution can carry correlated errors from camera intrinsics, board flatness, and corner detection that are not captured by a residual threshold. The validation workspace is also tiny: the Fig. 8 axes show a trajectory spanning roughly 4 cm x 10 cm x 3 cm. The 'sufficient motion' scenario therefore provides only limited excitation for estimating extrinsics, IMU biases, and the variable time-offset control points, which are exactly the quantities the method claims to improve. Only single values are reported, with no repeated trials; ARE (0.178 deg vs 0.2 deg target) and ATE (1.466 mm vs 2.0 mm target) have slim margins, so plausible reference error or run-to-run variation could change whether all four targets are met. The Section V-C benchmarking on public datasets is illustrative, not a validation of absolute GT accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MoCap2GT, a batch maximum-likelihood estimator that fuses 6-DoF marker-based MoCap pose measurements with IMU readings from the device under test to produce a high-precision ground-truth trajectory for SLAM benchmarking. The method consists of a linear initialization stage that recovers extrinsic rotation, translation, gravity, velocities, and initial time offset using screw-theory robust kernels and RANSAC, followed by nonlinear optimization over a factor graph. The MoCap factor is modeled via cubic B-spline interpolation on SE(3) with a time offset that is itself represented as a linear B-spline with free control points, allowing for clock-drift compensation. A degeneracy-aware rejection strategy gates the MoCap factor's contribution to spatiotemporal calibration parameters in low-rotation windows. The authors claim that, using a consumer-grade IMU and MoCap system, MoCap2GT is the first method to satisfy the accuracy targets ATE/ARE < 2 mm/0.2 deg and RTE/RRE < 0.4 mm/0.02 deg. Validation includes simulated data with controlled noise and a real-world experiment against a monocular bundle-adjustment reference, plus benchmarking applications on TUM-VI and EuRoC.","tokens_in":14458,"tokens_out":5123,"duration_ms":62197,"significance":"If the central accuracy claim is upheld, the paper would make a useful contribution to SLAM benchmarking by enabling reliable evaluation of rotation and inter-frame errors, which existing MoCap-based ground-truth pipelines do not provide at the stated precision. The formulation is well grounded: it builds on standard IMU preintegration and continuous-time B-splines, and the variable time-offset model is a sensible response to clock-scale drift. The manuscript also provides an open-source implementation and uses concrete, falsifiable accuracy targets. The main weakness is at the validation layer: the real-world experiment is a single short-range trajectory with a reference that is not independently calibrated, and the simulation uses an ORB-SLAM3 trajectory as the 'true' motion. These gaps, together with empirically set thresholds, leave the headline claim not yet firmly established.","major_comments":[{"comment":"The claimed first demonstration of meeting the Section I-A accuracy targets rests on a single real-world trajectory compared against a bundle-adjustment reference whose accuracy is inferred from a 0.1 px reprojection error via a theoretical conversion (0.05 mm and 0.003 deg at a viewing distance of 1 m). This conversion does not account for correlated error sources in monocular global BA, such as biases in camera intrinsics, calibration-board flatness, and corner detection. The trajectory shown in Fig. 8 spans only a few centimeters, and only single trial values are reported in Table III. With ATE margin of about 0.53 mm and ARE margin of about 0.022 deg, an independent reference (for example, a laser tracker or a second, higher-fidelity MoCap system) or repeated trials at multiple ranges and motion types is needed before the absolute-accuracy claim is established.","section":"Section V-B, Table III, Fig. 8"},{"comment":"The simulation uses an ORB-SLAM3 trajectory as the basis ground truth for generating synthetic IMU and MoCap measurements. Since an ORB-SLAM3 estimate has its own unknown trajectory error, the simulation cannot establish the absolute accuracy of MoCap2GT; it can only compare relative performance and noise robustness. To support the central claim that the method meets the specific error targets, the simulator should be driven by a trajectory with known ground truth, such as an analytically generated spline motion or a real trajectory measured by an independent high-accuracy device. The current simulation is therefore not sufficient evidence for the headline accuracy claim.","section":"Section V-A"},{"comment":"Several key parameters are set empirically without a reported sensitivity analysis: the kernel amplification factor mu = 5 in Eq. (10), the degeneracy window size w = 5 s, and the rotation threshold theta = 10 deg in Eq. (25). Since the method's robustness to motion degeneracy is presented as a contribution, it is important to show how the four output error metrics vary with these parameters, or to provide a data-driven justification for the chosen values. Without this, the reader cannot tell whether the single successful real-world result in Table III depends on parameters tuned to that particular short-range motion.","section":"Section IV-B3, Eq. (10), Eq. (25)"}],"minor_comments":[{"comment":"The paper explicitly acknowledges that the real-world setup limits the testable motion range to keep the board in view, but the conclusion states without qualification that the method meets the accuracy targets. The limitation should be reflected in the conclusion, or the authors should add experiments that cover longer and more varied trajectories.","section":"Section V-B"},{"comment":"The visualization of Jacobian directional errors would benefit from a colorbar and axis labels; in the current figure it is not clear how the 'directional error' is computed or what the color scale represents.","section":"Fig. 5"},{"comment":"The typographic device of green highlighting for target-satisfying cells will not be visible in grayscale print; please use a pattern or an explicit marker, and also state the margins with respect to the targets in the caption.","section":"Table III"},{"comment":"The cubic B-spline interpolates the raw MoCap positions, which still contain measurement noise; the claim that it 'reduces MoCap jitter' depends on the smoothing behavior of the spline basis. It would be helpful to report the residual statistics between the fitted spline and the raw MoCap data to quantify how much smoothing is actually applied.","section":"Section IV-B2, Eq. (19)"}],"recommendation":"major_revision","confidential_remarks":"The claim of being the first method to meet the stated accuracy targets is hard to verify from the manuscript alone; a direct numerical comparison with Vicon2GT in the same experimental setup would help. The paper fits the journal's scope, and there is no indication of problematic citation practices beyond the expected self-reference in the initialization module (Shu et al., 2024). The main concern is that the validation is too thin to support the headline claim as it stands."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a competent engineering contribution: robust linear initialization with screw-theory kernels in RANSAC, cubic B-spline SE(3) MoCap factors with variable time offset, and degeneracy-aware rejection. That combination is genuinely new relative to Vicon2GT, Kalibr-M, and HEC-Filter. The simulated experiments are reasonable, and the calibration-consistency comparison on TUM-VI is a fair way to rank methods when no reliable GT exists. The code is promised, which matters for reproducibility.\n\nThe soft spot is the real-world validation. The central claim—first method to meet ATE/ARE < 2 mm/0.2 deg and RTE/RRE < 0.4 mm/0.02 deg with consumer hardware—rests on Table III, and that table compares against a monocular global-BA reference whose 0.05 mm/0.003 deg accuracy is derived from a 0.1 px reprojection threshold, not verified. The trajectory spans only a few centimeters in Fig. 8. That is a short-range, low-excitation setup, which is exactly the regime where spatiotemporal calibration is hardest to validate. The margins to the targets are also slim: ARE is 0.178 deg against a 0.2 deg bound, and only single trials are reported. I do not think the accuracy claim is false; I think it is unproven. The paper itself acknowledges the limited range, which is honest, but the conclusion overstates what the evidence supports.\n\nAlso, several thresholds are empirically set (mu, w, theta, number of time-offset control points) without sensitivity analysis. The simulation uses an ORB-SLAM3 trajectory as ground truth for generating synthetic data, which is fine, but it is not independent validation.\n\nOn the positive side, the math is standard MLE factor-graph optimization, no circularity beyond mild self-citation in the initializer, and the comparison to existing methods is fair. The paper is clearly written and the method is sensible.\n\nWho is this for? Researchers building or using MoCap-based SLAM benchmarks. They will get a practical estimator and a clear description of the known limitations of existing GT pipelines. It deserves a serious referee, but the referee should push for stronger real-world validation: larger trajectories, repeated trials, and ideally an independent reference such as a laser tracker or a second MoCap system. I would accept it for peer review, not because the headline claim is proven, but because the work is useful and the validation gap is fixable.","headline":"Solid engineering paper with a plausible but unproven headline accuracy claim; the method is a sensible combination of known tools, worth reviewing, but the real-world validation is too thin to certify the 2 mm/0.2 deg targets.","tokens_in":14967,"tokens_out":645,"would_cite":false,"duration_ms":8327,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MoCap2GT claims that fusing marker-based motion capture with the device's own IMU produces ground-truth trajectories accurate enough to benchmark rotation and inter-frame SLAM errors, meeting target errors below 2 mm/0.2 deg and 0.4…","keywords":["motion capture","inertial measurement unit","ground truth trajectory estimation","SLAM benchmarking","spatiotemporal calibration","cubic B-spline on SE(3)","factor graph optimization","MoCap jitter mitigation"],"falsifier":"Collect a new sequence where a room-scale trajectory is measured simultaneously by the MoCap+IMU setup and by an independent reference with known sub-0.05 mm accuracy, such as a laser tracker; if MoCap2GT's ATE/ARE/RTE/RRE against that reference exceed the stated targets, the central claim is falsified.","tokens_in":13972,"feed_emoji":"🎯","tokens_out":8963,"duration_ms":92177,"temperature":0.7,"pith_summary":"MoCap2GT is a batch estimator that fuses marker-based motion-capture poses with the device's own IMU readings to produce the ground-truth trajectory for benchmarking visual-inertial SLAM. The paper's goal is to remove the two errors that make rotation and inter-frame metrics unreliable: spatiotemporal misalignment between the MoCap and the IMU clock/frame, and high-frequency MoCap jitter. By jointly estimating the extrinsic transform, a variable time offset, gravity alignment, and the IMU trajectory, the estimator outputs an IMU-centric trajectory that is drift-free and smooth. On real data collected with a consumer-grade IMU and MoCap system, the method reports ATE/ARE below 2 mm/0.2 deg and RTE/RRE below 0.4 mm/0.02 deg, the targets the paper sets as necessary for benchmarking state-of-the-art SLAM.","feed_headline":"Fusing MoCap and IMU hits SLAM ground-truth accuracy targets","feed_subtitle":"Joint optimization clears 2 mm/0.2 deg and 0.4 mm/0.02 deg error targets that official MoCap ground truth misses.","key_machinery":"The load-bearing object is a cumulative cubic B-spline pose parameterization on the $\\mathrm{SE}(3)$ manifold with a variable time offset, embedded as MoCap factors in a factor-graph maximum-likelihood problem. The B-spline interpolates the discrete, jittery MoCap poses with $C^2$ continuity, so each IMU-state residual queries a smooth pose; the time offset is itself a linear B-spline whose control points are estimated states, absorbing clock-scale drift between the two devices. Convergence is prepared by a linear initializer that uses screw-theory weighting and RANSAC to recover extrinsic rotation, translation, velocity, and gravity alignment without good priors. A degeneracy detector computes the maximum rotation in sliding windows and suppresses gradients of the spatiotemporal calibration parameters during low-excitation motion, preventing ill-conditioned updates.","core_discovery":"The paper's central claim is that a joint maximum-likelihood estimator over MoCap and IMU measurements can simultaneously solve spatiotemporal calibration and MoCap jitter mitigation, so the resulting ground truth is accurate enough to score rotation and inter-frame SLAM errors. The authors position this as the first consumer-grade setup to meet their stated accuracy targets, reporting ATE of 1.466 mm, ARE of 0.178 deg, RTE of 0.177 mm, and RRE of 0.013 deg in the sufficient-motion real-world scenario, with the estimator staying within the target bounds even under motion degradation. They further show that re-benchmarking a state-of-the-art visual-inertial SLAM system against MoCap2GT instead of official dataset ground truth changes reported relative rotation errors dramatically, which they take as evidence that official MoCap-based GT in public datasets is not accurate enough for rotation and inter-frame evaluation.","pith_inferences":["The accuracy target is demonstrated only inside a short-range, calibration-board volume; whether the same error levels persist over meter-scale, high-speed motion typical of SLAM evaluation is not tested here and would be the natural extension.","Because the method needs only raw MoCap and IMU streams, it could be applied retroactively to any existing dataset that stores both, giving an immediate upgrade path for benchmark ground truth without new data collection.","The variable time-offset model suggests a general recipe for any sensor pair with unsynchronized free-running clocks: estimate the mapping between clocks as parameters of the trajectory representation rather than a single scalar.","A reader wanting independent confirmation could replace the camera-based reference with a laser tracker or a second MoCap system; if the reference uncertainty is not at least an order of magnitude below the claimed errors, the headline numbers should be read as an upper bound."],"forward_implications":["If MoCap2GT's accuracy holds, SLAM benchmarks can report rotation and inter-frame errors (ARE/RTE/RRE) with the same confidence they now give absolute translation error (ATE).","Existing public datasets that publish raw MoCap and IMU data can be re-processed to produce cleaner ground truth; the paper does this for two widely used datasets and shows that official ground truth with jitter inflates relative rotation error.","The estimator's variable time-offset model keeps accuracy when the MoCap and IMU clocks drift, a regime in which fixed-offset baselines degrade sharply.","A visual-inertial SLAM system re-benchmarked on MoCap2GT ground truth yields materially different error numbers (e.g., relative rotation error on one room sequence drops from 0.362 deg to 0.064 deg), implying comparisons published against noisy GT may need revisiting."],"supporting_citations":[{"why":"Supplies the cumulative B-spline on Lie groups used to interpolate MoCap poses with C^2 continuity; this is the core measurement model of the paper.","marker":"[30]"},{"why":"The baseline batch estimator (IMU preintegration plus linear MoCap interpolation) that MoCap2GT extends and outperforms; its linear interpolation is replaced by cubic B-splines.","marker":"[14]"},{"why":"Provides the IMU preintegration formulation and noise propagation that the IMU factors are built from.","marker":"[28]"},{"why":"Supplies the screw-theory constraints used to weight the linear equations in the robust initializer and to define degeneracy and inlier metrics.","marker":"[29]"},{"why":"One of the two public datasets providing raw MoCap and IMU data plus official ground truth; the paper re-benchmarks a visual-inertial SLAM system against its new ground truth on these sequences.","marker":"[9]"},{"why":"The other public dataset; its official ground truth exhibits jitter and its median-filtering baseline is part of the comparison.","marker":"[10]"},{"why":"The correlation-based coarse time alignment and hand-eye calibration approach that motivates the initializer and reports calibration inaccuracies in public dataset ground truth.","marker":"[22]"},{"why":"The continuous-time calibration baseline whose MoCap branch is compared against; it assumes a fixed time offset, which the paper's variable offset improves on.","marker":"[27]"}],"fun_headline_variants":["MoCap-IMU joint optimization beats official SLAM ground truth","SLAM ground truth: MoCap plus IMU outperforms MoCap alone","MoCap2GT: high-precision ground truth for SLAM benchmarking","Fusing IMU with MoCap fixes jitter and calibration for SLAM GT","MoCap-IMU fusion yields ground truth that beats MoCap-only"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation against real-world reference trajectories assumes that a global bundle adjustment with reprojection error below 0.1 pixels yields reference poses accurate to about 0.05 mm and 0.003 deg at 1 m viewing distance; that conversion is theoretical, and the small test volume may not represent the motions real SLAM benchmarks require.","fun_headline_variants_meta":{"raw":{"variants":["MoCap-IMU joint optimization beats official SLAM ground truth","SLAM ground truth: MoCap plus IMU outperforms MoCap alone","MoCap2GT: high-precision ground truth for SLAM benchmarking","Fusing IMU with MoCap fixes jitter and calibration for SLAM GT","MoCap-IMU fusion yields ground truth that beats MoCap-only"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000756,"raw_usage":{"total_tokens":3374,"prompt_tokens":974,"completion_tokens":2400,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":2302}},"tokens_in":590,"tokens_out":2400,"duration_ms":17030,"temperature":1.0,"reasoning_tokens":2302,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:34:28.407958+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a new sequence where a room-scale trajectory is measured simultaneously by the MoCap+IMU setup and by an independent reference with known sub-0.05 mm accuracy, such as a laser tracker; if MoCap2GT's ATE/ARE/RTE/RRE against that reference exceed the stated targets, the central claim is falsified.","supporting_citations":[{"cited_title":"Efficient derivative computation for cumulative B-splines on Lie groups,","cited_arxiv_id":null,"evidence_quote":"Supplies the cumulative B-spline on Lie groups used to interpolate MoCap poses with C^2 continuity; this is the core measurement model of the paper."},{"cited_title":"Vicon2GT: Derivations and analysis,","cited_arxiv_id":null,"evidence_quote":"The baseline batch estimator (IMU preintegration plus linear MoCap interpolation) that MoCap2GT extends and outperforms; its linear interpolation is replaced by cubic B-splines."},{"cited_title":"VINS-Mono: A robust and versatile monocular visual-inertial state estimator,","cited_arxiv_id":null,"evidence_quote":"Provides the IMU preintegration formulation and noise propagation that the IMU factors are built from."},{"cited_title":"Chess–calibrating the hand-eye matrix with screw constraints and synchronization,","cited_arxiv_id":null,"evidence_quote":"Supplies the screw-theory constraints used to weight the linear equations in the robust initializer and to define degeneracy and inlier metrics."},{"cited_title":"The EuRoC micro aerial vehicle datasets,","cited_arxiv_id":null,"evidence_quote":"One of the two public datasets providing raw MoCap and IMU data plus official ground truth; the paper re-benchmarks a visual-inertial SLAM system against its new ground truth on these sequences."},{"cited_title":"The TUM VI benchmark for evaluating visual-inertial odometry,","cited_arxiv_id":null,"evidence_quote":"The other public dataset; its official ground truth exhibits jitter and its median-filtering baseline is part of the comparison."},{"cited_title":"A spatiotemporal hand- eye calibration for trajectory alignment in visual (-inertial) odometry evaluation,","cited_arxiv_id":null,"evidence_quote":"The correlation-based coarse time alignment and hand-eye calibration approach that motivates the initializer and reports calibration inaccuracies in public dataset ground truth."},{"cited_title":"Extending kalibr: Calibrating the extrinsics of multiple IMUs and of individual axes,","cited_arxiv_id":null,"evidence_quote":"The continuous-time calibration baseline whose MoCap branch is compared against; it assumes a fixed time offset, which the paper's variable offset improves on."}],"review_version":1}