{"id":"b4572c38-7113-4e7f-8ead-56870fcc38b4","arxiv_id":"2412.02506","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"ROVER is a 39-recording, multi-season, multi-sensor benchmark dataset for visual SLAM in park and garden environments, with benchmarks showing poor performance of most SLAM systems in low-light and high-vegetation conditions.","lead":"This paper introduces ROVER, a new outdoor dataset for testing visual SLAM systems in parks and gardens across all seasons and lighting conditions. It also benchmarks seven SLAM algorithms and finds that most struggle in low light and dense vegetation, especially in summer and autumn.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth accuracy is asserted, not demonstrated: 2–5 Hz total-station fixes, linearly interpolated to camera rate and transferred via CAD extrinsics, are the sole reference for ATE/RPE; unvalidated GT errors could change the reported rankings.","rationale":"The reader's weakest assumption is precisely the load-bearing point: the total-station prism positions, after CAD extrinsics and NTP synchronization, are assumed to yield ground truth accurate enough to support the reported ATE/RPE values and rankings. I agree with this assessment. The paper presents the Leica TS16's static accuracy specification as if it transfers directly to a dynamic, vibrating robot observed at 2–5 Hz, then interpolated to camera rate. No repeatability study, no independent reference, and no error analysis of the interpolation or lever-arm transform is provided. This matters because the headline benchmark numbers are close: the best configurations differ by tens of centimeters in mATE and small fractions of a meter in mRPE, while 10–20 cm of GT interpolation error is a realistic magnitude on uneven terrain. The seasonal and lighting conclusions are less exposed, since they compare the same method across conditions and the trends are large, but the absolute errors and rankings are not. I do not recommend REJECT: a dataset resource can remain valuable even if the numeric benchmark needs revision, and the central qualitative finding about low-light and high-vegetation difficulty is plausible and likely robust. CONDITIONAL, as the reader concluded, is the appropriate verdict: the paper should be accepted only with validation of the ground-truth pipeline and, if needed, corrected numbers. This concern does not move the verdict; it sharpens the condition already stated.","tokens_in":25196,"tokens_out":4747,"duration_ms":58390,"concrete_test":"Regenerate groundtruth.txt from the raw TS16 fixes (or, if raw fixes are unavailable, from the released ground truth by keeping only the original 2–5 Hz samples) using a physically motivated motion model instead of linear interpolation. Then rerun the Evo-based ATE/RPE pipeline of Section V-B. If any mATE or mRPE entry in Tables VII–XII changes by more than the smallest inter-method gap in its column, or if the summer/autumn versus winter/spring gap reverses for any method, the end-to-end GT pipeline is not validated and the benchmark numbers require revision. A complementary check is to place the robot statically at several surveyed prism positions and compare the released groundtruth.txt at camera timestamps against an independent second survey, quantifying CAD-extrinsic and synchronization errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's quantitative claims in Sections V-B and V-C rest entirely on groundtruth.txt. Section III-B3 describes a Leica TS16 tracking a prism at 2–5 Hz with 'millimeter-level precision', Section III-B4 aligns sensors to the prism using a CAD model, and Section III-C2 time-synchronizes via GNSS/NTP. The paper never validates the end-to-end chain: dynamic tracking accuracy of a prism mounted on a vibrating UGV, CAD-model extrinsics versus actual mount geometry, residual time offsets, and linear interpolation from 2–5 Hz fixes to 30 Hz camera timestamps. At 0.5 m/s, consecutive TS16 fixes are 0.1–0.25 m apart; during turns or on uneven terrain, linear interpolation can easily introduce decimeter-level errors. The best mATE values in Tables VII–XII are roughly 0.3–1.5 m, and several inter-method gaps are of the same magnitude as plausible GT error (e.g., ORB-SLAM3 RGBD 1.26 m vs. DROID-SLAM 2.32 m; OpenVINS 1.43 m vs. SVO 1.70 m). RPE per meter is even more exposed, since a few centimeters of local GT error on a one-meter baseline corresponds to several percent error. Without independent validation, 'millimeter-accurate' is a sensor specification, not a demonstrated property of the released trajectories, and the performance ordering could be partly an artifact of the reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ROVER is a multi-season visual SLAM dataset collected with a small UGV at five park and garden locations, providing 39 recordings over approximately 7.2 km and 311 minutes with monocular (PiCam), stereo/RGBD (RealSense D435i, T265), and inertial (internal and VN100) data. The paper describes the robotic platform, sensor calibration, time synchronization, dataset organization, and a benchmark of seven SLAM algorithms across supported sensor configurations, reporting ATE, RPE, and success rate after five trials per sequence. The main empirical claims are that RGBD and stereo-inertial configurations generally outperform monocular ones and that most systems degrade substantially in summer and autumn and under low-light conditions.","tokens_in":25491,"tokens_out":8546,"duration_ms":88736,"significance":"If the ground-truth trajectories are as accurate as claimed, ROVER would be a valuable complement to existing outdoor SLAM datasets: it combines multi-season coverage, multiple camera modalities, consistent perimeter scenarios across five semi-structured locations, and a public release of raw data with ROS conversion tooling. The benchmark also has a reproducible core, using open-source SLAM implementations with pretrained weights, Umeyama alignment, and the evo toolbox for ATE/RPE. The main scientific value, however, depends on the end-to-end accuracy of the total-station reference and on the statistical support for the reported performance orderings, both of which need strengthening.","major_comments":[{"comment":"The 'millimeter-accurate' claim is based on the Leica TS16 sensor specification, but the end-to-end reference trajectory is not validated. The TS16 tracks a prism at 2-5 Hz, and the released groundtruth.txt is obtained by linear interpolation to camera timestamps and by CAD-model extrinsics from the prism to each sensor. At 0.5 m/s, consecutive fixes are 0.1-0.25 m apart, so interpolation during turns or on uneven terrain can easily produce decimeter-level errors. The reported mATE values in Tables VII-XII are about 0.3-1.5 m for the best configurations, and several inter-method gaps (e.g., ORB-SLAM3 RGBD 1.26 m vs. DROID-SLAM 2.32 m in Table X; OpenVINS 1.43 m vs. SVO Pro 1.70 m in Table IX) are of the same order as the plausible ground-truth error. The paper should provide an independent validation of the reference (e.g., static repeatability tests, loop-closure residuals on the repeated perimeter rounds, or comparison with a higher-rate reference after careful alignment), or explicitly report per-trajectory ground-truth uncertainty and temper the 'millimeter-accurate' wording.","section":"III-B3, III-B4, III-C2, V-B"},{"comment":"Each sequence is run five times, but the paper reports only aggregated mATE/mRPE and success rate. Because the trajectory filter removes failed runs, the reported means are computed over a varying, selection-dependent subset of trials, and no spread (standard deviation, median, or per-trial values) is given. Without this information, it is impossible to tell whether the seasonal and lighting differences in Tables XI and XII exceed run-to-run variability. Please report per-sequence and per-trial results, or at least medians with interquartile ranges and the number of valid runs underlying each table cell.","section":"V-B"},{"comment":"The validity criteria (at least 80% temporal coverage and at least one pose per second) are arbitrary and do not include any accuracy or divergence condition, so a trajectory that drifts badly for up to 20% of the run is still counted as successful. The success-rate metric is therefore a coarse proxy for robustness, and the reported SR values should be accompanied by a sensitivity analysis or a justification of these thresholds.","section":"V-B"},{"comment":"The ground-truth reference contains only 3-DoF prism positions and no orientation, so the reported ATE/RPE metrics cannot evaluate rotational error. The paper does not state the pose relation used with evo (e.g., translation-only), so readers cannot know what the mRPE values include. Please specify this explicitly and discuss the impact on the comparison.","section":"III-B3, V-B"}],"minor_comments":[{"comment":"Table V lists '/t265/image_left' twice; the second entry should presumably be '/t265/image_right'. The IMU rates also appear inconsistent with Table II (D435i listed as 100/200 Hz in Table II but 300 Hz in Table V; T265 as 65/200 Hz but 265 Hz in Table V).","section":"Table V"},{"comment":"Several mRPE entries in Table VII use comma decimal separators (4,24; 1,73; 2,15) while the rest of the table uses decimal points; please make the formatting uniform.","section":"Table VII"},{"comment":"The contributions state 'multi-weather' coverage, but Table IV records only windy/sunny/cloudy/dusk/night/light conditions and no rain, snow, or fog; either add such sequences or soften the 'multi-weather' claim.","section":"Abstract and Introduction"},{"comment":"The qualitative trajectory plots in Figure 9 have no axis labels or scale bars, making it difficult to judge the reported drift magnitudes; please add them.","section":"Figure 9"},{"comment":"The sentence 'Following Zhang and Scaramuzza [70], SLAM methods rely solely on monocular systems, the trajectories are scaled...' is ungrammatical and should be rewritten.","section":"V-A"},{"comment":"The header of Table I is difficult to parse and the meaning of the x marks is not fully defined; please add a legend or restructure the columns so that the comparison is transparent.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The dataset fills a genuine gap and the benchmark is useful, but the lack of end-to-end ground-truth validation is a load-bearing weakness. The authors should be given the opportunity to add validation experiments and per-trial statistics; if they cannot provide such validation, the 'millimeter-accurate' claim and the performance rankings will need to be substantially qualified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: ROVER is a real contribution, not a repackaging. It offers 39 recordings across all four seasons, day/dusk/night lighting, monocular/stereo/RGBD plus IMUs, in five semi-structured park/garden locations, with total-station ground truth and open data and code. That combination genuinely does not exist in the cited prior work, and the benchmark does standard, reproducible things: Umeyama alignment, ATE/RPE from evo, success rate from TartanAir, open-source implementations with pretrained weights. The qualitative finding—most SLAM systems degrade in low-light and high-vegetation conditions, especially summer and autumn—is plausible and consistent with experience.\n\nThe soft spot is the ground truth. The paper calls the dataset \"millimeter-accurate,\" but that is the total station's spec, not a demonstrated property of the released trajectories. The chain is: 2–5 Hz prism fixes, linear interpolation to 30 Hz, sensor-to-prism transforms from a CAD model, and NTP synchronization. None of that is validated end-to-end. At 0.5 m/s, consecutive fixes are 0.1–0.25 m apart, so interpolation error alone could plausibly reach decimeters around turns. Several inter-method gaps in the tables are the same magnitude. That does not trivialize the dataset, but it does mean the precise ATE rankings should not be over-read. There are also no error bars or per-trial statistics even though each sequence was run five times, and the lighting and seasonal coverage is unbalanced across locations.\n\nI would push back on one thing in the stress-test note: the situation is not that the benchmark is meaningless. The broad conclusions are probably robust. But the authors need to either validate the ground truth or soften the quantitative claims. A simple check—a few sequences with a second reference, or at least a discussion of expected interpolation error—would strengthen the paper considerably.\n\nBottom line: send it to review. It is a dataset paper with shipped artifacts, and that is exactly what referees should evaluate. I would ask for ground-truth validation and per-trial stats before accepting. I am unlikely to cite it myself, but I would bring it to a reading group if anyone is working on outdoor SLAM.","headline":"A genuinely useful dataset for visual SLAM in gardens and parks, with a benchmark whose exact numbers are undercut by unvalidated total-station ground truth; still deserves peer review.","tokens_in":26040,"tokens_out":2409,"would_cite":false,"duration_ms":24549,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ROVER, a millimeter-accurate multi-season dataset of 39 park-and-garden recordings, shows that most visual SLAM systems degrade sharply in low light and dense vegetation, especially in summer and autumn.","keywords":["visual SLAM","multi-season dataset","benchmark","outdoor robotics","park and garden environments","low-light navigation","visual-inertial odometry","ground truth trajectory"],"falsifier":"An independent check would be to re-record a handful of ROVER routes with a second positioning system of equal or better precision and compare its trajectory to the released ground truth; if the disagreement reaches the same size as the accuracy gaps between top-ranked SLAM configurations, the rankings would not be supported.","tokens_in":1541,"feed_emoji":"🤖","tokens_out":1528,"duration_ms":73046,"temperature":0.7,"pith_summary":"This paper introduces ROVER, a benchmark dataset for visual SLAM (simultaneous localization and mapping with cameras) recorded in five outdoor campus, park, and garden locations across all four seasons and several lighting conditions, from daylight to night with artificial light. The authors claim the dataset is millimeter-accurate because robot positions come from a survey instrument tracking a reflector on the platform, and they use it to evaluate seven established visual and visual-inertial SLAM systems. Their central result is that most of these systems degrade sharply in low light and in dense vegetation, with the worst trajectory errors occurring in summer and autumn. If the benchmark is trustworthy, it provides a reusable testbed for long-term outdoor localization and shows that current visual SLAM methods still do not generalize well to semi-structured natural environments.","feed_headline":"Most visual SLAM systems fail in low light and dense greenery","feed_subtitle":"A millimeter-accurate park-and-garden dataset across all seasons gives researchers a testbed for outdoor localization.","key_machinery":"The carrying mechanism is the recording platform and its ground-truth chain: a small lawnmower-like robot with monocular, stereo fisheye, and RGBD cameras plus internal and external inertial sensors, time-synchronized via satellite-disciplined network time, with three-dimensional position ground truth provided by a robotic total station (a survey instrument that tracks a reflector to millimeter accuracy) at 2-5 Hz and transformed to each sensor through CAD-based extrinsics. The benchmark side is carried by a standardized evaluation protocol: trajectories are filtered for coverage and temporal density, aligned with a least-squares transformation (similarity transform for monocular, rigid transform for scale-observable configurations), and scored by root-mean-square absolute trajectory error, relative pose error per meter, and success rate averaged over five trials per sequence. This combination lets the authors attribute performance differences to season and lighting rather than to sensor setup.","core_discovery":"The paper's central claim is that ROVER is a millimeter-accurate, multi-season visual SLAM benchmark covering 39 recordings, 7.2 km of trajectory, and roughly 1 TB of sensor data, and that the accompanying evaluation demonstrates a clear environmental robustness gap: stereo-inertial and RGBD configurations perform well in favorable lighting and moderate vegetation, but most SLAM systems produce large absolute trajectory errors and lower success rates in low-light and high-vegetation conditions, particularly in summer and autumn. The benchmark also finds that adding inertial data to RGBD configurations does not help and can hurt, and that monocular systems suffer from scale ambiguity in larger open areas. The authors present the dataset as addressing a gap left by existing agricultural, forest, and urban multi-season datasets, which rarely combine monocular, stereo, and RGBD sensing with year-round park-and-garden recordings.","pith_inferences":["Beyond the paper: the reported seasonal degradation suggests a concrete stress test for future SLAM systems, namely scoring separately on the summer and autumn subset, since averaging across seasons can hide failure modes.","Beyond the paper: before using ROVER to rank algorithms, a user should verify the claimed millimeter accuracy on a few sequences with an independent reference, because the accuracy gaps between top systems are small enough that unvalidated ground-truth interpolation could influence the ordering.","Beyond the paper: the dataset's repeated routes across seasons could be repurposed for place recognition and re-localization experiments, not only trajectory evaluation.","Beyond the paper: if the lighting results generalize, a testable prediction is that active illumination plus exposure-adaptive preprocessing will improve feature-based SLAM at night more than adding more inertial sensors."],"forward_implications":["If ROVER is millimeter-accurate, it can serve as a common testbed for comparing visual SLAM systems in semi-structured outdoor environments, complementing indoor and urban benchmarks.","The finding that most systems degrade at night and in summer or autumn implies that robustness to low light and dense vegetation should be a first-class evaluation criterion for outdoor SLAM.","The result that stereo-inertial and RGBD configurations outperform monocular ones in these environments supports prioritizing depth-capable or inertial-fused setups for park-and-garden robots.","The observation that RGBD-inertial performs worse than RGBD-only suggests that adding IMU data is not a guaranteed improvement in this setting.","The repeated overlapping rounds of the Perimeter recording scenario enable multi-session mapping and loop-closure evaluation across seasons."],"supporting_citations":[{"why":"Provides the multi-sensor calibration toolbox used to obtain camera intrinsics, IMU biases, and camera-IMU extrinsics for the dataset.","marker":"[54]"},{"why":"Defines the ATE/RPE evaluation protocol and the timestamp-to-file format that ROVER adopts for its image and ground-truth files.","marker":"[55]"},{"why":"Supplies the least-squares alignment procedure used to map estimated trajectories onto ground truth before computing errors.","marker":"[69]"},{"why":"Justifies using a similarity transform for monocular trajectories to handle scale ambiguity and a rigid transform for scale-observable configurations.","marker":"[70]"},{"why":"Introduces the success-rate metric and trajectory filtering criteria that the benchmark adapts to quantify robustness across trials.","marker":"[30]"},{"why":"An open-source feature-based SLAM system whose mono, stereo, and RGBD configurations provide the main traditional baselines and the lowest overall mATE in the RGBD setting.","marker":"[58]"},{"why":"A deep dense-optical-flow SLAM system whose RGBD configuration performs best at night with external light and is a key deep-learning baseline.","marker":"[60]"},{"why":"A filter-based visual-inertial system whose stereo-inertial configuration shows the most consistent year-round performance in the benchmark.","marker":"[56]"}],"fun_headline_variants":["New SLAM benchmark: park dataset reveals seasonal failures","ROVER dataset: 39 recordings show SLAM's weak spots","Low light and dense vegetation trip up visual SLAM systems","Adding inertial data to RGBD SLAM doesn't help, dataset shows"],"cache_read_input_tokens":28160,"weakest_assumption_plain":"The whole benchmark ranking depends on the assumption that the ground-truth path, measured by a survey instrument tracking a reflector and then shifted to each camera's position, stays accurate enough after being interpolated to camera times to support error differences that are sometimes only a few centimeters apart.","fun_headline_variants_meta":{"raw":{"variants":["New SLAM benchmark: park dataset reveals seasonal failures","ROVER dataset: 39 recordings show SLAM's weak spots","Low light and dense vegetation trip up visual SLAM systems","Adding inertial data to RGBD SLAM doesn't help, dataset shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000941,"raw_usage":{"total_tokens":4055,"prompt_tokens":1010,"completion_tokens":3045,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":2974}},"tokens_in":626,"tokens_out":3045,"duration_ms":23229,"temperature":1.0,"reasoning_tokens":2974,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:22:21.200656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent check would be to re-record a handful of ROVER routes with a second positioning system of equal or better precision and compare its trajectory to the released ground truth; if the disagreement reaches the same size as the accuracy gaps between top-ranked SLAM configurations, the rankings would not be supported.","supporting_citations":[{"cited_title":"Unified temporal and spatial calibration for multi-sensor systems,","cited_arxiv_id":null,"evidence_quote":"Provides the multi-sensor calibration toolbox used to obtain camera intrinsics, IMU biases, and camera-IMU extrinsics for the dataset."},{"cited_title":"A tutorial on quantitative trajectory evaluation for visual(-inertial) odometry,","cited_arxiv_id":null,"evidence_quote":"Justifies using a similarity transform for monocular trajectories to handle scale ambiguity and a rigid transform for scale-observable configurations."},{"cited_title":"Tartanair: A dataset to push the limits of visual slam,","cited_arxiv_id":null,"evidence_quote":"Introduces the success-rate metric and trajectory filtering criteria that the benchmark adapts to quantify robustness across trials."}],"review_version":1}