REVIEW 4 major objections 6 minor 2 cited by
GO: The Great Outdoors Multimodal Dataset
T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper introduces the GO dataset, which it claims is the most comprehensive multimodal off-road dataset yet, pairing six sensor types with semantic labels and centimeter-precision GPS traces.
desk verdict GO dataset pairs thermal and radar with standard off-road sensors, but the 'most comprehensive' claim stumbles on a NIR/thermal mix-up and missing quality metrics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the GO dataset itself, defined by a single mobile robot platform carrying a 64-channel LiDAR, front-facing stereo cameras, a rear monocular camera, a thermal camera, a 360-degree millimeter-wave radar, an IMU/GPS unit, and an RTK GPS receiver accurate to about 1.4 cm. The mechanism that carries the comprehensiveness claim is the co-registration pipeline: Precision Time Protocol (PTP) synchronizes the cameras, LiDAR, and radar to a common high-precision clock, while the MSG-Cal multi-sensor graph-based calibration aligns LiDAR with the RGB and thermal cameras using a planar board, and radar is aligned with the LiDAR point cloud through a 3D model. Because the transforms are in place, semantic labels drawn on RGB images can be projected onto the thermal and LiDAR streams, converting separate sensors into one multimodal dataset.
What would settle it
A reader could falsify the central claim by downloading any route and measuring LiDAR-to-camera reprojection error and timestamp offsets; median reprojection error beyond a few pixels or inter-sensor drift beyond one frame's time would show the 'well-calibrated' premise fails, and the absence of NIR files in the download would show the table's NIR entry is unsupported.
Extended reading notes
Core claim
The paper's central claim is stated in the abstract: 'This dataset provides the most comprehensive set of data modalities and annotations compared to existing off-road datasets.' In the authors' comparison, GO is the only dataset marked for every sensor column—camera, stereo, LiDAR, NIR, radar, IMU, and GPS—while also carrying semantic labels. They document five teleoperated routes totaling 10.26 km and 98.60 minutes over paved areas, gravel trails, and forest terrain, with a 22-class ontology assembled from the RELLIS-3D and RUGD label sets. The annotation process is semi-automated: Segment Anything and OffSeg provide initial masks, human annotators correct boundaries and assign labels, and the RGB-thermal calibration is used to transfer labels to thermal imagery. The paper further claims centimeter-precision RTK GPS ground truth for trajectory evaluation.
Load-bearing premise
The claim that GO is the most comprehensive off-road dataset depends on the assumption that the sensors were actually calibrated and synchronized to the stated quality and that the semi-automated labels are accurate, but the paper reports no quantitative measurements of calibration error, synchronization latency, or label correctness.
Editorial extensions
If this is right
- Researchers can train and evaluate semantic segmentation, object detection, and SLAM on a single off-road platform whose LiDAR, camera, thermal, and radar data are time-synchronized and spatially aligned.
- The inclusion of thermal and radar data opens a direct route to studying perception in dust, fog, smoke, darkness, and other conditions where RGB cameras and LiDAR degrade.
- The centimeter-precision RTK GPS traces give odometry and SLAM algorithms a quantitative ground truth across 10.26 km of trails and forest, enabling direct drift measurement.
- The 22-class ontology drawn from RELLIS-3D and RUGD means models trained on GO can be evaluated against those earlier datasets without re-labeling.
- The repeated and overlapping routes, with Route 3 and Route 4 covering the same trail on different days, provide a built-in test for robustness to day-to-day appearance change.
Reading between the lines
- If the timing and calibration hold up in practice, the highest-value use of GO is probably multimodal fusion under degraded visibility—training radar-and-thermal traversability models and benchmarking them against camera-and-LiDAR baselines on the same routes.
- The paper's comparison table lists NIR as a GO modality while the sensor description names only a thermal camera; until the released files are inspected, users should not assume true near-infrared imagery exists.
- All five routes were collected in May at a single site, so season and weather diversity are absent; re-driving Route 2 and Route 4 across seasons would be a natural test of how well models trained on GO transfer to unseen conditions.
- The paper reports no quantitative calibration error, synchronization latency, or label-accuracy numbers, so the dataset's usability for cross-modal label transfer should be validated by users measuring reprojection error and label agreement on held-out frames.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Great Outdoors (GO) dataset, a multimodal off-road robotics dataset collected with a Clearpath Warthog platform. The dataset claims to include six sensor modalities (monocular/stereo RGB, thermal, LiDAR, radar, IMU/GPS) plus RTK GPS traces and 22-class semantic annotations, gathered over five routes totaling 10.26 km and 98.60 minutes. The abstract and introduction make the central claim that GO provides 'the most comprehensive set of data modalities and annotations compared to existing off-road datasets,' operationalized in Table I by check marks showing GO uniquely contains every listed modality. The paper also describes the calibration/synchronization pipeline and the semi-automated annotation process using SAM and OffSeg with human refinement.
Significance. If the claims are substantiated, GO would be a genuinely valuable community resource: it appears to be the only off-road dataset combining camera, stereo, thermal, LiDAR, radar, and IMU/GPS with semantic labels and RTK GPS traces, and it is publicly downloadable. The route design across trails, forest, and paved areas, with variable speeds and revisited trails, supports a range of benchmarks in segmentation, SLAM, and sensor fusion. The paper is clear in describing the sensor setup, route statistics, and annotation ontology, and it provides qualitative visualizations. However, the paper's headline contributions of 'high-quality annotations' and 'centimeter-precision GPS traces' are not backed by quantitative evaluation, and there is a direct inconsistency between the sensor list and Table I's NIR entry. These issues affect the central comprehensiveness and quality claims and require correction before the dataset can be reliably used by the community.
major comments (4)
- [Table I and Section II.A] There is a direct inconsistency between Table I and the sensor list. Table I gives GO a check mark under 'NIR', but Section II.A lists only an RGB camera (FLIR Blackfly S), a stereo pair (FLIR Blackfly S), and a thermal camera (FLIR Boson 640). The FLIR Boson 640 is a longwave infrared (LWIR) camera, not a near-infrared (NIR) camera, and no NIR camera is described anywhere in the sensor setup. Section III then refers to 'radar and NIR sensor data,' compounding the confusion. Since the abstract's 'most comprehensive' claim is operationalized by Table I, this is not a cosmetic wording issue. The authors must either confirm that an NIR stream exists in the released data and describe it, or remove the NIR check for GO and revise the comprehensiveness claim accordingly.
- [Section II.D] The contribution 'High-Quality Semantic Annotations' is not supported by any quantitative evidence. Section II.D describes a semi-automated pipeline using SAM and OffSeg with human refinement, but no annotation accuracy numbers, inter-annotator agreement, per-class IoU values on a held-out test set, or counts of annotated keyframes are reported. The class distribution in Fig. 5 shows imbalance but not quality. Without such validation, the claim of 'high-quality' annotations is an assertion, and users cannot assess whether the labels are suitable for benchmark use. The authors should add quantitative annotation statistics and, ideally, benchmark results for semantic segmentation on a split of the dataset.
- [Section II.B] The calibration and synchronization section is purely qualitative. It states that PTP is used for radar, LiDAR, and cameras, and that the thermal camera and GNSS/IMU use ROS timestamps, but it reports no calibration errors (e.g., reprojection error for the LiDAR-camera or camera-camera extrinsics), no temporal synchronization accuracy, and no verification of the claimed RTK accuracy. In particular, the abstract's 'centimeter-precision GPS traces' relies on the RTK receiver's maximum accuracy specification (1.4 cm), not on any evaluation of the actual collected trajectories. The authors should provide quantitative calibration results and an assessment of the GPS trajectory accuracy in the collected routes.
- [Section II.C] The dataset is described as 'large-scale' but no raw data counts are given. The paper reports total distance and duration (10.26 km, 98.60 minutes) and per-route times and speeds, but not the number of RGB images, stereo pairs, thermal frames, LiDAR scans, radar scans, or labeled keyframes. Without these numbers, the 'large-scale' claim is not verifiable and the dataset cannot be compared against other resources in terms of volume. Please include a table or paragraph listing the per-modality data counts and total storage size.
minor comments (6)
- [Fig. 2] Figure 2 has two items labeled '(e)': one for the thermal camera and one for the LiDAR/radar overlay. Please relabel the thermal image as (d) and the overlay as (e) or adjust the captions accordingly.
- [Section II.A] The sensor list uses inconsistent capitalization for LiDAR (e.g., 'LIDAR' appears in Table I and in the sensor bullet) and 'Robot' is written as 'Warthog mobile robot' while the platform is later called 'the Warthog.' Please standardize the terminology.
- [Section II.B] The phrase 'threshold filtered Radar data' is used in the Fig. 2 caption but the threshold is not defined anywhere in the paper. Brief clarification of what thresholding was applied would help.
- [References] Several references are incomplete, e.g., [1] (RELLIS-3D) lacks publication venue and year, and [12] (Segment Anything) lacks the conference/journal details. Please complete all bibliography entries.
- [Section II.A] The radar bullet says '400 input rotations per cycle' but does not explain what a 'cycle' means in this context; please clarify or remove the phrase.
- [Section III] In Section III, the phrase 'radar and NIR sensor data' again conflates thermal with NIR; this wording should be aligned with the actual sensor suite.
Circularity Check
No circular derivation; the dataset's contribution is an empirical comparison, not a result derived from its own assumptions.
full rationale
The GO paper is a dataset introduction with no fitted quantities, predictive model, or uniqueness theorem. Its central claim—'most comprehensive set of data modalities and annotations compared to existing off-road datasets'—is operationalized in Table I by listing publicly documented sensor modalities of GO and seven other datasets; this is an external comparison, not a quantity derived from the paper's own equations or fitted parameters. The self-citations (RELLIS-3D, RUGD, OffSeg) are used for ontology design and annotation initialization, and do not serve as the evidence for the comprehensiveness claim. Although Table I marks NIR for GO while Section II.A lists only a thermal (FLIR Boson 640) camera and no NIR camera, and Section III also refers to 'NIR sensor data,' this is an internal inconsistency about what the released data contain; it affects correctness of the comparison, but it is not circular. No step reduces to its own input by construction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The semi-automated annotation pipeline (SAM and OffSeg with human refinement) yields high-quality semantic labels.
- domain assumption PTP and ROS timestamp synchronization align all sensors with sufficient accuracy for multi-modal fusion.
- domain assumption The RTK GPS receiver provides centimeter-precision ground truth trajectories.
Cite this review
Pith. "Pith review of GO: The Great Outdoors Multimodal Dataset." pith.science (2026). https://pith.science/paper/QE572NRA
@misc{pith2026250119274,
author = {Pith},
title = {Pith review of: GO: The Great Outdoors Multimodal Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/QE572NRA}},
note = {Machine review of arXiv:2501.19274}
}
read the original abstract
The Great Outdoors (GO) dataset is a multi-modal annotated data resource aimed at advancing ground robotics research in unstructured environments. Existing off-road datasets often lack sensor diversity and exclude vital modalities like thermal and radar that are critical for operation in degraded conditions (e.g., low visibility or adverse weather). To address these gaps, we introduce a large-scale multimodal off-road dataset with six complementary sensor modalities, along with semantic annotations and GPS traces, to support tasks such as semantic segmentation, object detection, and SLAM. The diverse environmental conditions represented in the dataset present significant real-world challenges, which provide opportunities to develop more robust solutions to support the continued advancement of field robotics, autonomous exploration, and perception systems in natural environments. The dataset can be downloaded at: https://www.unmannedlab.org/the-great-outdoors-dataset/
Figures
Forward citations
Cited by 2 Pith papers
-
Learning Traversability-Aware Global Planners for Long Horizon Off-Road Navigation
Overhead multi-modal learning with PU human-trajectory supervision and LiDAR priors yields global off-road costmaps that nearly match human path length and sharply cut interventions versus local planners.
-
UAVScenes: A Multi-Modal Dataset for UAVs
UAVScenes adds frame-wise image and LiDAR semantic labels, reconstructed 6-DoF poses, and 3D maps to 120k frames of the MARS-LVIG dataset, with six benchmark tasks.
Reference graph
Works this paper leans on
- [1]
-
[2]
P. Mortimer, R. Hagmanns, M. Granero, T. Luettel, J. Petereit, and H.-J. Wuensche. The GOOSE Dataset for Perception in Unstructured Environments
-
[3]
R. Hagmanns, P. Mortimer, M. Granero, T. Luettel, and J. Petereit. Exca- vating in the Wild: The GOOSE-Ex Dataset for Semantic Segmentation
-
[4]
M. Sivaprakasam, P. Maheshwari, M. G. Castro, S. Triest, M. Nye, S. Willits, A. Saba, W. Wang, and S. Scherer. TartanDrive 2.0: More Modalities and Better Infrastructure to Further Self-Supervised Learning Research in Off-Road Driving Tasks
- [5]
-
[6]
M. Gadd, D. D. Martini, O. Bartlett, P. Murcutt, M. Towlson, M. Widojo, V . Mus ¸at, L. Robinson, E. Panagiotaki, G. Pramatarov, M. A. K ¨uhn, L. Marchegiani, P. Newman, and L. Kunze. OORD: The Oxford Offroad Radar Dataset
- [7]
-
[8]
P. Mortimer and H.-J. Wuensche. TAS-NIR: A VIS+NIR Dataset for Fine-grained Semantic Segmentation in Unstructured Outdoor Environ- ments
Show all 14 references
-
[9]
Mortimer and M
P. Mortimer and M. Maehlisch. Survey on Datasets for Perception in Unstructured Outdoor Environments
-
[10]
MSG-Cal: Multi-sensor graph-based calibration,
J. L. Owens, P. R. Osteen, and K. Daniilidis, “MSG-Cal: Multi-sensor graph-based calibration,” in 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 3660–3667
2015
-
[11]
A RUGD Dataset for Autonomous Navigation and Visual Perception in Unstructured Outdoor Environments,
M. Wigness, S. Eum, J. G. Rogers, D. Han, and H. Kwon, “A RUGD Dataset for Autonomous Navigation and Visual Perception in Unstructured Outdoor Environments,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 5000–5007
2019
-
[12]
Segment Anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Girshick, “Segment Anything,” pp. 4015–4026
-
[13]
Offseg: A semantic segmentation framework for off-road driving,
K. Viswanath, K. Singh, P. Jiang, P. B. Sujit, and S. Saripalli, “Offseg: A semantic segmentation framework for off-road driving,” in 2021 IEEE 17th International Conference on Automation Science and Engineering (CASE). IEEE, pp. 354–359
2021
-
[14]
D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine. ViNT: A Foundation Model for Visual Navigation. V. B IOGRAPHY SECTION Peng Jiang is a Postdoctoral Researcher in the Mike Walker ’66 Department of Mechanical Engineering at Texas A&M University...
2021
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.