REVIEW 12 cited by
One Million Scenes for Autonomous Driving: ONCE Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected data and incrementally self-training powerful recognition models have received increasing attention and may become the solutions of next-generation industry-level powerful and robust perception models in autonomous driving. However, the research community generally suffered from data inadequacy of those essential real-world scene data, which hampers the future exploration of fully/semi/self-supervised methods for 3D perception. In this paper, we introduce the ONCE (One millioN sCenEs) dataset for 3D object detection in the autonomous driving scenario. The ONCE dataset consists of 1 million LiDAR scenes and 7 million corresponding camera images. The data is selected from 144 driving hours, which is 20x longer than the largest 3D autonomous driving dataset available (e.g. nuScenes and Waymo), and it is collected across a range of different areas, periods and weather conditions. To facilitate future research on exploiting unlabeled data for 3D detection, we additionally provide a benchmark in which we reproduce and evaluate a variety of self-supervised and semi-supervised methods on the ONCE dataset. We conduct extensive analyses on those methods and provide valuable observations on their performance related to the scale of used data. Data, code, and more information are available at https://once-for-auto-driving.github.io/index.html.
Forward citations
Cited by 12 Pith papers
-
Orbis 2: A Hierarchical World Model for Driving
A hierarchical driving world model — planning in compressed DINO space at 2 Hz and rendering detailed frames at 10 Hz — achieves state-of-the-art long-horizon stability, steering response, and representation quality.
-
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.
-
ToosiCubix: Monocular 3D Cuboid Labeling via Vehicle Part Annotations
A monocular annotation method estimates vehicle position, orientation, and dimensions from user clicks on parts like wheels and badges, with accurate up-to-scale 8DoF results but limited full 9DoF accuracy.
-
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.
-
LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments
LiDARDustX releases 30,000 annotated LiDAR frames, over 80% of them dust-affected, and shows that dust degrades 3D detection average precision by 9.9 to 18.1 points across six detectors.
-
CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
Post-episode multi-agent debriefing lets LLM driving agents learn concise natural-language coordination protocols that avoid collisions and merge traffic, and distillation makes the policy fast enough for near-real-time use.
-
An interactive enhanced driving dataset for autonomous driving
A fused, interaction-labeled dataset of 7.31M driving segments with synthetic BEV videos and VQA pairs for training/evaluating driving VLMs.
-
Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
A single universal adversarial image perturbation can route different input semantics to different attacker-defined outputs in multimodal LLMs, with up to 66% success over five targets.
-
High-Fidelity Digital Twins for Bridging the Sim2Real Gap in LiDAR-Based ITS Perception
A digital twin of a real intersection can generate LiDAR training data that matches the target location, and a detector trained on it reported 4.8% higher car AP than a model trained on real data, though with more syn...
-
Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey
A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...
-
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
A closed-loop LLM-agent framework that auto-generates research ideas and code, reported to improve baseline performance on all 12 tasks it was tested on.
-
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.
Discussion (0). Sign in to comment.