REVIEW 17 cited by
Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Bird-eye-view (BEV) based methods have made great progress recently in multi-view 3D detection task. Comparing with BEV based methods, sparse based methods lag behind in performance, but still have lots of non-negligible merits. To push sparse 3D detection further, in this work, we introduce a novel method, named Sparse4D, which does the iterative refinement of anchor boxes via sparsely sampling and fusing spatial-temporal features. (1) Sparse 4D Sampling: for each 3D anchor, we assign multiple 4D keypoints, which are then projected to multi-view/scale/timestamp image features to sample corresponding features; (2) Hierarchy Feature Fusion: we hierarchically fuse sampled features of different view/scale, different timestamp and different keypoints to generate high-quality instance feature. In this way, Sparse4D can efficiently and effectively achieve 3D detection without relying on dense view transformation nor global attention, and is more friendly to edge devices deployment. Furthermore, we introduce an instance-level depth reweight module to alleviate the ill-posed issue in 3D-to-2D projection. In experiment, our method outperforms all sparse based methods and most BEV based methods on detection task in the nuScenes dataset.
Forward citations
Cited by 17 Pith papers
-
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.
-
Kerr-Schild Double Copy of the Randall-Sundrum Black String
Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.
-
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
Conditioning speed planning on the predicted path and relabeling synthetic cut-ins yields SOTA Bench2Drive scores (DS 89.07, SR 73.18%).
-
BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
A contrastive learning framework with instance-level and perspective-level losses consistently improves multiple BEV detection models on nuScenes by up to 2.4 mAP.
-
CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective
CoST unifies multi-agent and multi-time fusion into a single spatio-temporal space and transmits only dynamic object features, improving collaborative 3D detection accuracy while reducing bandwidth.
-
Monocular Semantic Scene Completion via Masked Recurrent Networks
Decomposing monocular semantic scene completion into a coarse stage plus a masked recurrent refinement network improves NYUv2 and SemanticKITTI completion and semantic IoU over prior monocular methods.
-
OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception
OnlineBEV achieves state-of-the-art 3D object detection on nuScenes by recurrently fusing bird's eye view features with motion-guided deformable attention and a heatmap consistency loss.
-
Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations
LiMA distills long-term multi-camera image features into LiDAR backbones and reports consistent gains on segmentation and detection benchmarks.
-
Coherent Online Road Topology Estimation and Reasoning with Standard-Definition Maps
Score jointly detects lane segments, road boundaries, and traffic elements, estimates lane topology, and associates traffic elements with lanes, using SD map priors and temporal fusion to reach state-of-the-art on Ope...
-
ODG: Occupancy Prediction Using Dual Gaussians
ODG uses separate static and dynamic Gaussian query sets, refined coarse-to-fine, plus rendering supervision, and reports state-of-the-art occupancy prediction on Occ3D-nuScenes and Occ3D-Waymo.
-
DriveCamSim: Generalizable Camera Simulation via Explicit Camera Modeling for Autonomous Driving
DriveCamSim uses explicit 3D-aware attention to generate multi-view driving video under new camera parameters and frame rates, trained on 2Hz nuScenes data.
-
DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model
A distillation framework with a ground-truth-annotation teacher, RL status optimization, and generative distribution interaction improves end-to-end planning collisions and closed-loop scores.
-
Occupancy Learning with Spatiotemporal Memory
ST-Occ improves 3D occupancy prediction for self-driving by storing a compact scene-level memory of past frames and conditioning current predictions on it with uncertainty-aware attention, gaining 3 mIoU over prior st...
-
Achieving Precise and Reliable Locomotion with Differentiable Simulation-Based System Identification
Estimating robot dynamics parameters from trajectory data alone inside a differentiable simulator, inside the reinforcement learning loop, is claimed to improve trajectory following in bipedal locomotion.
-
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
World4Drive couples multiple driving intentions with a latent world model to generate, score, and select trajectories, reporting state-of-the-art perception-free planning on nuScenes and NavSim.
-
OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving
Adding object-centric radiance-field rendering and height-aware opacity attention to the DualBEV detector improves camera-only 3D object detection on nuScenes by up to 2.0 mAP points.
-
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
DySS combines state-space feature learning with dynamic query merging and pruning to improve both accuracy and speed for camera-based 3D detection on nuScenes.
Discussion (0). Continue with ORCID to comment.