Pith. sign in

REVIEW 17 cited by

Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.10581 v2 pith:7OKT3673 submitted 2022-11-19 cs.CV

classification cs.CV
keywords detectionmethodssparsefeaturesdifferentmulti-viewsparse4danchor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bird-eye-view (BEV) based methods have made great progress recently in multi-view 3D detection task. Comparing with BEV based methods, sparse based methods lag behind in performance, but still have lots of non-negligible merits. To push sparse 3D detection further, in this work, we introduce a novel method, named Sparse4D, which does the iterative refinement of anchor boxes via sparsely sampling and fusing spatial-temporal features. (1) Sparse 4D Sampling: for each 3D anchor, we assign multiple 4D keypoints, which are then projected to multi-view/scale/timestamp image features to sample corresponding features; (2) Hierarchy Feature Fusion: we hierarchically fuse sampled features of different view/scale, different timestamp and different keypoints to generate high-quality instance feature. In this way, Sparse4D can efficiently and effectively achieve 3D detection without relying on dense view transformation nor global attention, and is more friendly to edge devices deployment. Furthermore, we introduce an instance-level depth reweight module to alleviate the ill-posed issue in 3D-to-2D projection. In experiment, our method outperforms all sparse based methods and most BEV based methods on detection task in the nuScenes dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 30 citations worldwide. Full citation record

  1. 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

    cs.CV 2025-07 conditional novelty 7.0 of 10

    3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.

  2. Kerr-Schild Double Copy of the Randall-Sundrum Black String

    hep-th 2026-04 unverdicted novelty 6.0 of 10

    Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.

  3. AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

    cs.RO 2026-01 unverdicted novelty 6.0 of 10

    Conditioning speed planning on the predicted path and relabeling synthetic cut-ins yields SOTA Bench2Drive scores (DS 89.07, SR 73.18%).

  4. BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A contrastive learning framework with instance-level and perspective-level losses consistently improves multiple BEV detection models on nuScenes by up to 2.4 mAP.

  5. CoST: Efficient Collaborative Perception From Unified Spatiotemporal Perspective

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CoST unifies multi-agent and multi-time fusion into a single spatio-temporal space and transmits only dynamic object features, improving collaborative 3D detection accuracy while reducing bandwidth.

  6. Monocular Semantic Scene Completion via Masked Recurrent Networks

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Decomposing monocular semantic scene completion into a coarse stage plus a masked recurrent refinement network improves NYUv2 and SemanticKITTI completion and semantic IoU over prior monocular methods.

  7. OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception

    cs.CV 2025-07 conditional novelty 6.0 of 10

    OnlineBEV achieves state-of-the-art 3D object detection on nuScenes by recurrently fusing bird's eye view features with motion-guided deformable attention and a heatmap consistency loss.

  8. Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LiMA distills long-term multi-camera image features into LiDAR backbones and reports consistent gains on segmentation and detection benchmarks.

  9. Coherent Online Road Topology Estimation and Reasoning with Standard-Definition Maps

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Score jointly detects lane segments, road boundaries, and traffic elements, estimates lane topology, and associates traffic elements with lanes, using SD map priors and temporal fusion to reach state-of-the-art on Ope...

  10. ODG: Occupancy Prediction Using Dual Gaussians

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ODG uses separate static and dynamic Gaussian query sets, refined coarse-to-fine, plus rendering supervision, and reports state-of-the-art occupancy prediction on Occ3D-nuScenes and Occ3D-Waymo.

  11. DriveCamSim: Generalizable Camera Simulation via Explicit Camera Modeling for Autonomous Driving

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DriveCamSim uses explicit 3D-aware attention to generate multi-view driving video under new camera parameters and frame rates, trained on 2Hz nuScenes data.

  12. DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A distillation framework with a ground-truth-annotation teacher, RL status optimization, and generative distribution interaction improves end-to-end planning collisions and closed-loop scores.

  13. Occupancy Learning with Spatiotemporal Memory

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ST-Occ improves 3D occupancy prediction for self-driving by storing a compact scene-level memory of past frames and conditioning current predictions on it with uncertainty-aware attention, gaining 3 mIoU over prior st...

  14. Achieving Precise and Reliable Locomotion with Differentiable Simulation-Based System Identification

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    Estimating robot dynamics parameters from trajectory data alone inside a differentiable simulator, inside the reinforcement learning loop, is claimed to improve trajectory following in bipedal locomotion.

  15. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    World4Drive couples multiple driving intentions with a latent world model to generate, score, and select trajectories, reporting state-of-the-art perception-free planning on nuScenes and NavSim.

  16. OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Adding object-centric radiance-field rendering and height-aware opacity attention to the DualBEV detector improves camera-only 3D object detection on nuScenes by up to 2.0 mAP points.

  17. DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DySS combines state-space feature learning with dynamic query merging and pruning to improve both accuracy and speed for camera-based 3D detection on nuScenes.

Pith tools