REVIEW 4 cited by
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras provide rich texture information and LiDAR offers precise 3D spatial data--relying on a single modality often leads to performance limitations. This paper introduces MV2DFusion, a multi-modal detection framework that integrates the strengths of both worlds through an advanced query-based fusion mechanism. By introducing an image query generator to align with image-specific attributes and a point cloud query generator, MV2DFusion effectively combines modality-specific object semantics without biasing toward one single modality. Then the sparse fusion process can be accomplished based on the valuable object semantics, ensuring efficient and accurate object detection across various scenarios. Our framework's flexibility allows it to integrate with any image and point cloud-based detectors, showcasing its adaptability and potential for future advancements. Extensive evaluations on the nuScenes and Argoverse2 datasets demonstrate that MV2DFusion achieves state-of-the-art performance, particularly excelling in long-range detection scenarios.
Forward citations
Cited by 4 Pith papers
-
Distilling Multi-modal Large Language Models for Autonomous Driving
DiMA jointly trains a vision-only planner with an LLM and uses auxiliary language, reconstruction, and scene editing tasks to improve planning on nuScenes while dropping the LLM at inference.
-
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
EVT achieves state-of-the-art 75.3% NDS on the nuScenes test set by using LiDAR-guided adaptive sampling and projection for image-to-BEV transformation, along with a geometry-aware transformer decoder.
-
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
RoCA, a Gaussian-process codebook over ego and agent tokens, improves cross-domain generalization and adaptation of end-to-end autonomous driving models without extra inference cost.
-
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
TimeAlign uses Swin-LSTM prediction and camera-guided combination of past and observed LiDAR BEV features to partially recover 3D detection accuracy under LiDAR time lag.
Discussion (0). Continue with ORCID to comment.