Pith. sign in

REVIEW 2 cited by

Multi-Camera Calibration Free BEV Representation for 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.17252 v1 pith:IPHP3EQY submitted 2022-10-31 cs.CV

classification cs.CV
keywords cameraattentioncalibrationparametersrepresentationachievescamera-drivencomputation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In advanced paradigms of autonomous driving, learning Bird's Eye View (BEV) representation from surrounding views is crucial for multi-task framework. However, existing methods based on depth estimation or camera-driven attention are not stable to obtain transformation under noisy camera parameters, mainly with two challenges, accurate depth prediction and calibration. In this work, we present a completely Multi-Camera Calibration Free Transformer (CFT) for robust BEV representation, which focuses on exploring implicit mapping, not relied on camera intrinsics and extrinsics. To guide better feature learning from image views to BEV, CFT mines potential 3D information in BEV via our designed position-aware enhancement (PA). Instead of camera-driven point-wise or global transformation, for interaction within more effective region and lower computation cost, we propose a view-aware attention which also reduces redundant computation and promotes converge. CFT achieves 49.7% NDS on the nuScenes detection task leaderboard, which is the first work removing camera parameters, comparable to other geometry-guided methods. Without temporal input and other modal information, CFT achieves second highest performance with a smaller image input 1600 * 640. Thanks to view-attention variant, CFT reduces memory and transformer FLOPs for vanilla attention by about 12% and 60%, respectively, with improved NDS by 1.0%. Moreover, its natural robustness to noisy camera parameters makes CFT more competitive.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection

    cs.CV 2026-08 conditional novelty 6.0 of 10

    NCGR adds a learned gated 2D offset to spatial-cross-attention sampling points, raising nuScenes NDS from 0.280 to 0.397 under five-camera extrinsic perturbation while preserving clean accuracy.

  2. Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    REO performs calibration-free 3D semantic occupancy prediction with vanilla cross-attention, auxiliary 2D/3D tasks, and query-based decoding, reporting large speedups and state-of-the-art benchmark numbers.

Pith tools