REVIEW 5 major objections 4 minor 72 references
NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection
T0 review · 5 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Gated 2D rectification of displaced reference points keeps BEV detection accurate when camera extrinsics are perturbed.
desk verdict The robustness gains are likely real, but the noise-conditional scalar cannot adapt at inference; the paper needs an augmentation-only baseline and a reinterpretation of its mechanism. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Gated reference-point rectification inside spatial cross-attention. The query-camera correction network G takes the BEV query and a camera-level condition vector, outputs one 2D offset per camera in normalized image coordinates, and the effective offset is delta* = g_i * delta, where g_i is the camera-level gate in [0,1]. The rectified anchor r + delta* is then passed to the untouched native deformable sampler. The second piece is the scheduling bridge: normalized rotation and translation perturbation magnitudes define training targets g_gt and q*, and during the middle of training the condition and gate are interpolated toward a learned camera scalar rho_i = 1 - sg(q_i), so at inference alp
What would settle it
Take the trained NCGR checkpoint, feed the validation set under clean extrinsics and under increasing perturbation bounds, and record the predicted camera scalar q_i and gate rho_i per camera. If q_i is statistically the same for clean and perturbed cameras, or if rho_i is constant, the blind per-camera adaptation claim is not realized. A complementary check: evaluate under rotation bounds beyond 15 degrees (for example 20–25 degrees); since the offset is bounded by s_delta = 0.1, the rectification should begin to fail and NDS should fall toward or below the baseline if the mechanism is genuin
Extended reading notes
Core claim
The central claim is that the damage from extrinsic noise in SCA-based BEV detection is a displaced base projection: the native deformable attention samples locally around a wrong anchor, so even content-adaptive offsets cannot recover the correct region. NCGR therefore rectifies the anchor itself. For each BEV query and camera, a shared correction network predicts a 2D offset bounded by a scaled tanh (s_delta = 0.1), a camera-level gate scales that offset, and the result is added to the reference point before the unmodified deformable sampler runs. The offset is trained through a teacher–student setup: the perturbed student branch with rectification is pushed toward the clean teacher branch
Load-bearing premise
The inference-time camera-level scalar, predicted from each camera's image feature, is assumed to carry useful per-camera gate and condition information even though the paper itself notes that perturbations modify projection metadata rather than image appearance; if that scalar does not respond to perturbation magnitude, blind inference becomes a fixed scale rather than per-camera adaptive rectification.
Editorial extensions
If this is right
- If the claim holds, SCA-based BEV detectors can tolerate dynamic multi-camera extrinsic drift without a separate calibration stage or LiDAR observations.
- The correction is anchor-level, not extrinsic-level: it adds about 0.7M parameters and roughly 5.7% latency (217.14 vs 205.40 ms/frame) rather than a full pose-estimation module.
- The advantage over CAPE grows as more cameras are perturbed and as rotation severity increases, with the largest NDS/mAP margins at 3–5 perturbed cameras.
- Clean-extrinsic performance stays near the baseline because the identity regularizer and the gate can suppress offsets on unperturbed cameras.
- Disabling the offset at inference drops NDS from 0.3969 to 0.3434 under the same five-camera stress test, tying the reported gain specifically to the rectification path.
Reading between the lines
- A natural next test is whether the learned camera scalar q_i reacts to perturbation magnitude per camera; if it stays constant, the blind-inference gate is a fixed scale and the per-camera adaptivity is not doing the work, although the bounded residual network could still help.
- Because the offset is bounded by s_delta = 0.1 and the training perturbation bounds are 15 degrees / 0.1 m, out-of-distribution shifts beyond those bounds would likely need a larger scale or iterative rectification; the reported sensitivity table already shows degradation at s_delta = 0.20.
- The same anchor-rectification idea should transfer to other projection-based view transformers, since it only changes where sampling is anchored, not the sampled features.
- One could test whether the BEV-consistency teacher is essential or whether the rectification alone, trained with the detection loss and perturbation-derived gate, would give most of the gain; the paper ablates only L_vanilla jointly.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes NCGR, an extension of BEVFormer that adds gated, query–camera 2D residual offsets to the SCA reference-point projection, together with a clean-teacher/perturbed-student BEV-consistency loss. During training, perturbation-derived condition and gate signals are gradually replaced, via a scheduled interpolation, by a learned camera-level scalar predicted from camera features, so that inference does not require perturbation metadata. On nuScenes with simulated dynamic and static extrinsic perturbations, NCGR reports large NDS gains over BEVFormer and CAPE under a five-camera stress test (0.3969 vs 0.2800 and 0.3323), while keeping clean-extrinsic NDS close to BEVFormer (0.520 vs 0.518). The paper includes controlled inference-time interventions, ablation of the consistency loss, evaluation-seed sensitivity, and efficiency measurements.
Significance. If the stated mechanism were correct, NCGR would be a practically useful and relatively lightweight way to make spatial-cross-attention detectors robust to camera extrinsic errors, a realistic failure mode for autonomous driving. The empirical package is unusually thorough: the perturbation protocol is specified, the offset-disabled and gate-closed interventions are internally consistent (Tables 2 and D.1), seed sensitivity is small (Table B.5), and runtime/memory overheads are quantified. However, the central conceptual claim—that the learned scalar provides per-camera, per-instance noise-conditional control at inference—is not supported by the evidence and is in fact contradicted by the paper's own statement that perturbations modify projection metadata rather than image appearance. The main results are therefore at risk of being explained by training-time perturbation augmentation plus query-content-based offsets rather than by adaptive rectification.
major comments (5)
- [Method Overview and Eqs. (4)–(7)] The learned scalar bq_i is supervised by Eq. (4) toward q*_i, a deterministic function of the true perturbation magnitudes, but its input F_i is the camera image feature, which is unaffected by extrinsic perturbations (as the paper concedes: “perturbations modify projection metadata rather than image appearance”). Hence the optimal predictor of q*_i given F_i is a conditional mean that does not depend on the actual perturbation realization; at inference, with α=1, c_i=ρ_i1_2 and g_i=ρ_i are therefore effectively constant per camera (or at most scene-dependent), not noise-conditional. The large NDS gain over BEVFormer cannot be attributed to per-camera adaptive rectification without further evidence. Please report the distribution of learned bq_i values across clean and perturbed cameras, and add a control in which bq_i is frozen to its training-set mean: if performance is unchanged, the
- [Tables 2 and D.1] The offset-disabled and gate-closed interventions only show that some nonzero effective offset is beneficial; they do not show that the offset is perturbation-corrective. Because the input to G is the BEV query and a constant c_i at inference, the offset δ_j,i is a query-content-dependent residual with no knowledge of the true displacement Δπ_i(p). The comparison with a model trained with the same perturbation augmentation but without the rectification branch (or with g_i fixed to a constant) is needed to separate the contribution of perturbation augmentation from the contribution of the learned offsets. Without such a baseline, the gain in the five-camera stress test is compatible with a fixed, content-based sampling modification rather than gated reference-point rectification.
- [Supplementary Section D.5, Fig. D.4] The layer-wise projection-error analysis is a direct test of the rectification mechanism, and it is weak: only Layer 6 reduces the mean distance to the clean reference (by 4.3%), while Layers 1–5 produce little reduction or small increases. If the main detection gain were caused by correcting displaced SCA base projections, one would expect a clearer and more consistent geometric effect across layers. The paper should report the cumulative effect of the multi-layer rectification path and explain why the large NDS improvement is accompanied by such a small and late projection-error reduction. As it stands, the evidence is more consistent with the offsets acting as an additional query-dependent sampling mechanism than with geometric correction.
- [Table B.4 and Experimental Setup] The rectification scale s_delta=0.10 is selected by sweeping s_delta on the same evaluation protocol used for the main comparison, and the sensitivity is large (e.g., clean NDS drops from 0.5199 to 0.3658 at s_delta=0.20). This is a form of test-set selection that can inflate the reported margins. The authors should either select s_delta on a held-out validation split or show that the main conclusions are stable across a range of s_delta values. The current seed-sensitivity analysis (Table B.5) is performed after this selection and does not address it.
- [Table 1 and Appendix C] The baselines BEVFormer and CAPE are trained with their standard configurations, while NCGR is trained with synthetic perturbations applied with probability 0.7. Since the comparison is specifically about robustness to perturbations, the absence of a perturbation-augmented BEVFormer baseline makes it impossible to separate the effect of NCGR's proposed components from the effect of simply training on the perturbation distribution. Please add a BEVFormer trained with the same perturbation augmentation (and, if possible, with the teacher–student consistency loss but without the rectification offsets) to isolate the contribution of the rectification module.
minor comments (4)
- [Experimental Setup] Typo: “mean! average orientation error” should be “mean average orientation error”.
- [Supplementary Table C.2] The caption “Static perturbation results…” is duplicated across the split table; the continuation should be labeled clearly.
- [Eq. (16)] The restored indices use δ^ℓ,b,j,i; check the placement of the ℓ superscript and whether the layer index is consistently defined across Eqs. (7)–(10) and (16).
- [Section D.3 / Fig. D.3] The discussion of nonzero corrections in the clean camera is reasonable but should be stated earlier and tied to the gate endpoint: if the gate g_i is zero for a clean camera, the effective offset δ*_j,i is zero regardless of δ_j,i; the visualization should clarify whether it shows raw or gated offsets.
Circularity Check
No significant circularity: the central claim is supported by external baselines and independent interventions; the camera-scalar limitation is a validation concern, not a circular derivation.
full rationale
No circularity found. The paper's robustness claim is established by matched external comparisons (BEVFormer, CAPE) on nuScenes under synthetic perturbations, and its internal ablations (offset disabled, gate closed, BEV-consistency removal) are independent tests of the mechanism. The closest thing to a circular construction is the camera scalar: Eq. (4) supervises q_i toward q*_i, which is derived from perturbation magnitudes, and at inference alpha=1 makes c_i and g_i functions of q_i. This is a supervised proxy, not a definitional equivalence. The paper explicitly concedes in the Method Overview that 'Because perturbations modify projection metadata rather than image appearance, the scalar provides a bounded operating point shaped by the training distribution, while the query--camera offset network predicts the spatial adjustment.' That concession is an identifiability/validation limitation for the 'noise-conditional' label -- the scalar may collapse to a constant -- but it does not make the derivation circular: the rectification offset is not fitted to the perturbation displacement, the BEV-consistency supervision uses a clean teacher, and the main empirical gains are measured against external baselines. No load-bearing self-citations or imported uniqueness claims appear in the reference list or method. Therefore score 0.
Assumptions & free parameters
free parameters (4)
- s_delta, maximum rectification offset scale =
0.10
- Loss weights lambda_v, lambda_clean, lambda_id, lambda_h, lambda_cam, lambda_j =
0.2, 0.25, 0.5, 0.15, 1.0, 0.1
- Alpha schedule interval =
0.3 to 0.7 of training progress
- Training perturbation distribution =
probability 0.7; 1 to 6 cameras; U(-15,15) degrees; U(-0.1,0.1) m
assumptions (5)
- domain assumption Pinhole camera model with SE(3) extrinsics is the exact sampling geometry used by SCA.
- domain assumption A single learned 2D offset per query-camera pair can approximately undo the image-plane displacement caused by arbitrary 6-DoF extrinsics perturbation.
- ad hoc to paper Teacher-student BEV consistency with a clean teacher is a valid training signal that does not distort detection learning.
- domain assumption The image features F_i contain enough information to predict the camera-health scalar q_i at inference.
- standard math Euler-angle parameterization of SO(3) and the norm bound ||alpha||_2 <= sqrt(3) sigma_r hold exactly.
Cite this review
Pith. "Pith review of NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection." pith.science (2026). https://pith.science/paper/52CU4GKD
@misc{pith2026260803895,
author = {Pith},
title = {Pith review of: NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/52CU4GKD}},
note = {Machine review of arXiv:2608.03895}
}
read the original abstract
Camera-based bird's-eye-view (BEV) 3D detection typically assumes accurate and fixed camera extrinsics. In detectors using spatial cross-attention (SCA), extrinsic perturbations displace the image-plane projections of BEV reference points, causing queries to sample features from incorrect regions and degrading detection performance. To address this failure mode, Noise-Conditional Gated Rectification (NCGR) is proposed to compensate for projection errors without explicitly estimating a full six-degree-of-freedom extrinsic correction. For each query-camera pair, a 2D rectification offset is predicted and modulated by a camera-level gate to rectify the base projection before native deformable sampling. During training, the perturbation-derived quantities used to construct the condition and gate are gradually replaced through scheduled interpolation by counterparts generated from an auxiliary scalar predicted from camera features. This transition enables blind inference without perturbation metadata. During training, a weight-shared clean-teacher/perturbed-student pair is used, and the rectification module is supervised by a BEV-consistency objective between the two branches. NCGR is evaluated on nuScenes with simulated dynamic and static extrinsic perturbations. In a five-camera dynamic stress test, NCGR achieves 39.69% NDS, compared with 28.00% for BEVFormer and 33.23% for CAPE. Under clean extrinsics, NCGR maintains performance comparable to that of BEVFormer.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Bai, Xuyang and Hu, Zeyu and Zhu, Xinge and Huang, Qingqiu and Chen, Yilun and Fu, Hongbo and Tai, Chiew-Lan , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[2]
Caesar, Holger and Bankiti, Varun and Lang, Alex H. and Vora, Sourabh and Liong, Venice Erin and Xu, Qiang and Krishnan, Anush and Pan, Yu and Baldan, Giancarlo and Beijbom, Oscar , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[3]
Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence , pages =
Chen, Zehui and Li, Zhenyu and Zhang, Shiquan and Fang, Liangji and Jiang, Qinhong and Zhao, Feng and Zhou, Bolei and Zhao, Hang , title =. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence , pages =. 2022 , doi =
work page 2022
-
[4]
International Conference on Learning Representations , year =
Chen, Zehui and Li, Zhenyu and Zhang, Shiquan and Fang, Liangji and Jiang, Qinhong and Zhao, Feng , title =. International Conference on Learning Representations , year =
-
[5]
Doll, Simon and Schulz, Richard and Schneider, Lukas and Benzin, Viviane and Enzweiler, Markus and Lensch, Hendrik P. A. , title =. European Conference on Computer Vision , pages =
-
[6]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Dong, Yinpeng and Kang, Caixin and Zhang, Jinlai and Zhu, Zijian and Wang, Yikai and Yang, Xiao and Su, Hang and Wei, Xingxing and Zhu, Jun , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[7]
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =
Fan, Siqi and Wang, Zhe and Huo, Xiaoliang and Wang, Yan and Liu, Jingjing , title =. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =. 2023 , doi =
work page 2023
-
[8]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
Show all 72 references
-
[9]
2021 , eprint =
Huang, Junjie and Huang, Guan and Zhu, Zheng and Ye, Yun and Du, Dalong , title =. 2021 , eprint =
2021
-
[10]
Karnik and Murthy, J
Iyer, Ganesh and Ram, R. Karnik and Murthy, J. Krishna and Krishna, K. Madhava , title =. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =
2018
-
[11]
2022 , eprint =
Jiang, Hongxiang and Meng, Wenming and Zhu, Hongmei and Zhang, Qian and Yin, Jihao , title =. 2022 , eprint =
2022
-
[12]
, title =
Klinghoffer, Tzofi and Philion, Jonah and Chen, Wenzheng and Litany, Or and Gojcic, Zan and Joo, Jungseock and Raskar, Ramesh and Fidler, Sanja and Alvarez, Jose M. , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[13]
2024 IEEE International Conference on Robotics and Automation (ICRA) , pages =
Lee, Yu-Chen and Chen, Kuan-Wen , title =. 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages =. 2024 , doi =
2024
-
[14]
Feature Pyramid Networks for Object Detection , booktitle =
Lin, Tsung-Yi and Doll. Feature Pyramid Networks for Object Detection , booktitle =
-
[15]
European Conference on Computer Vision , pages =
Li, Zhiqi and Wang, Wenhai and Li, Hongyang and Xie, Enze and Sima, Chonghao and Lu, Tong and Qiao, Yu and Dai, Jifeng , title =. European Conference on Computer Vision , pages =
-
[16]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Li, Yinhao and Ge, Zheng and Yu, Guanyi and Yang, Jinrong and Wang, Zengran and Shi, Yukang and Sun, Jianjian and Li, Zeming , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =
-
[17]
European Conference on Computer Vision , pages =
Liu, Yingfei and Wang, Tiancai and Zhang, Xiangyu and Sun, Jian , title =. European Conference on Computer Vision , pages =
-
[18]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Liu, Yingfei and Yan, Junjie and Jia, Fan and Li, Shuailin and Gao, Aqi and Wang, Tiancai and Zhang, Xiangyu , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[19]
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =
Liu, Zhijian and Tang, Haotian and Zhu, Sibo and Han, Song , title =. 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =. 2021 , doi =
2021
-
[20]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages =
Lv, Xudong and Wang, Boya and Dou, Ziwen and Ye, Dong and Wang, Shuo , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages =
-
[21]
European Conference on Computer Vision , pages =
Philion, Jonah and Fidler, Sanja , title =. European Conference on Computer Vision , pages =
-
[22]
2017 IEEE Intelligent Vehicles Symposium (IV) , pages =
Schneider, Nick and Piewak, Florian and Stiller, Christoph and Franke, Uwe , title =. 2017 IEEE Intelligent Vehicles Symposium (IV) , pages =
2017
-
[23]
European Conference on Computer Vision , pages =
Song, Ziying and Yang, Lei and Xu, Shaoqing and Liu, Lin and Xu, Dongyang and Jia, Caiyan and Jia, Feiyang and Wang, Li , title =. European Conference on Computer Vision , pages =
-
[24]
2025 IEEE International Conference on Imaging Systems and Techniques (IST) , pages =
Song, Zhihang and Yao, Dingyi and Ming, Ruibo and Peng, Lihui and Yao, Danya and Zhang, Yi , title =. 2025 IEEE International Conference on Imaging Systems and Techniques (IST) , pages =. 2025 , doi =
2025
-
[25]
Conference on Robot Learning , pages =
Wang, Yue and Guizilini, Vitor Campagnolo and Zhang, Tianyuan and Wang, Yilun and Zhao, Hang and Solomon, Justin , title =. Conference on Robot Learning , pages =
-
[26]
IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =
Xie, Shaoyuan and Kong, Lingdong and Zhang, Wenwei and Ren, Jiawei and Pan, Liang and Chen, Kai and Liu, Ziwei , title =. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =
-
[27]
2024 IEEE International Conference on Robotics and Automation (ICRA) , pages =
Xiao, Yuxuan and Li, Yao and Meng, Chengzhen and Li, Xingchen and Ji, Jianmin and Zhang, Yanyong , title =. 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages =. 2024 , doi =
2024
-
[28]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Xiong, Kaixin and Gong, Shi and Ye, Xiaoqing and Tan, Xiao and Wan, Ji and Ding, Errui and Wang, Jingdong and Bai, Xiang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[29]
, title =
Yang, Chenhongyi and Lin, Tianwei and Huang, Lichao and Crowley, Elliot J. , title =. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages =
2024
-
[30]
Sensors , volume =
Yang, Dongsheng and Fan, Xiaojie and Dong, Wei and Huang, Chaosheng and Li, Jun , title =. Sensors , volume =. 2024 , doi =
2024
-
[31]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Ye, Xin and Yaman, Burhaneddin and Cheng, Sheng and Tao, Feng and Mallik, Abhirup and Ren, Liu , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[32]
2025 , eprint =
Yuan, Weiduo and Li, Jerry and Yue, Justin and Shah, Divyank and Karydis, Konstantinos and Qiu, Hang , title =. 2025 , eprint =
2025
-
[33]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Zhou, Yunsong and He, Yuan and Zhu, Hongzi and Wang, Cheng and Li, Hongyang and Jiang, Qinhong , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[34]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Zhu, Xizhou and Hu, Han and Lin, Stephen and Dai, Jifeng , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[35]
International Conference on Learning Representations , year =
Zhu, Xizhou and Su, Weijie and Lu, Lewei and Li, Bin and Wang, Xiaogang and Dai, Jifeng , title =. International Conference on Learning Representations , year =
-
[36]
2026 , eprint =
Zhuo, Lifeng and Jin, Kefan and Liu, Zhe and Wang, Hesheng , title =. 2026 , eprint =
2026
-
[37]
Bai, X.; Hu, Z.; Zhu, X.; Huang, Q.; Chen, Y.; Fu, H.; and Tai, C.-L. 2022. TransFusion : Robust LiDAR -Camera Fusion for 3D Object Detection with Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1090--1099
2022
-
[38]
H.; Vora, S.; Liong, V
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020. nuScenes : A Multimodal Dataset for Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11621--11631
2020
-
[39]
Chen, Z.; Li, Z.; Zhang, S.; Fang, L.; Jiang, Q.; and Zhao, F. 2023. BEVDistill : Cross-Modal BEV Distillation for Multi-View 3D Object Detection. In International Conference on Learning Representations
2023
-
[40]
Chen, Z.; Li, Z.; Zhang, S.; Fang, L.; Jiang, Q.; Zhao, F.; Zhou, B.; and Zhao, H. 2022. AutoAlign : Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, 827--833
2022
-
[41]
Doll, S.; Schulz, R.; Schneider, L.; Benzin, V.; Enzweiler, M.; and Lensch, H. P. A. 2022. SpatialDETR : Robust Scalable Transformer-Based 3D Object Detection from Multi-View Camera Images with Global Cross-Sensor Attention. In European Conference on Computer Vision, 230--245....
2022
-
[42]
Dong, Y.; Kang, C.; Zhang, J.; Zhu, Z.; Wang, Y.; Yang, X.; Su, H.; Wei, X.; and Zhu, J. 2023. Benchmarking Robustness of 3D Object Detection to Common Corruptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1022--1032
2023
-
[43]
Fan, S.; Wang, Z.; Huo, X.; Wang, Y.; and Liu, J. 2023. Calibration-Free BEV Representation for Infrastructure Perception. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 9008--9013
2023
-
[44]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770--778
2016
-
[45]
Huang, J.; Huang, G.; Zhu, Z.; Ye, Y.; and Du, D. 2021. BEVDet : High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View. arXiv:2112.11790
2021 arXiv
-
[46]
K.; Murthy, J
Iyer, G.; Ram, R. K.; Murthy, J. K.; and Krishna, K. M. 2018. CalibNet : Geometrically Supervised Extrinsic Calibration Using 3D Spatial Transformer Networks. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1110--1117
2018
-
[47]
Jiang, H.; Meng, W.; Zhu, H.; Zhang, Q.; and Yin, J. 2022. Multi-Camera Calibration Free BEV Representation for 3D Object Detection. arXiv:2210.17252
2022 arXiv
-
[48]
Klinghoffer, T.; Philion, J.; Chen, W.; Litany, O.; Gojcic, Z.; Joo, J.; Raskar, R.; Fidler, S.; and Alvarez, J. M. 2023. Towards Viewpoint Robustness in Bird's Eye View Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8515--8524
2023
-
[49]
Lee, Y.-C.; and Chen, K.-W. 2024. LCCRAFT : LiDAR and Camera Calibration Using Recurrent All-Pairs Field Transforms Without Precise Initial Guess. In 2024 IEEE International Conference on Robotics and Automation (ICRA), 16669--16675
2024
-
[50]
Li, Y.; Ge, Z.; Yu, G.; Yang, J.; Wang, Z.; Shi, Y.; Sun, J.; and Li, Z. 2023. BEVDepth : Acquisition of Reliable Depth for Multi-View 3D Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence, 37(2): 1477--1485
2023
-
[51]
Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Qiao, Y.; and Dai, J. 2022. BEVFormer : Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers. In European Conference on Computer Vision, 1--18. Springer
2022
-
[52]
Lin, T.-Y.; Doll \'a r, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2117--2125
2017
-
[53]
Liu, Y.; Wang, T.; Zhang, X.; and Sun, J. 2022. PETR : Position Embedding Transformation for Multi-View 3D Object Detection. In European Conference on Computer Vision, 531--548. Springer
2022
-
[54]
Liu, Y.; Yan, J.; Jia, F.; Li, S.; Gao, A.; Wang, T.; and Zhang, X. 2023. PETRv2 : A Unified Framework for 3D Perception from Multi-Camera Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 3262--3272
2023
-
[55]
Liu, Z.; Tang, H.; Zhu, S.; and Han, S. 2021. SemAlign : Annotation-Free Camera- LiDAR Calibration with Semantic Alignment Loss. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 8845--8851
2021
-
[56]
Lv, X.; Wang, B.; Dou, Z.; Ye, D.; and Wang, S. 2021. LCCNet : LiDAR and Camera Self-Calibration Using Cost Volume Network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2894--2901
2021
-
[57]
Philion, J.; and Fidler, S. 2020. Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D . In European Conference on Computer Vision, 194--210. Springer
2020
-
[58]
Schneider, N.; Piewak, F.; Stiller, C.; and Franke, U. 2017. RegNet : Multimodal Sensor Registration Using Deep Neural Networks. In 2017 IEEE Intelligent Vehicles Symposium (IV), 1803--1810
2017
-
[59]
Song, Z.; Yang, L.; Xu, S.; Liu, L.; Xu, D.; Jia, C.; Jia, F.; and Wang, L. 2024. GraphBEV : Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection. In European Conference on Computer Vision, 347--366. Springer
2024
-
[60]
Song, Z.; Yao, D.; Ming, R.; Peng, L.; Yao, D.; and Zhang, Y. 2025. A Re-Calibration Method for Object Detection with Multimodal Alignment Bias in Autonomous Driving. In 2025 IEEE International Conference on Imaging Systems and Techniques (IST), 1--6
2025
-
[61]
C.; Zhang, T.; Wang, Y.; Zhao, H.; and Solomon, J
Wang, Y.; Guizilini, V. C.; Zhang, T.; Wang, Y.; Zhao, H.; and Solomon, J. 2022. DETR3D : 3D Object Detection from Multi-View Images via 3D -to- 2D Queries. In Conference on Robot Learning, 180--191. PMLR
2022
-
[62]
Xiao, Y.; Li, Y.; Meng, C.; Li, X.; Ji, J.; and Zhang, Y. 2024. CalibFormer : A Transformer-Based Automatic LiDAR -Camera Calibration Network. In 2024 IEEE International Conference on Robotics and Automation (ICRA), 16714--16720
2024
-
[63]
Xie, S.; Kong, L.; Zhang, W.; Ren, J.; Pan, L.; Chen, K.; and Liu, Z. 2025. Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous Driving. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(5): 3878--3894
2025
-
[64]
Xiong, K.; Gong, S.; Ye, X.; Tan, X.; Wan, J.; Ding, E.; Wang, J.; and Bai, X. 2023. CAPE : Camera View Position Embedding for Multi-View 3D Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21570--21579
2023
-
[65]
Yang, C.; Lin, T.; Huang, L.; and Crowley, E. J. 2024 a . WidthFormer : Toward Efficient Transformer-Based BEV View Transformation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 8457--8464
2024
-
[66]
Yang, D.; Fan, X.; Dong, W.; Huang, C.; and Li, J. 2024 b . Robust BEV 3D Object Detection for Vehicles with Tire Blow-Out. Sensors, 24(14): 4446
2024
-
[67]
Ye, X.; Yaman, B.; Cheng, S.; Tao, F.; Mallik, A.; and Ren, L. 2025. BEVDiffuser : Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1495--1504
2025
-
[68]
Yuan, W.; Li, J.; Yue, J.; Shah, D.; Karydis, K.; and Qiu, H. 2025. BEVCalib : LiDAR -Camera Calibration via Geometry-Guided Bird's-Eye View Representations. arXiv:2506.02587
2025 arXiv
-
[69]
Zhou, Y.; He, Y.; Zhu, H.; Wang, C.; Li, H.; and Jiang, Q. 2021. Monocular 3D Object Detection: An Extrinsic Parameter Free Approach. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7556--7566
2021
-
[70]
Zhu, X.; Hu, H.; Lin, S.; and Dai, J. 2019. Deformable ConvNets V2: More Deformable, Better Results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9308--9316
2019
-
[71]
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2021. Deformable DETR : Deformable Transformers for End-to-End Object Detection. In International Conference on Learning Representations
2021
-
[72]
Zhuo, L.; Jin, K.; Liu, Z.; and Wang, H. 2026. RESBev : Making BEV Perception More Robust. arXiv:2603.09529
2026 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.