Pith. sign in

REVIEW 3 major objections 7 minor 144 references

Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

T0 review · 3 major / 7 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read The paper argues that self-driving perception foundation models should be organized around four core capabilities—generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding—and that no current system i

desk verdict A useful capability-organized survey of foundation models for AD perception whose central 'novel taxonomy' claim is contradicted by its own Table 1 — worth a referee after cleanup, but not the field-defining framing it bills itself as. read the letter →

arxiv 2509.08302 v1 pith:5OLQUSI2 submitted 2025-09-10 cs.RO cs.CV

classification cs.ROcs.CV
keywords autonomousdrivingperceptionfoundationmodelscapabilitytaxonomygeneralizedknowledgespatialunderstandingmulti-sensorrobustnesstemporalself-supervisedlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that foundation models for self-driving perception should be organized around four capabilities—generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding—rather than around tasks or methods. For each capability it reviews the main technical routes: distillation and pseudo-labeling for generalized knowledge; volumetric models, neural rendering, and masked autoencoders for spatial understanding; cross-modal contrastive learning, knowledge distillation, multi-view consistency, and diffusion for multi-sensor robustness; and 4D prediction, diffusion world models, and temporal contrastive learning for temporal understanding. It claims that no existing system seamlessly integrates all four capabilities into a real-time framework. If the taxonomy is right, it gives developers a capability-driven checklist for building and evaluating perception models and names the integration of these capabilities as the central open problem.

What carries the argument

The load-bearing object is the four-capability taxonomy itself, plus the coverage matrix (Table 10) that assigns each surveyed method an X for each capability it provides. The taxonomy does the work of partitioning the literature: each of the four sections maps one capability to a family of techniques, such as knowledge distillation and pseudo-labeling for generalized knowledge; NeRF/3D Gaussian Splatting and masked autoencoders for spatial understanding; contrastive learning, distillation, multi-view consistency, and diffusion for multi-sensor robustness; and 4D forecasting and temporal contrastive learning for temporal understanding. The matrix then serves as the survey's evidence that cap

What would settle it

A concrete falsifier would be a deployed or benchmarked perception system that demonstrably operates in real time while satisfying all four capabilities on a held-out distribution-shift suite; observing such a system would falsify the paper's claim that no existing system integrates all four. Alternatively, a controlled study showing that improving temporal understanding necessarily changes multi-sensor robustness would falsify the taxonomy's separability assumption.

Watch

Extended reading notes

Core claim

The paper's central claim is that a capability-based taxonomy—generalized knowledge, spatial understanding, multi-sensor robustness, temporal understanding—captures what a perception foundation model must master to handle dynamic, long-tail driving, and that this taxonomy should steer research better than task- or method-based surveys do. To support this, the survey classifies recent methods into capability-oriented clusters: VFM/VLM/LLM adaptation mechanisms for generalized knowledge; volumetric models, NeRF/3D Gaussian Splatting rendering, and 3D masked autoencoders for spatial understanding; cross-modality contrastive learning, distillation, multi-view consistency, multi-modal masked auto

Load-bearing premise

The taxonomy assumes the four capabilities—generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding—are a complete and cleanly separable decomposition of autonomous-driving perception, so every method can be placed in distinct boxes.

Editorial extensions

If this is right

  • Researchers can use the four capabilities as a checklist when designing a perception foundation model, choosing whichever pillar is weakest for their deployment target.
  • The integration claim implies that hybrid systems—foundation models handling high-level reasoning at a reduced rate plus a conventional real-time pipeline—are an interim solution, not the end state.
  • Benchmarks should be built to isolate and stress each capability—for example, corruption benchmarks for multi-sensor robustness and accident-focused scenarios for generalized knowledge—rather than reporting only average-case mAP or IoU.
  • If no current system combines all four capabilities, progress depends on closing integration gaps: latency, calibration, synchronization, and the representation mismatch between dense spatial outputs and object-centric planning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's exhaustive-and-separable assumption is the part most worth testing: temporal understanding and multi-sensor robustness both lean on spatial alignment, so a capability overlap, rather than four independent axes, may be the real structure.
  • A quantitative analogue of Table 10—scoring degree of capability instead of binary X marks—would turn the survey's qualitative claim into a testable rubric and could reveal whether the four categories are actually independent.
  • The survey's benchmark table suggests a natural extension: build a benchmark suite that perturbs one capability axis at a time—sensor dropout, occlusion, unseen categories, temporal discontinuity—so capability-level comparisons can be made fairly across models.
  • The integration bottleneck also implies an opportunity: a model that does combine all four, even at modest performance, would be more informative for the field than another state-of-the-art result on a single task.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This manuscript surveys foundation models for autonomous-driving perception and organizes the literature around a proposed taxonomy of four core capabilities: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding. Each capability is motivated, illustrated with representative methods (e.g., VFM/VLM/LLM distillation, occupancy networks, neural rendering, cross-modal contrastive learning, temporal contrastive learning), and accompanied by challenges. The paper further claims that this capability-based taxonomy is novel and that no existing system integrates all four capabilities into a real-time framework. The final sections discuss benchmark limitations, real-time latency, data bias, regulatory issues, and hallucination risks.

Significance. If the taxonomy is accepted, the survey provides a useful organizational frame for a rapidly growing literature and a practical checklist for model development. Its strengths are the breadth of recent work covered, the structured tables (Tables 3, 6, 7, 9, 10), and the attention to deployment issues such as latency, benchmarks, and hallucination. No new quantitative results or code are presented, so the appropriate standard of assessment is internal consistency and accurate characterization of prior work. Measured against that standard, the paper needs revision: the central novelty claim is internally contradicted by Table 1, and the binary capability assignments in Table 10 are not fully aligned with the text.

major comments (3)
  1. [Section I and Table 1] The abstract and Section I claim a 'novel taxonomy' and state that existing surveys 'frequently overlook' multi-sensor robustness and spatial awareness. However, Table 1 lists survey [6] as covering all four capabilities, including exactly those two. This is a direct internal contradiction: either [6] already organizes around these capabilities, which undermines the novelty claim, or Table 1 mischaracterizes [6], which undermines the reliability of the comparison matrix. The paper provides no section-level comparison or criterion for what counts as 'coverage.' Please revise the claim or provide a precise distinction between method-based and capability-based coverage and show that [6] fails the latter.
  2. [Section VII.A and Table 10] The claim that no existing system integrates all four capabilities into a real-time framework rests on Table 10, but the row assignments are not reliable. For example, the text describes SEAL as combining generalized knowledge with multi-sensor robustness and temporal understanding, yet the Table 10 row 'SLidR [43] / SEAL [44]' appears to mark only two capabilities. The text and table must be aligned, and the coding rules for assigning X marks should be stated explicitly.
  3. [Section I and Table 10] The four capabilities are treated as cleanly separable binary attributes, but they are interdependent: cross-modal fusion requires spatial alignment, temporal modeling often depends on multi-view geometric consistency, and generalized knowledge is used to resolve ambiguities in 3D reconstruction. The paper does not define what 'covers a capability' means or how overlap is adjudicated. Since the 'no seamless integration' conclusion is the survey's central claim, I ask for a short scope definition for each capability and a reproducible coding rubric for Table 10.
minor comments (7)
  1. [Abstract vs. Section VI] The abstract uses 'temporal reasoning' while Section VI and the conclusion use 'temporal understanding.' Please use one term consistently.
  2. [Figure 4 caption] Typo: 'supervise a encoder' should be 'supervise an encoder.'
  3. [Section IV.A and V.E] Typos: 'V olumetric' should be 'Volumetric' in Section IV.A; 'defusion' should be 'diffusion' in Section V.E.
  4. [Section II.E] Capitalization: 'Therefore, They become foundational components' should be 'therefore, they become...'
  5. [Section II.B] The distillation loss equations are set with inline expressions for the softened probabilities, which is hard to read. Please display these as numbered equations.
  6. [Table 10] Rows that group multiple methods under one mark, e.g., 'SLidR [43] / SEAL [44]', obscure differences in capability coverage. Use separate rows or per-method marks.
  7. [Section VII] The survey does not state its own limitations, including the literature selection protocol and the coverage cutoff date. Adding a short limitations paragraph would improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's taxonomy is an interpretive frame; there are no fitted inputs, derived predictions, or load-bearing self-citations.

full rationale

This paper is a survey whose contribution is an organizing taxonomy (generalized knowledge, spatial understanding, multi-sensor robustness, temporal understanding), not a derivation of empirical predictions from first principles. There are no fitted parameters, no equations used to derive results, no predictions from fitted inputs, and no load-bearing self-citations: the authors do not cite their own prior work, and all cited systems are external published methods described by the survey. The four capabilities are introduced by definition in Section I and then used to structure the review; Table 10's X-marks are the authors' qualitative literature judgments, not quantities derived from data. Consequently, there is no chain by which an output reduces to an input by construction. The skeptical concern that Table 1 lists reference [6] as covering all four capabilities, conflicting with the Introduction's claim that prior surveys overlook multi-sensor robustness and spatial awareness, is an internal-consistency/novelty-support issue, not a circularity: resolving it requires checking [6]'s actual scope, not exhibiting an equation-level or citation-level loop. Likewise, the subjective coverage judgments in Table 10 are potential evidence-quality concerns but do not constitute circular reasoning. Thus no circularity is found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

The survey introduces no fitted numbers. Its conceptual content rests on accurate summarization of external work, the validity of the chosen capability decomposition, and the standard assumption that foundation-model pretraining transfers downstream. The four-capability taxonomy itself is an invented organizational entity with no independent falsifiable handle.

assumptions (3)
  • domain assumption The cited primary papers are accurately summarized, and each is a legitimate instance of the assigned capability.
    The entire survey stands on grouping external works into four categories; no re-implementation or benchmark comparison is provided.
  • ad hoc to paper The four capabilities are jointly exhaustive and largely orthogonal for AD perception.
    This is the paper's central framing, introduced in Section I and operationalized in Table 10; no evidence establishes exhaustiveness or orthogonality.
  • domain assumption Foundation-model pretraining on diverse data transfers to AD perception.
    Sections II and III assume this premise; it is standard in the literature but not verified by the survey.
invented entities (1)
  • Four-capability taxonomy
    purpose: Provides the survey's organizing structure: generalized knowledge, spatial understanding, multi-sensor robustness, temporal understanding.
    A conceptual framework asserted by the authors; no independent benchmark or analysis demonstrates that these four axes are the correct or complete decomposition of AD perception.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities." pith.science (2026). https://pith.science/paper/5OLQUSI2

@misc{pith2026250908302,
  author       = {Pith},
  title        = {Pith review of: Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5OLQUSI2}},
  note         = {Machine review of arXiv:2509.08302}
}
read the original abstract

Foundation models are revolutionizing autonomous driving perception, transitioning the field from narrow, task-specific deep learning models to versatile, general-purpose architectures trained on vast, diverse datasets. This survey examines how these models address critical challenges in autonomous perception, including limitations in generalization, scalability, and robustness to distributional shifts. The survey introduces a novel taxonomy structured around four essential capabilities for robust performance in dynamic driving environments: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal reasoning. For each capability, the survey elucidates its significance and comprehensively reviews cutting-edge approaches. Diverging from traditional method-centric surveys, our unique framework prioritizes conceptual design principles, providing a capability-driven guide for model development and clearer insights into foundational aspects. We conclude by discussing key challenges, particularly those associated with the integration of these capabilities into real-time, scalable systems, and broader deployment challenges related to computational demands and ensuring model reliability against issues like hallucinations and out-of-distribution failures. The survey also outlines crucial future research directions to enable the safe and effective deployment of foundation models in autonomous driving systems.

Figures

Figures reproduced from arXiv: 2509.08302 by the authors.

Figure 1
Figure 1. FIGURE 1 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIGURE 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIGURE 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: FIGURE 4 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: FIGURE 5 [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: FIGURE 6 [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: FIGURE 7 [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: FIGURE 8 [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 10
Figure 10. Figure 10: FIGURE 10 [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 12
Figure 12. Figure 12: FIGURE 12 [PITH_FULL_IMAGE:figures/full_fig_p018_12.png]
Figure 13
Figure 13. Figure 13: FIGURE 13 [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]
Figure 14
Figure 14. Figure 14: FIGURE 14 [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

144 extracted references · 47 canonical work pages

  1. [6]

    Forging vision foundation models for autonomous driving: Challenges, methodologies, and opportunities,

    X. Yan, H. Zhang, Y . Cai, J. Guo, W. Qiu, B. Gao, K. Zhou, Y . Zhao, H. Jin, J. Gaoet al., “Forging vision foundation models for autonomous driving: Challenges, methodologies, and opportunities,”arXiv preprint arXiv:2401.08045, 2024

  2. [43]

    Image-to-lidar self-supervised distilla- tion for autonomous driving data,

    C. Sautier, G. Puy, S. Gidaris, A. Boulch, A. Bursuc, and R. Marlet, “Image-to-lidar self-supervised distilla- tion for autonomous driving data,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9891–9901

  3. [44]

    Segment any point cloud sequences by distilling vision foundation models,

    Y . Liu, L. Kong, J. Cen, R. Chen, W. Zhang, L. Pan, K. Chen, and Z. Liu, “Segment any point cloud sequences by distilling vision foundation models,” Advances in Neural Information Processing Systems, vol. 36, pp. 37 193–37 229, 2023

  4. [1]

    3d object detection for autonomous driving: A survey,

    R. Qian, X. Lai, and X. Li, “3d object detection for autonomous driving: A survey,”Pattern Recognition, vol. 130, p. 108796, 2022

  5. [2]

    Vision-based semantic segmentation in scene understanding for autonomous driving: Recent achievements, challenges, and out- looks,

    K. Muhammad, T. Hussain, H. Ullah, J. Del Ser, M. Rezaei, N. Kumar, M. Hijji, P. Bellavista, and V . H. C. de Albuquerque, “Vision-based semantic segmentation in scene understanding for autonomous driving: Recent achievements, challenges, and out- looks,”IEEE Transactions on Intelligent Transporta- tion Systems, vol. 23, no. 12, pp. 22 694–22 715, 2022

  6. [3]

    A review of deep learning-based visual multi-object tracking algorithms for autonomous driving,

    S. Guo, S. Wang, Z. Yang, L. Wang, H. Zhang, P. Guo, Y . Gao, and J. Guo, “A review of deep learning-based visual multi-object tracking algorithms for autonomous driving,”Applied Sciences, vol. 12, no. 21, p. 10741, 2022

  7. [4]

    Towards long-tailed 3d detection,

    N. Peri, A. Dave, D. Ramanan, and S. Kong, “Towards long-tailed 3d detection,” inConference on Robot Learning. PMLR, 2023, pp. 1904–1915

  8. [5]

    On the opportuni- ties and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportuni- ties and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021

Show all 144 references
  1. [7]

    Applications of large scale foundation models for autonomous driving,

    Y . Huang, Y . Chen, and Z. Li, “Applications of large scale foundation models for autonomous driving,” arXiv preprint arXiv:2311.12144, 2023

  2. [8]

    A survey for foundation models in au- tonomous driving,

    H. Gao, Z. Wang, Y . Li, K. Long, M. Yang, and Y . Shen, “A survey for foundation models in au- tonomous driving,”arXiv preprint arXiv:2402.01105, 2024

  3. [9]

    Prospective role of foundation models in advancing autonomous vehicles,

    J. Wu, B. Gao, J. Gao, J. Yu, H. Chu, Q. Yu, X. Gong, Y . Chang, H. E. Tseng, H. Chenet al., “Prospective role of foundation models in advancing autonomous vehicles,”Research, vol. 7, p. 0399, 2024

  4. [10]

    Llm4drive: A survey of large language models for autonomous driving,

    Z. Yang, X. Jia, H. Li, and J. Yan, “Llm4drive: A survey of large language models for autonomous driving,”arXiv preprint arXiv:2311.01043, 2023

  5. [11]

    Vision language models in autonomous driving: A survey and outlook,

    X. Zhou, M. Liu, E. Yurtsever, B. L. Zagar, W. Zim- mer, H. Cao, and A. C. Knoll, “Vision language models in autonomous driving: A survey and outlook,” IEEE Transactions on Intelligent Vehicles, 2024

  6. [12]

    A simple framework for contrastive learning of vi- sual representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of vi- sual representations,” inInternational conference on machine learning. PmLR, 2020, pp. 1597–1607. 26 VOLUME 00, 2024

  7. [13]

    Momentum contrast for unsupervised visual repre- sentation learning,

    K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual repre- sentation learning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2020, pp. 9729–9738

  8. [14]

    Improved baselines with momentum contrastive learning,

    X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,”arXiv preprint arXiv:2003.04297, 2020

  9. [15]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 000–16 009

  10. [16]

    An image is worth 16x16 words: Transformers for image recogni- tion at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weis- senborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recogni- tion at scale,”arXiv preprint arXiv:2010.11929, 2020

  11. [17]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”arXiv preprint arXiv:1503.02531, 2015

  12. [18]

    Knowledge distillation: A survey,

    J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,”International Journal of Com- puter Vision, vol. 129, no. 6, pp. 1789–1819, 2021

  13. [19]

    Self- training with noisy student improves imagenet classi- fication,

    Q. Xie, M.-T. Luong, E. Hovy, and Q. V . Le, “Self- training with noisy student improves imagenet classi- fication,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 10 687–10 698

  14. [20]

    Bootstrap your own latent-a new approach to self- supervised learning,

    J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azaret al., “Bootstrap your own latent-a new approach to self- supervised learning,”Advances in neural information processing systems, vol. 33, pp. 21...

  15. [21]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 9650–9660

  16. [22]

    Structure-from- motion revisited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from- motion revisited,” inProceedings of the IEEE con- ference on computer vision and pattern recognition, 2016, pp. 4104–4113

  17. [23]

    Multi-view stereo: A tutorial,

    Y . Furukawa, C. Hern´andezet al., “Multi-view stereo: A tutorial,”Foundations and trends® in Computer Graphics and Vision, vol. 9, no. 1-2, pp. 1–148, 2015

  18. [24]

    Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,

    C. Cadena, L. Carlone, H. Carrillo, Y . Latif, D. Scara- muzza, J. Neira, I. Reid, and J. J. Leonard, “Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age,”IEEE Transactions on robotics, vol. 32, no. 6, pp. 1309– 1332, 2016

  19. [25]

    3d-r2n2: A unified approach for single and multi- view 3d object reconstruction,

    C. B. Choy, D. Xu, J. Gwak, K. Chen, and S. Savarese, “3d-r2n2: A unified approach for single and multi- view 3d object reconstruction,” inComputer vision– ECCV 2016: 14th European conference, amsterdam, the netherlands, October 11-14, 2016, proceedings, part VIII 14. Springer...

  20. [26]

    Predicting depth, surface normals and semantic labels with a common multi- scale convolutional architecture,

    D. Eigen and R. Fergus, “Predicting depth, surface normals and semantic labels with a common multi- scale convolutional architecture,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 2650–2658

  21. [27]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Bar- ron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM, vol. 65, no. 1, pp. 99– 106, 2021

  22. [28]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Dret- takis, “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  23. [29]

    Yolov3: An incremen- tal improvement,

    J. Redmon and A. Farhadi, “Yolov3: An incremen- tal improvement,”arXiv preprint arXiv:1804.02767, 2018

  24. [30]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  25. [31]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkatet al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023

  26. [32]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” inProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologie...

  27. [33]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.- Y . Loet al., “Segment anything,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026

  28. [34]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023

  29. [35]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PmLR, 2021, pp. 8748–8763

  30. [36]

    Grounding dino: Marrying dino with grounded pre-training for VOLUME 00, 2024 27 Authoret al.: Preparation of Papers for IEEE OPEN JOURNALS open-set object detection,

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Suet al., “Grounding dino: Marrying dino with grounded pre-training for VOLUME 00, 2024 27 Authoret al.: Preparation of Papers for IEEE OPEN JOURNALS open-set object detection,” inEuropean Conferen...

  31. [37]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020

  32. [38]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695

  33. [39]

    Video diffusion models,

    J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” Advances in Neural Information Processing Systems, vol. 35, pp. 8633–8646, 2022

  34. [40]

    3d shape generation and completion through point-voxel diffusion,

    L. Zhou, Y . Du, and J. Wu, “3d shape generation and completion through point-voxel diffusion,” inPro- ceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5826–5835

  35. [41]

    Diffusion-based signed distance fields for 3d shape generation,

    J. Shim, C. Kang, and K. Joo, “Diffusion-based signed distance fields for 3d shape generation,” inProceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 20 887–20 897

  36. [42]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Ka- plan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sas- try, A. Askellet al., “Language models are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  37. [45]

    Better call sal: Towards learning to segment anything in lidar,

    A. O ˇsep, T. Meinhardt, F. Ferroni, N. Peri, D. Ra- manan, and L. Leal-Taix ´e, “Better call sal: Towards learning to segment anything in lidar,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 71–90

  38. [46]

    Sam4udass: When sam meets unsupervised domain adaptive semantic segmentation in intelligent ve- hicles,

    W. Yan, Y . Qian, H. Zhuang, C. Wang, and M. Yang, “Sam4udass: When sam meets unsupervised domain adaptive semantic segmentation in intelligent ve- hicles,”IEEE Transactions on Intelligent Vehicles, vol. 9, no. 2, pp. 3396–3408, 2024

  39. [47]

    Occnerf: Self-supervised multi- camera occupancy prediction with neural radiance fields,

    C. Zhang, J. Yan, Y . Wei, J. Li, L. Liu, Y . Tang, Y . Duan, and J. Lu, “Occnerf: Self-supervised multi- camera occupancy prediction with neural radiance fields,”CoRR, 2023

  40. [48]

    Open 3d world in autonomous driving,

    X. Cheng and L. Li, “Open 3d world in autonomous driving,”arXiv preprint arXiv:2408.10880, 2024

  41. [49]

    Ovo: Open-vocabulary occupancy,

    Z. Tan, Z. Dong, C. Zhang, W. Zhang, H. Ji, and H. Li, “Ovo: Open-vocabulary occupancy,”arXiv preprint arXiv:2305.16133, 2023

  42. [50]

    Clip2scene: Towards label-efficient 3d scene understanding by clip,

    R. Chen, Y . Liu, L. Kong, X. Zhu, Y . Ma, Y . Li, Y . Hou, Y . Qiao, and W. Wang, “Clip2scene: Towards label-efficient 3d scene understanding by clip,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7020–7030

  43. [51]

    Vlm2scene: Self- supervised image-text-lidar learning with foundation models for autonomous driving scene understanding,

    G. Liao, J. Li, and X. Ye, “Vlm2scene: Self- supervised image-text-lidar learning with foundation models for autonomous driving scene understanding,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 3351–3359

  44. [52]

    Unsupervised 3d perception with 2d vision-language distillation for autonomous driv- ing,

    M. Najibi, J. Ji, Y . Zhou, C. R. Qi, X. Yan, S. Ettinger, and D. Anguelov, “Unsupervised 3d perception with 2d vision-language distillation for autonomous driv- ing,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8602– 8612

  45. [53]

    Opensight: A simple open-vocabulary framework for lidar-based object detection,

    H. Zhang, J. Xu, T. Tang, H. Sun, X. Yu, Z. Huang, and K. Yu, “Opensight: A simple open-vocabulary framework for lidar-based object detection,” inEu- ropean Conference on Computer Vision. Springer, 2024, pp. 1–19

  46. [54]

    Sam3d: Zero-shot 3d object detec- tion via segment anything model,

    D. Zhang, D. Liang, H. Yang, Z. Zou, X. Ye, Z. Liu, and X. Bai, “Sam3d: Zero-shot 3d object detec- tion via segment anything model,”arXiv preprint arXiv:2306.02245, 2023

  47. [55]

    Gpt- driver: Learning to drive with gpt,

    J. Mao, Y . Qian, J. Ye, H. Zhao, and Y . Wang, “Gpt- driver: Learning to drive with gpt,”arXiv preprint arXiv:2310.01415, 2023

  48. [56]

    Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,

    S. Wang, Z. Yu, X. Jiang, S. Lan, M. Shi, N. Chang, J. Kautz, Y . Li, and J. M. Alvarez, “Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,”arXiv preprint arXiv:2405.01533, 2024

  49. [57]

    Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

    Z. Xu, Y . Zhang, E. Xie, Z. Zhao, Y . Guo, K.-Y . K. Wong, Z. Li, and H. Zhao, “Drivegpt4: Interpretable end-to-end autonomous driving via large language model,”IEEE Robotics and Automation Letters, 2024

  50. [58]

    Dol- phins: Multimodal language model for driving,

    Y . Ma, Y . Cao, J. Sun, M. Pavone, and C. Xiao, “Dol- phins: Multimodal language model for driving,” in European Conference on Computer Vision. Springer, 2024, pp. 403–420

  51. [59]

    Emma: End-to- end multimodal model for autonomous driving,

    J.-J. Hwang, R. Xu, H. Lin, W.-C. Hung, J. Ji, K. Choi, D. Huang, T. He, P. Covington, B. Sapp, Y . Zhou, J. Guo, D. Anguelov, and M. Tan, “Emma: End-to- end multimodal model for autonomous driving,”arXiv preprint arXiv:2410.23262, 2024

  52. [60]

    Lidar-llm: Exploring the potential of large language models for 3d lidar understanding,

    S. Yang, J. Liu, R. Zhang, M. Pan, Z. Guo, X. Li, Z. Chen, P. Gao, H. Li, Y . Guoet al., “Lidar-llm: Exploring the potential of large language models for 3d lidar understanding,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 9, 2025, pp. 9247–9255

  53. [61]

    A survey on hallucination in large language mod- els: Principles, taxonomy, challenges, and open ques- tions,

    L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qinet al., 28 VOLUME 00, 2024 “A survey on hallucination in large language mod- els: Principles, taxonomy, challenges, and open ques- tions,”ACM Transactions on Information Systems, vol. 43, no. ...

  54. [62]

    Retrieval-augmented genera- tion for knowledge-intensive nlp tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschelet al., “Retrieval-augmented genera- tion for knowledge-intensive nlp tasks,”Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020

  55. [63]

    Self-rag: Learning to retrieve, generate, and critique through self-reflection,

    A. Asai, Z. Wu, Y . Wang, A. Sil, and H. Hajishirzi, “Self-rag: Learning to retrieve, generate, and critique through self-reflection,” 2024

  56. [64]

    Driving with llms: Fusing object-level vector modal- ity for explainable autonomous driving,

    L. Chen, O. Sinavski, J. H ¨unermann, A. Karnsund, A. J. Willmott, D. Birch, D. Maund, and J. Shotton, “Driving with llms: Fusing object-level vector modal- ity for explainable autonomous driving,” in2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2...

  57. [65]

    A survey on occupancy perception for autonomous driv- ing: The information fusion perspective,

    H. Xu, J. Chen, S. Meng, Y . Wang, and L.-P. Chau, “A survey on occupancy perception for autonomous driv- ing: The information fusion perspective,”Information Fusion, vol. 114, p. 102671, 2025

  58. [66]

    Neural vol- umetric world models for autonomous driving,

    Z. Huang, J. Zhang, and E. Ohn-Bar, “Neural vol- umetric world models for autonomous driving,” in European Conference on Computer Vision. Springer, 2024, pp. 195–213

  59. [67]

    Tri-perspective view for vision-based 3d semantic oc- cupancy prediction,

    Y . Huang, W. Zheng, Y . Zhang, J. Zhou, and J. Lu, “Tri-perspective view for vision-based 3d semantic oc- cupancy prediction,” inProceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2023, pp. 9223–9232

  60. [68]

    V oxformer: Sparse voxel transformer for camera-based 3d se- mantic scene completion,

    Y . Li, Z. Yu, C. Choy, C. Xiao, J. M. Alvarez, S. Fidler, C. Feng, and A. Anandkumar, “V oxformer: Sparse voxel transformer for camera-based 3d se- mantic scene completion,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 9087–9098

  61. [69]

    Occformer: Dual-path transformer for vision-based 3d semantic occupancy prediction,

    Y . Zhang, Z. Zhu, and D. Du, “Occformer: Dual-path transformer for vision-based 3d semantic occupancy prediction,” inProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, 2023, pp. 9433–9443

  62. [70]

    Fully sparse 3d occupancy prediction,

    H. Liu, Y . Chen, H. Wang, Z. Yang, T. Li, J. Zeng, L. Chen, H. Li, and L. Wang, “Fully sparse 3d occupancy prediction,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 54–71

  63. [71]

    Hybridocc: Nerf enhanced transformer-based multi- camera 3d occupancy prediction,

    X. Zhao, B. Chen, M. Sun, D. Yang, Y . Wang, X. Zhang, M. Li, D. Kou, X. Wei, and L. Zhang, “Hybridocc: Nerf enhanced transformer-based multi- camera 3d occupancy prediction,”IEEE Robotics and Automation Letters, 2024

  64. [72]

    Renderocc: Vision- centric 3d occupancy prediction with 2d rendering supervision,

    M. Pan, J. Liu, R. Zhang, P. Huang, X. Li, H. Xie, B. Wang, L. Liu, and S. Zhang, “Renderocc: Vision- centric 3d occupancy prediction with 2d rendering supervision,” in2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, pp. 12 404–12 411

  65. [73]

    S-nerf++: Autonomous driving simu- lation via neural reconstruction and generation,

    Y . Chen, J. Zhang, Z. Xie, W. Li, F. Zhang, J. Lu, and L. Zhang, “S-nerf++: Autonomous driving simu- lation via neural reconstruction and generation,”IEEE Transactions on Pattern Analysis and Machine Intel- ligence, 2025

  66. [74]

    Selfocc: Self-supervised vision-based 3d occupancy prediction,

    Y . Huang, W. Zheng, B. Zhang, J. Zhou, and J. Lu, “Selfocc: Self-supervised vision-based 3d occupancy prediction,” inProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024, pp. 19 946–19 956

  67. [75]

    Renderworld: World model with self-supervised 3d label,

    Z. Yan, W. Dong, Y . Shao, Y . Lu, L. Haiyang, J. Liu, H. Wang, Z. Wang, Y . Wang, F. Remondinoet al., “Renderworld: World model with self-supervised 3d label,”arXiv preprint arXiv:2409.11356, 2024

  68. [76]

    Gaussian- flowocc: Sparse and weakly supervised occupancy es- timation using gaussian splatting and temporal flow,

    S. Boeder, F. Gigengack, and B. Risse, “Gaussian- flowocc: Sparse and weakly supervised occupancy es- timation using gaussian splatting and temporal flow,” arXiv preprint arXiv:2502.17288, 2025

  69. [77]

    Street gaussians: Modeling dynamic urban scenes with gaussian splat- ting,

    Y . Yan, H. Lin, C. Zhou, W. Wang, H. Sun, K. Zhan, X. Lang, X. Zhou, and S. Peng, “Street gaussians: Modeling dynamic urban scenes with gaussian splat- ting,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 156–173

  70. [78]

    Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy prediction,

    Y . Huang, W. Zheng, Y . Zhang, J. Zhou, and J. Lu, “Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy prediction,” inEuropean Con- ference on Computer Vision. Springer, 2024, pp. 376– 393

  71. [79]

    Surroundocc: Multi-camera 3d occupancy pre- diction for autonomous driving,

    Y . Wei, L. Zhao, W. Zheng, Z. Zhu, J. Zhou, and J. Lu, “Surroundocc: Multi-camera 3d occupancy pre- diction for autonomous driving,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 21 729–21 740

  72. [80]

    Uno: Unsupervised occupancy fields for per- ception and forecasting,

    B. Agro, Q. Sykora, S. Casas, T. Gilles, and R. Ur- tasun, “Uno: Unsupervised occupancy fields for per- ception and forecasting,” inCVPR, 2024

  73. [81]

    Maeli: Masked au- toencoder for large-scale lidar point clouds,

    G. Krispel, D. Schinagl, C. Fruhwirth-Reisinger, H. Possegger, and H. Bischof, “Maeli: Masked au- toencoder for large-scale lidar point clouds,” inPro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 3383– 3392

  74. [82]

    Masked autoencoder for self-supervised pre-training on lidar point clouds,

    G. Hess, J. Jaxing, E. Svensson, D. Hagerman, C. Pe- tersson, and L. Svensson, “Masked autoencoder for self-supervised pre-training on lidar point clouds,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 350–359

  75. [83]

    Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving sce- narios,

    Z. Lin, Y . Wang, S. Qi, N. Dong, and M.-H. Yang, “Bev-mae: Bird’s eye view masked autoencoders for point cloud pre-training in autonomous driving sce- narios,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 3531– VOLUME 00, 2024 29 ...

  76. [84]

    Geo- mae: Masked geometric target prediction for self- supervised point cloud pre-training,

    X. Tian, H. Ran, Y . Wang, and H. Zhao, “Geo- mae: Masked geometric target prediction for self- supervised point cloud pre-training,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13 570–13 580

  77. [85]

    Openoccupancy: A large scale benchmark for surrounding semantic oc- cupancy perception,

    X. Wang, Z. Zhu, W. Xu, Y . Zhang, Y . Wei, X. Chi, Y . Ye, D. Du, J. Lu, and X. Wang, “Openoccupancy: A large scale benchmark for surrounding semantic oc- cupancy perception,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17 850–17 859

  78. [86]

    Occ3d: A large-scale 3d occu- pancy prediction benchmark for autonomous driving,

    X. Tian, T. Jiang, L. Yun, Y . Mao, H. Yang, Y . Wang, Y . Wang, and H. Zhao, “Occ3d: A large-scale 3d occu- pancy prediction benchmark for autonomous driving,” Advances in Neural Information Processing Systems, vol. 36, pp. 64 318–64 330, 2023

  79. [87]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”arXiv preprint arXiv:1807.03748, 2018

  80. [88]

    Cross-modal contrastive learning for domain adap- tation in 3d semantic segmentation,

    B. Xing, X. Ying, R. Wang, J. Yang, and T. Chen, “Cross-modal contrastive learning for domain adap- tation in 3d semantic segmentation,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, 2023, pp. 2974–2982

  81. [89]

    4d contrastive superflows are dense 3d representation learners,

    X. Xu, L. Kong, H. Shuai, W. Zhang, L. Pan, K. Chen, Z. Liu, and Q. Liu, “4d contrastive superflows are dense 3d representation learners,” inEuropean Con- ference on Computer Vision. Springer, 2024, pp. 58– 80

  82. [90]

    Superflow++: Enhanced spatiotemporal con- sistency for cross-modal data pretraining,

    ——, “Superflow++: Enhanced spatiotemporal con- sistency for cross-modal data pretraining,”arXiv e- prints, pp. arXiv–2503, 2025

  83. [91]

    Contrastalign: Toward robust bev feature alignment via contrastive learning for multi-modal 3d object detection,

    Z. Song, F. Jia, H. Pan, Y . Luo, C. Jia, G. Zhang, L. Liu, Y . Ji, L. Yang, and L. Wang, “Contrastalign: Toward robust bev feature alignment via contrastive learning for multi-modal 3d object detection,”arXiv preprint arXiv:2405.16873, 2024

  84. [92]

    Bevdistill: Cross-modal bev distillation for multi-view 3d object detection,

    Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao, “Bevdistill: Cross-modal bev distillation for multi-view 3d object detection,”arXiv preprint arXiv:2211.09386, 2022

  85. [93]

    Distill- bev: Boosting multi-camera 3d object detection with cross-modal knowledge distillation,

    Z. Wang, D. Li, C. Luo, C. Xie, and X. Yang, “Distill- bev: Boosting multi-camera 3d object detection with cross-modal knowledge distillation,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8637–8646

  86. [94]

    Geometric-aware pretraining for vision-centric 3d object detection,

    L. Huang, H. Wang, J. Zeng, S. Zhang, L. Cao, J. Yan, and H. Li, “Geometric-aware pretraining for vision-centric 3d object detection,”arXiv preprint arXiv:2304.03105, 2023

  87. [95]

    Unidis- till: A universal cross-modality knowledge distillation framework for 3d object detection in bird’s-eye view,

    S. Zhou, W. Liu, C. Hu, S. Zhou, and C. Ma, “Unidis- till: A universal cross-modality knowledge distillation framework for 3d object detection in bird’s-eye view,” inProceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, 2023, pp. 5116– 5125

  88. [96]

    Revisiting domain generalized stereo match- ing networks from a feature consistency perspec- tive,

    J. Zhang, X. Wang, X. Bai, C. Wang, L. Huang, Y . Chen, L. Gu, J. Zhou, T. Harada, and E. R. Han- cock, “Revisiting domain generalized stereo match- ing networks from a feature consistency perspec- tive,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  89. [97]

    Weakly supervised monocular 3d object detection using multi-view projection and direction consis- tency,

    R. Tao, W. Han, Z. Qiu, C.-z. Xu, and J. Shen, “Weakly supervised monocular 3d object detection using multi-view projection and direction consis- tency,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 482–17 492

  90. [98]

    Bevformer: learning bird’s-eye- view representation from lidar-camera via spatiotem- poral transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “Bevformer: learning bird’s-eye- view representation from lidar-camera via spatiotem- poral transformers,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  91. [99]

    Ega- depth: Efficient guided attention for self-supervised multi-camera depth estimation,

    Y . Shi, H. Cai, A. Ansari, and F. Porikli, “Ega- depth: Efficient guided attention for self-supervised multi-camera depth estimation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 119–129

  92. [100]

    Multimae: Multi-modal multi-task masked autoen- coders,

    R. Bachmann, D. Mizrahi, A. Atanov, and A. Zamir, “Multimae: Multi-modal multi-task masked autoen- coders,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 348–367

  93. [101]

    Pimae: Point cloud and image interactive masked autoencoders for 3d object detec- tion,

    A. Chen, K. Zhang, R. Zhang, Z. Wang, Y . Lu, Y . Guo, and S. Zhang, “Pimae: Point cloud and image interactive masked autoencoders for 3d object detec- tion,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 5291–5301

  94. [102]

    Unim 2 ae: Multi-modal masked autoencoders with unified 3d representation for 3d perception in autonomous driving,

    J. Zou, T. Huang, G. Yang, Z. Guo, T. Luo, C.- M. Feng, and W. Zuo, “Unim 2 ae: Multi-modal masked autoencoders with unified 3d representation for 3d perception in autonomous driving,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 296–313

  95. [103]

    Occgen: Generative multi-modal 3d occupancy prediction for autonomous driving,

    G. Wang, Z. Wang, P. Tang, J. Zheng, X. Ren, B. Feng, and C. Ma, “Occgen: Generative multi-modal 3d occupancy prediction for autonomous driving,” in European Conference on Computer Vision. Springer, 2024, pp. 95–112

  96. [104]

    Dif- fusion model for robust multi-sensor fusion in 3d object detection and bev segmentation,

    D.-T. Le, H. Shi, J. Cai, and H. Rezatofighi, “Dif- fusion model for robust multi-sensor fusion in 3d object detection and bev segmentation,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 232–249

  97. [105]

    Pys- ical informed driving world model,

    Z. Yang, X. Guo, C. Ding, C. Wang, and W. Wu, “Pys- ical informed driving world model,”arXiv preprint arXiv:2412.08410, 2024. 30 VOLUME 00, 2024

  98. [106]

    Visual point cloud forecasting enables scalable autonomous driv- ing,

    Z. Yang, L. Chen, Y . Sun, and H. Li, “Visual point cloud forecasting enables scalable autonomous driv- ing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 673–14 684

  99. [107]

    Uni- world: Autonomous driving pre-training via world models,

    C. Min, D. Zhao, L. Xiao, Y . Nie, and B. Dai, “Uni- world: Autonomous driving pre-training via world models,”arXiv preprint arXiv:2308.07234, 2023

  100. [108]

    Learning unsupervised world models for autonomous driving via discrete diffusion,

    L. Zhang, Y . Xiong, Z. Yang, S. Casas, R. Hu, and R. Urtasun, “Learning unsupervised world models for autonomous driving via discrete diffusion,”ICLR, 2024

  101. [109]

    Bev- world: A multimodal world model for autonomous driving via unified bev latent space,

    Y . Zhang, S. Gong, K. Xiong, X. Ye, X. Tan, F. Wang, J. Huang, H. Wu, and H. Wang, “Bev- world: A multimodal world model for autonomous driving via unified bev latent space,”arXiv preprint arXiv:2407.05679, 2024

  102. [110]

    Driving into the future: Multiview visual forecasting and planning with world model for autonomous driv- ing,

    Y . Wang, J. He, L. Fan, H. Li, Y . Chen, and Z. Zhang, “Driving into the future: Multiview visual forecasting and planning with world model for autonomous driv- ing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 749–14 759

  103. [111]

    Temporal consistent 3d lidar representation learning for semantic per- ception in autonomous driving,

    L. Nunes, L. Wiesmann, R. Marcuzzi, X. Chen, J. Behley, and C. Stachniss, “Temporal consistent 3d lidar representation learning for semantic per- ception in autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 5217–5228

  104. [112]

    Compass: Contrastive multimodal pretraining for autonomous systems,

    S. Ma, S. Vemprala, W. Wang, J. K. Gupta, Y . Song, D. McDufft, and A. Kapoor, “Compass: Contrastive multimodal pretraining for autonomous systems,” in 2022 IEEE/RSJ International Conference on Intelli- gent Robots and Systems (IROS). IEEE, 2022, pp. 1000–1007

  105. [113]

    Drivevlm: The convergence of autonomous driving and large vision- language models,

    X. Tian, J. Gu, B. Li, Y . Liu, Y . Wang, Z. Zhao, K. Zhan, P. Jia, X. Lang, and H. Zhao, “Drivevlm: The convergence of autonomous driving and large vision- language models,”arXiv preprint arXiv:2402.12289, 2024

  106. [114]

    Scalability in perception for autonomous driv- ing: Waymo open dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine et al., “Scalability in perception for autonomous driv- ing: Waymo open dataset,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2...

  107. [115]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inProceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2020, pp. 11 621–11 631

  108. [116]

    Argoverse: 3d tracking and forecasting with rich maps,

    M.-F. Chang, J. Lambert, P. Sangkloy, J. Singh, S. Bak, A. Hartnett, D. Wang, P. Carr, S. Lucey, D. Ramananet al., “Argoverse: 3d tracking and forecasting with rich maps,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 8748–8757

  109. [117]

    Are we ready for autonomous driving? the kitti vision benchmark suite,

    A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 3354–3361

  110. [118]

    Bdd100k: A diverse driving dataset for heterogeneous multitask learning,

    F. Yu, H. Chen, X. Wang, W. Xian, Y . Chen, F. Liu, V . Madhavan, and T. Darrell, “Bdd100k: A diverse driving dataset for heterogeneous multitask learning,” inProceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, 2020, pp. 2636– 2645

  111. [119]

    Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving,

    T. Wang, S. Kim, W. Ji, E. Xie, C. Ge, J. Chen, Z. Li, and L. Ping, “Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving,” arXiv preprint arXiv:2304.01168, 2023

  112. [120]

    Realistic corner case generation for autonomous ve- hicles with multimodal large language model,

    Q. Lu, M. Ma, X. Dai, X. Wang, and S. Feng, “Realistic corner case generation for autonomous ve- hicles with multimodal large language model,”arXiv preprint arXiv:2412.00243, 2024

  113. [121]

    Pesotif: A challenging visual dataset for perception sotif prob- lems in long-tail traffic scenarios,

    L. Peng, J. Li, W. Shao, and H. Wang, “Pesotif: A challenging visual dataset for perception sotif prob- lems in long-tail traffic scenarios,” in2023 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2023, pp. 1–8

  114. [122]

    Carla: An open urban driving simulator,

    A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun, “Carla: An open urban driving simulator,” inConference on robot learning. PMLR, 2017, pp. 1–16

  115. [123]

    Language prompt for autonomous driv- ing,

    D. Wu, W. Han, Y . Liu, T. Wang, C.-z. Xu, X. Zhang, and J. Shen, “Language prompt for autonomous driv- ing,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 8, 2025, pp. 8359– 8367

  116. [124]

    Anovox: A benchmark for multimodal anomaly detection in autonomous driv- ing,

    D. Bogdoll, I. Hamdard, L. N. R ¨oßler, F. Geisler, M. Bayram, F. Wang, J. Imhof, M. de Campos, A. Tabarov, Y . Yanget al., “Anovox: A benchmark for multimodal anomaly detection in autonomous driv- ing,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 206–223

  117. [125]

    Msc-bench: Benchmark- ing and analyzing multi-sensor corruption for driving perception,

    X. Hao, G. Liu, Y . Zhao, Y . Ji, M. Wei, H. Zhao, L. Kong, R. Yin, and Y . Liu, “Msc-bench: Benchmark- ing and analyzing multi-sensor corruption for driving perception,”arXiv preprint arXiv:2501.01037, 2025

  118. [126]

    Benchmarking robustness of 3d object detection to common corrup- tions,

    Y . Dong, C. Kang, J. Zhang, Z. Zhu, Y . Wang, X. Yang, H. Su, X. Wei, and J. Zhu, “Benchmarking robustness of 3d object detection to common corrup- tions,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1022–1032

  119. [127]

    Robo3d: Towards robust and reliable 3d perception against corruptions,

    L. Kong, Y . Liu, X. Li, R. Chen, W. Zhang, J. Ren, L. Pan, K. Chen, and Z. Liu, “Robo3d: Towards robust and reliable 3d perception against corruptions,” in Proceedings of the IEEE/CVF International Confer- VOLUME 00, 2024 31 Authoret al.: Preparation of Papers for IEEE OPEN J...

  120. [128]

    Robobev: Towards robust bird’s eye view perception under corruptions,

    S. Xie, L. Kong, W. Zhang, J. Ren, L. Pan, K. Chen, and Z. Liu, “Robobev: Towards robust bird’s eye view perception under corruptions,”arXiv preprint arXiv:2304.06719, 2023

  121. [129]

    Unity is strength? benchmarking the robustness of fusion-based 3d object detection against physical sensor attack,

    Z. Jin, X. Lu, B. Yang, Y . Cheng, C. Yan, X. Ji, and W. Xu, “Unity is strength? benchmarking the robustness of fusion-based 3d object detection against physical sensor attack,” inProceedings of the ACM Web Conference 2024, 2024, pp. 3031–3042

  122. [130]

    A benchmark for unsupervised anomaly detection in multi-agent trajectories,

    J. Wiederer, J. Schmidt, U. Kressel, K. Dietmayer, and V . Belagiannis, “A benchmark for unsupervised anomaly detection in multi-agent trajectories,” in2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2022, pp. 130– 137

  123. [131]

    A survey of model compression and acceleration for deep neural networks,

    Y . Cheng, D. Wang, P. Zhou, and T. Zhang, “A survey of model compression and acceleration for deep neural networks,”arXiv preprint arXiv:1710.09282, 2017

  124. [132]

    Dis- tilling step-by-step! outperforming larger language models with less training data and smaller model sizes,

    C.-Y . Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y . Fujii, A. Ratner, R. Krishna, C.-Y . Lee, and T. Pfister, “Dis- tilling step-by-step! outperforming larger language models with less training data and smaller model sizes,”arXiv preprint arXiv:2305.02301, 2023

  125. [133]

    Tinyvit: Fast pretraining distillation for small vision transformers,

    K. Wu, J. Zhang, H. Peng, M. Liu, B. Xiao, J. Fu, and L. Yuan, “Tinyvit: Fast pretraining distillation for small vision transformers,” inEuropean conference on computer vision. Springer, 2022, pp. 68–85

  126. [134]

    Distilling large vision-language model with out-of- distribution generalizability,

    X. Li, Y . Fang, M. Liu, Z. Ling, Z. Tu, and H. Su, “Distilling large vision-language model with out-of- distribution generalizability,” inProceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2023, pp. 2492–2503

  127. [135]

    Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,

    T. Dettmers, M. Lewis, Y . Belkada, and L. Zettle- moyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”Advances in neural information processing systems, vol. 35, pp. 30 318–30 332, 2022

  128. [136]

    Repq-vit: Scale reparameterization for post-training quantization of vision transformers,

    Z. Li, J. Xiao, L. Yang, and Q. Gu, “Repq-vit: Scale reparameterization for post-training quantization of vision transformers,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17 227–17 236

  129. [137]

    A simple and effective pruning approach for large language models,

    M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,”arXiv preprint arXiv:2306.11695, 2023

  130. [138]

    X-pruner: explainable prun- ing for vision transformers,

    L. Yu and W. Xiang, “X-pruner: explainable prun- ing for vision transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 24 355–24 363

  131. [139]

    Lpu: A latency-optimized and highly scalable processor for large language model inference,

    S. Moon, J.-H. Kim, J. Kim, S. Hong, J. Cha, M. Kim, S. Lim, G. Choi, D. Seo, J. Kimet al., “Lpu: A latency-optimized and highly scalable processor for large language model inference,”IEEE Micro, 2024

  132. [140]

    Prophet: Realizing a predictable real-time perception pipeline for autonomous vehicles,

    L. Liu, Z. Dong, Y . Wang, and W. Shi, “Prophet: Realizing a predictable real-time perception pipeline for autonomous vehicles,” in2022 IEEE Real-Time Systems Symposium (RTSS). IEEE, 2022, pp. 305– 317

  133. [141]

    Shallow- deep networks: Understanding and mitigating network overthinking,

    Y . Kaya, S. Hong, and T. Dumitras, “Shallow- deep networks: Understanding and mitigating network overthinking,” inInternational conference on machine learning. PMLR, 2019, pp. 3301–3310

  134. [142]

    Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos,

    H. Liu, Y . Teng, T. Lu, H. Wang, and L. Wang, “Sparsebev: High-performance sparse 3d object de- tection from multi-camera videos,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 18 580–18 590

  135. [143]

    Dynamicvit: Efficient vision transform- ers with dynamic token sparsification,

    Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.- J. Hsieh, “Dynamicvit: Efficient vision transform- ers with dynamic token sparsification,”Advances in neural information processing systems, vol. 34, pp. 13 937–13 949, 2021

  136. [144]

    Toward efficient inference for mixture of experts,

    H. Huang, N. Ardalani, A. Sun, L. Ke, S. Bhosale, H.-H. Lee, C.-J. Wu, and B. Lee, “Toward efficient inference for mixture of experts,”Advances in Neural Information Processing Systems, vol. 37, pp. 84 033– 84 059, 2024. 32 VOLUME 00, 2024

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.