Pith. sign in

REVIEW 4 major objections 5 minor 153 references

A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A task-agnostic survey organizes multi-sensor fusion perception for embodied AI into four technical families, independent of any single task or application domain.

desk verdict The task-agnostic organizing scheme is genuinely useful, but the timeline in Fig. 8 lists many methods that are never cited or discussed, so the survey's claim to be a rigorous map is not yet supported. read the letter →

arxiv 2506.19769 v1 pith:BNGFL3GV submitted 2025-06-24 cs.MM cs.AI

classification cs.MMcs.AI
keywords multi-sensorfusionembodiedAImulti-modalmulti-agentcollaborativeperceptiontime-seriesmultimodallargelanguagemodelsautonomousdrivingsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that multi-sensor fusion perception can be surveyed in a task-agnostic way, so that the same fusion techniques are visible to researchers working on any embodied-AI problem rather than only to those in autonomous driving or 3D detection. It argues that previous surveys are mostly task-specific or single-perspective and miss multi-agent fusion, time-series fusion, and large-model-based fusion. The authors organize the field into four technical families, covering the background, datasets, perception tasks, methods, and open challenges. If the survey is right, it gives researchers a shared vocabulary and a reference map for fusion across robotics, autonomous driving, and other embodied systems.

What carries the argument

The organizing device is a task-agnostic taxonomy built on the 'Agent-Sensor-Data-Model-Task' pipeline. The four categories are multi-modal fusion, multi-agent fusion, time-series fusion, and MM-LLM fusion, with time-series methods further split into dense query, sparse query, and hybrid query approaches. The taxonomy does the argument's work: it is what lets the survey claim that fusion techniques transfer across tasks rather than belonging to one application area.

What would settle it

Take a systematic sample of recent multi-sensor fusion papers from areas outside autonomous driving, such as thermal-visual pedestrian detection, visual-inertial odometry, or infrastructure-vehicle cooperation, and check whether each paper fits one of the four categories; any substantial cluster that does not fit would falsify the claim of task-agnostic coverage.

Watch

Extended reading notes

Core claim

The central claim is that the diversity of multi-sensor fusion perception can be presented independently of any downstream task, using four technical families: multi-modal fusion (at point, voxel, region, and multi-level), multi-agent fusion, time-series fusion (with dense, sparse, and hybrid query representations), and multimodal-LLM fusion (vision-language and vision-LiDAR-language). The paper argues that this organization lets any embodied-AI researcher, whatever their task, find the fusion technique relevant to them. It supports the claim by reviewing representative methods in each family, cataloging datasets and evaluation criteria, and discussing open challenges at the data, model, and application levels.

Load-bearing premise

The survey's usefulness assumes that the methods it selected are a representative sample of the whole field and that its four-category taxonomy really captures the variety of multi-sensor fusion research.

Editorial extensions

If this is right

  • Researchers in tasks outside autonomous driving can locate applicable fusion techniques through the task-agnostic categories.
  • The four-way split makes visible method families, such as multi-agent collaboration and time-series fusion, that single-task surveys tend to omit.
  • The dense-versus-sparse-versus-hybrid query taxonomy offers a direct way to compare efficiency and accuracy trade-offs in temporal fusion.
  • The challenge discussion points to concrete research targets, including synchronized multi-modal data augmentation, explainable fusion, and handling sparse radar or LiDAR data within multimodal LLMs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A cross-cutting axis the paper leaves implicit is fusion stage, such as data-level versus feature-level versus decision-level fusion; readers could combine that axis with the four families to locate methods even more precisely.
  • If the task-agnostic claim holds, specialized survey writers could reuse the four categories as a standard outline, letting method results accumulate across domains instead of being fragmented by task.
  • The paper does not describe its literature search or inclusion criteria, so absence of a method family from the survey should be read as an open question rather than proof that the family does not exist.
  • A natural test of the taxonomy is whether papers on less common sensor pairs, such as thermal-RGB pedestrian detection or visual-inertial odometry, can be placed cleanly into one of the four families.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This survey proposes a task-agnostic organization of multi-sensor fusion perception (MSFP) methods for embodied AI, structured around four technical views: multi-modal fusion (point-, voxel-, region-, and multi-level), multi-agent fusion, time-series fusion, and multimodal-LLM fusion. It provides background on sensors, datasets, and perception tasks, reviews representative methods in each category, and closes with challenges and future directions at the data, model, and application levels. The central claim is that existing surveys are too narrowly tied to single tasks (e.g., 3D object detection) or single perspectives (mainly multi-modal fusion), and that a task-agnostic, multi-perspective organization is needed for researchers across fields.

Significance. If the survey's organization were fully substantiated, it would fill a genuine gap: it covers time-series fusion and MM-LLM fusion as first-class categories, which most prior MSFP surveys do not, and it attempts a useful cross-task framing. The authors also gather relevant background on sensors and datasets, and the discussion of challenges (data quality, synchronization, explainability) is sensible. However, the paper's core value depends on the reliability of its method corpus and taxonomy, and the current manuscript does not yet demonstrate that reliability: the selection methodology is undocumented, several timeline entries are unreferenced, and key figures and tables contradict the text. These are fixable issues, but they are load-bearing for a survey claiming to be a comprehensive reference.

major comments (4)
  1. [Section I and Abstract] The paper claims a 'rigorous and detailed investigation' (Abstract) and 'a detailed investigation' (Section I), but it does not report the literature search strategy, inclusion/exclusion criteria, or the procedure by which the four-category taxonomy (multi-modal, multi-agent, time-series, MM-LLM) was derived. Without this methodology, the reader cannot determine whether the methods selected for Tables III–V and Figs. 7–8 are representative or exhaustive. Since the survey's stated value proposition is to provide a trustworthy task-agnostic map of MSFP methods, this missing documentation undermines the central claim of the paper.
  2. [Section V, Fig. 8] The timeline in Fig. 8 lists at least a dozen methods—PolarDETR, STS, MV-FCOS3D++, DORT, E-TMA, BridgeAD, SAD, RENet, Far3D, TLCFuse, PETR, and PETR v2—that are not cited in the reference list and are not discussed in the text. In addition, Fig. 8 contains an 'Others' group that has no counterpart in the three categories of Table IV, and no criterion is stated for what belongs in the taxonomy versus 'Others.' Because the survey claims comprehensiveness, these unreferenced entries and the unexplained extra category make the map unreliable and directly undercut the abstract's claim of a rigorous investigation.
  3. [Section V, Fig. 7] The first paragraph of Section V states 'Fig. 7 shows a simple pipeline of A2A fusion,' but Fig. 7 is placed in the time-series section and is captioned 'Framework overview of time series multi-sensor fusion network.' This is not an isolated typo: the confusion between agent-to-agent fusion (Section IV) and time-series fusion (Section V) matters because the distinction between these technical views is one of the paper's organizing axes. The figure reference and caption must be corrected so that each figure illustrates the correct category.
  4. [Section V, Table IV] Table IV's 'Methods' column is inconsistent with the text of Section V. The text discusses MUTR3D [97], PF-Track [98], FusionFormer [99], QTNet [101], and CRT-Fusion [102] as query-based time-series methods, but none of these appears in Table IV, while some table entries (e.g., SparseFusion3D) are discussed in the text. If the table is intended to be representative, that should be stated; if it is intended to be complete, these are omissions. In either case, the table cannot currently serve as a reliable quick-reference, which is the primary function of a survey table.
minor comments (5)
  1. [Section IV, first paragraph] The sentence 'we will focus on the multi-view fusion of agent-to-agent (A2A) collaborative perception' conflates multi-view fusion (multiple cameras on one agent) with multi-agent fusion (multiple agents sharing information). Please clarify or rephrase to avoid conflating these two distinct technical views.
  2. [References [89] and [100]] The DETR3D paper appears twice as separate references ([89] and [100]) with different formatting and slightly different author lists; this should be consolidated into a single reference.
  3. [Section III, Table III] Some methods listed in Table III (e.g., UVTR, SFD, E2E-MFD, MBNet) are not discussed in the body text, while several methods discussed in the text (e.g., PI-RCNN, FusionPainting, GraphAlign, VPFNet, VFF, AutoAlign, VoxelNextFusion, TransFusion, GAFF, RSDet, EPNet, DVF, LoGoNet, CAT-Det, SeaDATE, Fusion-Mamba) are not listed in the table. Please align the table with the text or state explicitly that the table is representative rather than exhaustive.
  4. [Fig. 8 and general figures] Fig. 8 is dense and the small font makes the method names difficult to read; consider a larger layout or a table-form timeline. Also, several figures (Figs. 2–10) are not explicitly referenced in the text at the point of discussion, which makes navigation harder.
  5. [Throughout] Capitalization and terminology are inconsistent in a few places, e.g., 'V oxel-level' in Table III and 'LIDAR' vs. 'LiDAR' in Section III-D; a careful proofread would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's organizational claims are descriptive syntheses of external literature, with no self-referential derivation or fitted prediction.

full rationale

This paper is a survey, not a derivation. It organizes existing multi-sensor fusion perception methods into descriptive taxonomies (multi-modal, multi-agent, time-series, MM-LLM fusion) and discusses challenges. There is no equation, no fitted parameter, and no predictive claim whose output is defined by its input. The authors cite their own prior work ([1]–[3]) only as general examples of AI progress in the introduction; those citations do not support the survey's structural claims and are not load-bearing. The absence of a documented literature-search protocol and the presence of timeline entries in Fig. 8 that are not discussed in the text are completeness and rigor concerns, but they are not circularity: an incomplete or unrepresentative survey is not a survey that derives its conclusions from its own assumptions. No step in the paper reduces to a self-citation, a renamed empirical pattern, or a fitted input called a prediction. The central organizational claim—that the survey is task-agnostic and broader than prior single-task surveys—is a comparative editorial assertion about scope, not a formal result requiring derivation. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

A survey contributes no free parameters or invented entities. It relies on the accuracy of its summaries, the validity of its taxonomy, and the representativeness of its selected methods, all of which are taken as assumptions rather than demonstrated.

assumptions (3)
  • domain assumption The descriptions of cited methods in Sections III-VI accurately reflect the original papers.
    The survey's utility depends on faithful summarization, but no verification against original implementations or evaluations is provided.
  • domain assumption The four-category taxonomy (multi-modal, multi-agent, time-series, MM-LLM) and subcategories (point, voxel, region, multi-level) are appropriate and non-overlapping.
    The classification scheme is asserted without justification or comparison to alternative taxonomies.
  • domain assumption The set of selected methods is representative of the field.
    No systematic search or inclusion criteria are described, so representativeness is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects." pith.science (2026). https://pith.science/paper/BNGFL3GV

@misc{pith2026250619769,
  author       = {Pith},
  title        = {Pith review of: A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BNGFL3GV}},
  note         = {Machine review of arXiv:2506.19769}
}
read the original abstract

Multi-sensor fusion perception (MSFP) is a key technology for embodied AI, which can serve a variety of downstream tasks (e.g., 3D object detection and semantic segmentation) and application scenarios (e.g., autonomous driving and swarm robotics). Recently, impressive achievements on AI-based MSFP methods have been reviewed in relevant surveys. However, we observe that the existing surveys have some limitations after a rigorous and detailed investigation. For one thing, most surveys are oriented to a single task or research field, such as 3D object detection or autonomous driving. Therefore, researchers in other related tasks often find it difficult to benefit directly. For another, most surveys only introduce MSFP from a single perspective of multi-modal fusion, while lacking consideration of the diversity of MSFP methods, such as multi-view fusion and time-series fusion. To this end, in this paper, we hope to organize MSFP research from a task-agnostic perspective, where methods are reported from various technical views. Specifically, we first introduce the background of MSFP. Next, we review multi-modal and multi-agent fusion methods. A step further, time-series fusion methods are analyzed. In the era of LLM, we also investigate multimodal LLM fusion methods. Finally, we discuss open challenges and future directions for MSFP. We hope this survey can help researchers understand the important progress in MSFP and provide possible insights for future research.

Figures

Figures reproduced from arXiv: 2506.19769 by the authors.

Figure 1
Figure 1. Overview of multi-sensor fusion perception pipeline. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 3
Figure 3. Overview of voxel-level fusion pipeline. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Overview of region-level fusion pipeline. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (4 more)
Figure 6
Figure 6. Figure 6: A simple agent-to-agent (A2A) fusion pipeline. [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 8
Figure 8. Figure 8: Timeline of time series fusion methods. cific tasks. Based on [76], [77], [91], [92], UniFusion [94] puts forward a unified temporal-spatial fusion framework, introducing the concept of virtual views. Historical frames are viewed as additional camera views with spatial…
Figure 9
Figure 9. Figure 9: Visual-Language based paradigm. OmniDrive [113], and NuInstruct enhance existing datasets by incorporating large language models (LLMs) to generate question-answer pairs covering perception, reasoning, and planning. Additionally, MAPLM [110] integrates multi-view image…
Figure 10
Figure 10. Figure 10: Visual-LiDAR-Language based paradigms. In paradigm (a), separate encoders are designed for radar and image, and then fused. Paradigm (b) fuses [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

153 extracted references · 44 canonical work pages

  1. [97]

    Hvdetfusion: A simple and robust camera-radar fusion framework,

    K. C. Lei, Z. Chen, S. Jia, and X. Zhang, “Hvdetfusion: A simple and robust camera-radar fusion framework,”arXiv preprint arXiv:2307.11323, 2023

  2. [98]

    Mutr3d: A multi- camera tracking framework via 3d-to-2d queries,

    T. Zhang, X. Chen, Y . Wang, Y . Wang, and H. Zhao, “Mutr3d: A multi- camera tracking framework via 3d-to-2d queries,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4537–4546

  3. [99]

    Standing between past and future: Spatio-temporal modeling for multi-camera 3d multi-object tracking,

    Z. Pang, J. Li, P. Tokmakov, D. Chen, S. Zagoruyko, and Y .-X. Wang, “Standing between past and future: Spatio-temporal modeling for multi-camera 3d multi-object tracking,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

  4. [101]

    Detr3d: 3d object detection from multi-view images via 3d-to-2d queries,

    Y . Wang, V . C. Guizilini, T. Zhang, Y . Wang, H. Zhao, and J. Solomon, “Detr3d: 3d object detection from multi-view images via 3d-to-2d queries,” inConference on Robot Learning. PMLR, 2022, pp. 180– 191

  5. [102]

    Query-based temporal fusion with explicit motion for 3d object detection,

    J. Hou, Z. Liu, Z. Zou, X. Ye, X. Baiet al., “Query-based temporal fusion with explicit motion for 3d object detection,”Advances in Neural Information Processing Systems, vol. 36, pp. 75 782–75 797, 2023

  6. [1]

    Dae-gan: Dynamic aspect-aware gan for text-to-image synthesis,

    S. Ruan, Y . Zhang, K. Zhang, Y . Fan, F. Tang, Q. Liu, and E. Chen, “Dae-gan: Dynamic aspect-aware gan for text-to-image synthesis,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 13 960–13 969

  7. [2]

    Cpws: Confident programmatic weak supervision for high-quality data labeling,

    S. Ruan, H. Liu, Z. Chen, B. Feng, K. Zhang, C. C. Cao, E. Chen, and L. Chen, “Cpws: Confident programmatic weak supervision for high-quality data labeling,”ACM Transactions on Information Systems, vol. 43, no. 4, pp. 1–26, 2025

  8. [3]

    Color enhanced cross correlation net for image sentiment analysis,

    S. Ruan, K. Zhang, L. Wu, T. Xu, Q. Liu, and E. Chen, “Color enhanced cross correlation net for image sentiment analysis,”IEEE Transactions on Multimedia, vol. 26, pp. 4097–4109, 2024

Show all 153 references
  1. [4]

    Embodied intelli- gence toward future smart manufacturing in the era of ai foundation model,

    L. Ren, J. Dong, S. Liu, L. Zhang, and L. Wang, “Embodied intelli- gence toward future smart manufacturing in the era of ai foundation model,”IEEE/ASME Transactions on Mechatronics, 2024

  2. [5]

    Embodied intel- ligence via learning and evolution,

    A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei, “Embodied intel- ligence via learning and evolution,”Nature communications, vol. 12, no. 1, p. 5721, 2021

  3. [6]

    Multi-modal 3d object detection in autonomous driving: a survey,

    Y . Wang, Q. Mao, H. Zhu, J. Deng, Y . Zhang, J. Ji, H. Li, and Y . Zhang, “Multi-modal 3d object detection in autonomous driving: a survey,” International Journal of Computer Vision, vol. 131, no. 8, pp. 2122– 2152, 2023

  4. [7]

    Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,

    L. Wang, X. Zhang, Z. Song, J. Bi, G. Zhang, H. Wei, L. Tang, L. Yang, J. Li, C. Jiaet al., “Multi-modal 3d object detection in autonomous driving: A survey and taxonomy,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 7, pp. 3781–3798, 2023

  5. [8]

    Multi-sensor fusion and cooperative perception for autonomous driv- ing: A review,

    C. Xiang, C. Feng, X. Xie, B. Shi, H. Lu, Y . Lv, M. Yang, and Z. Niu, “Multi-sensor fusion and cooperative perception for autonomous driv- ing: A review,”IEEE Intelligent Transportation Systems Magazine, 2023

  6. [9]

    Camera, lidar, and imu based multi-sensor fusion slam: A survey,

    J. Zhu, H. Li, and T. Zhang, “Camera, lidar, and imu based multi-sensor fusion slam: A survey,”Tsinghua Science and Technology, vol. 29, no. 2, pp. 415–429, 2023

  7. [10]

    Multi-modality 3d object detection in autonomous driving: A review,

    Y . Tang, H. He, Y . Wang, Z. Mao, and H. Wang, “Multi-modality 3d object detection in autonomous driving: A review,”Neurocomputing, p. 126587, 2023

  8. [11]

    Advancements in perception system with multi-sensor fusion for embodied agents,

    H. Du, L. Ren, Y . Wang, X. Cao, and C. Sun, “Advancements in perception system with multi-sensor fusion for embodied agents,” Information Fusion, p. 102859, 2024

  9. [12]

    A survey on the visual perception of humanoid robot,

    T. Bin, H. Yan, N. Wang, M. N. Nikoli ´c, J. Yao, and T. Zhang, “A survey on the visual perception of humanoid robot,”Biomimetic Intelligence and Robotics, p. 100197, 2024

  10. [13]

    Robustness-aware 3d object detection in autonomous driving: A review and outlook,

    Z. Song, L. Liu, F. Jia, Y . Luo, C. Jia, G. Zhang, L. Yang, and L. Wang, “Robustness-aware 3d object detection in autonomous driving: A review and outlook,”IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 11, pp. 15 407–15 436, 2024

  11. [14]

    Multimodal fusion and vision-language models: A survey for robot vision,

    X. Han, S. Chen, Z. Fu, Z. Feng, L. Fan, D. An, C. Wang, L. Guo, W. Meng, X. Zhanget al., “Multimodal fusion and vision-language models: A survey for robot vision,”arXiv preprint arXiv:2504.02477, 2025

  12. [15]

    Are we ready for autonomous driving? the kitti vision benchmark suite,

    A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in2012 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2012, pp. 3354– 3361

  13. [16]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11 621–11 631

  14. [17]

    Scalability in perception for autonomous driving: Waymo open dataset,

    P. Sun, H. Kretzschmar, A. Dotiwalla, C. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caineet al., “Scalability in perception for autonomous driving: Waymo open dataset,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020...

  15. [18]

    Cityscapes 3d: Dataset and benchmark for 9 dof vehicle detection,

    N. G ¨ahlert, N. Jourdan, M. Cordts, U. Franke, and J. Denzler, “Cityscapes 3d: Dataset and benchmark for 9 dof vehicle detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2020

  16. [19]

    Argoverse: 3d tracking and forecasting with rich maps,

    M.-F. Chang, J. Lambert, P. Sangkloy, J. Singh, S. Bak, A. Hartnett, D. Wang, P. Carr, S. Lucey, and D. Ramanan, “Argoverse: 3d tracking and forecasting with rich maps,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 8748–8757

  17. [20]

    A*3d: An autonomous driving dataset in challenging environments,

    Q.-H. Pham, P. Sevestre, R. S. Pahwa, H. Zhan, C. H. Pang, Y . Chen, A. Mustafa, V . Chandrasekhar, and J. Lin, “A*3d: An autonomous driving dataset in challenging environments,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 202...

  18. [21]

    The apolloscape dataset for autonomous driving,

    X. Huang, X. Cheng, Q. Geng, B. Cao, D. Zhou, P. Wang, Y . Lin, R. Yang, Z. Carmichael, C. Langet al., “The apolloscape dataset for autonomous driving,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 954– 960

  19. [22]

    All- in-one drive: A large-scale comprehensive perception dataset with high- density long-range point clouds,

    X. Weng, Y . Man, D. Cheng, J. Park, M. O’Toole, and K. Kitani, “All- in-one drive: A large-scale comprehensive perception dataset with high- density long-range point clouds,”arXiv preprint arXiv:2010.03180, 2020

  20. [23]

    The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes,

    S. Patil, F. P”atzold, C. H”ane, M. Tschentscher, and A. Knoll, “The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes,” in2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 9552–9558

  21. [24]

    The cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, H. Dahlkamp, A. Schuster, U. Franke, and S. Roth, “The cityscapes dataset for semantic urban scene understanding,” inProceedings of the IEEE/CVF Conference on Computer Vision 12 and Pattern Recognition (CVPR...

  22. [25]

    Pointfusion: Deep sensor fusion for 3d bounding box estimation,

    D. Xu, D. Anguelov, and A. Jain, “Pointfusion: Deep sensor fusion for 3d bounding box estimation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 244–253

  23. [26]

    Pointpainting: Sequential fusion for 3d object detection,

    S. V ora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4604–4612

  24. [27]

    Multimodal virtual point 3d de- tection,

    T. Yin, X. Zhou, and P. Kr ¨ahenb¨uhl, “Multimodal virtual point 3d de- tection,”Advances in Neural Information Processing Systems, vol. 34, pp. 16 494–16 507, 2021

  25. [28]

    Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection,

    Y . Li, A. W. Yu, T. Meng, B. Caine, J. Ngiam, D. Peng, J. Shen, Y . Lu, D. Zhou, Q. V . Leet al., “Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 17 182–17 191

  26. [29]

    Centerfusion: Center-based radar and camera fusion for 3d object detection,

    R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera fusion for 3d object detection,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021

  27. [30]

    Pointaugmenting: Cross- modal augmentation for 3d object detection,

    C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross- modal augmentation for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 11 794–11 803

  28. [31]

    Unifying voxel-based representation with transformer for 3d object detection,

    Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,”Advances in Neural Information Processing Systems, vol. 35, pp. 18 442–18 455, 2022

  29. [32]

    Sparse fuse dense: Towards high quality 3d detection with depth completion,

    X. Wu, L. Peng, H. Yang, L. Xie, C. Huang, C. Deng, H. Liu, and D. Cai, “Sparse fuse dense: Towards high quality 3d detection with depth completion,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5418–5427

  30. [33]

    Joint 3d proposal generation and object detection from view aggregation,

    J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” in2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2018, pp. 1–8

  31. [34]

    Roarnet: A robust 3d object detection based on region approximation refinement,

    K. Shin, Y . P. Kwon, and M. Tomizuka, “Roarnet: A robust 3d object detection based on region approximation refinement,” in2019 IEEE intelligent vehicles symposium (IV). IEEE, 2019, pp. 2510–2515

  32. [35]

    Weakly aligned cross-modal learning for multispectral pedestrian detection,

    L. Zhang, X. Zhu, X. Chen, X. Yang, Z. Lei, and Z. Liu, “Weakly aligned cross-modal learning for multispectral pedestrian detection,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5127–5137

  33. [36]

    E2e-mfd: Towards end-to-end synchronous multimodal fusion detection,

    J. Zhang, M. Cao, X. Yang, W. Xie, J. Lei, D. Li, W. Huang, and Y . Li, “E2e-mfd: Towards end-to-end synchronous multimodal fusion detection,”arXiv preprint arXiv:2403.09323, 2024

  34. [37]

    Mvx-net: Multimodal voxelnet for 3d object detection,

    V . A. Sindagi, Y . Zhou, and O. Tuzel, “Mvx-net: Multimodal voxelnet for 3d object detection,” in2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 7276–7282

  35. [38]

    Bridging the view disparity between radar and camera features for multi-modal fusion 3d object detection,

    T. Zhou, J. Chen, Y . Shi, K. Jiang, M. Yang, and D. Yang, “Bridging the view disparity between radar and camera features for multi-modal fusion 3d object detection,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 2, pp. 1523–1535, 2023

  36. [39]

    Improving multispectral pedestrian detection by addressing modality imbalance problems,

    K. Zhou, L. Chen, and X. Cao, “Improving multispectral pedestrian detection by addressing modality imbalance problems,” inComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVIII 16. Springer, 2020, pp. 787– 803

  37. [40]

    Multimodal object detection by channel switching and spatial attention,

    Y . Cao, J. Bin, J. Hamari, E. Blasch, and Z. Liu, “Multimodal object detection by channel switching and spatial attention,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2023, pp. 403–411

  38. [41]

    Pointnet: Deep learning on point sets for 3d classification and segmentation,

    C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 652–660

  39. [42]

    Pointnet++: Deep hierarchical feature learning on point sets in a metric space,

    C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,”Advances in neural information processing systems, vol. 30, 2017

  40. [43]

    Frustum pointnets for 3d object detection from rgb-d data,

    C. R. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum pointnets for 3d object detection from rgb-d data,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 918– 927

  41. [44]

    Pi-rcnn: An efficient multi-sensor 3d object detector with point-based attentive cont-conv fusion module,

    L. Xie, C. Xiang, Z. Yu, G. Xu, Z. Yang, D. Cai, and X. He, “Pi-rcnn: An efficient multi-sensor 3d object detector with point-based attentive cont-conv fusion module,” inProceedings of the AAAI conference on artificial intelligence, vol. 34, no. 07, 2020, pp. 12 460–12 467

  42. [45]

    Fusionpainting: Multimodal fusion with adaptive attention for 3d object detection,

    S. Xu, D. Zhou, J. Fang, J. Yin, Z. Bin, and L. Zhang, “Fusionpainting: Multimodal fusion with adaptive attention for 3d object detection,” in 2021 IEEE International Intelligent Transportation Systems Conference (ITSC). IEEE, 2021, pp. 3047–3054

  43. [46]

    Graphalign: Enhanc- ing accurate feature alignment by graph matching for multi-modal 3d object detection,

    Z. Song, H. Wei, L. Bai, L. Yang, and C. Jia, “Graphalign: Enhanc- ing accurate feature alignment by graph matching for multi-modal 3d object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3358–3369

  44. [47]

    Vpfnet: Improving 3d object detection with virtual point based lidar and stereo data fusion,

    H. Zhu, J. Deng, Y . Zhang, J. Ji, Q. Mao, H. Li, and Y . Zhang, “Vpfnet: Improving 3d object detection with virtual point based lidar and stereo data fusion,”IEEE Transactions on Multimedia, vol. 25, pp. 5291– 5304, 2022

  45. [48]

    V oxel field fusion for 3d object detection,

    Y . Li, X. Qi, Y . Chen, L. Wang, Z. Li, J. Sun, and J. Jia, “V oxel field fusion for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1120–1129

  46. [49]

    Autoalign: Pixel-instance feature aggregation for multi- modal 3d object detection,

    Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, F. Zhao, B. Zhou, and H. Zhao, “Autoalign: Pixel-instance feature aggregation for multi- modal 3d object detection,”arXiv preprint arXiv:2201.06493, 2022

  47. [50]

    Deformable feature aggregation for dynamic multi-modal 3d object detection,

    Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao, “Deformable feature aggregation for dynamic multi-modal 3d object detection,” in European conference on computer vision. Springer, 2022, pp. 628– 644

  48. [51]

    V ox- elnextfusion: A simple, unified and effective voxel fusion framework for multi-modal 3d object detection,

    Z. Song, G. Zhang, J. Xie, L. Liu, C. Jia, S. Xu, and Z. Wang, “V ox- elnextfusion: A simple, unified and effective voxel fusion framework for multi-modal 3d object detection,”arXiv preprint arXiv:2401.02702, 2024

  49. [52]

    Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,

    X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1090–1099

  50. [53]

    Learning cross-modal deep representations for robust pedestrian detection,

    D. Xu, W. Ouyang, E. Ricci, X. Wang, and N. Sebe, “Learning cross-modal deep representations for robust pedestrian detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5363–5371

  51. [54]

    Guided attentive feature fusion for multispectral pedestrian detection,

    H. Zhang, E. Fromont, S. Lef `evre, and B. Avignon, “Guided attentive feature fusion for multispectral pedestrian detection,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 72–80

  52. [55]

    Removal and selection: Improving rgb-infrared object detection via coarse-to-fine fusion,

    T. Zhao, M. Yuan, F. Jiang, N. Wang, and X. Wei, “Removal and selection: Improving rgb-infrared object detection via coarse-to-fine fusion,”arXiv preprint arXiv:2401.10731, 2024

  53. [56]

    Deep continuous fusion for multi-sensor 3d object detection,

    M. Liang, B. Yang, S. Wang, and R. Urtasun, “Deep continuous fusion for multi-sensor 3d object detection,” inProceedings of the European conference on computer vision (ECCV), 2018, pp. 641–656

  54. [57]

    Cross-modality 3d object detection,

    M. Zhu, C. Ma, P. Ji, and X. Yang, “Cross-modality 3d object detection,” inProceedings of the IEEE/CVF Winter conference on Applications of Computer Vision, 2021, pp. 3772–3781

  55. [58]

    Multi-task multi-sensor fusion for 3d object detection,

    M. Liang, B. Yang, Y . Chen, R. Hu, and R. Urtasun, “Multi-task multi-sensor fusion for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 7345–7353

  56. [59]

    Epnet: Enhancing point features with image semantics for 3d object detection,

    T. Huang, Z. Liu, X. Chen, and X. Bai, “Epnet: Enhancing point features with image semantics for 3d object detection,” inComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16. Springer, 2020, pp. 35–52

  57. [60]

    Epnet++: Cas- cade bi-directional fusion for multi-modal 3d object detection,

    Z. Liu, T. Huang, B. Li, X. Chen, X. Wang, and X. Bai, “Epnet++: Cas- cade bi-directional fusion for multi-modal 3d object detection,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 7, pp. 8324–8341, 2022

  58. [61]

    Dense voxel fusion for 3d object detection,

    A. Mahmoud, J. S. Hu, and S. L. Waslander, “Dense voxel fusion for 3d object detection,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 663–672

  59. [62]

    Logonet: Towards accurate 3d object detection with local-to-global cross-modal fusion,

    X. Li, T. Ma, Y . Hou, B. Shi, Y . Yang, Y . Liu, X. Wu, Q. Chen, Y . Li, Y . Qiaoet al., “Logonet: Towards accurate 3d object detection with local-to-global cross-modal fusion,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 ...

  60. [63]

    Cat-det: Contrastively augmented transformer for multi-modal 3d object detection,

    Y . Zhang, J. Chen, and D. Huang, “Cat-det: Contrastively augmented transformer for multi-modal 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 908–917

  61. [64]

    Sea- date: Remedy dual-attention transformer with semantic alignment via contrast learning for multimodal object detection,

    S. Dong, W. Xie, D. Yang, J. Tian, Y . Li, J. Zhang, and J. Lei, “Sea- date: Remedy dual-attention transformer with semantic alignment via contrast learning for multimodal object detection,”IEEE Transactions on Circuits and Systems for Video Technology, 2024. 13

  62. [65]

    Fusion-mamba for cross-modality object detection,

    W. Dong, H. Zhu, S. Lin, X. Luo, Y . Shen, X. Liu, J. Zhang, G. Guo, and B. Zhang, “Fusion-mamba for cross-modality object detection,” arXiv preprint arXiv:2404.09146, 2024

  63. [66]

    Cobevt: Cooper- ative bird’s eye view semantic segmentation with sparse transformers,

    R. Xu, Z. Tu, H. Xiang, W. Shao, B. Zhou, and J. Ma, “Cobevt: Cooper- ative bird’s eye view semantic segmentation with sparse transformers,” arXiv preprint arXiv:2207.02202, 2022

  64. [67]

    Collaboration helps camera overtake lidar in 3d detection,

    Y . Hu, Y . Lu, R. Xu, W. Xie, S. Chen, and Y . Wang, “Collaboration helps camera overtake lidar in 3d detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9243–9252

  65. [68]

    V2vnet: Vehicle-to-vehicle communication for joint perception and prediction,

    T.-H. Wang, S. Manivasagam, M. Liang, B. Yang, W. Zeng, and R. Ur- tasun, “V2vnet: Vehicle-to-vehicle communication for joint perception and prediction,” inComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II

  66. [69]

    Springer, 2020, pp. 605–621

  67. [70]

    Macp: Efficient model adaptation for cooperative perception,

    Y . Ma, J. Lu, C. Cui, S. Zhao, X. Cao, W. Ye, and Z. Wang, “Macp: Efficient model adaptation for cooperative perception,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 3373–3382

  68. [71]

    Hm-vit: Hetero-modal vehicle-to-vehicle cooperative perception with vision transformer,

    H. Xiang, R. Xu, and J. Ma, “Hm-vit: Hetero-modal vehicle-to-vehicle cooperative perception with vision transformer,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 284–295

  69. [72]

    Multi-agent collaborative perception via motion-aware robust communication network,

    S. Hong, Y . Liu, Z. Li, S. Li, and Y . He, “Multi-agent collaborative perception via motion-aware robust communication network,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 301–15 310

  70. [73]

    When2com: Multi-agent perception via communication graph grouping,

    Y .-C. Liu, J. Tian, N. Glaser, and Z. Kira, “When2com: Multi-agent perception via communication graph grouping,” inProceedings of the IEEE/CVF Conference on computer vision and pattern recognition, 2020, pp. 4106–4115

  71. [74]

    Who2com: Collaborative perception via learnable handshake com- munication,

    Y .-C. Liu, J. Tian, C.-Y . Ma, N. Glaser, C.-W. Kuo, and Z. Kira, “Who2com: Collaborative perception via learnable handshake com- munication,” in2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 6876–6883

  72. [75]

    How2comm: Communication-efficient and collaboration- pragmatic multi-agent perception,

    D. Yang, K. Yang, Y . Wang, J. Liu, Z. Xu, R. Yin, P. Zhai, and L. Zhang, “How2comm: Communication-efficient and collaboration- pragmatic multi-agent perception,”Advances in Neural Information Processing Systems, vol. 36, 2024

  73. [76]

    Communication- efficient collaborative perception via information filling with code- book,

    Y . Hu, J. Peng, S. Liu, J. Ge, S. Liu, and S. Chen, “Communication- efficient collaborative perception via information filling with code- book,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 481–15 490

  74. [77]

    Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y . Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” inProceedings of the European Conference on Computer Vision (ECCV), 2022

  75. [78]

    Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,

    C. Yang, Y . Chen, H. Tian, C. Tao, X. Zhu, Z. Zhang, G. Huang, H. Li, Y . Qiao, L. Luet al., “Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  76. [79]

    Exploring object- centric temporal modeling for efficient multi-view 3d object detection,

    S. Wang, Y . Liu, T. Wang, Y . Li, and X. Zhang, “Exploring object- centric temporal modeling for efficient multi-view 3d object detection,” in2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 3598–3608

  77. [80]

    Sparse4d: Multi-view 3d object detection with sparse spatial-temporal fusion,

    X. Lin, T. Lin, Z. Pei, L. Huang, and Z. Su, “Sparse4d: Multi-view 3d object detection with sparse spatial-temporal fusion,” 2022

  78. [81]

    Sparse4d v2: Recurrent temporal fusion with sparse model,

    X. Lin, T. Lin, Z. Pei, L. Huang, and Zhizhong Su, “Sparse4d v2: Recurrent temporal fusion with sparse model,”arXiv preprint arXiv:2305.14018, 2023

  79. [82]

    Sparse4d v3: Advancing end-to-end 3d detection and tracking,

    X. Lin, Z. Pei, T. Lin, L. Huang, and Z. Su, “Sparse4d v3: Advancing end-to-end 3d detection and tracking,” 2023

  80. [83]

    Sparsefusion3d: Sparse sensor fusion for 3d object detection by radar and camera in environmental perception,

    Z. Yu, W. Wan, M. Ren, X. Zheng, and Z. Fang, “Sparsefusion3d: Sparse sensor fusion for 3d object detection by radar and camera in environmental perception,”IEEE Transactions on Intelligent Vehicles, vol. 9, no. 1, pp. 1524–1536, 2024

  81. [84]

    Planning-oriented autonomous driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, L. Lu, X. Jia, Q. Liu, J. Dai, Y . Qiao, and H. Li, “Planning-oriented autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

  82. [85]

    Fusionad: Multi-modality fusion for prediction and planning tasks of autonomous driving,

    T. Ye, W. Jing, C. Hu, S. Huang, L. Gao, F. Li, J. Wang, K. Guo, W. Xiao, W. Mao, H. Zheng, K. Li, J. Chen, and K. Yu, “Fusionad: Multi-modality fusion for prediction and planning tasks of autonomous driving,”arXiv preprint arXiv:2308.01006, 2023

  83. [86]

    Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,

    Z. Lin, Z. Liu, Z. Xia, X. Wang, Y . Wang, S. Qi, Y . Dong, N. Dong, L. Zhang, and C. Zhu, “Rcbevdet: Radar-camera fusion in bird’s eye view for 3d object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 928–14 937

  84. [87]

    Vision-centric bev perception: A survey,

    Y . Ma, T. Wang, X. Bai, H. Yang, Y . Hou, Y . Wang, Y . Qiao, R. Yang, and X. Zhu, “Vision-centric bev perception: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10 978–10 997, 2024

  85. [88]

    End-to-end object detection with transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision. Springer, 2020, pp. 213– 229

  86. [89]

    Deformable DETR: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/forum?id=gZ9hCDWe6ke

  87. [90]

    Detr3d: 3d object detection from multi-view images via 3d- to-2d queries,

    Y . Wang, V . Guizilini, T. Zhang, Y . Wang, H. Zhao, , and J. M. Solomon, “Detr3d: 3d object detection from multi-view images via 3d- to-2d queries,” inThe Conference on Robot Learning (CoRL), 2021

  88. [91]

    Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly estimating depth,

    J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly estimating depth,” inProceedings of the European Conference on Computer Vision (ECCV), 2020

  89. [92]

    Bevdet4d: Exploit temporal cues in multi- camera 3d object detection,

    J. Huang and G. Huang, “Bevdet4d: Exploit temporal cues in multi- camera 3d object detection,”arXiv preprint arXiv:2203.17054, 2022

  90. [93]

    Bevdet: High- performance multi-camera 3d object detection in bird-eye-view,

    J. Huang, G. Huang, Z. Zhu, Y . Yun, D. Du, and X. Bai, “Bevdet: High- performance multi-camera 3d object detection in bird-eye-view,” in Proceedings of the European Conference on Computer Vision (ECCV), 2022

  91. [94]

    Beverse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving,

    Y . Zhang, Z. Zhu, W. Zheng, J. Huang, G. Huang, J. Zhou, and J. Lu, “Beverse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving,”arXiv preprint arXiv:2205.09743, 2022

  92. [95]

    Unifusion: Unified multi-view fusion transformer for spatial-temporal representation in bird’s-eye-view,

    Z. Qin, J. Chen, C. Chen, X. Chen, and X. Li, “Unifusion: Unified multi-view fusion transformer for spatial-temporal representation in bird’s-eye-view,” inProceedings of the IEEE/CVF International Con- ference on Computer Vision, 2023, pp. 8656–8665

  93. [96]

    Tbp-former: Learning temporal bird’s-eye-view pyramid for joint perception and prediction in vision-centric autonomous driving,

    S. Fang, Z. Wang, Y . Zhong, J. Ge, and S. Chen, “Tbp-former: Learning temporal bird’s-eye-view pyramid for joint perception and prediction in vision-centric autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, ...

  94. [100]

    Fusionformer: A concise unified feature fusion transformer for 3D pose estimation,

    Y . Cai, W. Zhang, Y . Wu, and C. Jin, “Fusionformer: A concise unified feature fusion transformer for 3D pose estimation,” inProceedings of the AAAI Conference on Artificial Intelligence, 2024

  95. [103]

    Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,

    J. Kim, M. Seong, and J. W. Choi, “Crt-fusion: Camera, radar, temporal fusion using motion information for 3d object detection,”Advances in Neural Information Processing Systems, vol. 37, pp. 108 625–108 648, 2024

  96. [104]

    Planning-oriented autonomous driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wanget al., “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 17 853–17 862

  97. [105]

    Crn: Camera radar net for accurate, robust, efficient 3d perception,

    Y . Kim, J. Shin, S. Kim, I.-J. Lee, J. W. Choi, and D. Kum, “Crn: Camera radar net for accurate, robust, efficient 3d perception,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 14

  98. [106]

    Sce2drivex: A generalized mllm framework for scene-to-drive learn- ing,

    R. Zhao, Q. Yuan, J. Li, H. Hu, Y . Li, C. Zheng, and F. Gao, “Sce2drivex: A generalized mllm framework for scene-to-drive learn- ing,”arXiv preprint arXiv:2502.14917, 2025

  99. [107]

    X-driver: Explainable autonomous driving with vision-language models,

    W. Liu, J. Zhang, B. Zheng, Y . Hu, Y . Lin, and Z. Zeng, “X-driver: Explainable autonomous driving with vision-language models,”arXiv preprint arXiv:2505.05098, 2025

  100. [108]

    Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,

    Z. Zhang, X. Li, Z. Xu, W. Peng, Z. Zhou, M. Shi, and S. Huang, “Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,”arXiv preprint arXiv:2504.00379, 2025

  101. [109]

    Safeauto: Knowledge-enhanced safe autonomous driving with multimodal foun- dation models,

    J. Zhang, X. Yang, T. Wang, Y . Yao, A. Petiushko, and B. Li, “Safeauto: Knowledge-enhanced safe autonomous driving with multimodal foun- dation models,”arXiv preprint arXiv:2503.00211, 2025

  102. [110]

    Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

    W. Wang, J. Xie, C. Hu, H. Zou, J. Fan, W. Tong, Y . Wen, S. Wu, H. Deng, Z. Liet al., “Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,”arXiv preprint arXiv:2312.09245, 2023

  103. [111]

    Maplm: A real-world large-scale vision- language benchmark for map and traffic scene understanding,

    X. Cao, T. Zhou, Y . Ma, W. Ye, C. Cui, K. Tang, Z. Cao, K. Liang, Z. Wang, J. M. Rehget al., “Maplm: A real-world large-scale vision- language benchmark for map and traffic scene understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  104. [112]

    Lidar-llm: Exploring the potential of large language models for 3d lidar understanding,

    S. Yang, J. Liu, R. Zhang, M. Pan, Z. Guo, X. Li, Z. Chen, P. Gao, H. Li, Y . Guoet al., “Lidar-llm: Exploring the potential of large language models for 3d lidar understanding,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 9, 2025, pp. 9247–9255

  105. [113]

    Drivelm: Driving with graph visual question answering,

    C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li, “Drivelm: Driving with graph visual question answering,”arXiv preprint arXiv:2312.14150, 2023

  106. [114]

    Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,

    S. Wang, Z. Yu, X. Jiang, S. Lan, M. Shi, N. Chang, J. Kautz, Y . Li, and J. M. Alvarez, “Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,”arXiv preprint arXiv:2405.01533, 2024

  107. [115]

    Drivevlm: The convergence of au- tonomous driving and large vision-language models,

    X. Tian, J. Gu, B. Li, Y . Liu, Y . Wang, Z. Zhao, K. Zhan, P. Jia, X. Lang, and H. Zhao, “Drivevlm: The convergence of au- tonomous driving and large vision-language models,”arXiv preprint arXiv:2402.12289, 2024

  108. [116]

    Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,

    M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, and L. Zhang, “Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,” inEuropean Conference on Computer Vision. Springer, 2025, pp. 292–308

  109. [117]

    Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,

    X. Ding, J. Han, H. Xu, X. Liang, W. Zhang, and X. Li, “Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 668–13 677

  110. [118]

    Embodied understanding of driving scenarios,

    Y . Zhou, L. Huang, Q. Bu, J. Zeng, T. Li, H. Qiu, H. Zhu, M. Guo, Y . Qiao, and H. Li, “Embodied understanding of driving scenarios,” inEuropean Conference on Computer Vision. Springer, 2025, pp. 129–148

  111. [119]

    Driving with llms: Fusing object- level vector modality for explainable autonomous driving,

    L. Chen, O. Sinavski, J. H ¨unermann, A. Karnsund, A. J. Willmott, D. Birch, D. Maund, and J. Shotton, “Driving with llms: Fusing object- level vector modality for explainable autonomous driving,” in2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 20...

  112. [120]

    Sensor data quality: A systematic review,

    H. Y . Teh, A. W. Kempa-Liehr, and K. I.-K. Wang, “Sensor data quality: A systematic review,”Journal of Big Data, vol. 7, no. 1, p. 11, 2020

  113. [121]

    Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,

    Z. Yang, H. Yang, Z. Pan, and L. Zhang, “Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,” inThe Twelfth International Conference on Learning Representations

  114. [122]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10 850–10 869, 2023

  115. [123]

    A survey of data quality measurement and monitoring tools,

    L. Ehrlinger and W. W ¨oß, “A survey of data quality measurement and monitoring tools,”Frontiers in big data, vol. 5, p. 850611, 2022

  116. [124]

    Predicting take-over time for autonomous driving with real-world data: Robust data augmentation, models, and evaluation,

    A. Rangesh, N. Deo, R. Greer, P. Gunaratne, and M. M. Trivedi, “Predicting take-over time for autonomous driving with real-world data: Robust data augmentation, models, and evaluation,”arXiv preprint arXiv:2107.12932, 2021

  117. [125]

    Polarmix: A general data augmentation technique for lidar point clouds,

    A. Xiao, J. Huang, D. Guan, K. Cui, S. Lu, and L. Shao, “Polarmix: A general data augmentation technique for lidar point clouds,”Advances in Neural Information Processing Systems, vol. 35, pp. 11 035–11 048, 2022

  118. [126]

    Camera-lidar extrinsic calibration via traffic signs,

    X. Yuan, Y . Xie, S. Wang, and T. Xiong, “Camera-lidar extrinsic calibration via traffic signs,” in2023 42nd Chinese Control Conference (CCC). IEEE, 2023, pp. 4679–4684

  119. [127]

    Sensor fusion for joint 3d object detection and semantic segmentation,

    G. P. Meyer, J. Charland, D. Hegde, A. Laddha, and C. Vallespi- Gonzalez, “Sensor fusion for joint 3d object detection and semantic segmentation,” inProceedings of the IEEE/CVF conference on com- puter vision and pattern recognition workshops, 2019, pp. 0–0

  120. [128]

    Perception and sensing for autonomous vehicles under adverse weather conditions: A survey,

    Y . Zhang, A. Carballo, H. Yang, and K. Takeda, “Perception and sensing for autonomous vehicles under adverse weather conditions: A survey,”ISPRS Journal of Photogrammetry and Remote Sensing, vol. 196, pp. 146–177, 2023

  121. [129]

    A survey of lidar and camera fusion enhancement,

    H. Zhong, H. Wang, Z. Wu, C. Zhang, Y . Zheng, and T. Tang, “A survey of lidar and camera fusion enhancement,”Procedia Computer Science, vol. 183, pp. 579–588, 2021

  122. [130]

    Bev perception for autonomous driving: State of the art and future perspectives,

    J. Zhao, J. Shi, and L. Zhuo, “Bev perception for autonomous driving: State of the art and future perspectives,”Expert Systems with Applica- tions, vol. 258, p. 125103, 2024

  123. [131]

    Bev-v2x: Cooperative birds-eye-view fusion and grid occupancy pre- diction via v2x-based data sharing,

    C. Chang, J. Zhang, K. Zhang, W. Zhong, X. Peng, S. Li, and L. Li, “Bev-v2x: Cooperative birds-eye-view fusion and grid occupancy pre- diction via v2x-based data sharing,”IEEE Transactions on Intelligent Vehicles, 2023

  124. [132]

    V oxel-based representation of 3d point clouds: Methods, applications, and its potential use in the construction industry,

    Y . Xu, X. Tong, and U. Stilla, “V oxel-based representation of 3d point clouds: Methods, applications, and its potential use in the construction industry,”Automation in Construction, vol. 126, p. 103675, 2021

  125. [133]

    Tempo- rally consistent enhancement of low-light videos via spatial-temporal compatible learning,

    L. Zhu, W. Yang, B. Chen, H. Zhu, X. Meng, and S. Wang, “Tempo- rally consistent enhancement of low-light videos via spatial-temporal compatible learning,”International Journal of Computer Vision, pp. 1–21, 2024

  126. [134]

    An adaptive approach to time synchro- nization for wireless sensors under extreme conditions,

    A. L ¨ubken and A. F ¨orster, “An adaptive approach to time synchro- nization for wireless sensors under extreme conditions,”IEEE Sensors Journal, 2024

  127. [135]

    Attention is all you need,

    A. Vaswani, “Attention is all you need,”Advances in Neural Informa- tion Processing Systems, 2017

  128. [136]

    Self- supervised representation learning: Introduction, advances, and chal- lenges,

    L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales, “Self- supervised representation learning: Introduction, advances, and chal- lenges,”IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 42–62, 2022

  129. [137]

    Mixing up contrastive learning: Self-supervised representation learn- ing for time series,

    K. Wickstrøm, M. Kampffmeyer, K. Ø. Mikalsen, and R. Jenssen, “Mixing up contrastive learning: Self-supervised representation learn- ing for time series,”Pattern Recognition Letters, vol. 155, pp. 54–61, 2022

  130. [138]

    mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

    Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang, “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 040–13 051

  131. [139]

    Tokenize the world into object-level knowledge to address long-tail events in autonomous driving,

    T. Tian, B. Li, X. Weng, Y . Chen, E. Schmerling, Y . Wang, B. Ivanovic, and M. Pavone, “Tokenize the world into object-level knowledge to address long-tail events in autonomous driving,” in8th Annual Conference on Robot Learning

  132. [140]

    Radarscenes: A real-world radar point cloud data set for automotive applications,

    O. Schumann, M. Hahn, N. Scheiner, F. Weishaupt, J. F. Tilly, J. Dick- mann, and C. W ¨ohler, “Radarscenes: A real-world radar point cloud data set for automotive applications,” in2021 IEEE 24th International Conference on Information Fusion (FUSION). IEEE, 2021, pp. 1–8

  133. [141]

    Deep learning for lidar point clouds in autonomous driving: A review,

    Y . Li, L. Ma, Z. Zhong, F. Liu, M. A. Chapman, D. Cao, and J. Li, “Deep learning for lidar point clouds in autonomous driving: A review,” IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 8, pp. 3412–3432, 2020

  134. [142]

    Radar-pointgnn: Graph based object recognition for unstructured radar point-cloud data,

    P. Svenningsson, F. Fioranelli, and A. Yarovoy, “Radar-pointgnn: Graph based object recognition for unstructured radar point-cloud data,” in 2021 IEEE Radar Conference (RadarConf21). IEEE, 2021, pp. 1–6

  135. [143]

    Estimating tree species composition from airborne laser scanning data using point-based deep learning models,

    B. A. Murray, N. C. Coops, L. Winiwarter, J. C. White, A. Dick, I. Barbeito, and A. Ragab, “Estimating tree species composition from airborne laser scanning data using point-based deep learning models,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 207, pp. 282–297, 2024

  136. [144]

    Unirag: Universal retrieval augmentation for multi-modal large language mod- els,

    S. Sharifymoghaddam, S. Upadhyay, W. Chen, and J. Lin, “Unirag: Universal retrieval augmentation for multi-modal large language mod- els,”arXiv preprint arXiv:2405.10311, 2024

  137. [145]

    A comprehensive study on self-learning methods and implications to autonomous driving,

    J. Xing, D. Wei, S. Zhou, T. Wang, Y . Huang, and H. Chen, “A comprehensive study on self-learning methods and implications to autonomous driving,”IEEE Transactions on Neural Networks and Learning Systems, 2024

  138. [146]

    An adaptive tinyml unsupervised online learning algorithm for driver behavior analysis,

    M. Silva, T. Medeiros, M. Azevedo, M. Medeiros, M. Themoteo, T. Gois, I. Silva, and D. G. Costa, “An adaptive tinyml unsupervised online learning algorithm for driver behavior analysis,” in2023 IEEE International Workshop on Metrology for Automotive (MetroAutomo- tive). IEEE, ...

  139. [147]

    Online learning: A comprehensive survey,

    S. C. Hoi, D. Sahoo, J. Lu, and P. Zhao, “Online learning: A comprehensive survey,”Neurocomputing, vol. 459, pp. 249–289, 2021. 15

  140. [148]

    A survey of zero-shot learning: Settings, methods, and applications,

    W. Wang, V . W. Zheng, H. Yu, and C. Miao, “A survey of zero-shot learning: Settings, methods, and applications,”ACM Transactions on Intelligent Systems and Technology (TIST), vol. 10, no. 2, pp. 1–37, 2019

  141. [149]

    Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,

    Y . Xian, C. H. Lampert, B. Schiele, and Z. Akata, “Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,”IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 9, pp. 2251–2265, 2018

  142. [150]

    Building trustworthy autonomous vehicles: The role of multi-sensor fusion and explainable ai (xai) in on- road and off-road scenarios,

    K. P. De Jong Yeong and J. Walsh, “Building trustworthy autonomous vehicles: The role of multi-sensor fusion and explainable ai (xai) in on- road and off-road scenarios,”Sensors and Electronic Instrumentation Advances, p. 145, 2024

  143. [151]

    Context-aware feature selection and classifi- cation

    J. Wang and M. Bilgic, “Context-aware feature selection and classifi- cation.” inIJCAI, 2023, pp. 4317–4325

  144. [152]

    A visual analytics system for improving attention-based traffic forecasting models,

    S. Jin, H. Lee, C. Park, H. Chu, Y . Tae, J. Choo, and S. Ko, “A visual analytics system for improving attention-based traffic forecasting models,”IEEE transactions on visualization and computer graphics, vol. 29, no. 1, pp. 1102–1112, 2022

  145. [153]

    At- tentionviz: A global view of transformer attention,

    C. Yeh, Y . Chen, A. Wu, C. Chen, F. Vi ´egas, and M. Wattenberg, “At- tentionviz: A global view of transformer attention,”IEEE Transactions on Visualization and Computer Graphics, 2023

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.