{"id":"bcfc46d6-8607-43e7-bf5f-57d13b4664cb","arxiv_id":"2506.19769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A task-agnostic survey of multi-sensor fusion perception methods for embodied AI, covering multi-modal, multi-agent, time-series, and multimodal large language model fusion.","lead":"This paper surveys multi-sensor fusion perception methods for embodied AI, organizing them into multi-modal, multi-agent, time-series, and multimodal large language model fusion. It offers a broad entry point for readers who want a task-agnostic overview of how cameras, LiDAR, and radar data are combined in robots and autonomous vehicles.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fig. 8's timeline lists numerous time-series fusion methods that are never cited or discussed, so the survey's core claim of a comprehensive, task-agnostic map after a 'rigorous and detailed investigation' is not yet supported.","rationale":"The reader's weakest assumption was that the selected methods are representative and that the four-category taxonomy accurately captures the diversity of MSFP, noting the absence of a search strategy. I partially agree, but I found a more concrete and more severe instantiation: Figure 8 names numerous time-series methods that are never cited or described. This is not just a missing methodology statement; it is an internal inconsistency between the survey's corpus and its own exhibits. If the timeline includes uncited methods, the comprehensiveness claim cannot be validated by a reader, and the task-agnostic map becomes unreliable for navigation. The issue is fixable by either adding the missing citations and discussions or removing the uncited entries and narrowing the claims, so I would not move to REJECT. The survey does provide useful background sections, clear tables for multi-modal fusion, and coverage of MM-LLM fusion methods that earlier surveys often omit; those strengths support keeping the manuscript available in conditional form. The authors' appended caveat about limited expertise further signals that this curation gap should be treated seriously rather than waved off. Because my concern reinforces the reader's CONDITIONAL verdict rather than changing it, I recommend leaving the verdict unchanged.","tokens_in":24866,"tokens_out":7009,"duration_ms":79344,"concrete_test":"Extract every method name from Fig. 8 and cross-check each against the reference list and the body text. Count how many names are cited and discussed, cited only, or neither cited nor discussed. Then independently verify whether the uncited names are legitimate time-series fusion methods from known venues via a literature search. If more than a small fraction (e.g., >20%) of Fig. 8's entries are absent from both text and references, the claim of a comprehensive, rigorously curated taxonomy is not supported. Additionally, check the Section V category definitions against each cited method to confirm that every method placed in 'dense query', 'sparse query', or 'hybrid query' actually uses the corresponding mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's main value proposition is that it gives researchers a reliable, task-agnostic map of MSFP methods. That claim rests on the corpus being representative and the categories being accurate. The internal evidence in Section V undermines this. Figure 8, the timeline of time-series fusion methods, names at least a dozen methods—PolarDETR, STS, MV-FCOS3D++, DORT, E-TMA, BridgeAD, SAD, RENet, Far3D, TLCFuse, PETR, and PETR v2—that do not appear in the reference list and are not discussed in the text. The accompanying taxonomy (Table IV) lists only three categories, yet the timeline also contains an 'Others' group, with no criterion stated for what belongs in the taxonomy versus 'Others'. A reader cannot determine whether the selection is representative or exhaustive, and the uncited entries suggest the method collection was assembled without a documented selection process. Separately, Section IV is titled 'Multi-Agent Fusion' but describes 'the multi-view fusion of agent-to-agent', and Fig. 7, placed in the time-series section, is captioned as an A2A fusion pipeline. These are not isolated typos; they indicate the organizing principle is applied unevenly. Since the survey is a reference work, the map itself must be trustworthy; the current version does not yet establish that.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a task-agnostic organization of multi-sensor fusion perception (MSFP) methods for embodied AI, structured around four technical views: multi-modal fusion (point-, voxel-, region-, and multi-level), multi-agent fusion, time-series fusion, and multimodal-LLM fusion. It provides background on sensors, datasets, and perception tasks, reviews representative methods in each category, and closes with challenges and future directions at the data, model, and application levels. The central claim is that existing surveys are too narrowly tied to single tasks (e.g., 3D object detection) or single perspectives (mainly multi-modal fusion), and that a task-agnostic, multi-perspective organization is needed for researchers across fields.","tokens_in":25072,"tokens_out":5368,"duration_ms":51621,"significance":"If the survey's organization were fully substantiated, it would fill a genuine gap: it covers time-series fusion and MM-LLM fusion as first-class categories, which most prior MSFP surveys do not, and it attempts a useful cross-task framing. The authors also gather relevant background on sensors and datasets, and the discussion of challenges (data quality, synchronization, explainability) is sensible. However, the paper's core value depends on the reliability of its method corpus and taxonomy, and the current manuscript does not yet demonstrate that reliability: the selection methodology is undocumented, several timeline entries are unreferenced, and key figures and tables contradict the text. These are fixable issues, but they are load-bearing for a survey claiming to be a comprehensive reference.","major_comments":[{"comment":"The paper claims a 'rigorous and detailed investigation' (Abstract) and 'a detailed investigation' (Section I), but it does not report the literature search strategy, inclusion/exclusion criteria, or the procedure by which the four-category taxonomy (multi-modal, multi-agent, time-series, MM-LLM) was derived. Without this methodology, the reader cannot determine whether the methods selected for Tables III–V and Figs. 7–8 are representative or exhaustive. Since the survey's stated value proposition is to provide a trustworthy task-agnostic map of MSFP methods, this missing documentation undermines the central claim of the paper.","section":"Section I and Abstract"},{"comment":"The timeline in Fig. 8 lists at least a dozen methods—PolarDETR, STS, MV-FCOS3D++, DORT, E-TMA, BridgeAD, SAD, RENet, Far3D, TLCFuse, PETR, and PETR v2—that are not cited in the reference list and are not discussed in the text. In addition, Fig. 8 contains an 'Others' group that has no counterpart in the three categories of Table IV, and no criterion is stated for what belongs in the taxonomy versus 'Others.' Because the survey claims comprehensiveness, these unreferenced entries and the unexplained extra category make the map unreliable and directly undercut the abstract's claim of a rigorous investigation.","section":"Section V, Fig. 8"},{"comment":"The first paragraph of Section V states 'Fig. 7 shows a simple pipeline of A2A fusion,' but Fig. 7 is placed in the time-series section and is captioned 'Framework overview of time series multi-sensor fusion network.' This is not an isolated typo: the confusion between agent-to-agent fusion (Section IV) and time-series fusion (Section V) matters because the distinction between these technical views is one of the paper's organizing axes. The figure reference and caption must be corrected so that each figure illustrates the correct category.","section":"Section V, Fig. 7"},{"comment":"Table IV's 'Methods' column is inconsistent with the text of Section V. The text discusses MUTR3D [97], PF-Track [98], FusionFormer [99], QTNet [101], and CRT-Fusion [102] as query-based time-series methods, but none of these appears in Table IV, while some table entries (e.g., SparseFusion3D) are discussed in the text. If the table is intended to be representative, that should be stated; if it is intended to be complete, these are omissions. In either case, the table cannot currently serve as a reliable quick-reference, which is the primary function of a survey table.","section":"Section V, Table IV"}],"minor_comments":[{"comment":"The sentence 'we will focus on the multi-view fusion of agent-to-agent (A2A) collaborative perception' conflates multi-view fusion (multiple cameras on one agent) with multi-agent fusion (multiple agents sharing information). Please clarify or rephrase to avoid conflating these two distinct technical views.","section":"Section IV, first paragraph"},{"comment":"The DETR3D paper appears twice as separate references ([89] and [100]) with different formatting and slightly different author lists; this should be consolidated into a single reference.","section":"References [89] and [100]"},{"comment":"Some methods listed in Table III (e.g., UVTR, SFD, E2E-MFD, MBNet) are not discussed in the body text, while several methods discussed in the text (e.g., PI-RCNN, FusionPainting, GraphAlign, VPFNet, VFF, AutoAlign, VoxelNextFusion, TransFusion, GAFF, RSDet, EPNet, DVF, LoGoNet, CAT-Det, SeaDATE, Fusion-Mamba) are not listed in the table. Please align the table with the text or state explicitly that the table is representative rather than exhaustive.","section":"Section III, Table III"},{"comment":"Fig. 8 is dense and the small font makes the method names difficult to read; consider a larger layout or a table-form timeline. Also, several figures (Figs. 2–10) are not explicitly referenced in the text at the point of discussion, which makes navigation harder.","section":"Fig. 8 and general figures"},{"comment":"Capitalization and terminology are inconsistent in a few places, e.g., 'V oxel-level' in Table III and 'LIDAR' vs. 'LiDAR' in Section III-D; a careful proofread would improve readability.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a survey whose main risk is not novelty but corpus reliability. The undocumented selection methodology and the internal inconsistencies (Fig. 8 unreferenced entries, Fig. 7 mislabeling, Table IV omissions) directly affect the trustworthiness of the map that the paper promises. I believe these are fixable within the scope of a revision: the authors can add a methodology subsection, complete the references, correct the figure–text alignment, and state the representativeness of each table. I do not see grounds for rejection, but the revision needs to be substantive rather than purely editorial. The authors' own earlier publications appear as general AI background ([1]–[3]) and are not used to support the central organizational claim, so there is no circularity concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the organizing scheme is real and useful. Most fusion surveys are pinned to 3D detection or autonomous driving, and this one deliberately steps back and groups methods by technical view — multi-modal, multi-agent, time-series, and MM-LLM. That framing is the paper's contribution, and it is a legitimate one for a survey. The MM-LLM section and the task-agnostic tables give newcomers something they will not find in the surveys listed in Table I.\n\nThe multi-modal fusion section does solid work: point, voxel, region, and multi-level is a standard but workable taxonomy, and the method descriptions are mostly accurate. The time-series section, organized by dense/sparse/hybrid query, is also a reasonable way to present that literature.\n\nThe soft spots are real, and one is load-bearing for a survey. Figure 8's timeline names a dozen-plus methods — PolarDETR, STS, MV-FCOS3D++, DORT, E-TMA, BridgeAD, SAD, RENet, Far3D, TLCFuse, PETR, PETR v2 — that are not in the reference list and not discussed in the text. For a survey claiming a 'rigorous and detailed investigation', a reader cannot check the selection or the coverage. That undermines the central value proposition: a trusted, task-agnostic map. It is fixable, but not cosmetic.\n\nThe other issues are smaller. Figure 7 is referenced in Section V as an A2A fusion pipeline but is captioned as the time-series framework; Section IV says 'multi-agent' but then narrows to 'multi-view fusion of agent-to-agent'; and the taxonomy in Table IV omits the 'Others' group that appears in Fig. 8. There is also no stated search methodology, inclusion/exclusion criteria, or justification for the taxonomy. The closing caveat about the authors' limited expertise is honest but does not belong in a formal survey.\n\nProportionate verdict: the survey is useful as an entry point, not yet reliable as a reference. With a documented selection process, the missing references restored, and the editorial errors cleaned up, it would serve the intended audience well. I would send it to peer review with a request for major revision, and I would read the revision.","headline":"The task-agnostic organizing scheme is genuinely useful, but the timeline in Fig. 8 lists many methods that are never cited or discussed, so the survey's claim to be a rigorous map is not yet supported.","tokens_in":25640,"tokens_out":2184,"would_cite":false,"duration_ms":20883,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A task-agnostic survey organizes multi-sensor fusion perception for embodied AI into four technical families, independent of any single task or application domain.","keywords":["multi-sensor fusion","embodied AI","multi-modal fusion","multi-agent collaborative perception","time-series fusion","multimodal large language models","autonomous driving","perception survey"],"falsifier":"Take a systematic sample of recent multi-sensor fusion papers from areas outside autonomous driving, such as thermal-visual pedestrian detection, visual-inertial odometry, or infrastructure-vehicle cooperation, and check whether each paper fits one of the four categories; any substantial cluster that does not fit would falsify the claim of task-agnostic coverage.","tokens_in":24640,"feed_emoji":"🧭","tokens_out":4083,"duration_ms":40867,"temperature":0.7,"pith_summary":"This paper tries to establish that multi-sensor fusion perception can be surveyed in a task-agnostic way, so that the same fusion techniques are visible to researchers working on any embodied-AI problem rather than only to those in autonomous driving or 3D detection. It argues that previous surveys are mostly task-specific or single-perspective and miss multi-agent fusion, time-series fusion, and large-model-based fusion. The authors organize the field into four technical families, covering the background, datasets, perception tasks, methods, and open challenges. If the survey is right, it gives researchers a shared vocabulary and a reference map for fusion across robotics, autonomous driving, and other embodied systems.","feed_headline":"Four fusion families map the field of embodied-AI perception","feed_subtitle":"Survey shows multi-modal, multi-agent, time-series, and LLM fusion techniques transfer across tasks, not just autonomous driving.","key_machinery":"The organizing device is a task-agnostic taxonomy built on the 'Agent-Sensor-Data-Model-Task' pipeline. The four categories are multi-modal fusion, multi-agent fusion, time-series fusion, and MM-LLM fusion, with time-series methods further split into dense query, sparse query, and hybrid query approaches. The taxonomy does the argument's work: it is what lets the survey claim that fusion techniques transfer across tasks rather than belonging to one application area.","core_discovery":"The central claim is that the diversity of multi-sensor fusion perception can be presented independently of any downstream task, using four technical families: multi-modal fusion (at point, voxel, region, and multi-level), multi-agent fusion, time-series fusion (with dense, sparse, and hybrid query representations), and multimodal-LLM fusion (vision-language and vision-LiDAR-language). The paper argues that this organization lets any embodied-AI researcher, whatever their task, find the fusion technique relevant to them. It supports the claim by reviewing representative methods in each family, cataloging datasets and evaluation criteria, and discussing open challenges at the data, model, and application levels.","pith_inferences":["A cross-cutting axis the paper leaves implicit is fusion stage, such as data-level versus feature-level versus decision-level fusion; readers could combine that axis with the four families to locate methods even more precisely.","If the task-agnostic claim holds, specialized survey writers could reuse the four categories as a standard outline, letting method results accumulate across domains instead of being fragmented by task.","The paper does not describe its literature search or inclusion criteria, so absence of a method family from the survey should be read as an open question rather than proof that the family does not exist.","A natural test of the taxonomy is whether papers on less common sensor pairs, such as thermal-RGB pedestrian detection or visual-inertial odometry, can be placed cleanly into one of the four families."],"forward_implications":["Researchers in tasks outside autonomous driving can locate applicable fusion techniques through the task-agnostic categories.","The four-way split makes visible method families, such as multi-agent collaboration and time-series fusion, that single-task surveys tend to omit.","The dense-versus-sparse-versus-hybrid query taxonomy offers a direct way to compare efficiency and accuracy trade-offs in temporal fusion.","The challenge discussion points to concrete research targets, including synchronized multi-modal data augmentation, explainable fusion, and handling sparse radar or LiDAR data within multimodal LLMs."],"supporting_citations":[{"why":"Represents the task-specific survey type the paper argues against: 3D object detection in autonomous driving.","marker":"[6]"},{"why":"Shows the single-perspective limitation by reviewing multi-sensor fusion mainly for autonomous driving cooperation.","marker":"[8]"},{"why":"Covers camera-LiDAR-IMU fusion only through the SLAM task, illustrating task-bounded organization.","marker":"[9]"},{"why":"Closest prior broad survey of multi-sensor fusion for embodied agents, which the paper's task-agnostic taxonomy extends.","marker":"[11]"},{"why":"Limits fusion discussion to humanoid-robot visual perception, another task-specific contrast.","marker":"[12]"},{"why":"Recent robot-vision survey combining multimodal fusion and vision-language models, showing the MM-LLM trend the paper includes.","marker":"[14]"}],"fun_headline_variants":["Four fusion families map embodied-AI perception","Task-agnostic survey: four fusion families for embodied AI","Survey breaks embodied-AI perception into four fusion views","One survey, four fusion lenses for embodied-AI perception"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness assumes that the methods it selected are a representative sample of the whole field and that its four-category taxonomy really captures the variety of multi-sensor fusion research.","fun_headline_variants_meta":{"raw":{"variants":["Four fusion families map embodied-AI perception","Task-agnostic survey: four fusion families for embodied AI","Survey breaks embodied-AI perception into four fusion views","One survey, four fusion lenses for embodied-AI perception"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1548,"prompt_tokens":939,"completion_tokens":609,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":545}},"tokens_in":555,"tokens_out":609,"duration_ms":6367,"temperature":1.0,"reasoning_tokens":545,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:24:27.387298+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a systematic sample of recent multi-sensor fusion papers from areas outside autonomous driving, such as thermal-visual pedestrian detection, visual-inertial odometry, or infrastructure-vehicle cooperation, and check whether each paper fits one of the four categories; any substantial cluster that does not fit would falsify the claim of task-agnostic coverage.","supporting_citations":[],"review_version":2}