{"id":"e8dfdf27-7fd8-4736-b0a1-ea3ce3d8f757","arxiv_id":"2505.04921","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey organizes multimodal reasoning research into a staged roadmap and proposes native large multimodal reasoning models that unify perception, generation, and agentic planning.","lead":"This paper surveys how AI systems that understand text, images, audio, and video have learned to reason, organizing roughly 700 studies into a historical roadmap. It proposes a future class of \"native\" multimodal reasoning models that would combine perception, generation, and planning in one unified architecture.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.1's anecdotal case studies do not establish that language-centric architecture causes omni-modal and agentic failures, so the N-LMRM motivation is not yet supported. The survey remains useful if this is framed as a hypothesis.","rationale":"I agree with the reader's weakest assumption. The central N-LMRM proposal is forward-looking, and the only empirical evidence offered for it is Section 4.1: a table of benchmark aggregates plus a handful of o3/o4-mini case studies. Neither is sufficient to attribute observed failures to a language-centric architecture, because the alternative explanations (tool limitations, data distribution, perceptual encoders, reward design) are not controlled away. I am not objecting to the survey's taxonomy or its speculative vision; I am objecting to presenting that vision as empirically grounded. The proposed concrete test is feasible with existing APIs and would settle whether the anecdotal failures are representative and whether they are specific to the language-centric substrate. If the test does not support the causal claim, the authors should label the N-LMRM direction as a hypothesis and soften the Section 4 opening. That would preserve the survey's value while making its evidentiary status clear. Since the reader already recommends CONDITIONAL, my read does not move the verdict; it sharpens the condition under which acceptance is appropriate.","tokens_in":53954,"tokens_out":3758,"duration_ms":39929,"concrete_test":"Replicate Section 4.1 with a predetermined protocol: select the three reported task families (finger counting, resume-PDF information extraction, puzzle solving), build 50 items per family, and run o3/o4-mini under three conditions: (i) full multimodal input, (ii) text-only description of the same stimulus, and (iii) multimodal input with external tools (e.g., image cropping/OCR). Have independent coders classify each error as perception, tool, reasoning, or hallucination. If the multimodal-vs-text gap is small, or if tool access eliminates most errors, the claimed language-centric bottleneck is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that multimodal reasoning is heading toward N-LMRMs, and the stated reason is that \"language-centric architectures impose critical constraints\" (Section 4 opening). The load-bearing support is Section 4.1, which presents two kinds of evidence: (a) selected omni-modal and agentic benchmark aggregates (Table 12, with GPT-4o 0.6% on BrowseComp, Claude 3.5 Sonnet 35% on WorldSense, etc.) and (b) a small set of o3/o4-mini case studies (finger counting, resume PDF parsing, puzzle solving). Neither establishes the causal claim. The benchmark numbers are not accompanied by a protocol, per-model score tables, or controls for task difficulty; BrowseComp is text-only, so it cannot show a multimodal-specific deficit. The case studies are hand-picked, with no sample sizes, no inter-rater coding, and no baseline. More importantly, even if every anecdote is accurate, the failure modes (visual grounding, PDF/tool handling, reward hacking during post-training) are not shown to be caused by a language-centric reasoning substrate; they could persist in any architecture. The leap from \"current models fail at these tasks\" to \"we need a natively non-language-centric architecture\" is therefore unsupported as it stands. This concern is load-bearing because it is the only empirical evidence for the paper's forward-looking proposal; the rest of the survey is taxonomy. The stage-count inconsistency (abstract: four-stage; Section 1: three-stage; Section 2: four-stage) is secondary but compounds the issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of large multimodal reasoning models (LMRMs). It proposes a developmental roadmap in which multimodal reasoning evolves from perception-driven modular systems, through language-centric short reasoning (System-1) and language-centric long reasoning (System-2), toward a prospective fourth stage called Native Large Multimodal Reasoning Models (N-LMRMs). The survey covers roughly 700 publications, organizes methods into detailed taxonomies, reviews multimodal RL-enhanced reasoning (O1-like and R1-like models), and catalogs datasets and benchmarks for understanding, generation, reasoning, and planning. The forward-looking section motivates N-LMRMs with benchmark aggregates and case studies of OpenAI o3 and o4-mini, and outlines capabilities and technical prospects for such models.","tokens_in":54263,"tokens_out":4064,"duration_ms":39123,"significance":"The survey is timely and broad, and its organization around a staged roadmap is a useful contribution to a rapidly growing field. The benchmark and dataset reorganization in Section 5, as well as the extensive tables of recent RL-based multimodal reasoning methods, provide a valuable reference. The N-LMRM concept is an interesting forward-looking synthesis that connects omni-modal understanding, generation, and agentic reasoning. However, the empirical support for the N-LMRM proposal is anecdotal, and several internal inconsistencies affect the clarity of the central contribution. The paper is likely to be useful to practitioners and researchers, but the forward-looking claims need to be reframed or supported more rigorously.","major_comments":[{"comment":"The opening of Section 4 states that 'language-centric architectures impose critical constraints,' and Section 4.2 says N-LMRMs are introduced 'based on the above experimental findings.' This is a load-bearing causal claim, but Section 4.1 does not provide sufficient evidence to support it. Table 12 lists benchmarks without a protocol, per-model score tables, or controls for task difficulty; for example, BrowseComp is a text-only web-browsing benchmark, so the reported GPT-4o accuracy of 0.6% cannot demonstrate a multimodal-specific deficit. The o3/o4-mini case studies in Figures 6-8 are hand-picked examples with no sample sizes, selection criteria, or baselines, and the observed failure modes (finger counting, PDF parsing, puzzle rationalization) could plausibly persist in any architecture, not specifically because of a language-centric design. The paper should either add systematic evidence, such as controlled comparisons across model families and modality conditions, or explicitly reframe the N-LMRM proposal as a hypothesis motivated by observed limitations rather than a conclusion established by these experiments. Additionally, Section 4.1 is titled 'Preliminary Study with o3 and o4-mini,' but only o3 results are reported; no o4-mini experiments appear in the text, despite the abstract mentioning 'experimental cases of OpenAI O3 and O4-mini.'","section":"Section 4 and Section 4.1"},{"comment":"The number of stages in the proposed roadmap is inconsistent across the paper. The abstract says the survey is 'organized around a four-stage developmental roadmap,' and Section 2 says 'we outline four key stages.' However, Section 1 states the roadmap is 'organized into three stages (Figure 2),' and the contributions list in Section 1 describes a 'three-stage roadmap.' Since the staged roadmap is the paper's central organizational contribution, this inconsistency is not merely cosmetic. The authors should use one consistent count throughout, for example by making explicit that Stages 1-3 are historical stages and Stage 4 is a prospective direction, and then aligning the abstract, introduction, Section 2, and the roadmap figure accordingly.","section":"Abstract, Section 1, Section 2, and Section 3"}],"minor_comments":[{"comment":"In Section 5.1.1, 'GQA' is cited as 'Ainslie et al., 2023,' but that reference is for 'GQA: Training Generalized Multi-Query Transformer Models,' not the visual question answering dataset. The correct citation is Hudson & Manning (2019), which does appear correctly in Section 3.1.2 and Table 14. The authors should correct this citation error throughout the manuscript.","section":"Section 5.1.1, Section 3.1.2, Table 14"},{"comment":"The heading 'Preliminary Study with o3 and o4-mini' promises results for both models, but the text only reports evaluations of o3. Either add o4-mini results or change the heading and the abstract's wording to refer only to o3.","section":"Section 4.1"},{"comment":"There are multiple typographical errors that should be fixed: 'Planing' instead of 'Planning' in Section 5's introductory paragraph, 'Omini-Modal' instead of 'Omni-Modal' in the Conclusion, 'syatem-2' instead of 'system-2' in the Section 3.3.3 takeaways, 'Trasnformer' instead of 'Transformer' in Section 3.1.1, and 'Operater' instead of 'Operator' in Section 4.2.","section":"Several sections"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a broad survey with a useful taxonomy, but the internal stage-count inconsistency and the thin empirical basis for the N-LMRM proposal should be addressed before publication. The GQA citation error may indicate a broader need for a reference audit. The o3/o4-mini case studies appear to be the authors' own evaluations without a described methodology; providing a detailed protocol or supplementary material would strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a broad and genuinely useful survey of multimodal reasoning models, covering roughly 700 works with a clear taxonomy and a benchmark/dataset catalog current to mid-2025. That alone makes it worth having on the shelf. The three-stage historical narrative (perception-driven modular, language-centric short reasoning, language-centric long reasoning) is a reasonable way to organize the literature, and the proposed fourth stage, N-LMRMs, gives the field a forward-looking label that may stick even if the specific architecture does not.\n\nThe concrete weaknesses are real but concentrated. First, the GQA citation is flatly wrong: the paper cites Ainslie et al. 2023 (the multi-query transformer paper) instead of Hudson & Manning 2019. That is an easy fix but an alarming one in a survey where citation accuracy is the product. Second, the stage count is inconsistent: the abstract and Section 2 announce four stages, while the introduction and contributions say three. A reader cannot tell whether N-LMRMs are stage four or a separate prospect section. Both should be reconciled.\n\nThe bigger issue is Section 4.1. The o3/o4-mini case studies are anecdotal: a handful of hand-picked examples, no protocol, no sample sizes, no baselines. They are presented as \"experimental findings\" that motivate replacing language-centric architectures with N-LMRMs, but they do not show that language-centricity causes the observed failures. The benchmark numbers in Table 12 are similarly unsupported--no per-model scores, no controls, and some are text-only tasks like BrowseComp that cannot demonstrate a multimodal-specific deficit. This is a load-bearing gap because it is the only empirical evidence for the paper's central proposal. The N-LMRM concept is a reasonable research agenda, but it is a hypothesis, not a finding.\n\nThe rest of the survey is solid. The RL-enhanced multimodal reasoning tables (R1-style methods) are extensive and current, the benchmark organization is sensible, and the taxonomy of MCoT variants is helpful. The self-citation burden is minimal and not a concern. The paper would benefit from a serious referee who can push the authors to either label Section 4.1 as illustrative case studies or replace it with a systematic evaluation. As a survey, it deserves peer review and likely publication after revision--readers in the multimodal reasoning community will cite it for the taxonomy and the benchmark list, even if they argue with the N-LMRM agenda.\n\nI would bring it to a reading group, but with the caveat that the forward-looking section is speculative. I would also cite it for the survey content, not for the empirical claims.","headline":"A useful, up-to-date survey with a forward-looking agenda whose empirical motivation is anecdotal; worth refereeing, but the N-LMRM case needs to be reframed as a hypothesis.","tokens_in":54878,"tokens_out":1384,"would_cite":true,"duration_ms":15964,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multimodal reasoning has evolved through four stages and now heads toward the Native Large Multimodal Reasoning Model (N-LMRM), where reasoning emerges from omni-modal perception and goal-driven interaction instead of being retrofitted…","keywords":["multimodal reasoning","large multimodal reasoning models","chain-of-thought","reinforcement learning","omni-modal understanding","agentic reasoning","research roadmap","survey"],"falsifier":"Run each of the paper's three o3/o4-mini failure modes — six-finger emoji counting, phone-number extraction from resume PDFs, and red-panda multimedia generation — on a systematic sample of roughly one hundred analogous instances per case, comparing the same models against a language-centric open model and a native omni-modal model matched for compute. If the language-centric models fail at about the same rate as the native model, or if the failures disappear once browsing tools and internet access are provided, the claim that language-centric architecture is the bottleneck would be refuted. A weaker decisive check: test whether accuracy on omni-modal benchmarks falls as the share of non-textual tokens inside the model's chain of thought rises.","tokens_in":53758,"feed_emoji":"🧠","tokens_out":10941,"duration_ms":85821,"temperature":0.7,"pith_summary":"This paper surveys roughly seven hundred publications to argue that multimodal reasoning develops in identifiable stages, and that the field is now at a turning point. Its central thesis is that today's large multimodal reasoning models (LMRMs) bolt perception onto a language model that thinks in text, and that this language-centric design has hit a ceiling on tasks that require combining many modalities, planning over long horizons, and interacting with dynamic environments. The survey organizes the literature into a four-stage roadmap — perception-driven modular reasoning, language-centric short (System-1) reasoning, language-centric long (System-2) reasoning, and a proposed fourth stage, native large multimodal reasoning models (N-LMRMs) — and presents benchmark evidence plus case studies of the o3 and o4-mini models to show where current systems fail. A sympathetic reader would care because the paper reframes the field's goal: not better text-based reasoning about images, but a single end-to-end architecture in which perception, generation, and reasoning are one native process.","feed_headline":"Multimodal reasoning is heading toward native omni-modal models","feed_subtitle":"A survey maps four stages from modular pipelines to reasoners that perceive, generate, and plan across every modality.","key_machinery":"The load-bearing construct is the four-stage developmental roadmap, which doubles as a taxonomy and as an argument. Each stage is defined by where reasoning resides: (1) modular reasoning networks and pretrained vision-language models, where reasoning is implicit in representation, alignment, and fusion; (2) language-centric System-1 reasoning, where short chains emerge from prompt-based MCoT, structural reasoning, and externally augmented reasoning; (3) language-centric System-2 reasoning, where long chains and planning come from cross-modal reasoning, O1-style models, and R1-style reinforcement learning; and (4) N-LMRMs, defined by two capabilities — Multimodal Agentic Reasoning (hierarchical planning, dynamic adaptation, embodied learning) and Omni-Modal Understanding and Generative Reasoning (unified representations, cross-modal synthesis, modality-agnostic inference). The roadmap does the argumentative work: by locating today's best models in Stage 3 and presenting Section 4.1's benchmark failures and o3/o4-mini case studies as symptoms of a language-centric ceiling, the survey converts a design preference into a diagnosed gap that Stage 4 is positioned to fill.","core_discovery":"The paper's central claim is that where reasoning lives in the architecture defines the era of multimodal AI. In Stage 1, reasoning was implicit, distributed across task-specific modules for representation, alignment, and fusion. In Stage 2, reasoning became explicit but shallow: language-centric models produced short, reactive chains through prompt-based Multimodal Chain-of-Thought (MCoT), structural reasoning, and external augmentation. In Stage 3, chains lengthened into deliberate System-2 thinking through cross-modal reasoning, O1-style long reasoning, and reinforcement learning (DPO and GRPO), typified by the R1 line of models. Drawing on omni-modal and agentic benchmarks and on hands-on case studies of the o3 and o4-mini models — six-finger emoji counting, resume PDF parsing, puzzle solving, multimedia generation — the paper argues that the language-centric paradigm cannot reach real-world utility, and introduces the Native Large Multimodal Reasoning Model (N-LMRM): a forward-looking architecture where reasoning natively emerges from omnimodal perception and interaction and from goal-driven cognition, combining Multimodal Agentic Reasoning with Omni-Modal Understanding and Generative Reasoning.","pith_inferences":["A controlled test the paper does not report: a native omni-modal model matched against a language-centric LMRM on the same data, compute, and RL budget, on the same omni-modal and agentic tasks; the roadmap predicts the native model wins, and that experiment would settle it.","The case studies imply a strong testable corollary: fabricated reasoning ('lying' rationales attached to correct answers) is a symptom of language-centric post-training and would shrink when reasoning is grounded in perceptual tokens.","Tool-augmented and search-augmented systems (deep-research-style agents) may reach N-LMRM-level task performance without a native architecture, which would weaken the necessity claim; the paper leaves this stopgap undiscussed as a rival.","If the roadmap generalizes, a similar staged trajectory should appear in audio-centric and embodied subfields — early modular perception, then language-mediated short reasoning, then long-horizon System-2 behavior — a pattern that could be checked against the literature the survey catalogs."],"forward_implications":["Future flagship multimodal systems would abandon the vision-encoder-plus-LLM design for unified representation spaces that treat text, image, audio, video, and sensor streams symmetrically.","Long-horizon agentic behavior — GUI navigation, embodied interaction, multi-tool chains — would become a core training objective rather than an outer wrapper, with reinforcement learning scaled across modalities.","Evaluation would shift from static question answering (MMMU, MathVista) toward omni-modal and interactive benchmarks, because the paper argues current benchmarks overstate real-world capability.","Interleaved multimodal chain-of-thought — reasoning traces that include image crops, audio snippets, or actions — would open a new axis of test-time compute scaling.","Reasoning training (O1/R1-style) would extend beyond text-heavy math and vision tasks to cross-modal generation and planning, where the model not only thinks but generates intermediate multimodal content."],"supporting_citations":[{"why":"Supplies the archetypal Stage-1 mechanism — dynamically assembled task-specific modules — which is the starting point of the roadmap.","marker":"(Andreas et al., 2016)"},{"why":"LLaVA anchors the language-centric paradigm of Stages 2-3 as the vision-encoder-plus-LLM architecture that instruction tuning improves.","marker":"(Liu et al., 2023a)"},{"why":"Multimodal-CoT grounds Stage 2's MCoT branch with the two-stage rationale-then-answer framework.","marker":"(Zhang et al., 2023g)"},{"why":"DeepSeek-R1 supplies Stage 3's reinforcement-learning formula (GRPO) and the promise of generalizable long-chain reasoning.","marker":"(DeepSeek-AI et al., 2025)"},{"why":"The o1 model is the reference point for the Multimodal-O1 long-reasoning line in Stage 3.","marker":"(Jaech et al., 2024)"},{"why":"The o3 and o4-mini release provides the case studies (finger counting, resume parsing, puzzle solving) that drive Section 4's case for N-LMRMs.","marker":"(OpenAI, 2025b)"},{"why":"OmniBench is the omni-modal benchmark whose low scores support the claim that current models cannot handle unified multimodal inputs.","marker":"(Li et al., 2024j)"},{"why":"OSWorld provides the agentic benchmark evidence for weak long-horizon, environment-interactive behavior.","marker":"(Xie et al., 2024a)"},{"why":"MMMU marks the static question-answering tasks current LMRMs handle well, used to contrast exam-style accuracy with real-world utility.","marker":"(Yue et al., 2024)"}],"fun_headline_variants":["Survey maps road from modular pipelines to native multimodal reasoning","Where reasoning lives defines each era of multimodal AI","From shallow chains to native omni-modal reasoning","Multimodal reasoning's future: native omni-modal models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument for replacing language-centric models rests on a small set of hand-picked failure examples (finger counting, resume PDF parsing, a puzzle) and aggregate benchmark scores, presented without a sampling protocol or a matched baseline; if those failures actually reflect task difficulty or missing tools rather than the language-centric design, the case for native multimodal reasoning models loses most of its force.","fun_headline_variants_meta":{"raw":{"variants":["Survey maps road from modular pipelines to native multimodal reasoning","Where reasoning lives defines each era of multimodal AI","From shallow chains to native omni-modal reasoning","Multimodal reasoning's future: native omni-modal models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1369,"prompt_tokens":1081,"completion_tokens":288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":697,"completion_tokens_details":{"reasoning_tokens":225}},"tokens_in":697,"tokens_out":288,"duration_ms":3037,"temperature":1.0,"reasoning_tokens":225,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:17:36.659991+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run each of the paper's three o3/o4-mini failure modes — six-finger emoji counting, phone-number extraction from resume PDFs, and red-panda multimedia generation — on a systematic sample of roughly one hundred analogous instances per case, comparing the same models against a language-centric open model and a native omni-modal model matched for compute. If the language-centric models fail at about the same rate as the native model, or if the failures disappear once browsing tools and internet access are provided, the claim that language-centric architecture is the bottleneck would be refuted. A weaker decisive check: test whether accuracy on omni-modal benchmarks falls as the share of non-textual tokens inside the model's chain of thought rises.","supporting_citations":[],"review_version":1}