{"id":"d43a3567-c439-4c32-8bb3-1fb4655bd6ea","arxiv_id":"2508.17298","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that classifies compositional visual reasoning methods into five stages, from prompt-based pipelines to unified agentic vision-language models, and catalogs associated benchmarks.","lead":"This paper reviews more than 260 studies on compositional visual reasoning, where AI models explain their steps before answering questions about images. It organizes the field into five developmental stages and catalogs benchmarks, aiming to serve as a reference for researchers building more transparent multimodal systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The five-stage 'paradigm shift' is asserted, not demonstrated: no dates, inclusion protocol, or stage-assignment rules are given, and at least one paper (CoF) is placed in two stages, so the central roadmap may be an organizing narrative rather than an empirical finding.","rationale":"The reader's weakest_assumption and my concern overlap: both target the five-stage taxonomy. I do not see a reason to change the conditional verdict, because the concern is about support for a central claim, not about internal mathematical soundness, and the paper contains no formal proof or quantitative validation of the progression. The survey has real strengths: a broad and current reference list, a useful benchmark catalog in Sec. 5, and a clear discussion of open problems in Sec. 6. Those stand even if the five stages are reframed as archetypes. The specific evidence I add is that stage assignment is not disciplined: CoF is explicitly used in both Stage IV and Stage V, and the chronological ordering is contradicted by examples like SEAL (CVPR 2024) being placed in the latest stage while contemporary tool-enhanced methods are placed in earlier stages. This does not destroy the survey's value, but it does mean the headline contribution, 'trace a five-stage paradigm shift,' is currently an assertion. The concrete test I propose, a date-and-assignment audit plus an inter-annotator reliability check, would settle whether the stages are empirically grounded. If the audit fails, the authors should either supply the missing selection protocol and assignment rules or soften the claim from 'paradigm shift' to 'organizational taxonomy.' Because this is a solvable presentation/support issue rather than a fundamental error, keeping the paper as CONDITIONAL is the right call; it could become ACCEPT after the audit is reported.","tokens_in":40434,"tokens_out":5809,"duration_ms":60028,"concrete_test":"Reconstruct the review corpus: for each of the 260+ reviewed papers, record the first-publication date, venue, and the stage(s) assigned in Section 4, then plot the date distribution per stage and check pairwise ordering (e.g., median date of Stage V vs Stage II). Independently, have two annotators reassign a stratified random sample of 50 papers to the five stages and measure inter-annotator agreement (Cohen's kappa) and the rate of multi-stage assignment. If Stage V papers are not systematically later than Stage II papers, or if kappa is below 0.6 or multi-assignment exceeds 20%, the 'paradigm shift' chronology and stage disjointness are not empirically supported; the survey should then reframe the stages as conceptual archetypes rather than a historical progression.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central novelty is the claim that it 'trace[s] a five-stage paradigm shift' in compositional visual reasoning, from prompt-enhanced language-centric pipelines to unified agentic VLMs. This claim is load-bearing because the paper presents itself as filling a gap by supplying a historical roadmap, not just a list of methods. The support for the ordering, however, is rhetorical: Section 4 motivates each stage with a stylized 'Question I–V' and Figure 3 draws a roadmap, but no bibliometric evidence, publication-date distribution, or explicit criterion for assigning a paper to a stage is provided. The stages are also not disjoint in the paper's own usage: CoF [78] is presented as a Stage IV visually grounded CoT VLM in Sec. 4.4.3 and again as a Stage V unified agentic VLM in Sec. 4.5.1; Visual Sketchpad [146] is a Stage III image-feedback tool-use method but is also a multi-step visual reasoning agent; SEAL [169], a Stage V exemplar, was published at CVPR 2024, concurrently with Stage II/III systems like LLaVA-Plus [140]. Without a rule for resolving overlaps or date evidence establishing the ordering, 'paradigm shift' is an organizing narrative rather than an empirical finding. The taxonomy may still be useful as a conceptual organization of architectures, but the specific claim of tracing a chronological shift is not currently supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of compositional visual reasoning (CVR) in the vision-language domain. It defines CVR, contrasts it with monolithic reasoning, and argues for its advantages in cognitive alignment, interpretability, generalization, and efficiency. The paper organizes recent methods into five developmental stages: prompt-enhanced language-centric pipelines, tool-enhanced LLMs, tool-enhanced VLMs, chain-of-thought VLMs, and unified agentic VLMs. It also catalogs benchmarks and evaluation metrics, and closes with insights, open challenges, and future directions. The authors claim to provide the first comprehensive, systematic review of 260+ papers from top venues between 2023 and 2025 and explicitly position the five-stage progression as a historical roadmap marking a 'paradigm shift.'","tokens_in":40764,"tokens_out":6369,"duration_ms":64276,"significance":"If the roadmap and taxonomy are validated, this survey would be a valuable reference for a rapidly growing area, bundling many recent systems into a structured framework and pointing to concrete open problems. The paper is clearly written and broad in coverage, with useful illustrations and a wide-ranging benchmark catalog. The authors are to be credited for assembling recent literature on tool-based, chain-of-thought, and agentic visual reasoning and for identifying challenges such as step-level evaluation and world-model integration. However, the survey's central contribution is the historical 'five-stage paradigm shift,' and as currently presented that claim is not supported by systematic evidence; the taxonomy may still be useful as a conceptual organization, but the specific 'historical roadmap' framing needs revision or additional evidence.","major_comments":[{"comment":"The central claim that the paper 'trace[s] a five-stage paradigm shift' is not supported by the evidence presented in the manuscript. No assignment rules are given for placing a paper into a stage, no publication-date distribution or bibliometric analysis establishes the proposed chronological ordering, and the stages are not disjoint in the paper's own usage: CoF [78] is described as a Stage IV visually grounded CoT VLM in Sec. 4.4.3 and again as a Stage V unified agentic VLM in Sec. 4.5.1, while Visual Sketchpad [146] appears as a Stage III image-feedback tool-use method in Sec. 4.3.3 despite also functioning as a multi-step visual reasoning agent. Because this roadmap is the survey's stated contribution beyond existing surveys, the authors should either provide explicit stage-assignment criteria and temporal evidence (e.g., a plot of publication dates of representative systems per stage) or explicitly reframe the five stages as a conceptual taxonomy of architectures rather than a chronological paradigm shift.","section":"Section 4 (Key Stages), including Figure 3"},{"comment":"The paper claims to 'systematically review 260+ papers from top venues,' but it reports no inclusion criteria, search strategy, screening process, or inter-annotator agreement for its taxonomy. This makes the 'comprehensive' and 'systematic' claims difficult to verify and leaves the survey open to selection bias; for example, the stage narratives in Sec. 4.2 and 4.3 rely heavily on the authors' own systems (HYDRA [29], DWIM [76], NA VER [137]), and no protocol is given that would explain how other relevant works were selected or excluded. A short methodology subsection describing the databases searched, the exact time window, inclusion/exclusion criteria, and the procedure for assigning works to stages would materially strengthen the survey and should be added before publication.","section":"Abstract and Section 1"}],"minor_comments":[{"comment":"The KL-divergence characterization of SFT versus RL is an oversimplification: supervised fine-tuning is better described as likelihood maximization (empirical forward KL), while RL fine-tuning such as RLHF typically optimizes a reward with a KL penalty to a reference policy rather than directly minimizing reverse KL. The authors should either provide a precise citation for this claim or soften the wording to 'interpreted as' to avoid misleading readers.","section":"Section 4.4.2"},{"comment":"TextVQA is cited as [212], but reference [212] in the bibliography is 'Towards Visual Dialog for Radiology' (Kovaleva et al.), not the TextVQA paper. Please replace it with the correct TextVQA citation and recheck all benchmark references in that section.","section":"Section 5.1"},{"comment":"The text refers to a 'Visual Abstractions Benchmark [190],' but reference [190] is 'What Makes a Maze Look Like a Maze?' by Hsu et al.; the dataset name and citation do not match, so the authors should verify the intended benchmark and its reference.","section":"Section 5.1"},{"comment":"Reference [168] for 'Visual Agents as Fast and Slow Thinkers' lists the venue as 'In The Thirteenth International Conference on Learning Representations, 2018,' but the thirteenth ICLR is in 2025; the year should be corrected.","section":"References"},{"comment":"The discussion of advances in evaluation says 'the next frontier lies in unifying these strengths' and mentions 'average output token length' in Sec. 5.2, but no reference or definition is given for the latter metric; please add a citation or brief explanation.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript makes a strong historical claim ('five-stage paradigm shift') that is currently supported only by narrative, and the absence of a methodology section is a concern for a survey that calls itself 'systematic.' The self-citation pattern in the stage narratives (HYDRA, DWIM, NA VER, and other works from the same groups) is noticeable and, while not disqualifying, should be examined by the editor when judging balance; adding a transparent selection protocol would also address this concern. The reference error for TextVQA in Section 5.1 is the kind of inaccuracy that can damage confidence in the benchmark catalog, so a careful pass over all citations is advised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Frankly, this survey is more useful as a reference work than as a historical account. The benchmark catalog and the organization of 260+ papers into a five-stage taxonomy give researchers a map that genuinely didn't exist before. The earlier surveys listed in Table 1 do cover multimodal reasoning or chain-of-thought, but none frame compositional visual reasoning as a distinct paradigm with its own evaluation landscape. If you need a quick orientation to this area, this is a good place to start.\n\nThe paper does several things well. The definitions are clear, the stage descriptions are readable, and the benchmark section is thorough — the 60+ datasets are grouped sensibly by task type, and the discussion of metrics is more concrete than most surveys. The open-challenges section is also honest, particularly the point that current benchmarks rarely evaluate intermediate reasoning steps, which is a real gap.\n\nThat said, the central claim that the field went through a \"five-stage paradigm shift\" is asserted, not shown. There is no reported search strategy, inclusion protocol, or inter-annotator agreement, so the corpus is not reproducible. More importantly, the stages are not disjoint in the paper's own usage: CoF is presented as both a Stage IV visually grounded CoT model and a Stage V unified agentic model; Visual Sketchpad is a Stage III tool-use method but is also a multi-step agent. And the chronology does not hold up — SEAL, a Stage V exemplar, came out at CVPR 2024 around the same time as Stage II/III systems like LLaVA-Plus. Without date distributions or a rule for assigning methods to stages, \"paradigm shift\" reads as a rhetorical device rather than a measured finding.\n\nThis is not fatal. The taxonomy still works as a conceptual organization of architectural ideas, and the authors could fix the problem by toning down the historical language and adding a short methodology section. The self-citations in the stage narrative are noticeable but not disqualifying; they are legitimate works and the survey does not depend on them.\n\nFor a reader, the value is real: it's a solid entry point and a handy benchmark index. I'd bring it to reading group, and I'd cite it. A serious referee should engage with it — but with a clear request to report the corpus construction and to separate the conceptual taxonomy from the unsupported chronological claim.","headline":"A genuinely useful reference survey with a solid benchmark catalog, but the five-stage 'paradigm shift' is a narrative device rather than an empirically demonstrated progression; worth engaging with after revisions.","tokens_in":41290,"tokens_out":2092,"would_cite":true,"duration_ms":21417,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Compositional visual reasoning is now a distinct paradigm with its own history, taxonomy, and benchmarks.","keywords":["compositional visual reasoning","visual question answering","vision-language models","chain-of-thought reasoning","tool-augmented language models","agentic vision-language models","visual grounding","benchmark survey"],"falsifier":"A chronological map of the 260+ surveyed papers by publication date, architecture, and stage assignment would settle the matter: if later-stage systems such as unified agentic vision-language models regularly predate earlier-stage systems, or if independent researchers cannot agree on stage labels for individual papers, then the five-stage progression is an imposed narrative rather than a discovered one.","tokens_in":40256,"feed_emoji":"🧩","tokens_out":5279,"duration_ms":51983,"temperature":0.7,"pith_summary":"This survey argues that a distinct research paradigm has emerged in multimodal AI: compositional visual reasoning, in which a model does not simply map an image-plus-question to an answer but first carries out explicit intermediate steps grounded in the image. The authors claim that no existing survey covers this rapidly growing area, and they fill that gap by systematically reviewing 260+ papers from 2023 to 2025, organizing them into five developmental stages, cataloging 60+ benchmarks, and extracting the field's insights, open challenges, and future directions. A reader should care because the paper supplies the shared vocabulary and roadmap that a young, fast-moving field needs to compare methods and set research priorities. The title captures the core commitment: explain before you answer.","feed_headline":"Five generations of visual reasoning, mapped in one survey","feed_subtitle":"A 260-paper review traces the shift from black-box answers to explicit, image-grounded reasoning steps.","key_machinery":"The organizing device is the five-stage taxonomy and roadmap of compositional visual reasoning paradigms, anchored by the formal definition of a compositional system as a sequence of grounded intermediate steps. The sequence $S = \\{s_1, \\dots, s_n\\}$ is the central object: it is the explicit explanation inserted between question and answer, and it may take the form of sub-questions, tool calls, scene graphs, chain-of-thought rationales, or visual manipulations such as cropping and zooming. The taxonomy does the work of partitioning 260+ papers into comparable families so that design choices and failure modes can be discussed at the level of paradigms rather than individual models. The survey also uses a benchmark catalog of 60+ datasets to connect each paradigm to the evidence used to judge it.","core_discovery":"The central claim is that visual reasoning has been undergoing a paradigm shift from monolithic models, which directly predict an answer from the image-question pair, to compositional systems that deliberately expose their reasoning. The survey formalizes compositional visual reasoning as any approach that maps a visual input $v$ and query $q$ to an answer $y$ through an intermediate structured representation $S = \\{s_1, \\dots, s_n\\}$, where each step may ground to objects, attributes, or relations, and the steps are executed sequentially or hierarchically. It then divides the recent literature into a five-stage progression: prompt-enhanced language-centric pipelines, tool-enhanced large language models, tool-enhanced vision-language models, chain-of-thought reasoning vision-language models, and unified agentic vision-language models. Each stage is characterized by its architectural design, strengths, and limitations, and the survey positions this progression as the field's historical roadmap.","pith_inferences":["Editorial inference: the five-stage ordering implies a falsifiable prediction that publication dates of the 260+ papers cluster by stage, with later stages peaking later; this can be checked bibliometrically.","Editorial inference: if step-level evaluation becomes standard, reported gains from chain-of-thought methods may shrink, because models can now be caught producing correct final answers from incorrect intermediate deductions.","Editorial inference: the monolithic-versus-compositional distinction suggests a testable hypothesis that forcing explicit grounding steps reduces hallucination rates more than prompt-based fixes do, measurable on benchmarks such as POPE or HallusionBench.","Editorial inference: a natural next step the survey names but does not formalize is a sixth stage, world-model-integrated agents that simulate hypothetical scenes during reasoning."],"forward_implications":["Papers in this area can now be located on a shared map: a method is a Stage I through Stage V system, making cross-paper comparisons more systematic.","The benchmark catalog gives researchers a ready-made test battery for grounding accuracy, chain-of-thought faithfulness, and high-resolution perception.","If the survey's evaluation critique is correct, future benchmark design should add step-level annotations rather than scoring only final answers.","The open-challenge list points to concrete next targets, including world-model integration, human-in-the-loop verification, and richer evaluation protocols.","The roadmap suggests the field's center of gravity is moving toward unified agentic vision-language models, so new work will likely build on those architectures."],"supporting_citations":[{"why":"Defines the canonical diagnostic setting for compositional language and elementary visual reasoning that the survey uses to motivate the field.","marker":"[5]"},{"why":"Provides a real-world compositional question-answering benchmark used throughout the survey for grounding and evaluation.","marker":"[73]"},{"why":"Exemplifies the program-based, training-free composition that anchors the tool-enhanced stages of the roadmap.","marker":"[72]"},{"why":"Represents code-generation compositional reasoning, a key baseline for the tool-enhanced paradigm.","marker":"[129]"},{"why":"Serves as the monolithic instruction-tuned vision-language model that compositional methods are contrasted against.","marker":"[9]"},{"why":"Supplies the frozen perception module used by many early prompt-decomposition pipelines.","marker":"[11]"},{"why":"Illustrates the chain-of-thought stage by injecting tool-generated reasoning traces into a vision-language model.","marker":"[100]"},{"why":"Exemplifies the agentic stage with visual working memory, targeted search, and zooming.","marker":"[169]"},{"why":"Demonstrates reinforcement-learning-controlled visual zooming used to ground multi-step reasoning in the agentic stage.","marker":"[78]"}],"fun_headline_variants":["Visual reasoning's five-stage shift, mapped","Explain first, then answer: the new visual reasoning","260+ papers, one roadmap for visual reasoning","From black-box to step-by-step visual inference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole roadmap depends on the assumption that the field really evolved through those five stages in that order, which the survey asserts without quantitative evidence that the stages are historically real, mutually distinct, and non-overlapping.","fun_headline_variants_meta":{"raw":{"variants":["Visual reasoning's five-stage shift, mapped","Explain first, then answer: the new visual reasoning","260+ papers, one roadmap for visual reasoning","From black-box to step-by-step visual inference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00031,"raw_usage":{"total_tokens":1806,"prompt_tokens":1020,"completion_tokens":786,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":727}},"tokens_in":636,"tokens_out":786,"duration_ms":8130,"temperature":1.0,"reasoning_tokens":727,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:05:28.844865+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A chronological map of the 260+ surveyed papers by publication date, architecture, and stage assignment would settle the matter: if later-stage systems such as unified agentic vision-language models regularly predate earlier-stage systems, or if independent researchers cannot agree on stage labels for individual papers, then the five-stage progression is an imposed narrative rather than a discovered one.","supporting_citations":[],"review_version":1}