Pith. sign in

REVIEW 2 major objections 5 minor 4 cited by

Explain Before You Answer: A Survey on Compositional Visual Reasoning

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Compositional visual reasoning is now a distinct paradigm with its own history, taxonomy, and benchmarks.

desk verdict A genuinely useful reference survey with a solid benchmark catalog, but the five-stage 'paradigm shift' is a narrative device rather than an empirically demonstrated progression; worth engaging with after revisions. read the letter →

arxiv 2508.17298 v3 pith:RN2WI3MQ submitted 2025-08-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords compositionalvisualreasoningquestionansweringvision-languagemodelschain-of-thoughttool-augmentedlanguageagenticgroundingbenchmarksurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that a distinct research paradigm has emerged in multimodal AI: compositional visual reasoning, in which a model does not simply map an image-plus-question to an answer but first carries out explicit intermediate steps grounded in the image. The authors claim that no existing survey covers this rapidly growing area, and they fill that gap by systematically reviewing 260+ papers from 2023 to 2025, organizing them into five developmental stages, cataloging 60+ benchmarks, and extracting the field's insights, open challenges, and future directions. A reader should care because the paper supplies the shared vocabulary and roadmap that a young, fast-moving field needs to compare methods and set research priorities. The title captures the core commitment: explain before you answer.

What carries the argument

The organizing device is the five-stage taxonomy and roadmap of compositional visual reasoning paradigms, anchored by the formal definition of a compositional system as a sequence of grounded intermediate steps. The sequence $S = \{s_1, \dots, s_n\}$ is the central object: it is the explicit explanation inserted between question and answer, and it may take the form of sub-questions, tool calls, scene graphs, chain-of-thought rationales, or visual manipulations such as cropping and zooming. The taxonomy does the work of partitioning 260+ papers into comparable families so that design choices and failure modes can be discussed at the level of paradigms rather than individual models. The survey also uses a benchmark catalog of 60+ datasets to connect each paradigm to the evidence used to judge it.

What would settle it

A chronological map of the 260+ surveyed papers by publication date, architecture, and stage assignment would settle the matter: if later-stage systems such as unified agentic vision-language models regularly predate earlier-stage systems, or if independent researchers cannot agree on stage labels for individual papers, then the five-stage progression is an imposed narrative rather than a discovered one.

Watch

Extended reading notes

Core claim

The central claim is that visual reasoning has been undergoing a paradigm shift from monolithic models, which directly predict an answer from the image-question pair, to compositional systems that deliberately expose their reasoning. The survey formalizes compositional visual reasoning as any approach that maps a visual input $v$ and query $q$ to an answer $y$ through an intermediate structured representation $S = \{s_1, \dots, s_n\}$, where each step may ground to objects, attributes, or relations, and the steps are executed sequentially or hierarchically. It then divides the recent literature into a five-stage progression: prompt-enhanced language-centric pipelines, tool-enhanced large language models, tool-enhanced vision-language models, chain-of-thought reasoning vision-language models, and unified agentic vision-language models. Each stage is characterized by its architectural design, strengths, and limitations, and the survey positions this progression as the field's historical roadmap.

Load-bearing premise

The whole roadmap depends on the assumption that the field really evolved through those five stages in that order, which the survey asserts without quantitative evidence that the stages are historically real, mutually distinct, and non-overlapping.

Editorial extensions

If this is right

  • Papers in this area can now be located on a shared map: a method is a Stage I through Stage V system, making cross-paper comparisons more systematic.
  • The benchmark catalog gives researchers a ready-made test battery for grounding accuracy, chain-of-thought faithfulness, and high-resolution perception.
  • If the survey's evaluation critique is correct, future benchmark design should add step-level annotations rather than scoring only final answers.
  • The open-challenge list points to concrete next targets, including world-model integration, human-in-the-loop verification, and richer evaluation protocols.
  • The roadmap suggests the field's center of gravity is moving toward unified agentic vision-language models, so new work will likely build on those architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the five-stage ordering implies a falsifiable prediction that publication dates of the 260+ papers cluster by stage, with later stages peaking later; this can be checked bibliometrically.
  • Editorial inference: if step-level evaluation becomes standard, reported gains from chain-of-thought methods may shrink, because models can now be caught producing correct final answers from incorrect intermediate deductions.
  • Editorial inference: the monolithic-versus-compositional distinction suggests a testable hypothesis that forcing explicit grounding steps reduces hallucination rates more than prompt-based fixes do, measurable on benchmarks such as POPE or HallusionBench.
  • Editorial inference: a natural next step the survey names but does not formalize is a sixth stage, world-model-integrated agents that simulate hypothetical scenes during reasoning.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a survey of compositional visual reasoning (CVR) in the vision-language domain. It defines CVR, contrasts it with monolithic reasoning, and argues for its advantages in cognitive alignment, interpretability, generalization, and efficiency. The paper organizes recent methods into five developmental stages: prompt-enhanced language-centric pipelines, tool-enhanced LLMs, tool-enhanced VLMs, chain-of-thought VLMs, and unified agentic VLMs. It also catalogs benchmarks and evaluation metrics, and closes with insights, open challenges, and future directions. The authors claim to provide the first comprehensive, systematic review of 260+ papers from top venues between 2023 and 2025 and explicitly position the five-stage progression as a historical roadmap marking a 'paradigm shift.'

Significance. If the roadmap and taxonomy are validated, this survey would be a valuable reference for a rapidly growing area, bundling many recent systems into a structured framework and pointing to concrete open problems. The paper is clearly written and broad in coverage, with useful illustrations and a wide-ranging benchmark catalog. The authors are to be credited for assembling recent literature on tool-based, chain-of-thought, and agentic visual reasoning and for identifying challenges such as step-level evaluation and world-model integration. However, the survey's central contribution is the historical 'five-stage paradigm shift,' and as currently presented that claim is not supported by systematic evidence; the taxonomy may still be useful as a conceptual organization, but the specific 'historical roadmap' framing needs revision or additional evidence.

major comments (2)
  1. [Section 4 (Key Stages), including Figure 3] The central claim that the paper 'trace[s] a five-stage paradigm shift' is not supported by the evidence presented in the manuscript. No assignment rules are given for placing a paper into a stage, no publication-date distribution or bibliometric analysis establishes the proposed chronological ordering, and the stages are not disjoint in the paper's own usage: CoF [78] is described as a Stage IV visually grounded CoT VLM in Sec. 4.4.3 and again as a Stage V unified agentic VLM in Sec. 4.5.1, while Visual Sketchpad [146] appears as a Stage III image-feedback tool-use method in Sec. 4.3.3 despite also functioning as a multi-step visual reasoning agent. Because this roadmap is the survey's stated contribution beyond existing surveys, the authors should either provide explicit stage-assignment criteria and temporal evidence (e.g., a plot of publication dates of representative systems per stage) or explicitly reframe the five stages as a conceptual taxonomy of architectures rather than a chronological paradigm shift.
  2. [Abstract and Section 1] The paper claims to 'systematically review 260+ papers from top venues,' but it reports no inclusion criteria, search strategy, screening process, or inter-annotator agreement for its taxonomy. This makes the 'comprehensive' and 'systematic' claims difficult to verify and leaves the survey open to selection bias; for example, the stage narratives in Sec. 4.2 and 4.3 rely heavily on the authors' own systems (HYDRA [29], DWIM [76], NA VER [137]), and no protocol is given that would explain how other relevant works were selected or excluded. A short methodology subsection describing the databases searched, the exact time window, inclusion/exclusion criteria, and the procedure for assigning works to stages would materially strengthen the survey and should be added before publication.
minor comments (5)
  1. [Section 4.4.2] The KL-divergence characterization of SFT versus RL is an oversimplification: supervised fine-tuning is better described as likelihood maximization (empirical forward KL), while RL fine-tuning such as RLHF typically optimizes a reward with a KL penalty to a reference policy rather than directly minimizing reverse KL. The authors should either provide a precise citation for this claim or soften the wording to 'interpreted as' to avoid misleading readers.
  2. [Section 5.1] TextVQA is cited as [212], but reference [212] in the bibliography is 'Towards Visual Dialog for Radiology' (Kovaleva et al.), not the TextVQA paper. Please replace it with the correct TextVQA citation and recheck all benchmark references in that section.
  3. [Section 5.1] The text refers to a 'Visual Abstractions Benchmark [190],' but reference [190] is 'What Makes a Maze Look Like a Maze?' by Hsu et al.; the dataset name and citation do not match, so the authors should verify the intended benchmark and its reference.
  4. [References] Reference [168] for 'Visual Agents as Fast and Slow Thinkers' lists the venue as 'In The Thirteenth International Conference on Learning Representations, 2018,' but the thirteenth ICLR is in 2025; the year should be corrected.
  5. [Section 5.3] The discussion of advances in evaluation says 'the next frontier lies in unifying these strengths' and mentions 'average output token length' in Sec. 5.2, but no reference or definition is given for the latter metric; please add a citation or brief explanation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey organizes the literature with a taxonomy and roadmap, and its claims do not reduce to fitted inputs, self-citation chains, or definitional equivalences.

full rationale

This is a survey paper with no mathematical derivations, fitted parameters, or predictive claims that could reduce to their inputs by construction. The central contribution is organizational: the paper defines compositional visual reasoning, reviews 260+ papers, and proposes a five-stage taxonomy in Section 4. The five-stage 'paradigm shift' is asserted rhetorically rather than proven empirically, and some stage assignments overlap (e.g., CoF appears in both Stage IV and Stage V), but an asserted or under-evidenced historical narrative is a correctness or scope concern, not circularity. The paper's self-citations (HYDRA, DWIM, NAVER, LEFT, and others) are used as representative examples of existing methods, not as load-bearing evidence for a theorem or uniqueness result; the taxonomy would stand even if those specific papers were replaced. There is no equation where an output is defined in terms of the input, no fitted value relabeled as a prediction, and no imported uniqueness theorem that forces the chosen organization. Accordingly, the survey is self-contained as a literature review and no circular step is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters or invented scientific entities. Its load-bearing assumptions are definitional and organizational: what counts as compositional visual reasoning and whether the five-stage taxonomy is a faithful representation of the literature. These are interpretive choices, not empirical inputs, and they are not independently grounded.

assumptions (3)
  • domain assumption Compositionality principle: the meaning of the whole is a function of the meanings of its parts.
    Invoked in Section 2.3 as the conceptual foundation for defining compositional visual reasoning.
  • ad hoc to paper Explicit intermediate reasoning steps are the defining characteristic of compositional visual reasoning.
    The paper defines CVR in terms of 'explicit thought steps' in Sections 1 and 2; this is the authors' framing rather than an established fact in the literature.
  • ad hoc to paper The five-stage progression reflects the actual chronological and architectural development of the field from 2023 to 2025.
    Section 4 presents the five stages as a historical roadmap, but no quantitative evidence is given that stages are chronologically distinct or that the ordering is the most useful one.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Explain Before You Answer: A Survey on Compositional Visual Reasoning." pith.science (2026). https://pith.science/paper/RN2WI3MQ

@misc{pith2026250817298,
  author       = {Pith},
  title        = {Pith review of: Explain Before You Answer: A Survey on Compositional Visual Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RN2WI3MQ}},
  note         = {Machine review of arXiv:2508.17298}
}
read the original abstract

Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference. While early surveys focus on monolithic vision-language models or general multimodal reasoning, a dedicated synthesis of the rapidly expanding compositional visual reasoning literature is still missing. We fill this gap with a comprehensive survey spanning 2023 to 2025 that systematically reviews 260+ papers from top venues (CVPR, ICCV, NeurIPS, ICML, ACL, etc.). We first formalize core definitions and describe why compositional approaches offer advantages in cognitive alignment, semantic fidelity, robustness, interpretability, and data efficiency. Next, we trace a five-stage paradigm shift: from prompt-enhanced language-centric pipelines, through tool-enhanced LLMs and tool-enhanced VLMs, to recently minted chain-of-thought reasoning and unified agentic VLMs, highlighting their architectural designs, strengths, and limitations. We then catalog 60+ benchmarks and corresponding metrics that probe compositional visual reasoning along dimensions such as grounding accuracy, chain-of-thought faithfulness, and high-resolution perception. Drawing on these analyses, we distill key insights, identify open challenges (e.g., limitations of LLM-based reasoning, hallucination, a bias toward deductive reasoning, scalable supervision, tool integration, and benchmark limitations), and outline future directions, including world-model integration, human-AI collaborative reasoning, and richer evaluation protocols. By offering a unified taxonomy, historical roadmap, and critical outlook, this survey aims to serve as a foundational reference and inspire the next generation of compositional visual reasoning research.

Figures

Figures reproduced from arXiv: 2508.17298 by the authors.

Figure 1
Figure 1. Illustration of visual reasoning tasks, highlighting the differences between monolithic and compositional approaches. Monolithic [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Key shift from monolithic reasoning to compositional reasoning. [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Roadmap of compositional visual reasoning models. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Overview of prompt-enhanced language-centric (Stage I) pipeline. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Tool-enhanced LLMs (Stage II) and VLMs (Stage III) pipeline. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Monolithic versus Chain-of-Thought (CoT) reasoning VLMs (Stage IV) pipeline. [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Illustration of unified agentic VLM (Stage V) inference. The model iteratively refines its understanding by appending intermediate [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Examples of benchmarks and datasets Diagnostic variants further introduce linguistic perturbations, attention supervision, or adversarial rephrasings to evaluate robustness and interpretability [40, 192]. Representative datasets include CLEVR [5], SHAPES [193], TaskMeA…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Execution Commitment of Vision-Language-Action Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    A3 reframes dynamic action chunk commitment in VLA models as self-speculative prefix verification, accepting the longest continuous sequence of actions that satisfies consensus-ordered conditional invariance and prefi...

  2. Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

    cs.AI 2026-04 unverdicted novelty 7.0 of 10

    Rule-VLN is the first large-scale benchmark injecting 177 regulatory categories into an urban environment, and the proposed SNRM module equips pre-trained VLN agents with zero-shot semantic reasoning and detour planni...

  3. Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training framework that adds a removable image-generation branch to multimodal LLMs improves visual understanding benchmarks with zero inference-time cost.

  4. VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

    cs.CV 2025-11 conditional novelty 5.0 of 10

    Fine-tuning Qwen2.5-VL on VisReason, a 489K-example multi-round visual chain-of-thought dataset (165K with pseudo-depth), modestly improves LLM-judged visual reasoning scores, with caveats about self-referential 3D ev...

Reference graph

Works this paper leans on

266 extracted references · 19 canonical work pages · cited by 4 Pith papers

  1. [78]

    Chain-of-Focus: Adaptive Visual Search and Zooming for Multimodal Reasoning via RL

    Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, et al. Chain-of-Focus: Adaptive Visual Search and Zooming for Multimodal Reasoning via RL. arXiv preprint arXiv:2505.15436,

  2. [146]

    Multimodal Foundation Models: From Specialists to General-Purpose Assistants

    Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, Jianfeng Gao, et al. Multimodal Foundation Models: From Specialists to General-Purpose Assistants. Foundations and Trends® in Computer Graphics and Vision, 16(1-2):1– 214, 2024. 12, 13

  3. [29]

    HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

    Fucai Ke, Zhixi Cai, Simindokht Jahangard, Weiqing Wang, Pari Delir Haghighi, and Hamid Rezatofighi. HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning. In European Conference on Computer Vision, pages 132–149. Springer, 2024. 3, 6, 7, 11, 21

  4. [76]

    DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

    Fucai Ke, Xingjian Leng, Zhixi Cai, Zaid Khan, Weiqing Wang, Pari Delir Haghighi, Hamid Rezatofighi, Manmohan Chandraker, et al. DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning. arXiv preprint arXiv:2503.19263, 2025. 6, 11, 16, 20, 21

  5. [137]

    NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

    Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Reza Haffari, Peter J Stuckey, and Hamid Rezatofighi. NA VER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning. arXiv preprint arXiv:2502.00372, 2025. 11, 18

  6. [1]

    A Benchmark for Compositional Visual Reasoning

    Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre. A Benchmark for Compositional Visual Reasoning. Advances in neural information processing systems, 35:29776–29788, 2022. 2, 5, 6, 7

  7. [2]

    Take A Step Back: Rethinking the Two Stages in Visual Reasoning

    Mingyu Zhang, Jiting Cai, Mingyu Liu, Yue Xu, Cewu Lu, and Yong-Lu Li. Take A Step Back: Rethinking the Two Stages in Visual Reasoning. In European Conference on Computer Vision, pages 124–141. Springer, 2024. 2

  8. [3]

    The Book of Why: The New Science of Cause and Effect

    M ´elanie Frappier. The Book of Why: The New Science of Cause and Effect. Science, 361:855 – 855, 2018

Show all 266 references
  1. [4]

    Inferring and Executing Programs for Visual Reasoning

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Inferring and Executing Programs for Visual Reasoning. In Proceedings of the IEEE international conference on computer vision, pages 2989–2998, 2017. 3, 6, 7

  2. [5]

    CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...

  3. [6]

    RA VEN: A Dataset for Relational and Analogical Visual REasoNing

    Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. RA VEN: A Dataset for Relational and Analogical Visual REasoNing. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5312–5322, 2019. 2, 17

  4. [7]

    Jakob Suchan, Mehul Bhatt, Przemyslaw Andrzej Walega, and Carl P. L. Schultz. Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects. ArXiv, abs/1712.00840, 2017

  5. [8]

    Maintaining Reasoning Consistency in Compositional Visual Question Answering

    Chenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu, and Qi Wu. Maintaining Reasoning Consistency in Compositional Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5099–5108, 2022. 2, 6

  6. [9]

    Visual Instruction Tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. Advances in neural information processing systems, 36:34892–34916, 2023. 2, 3, 5, 6, 7, 15

  7. [10]

    Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.arXiv preprint arXiv:2308.12966,

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.arXiv preprint arXiv:2308.12966,

  8. [11]

    BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In International conference on machine learning, pages 19730–19742. PMLR, 2023. 6, 7, 10

  9. [12]

    InstructBLIP: towards general-purpose vision-language models with instruction tuning

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: towards general-purpose vision-language models with instruction tuning. In Proceedings of the 37th International Conference on Neural ...

  10. [13]

    A Survey of Multimodel Large Language Models

    Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, and Ke Liu. A Survey of Multimodel Large Language Models. In Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering , pages 405–409, 2024. 2, 3

  11. [14]

    Interpretable Visual Reasoning: A Survey

    Feijuan He, Yaxian Wang, Xianglin Miao, and Xia Sun. Interpretable Visual Reasoning: A Survey. Image and Vision Computing, 112:104194, 2021. 2, 5

  12. [15]

    Thinking Machines: A Survey of LLM based Reasoning Strategies

    Dibyanayan Bandyopadhyay, Soham Bhattacharjee, and Asif Ekbal. Thinking Machines: A Survey of LLM based Reasoning Strategies. arXiv preprint arXiv:2503.10814, 2025

  13. [16]

    Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

    Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wanxiang Che. Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models. arXiv preprint arXiv:2503.09567, 2025. 2

  14. [17]

    Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba

    Yunsheng Ma, Wenqian Ye, Can Cui, Haiming Zhang, Shuo Xing, Fucai Ke, Jinhong Wang, Chenglin Miao, Jintai Chen, Hamid Rezatofighi, et al. Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba. In Proceedings of ...

  15. [18]

    NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UA V Search Missions.arXiv preprint arXiv:2409.10196, 2024

    Zhixi Cai, Cristian Rojas Cardenas, Kevin Leo, Chenyuan Zhang, Kal Backman, Hanbing Li, Boying Li, Mahsa Ghorbanali, Stavya Datta, Lizhen Qu, et al. NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UA V Search Missi...

  16. [19]

    UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning

    Jorge Fernandez Galarreta, Norman Kerle, and Markus Gerke. UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning. Natural hazards and earth system sciences, 15(6):1087–1101, 2015

  17. [20]

    A Study of Visual Reasoning in Medical Diagnosis

    Erika Rogers. A Study of Visual Reasoning in Medical Diagnosis. In Proceedings of the Eighteenth Annual Conference of the Cognitive Science Society, pages 213–218. Routledge, 2019

  18. [21]

    Medical Visual Question Answering via Conditional Reasoning

    Li-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen, and Xiao-Ming Wu. Medical Visual Question Answering via Conditional Reasoning. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2345–2354, 2020

  19. [22]

    Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

    Renhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao, Jiajun Wu, and Yang Gao. Programmatically Grounded, Compositionally Generalizable Robotic Manipulation. arXiv preprint arXiv:2304.13826, 2023. 3 23

  20. [23]

    Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

    Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...

  21. [24]

    When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,

    Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,

  22. [25]

    Challenges and Prospects in Vision and Language Research

    Kushal Kafle, Robik Shrestha, and Christopher Kanan. Challenges and Prospects in Vision and Language Research. Frontiers in Artificial Intelligence, 2, 2019

  23. [26]

    Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

    Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, and Hongxia Yang. Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasonin...

  24. [27]

    Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

    Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng, Shilong Liu, Liang-Yan Gui, Jan Kautz, Yu-Xiong Wang, and Zhiding Yu. Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 14268–14280, 2025. 15

  25. [28]

    Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

    Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 10632–10643, 2025. 3

  26. [30]

    Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

    Aleksandar Stani ´c, Sergi Caelles, and Michael Tschannen. Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers. Transactions on Machine Learning Research, 2024. 6, 7

  27. [31]

    Multimodal Chain-of- Thought Reasoning: A Comprehensive Survey

    Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. Multimodal Chain-of- Thought Reasoning: A Comprehensive Survey. arXiv preprint arXiv:2503.12605, 2025. 3, 4, 6

  28. [32]

    Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining

    Ugur Sahin, Hang Li, Qadeer Khan, Daniel Cremers, and V olker Tresp. Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5563–5573, 2024

  29. [33]

    Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

    Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muen- nighoff, Kyle Lo, Luca Soldaini, et al. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models. In Proceedings of the Compute...

  30. [34]

    Cascaded Mutual Modulation for Visual Reasoning

    Yiqun Yao, Jiaming Xu, Feng Wang, and Bo Xu. Cascaded Mutual Modulation for Visual Reasoning. In Conference on Empirical Methods in Natural Language Processing, 2018. 3

  31. [35]

    Compositional Chain-of-Thought Prompting for Large Mul- timodal Models

    Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional Chain-of-Thought Prompting for Large Mul- timodal Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14431,

  32. [36]

    Building Machines That Learn and Think Like People

    Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building Machines That Learn and Think Like People. Behavioral and brain sciences, 40:e253, 2017. 3, 6

  33. [37]

    Meta Module Network for Compositional Visual Reasoning

    Wenhu Chen, Zhe Gan, Linjie Li, Yu Cheng, William Wang, and Jingjing Liu. Meta Module Network for Compositional Visual Reasoning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 655–664, 2021. 5, 6

  34. [38]

    Improving Visual Reasoning Through Semantic Representation

    Wenfeng Zheng, Xiangjun Liu, Xubin Ni, Lirong Yin, and Bo Yang. Improving Visual Reasoning Through Semantic Representation. IEEE access, 9:91476–91486, 2021

  35. [39]

    What’s Left? Concept Grounding with Logic-Enhanced Foundation Models

    Joy Hsu, Jiayuan Mao, Josh Tenenbaum, and Jiajun Wu. What’s Left? Concept Grounding with Logic-Enhanced Foundation Models. Advances in Neural Information Processing Systems, 36:38798–38814, 2023. 3, 11

  36. [40]

    Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges

    Byeong Su Kim, Jieun Kim, Deokwoo Lee, and Beakcheol Jang. Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges. ACM Computing Surveys, 57(10):1–35, 2025. 3, 4, 17, 18, 21

  37. [41]

    Visual Question Answering: a State-of-the-Art Review

    Sruthy Manmadhan and Binsu C Kovoor. Visual Question Answering: a State-of-the-Art Review. Artificial Intelligence Review, 53(8):5705–5745, 2020

  38. [42]

    Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    Jiayi Kuang, Ying Shen, Jingyou Xie, Haohao Luo, Zhe Xu, Ronghao Li, Yinghui Li, Xianfeng Cheng, Xika Lin, and Yu Han. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey. ACM Computing Surveys, 57(8):1–36, 2025. 3, 4, 5, 20

  39. [43]

    Mind with Eyes: from Language Reasoning to Multimodal Reasoning

    Zhiyu Lin, Yifei Gao, Xian Zhao, Yunfan Yang, and Jitao Sang. Mind with Eyes: from Language Reasoning to Multimodal Reasoning. arXiv preprint arXiv:2503.18071, 2025. 3, 4, 6

  40. [44]

    A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge

    M Jaleed Khan, Filip Ilievski, John G Breslin, and Edward Curry. A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge. Neurosymbolic Artificial Intelligence, 1:NAI–240719, 2025. 3, 4, 6, 20

  41. [45]

    Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices

    Mikołaj Małki ´nski and Jacek Ma´ndziuk. Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices. ACM Computing Surveys, 57(7):1–36, 2025. 3, 4, 5 24

  42. [46]

    Large Multimodal Agents: A Survey

    Junlin Xie, Zhihong Chen, Ruifei Zhang, Xiang Wan, and Guanbin Li. Large Multimodal Agents: A Survey. arXiv preprint arXiv:2402.15116, 2024. 3, 4

  43. [47]

    From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities

    Md Farhan Ishmam, Md Sakib Hossain Shovon, Muhammad Firoz Mridha, and Nilanjan Dey. From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities. Information Fusion, 106:102270, 2024. 4, 5, 7

  44. [48]

    Robust Visual Question Answering: Datasets, Methods, and Future Challenges

    Jie Ma, Pinghui Wang, Dechen Kong, Zewei Wang, Jun Liu, Hongbin Pei, and Junzhou Zhao. Robust Visual Question Answering: Datasets, Methods, and Future Challenges. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 4, 6, 21

  45. [49]

    Perception, reason, think, and plan: A survey on large multimodal reasoning models.arXiv preprint arXiv:2505.04921,

    Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. Perception, reason, think, and plan: A survey on large multimodal reasoning models.arXiv preprint arXiv:2505.04921,

  46. [50]

    How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model

    Shezheng Song, Xiaopeng Li, Shasha Li, Shan Zhao, Jie Yu, Jun Ma, Xiaoguang Mao, Weimin Zhang, and Meng Wang. How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model. IEEE Transactions on Knowledge and Data Engineering, 2025. 4

  47. [51]

    Tool learning with large language models: a survey

    Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: a survey. Frontiers of Computer Science, 19(8):198343, 2025. 4

  48. [52]

    Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

    Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9199...

  49. [53]

    Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

    Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. Multimodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing, 14(3):478–493, 2020. 5

  50. [54]

    Transformation Driven Visual Reasoning

    Xin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo, and Xueqi Cheng. Transformation Driven Visual Reasoning. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 6903–6912, 2021. 5

  51. [55]

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 5

  52. [56]

    Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction

    Wentao He, Jianfeng Ren, Ruibin Bai, and Xudong Jiang. Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction. Pattern Recognition, 160:111151, 2025. 5

  53. [57]

    Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning

    Wentao He, Jialu Zhang, Jianfeng Ren, Ruibin Bai, and Xudong Jiang. Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 22–30, 2023. 5

  54. [58]

    You Only Look Once: Unified, Real-Time Object Detection

    Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016. 5

  55. [59]

    Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering

    Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering. In Proceedings of the IEEE international conference on computer vision, pages 1821–1830, 2017. 5, 6

  56. [60]

    Bilinear Attention Networks

    Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear Attention Networks. Advances in neural information processing systems, 31, 2018. 6

  57. [61]

    Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering

    Duy-Kien Nguyen and Takayuki Okatani. Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 6087–6096, 2018. 5

  58. [62]

    Deep Modular Co-Attention Networks for Visual Question Answering

    Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep Modular Co-Attention Networks for Visual Question Answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281–6290, 2019. 6

  59. [63]

    ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Advances in neural information processing systems, 32, 2019. 6

  60. [64]

    Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Supervision. In International conference on machine learning, pa...

  61. [65]

    UNITER: UNiversal Image-TExt Representation Learning

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: UNiversal Image-TExt Representation Learning. In European conference on computer vision, pages 104–120. Springer, 2020. 6

  62. [66]

    Flamingo: a Visual Language Model for Few-Shot Learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a Visual Language Model for Few-Shot Learning. Advances in neural information processing systems, 35:23716...

  63. [67]

    MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In The Twelfth International Conference on Learning Representations, 2024. 6

  64. [68]

    Otter: A Multi-Modal Model With In-Context Instruction Tuning

    Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Otter: A Multi-Modal Model With In-Context Instruction Tuning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 6, 7 25

  65. [69]

    CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023

    Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna. CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023. 6, 16

  66. [70]

    Logics and Languages

    Max J Cresswell. Logics and Languages. Routledge, 2016. 6

  67. [71]

    From Recognition to Cognition: Visual Commonsense Reasoning

    Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. From Recognition to Cognition: Visual Commonsense Reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731, 2019. 6, 7, 17

  68. [72]

    Visual Programming: Compositional Visual Reasoning without Training

    Tanmay Gupta and Aniruddha Kembhavi. Visual Programming: Compositional Visual Reasoning without Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962, 2023. 7, 11

  69. [73]

    GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

    Drew A Hudson and Christopher D Manning. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 6700–6709, 2019. 7, 14, 16, 18

  70. [74]

    Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

    Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. Advances in Neural Information Processing Systems , 36,

  71. [75]

    Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions

    Huaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu, Yuke Zhu, and Anima Anandkumar. Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19056–19065, 2022. 6, 7

  72. [77]

    Iterated Learning Improves Compositionality in Large Vision-Language Models

    Chenhao Zheng, Jieyu Zhang, Aniruddha Kembhavi, and Ranjay Krishna. Iterated Learning Improves Compositionality in Large Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13785–13795, June 2024. 6

  73. [79]

    Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

    Di Zhang, Jingdi Lei, Junxian Li, Xunzhi Wang, Yujie Liu, Zonglin Yang, Jiatong Li, Weida Wang, Suorong Yang, Jianbo Wu, et al. Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages ...

  74. [80]

    Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, and David Acuna. Can Large Vision-Language Models Correct Semantic Ground- ing Errors By Themselves? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 14667–14678, 2025. 6

  75. [81]

    Visual Compositional Learning for Human-Object Interaction Detection

    Zhi Hou, Xiaojiang Peng, Yu Qiao, and Dacheng Tao. Visual Compositional Learning for Human-Object Interaction Detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16, pages 584–600. Springer, 2020. 6

  76. [82]

    Learn- ing Visual Composition through Improved Semantic Guidance

    Austin Stone, Hagen Soltau, Robert Geirhos, Xi Yi, Ye Xia, Bingyi Cao, Kaifeng Chen, Abhijit Ogale, and Jonathon Shlens. Learn- ing Visual Composition through Improved Semantic Guidance. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 3740–3750, 2025

  77. [83]

    Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage

    Andy Zeng, Maria Attarian, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, et al. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage. In The Eleventh International Co...

  78. [84]

    Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models

    Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi, and Ranjay Krishna. Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models. In IEEE/CVF Conference ...

  79. [85]

    Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

    Jiaxing Chen, Yuxuan Liu, Dehu Li, Xiang An, Weimo Deng, Ziyong Feng, Yongle Zhao, and Yin Xie. Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models. arXiv preprint arXiv:2403.19322, 2024. 7, 12, 13, 20, 21

  80. [86]

    DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage

    Joy Hsu, Jiayuan Mao, and Jiajun Wu. DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage. Transactions on Machine Learning Research, 2023. 7

  81. [87]

    Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning

    Shi Chen and Qi Zhao. Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6736–6745, 2023. 7

  82. [88]

    VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

    Zejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang, and Zhongyu Wei. VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computation...

  83. [89]

    Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement

    Zaid Khan, Vijay Kumar BG, Samuel Schulter, Yun Fu, and Manmohan Chandraker. Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14344–1...

  84. [90]

    From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

    Chuanqi Cheng, Jian Guan, Wei Wu, and Rui Yan. From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4941–4957, 2024. 12

  85. [91]

    Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

    Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu, Yan Ma, Xiaoye Qu, Jiaqi Liu, Yanshu Li, Kaide Zeng, Zhengyuan Yang, et al. Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers. arXiv preprint arXiv:2506.23918,

  86. [92]

    Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

    Qiji Zhou, Ruochen Zhou, Zike Hu, Panzhong Lu, Siyang Gao, and Yue Zhang. Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models. arXiv preprint arXiv:2405.13872, 2024. 7, 12, 18, 20

  87. [93]

    The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). arXiv preprint arXiv:2309.17421, 9(1):1, 2023

  88. [94]

    Grounded Chain-of- Thought for Multimodal Large Language Models

    Qiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang, Baiyang Song, Xiaoshuai Sun, and Rongrong Ji. Grounded Chain-of- Thought for Multimodal Large Language Models. arXiv preprint arXiv:2503.12799, 2025. 7, 14, 19, 20, 21

  89. [95]

    NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

    Baiqi Li, Zhiqiu Lin, Wenxuan Peng, Jean de Dieu Nyandwi, Daniel Jiang, Zixian Ma, Simran Khanuja, Ranjay Krishna, Graham Neubig, and Deva Ramanan. NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, ...

  90. [96]

    BLINK: Multimodal Large Language Models Can See but Not Perceive

    Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. BLINK: Multimodal Large Language Models Can See but Not Perceive. In European Conference on Computer Vision, pages 148–166. Springer, 2024. 7

  91. [97]

    ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

    Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora. ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty. Advances in Neural Information Processing Systems, 37:86004–86047, 2024. 7

  92. [98]

    Visual Reasoning by Progressive Module Networks

    Seung Wook Kim, Makarand Tapaswi, and Sanja Fidler. Visual Reasoning by Progressive Module Networks. In International Conference on Learning Representations, 2018. 7

  93. [99]

    Predicate Hierarchies Improve Few-Shot State Classification

    Emily Jin, Joy Hsu, and Jiajun Wu. Predicate Hierarchies Improve Few-Shot State Classification. arXiv preprint arXiv:2502.12481,

  94. [100]

    Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

    Yushi Hu, Otilia Stretcu, Chun-Ta Lu, Krishnamurthy Viswanathan, Kenji Hata, Enming Luo, Ranjay Krishna, and Ariel Fuxman. Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Compute...

  95. [101]

    Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

    Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng, Fan Yang, Ming Tang, and Jinqiao Wang. Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models. arXiv preprint arXiv:2505.20753, 2025. 7, 14, 16, 19, 20

  96. [102]

    Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv preprint arXiv:2409.12191,

  97. [103]

    Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

    Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4818–...

  98. [104]

    LLaV A-NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. LLaV A-NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024. 15

  99. [105]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024

  100. [106]

    DeepSeek-VL: Towards Real-World Vision-Language Understanding.arXiv preprint arXiv:2403.05525, 2024

    Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, et al. DeepSeek-VL: Towards Real-World Vision-Language Understanding.arXiv preprint arXiv:2403.05525, 2024

  101. [107]

    GPT-4 Technical Report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774, 2023

  102. [108]

    mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

    Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality. arXiv preprint arXiv:2304.14178,

  103. [109]

    SpatialVLM: Endowing Vision- Language Models with Spatial Reasoning Capabilities

    Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. SpatialVLM: Endowing Vision- Language Models with Spatial Reasoning Capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14455–144...

  104. [110]

    Visual Spatial Reasoning

    Fangyu Liu, Guy Emerson, and Nigel Collier. Visual Spatial Reasoning. Transactions of the Association for Computational Linguistics, 11:635–651, 2023

  105. [111]

    CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations.arXiv preprint arXiv:2402.04236, 2024

    Ji Qi, Ming Ding, Weihan Wang, Yushi Bai, Qingsong Lv, Wenyi Hong, Bin Xu, Lei Hou, Juanzi Li, Yuxiao Dong, et al. CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations.arXiv preprint arXiv:2402.04236, 2024. 8, 15, 16, 19, 20, 21 27

  106. [112]

    Program Synthesis with Large Language Models

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program Synthesis with Large Language Models. arXiv preprint arXiv:2108.07732, 2021. 8

  107. [113]

    Large Language Models are Zero-Shot Reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large Language Models are Zero-Shot Reasoners. Advances in neural information processing systems, 35:22199–22213, 2022

  108. [114]

    Training Verifiers to Solve Math Word Problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021. 8

  109. [115]

    DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multi- modal Reasoning in Language Models

    Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multi- modal Reasoning in Language Models. Advances in Neural Information Processing Systems, 36:5168–5191, 2023. 10

  110. [116]

    ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

    Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny. ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions. Transactions on Machine Learning Research, 2024. 10

  111. [117]

    Ide- alGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

    Haoxuan You, Rui Sun, Zhecan Wang, Long Chen, Gengyu Wang, Hammad Ayyubi, Kai-Wei Chang, and Shih-Fu Chang. Ide- alGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pa...

  112. [118]

    Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use

    Imad Eddine Toubal, Aditya Avinash, Neil Gordon Alldrin, Jan Dlabal, Wenlei Zhou, Enming Luo, Otilia Stretcu, Hao Xiong, Chun-Ta Lu, Howard Zhou, et al. Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use. InProceedings o...

  113. [119]

    Large Language Models are Visual Reasoning Coordinators

    Liangyu Chen, Bo Li, Sheng Shen, Jingkang Yang, Chunyuan Li, Kurt Keutzer, Trevor Darrell, and Ziwei Liu. Large Language Models are Visual Reasoning Coordinators. Advances in Neural Information Processing Systems, 36:70115–70140, 2023. 10

  114. [120]

    PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3

    Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo. PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2963–2975, 2023. 10

  115. [121]

    Multimodal Chain-of-Thought Reasoning in Language Models

    Zhuosheng Zhang, Aston Zhang, Mu Li, hai zhao, George Karypis, and Alex Smola. Multimodal Chain-of-Thought Reasoning in Language Models. Transactions on Machine Learning Research, 2024. 10

  116. [122]

    AdGPT: Explore Meaningful Advertising with ChatGPT

    Jiannan Huang, Mengxue Qu, Longfei Li, and Yunchao Wei. AdGPT: Explore Meaningful Advertising with ChatGPT. ACM Trans. Multimedia Comput. Commun. Appl., 21(4), April 2025. 10

  117. [123]

    Visually Descriptive Language Model for Vector Graphics Reasoning

    Zhenhailong Wang, Joy Hsu, Xingyao Wang, Kuan-Hao Huang, Manling Li, Jiajun Wu, and Heng Ji. Visually Descriptive Language Model for Vector Graphics Reasoning. arXiv preprint arXiv:2404.06479, 2024. 10

  118. [124]

    Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

    Hulingxiao He, Geng Li, Zijun Geng, Jinglin Xu, and Yuxin Peng. Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models. In The Thirteenth International Conference on Learning Representations ,

  119. [125]

    Visual Chain-of- Thought Prompting for Knowledge-Based Visual Reasoning

    Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Zhiqing Sun, Dan Gutfreund, and Chuang Gan. Visual Chain-of- Thought Prompting for Knowledge-Based Visual Reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, pages 1254–1262, 2024. 10, 19

  120. [126]

    VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things

    Yaoyao Zhong, Mengshi Qi, Rui Wang, Yuhan Qiu, Yang Zhang, and Huadong Ma. VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 39, pages 10680–10688, 2025. 11

  121. [127]

    GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

    Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction. Advances in Neural Information Processing Systems, 36:71995–72007, 2023. 11, 13

  122. [128]

    Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

    Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models. arXiv preprint arXiv:2303.04671, 2023. 11

  123. [129]

    ViperGPT: Visual Inference via Python Execution for Reasoning.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11854–11864, 2023

    D’idac Sur’is, Sachit Menon, and Carl V ondrick. ViperGPT: Visual Inference via Python Execution for Reasoning.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11854–11864, 2023. 11

  124. [130]

    HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. Advances in Neural Information Processing Systems, 36:38154–38180, 2023. 11

  125. [131]

    CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets

    Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung, Hao Peng, and Heng Ji. CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets. In The Twelfth International Conference on Learning Representations, 2024. 11

  126. [132]

    MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

    Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action. arXiv preprint arXiv:2303.11381, 2023. 11, 21

  127. [133]

    ContextualCoder: Adaptive In-context Prompting for Programmatic Visual Question Answering

    Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex C Kot, and Koichi Shinoda. ContextualCoder: Adaptive In-context Prompting for Programmatic Visual Question Answering. IEEE Transactions on Multimedia, 2025. 11

  128. [134]

    InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

    Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang, Yi Wang, Shoufa Chen, Qinglong Zhang, Zeqiang Lai, Yang Yang, Qingyun Li, et al. InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language. arXiv preprint arXiv:2305.05662, 2023. 11

  129. [135]

    Can Visual Scratchpads With Diagrammatic Abstractions Augment LLM Reasoning? In Proceedings on, pages 21–28

    Joy Hsu, Gabriel Poesia, Jiajun Wu, and Noah Goodman. Can Visual Scratchpads With Diagrammatic Abstractions Augment LLM Reasoning? In Proceedings on, pages 21–28. PMLR, 2023. 11 28

  130. [136]

    CLOV A: A Closed-LOop Visual Assistant with Tool Usage and Update

    Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wenjuan Han, Song-Chun Zhu, and Qing Li. CLOV A: A Closed-LOop Visual Assistant with Tool Usage and Update. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13258–13268, 2024. 11, 16, 21

  131. [138]

    SYNAPSE: SYmbolic Neural-Aided Preference Synthesis Engine

    Sadanand Modak, Noah Tobias Patton, Isil Dillig, and Joydeep Biswas. SYNAPSE: SYmbolic Neural-Aided Preference Synthesis Engine. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 27529–27537, 2025. 11

  132. [139]

    ViUniT: Visual Unit Tests for More Robust Visual Programming

    Artemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, and Juan Carlos Niebles. ViUniT: Visual Unit Tests for More Robust Visual Programming. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24646–2...

  133. [140]

    LLaV A-Plus: Learning to Use Tools for Creating Multimodal Agents

    Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. LLaV A-Plus: Learning to Use Tools for Creating Multimodal Agents. In European Conference on Computer Vision, pages 126–

  134. [141]

    Synthesize Step-by-Step: Tools Templates and LLMs as Data Gener- ators for Reasoning-Based Chart VQA

    Zhuowan Li, Bhavan Jasani, Peng Tang, and Shabnam Ghadar. Synthesize Step-by-Step: Tools Templates and LLMs as Data Gener- ators for Reasoning-Based Chart VQA. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13613–13623, 2024. 12

  135. [142]

    12, 16, 21

    Springer, 2024. 12, 16, 21

  136. [143]

    OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

    Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Juntao Li, Xiaoye Qu, et al. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning. arXiv preprint arXiv:2505.08617, 2025. 12

  137. [144]

    VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

    Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, Chengxiang Zhai, and Klara Nahrstedt. VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use. arXiv preprint arXiv:2505.19255, 2025. 12, 13

  138. [145]

    Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

    Hao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua, and Shuicheng Yan. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing. arXiv preprint arXiv:2412.19806, 2024. 12, 21

  139. [147]

    Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

    Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models. Advances in Neural Information Processing Systems, 37:139348–139379, 2024. 13

  140. [148]

    Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination

    Syeda Nahida Akter, Aman Madaan, Sangwu Lee, Yiming Yang, and Eric Nyberg. Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination. ArXiv, abs/2401.08025, 2024. 13

  141. [149]

    LATTE: Learning to Reason with Vision Specialists

    Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, et al. LATTE: Learning to Reason with Vision Specialists. In The 2025 Conference on Empirical Methods in Natural Language Processing, 2025. 13

  142. [150]

    CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

    Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al. CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models. In Proceedings of the Computer Vision and Pattern Recognition Confere...

  143. [151]

    Creative Agents: Empowering Agents with Imagination for Creative Tasks

    Penglin Cai, Chi Zhang, Yuhui Fu, Haoqi Yuan, and Zongqing Lu. Creative Agents: Empowering Agents with Imagination for Creative Tasks. In The 41st Conference on Uncertainty in Artificial Intelligence, 2025. 13

  144. [152]

    OpenAI o1 System Card

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. OpenAI o1 System Card. arXiv preprint arXiv:2412.16720, 2024. 13, 14

  145. [153]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948,

  146. [154]

    LLaV A-CoT: Let Vision Language Models Reason Step-by- Step, 2025

    Guowei Xu, Peng Jin, Hao Li, Yibing Song, Lichao Sun, and Li Yuan. LLaV A-CoT: Let Vision Language Models Reason Step-by- Step, 2025. 13

  147. [155]

    Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning

    Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Minglan Lin, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang. Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning. arXiv preprint arXiv:2503.20752, 2025. 14, 21

  148. [156]

    Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

    Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models. arXiv preprint arXiv:2503.06749, 2025. 14, 19, 21

  149. [157]

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300,

  150. [158]

    DeepPer- ception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding

    Xinyu Ma, Ziyang Ding, Zhicong Luo, Chi Chen, Zonghao Guo, Derek F Wong, Xiaoyi Feng, and Maosong Sun. DeepPer- ception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding. arXiv preprint arXiv:2503.12797, 2025. 14 29

  151. [159]

    G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

    Liang Chen, Hongcheng Gao, Tianyu Liu, Zhiqi Huang, Flood Sung, Xinyu Zhou, Yuxin Wu, and Baobao Chang. G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning. arXiv preprint arXiv:2505.13426 ,

  152. [160]

    Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning

    Meng Cao, Haoze Zhao, Can Zhang, Xiaojun Chang, Ian Reid, and Xiaodan Liang. Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning. arXiv preprint arXiv:2505.20272, 2025. 14, 16, 19, 20

  153. [161]

    On a Connection Between Imitation Learning and RLHF

    Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, and Vasant G Honavar. On a Connection Between Imitation Learning and RLHF. arXiv preprint arXiv:2503.05079, 2025. 14

  154. [162]

    Beyond Reverse KL: Generalizing Direct Preference Opti- mization with Diverse Divergence Constraints

    Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen. Beyond Reverse KL: Generalizing Direct Preference Opti- mization with Diverse Divergence Constraints. In The Twelfth International Conference on Learning Representations, 2024

  155. [163]

    SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

    Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V Le, Sergey Levine, and Yi Ma. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training. arXiv preprint arXiv:2501.17161 ,

  156. [164]

    OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , pages 3195–3204, 2019. 14, 17, 18

  157. [165]

    A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge. In European conference on computer vision, pages 146–162. Springer, 2022. 14, 17, 18

  158. [166]

    Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

    Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G Shapiro, and Ranjay Krishna. Perception Tokens Enhance Visual Reasoning in Multimodal Language Models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 3836–3845, 2025. 14

  159. [167]

    Visually Interpretable Subtask Reasoning for Visual Question Answering

    Yu Cheng, Arushi Goel, and Hakan Bilen. Visually Interpretable Subtask Reasoning for Visual Question Answering. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 2760–2780, 2025. 14

  160. [168]

    CoReS: Orchestrating the Dance of Reasoning and Segmentation

    Xiaoyi Bao, Siyang Sun, Shuailei Ma, Kecheng Zheng, Yuxin Guo, Guosheng Zhao, Yun Zheng, and Xingang Wang. CoReS: Orchestrating the Dance of Reasoning and Segmentation. In European Conference on Computer Vision, pages 187–204. Springer,

  161. [169]

    Visual Agents as Fast and Slow Thinkers

    Guangyan Sun, Mingyu Jin, Zhenting Wang, Cheng-Long Wang, Siqi Ma, Qifan Wang, Tong Geng, Ying Nian Wu, Yongfeng Zhang, and Dongfang Liu. Visual Agents as Fast and Slow Thinkers. In The Thirteenth International Conference on Learning Representations, 2018. 15, 16, 19, 20, 21

  162. [170]

    V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs

    Penghao Wu and Saining Xie. V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094, 2024. 15, 16, 17, 18, 19, 21

  163. [171]

    Divide, Conquer and Com- bine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

    Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, Wei Yu, and Dacheng Tao. Divide, Conquer and Com- bine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models. In Proceedings of the AAAI Conference on Artificial...

  164. [172]

    Segment Anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexan- der C Berg, Wan-Yen Lo, et al. Segment Anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023. 15

  165. [173]

    Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection. In European Conference on Computer Vision, pages 38–55. Springer...

  166. [174]

    ZoomEye: En- hancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

    Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, and Jianwei Yin. ZoomEye: En- hancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. arXiv preprint arXiv:2411.16044, 2024. 15, 16, 19

  167. [175]

    Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

    Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning. Advances in Neural Information Processing Systems, ...

  168. [176]

    Insight-V: Exploring Long- Chain Visual Reasoning with Multimodal Large Language Models

    Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. Insight-V: Exploring Long- Chain Visual Reasoning with Multimodal Large Language Models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 9062–9072, 2025. 15

  169. [177]

    DeepEyes: Incentivizing ”Thinking with Images” via Reinforcement Learning

    Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. DeepEyes: Incentivizing ”Thinking with Images” via Reinforcement Learning. ArXiv, abs/2505.14362, 2025. 15

  170. [178]

    GeReA: Question-Aware Prompt Captions for Knowledge- based Visual Question Answering

    Ziyu Ma, Shutao Li, Bin Sun, Jianfei Cai, Zuxiang Long, and Fuyan Ma. GeReA: Question-Aware Prompt Captions for Knowledge- based Visual Question Answering. arXiv preprint arXiv:2402.02503, 2024. 15

  171. [179]

    Machine Mental Imagery: Empower Multimodal Reason- ing with Latent Visual Tokens

    Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, and Chuang Gan. Machine Mental Imagery: Empower Multimodal Reason- ing with Latent Visual Tokens. ArXiv, abs/2506.17218, 2025. 16

  172. [180]

    From Foresight to Forethought: VLM-in-the-loop policy steering via latent alignment

    Yilin Wu, Ran Tian, Gokul Swamy, and Andrea Bajcsy. From Foresight to Forethought: VLM-in-the-loop policy steering via latent alignment. In ICLR 2025 Workshop on World Models: Understanding, Modelling and Scaling, 2025. 16 30

  173. [181]

    LISA: Reasoning Segmentation via Large Language Model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmentation via Large Language Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 9579–9589,

  174. [182]

    MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

    Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Sca...

  175. [183]

    VQA: Visual Question Answering

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. VQA: Visual Question Answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015. 16, 21

  176. [184]

    Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 16

  177. [185]

    Microsoft COCO: Common Objects in Context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedin...

  178. [186]

    Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International journal of computer ...

  179. [187]

    Visual7W: Grounded Question Answering in Images

    Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7W: Grounded Question Answering in Images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4995–5004, 2016. 16

  180. [188]

    Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

    Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 4971–4980, 2018. 16

  181. [189]

    TallyQA: Answering Complex Counting Questions

    Manoj Acharya, Kushal Kafle, and Christopher Kanan. TallyQA: Answering Complex Counting Questions. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8076–8084, 2019. 16

  182. [190]

    Cola: A Benchmark for Compositional Text-to-image Retrieval

    Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna, and Kate Saenko. Cola: A Benchmark for Compositional Text-to-image Retrieval. Advances in Neural Information Processing Systems, 36:46433–46445, 2023. 16

  183. [191]

    What Makes a Maze Look Like a Maze? In International Conference on Learning Representations (ICLR), 2025

    Joy Hsu, Jiayuan Mao, Joshua B Tenenbaum, Noah D Goodman, and Jiajun Wu. What Makes a Maze Look Like a Maze? In International Conference on Learning Representations (ICLR), 2025. 16

  184. [192]

    SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

    Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality. Advances in neural information processing systems, 36:31096–31116, 2023. 16, 21

  185. [193]

    A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

    Raihan Kabir, Naznin Haque, Md Saiful Islam, et al. A Comprehensive Survey on Visual Question Answering Datasets and Algorithms. arXiv preprint arXiv:2411.11150, 2024. 17

  186. [194]

    Neural Module Networks

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural Module Networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39–48, 2016. 17

  187. [195]

    Task Me Anything

    Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. Task Me Anything. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 17

  188. [196]

    Comparing Machines and Humans on a Visual Categorization Test

    Franc ¸ois Fleuret, Ting Li, Charles Dubout, Emma K Wampler, Steven Yantis, and Donald Geman. Comparing Machines and Humans on a Visual Categorization Test. Proceedings of the National Academy of Sciences, 108:17621 – 17625, 2011. 17

  189. [197]

    O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron C

    Dzmitry Bahdanau, Harm de Vries, Timothy J. O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron C. Courville. CLOSURE: Assessing Systematic Generalization of CLEVR Models. ArXiv, abs/1912.05783, 2019. 17

  190. [198]

    CURI: A Benchmark for Productive Concept Learning Under Uncertainty

    Ramakrishna Vedantam, Arthur Szlam, Maximillian Nickel, Ari Morcos, and Brenden M Lake. CURI: A Benchmark for Productive Concept Learning Under Uncertainty. In International Conference on Machine Learning, pages 10519–10529. PMLR, 2021. 17

  191. [199]

    CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions

    Runtao Liu, Chenxi Liu, Yutong Bai, and Alan L Yuille. CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4185–4194, 2019. 17

  192. [200]

    CLEVR-XAI: A Benchmark Dataset for the Ground Truth Evaluation of Neural Network Explanations

    Leila Arras, Ahmed Osman, and Wojciech Samek. CLEVR-XAI: A Benchmark Dataset for the Ground Truth Evaluation of Neural Network Explanations. Information Fusion, 81:14–40, 2022. 17

  193. [201]

    QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

    Zechen Li and Anders Søgaard. QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning. arXiv preprint arXiv:2205.03075, 2022. 17

  194. [202]

    Sophia Koepke, Hendrik P

    Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch, and Zeynep Akata. CLEVR-X: A Visual Reasoning Dataset for Natural Language Explanations. ArXiv, abs/2204.02380, 2022. 17

  195. [203]

    Satwik Kottur, Jos ´e M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. CLEVR-Dialog: A Diagnostic Dataset for Multi-Round Reasoning in Visual Dialog. In North American Chapter of the Association for Computational Linguistics, 2019. 17 31

  196. [204]

    Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

    Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. Multimodal Explanations: Justifying Decisions and Pointing to the Evidence. In Proceedings of the IEEE conference on computer vision and pattern recognition, pa...

  197. [205]

    Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions? Computer Vision and Image Understanding, 163:90–100, 2017

    Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions? Computer Vision and Image Understanding, 163:90–100, 2017. 17

  198. [206]

    Cycle-Consistency for Robust Visual Question Answering

    Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh. Cycle-Consistency for Robust Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6649–6658, 2019. 17

  199. [207]

    Adam Santoro, Felix Hill, David G. T. Barrett, Ari S. Morcos, and Timothy P. Lillicrap. Measuring Abstract Reasoning in Neural Networks. ArXiv, abs/1807.04225, 2018. 17

  200. [208]

    KANDINSKYPatterns–An experimental exploration environment for Pattern Analysis and Machine Intelligence

    Andreas Holzinger, Anna Saranti, and Heimo Mueller. KANDINSKYPatterns–An experimental exploration environment for Pattern Analysis and Machine Intelligence. arXiv preprint arXiv:2103.00519, 2021. 17

  201. [209]

    Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and Reasoning

    Weili Nie, Zhiding Yu, Lei Mao, Ankit B Patel, Yuke Zhu, and Anima Anandkumar. Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and Reasoning. Advances in Neural Information Processing Systems, 33:16468–16480, 2020. 17

  202. [210]

    FVQA: Fact-Based Visual Question Answering

    Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. FVQA: Fact-Based Visual Question Answering. IEEE transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017. 17

  203. [211]

    Explicit Knowledge-based Reasoning for Visual Question Answering

    Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. Explicit Knowledge-based Reasoning for Visual Question Answering. arXiv preprint arXiv:1511.02570, 2015. 17

  204. [212]

    Shan, and Xilin Chen

    Difei Gao, Ruiping Wang, S. Shan, and Xilin Chen. CRIC: A VQA Dataset for Compositional Reasoning on Vision and Common- sense. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:5561–5578, 2019. 17

  205. [213]

    Towards Visual Dialog for Radiology

    Olga Kovaleva, Chaitanya Shivade, Satyananda Kashyap, Karina Kanjaria, Joy Wu, Deddeh Ballah, Adam Coy, Alexandros Karar- gyris, Yufan Guo, David Beymer Beymer, et al. Towards Visual Dialog for Radiology. In Proceedings of the 19th SIGBioMed workshop on biomedical language pro...

  206. [214]

    Scene Text Visual Question Answering

    Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marc ¸al Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene Text Visual Question Answering. InProceedings of the IEEE/CVF international conference on computer vision, pages 4291– 4301, 2019. 17, 18

  207. [215]

    DocVQA: A Dataset for VQA on Document Images

    Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. DocVQA: A Dataset for VQA on Document Images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 17

  208. [216]

    A Diagram is Worth a Dozen Images

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A Diagram is Worth a Dozen Images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 235–251. ...

  209. [217]

    A Dataset of Clinically Generated Visual Questions and Answers About Radiology Images

    Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A Dataset of Clinically Generated Visual Questions and Answers About Radiology Images. Scientific data, 5(1):1–10, 2018. 18

  210. [218]

    PathVQA: 30000+ Questions for Medical Visual Question Answering

    Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. PathVQA: 30000+ Questions for Medical Visual Question Answering. arXiv preprint arXiv:2003.10286, 2020. 18

  211. [219]

    VizWiz Grand Challenge: Answering Visual Questions From Blind People

    Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. VizWiz Grand Challenge: Answering Visual Questions From Blind People. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–36...

  212. [220]

    WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

    Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Yutong Wang, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, et al. WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Questi...

  213. [221]

    ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Find- ings of the Association for Computati...

  214. [222]

    Visual Question Answering on 360deg Images

    Shih-Han Chou, Wei-Lun Chao, Wei-Sheng Lai, Min Sun, and Ming-Hsuan Yang. Visual Question Answering on 360deg Images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 1607–1616, 2020. 18

  215. [223]

    Geoclidean: Few-Shot Generalization in Euclidean Geometry

    Joy Hsu, Jiajun Wu, and Noah Goodman. Geoclidean: Few-Shot Generalization in Euclidean Geometry. Advances in Neural Information Processing Systems, 35:39007–39019, 2022. 18

  216. [224]

    MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. In The Twelfth International Conference on Learning...

  217. [225]

    MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024

    Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024. 18

  218. [226]

    MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yux- uan Sun, et al. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. In Proceedings of the IEEE/CVF Conference on...

  219. [227]

    SEED-Bench: Benchmarking Multimodal Large Language Models

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. SEED-Bench: Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13299–13308, June 2024. 18

  220. [228]

    MMBench: Is Your Multi-modal Model an All-Around Player? In European conference on computer vision, pages 216–233

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. MMBench: Is Your Multi-modal Model an All-Around Player? In European conference on computer vision, pages 216–233. Springer, 2024. 18

  221. [229]

    Improved Baselines with Visual Instruction Tuning, 2023

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved Baselines with Visual Instruction Tuning, 2023. 18

  222. [230]

    Nov- Phy: A Physical Reasoning Benchmark for Open-World AI Systems

    Vimukthini Pinto, Chathura Gamage, Cheng Xue, Peng Zhang, Ekaterina Nikonova, Matthew Stephenson, and Jochen Renz. Nov- Phy: A Physical Reasoning Benchmark for Open-World AI Systems. Artificial Intelligence, 336:104198, 2024. 18

  223. [231]

    M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

    Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought. arXiv preprint arXiv:2405.16473, 2024. 18

  224. [232]

    VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

    Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, and Ludwig Schimdt. VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use. In Proceedings of the 37th International Conference...

  225. [233]

    Evaluating Object Hallucination in Large Vision-Language Models

    Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating Object Hallucination in Large Vision-Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 292–305, 2023. 18

  226. [234]

    HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

    Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al. HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models. In Proce...

  227. [235]

    When’YES’Meets’ BUT’: Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? arXiv preprint arXiv:2503.23137, 2025

    Tuo Liang, Zhe Hu, Jing Li, Hao Zhang, Yiren Lu, Yunlai Zhou, Yiran Qiao, Disheng Liu, Jeirui Peng, Jing Ma, et al. When’YES’Meets’ BUT’: Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? arXiv preprint arXiv:2503.23137, 2025. 18

  228. [236]

    Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing

    Vedika Agarwal, Rakshith Shetty, and Mario Fritz. Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 9687–9695, 2019. 18

  229. [237]

    Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. Advances in Neural Information Processing Systems, 35:2507–25...

  230. [238]

    Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression

    Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 658–6...

  231. [239]

    BLEU: a Method for Automatic Evaluation of Machine Translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: a Method for Automatic Evaluation of Machine Translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Com- putational Linguistics,...

  232. [240]

    ROUGE: A Package for Automatic Evaluation of Summaries

    Chin-Yew Lin. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. 18, 19

  233. [241]

    CLIPScore: A Reference-free Evaluation Metric for Image Captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical ...

  234. [242]

    OpenAI o3 and o4-mini System Cards, 2025

    OpenAI. OpenAI o3 and o4-mini System Cards, 2025. System Cards for OpenAI’s o3 and o4-mini models. 18, 21

  235. [243]

    Reasoning with Language Model is Planning with World Model

    Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with Language Model is Planning with World Model. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 20

  236. [244]

    Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks

    Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Paul Saldyt, and Anil B Murthy. Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks. In Forty-first International Confer- ence on Machine Learning, 2024. 20

  237. [245]

    Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts

    Sukai Huang, Nir Lipovetzky, and Trevor Cohn. Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 26542–26550, 2025. 20

  238. [246]

    The empirical case for two systems of reasoning

    Steven A Sloman. The empirical case for two systems of reasoning. Psychological bulletin, 119(1):3, 1996

  239. [247]

    PromptA- gent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization

    Xinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai, Haotian Luo, Jiayou Zhang, Nebojsa Jojic, Eric Xing, and Zhiting Hu. PromptA- gent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization. InThe Twelfth International Conference on Learning Representations...

  240. [248]

    Survey on Evaluation of LLM-based Agents

    Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. Survey on Evaluation of LLM-based Agents. arXiv preprint arXiv:2503.16416, 2025. 20

  241. [249]

    LASP: Surveying the State-of-the-Art in Large Language Model- Assisted AI Planning

    Haoming Li, Zhaoliang Chen, Jonathan Zhang, and Fei Liu. LASP: Surveying the State-of-the-Art in Large Language Model- Assisted AI Planning. arXiv preprint arXiv:2409.01806, 2024. 20

  242. [250]

    Deductive Reasoning

    Philip N Johnson-Laird. Deductive Reasoning. Annual review of psychology, 50(1):109–135, 1999. 20

  243. [251]

    Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

    Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merri ¨enboer, Armand Joulin, and Tomas Mikolov. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks. arXiv preprint arXiv:1502.05698, 2015. 20

  244. [252]

    Natural Language Reasoning, A Survey

    Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. Natural Language Reasoning, A Survey. ACM Computing Surveys, 56(12):1–39, 2024. 20

  245. [253]

    Introspective Learning : A Two-Stage approach for Inference in Neural Networks

    Mohit Prabhushankar and Ghassan AlRegib. Introspective Learning : A Two-Stage approach for Inference in Neural Networks. Advances in Neural Information Processing Systems, 35:12126–12140, 2022. 20

  246. [254]

    A Survey on Neural-Symbolic Learning Systems

    Dongran Yu, Bo Yang, Da Liu, Hui Wang, and Shirui Pan. A Survey on Neural-Symbolic Learning Systems. Neural networks : the official journal of the International Neural Network Society, 166:105–126, 2021. 20

  247. [255]

    Neural Analogical Matching

    Maxwell Crouse, Constantine Nakos, Ibrahim Abdelaziz, and Ken Forbus. Neural Analogical Matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 809–817, 2021. 21

  248. [256]

    van Rooij

    Mark Blokpoel, Todd Wareham, W.F.G Pim Haselager, Ivan Toni, and I.J.E.I. van Rooij. Deep Analogical Inference as the Origin of Hypotheses. J. Probl. Solving, 11, 2019. 21

  249. [257]

    TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.Advances in neural information processing systems, 37:32731–32760, 2024

    Maitreya Patel, Naga Sai Abhiram Kusumba, Sheng Cheng, Changhoon Kim, Tejas Gokhale, Chitta Baral, et al. TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.Advances in neural information processing systems, 37:32731–32760, 2024. 21

  250. [258]

    COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning

    Xindi Wu, Hee Seung Hwang, Polina Kirichenko, and Olga Russakovsky. COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning. In Synthetic Data for Computer Vision Workshop@ CVPR 2025, 2025. 21

  251. [259]

    ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

    Jieyu Zhang, Le Xue, Linxin Song, Jun Wang, Weikai Huang, Manli Shu, An Yan, Zixian Ma, Juan Carlos Niebles, Silvio Savarese, et al. ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models. arXiv preprint arXiv:2412.07012, 2024. 21

  252. [260]

    JRDB-PanoTrack: An Open- world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

    Duy Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi, Ian Reid, Jianfei Cai, and Hamid Rezatofighi. JRDB-PanoTrack: An Open- world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...

  253. [261]

    JRDB-Act: A Large-Scale Dataset for Spatio- Temporal Action, Social Group and Activity Detection

    Mahsa Ehsanpour, Fatemeh Saleh, Silvio Savarese, Ian Reid, and Hamid Rezatofighi. JRDB-Act: A Large-Scale Dataset for Spatio- Temporal Action, Social Group and Activity Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20983...

  254. [262]

    m & m’s: A Benchmark to Evaluate Tool-Use for m ulti-step m ulti-modal Tasks

    Zixian Ma, Weikai Huang, Jieyu Zhang, Tanmay Gupta, and Ranjay Krishna. m & m’s: A Benchmark to Evaluate Tool-Use for m ulti-step m ulti-modal Tasks. In European Conference on Computer Vision, pages 18–34. Springer, 2024. 21

  255. [263]

    AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

    Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou. AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn. arXiv preprint arXiv:2306.08640, 2023. 21

  256. [264]

    Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition

    Xinyu Xi, Hua Yang, Shentai Zhang, Yijie Liu, Sijin Sun, and Xiuju Fu. Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition. arXiv preprint arXiv:2503.06978, 2025. 21

  257. [265]

    Deciphering the Role of Representation Disentan- glement: Investigating Compositional Generalization in CLIP Models

    Reza Abbasi, Mohammad Hossein Rohban, and Mahdieh Soleymani Baghshah. Deciphering the Role of Representation Disentan- glement: Investigating Compositional Generalization in CLIP Models. In European Conference on Computer Vision, pages 35–50. Springer, 2024. 21

  258. [266]

    VISCO: Benchmarking Fine- Grained Critique and Correction Towards Self-Improvement in Visual Reasoning

    Xueqing Wu, Yuheng Ding, Bingxuan Li, Pan Lu, Da Yin, Kai-Wei Chang, and Nanyun Peng. VISCO: Benchmarking Fine- Grained Critique and Correction Towards Self-Improvement in Visual Reasoning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9527–95...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.