REVIEW 2 major objections 5 minor 4 cited by
Explain Before You Answer: A Survey on Compositional Visual Reasoning
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Compositional visual reasoning is now a distinct paradigm with its own history, taxonomy, and benchmarks.
desk verdict A genuinely useful reference survey with a solid benchmark catalog, but the five-stage 'paradigm shift' is a narrative device rather than an empirically demonstrated progression; worth engaging with after revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the five-stage taxonomy and roadmap of compositional visual reasoning paradigms, anchored by the formal definition of a compositional system as a sequence of grounded intermediate steps. The sequence $S = \{s_1, \dots, s_n\}$ is the central object: it is the explicit explanation inserted between question and answer, and it may take the form of sub-questions, tool calls, scene graphs, chain-of-thought rationales, or visual manipulations such as cropping and zooming. The taxonomy does the work of partitioning 260+ papers into comparable families so that design choices and failure modes can be discussed at the level of paradigms rather than individual models. The survey also uses a benchmark catalog of 60+ datasets to connect each paradigm to the evidence used to judge it.
What would settle it
A chronological map of the 260+ surveyed papers by publication date, architecture, and stage assignment would settle the matter: if later-stage systems such as unified agentic vision-language models regularly predate earlier-stage systems, or if independent researchers cannot agree on stage labels for individual papers, then the five-stage progression is an imposed narrative rather than a discovered one.
Extended reading notes
Core claim
The central claim is that visual reasoning has been undergoing a paradigm shift from monolithic models, which directly predict an answer from the image-question pair, to compositional systems that deliberately expose their reasoning. The survey formalizes compositional visual reasoning as any approach that maps a visual input $v$ and query $q$ to an answer $y$ through an intermediate structured representation $S = \{s_1, \dots, s_n\}$, where each step may ground to objects, attributes, or relations, and the steps are executed sequentially or hierarchically. It then divides the recent literature into a five-stage progression: prompt-enhanced language-centric pipelines, tool-enhanced large language models, tool-enhanced vision-language models, chain-of-thought reasoning vision-language models, and unified agentic vision-language models. Each stage is characterized by its architectural design, strengths, and limitations, and the survey positions this progression as the field's historical roadmap.
Load-bearing premise
The whole roadmap depends on the assumption that the field really evolved through those five stages in that order, which the survey asserts without quantitative evidence that the stages are historically real, mutually distinct, and non-overlapping.
Editorial extensions
If this is right
- Papers in this area can now be located on a shared map: a method is a Stage I through Stage V system, making cross-paper comparisons more systematic.
- The benchmark catalog gives researchers a ready-made test battery for grounding accuracy, chain-of-thought faithfulness, and high-resolution perception.
- If the survey's evaluation critique is correct, future benchmark design should add step-level annotations rather than scoring only final answers.
- The open-challenge list points to concrete next targets, including world-model integration, human-in-the-loop verification, and richer evaluation protocols.
- The roadmap suggests the field's center of gravity is moving toward unified agentic vision-language models, so new work will likely build on those architectures.
Reading between the lines
- Editorial inference: the five-stage ordering implies a falsifiable prediction that publication dates of the 260+ papers cluster by stage, with later stages peaking later; this can be checked bibliometrically.
- Editorial inference: if step-level evaluation becomes standard, reported gains from chain-of-thought methods may shrink, because models can now be caught producing correct final answers from incorrect intermediate deductions.
- Editorial inference: the monolithic-versus-compositional distinction suggests a testable hypothesis that forcing explicit grounding steps reduces hallucination rates more than prompt-based fixes do, measurable on benchmarks such as POPE or HallusionBench.
- Editorial inference: a natural next step the survey names but does not formalize is a sixth stage, world-model-integrated agents that simulate hypothetical scenes during reasoning.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of compositional visual reasoning (CVR) in the vision-language domain. It defines CVR, contrasts it with monolithic reasoning, and argues for its advantages in cognitive alignment, interpretability, generalization, and efficiency. The paper organizes recent methods into five developmental stages: prompt-enhanced language-centric pipelines, tool-enhanced LLMs, tool-enhanced VLMs, chain-of-thought VLMs, and unified agentic VLMs. It also catalogs benchmarks and evaluation metrics, and closes with insights, open challenges, and future directions. The authors claim to provide the first comprehensive, systematic review of 260+ papers from top venues between 2023 and 2025 and explicitly position the five-stage progression as a historical roadmap marking a 'paradigm shift.'
Significance. If the roadmap and taxonomy are validated, this survey would be a valuable reference for a rapidly growing area, bundling many recent systems into a structured framework and pointing to concrete open problems. The paper is clearly written and broad in coverage, with useful illustrations and a wide-ranging benchmark catalog. The authors are to be credited for assembling recent literature on tool-based, chain-of-thought, and agentic visual reasoning and for identifying challenges such as step-level evaluation and world-model integration. However, the survey's central contribution is the historical 'five-stage paradigm shift,' and as currently presented that claim is not supported by systematic evidence; the taxonomy may still be useful as a conceptual organization, but the specific 'historical roadmap' framing needs revision or additional evidence.
major comments (2)
- [Section 4 (Key Stages), including Figure 3] The central claim that the paper 'trace[s] a five-stage paradigm shift' is not supported by the evidence presented in the manuscript. No assignment rules are given for placing a paper into a stage, no publication-date distribution or bibliometric analysis establishes the proposed chronological ordering, and the stages are not disjoint in the paper's own usage: CoF [78] is described as a Stage IV visually grounded CoT VLM in Sec. 4.4.3 and again as a Stage V unified agentic VLM in Sec. 4.5.1, while Visual Sketchpad [146] appears as a Stage III image-feedback tool-use method in Sec. 4.3.3 despite also functioning as a multi-step visual reasoning agent. Because this roadmap is the survey's stated contribution beyond existing surveys, the authors should either provide explicit stage-assignment criteria and temporal evidence (e.g., a plot of publication dates of representative systems per stage) or explicitly reframe the five stages as a conceptual taxonomy of architectures rather than a chronological paradigm shift.
- [Abstract and Section 1] The paper claims to 'systematically review 260+ papers from top venues,' but it reports no inclusion criteria, search strategy, screening process, or inter-annotator agreement for its taxonomy. This makes the 'comprehensive' and 'systematic' claims difficult to verify and leaves the survey open to selection bias; for example, the stage narratives in Sec. 4.2 and 4.3 rely heavily on the authors' own systems (HYDRA [29], DWIM [76], NA VER [137]), and no protocol is given that would explain how other relevant works were selected or excluded. A short methodology subsection describing the databases searched, the exact time window, inclusion/exclusion criteria, and the procedure for assigning works to stages would materially strengthen the survey and should be added before publication.
minor comments (5)
- [Section 4.4.2] The KL-divergence characterization of SFT versus RL is an oversimplification: supervised fine-tuning is better described as likelihood maximization (empirical forward KL), while RL fine-tuning such as RLHF typically optimizes a reward with a KL penalty to a reference policy rather than directly minimizing reverse KL. The authors should either provide a precise citation for this claim or soften the wording to 'interpreted as' to avoid misleading readers.
- [Section 5.1] TextVQA is cited as [212], but reference [212] in the bibliography is 'Towards Visual Dialog for Radiology' (Kovaleva et al.), not the TextVQA paper. Please replace it with the correct TextVQA citation and recheck all benchmark references in that section.
- [Section 5.1] The text refers to a 'Visual Abstractions Benchmark [190],' but reference [190] is 'What Makes a Maze Look Like a Maze?' by Hsu et al.; the dataset name and citation do not match, so the authors should verify the intended benchmark and its reference.
- [References] Reference [168] for 'Visual Agents as Fast and Slow Thinkers' lists the venue as 'In The Thirteenth International Conference on Learning Representations, 2018,' but the thirteenth ICLR is in 2025; the year should be corrected.
- [Section 5.3] The discussion of advances in evaluation says 'the next frontier lies in unifying these strengths' and mentions 'average output token length' in Sec. 5.2, but no reference or definition is given for the latter metric; please add a citation or brief explanation.
Circularity Check
No significant circularity: the survey organizes the literature with a taxonomy and roadmap, and its claims do not reduce to fitted inputs, self-citation chains, or definitional equivalences.
full rationale
This is a survey paper with no mathematical derivations, fitted parameters, or predictive claims that could reduce to their inputs by construction. The central contribution is organizational: the paper defines compositional visual reasoning, reviews 260+ papers, and proposes a five-stage taxonomy in Section 4. The five-stage 'paradigm shift' is asserted rhetorically rather than proven empirically, and some stage assignments overlap (e.g., CoF appears in both Stage IV and Stage V), but an asserted or under-evidenced historical narrative is a correctness or scope concern, not circularity. The paper's self-citations (HYDRA, DWIM, NAVER, LEFT, and others) are used as representative examples of existing methods, not as load-bearing evidence for a theorem or uniqueness result; the taxonomy would stand even if those specific papers were replaced. There is no equation where an output is defined in terms of the input, no fitted value relabeled as a prediction, and no imported uniqueness theorem that forces the chosen organization. Accordingly, the survey is self-contained as a literature review and no circular step is present.
Assumptions & free parameters
assumptions (3)
- domain assumption Compositionality principle: the meaning of the whole is a function of the meanings of its parts.
- ad hoc to paper Explicit intermediate reasoning steps are the defining characteristic of compositional visual reasoning.
- ad hoc to paper The five-stage progression reflects the actual chronological and architectural development of the field from 2023 to 2025.
Cite this review
Pith. "Pith review of Explain Before You Answer: A Survey on Compositional Visual Reasoning." pith.science (2026). https://pith.science/paper/RN2WI3MQ
@misc{pith2026250817298,
author = {Pith},
title = {Pith review of: Explain Before You Answer: A Survey on Compositional Visual Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/RN2WI3MQ}},
note = {Machine review of arXiv:2508.17298}
}
read the original abstract
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference. While early surveys focus on monolithic vision-language models or general multimodal reasoning, a dedicated synthesis of the rapidly expanding compositional visual reasoning literature is still missing. We fill this gap with a comprehensive survey spanning 2023 to 2025 that systematically reviews 260+ papers from top venues (CVPR, ICCV, NeurIPS, ICML, ACL, etc.). We first formalize core definitions and describe why compositional approaches offer advantages in cognitive alignment, semantic fidelity, robustness, interpretability, and data efficiency. Next, we trace a five-stage paradigm shift: from prompt-enhanced language-centric pipelines, through tool-enhanced LLMs and tool-enhanced VLMs, to recently minted chain-of-thought reasoning and unified agentic VLMs, highlighting their architectural designs, strengths, and limitations. We then catalog 60+ benchmarks and corresponding metrics that probe compositional visual reasoning along dimensions such as grounding accuracy, chain-of-thought faithfulness, and high-resolution perception. Drawing on these analyses, we distill key insights, identify open challenges (e.g., limitations of LLM-based reasoning, hallucination, a bias toward deductive reasoning, scalable supervision, tool integration, and benchmark limitations), and outline future directions, including world-model integration, human-AI collaborative reasoning, and richer evaluation protocols. By offering a unified taxonomy, historical roadmap, and critical outlook, this survey aims to serve as a foundational reference and inspire the next generation of compositional visual reasoning research.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 4 Pith papers
-
Dynamic Execution Commitment of Vision-Language-Action Models
A3 reframes dynamic action chunk commitment in VLA models as self-speculative prefix verification, accepting the longest continuous sequence of actions that satisfies consensus-ordered conditional invariance and prefi...
-
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
Rule-VLN is the first large-scale benchmark injecting 177 regulatory categories into an urban environment, and the proposed SNRM module equips pre-trained VLN agents with zero-shot semantic reasoning and detour planni...
-
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
A training framework that adds a removable image-generation branch to multimodal LLMs improves visual understanding benchmarks with zero inference-time cost.
-
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Fine-tuning Qwen2.5-VL on VisReason, a 489K-example multi-round visual chain-of-thought dataset (165K with pseudo-depth), modestly improves LLM-judged visual reasoning scores, with caveats about self-referential 3D ev...
Reference graph
Works this paper leans on
-
[78]
Chain-of-Focus: Adaptive Visual Search and Zooming for Multimodal Reasoning via RL
Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, et al. Chain-of-Focus: Adaptive Visual Search and Zooming for Multimodal Reasoning via RL. arXiv preprint arXiv:2505.15436,
-
[146]
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, Jianfeng Gao, et al. Multimodal Foundation Models: From Specialists to General-Purpose Assistants. Foundations and Trends® in Computer Graphics and Vision, 16(1-2):1– 214, 2024. 12, 13
2024
-
[29]
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
Fucai Ke, Zhixi Cai, Simindokht Jahangard, Weiqing Wang, Pari Delir Haghighi, and Hamid Rezatofighi. HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning. In European Conference on Computer Vision, pages 132–149. Springer, 2024. 3, 6, 7, 11, 21
2024
-
[76]
Fucai Ke, Xingjian Leng, Zhixi Cai, Zaid Khan, Weiqing Wang, Pari Delir Haghighi, Hamid Rezatofighi, Manmohan Chandraker, et al. DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning. arXiv preprint arXiv:2503.19263, 2025. 6, 11, 16, 20, 21
arXiv 2025
-
[137]
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
Zhixi Cai, Fucai Ke, Simindokht Jahangard, Maria Garcia de la Banda, Reza Haffari, Peter J Stuckey, and Hamid Rezatofighi. NA VER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning. arXiv preprint arXiv:2502.00372, 2025. 11, 18
work page Pith review arXiv 2025
-
[1]
A Benchmark for Compositional Visual Reasoning
Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre. A Benchmark for Compositional Visual Reasoning. Advances in neural information processing systems, 35:29776–29788, 2022. 2, 5, 6, 7
2022
-
[2]
Take A Step Back: Rethinking the Two Stages in Visual Reasoning
Mingyu Zhang, Jiting Cai, Mingyu Liu, Yue Xu, Cewu Lu, and Yong-Lu Li. Take A Step Back: Rethinking the Two Stages in Visual Reasoning. In European Conference on Computer Vision, pages 124–141. Springer, 2024. 2
2024
-
[3]
The Book of Why: The New Science of Cause and Effect
M ´elanie Frappier. The Book of Why: The New Science of Cause and Effect. Science, 361:855 – 855, 2018
2018
Show all 266 references
-
[4]
Inferring and Executing Programs for Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Inferring and Executing Programs for Visual Reasoning. In Proceedings of the IEEE international conference on computer vision, pages 2989–2998, 2017. 3, 6, 7
2017
-
[5]
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...
2017
-
[6]
RA VEN: A Dataset for Relational and Analogical Visual REasoNing
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. RA VEN: A Dataset for Relational and Analogical Visual REasoNing. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5312–5322, 2019. 2, 17
2019
-
[7]
Jakob Suchan, Mehul Bhatt, Przemyslaw Andrzej Walega, and Carl P. L. Schultz. Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects. ArXiv, abs/1712.00840, 2017
2017 arXiv
-
[8]
Maintaining Reasoning Consistency in Compositional Visual Question Answering
Chenchen Jing, Yunde Jia, Yuwei Wu, Xinyu Liu, and Qi Wu. Maintaining Reasoning Consistency in Compositional Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5099–5108, 2022. 2, 6
2022
-
[9]
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. Advances in neural information processing systems, 36:34892–34916, 2023. 2, 3, 5, 6, 7, 15
2023
-
[10]
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.arXiv preprint arXiv:2308.12966,
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.arXiv preprint arXiv:2308.12966,
-
[11]
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In International conference on machine learning, pages 19730–19742. PMLR, 2023. 6, 7, 10
2023
-
[12]
InstructBLIP: towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: towards general-purpose vision-language models with instruction tuning. In Proceedings of the 37th International Conference on Neural ...
2023
-
[13]
A Survey of Multimodel Large Language Models
Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, and Ke Liu. A Survey of Multimodel Large Language Models. In Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering , pages 405–409, 2024. 2, 3
2024
-
[14]
Interpretable Visual Reasoning: A Survey
Feijuan He, Yaxian Wang, Xianglin Miao, and Xia Sun. Interpretable Visual Reasoning: A Survey. Image and Vision Computing, 112:104194, 2021. 2, 5
2021
-
[15]
Thinking Machines: A Survey of LLM based Reasoning Strategies
Dibyanayan Bandyopadhyay, Soham Bhattacharjee, and Asif Ekbal. Thinking Machines: A Survey of LLM based Reasoning Strategies. arXiv preprint arXiv:2503.10814, 2025
2025 arXiv
-
[16]
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wanxiang Che. Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models. arXiv preprint arXiv:2503.09567, 2025. 2
2025 arXiv
-
[17]
Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba
Yunsheng Ma, Wenqian Ye, Can Cui, Haiming Zhang, Shuo Xing, Fucai Ke, Jinhong Wang, Chenglin Miao, Jintai Chen, Hamid Rezatofighi, et al. Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba. In Proceedings of ...
2025
-
[18]
NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UA V Search Missions.arXiv preprint arXiv:2409.10196, 2024
Zhixi Cai, Cristian Rojas Cardenas, Kevin Leo, Chenyuan Zhang, Kal Backman, Hanbing Li, Boying Li, Mahsa Ghorbanali, Stavya Datta, Lizhen Qu, et al. NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UA V Search Missi...
2024 arXiv
-
[19]
UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning
Jorge Fernandez Galarreta, Norman Kerle, and Markus Gerke. UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning. Natural hazards and earth system sciences, 15(6):1087–1101, 2015
2015
-
[20]
A Study of Visual Reasoning in Medical Diagnosis
Erika Rogers. A Study of Visual Reasoning in Medical Diagnosis. In Proceedings of the Eighteenth Annual Conference of the Cognitive Science Society, pages 213–218. Routledge, 2019
2019
-
[21]
Medical Visual Question Answering via Conditional Reasoning
Li-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen, and Xiao-Ming Wu. Medical Visual Question Answering via Conditional Reasoning. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2345–2354, 2020
2020
-
[22]
Programmatically Grounded, Compositionally Generalizable Robotic Manipulation
Renhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao, Jiajun Wu, and Yang Gao. Programmatically Grounded, Compositionally Generalizable Robotic Manipulation. arXiv preprint arXiv:2304.13826, 2023. 3 23
2023 arXiv
-
[23]
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...
2022
-
[24]
When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,
-
[25]
Challenges and Prospects in Vision and Language Research
Kushal Kafle, Robik Shrestha, and Christopher Kanan. Challenges and Prospects in Vision and Language Research. Frontiers in Artificial Intelligence, 2, 2019
2019
-
[26]
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
Yiqi Wang, Wentao Chen, Xiaotian Han, Xudong Lin, Haiteng Zhao, Yongfei Liu, Bohan Zhai, Jianbo Yuan, Quanzeng You, and Hongxia Yang. Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasonin...
2024 arXiv
-
[27]
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng, Shilong Liu, Liang-Yan Gui, Jan Kautz, Yu-Xiong Wang, and Zhiding Yu. Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 14268–14280, 2025. 15
2025
-
[28]
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 10632–10643, 2025. 3
2025
-
[30]
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
Aleksandar Stani ´c, Sergi Caelles, and Michael Tschannen. Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers. Transactions on Machine Learning Research, 2024. 6, 7
2024
-
[31]
Multimodal Chain-of- Thought Reasoning: A Comprehensive Survey
Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. Multimodal Chain-of- Thought Reasoning: A Comprehensive Survey. arXiv preprint arXiv:2503.12605, 2025. 3, 4, 6
2025 arXiv
-
[32]
Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining
Ugur Sahin, Hang Li, Qadeer Khan, Daniel Cremers, and V olker Tresp. Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5563–5573, 2024
2024
-
[33]
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muen- nighoff, Kyle Lo, Luca Soldaini, et al. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models. In Proceedings of the Compute...
2025
-
[34]
Cascaded Mutual Modulation for Visual Reasoning
Yiqun Yao, Jiaming Xu, Feng Wang, and Bo Xu. Cascaded Mutual Modulation for Visual Reasoning. In Conference on Empirical Methods in Natural Language Processing, 2018. 3
2018
-
[35]
Compositional Chain-of-Thought Prompting for Large Mul- timodal Models
Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional Chain-of-Thought Prompting for Large Mul- timodal Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14431,
-
[36]
Building Machines That Learn and Think Like People
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building Machines That Learn and Think Like People. Behavioral and brain sciences, 40:e253, 2017. 3, 6
2017
-
[37]
Meta Module Network for Compositional Visual Reasoning
Wenhu Chen, Zhe Gan, Linjie Li, Yu Cheng, William Wang, and Jingjing Liu. Meta Module Network for Compositional Visual Reasoning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 655–664, 2021. 5, 6
2021
-
[38]
Improving Visual Reasoning Through Semantic Representation
Wenfeng Zheng, Xiangjun Liu, Xubin Ni, Lirong Yin, and Bo Yang. Improving Visual Reasoning Through Semantic Representation. IEEE access, 9:91476–91486, 2021
2021
-
[39]
What’s Left? Concept Grounding with Logic-Enhanced Foundation Models
Joy Hsu, Jiayuan Mao, Josh Tenenbaum, and Jiajun Wu. What’s Left? Concept Grounding with Logic-Enhanced Foundation Models. Advances in Neural Information Processing Systems, 36:38798–38814, 2023. 3, 11
2023
-
[40]
Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges
Byeong Su Kim, Jieun Kim, Deokwoo Lee, and Beakcheol Jang. Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges. ACM Computing Surveys, 57(10):1–35, 2025. 3, 4, 17, 18, 21
2025
-
[41]
Visual Question Answering: a State-of-the-Art Review
Sruthy Manmadhan and Binsu C Kovoor. Visual Question Answering: a State-of-the-Art Review. Artificial Intelligence Review, 53(8):5705–5745, 2020
2020
-
[42]
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
Jiayi Kuang, Ying Shen, Jingyou Xie, Haohao Luo, Zhe Xu, Ronghao Li, Yinghui Li, Xianfeng Cheng, Xika Lin, and Yu Han. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey. ACM Computing Surveys, 57(8):1–36, 2025. 3, 4, 5, 20
2025
-
[43]
Mind with Eyes: from Language Reasoning to Multimodal Reasoning
Zhiyu Lin, Yifei Gao, Xian Zhao, Yunfan Yang, and Jitao Sang. Mind with Eyes: from Language Reasoning to Multimodal Reasoning. arXiv preprint arXiv:2503.18071, 2025. 3, 4, 6
2025 arXiv
-
[44]
A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge
M Jaleed Khan, Filip Ilievski, John G Breslin, and Edward Curry. A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge. Neurosymbolic Artificial Intelligence, 1:NAI–240719, 2025. 3, 4, 6, 20
2025
-
[45]
Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices
Mikołaj Małki ´nski and Jacek Ma´ndziuk. Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices. ACM Computing Surveys, 57(7):1–36, 2025. 3, 4, 5 24
2025
-
[46]
Large Multimodal Agents: A Survey
Junlin Xie, Zhihong Chen, Ruifei Zhang, Xiang Wan, and Guanbin Li. Large Multimodal Agents: A Survey. arXiv preprint arXiv:2402.15116, 2024. 3, 4
2024 arXiv
-
[47]
From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities
Md Farhan Ishmam, Md Sakib Hossain Shovon, Muhammad Firoz Mridha, and Nilanjan Dey. From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities. Information Fusion, 106:102270, 2024. 4, 5, 7
2024
-
[48]
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
Jie Ma, Pinghui Wang, Dechen Kong, Zewei Wang, Jun Liu, Hongbin Pei, and Junzhou Zhao. Robust Visual Question Answering: Datasets, Methods, and Future Challenges. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 4, 6, 21
2024
-
[49]
Perception, reason, think, and plan: A survey on large multimodal reasoning models.arXiv preprint arXiv:2505.04921,
Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al. Perception, reason, think, and plan: A survey on large multimodal reasoning models.arXiv preprint arXiv:2505.04921,
-
[50]
How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model
Shezheng Song, Xiaopeng Li, Shasha Li, Shan Zhao, Jie Yu, Jun Ma, Xiaoguang Mao, Weimin Zhang, and Meng Wang. How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model. IEEE Transactions on Knowledge and Data Engineering, 2025. 4
2025
-
[51]
Tool learning with large language models: a survey
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: a survey. Frontiers of Computer Science, 19(8):198343, 2025. 4
2025
-
[52]
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9199...
2025
-
[53]
Multimodal Intelligence: Representation Learning, Information Fusion, and Applications
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. Multimodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing, 14(3):478–493, 2020. 5
2020
-
[54]
Transformation Driven Visual Reasoning
Xin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo, and Xueqi Cheng. Transformation Driven Visual Reasoning. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 6903–6912, 2021. 5
2021
-
[55]
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 5
2016
-
[56]
Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction
Wentao He, Jianfeng Ren, Ruibin Bai, and Xudong Jiang. Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction. Pattern Recognition, 160:111151, 2025. 5
2025
-
[57]
Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning
Wentao He, Jialu Zhang, Jianfeng Ren, Ruibin Bai, and Xudong Jiang. Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 22–30, 2023. 5
2023
-
[58]
You Only Look Once: Unified, Real-Time Object Detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016. 5
2016
-
[59]
Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering. In Proceedings of the IEEE international conference on computer vision, pages 1821–1830, 2017. 5, 6
2017
-
[60]
Bilinear Attention Networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear Attention Networks. Advances in neural information processing systems, 31, 2018. 6
2018
-
[61]
Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering
Duy-Kien Nguyen and Takayuki Okatani. Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 6087–6096, 2018. 5
2018
-
[62]
Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep Modular Co-Attention Networks for Visual Question Answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281–6290, 2019. 6
2019
-
[63]
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Advances in neural information processing systems, 32, 2019. 6
2019
-
[64]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Supervision. In International conference on machine learning, pa...
2021
-
[65]
UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: UNiversal Image-TExt Representation Learning. In European conference on computer vision, pages 104–120. Springer, 2020. 6
2020
-
[66]
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a Visual Language Model for Few-Shot Learning. Advances in neural information processing systems, 35:23716...
2022
-
[67]
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In The Twelfth International Conference on Learning Representations, 2024. 6
2024
-
[68]
Otter: A Multi-Modal Model With In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Otter: A Multi-Modal Model With In-Context Instruction Tuning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 6, 7 25
2025
-
[69]
CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna. CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023. 6, 16
2023
-
[70]
Logics and Languages
Max J Cresswell. Logics and Languages. Routledge, 2016. 6
2016
-
[71]
From Recognition to Cognition: Visual Commonsense Reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. From Recognition to Cognition: Visual Commonsense Reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731, 2019. 6, 7, 17
2019
-
[72]
Visual Programming: Compositional Visual Reasoning without Training
Tanmay Gupta and Aniruddha Kembhavi. Visual Programming: Compositional Visual Reasoning without Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962, 2023. 7, 11
2023
-
[73]
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Drew A Hudson and Christopher D Manning. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 6700–6709, 2019. 7, 14, 16, 18
2019
-
[74]
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. Advances in Neural Information Processing Systems , 36,
-
[75]
Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions
Huaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu, Yuke Zhu, and Anima Anandkumar. Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19056–19065, 2022. 6, 7
2022
-
[77]
Iterated Learning Improves Compositionality in Large Vision-Language Models
Chenhao Zheng, Jieyu Zhang, Aniruddha Kembhavi, and Ranjay Krishna. Iterated Learning Improves Compositionality in Large Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13785–13795, June 2024. 6
2024
-
[79]
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
Di Zhang, Jingdi Lei, Junxian Li, Xunzhi Wang, Yujie Liu, Zonglin Yang, Jiatong Li, Weida Wang, Suorong Yang, Jianbo Wu, et al. Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages ...
2025
-
[80]
Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, and David Acuna. Can Large Vision-Language Models Correct Semantic Ground- ing Errors By Themselves? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 14667–14678, 2025. 6
2025
-
[81]
Visual Compositional Learning for Human-Object Interaction Detection
Zhi Hou, Xiaojiang Peng, Yu Qiao, and Dacheng Tao. Visual Compositional Learning for Human-Object Interaction Detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16, pages 584–600. Springer, 2020. 6
2020
-
[82]
Learn- ing Visual Composition through Improved Semantic Guidance
Austin Stone, Hagen Soltau, Robert Geirhos, Xi Yi, Ye Xia, Bingyi Cao, Kaifeng Chen, Abhijit Ogale, and Jonathon Shlens. Learn- ing Visual Composition through Improved Semantic Guidance. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 3740–3750, 2025
2025
-
[83]
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage
Andy Zeng, Maria Attarian, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, et al. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage. In The Eleventh International Co...
2023
-
[84]
Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models
Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi, and Ranjay Krishna. Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models. In IEEE/CVF Conference ...
2025
-
[85]
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
Jiaxing Chen, Yuxuan Liu, Dehu Li, Xiang An, Weimo Deng, Ziyong Feng, Yongle Zhao, and Yin Xie. Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models. arXiv preprint arXiv:2403.19322, 2024. 7, 12, 13, 20, 21
2024 arXiv
-
[86]
DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage
Joy Hsu, Jiayuan Mao, and Jiajun Wu. DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage. Transactions on Machine Learning Research, 2023. 7
2023
-
[87]
Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning
Shi Chen and Qi Zhao. Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6736–6745, 2023. 7
2023
-
[88]
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
Zejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang, and Zhongyu Wei. VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computation...
2025
-
[89]
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
Zaid Khan, Vijay Kumar BG, Samuel Schulter, Yun Fu, and Manmohan Chandraker. Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14344–1...
2024
-
[90]
From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis
Chuanqi Cheng, Jian Guan, Wei Wu, and Rui Yan. From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4941–4957, 2024. 12
2024
-
[91]
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu, Yan Ma, Xiaoye Qu, Jiaqi Liu, Yanshu Li, Kaide Zeng, Zhengyuan Yang, et al. Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers. arXiv preprint arXiv:2506.23918,
-
[92]
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
Qiji Zhou, Ruochen Zhou, Zike Hu, Panzhong Lu, Siyang Gao, and Yue Zhang. Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models. arXiv preprint arXiv:2405.13872, 2024. 7, 12, 18, 20
2024 arXiv
-
[93]
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). arXiv preprint arXiv:2309.17421, 9(1):1, 2023
2023 arXiv
-
[94]
Grounded Chain-of- Thought for Multimodal Large Language Models
Qiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang, Baiyang Song, Xiaoshuai Sun, and Rongrong Ji. Grounded Chain-of- Thought for Multimodal Large Language Models. arXiv preprint arXiv:2503.12799, 2025. 7, 14, 19, 20, 21
2025 arXiv
-
[95]
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Baiqi Li, Zhiqiu Lin, Wenxuan Peng, Jean de Dieu Nyandwi, Daniel Jiang, Zixian Ma, Simran Khanuja, Ranjay Krishna, Graham Neubig, and Deva Ramanan. NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, ...
2024
-
[96]
BLINK: Multimodal Large Language Models Can See but Not Perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. BLINK: Multimodal Large Language Models Can See but Not Perceive. In European Conference on Computer Vision, pages 148–166. Springer, 2024. 7
2024
-
[97]
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora. ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty. Advances in Neural Information Processing Systems, 37:86004–86047, 2024. 7
2024
-
[98]
Visual Reasoning by Progressive Module Networks
Seung Wook Kim, Makarand Tapaswi, and Sanja Fidler. Visual Reasoning by Progressive Module Networks. In International Conference on Learning Representations, 2018. 7
2018
-
[99]
Predicate Hierarchies Improve Few-Shot State Classification
Emily Jin, Joy Hsu, and Jiajun Wu. Predicate Hierarchies Improve Few-Shot State Classification. arXiv preprint arXiv:2502.12481,
-
[100]
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
Yushi Hu, Otilia Stretcu, Chun-Ta Lu, Krishnamurthy Viswanathan, Kenji Hata, Enming Luo, Ranjay Krishna, and Ariel Fuxman. Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Compute...
2024
-
[101]
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng, Fan Yang, Ming Tang, and Jinqiao Wang. Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models. arXiv preprint arXiv:2505.20753, 2025. 7, 14, 16, 19, 20
2025 arXiv
-
[102]
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv preprint arXiv:2409.12191,
-
[103]
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4818–...
2024
-
[104]
LLaV A-NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. LLaV A-NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024. 15
2024
-
[105]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024
2024 arXiv
-
[106]
DeepSeek-VL: Towards Real-World Vision-Language Understanding.arXiv preprint arXiv:2403.05525, 2024
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, et al. DeepSeek-VL: Towards Real-World Vision-Language Understanding.arXiv preprint arXiv:2403.05525, 2024
2024 arXiv
-
[107]
GPT-4 Technical Report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[108]
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality. arXiv preprint arXiv:2304.14178,
-
[109]
SpatialVLM: Endowing Vision- Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. SpatialVLM: Endowing Vision- Language Models with Spatial Reasoning Capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14455–144...
2024
-
[110]
Visual Spatial Reasoning
Fangyu Liu, Guy Emerson, and Nigel Collier. Visual Spatial Reasoning. Transactions of the Association for Computational Linguistics, 11:635–651, 2023
2023
-
[111]
CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations.arXiv preprint arXiv:2402.04236, 2024
Ji Qi, Ming Ding, Weihan Wang, Yushi Bai, Qingsong Lv, Wenyi Hong, Bin Xu, Lei Hou, Juanzi Li, Yuxiao Dong, et al. CogCoM: Train Large Vision-Language Models Diving into Details through Chain of Manipulations.arXiv preprint arXiv:2402.04236, 2024. 8, 15, 16, 19, 20, 21 27
2024 arXiv
-
[112]
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program Synthesis with Large Language Models. arXiv preprint arXiv:2108.07732, 2021. 8
2021 arXiv
-
[113]
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large Language Models are Zero-Shot Reasoners. Advances in neural information processing systems, 35:22199–22213, 2022
2022
-
[114]
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021. 8
-
[115]
DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multi- modal Reasoning in Language Models
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multi- modal Reasoning in Language Models. Advances in Neural Information Processing Systems, 36:5168–5191, 2023. 10
2023
-
[116]
ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions
Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny. ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions. Transactions on Machine Learning Research, 2024. 10
2024
-
[117]
Ide- alGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
Haoxuan You, Rui Sun, Zhecan Wang, Long Chen, Gengyu Wang, Hammad Ayyubi, Kai-Wei Chang, and Shih-Fu Chang. Ide- alGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pa...
2023
-
[118]
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
Imad Eddine Toubal, Aditya Avinash, Neil Gordon Alldrin, Jan Dlabal, Wenlei Zhou, Enming Luo, Otilia Stretcu, Hao Xiong, Chun-Ta Lu, Howard Zhou, et al. Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use. InProceedings o...
-
[119]
Large Language Models are Visual Reasoning Coordinators
Liangyu Chen, Bo Li, Sheng Shen, Jingkang Yang, Chunyuan Li, Kurt Keutzer, Trevor Darrell, and Ziwei Liu. Large Language Models are Visual Reasoning Coordinators. Advances in Neural Information Processing Systems, 36:70115–70140, 2023. 10
2023
-
[120]
PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo. PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2963–2975, 2023. 10
2023
-
[121]
Multimodal Chain-of-Thought Reasoning in Language Models
Zhuosheng Zhang, Aston Zhang, Mu Li, hai zhao, George Karypis, and Alex Smola. Multimodal Chain-of-Thought Reasoning in Language Models. Transactions on Machine Learning Research, 2024. 10
2024
-
[122]
AdGPT: Explore Meaningful Advertising with ChatGPT
Jiannan Huang, Mengxue Qu, Longfei Li, and Yunchao Wei. AdGPT: Explore Meaningful Advertising with ChatGPT. ACM Trans. Multimedia Comput. Commun. Appl., 21(4), April 2025. 10
2025
-
[123]
Visually Descriptive Language Model for Vector Graphics Reasoning
Zhenhailong Wang, Joy Hsu, Xingyao Wang, Kuan-Hao Huang, Manling Li, Jiajun Wu, and Heng Ji. Visually Descriptive Language Model for Vector Graphics Reasoning. arXiv preprint arXiv:2404.06479, 2024. 10
2024 arXiv
-
[124]
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
Hulingxiao He, Geng Li, Zijun Geng, Jinglin Xu, and Yuxin Peng. Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models. In The Thirteenth International Conference on Learning Representations ,
-
[125]
Visual Chain-of- Thought Prompting for Knowledge-Based Visual Reasoning
Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Zhiqing Sun, Dan Gutfreund, and Chuang Gan. Visual Chain-of- Thought Prompting for Knowledge-Based Visual Reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, pages 1254–1262, 2024. 10, 19
2024
-
[126]
VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things
Yaoyao Zhong, Mengshi Qi, Rui Wang, Yuhan Qiu, Yang Zhang, and Huadong Ma. VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 39, pages 10680–10688, 2025. 11
2025
-
[127]
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction. Advances in Neural Information Processing Systems, 36:71995–72007, 2023. 11, 13
2023
-
[128]
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models. arXiv preprint arXiv:2303.04671, 2023. 11
2023 arXiv
-
[129]
ViperGPT: Visual Inference via Python Execution for Reasoning.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11854–11864, 2023
D’idac Sur’is, Sachit Menon, and Carl V ondrick. ViperGPT: Visual Inference via Python Execution for Reasoning.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11854–11864, 2023. 11
2023
-
[130]
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. Advances in Neural Information Processing Systems, 36:38154–38180, 2023. 11
2023
-
[131]
CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets
Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung, Hao Peng, and Heng Ji. CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets. In The Twelfth International Conference on Learning Representations, 2024. 11
2024
-
[132]
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action. arXiv preprint arXiv:2303.11381, 2023. 11, 21
2023 arXiv
-
[133]
ContextualCoder: Adaptive In-context Prompting for Programmatic Visual Question Answering
Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex C Kot, and Koichi Shinoda. ContextualCoder: Adaptive In-context Prompting for Programmatic Visual Question Answering. IEEE Transactions on Multimedia, 2025. 11
2025
-
[134]
InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang, Yi Wang, Shoufa Chen, Qinglong Zhang, Zeqiang Lai, Yang Yang, Qingyun Li, et al. InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language. arXiv preprint arXiv:2305.05662, 2023. 11
2023 arXiv
-
[135]
Can Visual Scratchpads With Diagrammatic Abstractions Augment LLM Reasoning? In Proceedings on, pages 21–28
Joy Hsu, Gabriel Poesia, Jiajun Wu, and Noah Goodman. Can Visual Scratchpads With Diagrammatic Abstractions Augment LLM Reasoning? In Proceedings on, pages 21–28. PMLR, 2023. 11 28
2023
-
[136]
CLOV A: A Closed-LOop Visual Assistant with Tool Usage and Update
Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wenjuan Han, Song-Chun Zhu, and Qing Li. CLOV A: A Closed-LOop Visual Assistant with Tool Usage and Update. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13258–13268, 2024. 11, 16, 21
2024
-
[138]
SYNAPSE: SYmbolic Neural-Aided Preference Synthesis Engine
Sadanand Modak, Noah Tobias Patton, Isil Dillig, and Joydeep Biswas. SYNAPSE: SYmbolic Neural-Aided Preference Synthesis Engine. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 27529–27537, 2025. 11
2025
-
[139]
ViUniT: Visual Unit Tests for More Robust Visual Programming
Artemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, and Juan Carlos Niebles. ViUniT: Visual Unit Tests for More Robust Visual Programming. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24646–2...
2025
-
[140]
LLaV A-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. LLaV A-Plus: Learning to Use Tools for Creating Multimodal Agents. In European Conference on Computer Vision, pages 126–
-
[141]
Synthesize Step-by-Step: Tools Templates and LLMs as Data Gener- ators for Reasoning-Based Chart VQA
Zhuowan Li, Bhavan Jasani, Peng Tang, and Shabnam Ghadar. Synthesize Step-by-Step: Tools Templates and LLMs as Data Gener- ators for Reasoning-Based Chart VQA. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13613–13623, 2024. 12
2024
-
[142]
12, 16, 21
Springer, 2024. 12, 16, 21
2024
-
[143]
OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Juntao Li, Xiaoye Qu, et al. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning. arXiv preprint arXiv:2505.08617, 2025. 12
2025 arXiv
-
[144]
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, Chengxiang Zhai, and Klara Nahrstedt. VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use. arXiv preprint arXiv:2505.19255, 2025. 12, 13
2025
-
[145]
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
Hao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua, and Shuicheng Yan. Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing. arXiv preprint arXiv:2412.19806, 2024. 12, 21
2024 arXiv
-
[147]
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models. Advances in Neural Information Processing Systems, 37:139348–139379, 2024. 13
2024
-
[148]
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
Syeda Nahida Akter, Aman Madaan, Sangwu Lee, Yiming Yang, and Eric Nyberg. Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination. ArXiv, abs/2401.08025, 2024. 13
2024 arXiv
-
[149]
LATTE: Learning to Reason with Vision Specialists
Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, et al. LATTE: Learning to Reason with Vision Specialists. In The 2025 Conference on Empirical Methods in Natural Language Processing, 2025. 13
2025
-
[150]
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al. CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models. In Proceedings of the Computer Vision and Pattern Recognition Confere...
2025
-
[151]
Creative Agents: Empowering Agents with Imagination for Creative Tasks
Penglin Cai, Chi Zhang, Yuhui Fu, Haoqi Yuan, and Zongqing Lu. Creative Agents: Empowering Agents with Imagination for Creative Tasks. In The 41st Conference on Uncertainty in Artificial Intelligence, 2025. 13
2025
-
[152]
OpenAI o1 System Card
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. OpenAI o1 System Card. arXiv preprint arXiv:2412.16720, 2024. 13, 14
2024 arXiv
-
[153]
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948,
-
[154]
LLaV A-CoT: Let Vision Language Models Reason Step-by- Step, 2025
Guowei Xu, Peng Jin, Hao Li, Yibing Song, Lichao Sun, and Li Yuan. LLaV A-CoT: Let Vision Language Models Reason Step-by- Step, 2025. 13
2025
-
[155]
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning
Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Minglan Lin, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang. Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning. arXiv preprint arXiv:2503.20752, 2025. 14, 21
2025
-
[156]
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models. arXiv preprint arXiv:2503.06749, 2025. 14, 19, 21
2025 arXiv
-
[157]
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300,
-
[158]
DeepPer- ception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding
Xinyu Ma, Ziyang Ding, Zhicong Luo, Chi Chen, Zonghao Guo, Derek F Wong, Xiaoyi Feng, and Maosong Sun. DeepPer- ception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding. arXiv preprint arXiv:2503.12797, 2025. 14 29
2025
-
[159]
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
Liang Chen, Hongcheng Gao, Tianyu Liu, Zhiqi Huang, Flood Sung, Xinyu Zhou, Yuxin Wu, and Baobao Chang. G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning. arXiv preprint arXiv:2505.13426 ,
-
[160]
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
Meng Cao, Haoze Zhao, Can Zhang, Xiaojun Chang, Ian Reid, and Xiaodan Liang. Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning. arXiv preprint arXiv:2505.20272, 2025. 14, 16, 19, 20
2025
-
[161]
On a Connection Between Imitation Learning and RLHF
Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, and Vasant G Honavar. On a Connection Between Imitation Learning and RLHF. arXiv preprint arXiv:2503.05079, 2025. 14
2025 arXiv
-
[162]
Beyond Reverse KL: Generalizing Direct Preference Opti- mization with Diverse Divergence Constraints
Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen. Beyond Reverse KL: Generalizing Direct Preference Opti- mization with Diverse Divergence Constraints. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[163]
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V Le, Sergey Levine, and Yi Ma. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training. arXiv preprint arXiv:2501.17161 ,
-
[164]
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition , pages 3195–3204, 2019. 14, 17, 18
2019
-
[165]
A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge. In European conference on computer vision, pages 146–162. Springer, 2022. 14, 17, 18
2022
-
[166]
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G Shapiro, and Ranjay Krishna. Perception Tokens Enhance Visual Reasoning in Multimodal Language Models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 3836–3845, 2025. 14
2025
-
[167]
Visually Interpretable Subtask Reasoning for Visual Question Answering
Yu Cheng, Arushi Goel, and Hakan Bilen. Visually Interpretable Subtask Reasoning for Visual Question Answering. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 2760–2780, 2025. 14
2025
-
[168]
CoReS: Orchestrating the Dance of Reasoning and Segmentation
Xiaoyi Bao, Siyang Sun, Shuailei Ma, Kecheng Zheng, Yuxin Guo, Guosheng Zhao, Yun Zheng, and Xingang Wang. CoReS: Orchestrating the Dance of Reasoning and Segmentation. In European Conference on Computer Vision, pages 187–204. Springer,
-
[169]
Visual Agents as Fast and Slow Thinkers
Guangyan Sun, Mingyu Jin, Zhenting Wang, Cheng-Long Wang, Siqi Ma, Qifan Wang, Tong Geng, Ying Nian Wu, Yongfeng Zhang, and Dongfang Liu. Visual Agents as Fast and Slow Thinkers. In The Thirteenth International Conference on Learning Representations, 2018. 15, 16, 19, 20, 21
2018
-
[170]
V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Penghao Wu and Saining Xie. V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094, 2024. 15, 16, 17, 18, 19, 21
2024
-
[171]
Divide, Conquer and Com- bine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, Wei Yu, and Dacheng Tao. Divide, Conquer and Com- bine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models. In Proceedings of the AAAI Conference on Artificial...
2025
-
[172]
Segment Anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexan- der C Berg, Wan-Yen Lo, et al. Segment Anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023. 15
2023
-
[173]
Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection. In European Conference on Computer Vision, pages 38–55. Springer...
2024
-
[174]
ZoomEye: En- hancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, and Jianwei Yin. ZoomEye: En- hancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. arXiv preprint arXiv:2411.16044, 2024. 15, 16, 19
2024 arXiv
-
[175]
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning. Advances in Neural Information Processing Systems, ...
2024
-
[176]
Insight-V: Exploring Long- Chain Visual Reasoning with Multimodal Large Language Models
Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. Insight-V: Exploring Long- Chain Visual Reasoning with Multimodal Large Language Models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 9062–9072, 2025. 15
2025
-
[177]
DeepEyes: Incentivizing ”Thinking with Images” via Reinforcement Learning
Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao, Guohai Xu, Le Yang, Chao Shen, and Xing Yu. DeepEyes: Incentivizing ”Thinking with Images” via Reinforcement Learning. ArXiv, abs/2505.14362, 2025. 15
2025 arXiv
-
[178]
GeReA: Question-Aware Prompt Captions for Knowledge- based Visual Question Answering
Ziyu Ma, Shutao Li, Bin Sun, Jianfei Cai, Zuxiang Long, and Fuyan Ma. GeReA: Question-Aware Prompt Captions for Knowledge- based Visual Question Answering. arXiv preprint arXiv:2402.02503, 2024. 15
2024 arXiv
-
[179]
Machine Mental Imagery: Empower Multimodal Reason- ing with Latent Visual Tokens
Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, and Chuang Gan. Machine Mental Imagery: Empower Multimodal Reason- ing with Latent Visual Tokens. ArXiv, abs/2506.17218, 2025. 16
2025 arXiv
-
[180]
From Foresight to Forethought: VLM-in-the-loop policy steering via latent alignment
Yilin Wu, Ran Tian, Gokul Swamy, and Andrea Bajcsy. From Foresight to Forethought: VLM-in-the-loop policy steering via latent alignment. In ICLR 2025 Workshop on World Models: Understanding, Modelling and Scaling, 2025. 16 30
2025
-
[181]
LISA: Reasoning Segmentation via Large Language Model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmentation via Large Language Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 9579–9589,
-
[182]
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Sca...
2024
-
[183]
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. VQA: Visual Question Answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015. 16, 21
2015
-
[184]
Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 16
2017
-
[185]
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedin...
2014
-
[186]
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International journal of computer ...
2017
-
[187]
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7W: Grounded Question Answering in Images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4995–5004, 2016. 16
2016
-
[188]
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 4971–4980, 2018. 16
2018
-
[189]
TallyQA: Answering Complex Counting Questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan. TallyQA: Answering Complex Counting Questions. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8076–8084, 2019. 16
2019
-
[190]
Cola: A Benchmark for Compositional Text-to-image Retrieval
Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna, and Kate Saenko. Cola: A Benchmark for Compositional Text-to-image Retrieval. Advances in Neural Information Processing Systems, 36:46433–46445, 2023. 16
2023
-
[191]
What Makes a Maze Look Like a Maze? In International Conference on Learning Representations (ICLR), 2025
Joy Hsu, Jiayuan Mao, Joshua B Tenenbaum, Noah D Goodman, and Jiajun Wu. What Makes a Maze Look Like a Maze? In International Conference on Learning Representations (ICLR), 2025. 16
2025
-
[192]
SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality. Advances in neural information processing systems, 36:31096–31116, 2023. 16, 21
2023
-
[193]
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
Raihan Kabir, Naznin Haque, Md Saiful Islam, et al. A Comprehensive Survey on Visual Question Answering Datasets and Algorithms. arXiv preprint arXiv:2411.11150, 2024. 17
2024 arXiv
-
[194]
Neural Module Networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural Module Networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39–48, 2016. 17
2016
-
[195]
Task Me Anything
Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. Task Me Anything. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 17
2024
-
[196]
Comparing Machines and Humans on a Visual Categorization Test
Franc ¸ois Fleuret, Ting Li, Charles Dubout, Emma K Wampler, Steven Yantis, and Donald Geman. Comparing Machines and Humans on a Visual Categorization Test. Proceedings of the National Academy of Sciences, 108:17621 – 17625, 2011. 17
2011
-
[197]
O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron C
Dzmitry Bahdanau, Harm de Vries, Timothy J. O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron C. Courville. CLOSURE: Assessing Systematic Generalization of CLEVR Models. ArXiv, abs/1912.05783, 2019. 17
1912 arXiv
-
[198]
CURI: A Benchmark for Productive Concept Learning Under Uncertainty
Ramakrishna Vedantam, Arthur Szlam, Maximillian Nickel, Ari Morcos, and Brenden M Lake. CURI: A Benchmark for Productive Concept Learning Under Uncertainty. In International Conference on Machine Learning, pages 10519–10529. PMLR, 2021. 17
2021
-
[199]
CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions
Runtao Liu, Chenxi Liu, Yutong Bai, and Alan L Yuille. CLEVR-Ref+: Diagnosing Visual Reasoning With Referring Expressions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4185–4194, 2019. 17
2019
-
[200]
CLEVR-XAI: A Benchmark Dataset for the Ground Truth Evaluation of Neural Network Explanations
Leila Arras, Ahmed Osman, and Wojciech Samek. CLEVR-XAI: A Benchmark Dataset for the Ground Truth Evaluation of Neural Network Explanations. Information Fusion, 81:14–40, 2022. 17
2022
-
[201]
QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning
Zechen Li and Anders Søgaard. QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning. arXiv preprint arXiv:2205.03075, 2022. 17
2022 arXiv
-
[202]
Sophia Koepke, Hendrik P
Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch, and Zeynep Akata. CLEVR-X: A Visual Reasoning Dataset for Natural Language Explanations. ArXiv, abs/2204.02380, 2022. 17
2022 arXiv
-
[203]
Satwik Kottur, Jos ´e M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. CLEVR-Dialog: A Diagnostic Dataset for Multi-Round Reasoning in Visual Dialog. In North American Chapter of the Association for Computational Linguistics, 2019. 17 31
2019
-
[204]
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. Multimodal Explanations: Justifying Decisions and Pointing to the Evidence. In Proceedings of the IEEE conference on computer vision and pattern recognition, pa...
2018
-
[205]
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions? Computer Vision and Image Understanding, 163:90–100, 2017
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions? Computer Vision and Image Understanding, 163:90–100, 2017. 17
2017
-
[206]
Cycle-Consistency for Robust Visual Question Answering
Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh. Cycle-Consistency for Robust Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6649–6658, 2019. 17
2019
-
[207]
Adam Santoro, Felix Hill, David G. T. Barrett, Ari S. Morcos, and Timothy P. Lillicrap. Measuring Abstract Reasoning in Neural Networks. ArXiv, abs/1807.04225, 2018. 17
2018 arXiv
-
[208]
KANDINSKYPatterns–An experimental exploration environment for Pattern Analysis and Machine Intelligence
Andreas Holzinger, Anna Saranti, and Heimo Mueller. KANDINSKYPatterns–An experimental exploration environment for Pattern Analysis and Machine Intelligence. arXiv preprint arXiv:2103.00519, 2021. 17
2021 arXiv
-
[209]
Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and Reasoning
Weili Nie, Zhiding Yu, Lei Mao, Ankit B Patel, Yuke Zhu, and Anima Anandkumar. Bongard-LOGO: A New Benchmark for Human-Level Concept Learning and Reasoning. Advances in Neural Information Processing Systems, 33:16468–16480, 2020. 17
2020
-
[210]
FVQA: Fact-Based Visual Question Answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. FVQA: Fact-Based Visual Question Answering. IEEE transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017. 17
2017
-
[211]
Explicit Knowledge-based Reasoning for Visual Question Answering
Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. Explicit Knowledge-based Reasoning for Visual Question Answering. arXiv preprint arXiv:1511.02570, 2015. 17
2015 arXiv
-
[212]
Shan, and Xilin Chen
Difei Gao, Ruiping Wang, S. Shan, and Xilin Chen. CRIC: A VQA Dataset for Compositional Reasoning on Vision and Common- sense. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:5561–5578, 2019. 17
2019
-
[213]
Towards Visual Dialog for Radiology
Olga Kovaleva, Chaitanya Shivade, Satyananda Kashyap, Karina Kanjaria, Joy Wu, Deddeh Ballah, Adam Coy, Alexandros Karar- gyris, Yufan Guo, David Beymer Beymer, et al. Towards Visual Dialog for Radiology. In Proceedings of the 19th SIGBioMed workshop on biomedical language pro...
2020
-
[214]
Scene Text Visual Question Answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marc ¸al Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene Text Visual Question Answering. InProceedings of the IEEE/CVF international conference on computer vision, pages 4291– 4301, 2019. 17, 18
2019
-
[215]
DocVQA: A Dataset for VQA on Document Images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. DocVQA: A Dataset for VQA on Document Images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 17
2021
-
[216]
A Diagram is Worth a Dozen Images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A Diagram is Worth a Dozen Images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 235–251. ...
2016
-
[217]
A Dataset of Clinically Generated Visual Questions and Answers About Radiology Images
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A Dataset of Clinically Generated Visual Questions and Answers About Radiology Images. Scientific data, 5(1):1–10, 2018. 18
2018
-
[218]
PathVQA: 30000+ Questions for Medical Visual Question Answering
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. PathVQA: 30000+ Questions for Medical Visual Question Answering. arXiv preprint arXiv:2003.10286, 2020. 18
2003 arXiv
-
[219]
VizWiz Grand Challenge: Answering Visual Questions From Blind People
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. VizWiz Grand Challenge: Answering Visual Questions From Blind People. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–36...
2018
-
[220]
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Yutong Wang, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, et al. WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Questi...
-
[221]
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Find- ings of the Association for Computati...
2022
-
[222]
Visual Question Answering on 360deg Images
Shih-Han Chou, Wei-Lun Chao, Wei-Sheng Lai, Min Sun, and Ming-Hsuan Yang. Visual Question Answering on 360deg Images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 1607–1616, 2020. 18
2020
-
[223]
Geoclidean: Few-Shot Generalization in Euclidean Geometry
Joy Hsu, Jiajun Wu, and Noah Goodman. Geoclidean: Few-Shot Generalization in Euclidean Geometry. Advances in Neural Information Processing Systems, 35:39007–39019, 2022. 18
2022
-
[224]
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. In The Twelfth International Conference on Learning...
2024
-
[225]
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024. 18
2024
-
[226]
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yux- uan Sun, et al. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. In Proceedings of the IEEE/CVF Conference on...
2024
-
[227]
SEED-Bench: Benchmarking Multimodal Large Language Models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. SEED-Bench: Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13299–13308, June 2024. 18
2024
-
[228]
MMBench: Is Your Multi-modal Model an All-Around Player? In European conference on computer vision, pages 216–233
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. MMBench: Is Your Multi-modal Model an All-Around Player? In European conference on computer vision, pages 216–233. Springer, 2024. 18
2024
-
[229]
Improved Baselines with Visual Instruction Tuning, 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved Baselines with Visual Instruction Tuning, 2023. 18
2023
-
[230]
Nov- Phy: A Physical Reasoning Benchmark for Open-World AI Systems
Vimukthini Pinto, Chathura Gamage, Cheng Xue, Peng Zhang, Ekaterina Nikonova, Matthew Stephenson, and Jochen Renz. Nov- Phy: A Physical Reasoning Benchmark for Open-World AI Systems. Artificial Intelligence, 336:104198, 2024. 18
2024
-
[231]
M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought. arXiv preprint arXiv:2405.16473, 2024. 18
2024 arXiv
-
[232]
VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use
Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, and Ludwig Schimdt. VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use. In Proceedings of the 37th International Conference...
2023
-
[233]
Evaluating Object Hallucination in Large Vision-Language Models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating Object Hallucination in Large Vision-Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 292–305, 2023. 18
2023
-
[234]
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al. HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models. In Proce...
2024
-
[235]
When’YES’Meets’ BUT’: Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? arXiv preprint arXiv:2503.23137, 2025
Tuo Liang, Zhe Hu, Jing Li, Hao Zhang, Yiren Lu, Yunlai Zhou, Yiran Qiao, Disheng Liu, Jeirui Peng, Jing Ma, et al. When’YES’Meets’ BUT’: Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? arXiv preprint arXiv:2503.23137, 2025. 18
2025 arXiv
-
[236]
Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing
Vedika Agarwal, Rakshith Shetty, and Mario Fritz. Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic Editing. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 9687–9695, 2019. 18
2020
-
[237]
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. Advances in Neural Information Processing Systems, 35:2507–25...
2022
-
[238]
Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 658–6...
2019
-
[239]
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: a Method for Automatic Evaluation of Machine Translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Com- putational Linguistics,...
2002
-
[240]
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. 18, 19
2004
-
[241]
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical ...
2021
-
[242]
OpenAI o3 and o4-mini System Cards, 2025
OpenAI. OpenAI o3 and o4-mini System Cards, 2025. System Cards for OpenAI’s o3 and o4-mini models. 18, 21
2025
-
[243]
Reasoning with Language Model is Planning with World Model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with Language Model is Planning with World Model. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 20
2023
-
[244]
Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Paul Saldyt, and Anil B Murthy. Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks. In Forty-first International Confer- ence on Machine Learning, 2024. 20
2024
-
[245]
Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts
Sukai Huang, Nir Lipovetzky, and Trevor Cohn. Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 26542–26550, 2025. 20
2025
-
[246]
The empirical case for two systems of reasoning
Steven A Sloman. The empirical case for two systems of reasoning. Psychological bulletin, 119(1):3, 1996
1996
-
[247]
PromptA- gent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization
Xinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai, Haotian Luo, Jiayou Zhang, Nebojsa Jojic, Eric Xing, and Zhiting Hu. PromptA- gent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization. InThe Twelfth International Conference on Learning Representations...
2024
-
[248]
Survey on Evaluation of LLM-based Agents
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. Survey on Evaluation of LLM-based Agents. arXiv preprint arXiv:2503.16416, 2025. 20
2025 arXiv
-
[249]
LASP: Surveying the State-of-the-Art in Large Language Model- Assisted AI Planning
Haoming Li, Zhaoliang Chen, Jonathan Zhang, and Fei Liu. LASP: Surveying the State-of-the-Art in Large Language Model- Assisted AI Planning. arXiv preprint arXiv:2409.01806, 2024. 20
2024 arXiv
-
[250]
Deductive Reasoning
Philip N Johnson-Laird. Deductive Reasoning. Annual review of psychology, 50(1):109–135, 1999. 20
1999
-
[251]
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merri ¨enboer, Armand Joulin, and Tomas Mikolov. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks. arXiv preprint arXiv:1502.05698, 2015. 20
2015 arXiv
-
[252]
Natural Language Reasoning, A Survey
Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. Natural Language Reasoning, A Survey. ACM Computing Surveys, 56(12):1–39, 2024. 20
2024
-
[253]
Introspective Learning : A Two-Stage approach for Inference in Neural Networks
Mohit Prabhushankar and Ghassan AlRegib. Introspective Learning : A Two-Stage approach for Inference in Neural Networks. Advances in Neural Information Processing Systems, 35:12126–12140, 2022. 20
2022
-
[254]
A Survey on Neural-Symbolic Learning Systems
Dongran Yu, Bo Yang, Da Liu, Hui Wang, and Shirui Pan. A Survey on Neural-Symbolic Learning Systems. Neural networks : the official journal of the International Neural Network Society, 166:105–126, 2021. 20
2021
-
[255]
Neural Analogical Matching
Maxwell Crouse, Constantine Nakos, Ibrahim Abdelaziz, and Ken Forbus. Neural Analogical Matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 809–817, 2021. 21
2021
-
[256]
van Rooij
Mark Blokpoel, Todd Wareham, W.F.G Pim Haselager, Ivan Toni, and I.J.E.I. van Rooij. Deep Analogical Inference as the Origin of Hypotheses. J. Probl. Solving, 11, 2019. 21
2019
-
[257]
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.Advances in neural information processing systems, 37:32731–32760, 2024
Maitreya Patel, Naga Sai Abhiram Kusumba, Sheng Cheng, Changhoon Kim, Tejas Gokhale, Chitta Baral, et al. TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.Advances in neural information processing systems, 37:32731–32760, 2024. 21
2024
-
[258]
COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning
Xindi Wu, Hee Seung Hwang, Polina Kirichenko, and Olga Russakovsky. COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning. In Synthetic Data for Computer Vision Workshop@ CVPR 2025, 2025. 21
2025
-
[259]
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
Jieyu Zhang, Le Xue, Linxin Song, Jun Wang, Weikai Huang, Manli Shu, An Yan, Zixian Ma, Juan Carlos Niebles, Silvio Savarese, et al. ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models. arXiv preprint arXiv:2412.07012, 2024. 21
2024 arXiv
-
[260]
JRDB-PanoTrack: An Open- world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
Duy Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi, Ian Reid, Jianfei Cai, and Hamid Rezatofighi. JRDB-PanoTrack: An Open- world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2024
-
[261]
JRDB-Act: A Large-Scale Dataset for Spatio- Temporal Action, Social Group and Activity Detection
Mahsa Ehsanpour, Fatemeh Saleh, Silvio Savarese, Ian Reid, and Hamid Rezatofighi. JRDB-Act: A Large-Scale Dataset for Spatio- Temporal Action, Social Group and Activity Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20983...
2022
-
[262]
m & m’s: A Benchmark to Evaluate Tool-Use for m ulti-step m ulti-modal Tasks
Zixian Ma, Weikai Huang, Jieyu Zhang, Tanmay Gupta, and Ranjay Krishna. m & m’s: A Benchmark to Evaluate Tool-Use for m ulti-step m ulti-modal Tasks. In European Conference on Computer Vision, pages 18–34. Springer, 2024. 21
2024
-
[263]
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou. AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn. arXiv preprint arXiv:2306.08640, 2023. 21
2023 arXiv
-
[264]
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition
Xinyu Xi, Hua Yang, Shentai Zhang, Yijie Liu, Sijin Sun, and Xiuju Fu. Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition. arXiv preprint arXiv:2503.06978, 2025. 21
2025 arXiv
-
[265]
Deciphering the Role of Representation Disentan- glement: Investigating Compositional Generalization in CLIP Models
Reza Abbasi, Mohammad Hossein Rohban, and Mahdieh Soleymani Baghshah. Deciphering the Role of Representation Disentan- glement: Investigating Compositional Generalization in CLIP Models. In European Conference on Computer Vision, pages 35–50. Springer, 2024. 21
2024
-
[266]
VISCO: Benchmarking Fine- Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
Xueqing Wu, Yuheng Ding, Bingxuan Li, Pan Lu, Da Yin, Kai-Wei Chang, and Nanyun Peng. VISCO: Benchmarking Fine- Grained Critique and Correction Towards Self-Improvement in Visual Reasoning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9527–95...
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.