REVIEW 2 major objections 3 minor 22 cited by
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
T0 review · 2 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Multimodal reasoning has evolved through four stages and now heads toward the Native Large Multimodal Reasoning Model (N-LMRM), where reasoning emerges from omni-modal perception and goal-driven interaction instead of being retrofitted…
desk verdict A useful, up-to-date survey with a forward-looking agenda whose empirical motivation is anecdotal; worth refereeing, but the N-LMRM case needs to be reframed as a hypothesis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing construct is the four-stage developmental roadmap, which doubles as a taxonomy and as an argument. Each stage is defined by where reasoning resides: (1) modular reasoning networks and pretrained vision-language models, where reasoning is implicit in representation, alignment, and fusion; (2) language-centric System-1 reasoning, where short chains emerge from prompt-based MCoT, structural reasoning, and externally augmented reasoning; (3) language-centric System-2 reasoning, where long chains and planning come from cross-modal reasoning, O1-style models, and R1-style reinforcement learning; and (4) N-LMRMs, defined by two capabilities — Multimodal Agentic Reasoning (hierarchical planning, dynamic adaptation, embodied learning) and Omni-Modal Understanding and Generative Reasoning (unified representations, cross-modal synthesis, modality-agnostic inference). The roadmap does the argumentative work: by locating today's best models in Stage 3 and presenting Section 4.1's benchmark failures and o3/o4-mini case studies as symptoms of a language-centric ceiling, the survey converts a design preference into a diagnosed gap that Stage 4 is positioned to fill.
What would settle it
Run each of the paper's three o3/o4-mini failure modes — six-finger emoji counting, phone-number extraction from resume PDFs, and red-panda multimedia generation — on a systematic sample of roughly one hundred analogous instances per case, comparing the same models against a language-centric open model and a native omni-modal model matched for compute. If the language-centric models fail at about the same rate as the native model, or if the failures disappear once browsing tools and internet access are provided, the claim that language-centric architecture is the bottleneck would be refuted. A weaker decisive check: test whether accuracy on omni-modal benchmarks falls as the share of non-textual tokens inside the model's chain of thought rises.
Extended reading notes
Core claim
The paper's central claim is that where reasoning lives in the architecture defines the era of multimodal AI. In Stage 1, reasoning was implicit, distributed across task-specific modules for representation, alignment, and fusion. In Stage 2, reasoning became explicit but shallow: language-centric models produced short, reactive chains through prompt-based Multimodal Chain-of-Thought (MCoT), structural reasoning, and external augmentation. In Stage 3, chains lengthened into deliberate System-2 thinking through cross-modal reasoning, O1-style long reasoning, and reinforcement learning (DPO and GRPO), typified by the R1 line of models. Drawing on omni-modal and agentic benchmarks and on hands-on case studies of the o3 and o4-mini models — six-finger emoji counting, resume PDF parsing, puzzle solving, multimedia generation — the paper argues that the language-centric paradigm cannot reach real-world utility, and introduces the Native Large Multimodal Reasoning Model (N-LMRM): a forward-looking architecture where reasoning natively emerges from omnimodal perception and interaction and from goal-driven cognition, combining Multimodal Agentic Reasoning with Omni-Modal Understanding and Generative Reasoning.
Load-bearing premise
The argument for replacing language-centric models rests on a small set of hand-picked failure examples (finger counting, resume PDF parsing, a puzzle) and aggregate benchmark scores, presented without a sampling protocol or a matched baseline; if those failures actually reflect task difficulty or missing tools rather than the language-centric design, the case for native multimodal reasoning models loses most of its force.
Editorial extensions
If this is right
- Future flagship multimodal systems would abandon the vision-encoder-plus-LLM design for unified representation spaces that treat text, image, audio, video, and sensor streams symmetrically.
- Long-horizon agentic behavior — GUI navigation, embodied interaction, multi-tool chains — would become a core training objective rather than an outer wrapper, with reinforcement learning scaled across modalities.
- Evaluation would shift from static question answering (MMMU, MathVista) toward omni-modal and interactive benchmarks, because the paper argues current benchmarks overstate real-world capability.
- Interleaved multimodal chain-of-thought — reasoning traces that include image crops, audio snippets, or actions — would open a new axis of test-time compute scaling.
- Reasoning training (O1/R1-style) would extend beyond text-heavy math and vision tasks to cross-modal generation and planning, where the model not only thinks but generates intermediate multimodal content.
Reading between the lines
- A controlled test the paper does not report: a native omni-modal model matched against a language-centric LMRM on the same data, compute, and RL budget, on the same omni-modal and agentic tasks; the roadmap predicts the native model wins, and that experiment would settle it.
- The case studies imply a strong testable corollary: fabricated reasoning ('lying' rationales attached to correct answers) is a symptom of language-centric post-training and would shrink when reasoning is grounded in perceptual tokens.
- Tool-augmented and search-augmented systems (deep-research-style agents) may reach N-LMRM-level task performance without a native architecture, which would weaken the necessity claim; the paper leaves this stopgap undiscussed as a rival.
- If the roadmap generalizes, a similar staged trajectory should appear in audio-centric and embodied subfields — early modular perception, then language-mediated short reasoning, then long-horizon System-2 behavior — a pattern that could be checked against the literature the survey catalogs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of large multimodal reasoning models (LMRMs). It proposes a developmental roadmap in which multimodal reasoning evolves from perception-driven modular systems, through language-centric short reasoning (System-1) and language-centric long reasoning (System-2), toward a prospective fourth stage called Native Large Multimodal Reasoning Models (N-LMRMs). The survey covers roughly 700 publications, organizes methods into detailed taxonomies, reviews multimodal RL-enhanced reasoning (O1-like and R1-like models), and catalogs datasets and benchmarks for understanding, generation, reasoning, and planning. The forward-looking section motivates N-LMRMs with benchmark aggregates and case studies of OpenAI o3 and o4-mini, and outlines capabilities and technical prospects for such models.
Significance. The survey is timely and broad, and its organization around a staged roadmap is a useful contribution to a rapidly growing field. The benchmark and dataset reorganization in Section 5, as well as the extensive tables of recent RL-based multimodal reasoning methods, provide a valuable reference. The N-LMRM concept is an interesting forward-looking synthesis that connects omni-modal understanding, generation, and agentic reasoning. However, the empirical support for the N-LMRM proposal is anecdotal, and several internal inconsistencies affect the clarity of the central contribution. The paper is likely to be useful to practitioners and researchers, but the forward-looking claims need to be reframed or supported more rigorously.
major comments (2)
- [Section 4 and Section 4.1] The opening of Section 4 states that 'language-centric architectures impose critical constraints,' and Section 4.2 says N-LMRMs are introduced 'based on the above experimental findings.' This is a load-bearing causal claim, but Section 4.1 does not provide sufficient evidence to support it. Table 12 lists benchmarks without a protocol, per-model score tables, or controls for task difficulty; for example, BrowseComp is a text-only web-browsing benchmark, so the reported GPT-4o accuracy of 0.6% cannot demonstrate a multimodal-specific deficit. The o3/o4-mini case studies in Figures 6-8 are hand-picked examples with no sample sizes, selection criteria, or baselines, and the observed failure modes (finger counting, PDF parsing, puzzle rationalization) could plausibly persist in any architecture, not specifically because of a language-centric design. The paper should either add systematic evidence, such as controlled comparisons across model families and modality conditions, or explicitly reframe the N-LMRM proposal as a hypothesis motivated by observed limitations rather than a conclusion established by these experiments. Additionally, Section 4.1 is titled 'Preliminary Study with o3 and o4-mini,' but only o3 results are reported; no o4-mini experiments appear in the text, despite the abstract mentioning 'experimental cases of OpenAI O3 and O4-mini.'
- [Abstract, Section 1, Section 2, and Section 3] The number of stages in the proposed roadmap is inconsistent across the paper. The abstract says the survey is 'organized around a four-stage developmental roadmap,' and Section 2 says 'we outline four key stages.' However, Section 1 states the roadmap is 'organized into three stages (Figure 2),' and the contributions list in Section 1 describes a 'three-stage roadmap.' Since the staged roadmap is the paper's central organizational contribution, this inconsistency is not merely cosmetic. The authors should use one consistent count throughout, for example by making explicit that Stages 1-3 are historical stages and Stage 4 is a prospective direction, and then aligning the abstract, introduction, Section 2, and the roadmap figure accordingly.
minor comments (3)
- [Section 5.1.1, Section 3.1.2, Table 14] In Section 5.1.1, 'GQA' is cited as 'Ainslie et al., 2023,' but that reference is for 'GQA: Training Generalized Multi-Query Transformer Models,' not the visual question answering dataset. The correct citation is Hudson & Manning (2019), which does appear correctly in Section 3.1.2 and Table 14. The authors should correct this citation error throughout the manuscript.
- [Section 4.1] The heading 'Preliminary Study with o3 and o4-mini' promises results for both models, but the text only reports evaluations of o3. Either add o4-mini results or change the heading and the abstract's wording to refer only to o3.
- [Several sections] There are multiple typographical errors that should be fixed: 'Planing' instead of 'Planning' in Section 5's introductory paragraph, 'Omini-Modal' instead of 'Omni-Modal' in the Conclusion, 'syatem-2' instead of 'system-2' in the Section 3.3.3 takeaways, 'Trasnformer' instead of 'Transformer' in Section 3.1.1, and 'Operater' instead of 'Operator' in Section 4.2.
Circularity Check
No significant circularity: the survey's roadmap and N-LMRM proposal are organizational and motivational, not a derivation whose output equals its input by construction.
full rationale
This paper is a survey and taxonomy. Its central contributions are a four-stage developmental roadmap for multimodal reasoning and a forward-looking concept, N-LMRMs, defined as models where reasoning 'natively emerges from omnimodal perception and interaction, and goal-driven cognition'. Neither contribution is obtained by fitting a parameter to data and then predicting that same data, nor by defining a term in terms of the conclusion it is meant to support. The N-LMRM motivation is grounded in external benchmark results (e.g., GPT-4o's 0.6% on BrowseComp, Claude 3.5 Sonnet's 35% on WorldSense) and a few o3/o4-mini case studies; those examples may be non-representative or insufficiently controlled, but that is an evidence-quality concern, not circularity, because the paper does not fit anything to those examples nor define N-LMRMs as 'the class of models that pass these anecdotes'. The assertion that 'language-centric architectures impose critical constraints' is supported by external citations (Kumar et al., 2025; Pfister & Jud, 2025) and functions as an interpretive hypothesis rather than a theorem derived from the authors' own prior work. There is no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The stage-count inconsistency (four-stage in the abstract and Section 2 versus three-stage in Section 1) is an internal-consistency flaw, not circular reasoning. On the evidence available, the survey's derivational chain is self-contained and non-circular.
Assumptions & free parameters
assumptions (3)
- domain assumption The true development of multimodal reasoning is captured by the proposed three/four-stage roadmap.
- domain assumption The benchmarks and case studies in Section 4.1 are representative of real-world LMRM limitations.
- domain assumption Language-centric architecture is the main bottleneck preventing deeper multimodal reasoning.
invented entities (1)
-
Native Large Multimodal Reasoning Model (N-LMRM)
Cite this review
Pith. "Pith review of Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models." pith.science (2026). https://pith.science/paper/SNBHDPWV
@misc{pith2026250504921,
author = {Pith},
title = {Pith review of: Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/SNBHDPWV}},
note = {Machine review of arXiv:2505.04921}
}
read the original abstract
Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasingly operate in open, uncertain, and multimodal environments, reasoning becomes essential for enabling robust and adaptive behavior. Large Multimodal Reasoning Models (LMRMs) have emerged as a promising paradigm, integrating modalities such as text, images, audio, and video to support complex reasoning capabilities and aiming to achieve comprehensive perception, precise understanding, and deep reasoning. As research advances, multimodal reasoning has rapidly evolved from modular, perception-driven pipelines to unified, language-centric frameworks that offer more coherent cross-modal understanding. While instruction tuning and reinforcement learning have improved model reasoning, significant challenges remain in omni-modal generalization, reasoning depth, and agentic behavior. To address these issues, we present a comprehensive and structured survey of multimodal reasoning research, organized around a four-stage developmental roadmap that reflects the field's shifting design philosophies and emerging capabilities. First, we review early efforts based on task-specific modules, where reasoning was implicitly embedded across stages of representation, alignment, and fusion. Next, we examine recent approaches that unify reasoning into multimodal LLMs, with advances such as Multimodal Chain-of-Thought (MCoT) and multimodal reinforcement learning enabling richer and more structured reasoning chains. Finally, drawing on empirical insights from challenging benchmarks and experimental cases of OpenAI O3 and O4-mini, we discuss the conceptual direction of native large multimodal reasoning models (N-LMRMs), which aim to support scalable, agentic, and adaptive reasoning and planning in complex, real-world environments.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 22 Pith papers
-
Where, What, Why: Towards Explainable Driver Attention Prediction
W3DA adds semantic and causal labels to four driver gaze datasets, and the LLada model predicts attention maps, attended semantics, and reasons in one end-to-end system.
-
Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
UAV-DualCog evaluates multimodal LLMs on self-state and environment-state reasoning in aerial images and videos, and shows current models are unreliable at spatial and temporal grounding.
-
Mixture of Cognitive Experts in Large Vision-Language Models
Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.
-
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...
-
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
A confidence-based when-to-look gate plus semantic-guided attention where-to-look module improves training-free fine-grained visual reasoning accuracy and reduces inference cost versus search-based baselines.
-
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Argos is an agentic verifier that adaptively picks scoring functions to evaluate accuracy, localization, and reasoning quality, enabling stronger multimodal RL training for AI agents.
-
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
VL-Cogito, trained with progressive curriculum RL, online difficulty weighting, and dynamic length rewards, matches or beats prior reasoning MLLMs on ten multimodal benchmarks.
-
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
BMMR provides a 110k-question bilingual, multimodal, college-level dataset across 300 subjects where state-of-the-art models score at most about 50%.
-
MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?
MATP-BENCH pairs 1,056 multimodal math problems with formal theorem statements in Lean 4, Coq, and Isabelle; the strongest tested model solves only 5.68% of Lean 4 end-to-end proving tasks at pass@10.
-
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
EffiVLM-Bench is a benchmark study showing token compression is task- and model-dependent, KV cache methods are more loyal, and parameter compression preserves accuracy better at typical ratios.
-
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
VisCRA jailbreaks multimodal LLMs by masking the most harmful image region and using a two-stage reasoning prompt to make the model infer and then comply.
-
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
VerIPO interleaves GRPO, a verifier that curates preference pairs from rollouts, and DPO to steadily improve accuracy and chain-of-thought consistency in video LLMs.
-
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
EarthSE provides a two-level QA benchmark and an open-ended dialogue benchmark for Earth science and shows current LLMs perform poorly on both.
-
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
DC-Leap accelerates diffusion LLM decoding by verifying contiguous token spans at a lower confidence threshold and using high-confidence future drafts as look-ahead context, achieving up to 53x speedup with comparable...
-
Explain Before You Answer: A Survey on Compositional Visual Reasoning
A survey that classifies compositional visual reasoning methods into five stages, from prompt-based pipelines to unified agentic vision-language models, and catalogs associated benchmarks.
-
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
An LLM-powered multi-agent Copilot retrieves and constructs ComfyUI workflows, reporting at least 88.5% recall on its own test set and 85.9% online acceptance of proposed workflows.
-
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.
-
AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives
A survey that organizes LLM and AI reasoning methods into a taxonomy and maps them onto the physical, link, network, transport, and application layers of wireless networks.
-
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.
-
Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
A survey that organizes agentic AI for 6G edge networks into four pillars, compactness, efficiency, knowledge and reasoning, and migration, and illustrates them with prior case studies.
-
Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
A survey of plane geometry problem solving that classifies methods into an encoder-decoder framework and analyzes hallucination and data leakage in current benchmarks.
-
Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?
A position paper arguing that vision foundation models need dynamic relational graphs for relational reasoning, with evidence drawn from the author's own prior action recognition and tumor segmentation systems.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Youtube-8m: A large-scale video classification benchmark, 2016
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark, 2016. URL https://arxiv.org/abs/1609.08675
arXiv 2016
-
[3]
A vision centric remote sensing benchmark, 2025
Abduljaleel Adejumo, Faegheh Yeganli, Clifford Broni-bediako, Aoran Xiao, Naoto Yokoya, and Mennatullah Siam. A vision centric remote sensing benchmark, 2025. URL https://arxiv.org/abs/2503.15816
arXiv 2025
-
[4]
Denk, Zal \' a n Borsos, Jesse H
Andrea Agostinelli, Timo I. Denk, Zal \' a n Borsos, Jesse H. Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matthew Sharifi, Neil Zeghidour, and Christian Havn Frank. Musiclm: Generating music from text. CoRR, abs/2301.11325, 2023. doi:10.48550/ARXIV.2301.11325. URL https://doi.org/10.48550/arXiv.2301.11325
-
[5]
Ming-omni: A unified multimodal model for perception and generation, 2025
Inclusion AI, Biao Gong, Cheng Zou, Chuanyang Zheng, Chunluan Zhou, Canxiang Yan, Chunxiang Jin, Chunjie Shen, Dandan Zheng, Fudong Wang, Furong Xu, GuangMing Yao, Jun Zhou, Jingdong Chen, Jianxin Sun, Jiajia Liu, Jianjiang Zhu, Jun Peng, Kaixiang Ji, Kaiyou Song, Kaimeng Ren, Libin Wang, Lixiang Ru, Lele Xie, Longhua Tan, Lyuxin Xue, Lan Wang, Mochen Bai...
arXiv 2025
-
[6]
GQA: training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee - Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebr \' o n, and Sumit Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapo...
-
[7]
SBVQA 2.0 : Robust end-to-end speech-based visual question answering for open-ended questions
Faris Alasmary and Saad Al-Ahmadi. SBVQA 2.0 : Robust end-to-end speech-based visual question answering for open-ended questions. IEEE Access, 11: 0 140967--140980, 2023. doi:10.1109/ACCESS.2023.3339537
arXiv 2023
-
[8]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. ArXiv preprint, abs/2204.14198, 2022. URL https://arxiv.org/abs/2204.14198
arXiv 2022
Show all 298 references
-
[9]
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel - Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019...
2019
-
[10]
OpenLEAF : A novel benchmark for open-domain interleaved image-text generation
Jie An, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Lijuan Wang, and Jiebo Luo. OpenLEAF : A novel benchmark for open-domain interleaved image-text generation. In Proceedings of the 32nd ACM International Conference on Multimedia (MM'24), pp.\ 11137--1114...
2024
-
[11]
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 607...
2018
-
[12]
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 39--48, 2016
2016
-
[13]
Introducing the model context protocol, April 2025
Anthropic. Introducing the model context protocol, April 2025. URL https://www.anthropic.com/news/model-context-protocol. Anthropic News. Accessed: 2025-04-17
2025
-
[14]
Lawrence Zitnick, and Devi Parikh
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA : V isual Q uestion A nswering. In International Conference on Computer Vision (ICCV), 2015
2015
-
[15]
Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu. Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M...
2024
-
[16]
Tyers, and Gregor Weber
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. In Nicoletta Calzolari, Fr \' e d \' e ric B \' e chet, Philippe ...
2020
-
[17]
Genesis: A universal and generative physics engine for robotics and beyond, 2024
Genesis Authors. Genesis: A universal and generative physics engine for robotics and beyond, 2024
2024
-
[18]
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023. URL https://arxiv.org/abs/2308.12966
2023 arXiv
-
[19]
Univg-r1: Reasoning guided universal visual grounding with reinforcement learning, 2025
Sule Bai, Mingxing Li, Yong Liu, Jing Tang, Haoji Zhang, Lei Sun, Xiangxiang Chu, and Yansong Tang. Univg-r1: Reasoning guided universal visual grounding with reinforcement learning, 2025. URL https://arxiv.org/abs/2505.14231
2025 arXiv
-
[20]
Frozen in time: A joint video and image encoder for end-to-end retrieval, 2022
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval, 2022. URL https://arxiv.org/abs/2104.00650
2022 arXiv
-
[21]
Videophy: Evaluating physical commonsense for video generation
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai - Wei Chang, and Aditya Grover. Videophy: Evaluating physical commonsense for video generation. CoRR, abs/2406.03520, 2024. doi:10.48550/ARXIV.2406.03520. URL htt...
-
[22]
Beit: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. Beit: BERT pre-training of image transformers. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net, 2022 a . URL https://openreview.net/forum?id=p-BhZSz59o4
2022
-
[23]
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. Advances in Neural Information Processing Systems, 35: 0 32897--32912, 2022 b
2022
-
[24]
Aquallm: audio question answering data generation using large language models
Swarup Ranjan Behera, Krishna Mohan Injeti, Jaya Sai Kiran Patibandla, Praveen Kumar Pokala, and Balakrishna Reddy Pailla. Aquallm: audio question answering data generation using large language models. arXiv preprint arXiv:2312.17343, 2023
2023 arXiv
-
[25]
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washing...
2013
-
[26]
Why reasoning matters? a survey of advancements in multimodal reasoning (v1)
Jing Bi, Susan Liang, Xiaofei Zhou, Pinxin Liu, Junjia Guo, Yunlong Tang, Luchuan Song, Chao Huang, Guangyu Sun, Jinxi He, et al. Why reasoning matters? a survey of advancements in multimodal reasoning (v1). arXiv preprint arXiv:2504.03151, 2025
2025
-
[27]
Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks
L \'e o Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault de Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks. Advances in Neural Info...
2024
-
[28]
Windows agent arena: Evaluating multi-modal OS agents at scale
Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, and Zack Hui. Windows agent arena: Evaluating multi-modal OS agents at scale. CoRR, abs/2409.08264, 2024. doi:10.4855...
-
[29]
Clime: Evaluating multimodal climate discourse on social media and the climate alignment quotient (caq), 2025
Abhilekh Borah, Hasnat Md Abdullah, Kangda Wei, and Ruihong Huang. Clime: Evaluating multimodal climate discourse on social media and the climate alignment quotient (caq), 2025. URL https://arxiv.org/abs/2504.03906
2025 arXiv
-
[30]
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , pp.\ 18392--18402. IEEE , 2023. doi:10....
2023
-
[31]
AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng. AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline. In 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Asses...
2017
-
[32]
Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M
James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano, Sanket Rajan Gupte, Jesus G. Galaz-Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M. Hasan, Alexandra Johannesson, William D. Leineweber, Malvika G Nair, Ridhi Yarlagadda, Connor Z...
2025 arXiv
-
[33]
Murel: Multimodal relational reasoning for visual question answering
Remi Cadene, Hedi Ben-Younes, Matthieu Cord, and Nicolas Thome. Murel: Multimodal relational reasoning for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 1989--1998, 2019
1989
-
[34]
Coco-stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. arXiv preprint arXiv:1612.03716, 2016
2016 arXiv
-
[35]
Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models, 2025
Huanqia Cai, Yijun Yang, and Winston Hu. Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models, 2025
2025
-
[36]
AMEX: android multi-annotation expo dataset for mobile GUI agents
Yuxiang Chai, Siyuan Huang, Yazhe Niu, Han Xiao, Liang Liu, Dingyu Zhang, Peng Gao, Shuai Ren, and Hongsheng Li. AMEX: android multi-annotation expo dataset for mobile GUI agents. CoRR, abs/2407.17490, 2024. doi:10.48550/ARXIV.2407.17490. URL https://doi.org/10.48550/arXiv.2407.17490
-
[37]
Chateval: Towards better LLM -based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. Chateval: Towards better LLM -based evaluators through multi-agent debate. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.ne...
2024
-
[38]
Extremeaigc: Benchmarking lmm vulnerability to ai-generated extremist content, 2025
Bhavik Chandna, Mariam Aboujenane, and Usman Naseem. Extremeaigc: Benchmarking lmm vulnerability to ai-generated extremist content, 2025. URL https://arxiv.org/abs/2503.09964
2025 arXiv
-
[39]
A survey of data synthesis approaches
Hsin-Yu Chang, Pei-Yu Chen, Tun-Hsiang Chou, Chang-Sheng Kao, Hsuan-Yun Yu, Yen-Ting Lin, and Yun-Nung Chen. A survey of data synthesis approaches. arXiv preprint arXiv:2407.03672, 2024
2024 arXiv
-
[40]
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 3558--3568, 2021
2021
-
[41]
Mllm-as-a-judge: assessing multimodal llm-as-a-judge with vision-language benchmark
Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. Mllm-as-a-judge: assessing multimodal llm-as-a-judge with vision-language benchmark. In Proceedings of the 41st International Conference on Machine Le...
2024
-
[42]
GUI-WORLD: A dataset for gui-oriented multimodal llm-based agents
Dongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang, Liuyi Chen, Yilin Bai, Zhigang He, Chenlong Wang, Huichi Zhou, Yiqiang Li, Tianshuo Zhou, Yue Yu, Chujie Gao, Qihui Zhang, Yi Gui, Zhen Li, Yao Wan, Pan Zhou, Jianfeng Gao, and Lichao Sun. GUI-WORLD: A dataset for gui-oriented ...
-
[43]
Mathflow: Enhancing the perceptual flow of mllms for visual mathematical problems, 2025 a
Felix Chen, Hangjie Yuan, Yunqiu Xu, Tao Feng, Jun Cen, Pengwei Liu, Zeying Huang, and Yi Yang. Mathflow: Enhancing the perceptual flow of mllms for visual mathematical problems, 2025 a . URL https://arxiv.org/abs/2503.16549
2025 arXiv
-
[44]
Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos
Houlun Chen, Xin Wang, Hong Chen, Zihan Song, Jia Jia, and Wenwu Zhu. Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos. arXiv preprint arXiv:2312.17117, 2023 a
2023 arXiv
-
[45]
Xing, and Liang Lin
Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric P. Xing, and Liang Lin. Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning, 2022 a . URL https://arxiv.org/abs/2105.14517
2022 arXiv
-
[46]
Spa-bench: A comprehensive benchmark for smartphone agent evaluation
Jingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang, Gongwei Chen, Zhihao Wu, Li Yixing, Xurui Zhou, Weiwen Liu, Shuai Wang, et al. Spa-bench: A comprehensive benchmark for smartphone agent evaluation. In NeurIPS 2024 Workshop on Open-World Agents, 2024 c
2024
-
[47]
R e C oncile: Round-table conference improves reasoning via consensus among diverse LLM s
Justin Chen, Swarnadeep Saha, and Mohit Bansal. R e C oncile: Round-table conference improves reasoning via consensus among diverse LLM s. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Lingu...
2024 doi
-
[48]
Shikra: Unleashing multimodal llm's referential dialogue magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm's referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023 b
2023 arXiv
-
[49]
G1: Bootstrapping perception and reasoning abilities of vision-language model via reinforcement learning, 2025 b
Liang Chen, Hongcheng Gao, Tianyu Liu, Zhiqi Huang, Flood Sung, Xinyu Zhou, Yuxin Wu, and Baobao Chang. G1: Bootstrapping perception and reasoning abilities of vision-language model via reinforcement learning, 2025 b . URL https://arxiv.org/abs/2505.13426
2025 arXiv
-
[50]
R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3
Liang Chen, Lei Li, Haozhe Zhao, Yifan Song, and Vinci. R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3. https://github.com/Deep-Agent/R1-V, 2025 c . Accessed: 2025-02-02
2025
-
[51]
Omnixr: Evaluating omni-modality language models on reasoning across modalities
Lichang Chen, Hexiang Hu, Mingda Zhang, Yiwen Chen, Zifeng Wang, Yandong Li, Pranav Shyam, Tianyi Zhou, Heng Huang, Ming-Hsuan Yang, et al. Omnixr: Evaluating omni-modality language models on reasoning across modalities. arXiv preprint arXiv:2410.12219, 2024 e
-
[52]
Are we on the right way for evaluating large vision-language models? In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao. Are we on the right way for evaluating large vision-language models? In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paq...
2024
-
[53]
Websrc: A dataset for web-based structural reading comprehension
Lu Chen, Xingyu Chen, Zihan Zhao, Danyang Zhang, Jiabao Ji, Ao Luo, Yuxuan Xiong, and Kai Yu. Websrc: A dataset for web-based structural reading comprehension. CoRR, abs/2101.09465, 2021. URL https://arxiv.org/abs/2101.09465
2021 arXiv
-
[54]
Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective
Meiqi Chen, Yixin Cao, Yan Zhang, and Chaochao Lu. Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Findings of the Association for Computational Linguistics: EMNL...
2024 doi
-
[55]
Rbf++: Quantifying and optimizing reasoning boundaries across measurable and unmeasurable capabilities for chain-of-thought reasoning, 2025 d
Qiguang Chen, Libo Qin, Jinhao Liu, Yue Liao, Jiaqi Wang, Jingxuan Zhou, and Wanxiang Che. Rbf++: Quantifying and optimizing reasoning boundaries across measurable and unmeasurable capabilities for chain-of-thought reasoning, 2025 d . URL https://arxiv.org/abs/2505.13307
2025 arXiv
-
[56]
Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning
Shuang Chen, Yue Guo, Zhaochen Su, Yafu Li, Yulun Wu, Jiacheng Chen, Jiayu Chen, Weijie Wang, Xiaoye Qu, and Yu Cheng. Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning. arXiv preprint arXiv:2506.04207, 2025 e
2025
-
[57]
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022 b
2022 arXiv
-
[58]
Chart-hqa: A benchmark for hypothetical question answering in charts, 2025 f
Xiangnan Chen, Yuancheng Fang, Qian Xiao, Juncheng Li, Jun Lin, Siliang Tang, Yi Yang, and Yueting Zhuang. Chart-hqa: A benchmark for hypothetical question answering in charts, 2025 f . URL https://arxiv.org/abs/2503.04095
2025 arXiv
-
[59]
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811, 2025 g
2025 arXiv
-
[60]
Videovista-culturallingo: 360 ^ horizons-bridging cultures, languages, and domains in video comprehension, 2025 h
Xinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, and Min Zhang. Videovista-culturallingo: 360 ^ horizons-bridging cultures, languages, and domains in video comprehension, 2025 h . URL https://arxiv.org/abs/2504.17821
2025 arXiv
-
[61]
Rm-r1: Reward modeling as reasoning, 2025 i
Xiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin, Cheng Qian, Yu Wang, Hongru Wang, Yu Zhang, Denghui Zhang, Tong Zhang, Hanghang Tong, and Heng Ji. Rm-r1: Reward modeling as reasoning, 2025 i . URL https://arxiv.org/abs/2505.02387
2025
-
[62]
UNITER: universal image-text representation learning
Yen - Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: universal image-text representation learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan - Michael Frahm (eds.), Computer Vision - ECCV 2020 - 16th Eu...
2020 doi
-
[63]
Exploring the effect of reinforcement learning on video understanding: Insights from seed-bench-r1, 2025 j
Yi Chen, Yuying Ge, Rui Wang, Yixiao Ge, Lu Qiu, Ying Shan, and Xihui Liu. Exploring the effect of reinforcement learning on video understanding: Insights from seed-bench-r1, 2025 j . URL https://arxiv.org/abs/2503.24376
2025 arXiv
- [64]
-
[65]
R1-code-interpreter: Training llms to reason with code via supervised and reinforcement learning, 2025 k
Yongchao Chen, Yueying Liu, Junwei Zhou, Yilun Hao, Jingquan Wang, Yang Zhang, and Chuchu Fan. R1-code-interpreter: Training llms to reason with code via supervised and reinforcement learning, 2025 k . URL https://arxiv.org/abs/2505.21668
2025
-
[66]
Octavius: Mitigating task interference in MLLM s via lo RA -moe
Zeren Chen, Ziqin Wang, Zhen Wang, Huayang Liu, Zhenfei Yin, Si Liu, Lu Sheng, Wanli Ouyang, and Jing Shao. Octavius: Mitigating task interference in MLLM s via lo RA -moe. In The Twelfth International Conference on Learning Representations, 2024 i . URL https://openreview.net...
2024
-
[67]
Visrl: Intention-driven visual perception via reinforced reasoning
Zhangquan Chen, Xufang Luo, and Dongsheng Li. Visrl: Intention-driven visual perception via reinforced reasoning. arXiv preprint arXiv:2503.07523, 2025 l
2025 arXiv
-
[68]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vis...
2024
-
[69]
See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning
Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Hao Zhang, and Chuang Gan. See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning. arXiv preprint arXiv:2301.05226, 2023 c
2023 arXiv
-
[70]
Unmasking deceptive visuals: Benchmarking multimodal large language models on misleading chart question answering, 2025 m
Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin, Rui Sheng, and Huamin Qu. Unmasking deceptive visuals: Benchmarking multimodal large language models on misleading chart question answering, 2025 m . URL https://arxiv.org/abs/2503.18172
2025
-
[71]
An adaptive framework for generating systematic explanatory answer in online q&a platforms
Ziyang Chen, Xiaobin Wang, Yong Jiang, Jinzhi Liao, Pengjun Xie, Fei Huang, and Xiang Zhao. An adaptive framework for generating systematic explanatory answer in online q&a platforms. arXiv preprint arXiv:2410.17694, 2024 k
2024 arXiv
-
[72]
From the least to the most: Building a plug-and-play visual reasoner via data synthesis
Chuanqi Cheng, Jian Guan, Wei Wu, and Rui Yan. From the least to the most: Building a plug-and-play visual reasoner via data synthesis. arXiv preprint arXiv:2406.19934, 2024 a
2024 arXiv
-
[73]
Seeclick: Harnessing gui grounding for advanced visual gui agents
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents. arXiv preprint arXiv:2401.10935, 2024 b
2024 arXiv
-
[74]
Embodiedeval: Evaluate multimodal llms as embodied agents
Zhili Cheng, Yuge Tu, Ran Li, Shiqi Dai, Jinyi Hu, Shengding Hu, Jiahao Li, Yang Shi, Tianyu Yu, Weize Chen, et al. Embodiedeval: Evaluate multimodal llms as embodied agents. arXiv preprint arXiv:2501.11858, 2025 a
2025 arXiv
-
[75]
Comt: A novel benchmark for chain of multi-modal thought on large vision-language models, 2025 b
Zihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei, Xiaocheng Feng, Wanxiang Che, Min Li, and Libo Qin. Comt: A novel benchmark for chain of multi-modal thought on large vision-language models, 2025 b . URL https://arxiv.org/abs/2412.12932
2025 arXiv
-
[76]
Anole: An open, autoregressive, native large multimodal models for interleaved image-text generation
Ethan Chern, Jiadi Su, Yan Ma, and Pengfei Liu. Anole: An open, autoregressive, native large multimodal models for interleaved image-text generation. arXiv preprint arXiv:2407.06135, 2024
2024 arXiv
-
[77]
Eva: An embodied world model for future video anticipation
Xiaowei Chi, Hengyuan Zhang, Chun-Kai Fan, Xingqun Qi, Rongyu Zhang, Anthony Chen, Chi-min Chan, Wei Xue, Wenhan Luo, Shanghang Zhang, and Yike Guo. Eva: An embodied world model for future video anticipation. arXiv preprint arXiv:2410.15461, 2024
-
[78]
Kitani, and Laszlo A
Rohan Choudhury, Koichiro Niinuma, Kris M. Kitani, and Laszlo A. Jeni. Zero-shot video question answering with procedural programs. In European Conference on Computer Vision (ECCV), pp.\ 1113--1130. Springer, 2024
2024
-
[79]
Merit: Multilingual semantic retrieval with interleaved multi-condition query, 2025 a
Wei Chow, Yuan Gao, Linfeng Li, Xian Wang, Qi Xu, Hang Song, Lingdong Kong, Ran Zhou, Yi Zeng, Yidong Cai, Botian Jiang, Shilin Xu, Jiajun Zhang, Minghui Qiu, Xiangtai Li, Tianshu Yang, Siliang Tang, and Juncheng Li. Merit: Multilingual semantic retrieval with interleaved mult...
2025
-
[80]
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, and Yue Wang. Physbench: Benchmarking and enhancing vision-language models for physical world understanding. CoRR, abs/2501.16411, 2025 b . doi:10.48550/ARXIV.2501.16411. URL https://doi.org/10.48550/arXiv.2501.16411
- [81]
-
[82]
Le, Sergey Levine, and Yi Ma
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, and Yi Ma. Sft memorizes, rl generalizes: A comparative study of foundation model post-training, 2025. URL https://arxiv.org/abs/2501.17161
2025 arXiv
-
[83]
Don't look only once: Towards multimodal interactive reasoning with selective visual revisitation, 2025
Jiwan Chung, Junhyeok Kim, Siyeol Kim, Jaeyoung Lee, Min Soo Kim, and Youngjae Yu. Don't look only once: Towards multimodal interactive reasoning with selective visual revisitation, 2025. URL https://arxiv.org/abs/2505.18842
2025 arXiv
-
[84]
FLEURS: few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. FLEURS: few-shot learning evaluation of universal representations of speech. In IEEE Spoken Language Technology Workshop, SLT 2022, Doha, Qatar, Januar...
2022
-
[85]
Receive, reason, and react: Drive as you say, with large language models in autonomous vehicles
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang. Receive, reason, and react: Drive as you say, with large language models in autonomous vehicles. IEEE Intelligent Transportation Systems Magazine, 2024
2024
-
[86]
Draw with thought: Unleashing multimodal reasoning for scientific diagram generation, 2025
Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, and Zhenglong Ding. Draw with thought: Unleashing multimodal reasoning for scientific diagram generation, 2025. URL https://arxiv.org/abs/2504.09479
2025
-
[87]
Vard: Efficient and dense fine-tuning for diffusion models with value-based rl, 2025 a
Fengyuan Dai, Zifeng Zhuang, Yufei Huang, Siteng Huang, Bangyan Liao, Donglin Wang, and Fajie Yuan. Vard: Efficient and dense fine-tuning for diffusion models with value-based rl, 2025 a . URL https://arxiv.org/abs/2505.15791
2025 arXiv
-
[88]
Physicsarena: The first multimodal physics reasoning benchmark exploring variable, process, and solution dimensions, 2025 b
Song Dai, Yibo Yan, Jiamin Su, Dongfang Zihao, Yubo Gao, Yonghua Hei, Jungang Li, Junyan Zhang, Sicheng Tao, Zhuoran Gao, and Xuming Hu. Physicsarena: The first multimodal physics reasoning benchmark exploring variable, process, and solution dimensions, 2025 b . URL https://ar...
2025 arXiv
-
[89]
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. URL https://arxiv.org/abs/2305.06500
2023 arXiv
-
[90]
Reinforcing video reasoning with focused thinking, 2025
Jisheng Dang, Jingze Wu, Teng Wang, Xuanhui Lin, Nannan Zhu, Hongbo Chen, Wei-Shi Zheng, Meng Wang, and Tat-Seng Chua. Reinforcing video reasoning with focused thinking, 2025. URL https://arxiv.org/abs/2505.24718
2025 arXiv
-
[91]
Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models
Rocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Iliyanov Dimitrov, Ivan Koychev, and Preslav Nakov. Exams-v: A multi-discipline multilingual multimodal exam benchmark for evaluating vision language models. arXiv preprint arXiv:2403.10378, 2024
2024 arXiv
- [92]
- [93]
-
[94]
Procthor: Large-scale embodied AI using procedural generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied AI using procedural generation. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, ...
2022
-
[95]
Rico: A mobile app dataset for building data-driven design applications
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar. Rico: A mobile app dataset for building data-driven design applications. In Proceedings of the 30th Annual Symposium on User Interface Software and Technolo...
2017
-
[96]
Emerging properties in unified multimodal pretraining
Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Shi Guang, and Haoqi Fan. Emerging properties in unified multimodal pretraining. CoRR, abs/2505.14683, 2025 a . doi:10.48550/ARXIV.2505.14683. URL https://doi....
-
[97]
Boosting the generalization and reasoning of vision language models with curriculum reinforcement learning, 2025 b
Huilin Deng, Ding Zou, Rui Ma, Hongchen Luo, Yang Cao, and Yu Kang. Boosting the generalization and reasoning of vision language models with curriculum reinforcement learning, 2025 b . URL https://arxiv.org/abs/2503.07065
2025 arXiv
-
[98]
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samual Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Pr...
2023
-
[99]
Redcaps: Web-curated image-text data created by the people, for the people
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson. Redcaps: Web-curated image-text data created by the people, for the people. In Joaquin Vanschoren and Sai - Kit Yeung (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1...
2021
-
[100]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Lan...
2019 doi
-
[101]
Write and paint: Generative vision-language models are unified modal learners
Shizhe Diao, Wangchunshu Zhou, Xinsong Zhang, and Jiawei Wang. Write and paint: Generative vision-language models are unified modal learners. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023. ...
2023
-
[102]
Soundmind: Rl-incentivized logic reasoning for audio-language models
Xingjian Diao, Chunhui Zhang, Keyi Kong, Weiyi Wu, Chiyu Ma, Zhongyu Ouyang, Peijun Qing, Soroush Vosoughi, and Jiang Gui. Soundmind: Rl-incentivized logic reasoning for audio-language models. arXiv preprint arXiv:2506.12935, 2025
2025
-
[103]
Hallu-pi: Evaluating hallucination in multi-modal large language models within perturbed inputs
Peng Ding, Jingyu Wu, Jun Kuang, Dan Ma, Xuezhi Cao, Xunliang Cai, Shi Chen, Jiajun Chen, and Shujian Huang. Hallu-pi: Evaluating hallucination in multi-modal large language models within perturbed inputs. In Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prabhakaran, Susanne...
2024
-
[104]
Mm-ifengine: Towards multimodal instruction following, 2025
Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, and Jiaqi Wang. Mm-ifengine: Towards multimodal instruction following, 2025. URL https://arxiv.org/abs/2504.07957
2025 arXiv
-
[105]
Progressive multimodal reasoning via active retrieval
Guanting Dong, Chenghao Zhang, Mengjie Deng, Yutao Zhu, Zhicheng Dou, and Ji-Rong Wen. Progressive multimodal reasoning via active retrieval. arXiv preprint arXiv:2412.14835, 2024 a
2024 arXiv
-
[106]
Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. arXiv preprint arXiv:2411.14432, 2024 b
2024 arXiv
-
[107]
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al. An empirical study of training end-to-end vision-and-language transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...
2022
-
[108]
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. ArXiv preprint, abs/2303.03378, 2023. URL https://arxiv.org/abs/2303.03378
2023 arXiv
-
[109]
Clotho: an audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. Clotho: an audio captioning dataset. In 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020 , pp.\ 736--740. IEEE , 2020. doi:10.1109/ICASSP40776....
2020
-
[110]
Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning, 2025 a
Chengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang, Linjiang Huang, Xingyu Zeng, Hongsheng Li, and Xihui Liu. Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning, 2025 a . URL https://arxiv.org/abs/2505.17022
2025 arXiv
-
[111]
Worldscore: A unified evaluation benchmark for world generation
Haoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei, and Jiajun Wu. Worldscore: A unified evaluation benchmark for world generation. arXiv preprint arXiv:2504.00983, 2025 b
2025
-
[112]
Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan
Jesse H. Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan. Neural audio synthesis of musical notes with wavenet autoencoders. In Doina Precup and Yee Whye Teh (eds.), Proceedings of the 34th International Conference on Mac...
2017
-
[113]
Kosiorek, Oiwi Parker Jones, and Ingmar Posner
Martin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, and Ingmar Posner. Genesis: Generative scene inference and sampling with object-centric latent representations. arXiv preprint arXiv:1907.13052, 2019
1907 arXiv
-
[114]
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang. Heterogeneous memory enhanced multimodal attention model for video question answering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 20...
2019
-
[115]
Sophiavl-r1: Reinforcing mllms reasoning with thinking reward, 2025 a
Kaixuan Fan, Kaituo Feng, Haoming Lyu, Dongzhan Zhou, and Xiangyu Yue. Sophiavl-r1: Reinforcing mllms reasoning with thinking reward, 2025 a . URL https://arxiv.org/abs/2505.17018
2025
-
[116]
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De - An Huang, Yuke Zhu, and Anima Anandkumar. Minedojo: Building open-ended embodied agents with internet-scale knowledge. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, ...
2022
-
[117]
Grit: Teaching mllms to think with images, 2025 b
Yue Fan, Xuehai He, Diji Yang, Kaizhi Zheng, Ching-Chen Kuo, Yuting Zheng, Sravana Jyothi Narayanaraju, Xinze Guan, and Xin Eric Wang. Grit: Teaching mllms to think with images, 2025 b . URL https://arxiv.org/abs/2505.15879
2025 arXiv
-
[118]
Chestx-reasoner: Advancing radiology foundation models with reasoning through step-by-step verification
Ziqing Fan, Cheng Liang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Chestx-reasoner: Advancing radiology foundation models with reasoning through step-by-step verification. arXiv preprint arXiv:2504.20930, 2025 c
2025 arXiv
-
[119]
Video-of-thought: Step-by-step video reasoning from perception to cognition
Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Forty-first International Conference on Machine Learning, 2024
2024
-
[120]
Retool: Reinforcement learning for strategic tool use in llms, 2025 a
Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang, Yujia Qin, Baoquan Zhong, Chengquan Jiang, Jinxin Chi, and Wanjun Zhong. Retool: Reinforcement learning for strategic tool use in llms, 2025 a . URL https://arxiv.org/abs/2504.11536
2025 arXiv
-
[121]
Video-r1: Reinforcing video reasoning in mllms
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Benyou Wang, and Xiangyu Yue. Video-r1: Reinforcing video reasoning in mllms. arXiv preprint arXiv:2503.21776, 2025 b
2025 arXiv
-
[122]
Visualsphinx: Large-scale synthetic vision logic puzzles for rl
Yichen Feng, Zhangchen Xu, Fengqing Jiang, Yuetai Li, Bhaskar Ramasubramanian, Luyao Niu, Bill Yuchen Lin, and Radha Poovendran. Visualsphinx: Large-scale synthetic vision logic puzzles for rl. arXiv preprint arXiv:2505.23977, 2025 c
2025 arXiv
-
[123]
Wikimixqa: A multimodal benchmark for question answering over tables and charts, 2025
Negar Foroutan, Angelika Romanou, Matin Ansaripour, Julian Martin Eisenschlos, Karl Aberer, and Rémi Lebret. Wikimixqa: A multimodal benchmark for question answering over tables and charts, 2025. URL https://arxiv.org/abs/2506.15594
2025 arXiv
-
[124]
Causalvqa: A physically grounded causal reasoning benchmark for video models, 2025
Aaron Foss, Chloe Evans, Sasha Mitts, Koustuv Sinha, Ammar Rizvi, and Justine T Kao. Causalvqa: A physically grounded causal reasoning benchmark for video models, 2025
2025
-
[125]
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, and Xing Sun. Video-mme: The first-ever compreh...
-
[126]
GPTS core: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. GPTS core: Evaluate as you desire. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language...
2024 doi
-
[127]
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv preprint arXiv:1606.01847, 2016
2016 arXiv
-
[128]
Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander J
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Muss...
2023
-
[129]
Feigelis, Daniel Bear, Dan Gutfreund, David D
Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, Kuno Kim, Elias Wang, Michael Lingelbach, Aidan Curtis, Kevin T. Feigelis, Daniel Bear, Dan Gutfreund, David ...
2021
-
[130]
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou. Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn. arXiv preprint arXiv:2306.08640, 2023
2023 arXiv
-
[131]
Exploring hallucination of large multimodal models in video understanding: Benchmark, analysis and mitigation, 2025 a
Hongcheng Gao, Jiashu Qu, Jingyi Tang, Baolong Bi, Yue Liu, Hongyu Chen, Li Liang, Li Su, and Qingming Huang. Exploring hallucination of large multimodal models in video understanding: Benchmark, analysis and mitigation, 2025 a . URL https://arxiv.org/abs/2503.19622
2025 arXiv
-
[132]
Interleaved-modal chain-of-thought
Jun Gao, Yongqi Li, Ziqiang Cao, and Wenjie Li. Interleaved-modal chain-of-thought. arXiv preprint arXiv:2411.19488, 2024 a
2024 arXiv
-
[133]
Pm4bench: A parallel multilingual multi-modal multi-task benchmark for large vision language model, 2025 b
Junyuan Gao, Jiahe Song, Jiang Wu, Runchuan Zhu, Guanlin Shen, Shasha Wang, Xingjian Wei, Haote Yang, Songyang Zhang, Weijia Li, Bin Wang, Dahua Lin, Lijun Wu, and Conghui He. Pm4bench: A parallel multilingual multi-modal multi-task benchmark for large vision language model, 2...
2025
-
[134]
Benchmarking open-ended audio dialogue understanding for large audio-language models
Kuofeng Gao, Shu - Tao Xia, Ke Xu, Philip Torr, and Jindong Gu. Benchmarking open-ended audio dialogue understanding for large audio-language models. CoRR, abs/2412.05167, 2024 b . doi:10.48550/ARXIV.2412.05167. URL https://doi.org/10.48550/arXiv.2412.05167
-
[135]
Uishift: Enhancing vlm-based gui agents through self-supervised reinforcement learning
Longxi Gao, Li Zhang, and Mengwei Xu. Uishift: Enhancing vlm-based gui agents through self-supervised reinforcement learning. arXiv preprint arXiv:2505.12493, 2025 c
2025
-
[136]
Cantor: Inspiring multimodal chain-of-thought of mllm
Timin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu, Yunhang Shen, Yan Zhang, Shengchuan Zhang, Xiawu Zheng, Xing Sun, Liujuan Cao, et al. Cantor: Inspiring multimodal chain-of-thought of mllm. In Proceedings of the 32nd ACM International Conference on Multimedia, pp.\ 9096--91...
2024
-
[137]
Towards generalizable vision-language robotic manipulation: A benchmark and llm-guided 3d policy
Ricardo Garcia, Shizhe Chen, and Cordelia Schmid. Towards generalizable vision-language robotic manipulation: A benchmark and llm-guided 3d policy. CoRR, abs/2410.01345, 2024. doi:10.48550/ARXIV.2410.01345. URL https://doi.org/10.48550/arXiv.2410.01345
-
[138]
Adve, and Yu-Xiong Wang
Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, and Yu-Xiong Wang. Agmmu: A comprehensive agricultural multimodal understanding and reasoning benchmark, 2025. URL https://arxiv.org/abs/2504.10568
2025 arXiv
-
[139]
Chain of thought prompt tuning in vision language models
Jiaxin Ge, Hongyin Luo, Siyuan Qian, Yulu Gan, Jie Fu, and Shanghang Zhang. Chain of thought prompt tuning in vision language models. arXiv preprint arXiv:2304.07919, 2023
2023 arXiv
-
[140]
Seed-data-edit technical report: A hybrid dataset for instructional image editing
Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge, and Ying Shan. Seed-data-edit technical report: A hybrid dataset for instructional image editing. arXiv preprint arXiv:2405.04007, 2024
2024 arXiv
-
[141]
Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos
Tiantian Geng, Jinrui Zhang, Qingni Wang, Teng Wang, Jinming Duan, and Feng Zheng. Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos. arXiv preprint arXiv:2411.19772, 2024
2024 arXiv
-
[142]
Geneval: An object-focused framework for evaluating text-to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Syst...
2023
-
[143]
Multimodal-tot
Kye Gomez. Multimodal-tot. https://github.com/kyegomez/MultiModal-ToT, 2023
2023
-
[144]
Space-10: A comprehensive benchmark for multimodal large language models in compositional spatial intelligence, 2025
Ziyang Gong, Wenhao Li, Oliver Ma, Songyuan Li, Jiayi Ji, Xue Yang, Gen Luo, Junchi Yan, and Rongrong Ji. Space-10: A comprehensive benchmark for multimodal large language models in compositional spatial intelligence, 2025. URL https://arxiv.org/abs/2506.07966
2025
-
[145]
Introducing gemini 2.0: our new ai model for the agentic era, 2024
Google. Introducing gemini 2.0: our new ai model for the agentic era, 2024. URL https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/#ceo-message
2024
-
[146]
Navigating the digital world as humans do: Universal visual grounding for GUI agents
Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for GUI agents. In The Thirteenth International Conference on Learning Representations, 2025 a . URL https://open...
2025
-
[147]
Perceptual decoupling for scalable multi-modal reasoning via reward-optimized captioning
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Xin Jin, Zhenguo Li, James T Kwok, and Yu Zhang. Perceptual decoupling for scalable multi-modal reasoning via reward-optimized captioning. arXiv preprint arXiv:2506.04559, 2025 b
2025
-
[148]
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers - Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, J...
2017 doi
-
[149]
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark, 2022
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Minzhe Niu, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang, Xin Jiang, Chunjing Xu, and Hang Xu. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark, 2022. URL https://arxiv.org/abs/2202.06767
2022 arXiv
-
[150]
Breaking the modality barrier: Universal embedding learning with multimodal llms, 2025
Tiancheng Gu, Kaicheng Yang, Ziyong Feng, Xingjun Wang, Yanzhao Zhang, Dingkun Long, Yingda Chen, Weidong Cai, and Jiankang Deng. Breaking the modality barrier: Universal embedding learning with multimodal llms, 2025. URL https://arxiv.org/abs/2504.17432
2025
-
[151]
M2-omni: Advancing omni-mllm for comprehensive modality support with competitive performance
Qingpei Guo, Kaiyou Song, Zipeng Feng, Ziping Ma, Qinglong Zhang, Sirui Gao, Xuzheng Yu, Yunxiao Sun, Jingdong Chen, Ming Yang, et al. M2-omni: Advancing omni-mllm for comprehensive modality support with competitive performance. arXiv preprint arXiv:2502.18778, 2025
2025 arXiv
-
[152]
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi. Visual programming: Compositional visual reasoning without training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 14953--14962, 2023
2023
-
[153]
An open-source software toolkit & benchmark suite for the evaluation and adaptation of multimodal action models, 2025
Pranav Guruprasad, Yangyue Wang, Sudipta Chowdhury, Jaewoo Song, and Harshvardhan Sikka. An open-source software toolkit & benchmark suite for the evaluation and adaptation of multimodal action models, 2025. URL https://arxiv.org/abs/2506.09172
2025 arXiv
-
[154]
Virology capabilities test (vct): A multimodal virology qa benchmark, 2025
Jasper Götting, Pedro Medeiros, Jon G Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe. Virology capabilities test (vct): A multimodal virology qa benchmark, 2025. URL https://arxiv.org/abs/2504.16137
2025 arXiv
-
[155]
Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning, 2025
Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, and Yu-Gang Jiang. Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning, 2025. URL https://arxiv.org/abs/2506.03596
2025
-
[156]
Language models are general-purpose interfaces
Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei. Language models are general-purpose interfaces. CoRR, abs/2206.06336, 2022. doi:10.48550/ARXIV.2206.06336. URL https://doi.org/10.48550/arXiv.2206.06336
-
[157]
Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, and Wenhu Chen. Videoscore: Building auto...
2024
-
[158]
Zhitao He, Zongwei Lyu, Dazhong Chen, Dadi Guo, and Yi R. Fung. Matp-bench: Can mllm be a good automated theorem prover for multimodal problems?, 2025. URL https://arxiv.org/abs/2506.06034
2025 arXiv
-
[159]
Tuomo Hiippala, Malihe Alikhani, Jonas Haverinen, Timo Kalliokoski, Evanfiya Logacheva, Serafina Orekhova, Aino Tuomainen, Matthew Stone, and John A. Bateman. AI2D-RST: a multimodal corpus of 1000 primary school science diagrams. Lang. Resour. Evaluation, 55 0 (3): 0 661--688,...
2021 doi
-
[160]
Let's think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of-thought
Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, and William Yang Wang. Let's think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of-thought. arXiv preprint arXiv:2305.13...
2023 arXiv
-
[161]
Worldsense: Evaluating real-world omnimodal understanding for multimodal llms
Jack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, and Weidi Xie. Worldsense: Evaluating real-world omnimodal understanding for multimodal llms. arXiv preprint arXiv:2502.04326, 2025
2025 arXiv
-
[162]
Cogagent: A visual language model for gui agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, et al. Cogagent: A visual language model for gui agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 14281--14...
2024
-
[163]
Tdbench: Benchmarking vision-language models in understanding top-down images, 2025
Kaiyuan Hou, Minghui Zhao, Lilin Xu, Yuang Fan, and Xiaofan Jiang. Tdbench: Benchmarking vision-language models in understanding top-down images, 2025. URL https://arxiv.org/abs/2504.03748
2025
-
[164]
GAIA-1: A generative world model for autonomous driving
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. GAIA-1: A generative world model for autonomous driving. CoRR, abs/2309.17080, 2023. doi:10.48550/ARXIV.2309.17080. URL https://doi.org/10.48550/arXiv.2309.17080
-
[165]
Mcitebench: A benchmark for multimodal citation text generation in mllms, 2025 a
Caiyu Hu, Yikai Zhang, Tinghui Zhu, Yiwei Ye, and Yanghua Xiao. Mcitebench: A benchmark for multimodal citation text generation in mllms, 2025 a . URL https://arxiv.org/abs/2503.02589
2025 arXiv
-
[166]
Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos
Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Yuanhan Zhang, Xiang Yue, Bo Li, and Ziwei Liu. Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos. CoRR, abs/2501.13826, 2025 b . doi:10.48550/ARXIV.2501.13826. URL https://doi.org/10.48550/arXiv....
-
[167]
Unit: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh. Unit: Multimodal multitask learning with a unified transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pp.\ 1439--1449, 2021
2021
-
[168]
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. Learning to reason: End-to-end module networks for visual question answering. In Proceedings of the IEEE international conference on computer vision, pp.\ 804--813, 2017
2017
-
[169]
ELLA: equip diffusion models with LLM for enhanced semantic alignment
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. ELLA: equip diffusion models with LLM for enhanced semantic alignment. CoRR, abs/2403.05135, 2024 a . doi:10.48550/ARXIV.2403.05135. URL https://doi.org/10.48550/arXiv.2403.05135
-
[170]
Visual sketchpad: Sketching as a visual chain of thought for multimodal language models
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. arXiv preprint arXiv:2406.09403, 2024 b
2024 arXiv
-
[172]
T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation, 2025 a
Kaiyi Huang, Chengqi Duan, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation, 2025 a . URL https://arxiv.org/abs/2307.06350
2025 arXiv
-
[173]
Enhance mobile agents thinking process via iterative preference learning
Kun Huang, Weikai Xu, Yuxuan Liu, Quandong Wang, Pengzhi Gao, Wei Liu, Jian Luan, Bin Wang, and Bo An. Enhance mobile agents thinking process via iterative preference learning. arXiv preprint arXiv:2505.12299, 2025 b
2025
-
[174]
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Nils Johan Bertil Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. Language is not all you ...
2023
-
[175]
Vision-r1: Incentivizing reasoning capability in multimodal large language models
Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Yao Hu, and Shaohui Lin. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749, 2025 c
2025 arXiv
-
[176]
Key-point-driven data synthesis with its enhancement on mathematical reasoning
Yiming Huang, Xiao Liu, Yeyun Gong, Zhibin Gou, Yelong Shen, Nan Duan, and Weizhu Chen. Key-point-driven data synthesis with its enhancement on mathematical reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pp.\ 24176--24184, 2025 d
2025
-
[177]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Alice Oh, Tristan N...
2023
-
[178]
Visualtoolagent (vista): A reinforcement learning framework for visual tool selection, 2025 e
Zeyi Huang, Yuyang Ji, Anirudh Sundara Rajan, Zefan Cai, Wen Xiao, Junjie Hu, and Yong Jae Lee. Visualtoolagent (vista): A reinforcement learning framework for visual tool selection, 2025 e . URL https://arxiv.org/abs/2505.20289
2025 arXiv
-
[179]
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849, 2020
2004 arXiv
-
[181]
Vbench++: Comprehensive and versatile benchmark suite for video generative models
Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying - Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench++: Comprehensive and versatile benchmark suite for...
-
[182]
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning. Compositional attention networks for machine reasoning. arXiv preprint arXiv:1803.03067, 2018
2018 arXiv
-
[183]
Hudson and Christopher D
Drew A. Hudson and Christopher D. Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , pp.\ 6700--6709. Computer Visio...
2019
-
[184]
Hq-edit: A high-quality dataset for instruction-based image editing
Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Heng Wang, Peng Wang, Yuyin Zhou, and Cihang Xie. Hq-edit: A high-quality dataset for instruction-based image editing. CoRR, abs/2404.09990, 2024. doi:10.48550/ARXIV.2404.09990. URL https://doi.org/10.48550/arXiv.2404.09990
-
[185]
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[186]
Multimodal learning and reasoning for visual question answering
Ilija Ilievski and Jiashi Feng. Multimodal learning and reasoning for visual question answering. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Advances in Neural Information Processing Systems...
2017
-
[187]
Openai o1 system card
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El - Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennett, Ananya...
-
[188]
Wavreward: Spoken dialogue models with generalist reward evaluators
Shengpeng Ji, Tianle Liang, Yangzhuo Li, Jialong Zuo, Minghui Fang, Jinzheng He, Yifu Chen, Zhengqing Liu, Ziyue Jiang, Xize Cheng, et al. Wavreward: Spoken dialogue models with generalist reward evaluators. arXiv preprint arXiv:2505.09558, 2025
2025
-
[189]
Le, Yunhsuan Sung, Zhen Li, and Tom Duerig
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision, 2021. URL https://arxiv.org/abs/2102.05918
2021 arXiv
-
[190]
Dcot: Dual chain-of-thought prompting for large multimodal models
Zixi Jia, Jiqiang Liu, Hexiao Li, Qinghua Liu, and Hongbin Gao. Dcot: Dual chain-of-thought prompting for large multimodal models. In The 16th Asian Conference on Machine Learning (Conference Track), 2024
2024
-
[191]
Bootstrapping vision-language learning with decoupled language pre-training
Yiren Jian, Chongyang Gao, and Soroush Vosoughi. Bootstrapping vision-language learning with decoupled language pre-training. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36...
2023
-
[192]
Vlm-r ^3 : Region recognition, reasoning, and refinement for enhanced multimodal chain-of-thought, 2025 a
Chaoya Jiang, Yongrui Heng, Wei Ye, Han Yang, Haiyang Xu, Ming Yan, Ji Zhang, Fei Huang, and Shikun Zhang. Vlm-r ^3 : Region recognition, reasoning, and refinement for enhanced multimodal chain-of-thought, 2025 a . URL https://arxiv.org/abs/2505.16192
2025 arXiv
-
[193]
T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot, 2025 b
Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, and Hongsheng Li. T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot, 2025 b . URL https://arxiv.org/abs/2505.00703
2025 arXiv
-
[194]
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han. Reasoning with heterogeneous graph alignment for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp.\ 11109--11116, 2020
2020
-
[195]
Rex-thinker: Grounded object referring via chain-of-thought reasoning, 2025 c
Qing Jiang, Xingyu Chen, Zhaoyang Zeng, Junzhi Yu, and Lei Zhang. Rex-thinker: Grounded object referring via chain-of-thought reasoning, 2025 c . URL https://arxiv.org/abs/2506.04034
2025 arXiv
-
[196]
VIMA: general robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei - Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. VIMA: general robot manipulation with multimodal prompts. CoRR, abs/2210.03094, 2022. doi:10.48550/ARXIV.2210.03094. URL https://doi.org/...
-
[197]
Unitoken: Harmonizing multimodal understanding and generation through unified visual encoding, 2025
Yang Jiao, Haibo Qiu, Zequn Jie, Shaoxiang Chen, Jingjing Chen, Lin Ma, and Yu-Gang Jiang. Unitoken: Harmonizing multimodal understanding and generation through unified visual encoding, 2025. URL https://arxiv.org/abs/2504.04423
2025 arXiv
-
[198]
Unified language-vision pretraining in LLM with dynamic discrete visual tokenization
Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Quzhe Huang, Bin Chen, Chengru Song, Dai Meng, Di Zhang, Wenwu Ou, Kun Gai, and Yadong Mu. Unified language-vision pretraining in LLM with dynamic discrete visual tokenization. In The Twelfth International Conferen...
2024
-
[199]
Lawrence Zitnick, and Ross Girshick
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016. URL https://arxiv.org/abs/1612.06890
2016 arXiv
-
[200]
Weld, and Luke Zettlemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Regina Barzilay and Min - Yen Kan (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational L...
2017 doi
-
[201]
Kadhim, Lei Jiao, Rishad Shafik, and Ole-Christoffer Granmo
Ahmed K. Kadhim, Lei Jiao, Rishad Shafik, and Ole-Christoffer Granmo. Omni tm-ae: A scalable and interpretable embedding model using the full tsetlin machine state space, 2025. URL https://arxiv.org/abs/2505.16386
2025 arXiv
-
[202]
Visual question answering: Datasets, algorithms, and future challenges
Kushal Kafle and Christopher Kanan. Visual question answering: Datasets, algorithms, and future challenges. CoRR, abs/1610.01465, 2016. URL http://arxiv.org/abs/1610.01465
2016 arXiv
-
[203]
An analysis of visual question answering algorithms, 2017
Kushal Kafle and Christopher Kanan. An analysis of visual question answering algorithms, 2017. URL https://arxiv.org/abs/1703.09684
2017 arXiv
-
[204]
Price, Scott Cohen, and Christopher Kanan
Kushal Kafle, Brian L. Price, Scott Cohen, and Christopher Kanan. DVQA: understanding data visualizations via question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018 , pp.\ 5648--5656. Compute...
2018
-
[205]
Thinking, fast and slow
Daniel Kahneman. Thinking, fast and slow. macmillan, 2011
2011
-
[206]
Do you see me : A multidimensional benchmark for evaluating visual perception in multimodal llms, 2025
Aditya Kanade and Tanuja Ganu. Do you see me : A multidimensional benchmark for evaluating visual perception in multimodal llms, 2025. URL https://arxiv.org/abs/2506.02022
2025
-
[207]
Hssbench: Benchmarking humanities and social sciences ability for multimodal large language models, 2025
Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, and Xuelong Li. Hssbench: Benchmarking humanities and social sciences abi...
2025 arXiv
-
[208]
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, and Ruslan Salakhutdinov. Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web. In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Ru...
2024
-
[209]
Hydra: A hyper agent for dynamic compositional visual reasoning
Fucai Ke, Zhixi Cai, Simindokht Jahangard, Weiqing Wang, Pari Delir Haghighi, and Hamid Rezatofighi. Hydra: A hyper agent for dynamic compositional visual reasoning. In European Conference on Computer Vision, pp.\ 132--149. Springer, 2024
2024
-
[210]
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Min Joon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (eds.), Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherla...
2016 doi
-
[211]
Ragar, your falsehood radar: Rag-augmented reasoning for political fact-checking using multimodal large language models
M Abdul Khaliq, P Chang, M Ma, Bernhard Pflugfelder, and F Mileti \'c . Ragar, your falsehood radar: Rag-augmented reasoning for political fact-checking using multimodal large language models. arXiv preprint arXiv:2404.12065, 2024
2024 arXiv
-
[212]
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. Audiocaps: Generating captions for audios in the wild. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computati...
2019 doi
-
[213]
Videocomp: Advancing fine-grained compositional and temporal alignment in video-text models, 2025
Dahun Kim, AJ Piergiovanni, Ganesh Mallya, and Anelia Angelova. Videocomp: Advancing fine-grained compositional and temporal alignment in video-text models, 2025. URL https://arxiv.org/abs/2504.03970
2025 arXiv
-
[214]
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear attention networks. Advances in neural information processing systems, 31, 2018
2018
-
[215]
GVCCI: lifelong learning of visual grounding for language-guided robotic manipulation
Junghyun Kim, Gi - Cheon Kang, Jaein Kim, Suyeon Shin, and Byoung - Tak Zhang. GVCCI: lifelong learning of visual grounding for language-guided robotic manipulation. In IROS , pp.\ 952--959, 2023. doi:10.1109/IROS55552.2023.10342021. URL https://doi.org/10.1109/IROS55552.2023.10342021
2023
-
[216]
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024
2024 arXiv
-
[217]
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv preprint arXiv:2401.13649, 2024
2024 arXiv
-
[218]
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural Information Processing Systems 35: Annual ...
2022
-
[219]
cadrille: Multi-modal cad reconstruction with online reinforcement learning
Maksim Kolodiazhnyi, Denis Tarasov, Dmitrii Zhemchuzhnikov, Alexander Nikulin, Ilya Zisman, Anna Vorontsova, Anton Konushin, Vladislav Kurenkov, and Danila Rukhovich. cadrille: Multi-modal cad reconstruction with online reinforcement learning. arXiv preprint arXiv:2505.22914, 2025
2025
-
[220]
AI2-THOR: an interactive 3d environment for visual AI
Eric Kolve, Roozbeh Mottaghi, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. AI2-THOR: an interactive 3d environment for visual AI . CoRR, abs/1712.05474, 2017. URL http://arxiv.org/abs/1712.05474
2017 arXiv
-
[221]
Ross, Bryan Seybold, and Lu Jiang
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...
-
[222]
Shamma, Michael S
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. Visual genome: Connecting language and vision using crowdsourced dense image annotations, 20...
2016 arXiv
-
[223]
Overthinking: Slowdown attacks on reasoning llms
Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, and Eugene Bagdasarian. Overthinking: Slowdown attacks on reasoning llms. arXiv preprint arXiv:2502.02542, 2025
2025
-
[224]
Multimodal open r1, 2025
EvolvingLMMs Lab. Multimodal open r1, 2025. URL https://github.com/EvolvingLMMs-Lab/open-r1-multimodal. Accessed: 2025-02-28
2025
-
[225]
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 9579--9589, 2024
2024
-
[226]
Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xiaofeng Yang. Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models. arXiv preprint arXiv:2503.13939, 2025
2025
-
[227]
Finlmm-r1: Enhancing financial reasoning in lmm through scalable data and reward design, 2025
Kai Lan, Jiayong Zhu, Jiangtong Li, Dawei Cheng, Guang Chen, and Changjun Jiang. Finlmm-r1: Enhancing financial reasoning in lmm through scalable data and reward design, 2025. URL https://arxiv.org/abs/2506.13066
2025 arXiv
-
[228]
Refocus: Reinforcement-guided frame optimization for contextual understanding, 2025
Hosu Lee, Junho Kim, Hyunjun Kim, and Yong Man Ro. Refocus: Reinforcement-guided frame optimization for contextual understanding, 2025. URL https://arxiv.org/abs/2506.01274
2025 arXiv
-
[229]
Multimodal reasoning with multimodal knowledge graph
Junlin Lee, Yequan Wang, Jing Li, and Min Zhang. Multimodal reasoning with multimodal knowledge graph. arXiv preprint arXiv:2406.02030, 2024
2024 arXiv
-
[230]
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. Tvqa: Localized, compositional video question answering. arXiv preprint arXiv:1809.01696, 2018
2018 arXiv
-
[231]
Godbench: A benchmark for multimodal large language models in video comment art, 2025
Yiming Lei, Chenkai Zhang, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu, and Yunhong Wang. Godbench: A benchmark for multimodal large language models in video comment art, 2025. URL https://arxiv.org/abs/2505.11436
2025 arXiv
-
[232]
The curse of multi-modalities: Evaluating hallucinations of large multimodal models across language, visual, and audio, 2024
Sicong Leng, Yun Xing, Zesen Cheng, Yang Zhou, Hang Zhang, Xin Li, Deli Zhao, Shijian Lu, Chunyan Miao, and Lidong Bing. The curse of multi-modalities: Evaluating hallucinations of large multimodal models across language, visual, and audio, 2024. URL https://arxiv.org/abs/2410.12787
2024 arXiv
-
[233]
Genai-bench: Evaluating and improving compositional text-to-visual generation
Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Tiffany Ling, Xide Xia, Pengchuan Zhang, Graham Neubig, and Deva Ramanan. Genai-bench: Evaluating and improving compositional text-to-visual generation. CoRR, abs/2406.13743, 2024 a . doi:10.48550/ARXIV.2406....
-
[234]
Naturalbench: Evaluating vision-language models on natural adversarial samples
Baiqi Li, Zhiqiu Lin, Wenxuan Peng, Jean de Dieu Nyandwi, Daniel Jiang, Zixian Ma, Simran Khanuja, Ranjay Krishna, Graham Neubig, and Deva Ramanan. Naturalbench: Evaluating vision-language models on natural adversarial samples. In Amir Globersons, Lester Mackey, Danielle Belgr...
2024
-
[235]
Otter: A multi-modal model with in-context instruction tuning, 2023 a
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning, 2023 a . URL https://arxiv.org/abs/2305.03726
2023 arXiv
-
[236]
Seed-bench: Benchmarking multimodal llms with generative comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. CoRR, abs/2307.16125, 2023 b . doi:10.48550/ARXIV.2307.16125. URL https://doi.org/10.48550/arXiv.2307.16125
-
[237]
Megrez-omni technical report
Boxun Li, Yadong Li, Zhiyuan Li, Congyi Liu, Weilin Liu, Guowei Niu, Zheyue Tan, Haiyang Xu, Zhuyu Yao, Tao Yuan, et al. Megrez-omni technical report. arXiv preprint arXiv:2502.15803, 2025 a
2025 arXiv
-
[238]
Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei - Fei, and Silvio Savarese
Chengshu Li, Fei Xia, Roberto Mart \' n - Mart \' n, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, Andrey Kurenkov, C. Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei - Fei, and Silvio Savarese. igibson 2.0: Object-centric s...
2021 arXiv
-
[239]
Matthews, Ivan Villa - Renteria, Jerry Huayang Tang, Claire Tang, Fei Xia, Yunzhu Li, Silvio Savarese, Hyowon Gweon, C
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart \' n - Mart \' n, Chen Wang, Gabrael Levine, Wensi Ai, Benjamin Jose Martinez, Hang Yin, Michael Lingelbach, Minjune Hwang, Ayano Hiranaka, Sujay Garlanka, Arman Aydin, Sharon Lee, Jiankai Sun...
-
[240]
Imagine while reasoning in space: Multimodal visualization-of-thought
Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vuli \'c , and Furu Wei. Imagine while reasoning in space: Multimodal visualization-of-thought. arXiv preprint arXiv:2501.07542, 2025 b
2025 arXiv
-
[241]
Gonzalez, Ion Stoica, Song Han, and Yao Lu
Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang, Hongxu Yin, Joseph E. Gonzalez, Ion Stoica, Song Han, and Yao Lu. Worldmodelbench: Judging video generation models as world models. CoRR, abs/2502.20694, 2025 c . doi:10.48550/...
-
[242]
Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image generation, 2024 d
Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet, Linmiao Xu, and Suhail Doshi. Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image generation, 2024 d
2024
-
[243]
Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering
Gang Li, Jizhong Liu, Heinrich Dinkel, Yadong Niu, Junbo Zhang, and Jian Luan. Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering. arXiv preprint arXiv:2503.11197, 2025 d
2025 arXiv
-
[244]
Avqa-cot: When cot meets question answering in audio-visual scenarios
Guangyao Li, Henghui Du, and Di Hu. Avqa-cot: When cot meets question answering in audio-visual scenarios. In CVPR Workshops, 2024 e
2024
-
[245]
CMMLU: measuring massive multitask language understanding in chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. CMMLU: measuring massive multitask language understanding in chinese. In Lun - Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Findings of the Association for Computational ...
2024 doi
-
[246]
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34: 0 9694--9705, 2021 b
2021
-
[247]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, pp.\ 12888--12900. PMLR, 2022
2022
-
[248]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.)...
2023
-
[249]
Muep: A multimodal benchmark for embodied planning with foundation models
Kanxue Li, Baosheng Yu, Qi Zheng, Yibing Zhan, Yuhui Zhang, Tianle Zhang, Yijun Yang, Yue Chen, Lei Sun, Qiong Cao, Li Shen, Lusong Li, Dapeng Tao, and Xiaodong He. Muep: A multimodal benchmark for embodied planning with foundation models. In Proceedings of the Thirty-Third In...
2024
-
[250]
Cpseg: Finer-grained image semantic segmentation via chain-of-thought language prompting
Lei Li. Cpseg: Finer-grained image semantic segmentation via chain-of-thought language prompting. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.\ 513--522, 2024
2024
-
[251]
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019
1908 arXiv
-
[252]
Towards visual text grounding of multimodal large language model, 2025 e
Ming Li, Ruiyi Zhang, Jian Chen, Jiuxiang Gu, Yufan Zhou, Franck Dernoncourt, Wanrong Zhu, Tianyi Zhou, and Tong Sun. Towards visual text grounding of multimodal large language model, 2025 e . URL https://arxiv.org/abs/2504.04974
2025
-
[253]
PRD : Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. PRD : Peer rank and discussion improve large language model based evaluations. Transactions on Machine Learning Research, 2024 h . ISSN 2835-8856. URL https://openreview.net/forum?id=YVD1QqWRaj
2024
-
[254]
Truth in the few: High-value data selection for efficient multi-modal reasoning
Shenshen Li, Kaiyuan Deng, Lei Wang, Hao Yang, Chong Peng, Peng Yan, Fumin Shen, Heng Tao Shen, and Xing Xu. Truth in the few: High-value data selection for efficient multi-modal reasoning. arXiv preprint arXiv:2506.04755, 2025 f
2025
-
[255]
Gdi-bench: A benchmark for general document intelligence with vision and reasoning decoupling, 2025 g
Siqi Li, Yufan Shen, Xiangnan Chen, Jiayi Chen, Hengwei Ju, Haodong Duan, Song Mao, Hongbin Zhou, Bo Zhang, Bin Fu, Pinlong Cai, Licheng Wen, Botian Shi, Yong Liu, Xinyu Cai, and Yu Qiao. Gdi-bench: A benchmark for general document intelligence with vision and reasoning decoup...
2025 arXiv
-
[256]
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
Xiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng, Yuxing Long, Yan Shen, Renrui Zhang, Jiaming Liu, and Hao Dong. Manipllm: Embodied multimodal large language model for object-centric robotic manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patter...
2024
-
[257]
Search-o1: Agentic search-enhanced large reasoning models
Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. Search-o1: Agentic search-enhanced large reasoning models. arXiv preprint arXiv:2501.05366, 2025 h
2025 arXiv
-
[258]
Multimodal coreference resolution for chinese social media dialogues: Dataset and benchmark approach, 2025 i
Xingyu Li, Chen Gong, and Guohong Fu. Multimodal coreference resolution for chinese social media dialogues: Dataset and benchmark approach, 2025 i . URL https://arxiv.org/abs/2504.14321
2025 arXiv
-
[259]
Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning, 2025 j
Xinhao Li, Ziang Yan, Desen Meng, Lu Dong, Xiangyu Zeng, Yinan He, Yali Wang, Yu Qiao, Yi Wang, and Limin Wang. Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning, 2025 j . URL https://arxiv.org/abs/2504.06958
2025 arXiv
-
[260]
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan - Mi...
2020
-
[261]
Baichuan-omni-1.5 technical report
Yadong Li, Jun Liu, Tao Zhang, Song Chen, Tianpeng Li, Zehuan Li, Lijun Liu, Lingfeng Ming, Guosheng Dong, Da Pan, et al. Baichuan-omni-1.5 technical report. arXiv preprint arXiv:2501.15368, 2025 k
2025
-
[262]
Omnibench: Towards the future of universal omni-language models
Yizhi Li, Ge Zhang, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Jiaheng Liu, Zekun Wang, Jian Yang, et al. Omnibench: Towards the future of universal omni-language models. arXiv preprint arXiv:2409.15272, 2024 j
2024
-
[263]
Sti-bench: Are mllms ready for precise spatial-temporal world understanding?, 2025 l
Yun Li, Yiming Zhang, Tao Lin, XiangRui Liu, Wenxiao Cai, Zheng Liu, and Bo Zhao. Sti-bench: Are mllms ready for precise spatial-temporal world understanding?, 2025 l . URL https://arxiv.org/abs/2503.23765
2025 arXiv
-
[264]
A multi-modal context reasoning approach for conditional inference on joint textual and visual clues
Yunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding, Lin Ma, and Min Zhang. A multi-modal context reasoning approach for conditional inference on joint textual and visual clues. arXiv preprint arXiv:2305.04530, 2023 d
2023 arXiv
-
[265]
Videovista: A versatile benchmark for video understanding and reasoning
Yunxin Li, Xinyu Chen, Baotian Hu, Longyue Wang, Haoyuan Shi, and Min Zhang. Videovista: A versatile benchmark for video understanding and reasoning. CoRR, abs/2406.11303, 2024 k . doi:10.48550/ARXIV.2406.11303. URL https://doi.org/10.48550/arXiv.2406.11303
-
[266]
Lmeye: An interactive perception network for large language models
Yunxin Li, Baotian Hu, Xinyu Chen, Lin Ma, Yong Xu, and Min Zhang. Lmeye: An interactive perception network for large language models. IEEE Trans. Multim. , 26: 0 10952--10964, 2024 l . doi:10.1109/TMM.2024.3428317. URL https://doi.org/10.1109/TMM.2024.3428317
2024
-
[267]
Veripo: Cultivating long reasoning in video-llms via verifier-gudied iterative policy optimization, 2025 m
Yunxin Li, Xinyu Chen, Zitao Li, Zhenyu Liu, Longyue Wang, Wenhan Luo, Baotian Hu, and Min Zhang. Veripo: Cultivating long reasoning in video-llms via verifier-gudied iterative policy optimization, 2025 m . URL https://arxiv.org/abs/2505.19000
2025 arXiv
-
[268]
Uni-moe: Scaling unified multimodal llms with mixture of experts
Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma, and Min Zhang. Uni-moe: Scaling unified multimodal llms with mixture of experts. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025 n
2025
-
[269]
Vision matters: Simple visual perturbations can boost multimodal math reasoning
Yuting Li, Lai Wei, Kaipeng Zheng, Jingyuan Huang, Linghe Kong, Lichao Sun, and Weiran Huang. Vision matters: Simple visual perturbations can boost multimodal math reasoning. arXiv preprint arXiv:2506.09736, 2025 o
2025
-
[270]
Iuzzolino, Brett D
Yuxuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads, Asli Celikyilmaz, and Karl Ridgeway. Egotom: Benchmarking theory of mind reasoning from egocentric videos, 2025 p . URL https://arxiv.org/abs/2503.22152
2025 arXiv
-
[271]
Vocot: Unleashing visually grounded multi-step reasoning in large multi-modal models
Zejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang, and Zhongyu Wei. Vocot: Unleashing visually grounded multi-step reasoning in large multi-modal models. arXiv preprint arXiv:2405.16919, 2024 m
2024 arXiv
-
[272]
Enhancing advanced visual reasoning ability of large language models
Zhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang, Tengfei Xue, and Weidong Cai. Enhancing advanced visual reasoning ability of large language models. arXiv preprint arXiv:2409.13980, 2024 n
2024 arXiv
-
[273]
Star-r1: Spatial transformation reasoning by reinforcing multimodal llms, 2025 q
Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, and Wenbing Huang. Star-r1: Spatial transformation reasoning by reinforcing multimodal llms, 2025 q . URL https://arxiv.org/abs/2505.15804
2025 arXiv
-
[274]
Explainable multimodal emotion reasoning
Zheng Lian, Licai Sun, Mingyu Xu, Haiyang Sun, Ke Xu, Zhuofan Wen, Shun Chen, Bin Liu, and Jianhua Tao. Explainable multimodal emotion reasoning. CoRR, 2023
2023
-
[275]
ABSE val: An agent-based framework for script evaluation
Sirui Liang, Baoli Zhang, Jun Zhao, and Kang Liu. ABSE val: An agent-based framework for script evaluation. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 12418--12434, M...
2024 doi
-
[276]
Memory-driven multimodal chain of thought for embodied long-horizon task planning
Xiwen Liang, Min Lin, Weiqi Ruan, Yuecheng Liu, Yuzheng Zhuang, and Xiaodan Liang. Memory-driven multimodal chain of thought for embodied long-horizon task planning. Openreview, 2025 a
2025
-
[277]
Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness, 2025 b
Yijun Liang, Ming Li, Chenrui Fan, Ziyue Li, Dang Nguyen, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, and Tianyi Zhou. Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness, 2025 b . U...
2025
-
[278]
Modomodo: Multi-domain data mixtures for multimodal llm reinforcement learning
Yiqing Liang, Jielin Qiu, Wenhao Ding, Zuxin Liu, James Tompkin, Mengdi Xu, Mengzhou Xia, Zhengzhong Tu, Laixi Shi, and Jiacheng Zhu. Modomodo: Multi-domain data mixtures for multimodal llm reinforcement learning. arXiv preprint arXiv:2505.24871, 2025 c
2025 arXiv
-
[279]
Improved visual-spatial reasoning via r1-zero-like training, 2025
Zhenyi Liao, Qingsong Xie, Yanhao Zhang, Zijian Kong, Haonan Lu, Zhenyu Yang, and Zhijie Deng. Improved visual-spatial reasoning via r1-zero-like training, 2025. URL https://arxiv.org/abs/2504.00883
2025 arXiv
-
[280]
Visescape: A benchmark for evaluating exploration-driven decision-making in virtual escape rooms, 2025
Seungwon Lim, Sungwoong Kim, Jihwan Yu, Sungjae Lee, Jiwan Chung, and Youngjae Yu. Visescape: A benchmark for evaluating exploration-driven decision-making in virtual escape rooms, 2025. URL https://arxiv.org/abs/2503.14427
2025 arXiv
-
[281]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C
Tsung - Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In David J. Fleet, Tom \' a s Pajdla, Bernt Schiele, and Tinne Tuytelaars (eds.), Computer Vision -...
2014 doi
-
[282]
Investigating inference-time scaling for chain of multi-modal thought: A preliminary study
Yujie Lin, Ante Wang, Moye Chen, Jingyao Liu, Hao Liu, Jinsong Su, and Xinyan Xiao. Investigating inference-time scaling for chain of multi-modal thought: A preliminary study. arXiv preprint arXiv:2502.11514, 2025 a
2025 arXiv
-
[283]
Why we feel: Breaking boundaries in emotional reasoning with multimodal large language models, 2025 b
Yuxiang Lin, Jingdong Sun, Zhi-Qi Cheng, Jue Wang, Haomin Liang, Zebang Cheng, Yifei Dong, Jun-Yan He, Xiaojiang Peng, and Xian-Sheng Hua. Why we feel: Breaking boundaries in emotional reasoning with multimodal large language models, 2025 b . URL https://arxiv.org/abs/2504.07521
2025 arXiv
-
[284]
Webuibench: A comprehensive benchmark for evaluating multimodal large language models in webui-to-code, 2025 c
Zhiyu Lin, Zhengda Zhou, Zhiyuan Zhao, Tianrui Wan, Yilun Ma, Junyu Gao, and Xuelong Li. Webuibench: A comprehensive benchmark for evaluating multimodal large language models in webui-to-code, 2025 c . URL https://arxiv.org/abs/2506.07818
2025 arXiv
-
[285]
Clotho-aqa: A crowdsourced dataset for audio question answering
Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, and Tuomas Virtanen. Clotho-aqa: A crowdsourced dataset for audio question answering. In 30th European Signal Processing Conference, EUSIPCO 2022, Belgrade, Serbia, August 29 - Sept. 2, 2022 , pp.\ 1140--1144. IE...
2022
-
[286]
More thinking, less seeing? assessing amplified hallucination in multimodal reasoning models, 2025 a
Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou, Xin Eric Wang, Yuyin Zhou, and Sheng Liu. More thinking, less seeing? assessing amplified hallucination in multimodal reasoning models, 2025 a . URL https://arxiv.org/abs/2505.21523
2025 arXiv
-
[287]
World model on million-length video and language with blockwise ringattention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with blockwise ringattention. CoRR, abs/2402.08268, 2024 a . doi:10.48550/ARXIV.2402.08268. URL https://doi.org/10.48550/arXiv.2402.08268
-
[288]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information ...
2023
-
[289]
Flow-grpo: Training flow matching models via online rl
Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl. arXiv preprint arXiv:2505.05470, 2025 b
2025 arXiv
-
[290]
Towards unified referring expression segmentation across omni-level visual target granularities, 2025 c
Jing Liu, Wenxuan Wang, Yisi Zhang, Yepeng Tang, Xingjian He, Longteng Guo, Tongtian Yue, and Xinlong Wang. Towards unified referring expression segmentation across omni-level visual target granularities, 2025 c . URL https://arxiv.org/abs/2504.01954
2025 arXiv
-
[291]
Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding? CoRR, abs/2404.05955, 2024 b
Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue. Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding? CoRR, abs/2404.05955, 2024 b . doi:10.48550/ARXIV.2404.05955. URL https://doi.org/10.48550...
-
[292]
Kokushimd-10: Benchmark for evaluating large language models on ten japanese national healthcare licensing examinations, 2025 d
Junyu Liu, Kaiqi Yan, Tianyang Wang, Qian Niu, Momoko Nagai-Tanima, and Tomoki Aoyama. Kokushimd-10: Benchmark for evaluating large language models on ten japanese national healthcare licensing examinations, 2025 d . URL https://arxiv.org/abs/2506.11114
2025 arXiv
-
[293]
Holistic evaluation for interleaved text-and-image generation
Minqian Liu, Zhiyang Xu, Zihao Lin, Trevor Ashby, Joy Rimchala, Jiaxin Zhang, and Lifu Huang. Holistic evaluation for interleaved text-and-image generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 15489--15507, 2024 c
2024
-
[294]
X-reasoner: Towards generalizable reasoning across modalities and domains, 2025 e
Qianchu Liu, Sheng Zhang, Guanghui Qin, Timothy Ossowski, Yu Gu, Ying Jin, Sid Kiblawi, Sam Preston, Mu Wei, Paul Vozila, Tristan Naumann, and Hoifung Poon. X-reasoner: Towards generalizable reasoning across modalities and domains, 2025 e . URL https://arxiv.org/abs/2505.03981
2025 arXiv
-
[295]
Audio-visual event localization on portrait mode short videos, 2025 f
Wuyang Liu, Yi Chai, Yongpeng Yan, and Yanzhen Ren. Audio-visual event localization on portrait mode short videos, 2025 f . URL https://arxiv.org/abs/2504.06884
2025 arXiv
-
[296]
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688, 2023 b
2023 arXiv
-
[297]
Visualagentbench: Towards large multimodal models as visual foundation agents
Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, Jiadai Sun, Xinyue Yang, Yu Yang, Zehan Qi, Shuntian Yao, Xueqiao Sun, Siyi Cheng, Qinkai Zheng, Hao Yu, Hanchen Zhang, Wenyi Hong, Ming Ding, Lihang Pan, Xia...
-
[298]
Evalcrafter: Benchmarking and evaluating large video generation models
Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024...
2024
-
[299]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. Mmbench: Is your multi-modal model an all-around player? In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sat...
2024
-
[300]
Omnidiff: A comprehensive benchmark for fine-grained image difference captioning, 2025 g
Yuan Liu, Saihui Hou, Saijie Hou, Jiabao Du, Shibei Meng, and Yongzhen Huang. Omnidiff: A comprehensive benchmark for fine-grained image difference captioning, 2025 g . URL https://arxiv.org/abs/2503.11093
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.