Pith. sign in

Paper Citation Record · LEDGER

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

As of 20 August 2026, this Paper Citation Record lists 100 of 298 outbound references and 35 inbound Pith citation observations for arXiv:2505.04921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04921 v2

Coverage vector

measured 100 of 298 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.332122Z

measured 135 of 135 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:17.693163Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:28.049569Z

Reference resolution

100 of 298 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0ad07dd-b816-42cd-a1b7-e83469af7953 · outbound

This paper cites write newline.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.811188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.811188Z digest=sha256:33896b93a7463d47cf2ad76dda7bfc06b4015d8246c7a7e7e1fd5256e548af5a

Observation 91a1291b-d854-42a9-8163-7475fd2cac57 · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.818647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.818647Z digest=sha256:0f8f6a0929c86407fa92dc0dca1883ed90f5a3b78d96cc9bdf316aebb3d33dc4

Observation 0753eab9-8266-4a2b-b1dd-2e43726b9634 · outbound

This paper cites A Vision Centric Remote Sensing Benchmark.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models A Vision Centric Remote Sensing Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.826609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.826609Z digest=sha256:d1974d88ba33329a435be22ef04e3f23c5d2ee9c825a7a14dea59b625ce6e4b1

Observation de9d64bd-e2b2-49c6-99b7-3b8d982bd2d0 · outbound

This paper cites MusicLM: Generating Music From Text.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models MusicLM: Generating Music From Text

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.832194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.832194Z digest=sha256:23bfbb10cb058f11061f7945de9b8de397ee891a5b57c8ad439ceedef3cc889c

Observation f5a44a29-2cb8-4c17-8dc1-78f83a0dfabf · outbound

This paper cites Ming-Omni: A Unified Multimodal Model for Perception and Generation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.838092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.838092Z digest=sha256:1d003ad88163d15fac541cc337bf16427ed606030311844dfc3d99511b827f03

Observation ffe5edc0-4eb4-41c1-b128-1c7e31112911 · outbound

This paper cites GQA: training generalized multi-query transformer models from multi-head checkpoints.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GQA: training generalized multi-query transformer models from multi-head checkpoints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.843577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.843577Z digest=sha256:5188f68baea5fecf44a6897130ca3010d10971cea80736b11a5f001a2919fb20

Observation ed9d90f8-c783-4f1d-aa89-6d2db0ec3bfd · outbound

This paper cites SBVQA 2.0 : Robust end-to-end speech-based visual question answering for open-ended questions.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models SBVQA 2.0 : Robust end-to-end speech-based visual question answering for open-ended questions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.849201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.849201Z digest=sha256:adedb49947fa7e3dfe0b0c474c61cfbca73e7e6fd132dba0cfcfede24832965e

Observation 8abd5d5b-f358-4d7d-8924-dc5807925307 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.854968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.854968Z digest=sha256:44ccd928a90184a7bcd84ced588f4e5887beb437c59286981cfd46b26b7027f0

Observation 05ac0287-a6b4-4f15-b568-451ae72353fe · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Mathqa: Towards interpretable math word problem solving with operation-based formalisms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.860328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.860328Z digest=sha256:c1796d8ec5b34ae4c1178b9767925596c1104a8092d2d9ce91da68eda0dd91ec

Observation 69092a75-45bc-478a-b916-2ec903181ed0 · outbound

This paper cites OpenLEAF : A novel benchmark for open-domain interleaved image-text generation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models OpenLEAF : A novel benchmark for open-domain interleaved image-text generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.864991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.864991Z digest=sha256:34d46660ff9d0fe2a05da6ed92cb2592da7c7a6f46bc15726f22fb330690f9c3

Observation 59038a36-2311-4332-ba1e-e79d8844ac9a · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Bottom-up and top-down attention for image captioning and visual question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.869810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.869810Z digest=sha256:37f5bf94ef9ef0662b67f96bbd1bdc95f97e1ee62d612fd19339902d7bcd7add

Observation 2eedefe3-1a12-4dc9-8a70-22f215f4aada · outbound

This paper cites Neural module networks.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Neural module networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.874907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.874907Z digest=sha256:e5123236313f112b4baa6619b694c583e84bdd5e85eece9813d45236828693b9

Observation a817012d-de77-4ca1-b6f3-f9aba882d68a · outbound

This paper cites Introducing the model context protocol, April 2025.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Introducing the model context protocol, April 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.879763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.879763Z digest=sha256:d39097cd15f3ca23059fcfd6530f697a3f2db637be020e4b1d7e8fad2def2811

Observation c716147e-05df-4381-9b11-b5436ed5fd19 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Lawrence Zitnick, and Devi Parikh

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.884218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.884218Z digest=sha256:0313fa58d9baa0661aef71a7045f8000c432403f9bf17aa37c8952a68b7e40da

Observation 5b5d3223-0fb7-4f4c-a11b-4b1610fb0632 · outbound

This paper cites Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.889316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.889316Z digest=sha256:d0d907ca35ac35a0c0c7d28f8bd3f627af9549fa0df80b645b6d9a982b3169c6

Observation fc5452d0-9618-4e2c-a68b-8eece14e1ca4 · outbound

This paper cites Tyers, and Gregor Weber.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Tyers, and Gregor Weber

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.894727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.894727Z digest=sha256:29258f892c4f6aef851c168d3cf14fd9426e1fa056dfaf064ab8877db66cb32e

Observation 95f4518d-20c2-415c-81e3-4ae46eb773da · outbound

This paper cites Genesis: A universal and generative physics engine for robotics and beyond, 2024.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Genesis: A universal and generative physics engine for robotics and beyond, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.899715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.899715Z digest=sha256:167f9999f5d42bba393e4fcccf9122787326377f7d1e7a557e3b2094a7af4465

Observation cc538788-283e-4e4d-a4d7-3921e53b85e3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.905247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.905247Z digest=sha256:9a547f51393d2a92a160687d354fd3d92a83c179ba1eb4a346bd2f3103006284

Observation da5cbcd8-03ad-48e7-b678-bd2f4775f3bb · outbound

This paper cites UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.910364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.910364Z digest=sha256:c456727e5807533787eded09a01e3a44d6fd4211dfedebd9ed9f14f1c7961900

Observation 2197da46-132a-4116-ac63-ac48313c804b · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.915560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.915560Z digest=sha256:db86f9ed8da23978e717e7f3cccddf788aae35b8ffb4b38c27a4a03e0b994fba

Observation 91539f90-b20f-4bda-939a-01309eead3a5 · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.921526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.921526Z digest=sha256:a5cce32c7410dc336e726a679d35aefa08e19fb40f02b462bd60e51bdda97807

Observation 4d1d1cdc-4647-4914-a32c-e7eb81253f22 · outbound

This paper cites Beit: BERT pre-training of image transformers.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Beit: BERT pre-training of image transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.926461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.926461Z digest=sha256:1b06c48bbbb399223d6ef3030571e5baf2c834b23ecca1dcd3e174a017ff712a

Observation d77300af-b664-426e-b247-c168a4b15752 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.932120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.932120Z digest=sha256:75978797bd7649d7a0f9badc11d33a39dff7b150e13755d047edb7d803735d37

Observation 18dbab3f-4647-4ec4-b20c-2af90bfec365 · outbound

This paper cites AQUALLM: Audio Question Answering Data Generation Using Large Language Models.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models AQUALLM: Audio Question Answering Data Generation Using Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.937028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.937028Z digest=sha256:961ec63fd407003b7882f98a51b3b13b9d37cb97f57d75cc098d52f3c1441b44

Observation c994d44a-4fab-4b12-b5b0-9fba4236c8ff · outbound

This paper cites Semantic parsing on freebase from question-answer pairs.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Semantic parsing on freebase from question-answer pairs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.942359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.942359Z digest=sha256:67dde766f45e5c86754700d79716646b3817ae0a32f989ce141a450b40507804

Observation 40e552fa-81bc-447d-802c-813a4f8fa3f2 · outbound

This paper cites Why reasoning matters? a survey of advancements in multimodal reasoning (v1).

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Why reasoning matters? a survey of advancements in multimodal reasoning (v1)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.947035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.947035Z digest=sha256:e27c4da2a44fca00072d3ad17b221389aad01ab359627b5e1dce705085737486

Observation 1998264f-89a4-4022-aa1c-54cb5b69af54 · outbound

This paper cites Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.952240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.952240Z digest=sha256:9ef7970f6a372e3916ef88a12203a62afe222600ee409de9bcc7762017940a0f

Observation c313bb05-9935-4fb3-8e46-8c72e953fb01 · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.957074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.957074Z digest=sha256:d99ad47af4316e67d0adc8773a4b060f36c8f4c052f4ffb156d606493485c8aa

Observation de9d40b0-7bec-41af-8995-87d8e10e5cee · outbound

This paper cites CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ).

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.962158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.962158Z digest=sha256:69a7e8bd6d0eb8d85fefc423bdc6eb9af08e51e076e992f42a5e13c9741854c0

Observation a5bcfe1b-48ff-42f7-963d-9576cd04ca1b · outbound

This paper cites an unresolved cited work.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.967316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.967316Z digest=sha256:1f317753c6f6f2db64e4cc86a75f524b42ad9cc87f18fac9aed1779682e801a3

Observation 5e149d3c-4e02-4e76-bd59-ccf9e2f6b061 · outbound

This paper cites AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.972426Z digest=sha256:29c61a58982fda504e1672de4af0e1123adb804b5eb4658af8a136934258f3cf

Observation a08b21d3-4451-49e1-b3b9-f018d1a7b4fd · outbound

This paper cites MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.977588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.977588Z digest=sha256:bc9af99c982382f08d94fdc9e0a99118f919112153ab9ed3c02356dd31bd0dce

Observation 9db43788-3578-4b3a-8e45-9b8e83fcb1f3 · outbound

This paper cites Murel: Multimodal relational reasoning for visual question answering.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Murel: Multimodal relational reasoning for visual question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.983047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.983047Z digest=sha256:b8ee6758f241587593b20adf1b5318ea38cdec82ca4c9969fa259d57a1391c53

Observation 5e409cc1-cd3f-4673-b624-203c3e6f8f78 · outbound

This paper cites COCO-Stuff: Thing and Stuff Classes in Context.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models COCO-Stuff: Thing and Stuff Classes in Context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.988107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.988107Z digest=sha256:c95f7e06ef860749d9f1572ccbecca14707411efc2b092a182db3fc26e8055a2

Observation d731f108-1802-49d9-be03-ddc3c4affbc3 · outbound

This paper cites Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models, 2025.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Mm-iq: Benchmarking human-like abstraction and reasoning in multimodal models, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.993564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.993564Z digest=sha256:08a49aace49d29c6a2ec1d2b80a57c6bf2fd549bd557d0bd5efda802da84116b

Observation 759aab37-0d16-4737-bd57-2539dace4745 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:11.998414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:11.998414Z digest=sha256:13b50ac8dc74545d677c8635516245ee2b18d6d63006f4f79cf4ae61c92d3803

Observation 5c5d6b69-2408-4cde-b9db-40bf07644f0e · outbound

This paper cites Chateval: Towards better LLM -based evaluators through multi-agent debate.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Chateval: Towards better LLM -based evaluators through multi-agent debate

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.003791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.003791Z digest=sha256:ab3dd96d4ca542ecdc4f3d1f703ef885a7a19664b4d8ce050e75c201a813ea54

Observation 843df601-43d7-4e6c-85ee-a936f988832f · outbound

This paper cites ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.008443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.008443Z digest=sha256:cc9afd46e95c5ff2df5ec4ced050c7089583885858f2a09149bf03a6d9051833

Observation dc2f9211-1469-4310-9e56-6077451e8b22 · outbound

This paper cites A Survey of Data Synthesis Approaches.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models A Survey of Data Synthesis Approaches

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.013775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.013775Z digest=sha256:d3b0c73c3bafada09166c754cba97247c5f282738ebe075d70ea2123929d3143

Observation b3354fd1-0959-42c2-a3b6-2e3fb93ab611 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.018928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.018928Z digest=sha256:7ecc41f41e78c46958cc8257c864ed149efbecf9a0b92ac5945f26274df346fb

Observation 41bc3055-6c39-4f7e-98fa-81938b173764 · outbound

This paper cites Mllm-as-a-judge: assessing multimodal llm-as-a-judge with vision-language benchmark.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Mllm-as-a-judge: assessing multimodal llm-as-a-judge with vision-language benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.024063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.024063Z digest=sha256:e09ee4cdb9233f933a948c27f0f93e0a5957e6fc1b31a3aae059b243e9d5277a

Observation cd0caa97-4d37-457c-9657-47a2f6a4da83 · outbound

This paper cites GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.028976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.028976Z digest=sha256:2cd3ab28792f7111d102e20700b041b73fda52595f6c1d082704fb89bc9ae5e3

Observation 5cdf7f91-a350-40bd-a4fc-25d56ad3c9df · outbound

This paper cites MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.034219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.034219Z digest=sha256:736155aede234bfe245bf08e16e1295dcb0e8b365f00dc277fb53919e8e8fbde

Observation ebb9cdb0-ca7f-4cec-967b-cf93217f955d · outbound

This paper cites Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.039359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.039359Z digest=sha256:9a6d59df75b710a8d01583b421e06dcc6b344d455e93456b3ee77978eb8f5354

Observation 415db557-b096-48c8-ae49-f1f428093b02 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.044445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.044445Z digest=sha256:0d762f7fceebf1ff1a494d901703ff37dc91538e9f62e7148f320f8a1837923e

Observation 5a46221f-8bfb-495c-bdfe-71b6c67920b7 · outbound

This paper cites Spa-bench: A comprehensive benchmark for smartphone agent evaluation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Spa-bench: A comprehensive benchmark for smartphone agent evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.049550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.049550Z digest=sha256:01079cd22b479c71c64eacaf9319637e393a581b871af086d63dd4f627fd4843

Observation 2c4c9ce8-7edb-4e7b-b2d0-29d1a29877c5 · outbound

This paper cites R e C oncile: Round-table conference improves reasoning via consensus among diverse LLM s.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models R e C oncile: Round-table conference improves reasoning via consensus among diverse LLM s

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.054756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.054756Z digest=sha256:063a1bc54f95e0e7ebac9e01a9eab8edb27e2cb4204f55b30f8ac0eebeac80bd

Observation f1e0544c-5e10-455f-940c-650e05412a27 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.059706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.059706Z digest=sha256:dde3dd12dd96fff4314b029942d890c63ff4b453f5e63a5a14c30f09d7704e7e

Observation b6a2d442-edd3-43fe-b415-f99bd643a4e8 · outbound

This paper cites G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.064728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.064728Z digest=sha256:ab00e2f293ad99c9c3f6b216bc3f88c74988497176369f742c00a5c334d2789f

Observation cd3ccec9-a395-4171-b19e-c7ecafe883fc · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.070037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.070037Z digest=sha256:0c0806e5ad836317bdde0b32a847cd40246c1f83570bd45e5d981b4fc93b4d64

Observation 4c956736-d011-4c89-800f-6cbb2261f431 · outbound

This paper cites OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.075122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.075122Z digest=sha256:d7b078789efd56da82a6d343850fffbaf3997266cc2802e62a226c03acee0032

Observation 5532dfc8-24b9-429c-aeb4-0df0c37ea199 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Are we on the right way for evaluating large vision-language models? In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.080492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.080492Z digest=sha256:d7c89c4c439e9ab53fe0bd15c0607391204fea99f041f6bfe25055e6c51686f0

Observation 0aeccccd-7d59-4782-8178-f010a22f4db9 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.085421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.085421Z digest=sha256:b93f47808a4d6271c19ba8f02d0ce1b7257d461644636aad58134210c5537b81

Observation 0bb90560-1470-4def-8e7f-4d5588cf6454 · outbound

This paper cites Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.090626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.090626Z digest=sha256:5c0eb6c0267daea6e47ebc412ba58de8539942d5a69c151ca2882259727ac676

Observation 8b8b8fb4-4e33-46f0-bc71-d309f17be26f · outbound

This paper cites RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.095738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.095738Z digest=sha256:fce688a04c745569e391c6f2e3d01e2ac09d9035ca835523ec04ff388a8dc828

Observation 8d3b8de2-eaee-40ec-92dc-6ad72db75456 · outbound

This paper cites Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.100895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.100895Z digest=sha256:5f62f1bca27b86e4970a1f4d789acef8435b7e9b13339e42c3aeba938fde0ad1

Observation f723130b-be50-4c62-b4ac-ab2253118efa · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.106242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.106242Z digest=sha256:c7938178e98666fc5fae13fca06fb404bd17150e06dbb86c5491294e08f79307

Observation 948ca7e9-a83d-4938-a53e-30e0cd99ff5f · outbound

This paper cites Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.111259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.111259Z digest=sha256:2faa46a7cb71736a66b6054cc33bfb2ecd0cc64dd6d2c3eb8896680a9a66df6c

Observation 6cc0b1a0-2e9d-4d2a-ab1e-73a7267e8945 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.116229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.116229Z digest=sha256:a232ff78d699621c990e99a153b6828865fe59f220a3ddc987e7056947fa6701

Observation 23230567-8d6d-4e3e-bfb5-008657d004eb · outbound

This paper cites VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.121274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.121274Z digest=sha256:f5637b5593bd43109793340042111e68b6c4db7eef44a542c5ed29d978575e56

Observation 615e35ec-a315-4034-813a-b29fd827f967 · outbound

This paper cites Rm-r1: Reward modeling as reasoning, 2025 i.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Rm-r1: Reward modeling as reasoning, 2025 i

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.126392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.126392Z digest=sha256:fc7e0728926061504ded22b7270c1ac25d8a608db8b36f83ffa058015b15a63d

Observation 7618190d-598d-4a36-8b72-5873074a2736 · outbound

This paper cites UNITER: universal image-text representation learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models UNITER: universal image-text representation learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.131136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.131136Z digest=sha256:9ea2e4b3e6f262593c825e33feccd5d8384071f2190002a22c29e85187a0f6b2

Observation ae49687e-b698-47a0-b991-196655037c25 · outbound

This paper cites Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.136721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.136721Z digest=sha256:b54ef69e541dd1026f9c25b3862535a91b7f25111a67af2c1089b3763f563818

Observation 3949f0ac-02e9-4c65-b970-85d7ace68290 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.141861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.141861Z digest=sha256:eeaeac1cf87f4f98be2b3533f8272803954efa445ea3ce9a36d7823156b86b0b

Observation aa86a808-9685-4bec-8970-a7443441b959 · outbound

This paper cites R1-code-interpreter: Training llms to reason with code via supervised and reinforcement learning, 2025 k.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models R1-code-interpreter: Training llms to reason with code via supervised and reinforcement learning, 2025 k

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.148590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.148590Z digest=sha256:b3b82bc67e8fb4b3a4d8a4aec30f60a83745bdd604775fc41cf5338d3ae539b0

Observation 6316e50a-17b5-4d4e-b912-a6adfd5d410f · outbound

This paper cites Octavius: Mitigating task interference in MLLM s via lo RA -moe.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Octavius: Mitigating task interference in MLLM s via lo RA -moe

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.153731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.153731Z digest=sha256:c1f53e0927c3e11ef32a054f1373a625abdec634264693979fb21781f14db160

Observation 1ebe6319-e15f-4b5a-a07e-37c37ce79dea · outbound

This paper cites VisRL: Intention-Driven Visual Perception via Reinforced Reasoning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.158630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.158630Z digest=sha256:094b0424336e8ea8284fee5a46b753b63830f066169b0a09df259d159c06b92c

Observation 151ec9df-9d1d-428e-9e0b-572b5d2a9a45 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.163846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.163846Z digest=sha256:3942e7aa4597ed12811764d5de8e456d24c3dce7035bcc48429f199a3edb57ce

Observation ef3516b1-70a7-4167-a7cb-03482615c18c · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.168772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.168772Z digest=sha256:0408415da2c5c8a5bbd1a62ae70079a6ea05000ed712a8e2cbe98cbc38f944ff

Observation dc29b802-f4b1-4b5d-9609-bb59cb2c2ae9 · outbound

This paper cites Unmasking deceptive visuals: Benchmarking multimodal large language models on misleading chart question answering, 2025 m.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Unmasking deceptive visuals: Benchmarking multimodal large language models on misleading chart question answering, 2025 m

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.174029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.174029Z digest=sha256:27f1312b6aab947e67437ed62e79040579eddcb3fb261955d75642af48194f6c

Observation 5d3dbfac-d71b-4238-9b60-ff90448c1bfa · outbound

This paper cites An Adaptive Framework for Generating Systematic Explanatory Answer in Online Q&A Platforms.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models An Adaptive Framework for Generating Systematic Explanatory Answer in Online Q&A Platforms

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.179725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.179725Z digest=sha256:d45ac8c5e0bcddb1656085d29558364f35d3034a53ef16a8f9cc4d776dd71a5d

Observation 73cb355d-1519-4a5c-be06-ae8aa85dfec1 · outbound

This paper cites From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.184814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.184814Z digest=sha256:7f49adcfc785bd307c9d8430965545d21dbfaef330fe3a4d913786317d9f5da9

Observation 752549d0-f078-4e77-8439-63ece57c62eb · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.190411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.190411Z digest=sha256:d1a7687c4aa97ad97242b91146ef79ed59b87bcea515049c222e95021ebf678a

Observation 91d461a9-d926-4de3-9adf-cd58371a0fab · outbound

This paper cites EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.195994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.195994Z digest=sha256:950a142d29cfc4e6c81a558c1696c3935e5996bb5d709a4dd905c365e9eb1bed

Observation fe5c8ca4-39fd-4ff4-9ba4-a7f26e3f0b0c · outbound

This paper cites CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.201133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.201133Z digest=sha256:6006a9c825fcf6a38155ad683c561f1a3ab514e679266875c363419884d32eec

Observation bc66f83b-0978-4db4-8cbd-fa239220af94 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.206435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.206435Z digest=sha256:c9c6d6c9a8a0a4a916fe7b4e7de6dd170c4211c2b64607226b3d0752cdd877e6

Observation 763c1396-5dab-493f-b20e-cf6859d3b921 · outbound

This paper cites EVA: An Embodied World Model for Future Video Anticipation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models EVA: An Embodied World Model for Future Video Anticipation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.211559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.211559Z digest=sha256:24de6318f40e1f435d3da1a070d33adce556ffdd02f17c00a980f902f7f857de

Observation 5e89c04b-c750-406b-a36f-f7790b43f911 · outbound

This paper cites Kitani, and Laszlo A.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Kitani, and Laszlo A

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.216890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.216890Z digest=sha256:d7671c618beabd0711212d0eea4bdbb7b09c015bf4d91be36546424afdd87d3d

Observation 46127784-700b-42f2-974b-4ce9269e11c2 · outbound

This paper cites Merit: Multilingual semantic retrieval with interleaved multi-condition query, 2025 a.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Merit: Multilingual semantic retrieval with interleaved multi-condition query, 2025 a

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.222126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.222126Z digest=sha256:489e697999a7f0bfc0aa07a24ea0e7c35208459d010ddac3a4bc40122030eecf

Observation d0cd0ca3-4dbc-49ee-ac8c-52c7c301d3b4 · outbound

This paper cites PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.227048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.227048Z digest=sha256:3186bf8cb42e4a16322b912e69bf573827678caa3b8b97b2854e6e1964335af6

Observation c6cc558b-797b-4242-9bba-ba500745fcc9 · outbound

This paper cites Agents Thinking Fast and Slow: A Talker-Reasoner Architecture.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Agents Thinking Fast and Slow: A Talker-Reasoner Architecture

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.232119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.232119Z digest=sha256:277e4674525fb0818cde248c79055989f7e2be190aaddd5f60b1006ee818e4e2

Observation 10de2b53-c208-485a-95e6-54b120226f6b · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.237179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.237179Z digest=sha256:cda821cfa899e024c7cd8d8670ba5f91c8b251cc609fd3b666334fe1d8ec1501

Observation 125a5fe9-59da-46af-9599-fa093e32d7ce · outbound

This paper cites v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.242700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.242700Z digest=sha256:9625f0f64c299c48e52b04067f5b3516edef0013b102ce0e00eec3e17c75c0b9

Observation dbf2930a-9688-4b2a-a806-79e1e51bf814 · outbound

This paper cites FLEURS: few-shot learning evaluation of universal representations of speech.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models FLEURS: few-shot learning evaluation of universal representations of speech

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.247690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.247690Z digest=sha256:5a91fb435246a75f5c75feb3717347af268ba5ffd29a9194e897e28d367991e5

Observation 10d53e3e-6673-4830-9865-b696b61ad70b · outbound

This paper cites Receive, reason, and react: Drive as you say, with large language models in autonomous vehicles.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Receive, reason, and react: Drive as you say, with large language models in autonomous vehicles

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.252434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.252434Z digest=sha256:2f37bd69bca303053ef435e34ed091e4cabbb1f946a7e0375f5c1329f3a24992

Observation c0618d94-3641-42ef-bb9f-c18491a1050a · outbound

This paper cites Draw with thought: Unleashing multimodal reasoning for scientific diagram generation, 2025.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Draw with thought: Unleashing multimodal reasoning for scientific diagram generation, 2025

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.257157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.257157Z digest=sha256:ac44538495b486f6c8f7502e0d4f8ebe093dfd085d3cbf6bd1725a9799cae1fd

Observation d255a3be-70f8-42d9-8d24-f07cd6d4d8b4 · outbound

This paper cites VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.263047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.263047Z digest=sha256:e0557ed5885898e2270e09efe1a233209fe8062bf5410407c58d5b151bc6d482

Observation 4c52a636-c6d5-4745-84ab-235d20a4687a · outbound

This paper cites PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.267860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.267860Z digest=sha256:d2c12f2a3b4deda947f261d51092856238fafc7c05b9e2c1244050fa4c957f76

Observation 986e10f2-9de2-4497-b7d3-3128d2a2a072 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.273019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.273019Z digest=sha256:77e1dd82eac45aa7136bff8651ec096f3ffc4dfeef2a1a6a1dcdfdc45622a604

Observation 20a21313-389b-424e-95f4-d4bf0c82cd26 · outbound

This paper cites Reinforcing Video Reasoning with Focused Thinking.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Reinforcing Video Reasoning with Focused Thinking

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.278852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.278852Z digest=sha256:5512581148ddb1f6182d18eb2b2aeaab6a72c68b2e836acfa4c8ae18971d495a

Observation f43cbe3c-13a3-4a2c-bfb5-e690d28dea38 · outbound

This paper cites EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.284324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.284324Z digest=sha256:e04ff19431deb8590d793ff15b9f1c1ba3988469d11d5b7dbb336f0c3389a185

Observation 35ebd730-a49a-4bda-80e5-60a0cd2c4f35 · outbound

This paper cites System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.289641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.289641Z digest=sha256:193e2a86676b47ab7b10b521e171c62c46164228dd9db61353bd4120ef35b676

Observation 39b6ebd5-6118-4dbb-93e6-e6301b0e500d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.296251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.296251Z digest=sha256:e0fad4a7c417db059949bce0ceb9c2597442c80502a729ef4e3ccb36a05712d7

Observation 8bf1c400-807f-4a54-be5e-3c94037cdbc9 · outbound

This paper cites Procthor: Large-scale embodied AI using procedural generation.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Procthor: Large-scale embodied AI using procedural generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.301556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.301556Z digest=sha256:61e2e9ffb43e698095bff5be11573a0be04bac8235f11bfe61d50be8a146db42

Observation 53c76fd7-204a-4c8a-bfbd-3ea0461ca7f4 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Rico: A mobile app dataset for building data-driven design applications

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.306739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.306739Z digest=sha256:a40fcb5a92a27c87f7702260fef0a79820cf58c277c799686a71aee69c850b3b

Observation fd4b099d-9a83-4303-9f6b-2e0043ff5b8f · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Emerging Properties in Unified Multimodal Pretraining

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.311953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.311953Z digest=sha256:2ba958dbe58e5034826cb120e2de2d6652619e571509b7cfbadc2de99442bc18

Observation 69ec2b47-087f-4ba7-b52a-c4cc0a292bf4 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.317779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.317779Z digest=sha256:5110d7fa6d60da2571cc7a44abd4e204c110dd8edd6d50e775d872f0d38320b4

Observation f7b3b485-8d68-4d03-91f3-71911915551f · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Mind2web: Towards a generalist agent for the web

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.322686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.322686Z digest=sha256:06a28d063738c5854027f032684e9b9d25e83f0b9c43c74a0be51617d8b5f6f0

Observation bddc961a-5733-46db-adb5-149c0782994e · outbound

This paper cites Redcaps: Web-curated image-text data created by the people, for the people.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Redcaps: Web-curated image-text data created by the people, for the people

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.327456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.327456Z digest=sha256:cf56a60e9f7a3db8bd2a4ea1fa6044bbff61edd3eada2d15ec75bec6d327605d

Observation 27386cb4-1482-437f-a6d6-c4bb237a47c8 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.332122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.332122Z digest=sha256:54f7b33d93b9cf3e4625117e35f09a86b44b537dbb493e6fa35b8315ad2fb8e6

Pith citing papers

Observation 46b28112-0884-4205-b1a4-760b88ce13e1 · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.543445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:54.543445Z digest=sha256:93d35a74f48031c5762be19a33fd95786ae2142e59138c5797c6866547afcc52

Observation fe3cb26d-1bf5-4cb4-8269-3a815cbc21c2 · inbound

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs cites this paper.

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:20.798959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:20.798959Z digest=sha256:bdfde0d7eaf9c1796a8cac2f01a28a7691e82b25201d88b1491dc899d45fcbe4

Observation 5564ec1e-1a3c-4ca7-9f33-586185119adb · inbound

VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization cites this paper.

VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:36.210901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:36.210901Z digest=sha256:52f2179dd965e00e94d3c5b489b9a957731f2a8b865fa8e7a8ec6a3e9a005233

Observation 7e656524-aeba-4576-81d9-416574124109 · inbound

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models cites this paper.

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:18.442382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:18.442382Z digest=sha256:f0162b465ff6678d720bbde3e55e2fac3508a38f1d659bfba224743e68b2986e

Observation ac921fad-726d-400e-9f44-58e6cc0506de · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:05.297228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:05.297228Z digest=sha256:8de39bbc0c2e27c144ec9743e0b998caa531d84e03d083a4fe20a46872a30c79

Observation b9c93a89-4802-402b-93ce-6ad3eb38d705 · inbound

ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development cites this paper.

ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:32:35.405937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:32:35.405937Z digest=sha256:c81f5944c3cd64f9461eda219f043c764e29cf5f43ee5d40f21db7bfb48362e9

Observation 1330ca35-564b-4e96-b1b6-6053bf94d57a · inbound

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems? cites this paper.

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems? Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:08:24.242469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:08:24.242469Z digest=sha256:f2ecffd05b51db86b12efc2fc54ec463f4d906721374fca9d200fbe6212812c4

Observation aaf87cec-3555-4a75-a339-345378baf007 · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:00.550341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:00.550341Z digest=sha256:6f395dcf2dad54b859dbed6049347f2e08e4d55e3f29323ad72a86d372a28b45

Observation 8bb93239-230e-44fd-a208-d7405e69ce61 · inbound

Where, What, Why: Towards Explainable Driver Attention Prediction cites this paper.

Where, What, Why: Towards Explainable Driver Attention Prediction Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:44.397750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:44.397750Z digest=sha256:5825136f141f05d6fb31989c891843ad42acd9d5f45ec6075cc69d57dac71955

Observation 814105ab-5d6c-4ee3-92ab-bf4575a160b2 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.633699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.633699Z digest=sha256:7cebfe93f85f9fce0f7f0cf0ca53b50b9091a818b825d33769e521594e65d294

Observation 9d919eb4-5b50-4b83-a92f-34d7ca3bfdd8 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:13.370466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:13.370466Z digest=sha256:acf3a8c81e1e56146b725f92a6136aef02db9a377a2d0799c82616318d846e9e

Observation ba5105d8-c291-475e-bc8e-29ac4f2cd677 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.693163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.693163Z digest=sha256:65cb1260a6fbaab8195f5d2656d7c938917f3714b2be6ecae54ccd4332dbed00

Observation bffba20d-c726-44ca-9288-01e43290b833 · inbound

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models? cites this paper.

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models? Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:57.845882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:57.845882Z digest=sha256:a38b1cfad79ee4021f6487037deb8506b4101e9f49a2eb590e08d86644541e8e

Observation e45b5ec3-a23d-49ae-98a7-e04075498a79 · inbound

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions cites this paper.

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-05T16:20:59.685324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:20:59.685324Z digest=sha256:e93fd131594a645072a74708b0be67dd1eb0bd2fb2332690a6c60c2774980ce0

Observation 194346aa-9f27-42e6-b08f-ae432e5de4b5 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.493981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.493981Z digest=sha256:4f50046ceaa0efc23f6d75a63655ef523948893cb812064bfebaaf5802a59d5f

Observation d65d6905-acca-4367-9f24-339f66b6618b · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:25.492417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:25.492417Z digest=sha256:bdd5353bda78ddd6cae5f377dcd7c6b9a01979e70e385f9bd711279ccaa1332f

Observation 616ab416-5f6c-42a8-bd47-e9155f0ad6a5 · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:32:36.361359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:ac496414a102b265d429a7114a125599f30662c5fef74cc31c4d75f091b67cd2

Observation 36a7b0d3-5897-4037-a822-8ca9115891bd · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:11:30.087068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:39341526b5e1566db4914c7f435de679d2504db0c42d98309ea9a0aab205fe4b

Observation ec5d3588-2cf9-4939-b169-9d4cc35722ef · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.298957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.298957Z digest=sha256:676cfa0ce5f987445d3ab2b5a685c7516db3d78d835bd3069993f2da624206f9

Observation 6cd204ca-3d25-4713-9dd6-08f368f410ad · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:56.951607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:56.951607Z digest=sha256:568b9cd7bb503b7ff7c4e4bd62e8e379baed3d9a0836547afbe90d86ba3d785d

Observation 6d1c70cb-d8a9-40b9-b5b8-ec7df293a230 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.552354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:dca79066acd92bc1e8d42b6d3bc67a5023b5dadf7c3d3deb0d2740fe153aeb8e

Observation 438d5b62-a28e-4181-bc8a-a59a304a203a · inbound

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models cites this paper.

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:05.744774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:15:04.261855Z digest=sha256:74e7ef7f1fee26e85cb8520d83c6f93502b613bfceeed2eb1cae862412bd8e7b

Observation 65ac6982-1253-44bc-8b74-f1915ca8ca32 · inbound

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors cites this paper.

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.578653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T04:32:04.547197Z digest=sha256:7fbaf5fd6190cd8a54cf5f1a569f1a8e9de3dcbb7a655da4a8fa4b05e5e1557e

Observation c0c8f5af-aa88-4182-a8e6-315ab58e9420 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.421527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:a6e19d99fc393f7b58adbe7a948918c8d9c46e475ac3c925ce308eed1bf292b5

Observation 5fc18d1e-fa3c-4bc0-9b50-8b7785c4f707 · inbound

NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference cites this paper.

NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:00:16.019653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T02:57:04.813106Z digest=sha256:df8a0775cdbdff24e276d0d77a30af391aacde763f87b01e34aa31d3fe0b1267

Observation f2fe28ed-01c0-4b75-87bb-6b6533eebe8f · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.936877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:3cb7af958e159030e49d776a99de8a4632cd9b6110423f77561e038f4aebff73

Observation 86f6ae4c-1b6e-4e9b-82de-283950f3170c · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.662917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:55f6bfb4c9b91453911b8edd6dc3bc3948c108431bc7ac0b7358ad346cbdec7b

Observation 50254a08-a76b-4ad7-94c6-df15f8d46f46 · inbound

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning cites this paper.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:29.018395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T13:38:01.819121Z digest=sha256:13e52d81db5be8f436c3b79712ccfbc62c6de6821308eb6df07beed2de386ee0

Observation 3db0dc68-3a88-4d5a-b518-32cb56fcfb3b · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.037520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c8fe8ff66c1b8e2a57e7084f5841506ed19b888d636f7f26fa0b3e9f455e7e9a

Observation f137f6d5-c43a-4677-9922-2d3efa2a795b · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.347654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:80bc1dfd4fbdd82e822d3840c2bc71f93722b0316b49bb58f28d954af0cdf4a0

Observation 815cbf74-c26e-40c7-9108-d33440545e1c · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:29:53.387832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:bc0808c1fb1cf4c7809bf30ce16bb20d742df449ecb5eed91fd39b1b763b6327

Observation 9729d942-8013-4e75-8fe9-66e95fc67465 · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.050948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a5642169db4b256671bbf129404e0f626527819d382f83cbff77149f4ea4ac17

Observation 39982139-c36d-4678-a8ab-a2e0f5789ddd · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:5925c18f69b0e3706908ee7f8153d2df98aa7f09a7c8a8fad53ed35b2e2a6814

Observation d79251d8-84ac-45b3-bd19-4b1fb35c38a3 · inbound

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs cites this paper.

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:19.564996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:19.564996Z digest=sha256:146b7171fdf090d524cce75b656fbc02a815dc333cf1d8941e962ba6cfd3573d

Observation 6f897abd-8868-4151-b05f-2845cbbfeee4 · inbound

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding cites this paper.

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T13:42:25.191504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:42:25.191504Z digest=sha256:f3c99516b5da548bc3bc3d3b3daec3a45db46842cd7d1b694a471801146ae454