Pith. sign in

Paper Citation Record · LEDGER

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

As of 17 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 3 inbound Pith citation observations for arXiv:2505.20753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20753 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:33.769358Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:17.886378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.004072Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0c5ce36-7756-4d7f-99c4-076090780625 · outbound

This paper cites Tallyqa: Answering complex counting questions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tallyqa: Answering complex counting questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.496274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.496274Z digest=sha256:01397ac4c98bf522b62afd9b33a16a3d0d61c9f2d1a55de80835a71ff2673ad2

Observation 378366f4-92e8-4dfe-8de8-9cc157e94372 · outbound

This paper cites Macmillan, 2005.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Macmillan, 2005

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.548421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.548421Z digest=sha256:8e015fd4e50f162f67e8610bed64fbfd7e8abe86c79802fd11648664132e4b25

Observation fdf90947-2c0e-48fd-b9a6-444821b05537 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.693636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.693636Z digest=sha256:394e70a8a55cc3b5a59e61f91b96145fd8a69a58c74b6dbeaba2b1bd51a694d1

Observation 31ab5a2e-056d-4f9d-ab3c-b2824f777881 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Graph of thoughts: Solving elaborate problems with large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.758407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.758407Z digest=sha256:1cf7ed2b866bcd946fda1996bf700cd1b90e3a872eec7b89bb8b1fe635cb8765

Observation c7367a42-a8d2-4aba-b715-d790901e698b · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vizwiz: nearly real-time answers to visual questions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.844787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.844787Z digest=sha256:142d5d8a1fc6be028aa311fbfe6f883a543119a9eba7737b2e2b1af06e947355

Observation 27d52929-5863-4763-8f58-0d862d6a24ce · outbound

This paper cites Due: End-to-end document understanding benchmark.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Due: End-to-end document understanding benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.894399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.894399Z digest=sha256:6241ae249ea73ba4ad908e878c3bc0267ebc4b3e80c89ee7fb9da0fde21fb342

Observation a14884bd-d4e2-49a5-babe-6df5e63e82e7 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.976064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.976064Z digest=sha256:f5cc25f9a200a2bdbad44da5290f250e589e1aea8df035c76ffca8294f179bb0

Observation 6f9a441d-3597-404b-865a-d2b17460588e · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.043444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.043444Z digest=sha256:c501607adf6cb1dcb34ad0bdf18fb1ba45419461751c7e4726bea17c260fce05

Observation bda90484-709f-472d-b9fa-185e123158c1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.114407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.114407Z digest=sha256:d49e35442ac216899d418ab6be95abb936641ec1dbbd1b25917787876dbd5c41

Observation 65924100-3232-4961-af83-da8e2ddea302 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.174756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.174756Z digest=sha256:66eb3d0aaf046c5a73c147da039c894f01af4c6e95d84c1325dbdcf29d77bcb8

Observation 4db8759c-387c-4251-be4e-56fc6039b0cb · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.255910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.255910Z digest=sha256:117c245cb64636b4c4a55330e9cbc7cd37c601b812e03b65a610a270f84e4499

Observation 52b4758f-07a6-4826-a16b-2257cfbfeaff · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.325250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.325250Z digest=sha256:6b7639532be258e33e856de390ca35bfbcc8e3f03a31669c3a383d7e7e915b20

Observation 7a573766-a33e-4c3b-a218-901eee60fb5d · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.400173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.400173Z digest=sha256:95a5d0f2fb3b0e425477d1fd8d56951893e41a02f3536c9552ad0ff83a675e4c

Observation 417f4e4c-8ac4-44a4-ab4c-d8069a1cf6ed · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.469521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.469521Z digest=sha256:ae968130b49800ddb93146797835abb48bf9114cd0fd664de48a657e86d99079

Observation 955f3862-cd6d-425b-9487-cacf88b4967d · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.529695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.529695Z digest=sha256:71a43d3d4e4a52e54f8b221addd1e0d6557614f6047a9425c7616f0d9f3e3b1f

Observation df4562eb-e1eb-4298-ab9d-ce26fb8d5219 · outbound

This paper cites Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.586311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.586311Z digest=sha256:20ba8f5d5ec15c6376d17ce6ee80f95d8f93799270e762269d515b456759d6cc

Observation 580d697d-f7a0-417f-9820-53c92980fb3e · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.636978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.636978Z digest=sha256:4ccde8a9381eff59ff1174edfdb57d76bbbd2dbcc8ae9c9799aca4d0938d7e7d

Observation 9d33b65c-aaa8-43f7-ac12-8c5fefa27395 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.701415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.701415Z digest=sha256:9d1c1762e8cb8cb537b2e97cda18ec94d6228a2caa388022b8c3747834918c37

Observation 92a0534e-ba85-4e6b-94bb-03596443725c · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual programming: Compositional visual reasoning without training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.754461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.754461Z digest=sha256:8710626ff3a108b101ab91bc72d8163bc95b06583b34edef1bca4cd90b68663e

Observation 64bfc413-9da2-4864-9bc9-332813dd0d7f · outbound

This paper cites GREC: Generalized Referring Expression Comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models GREC: Generalized Referring Expression Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.818723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.818723Z digest=sha256:a6ac3ebc2a90651a6e292f5af2ce929c5ef7ab0e78350c8eaa7edc6619f00f25

Observation 608935c6-b7d3-4b0e-9d73-7fb1f6999ddf · outbound

This paper cites spaCy: Industrial-strength Natural Language Processing in Python.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models spaCy: Industrial-strength Natural Language Processing in Python

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.893626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.893626Z digest=sha256:00b051dadcd4ac3c63ad4dbe209f1a968ae9651eb4725cdb5bc2f6123c449faf

Observation adcf352b-ff59-4a77-a590-39a1900d31a6 · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.001252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.001252Z digest=sha256:1e149708702adda27878564af70c105ced26d5b47b8972dcd07cb30a4e2a19b2

Observation 5ca6894f-45a7-4e25-b8b5-4e6c8779bbe1 · outbound

This paper cites Hudson and Christopher D.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Hudson and Christopher D

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.112689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.112689Z digest=sha256:eb3061f63feb8798ed956fad580796488ec1c1d2af316735eb4318028d7eac07

Observation e40686fc-98d3-4eb9-b2aa-989e442e7940 · outbound

This paper cites Vcoder: Versatile vision encoders for multimodal large lan- guage models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vcoder: Versatile vision encoders for multimodal large lan- guage models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.187118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.187118Z digest=sha256:9dfdd5331e546b06683b1b7623aa3caa8023584fcdc39e98688b2668215cd114

Observation 5a3bc90a-502a-4109-a130-f4f4218e71e7 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.232788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.232788Z digest=sha256:54962483fc87a2097d14764bc9e48766a43369c58959a32d0e861cf251cee9af

Observation fc28b69b-2bdd-4f01-92cb-1e75ae9a1463 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Dvqa: Understanding data visualizations via question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.291645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.291645Z digest=sha256:9eaeaa6b70f65a178de1d766e2d7566f6f1a3ee0d7ceda6064cdcc19faefea53

Observation 7e8d1255-7b28-40ec-9aeb-c7560ea16bc5 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.375712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.375712Z digest=sha256:1253108ca5d7d8bce3970629e2293d4b622d7e8e3a503e48c8b92cea4148fe2d

Observation 6c46aa50-7e9e-497d-8cdd-58f9336c2bdf · outbound

This paper cites A diagram is worth a dozen images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A diagram is worth a dozen images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.469500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.469500Z digest=sha256:633658de4d28363e0fe902d9ec809a60467b20a1ff0956bbc63e63eff2c5f3d5

Observation b88d2eda-89bf-441f-a912-3f178702cb78 · outbound

This paper cites Ocr-free document understanding transformer.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-free document understanding transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.575290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.575290Z digest=sha256:370adb07dba225a3a93a7fe1cb180a3517a1213fa45c2cd542fb85fc5c1e40c7

Observation fde02abc-151c-4d07-8be9-3d5787e222d9 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.660540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.660540Z digest=sha256:2fd17474f4cfb99e3963fcc08f94c4d54c91d5f85b75caca5e9c2c04847f75a4

Observation 8733d8a4-d5f7-4e6d-8edb-c83824827eb1 · outbound

This paper cites Shamma, Michael S.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shamma, Michael S

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.781128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.781128Z digest=sha256:37351d78c2ad54137fb41097d4b3223b97301fb2ffb3bfbc862dca70aaa1eb8f

Observation d264e598-dba1-4a0c-8182-39169e1c6db4 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.886226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.886226Z digest=sha256:8d758ecff7c026c46cbee67c2770172042cdec19531fa36e0fa1cda0e265f9a5

Observation 435ca4ef-4f62-4aeb-99fb-5769e0e1ef27 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.997699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.997699Z digest=sha256:8d81365acf1db35f5140097d8210b96ba797d17f51f3daf477dfb6e9727d38ee

Observation aaa4e6ba-fc9f-48f5-9b0b-facb283dbcf0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.107442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.107442Z digest=sha256:9d43f67ddb73de39ecdd9ae375f6bb05f2945ad183c0c64fa2af886440a33f1f

Observation a55e395e-5397-4627-8430-f5d8ea9f42d7 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.204884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.204884Z digest=sha256:28b37faa888cf41d2b30e492bf89449577c0337b37fd3b22b7a295abf1b88445

Observation 4f1502bd-9178-44d1-928a-fb4d474499b0 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.323235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.323235Z digest=sha256:f913127833c9a0d42c5859cf745a10fe05bb4fc392e8c5cd30946d8e1c8a3043

Observation 8473446e-f3dc-4a6b-9da8-71901626e012 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:42.125743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:22.424312Z digest=sha256:6f0ef51e3a35d963c40e3bd176fc9b5a56ada639f6a5227e3f3b427beb70fab2

Observation c11fef41-281f-476d-8165-e17a14aa3f49 · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Openorca: An open dataset of gpt augmented flan reasoning traces

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.848288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:22.545417Z digest=sha256:c42bc2f13c015d1c5feb91a162ed57d846413a54d46ce387985a8badfce35256

Observation a2e8eb45-df10-44a8-8d88-9f589bc5a246 · outbound

This paper cites Microsoft coco: Common objects in context.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft coco: Common objects in context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.675306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.675306Z digest=sha256:36ac9e58d14e7e6f72fbab2b66d8b46c4730b6fd4c6f2ba4279495cb707afc99

Observation 32adae86-aba9-41c4-98a9-81111160a3a0 · outbound

This paper cites Sphinx-x: Scaling data and parameters for a family of multi-modal large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Sphinx-x: Scaling data and parameters for a family of multi-modal large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.608715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:22.820414Z digest=sha256:82a80ca739e3e0280be54915b006ae440e06843040b866e58812941b55960d86

Observation 82437ce4-a71e-489f-8774-159e3913f8a1 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.964209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.964209Z digest=sha256:7d1ecd67bb23ce2b2ec46807cd38612d8dacf63edeb2cd6006f7a955f55f0df4

Observation 2e9833d0-f52b-4065-8c01-1211b45fa3a0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Improved baselines with visual instruction tuning, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.130259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.130259Z digest=sha256:9d86bfc881cc6a20e75487c0d7e4539e356816834d5d213b3b6b1693e7766d0b

Observation 63cea63f-afb5-4efd-bb9f-2b4d68737187 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.373071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.373071Z digest=sha256:a2bd16398a750cc24d1b77e7fb06a4c31cf5db60a558115e5d995a60be2e5165

Observation 9b47bb85-f4b1-4233-ab38-4ddb71c2176b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.500464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.500464Z digest=sha256:464f718b459b840e2d7d5de8edf9e72c25c7e8a8b2887cb5ac97afd33f181ca1

Observation 05240b21-1ac2-4421-a709-8474e902b52d · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.608745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.608745Z digest=sha256:c2073ff3a37000c8d434c420337b69c2ef35ec5e10f9557d5a7bbb73f8eeed1c

Observation ea90225d-36a0-4628-b4d1-b53b74070a7d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.686995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.686995Z digest=sha256:1c2c27a3141d9e12a7da7dee5623e5011e23493f9095aa1b3982da1ebf1459c0

Observation e2adaeff-cf3f-4379-9212-7a1a5c377f61 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.781434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.781434Z digest=sha256:7629c84ff220fdc76af6a30f86660cba961a7a933d7ee3e8a9186c447721804e

Observation 1c56508f-e269-44b5-96d9-07a2bec943ec · outbound

This paper cites Decoupled Weight Decay Regularization.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Decoupled Weight Decay Regularization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.884963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.884963Z digest=sha256:e37aaeed7c779047596bb9a0bf4b6c3ba14b7186d6b6d5143c766a1ec34a9ae6

Observation 2d5289b8-a2f0-4180-8ac6-af8bf1b6a8b7 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.995606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.995606Z digest=sha256:8a90fb815ae44390606706203f8c5bf87c3f1649e5b58c14d35d7f493b0144d6

Observation a68b5598-0cb6-4120-bbc7-4155901aeac1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.328177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:24.139378Z digest=sha256:45eb2bf17353ebda6b9f42697386b6a3f9bf827e92d88aff5f3f58a7cb83acfa

Observation 93c279a6-7aaf-40bb-99c5-b7769f486fd0 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.296339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.296339Z digest=sha256:3b0e68faea91351b892c00d5c082d6273373910f05537cfafa9a80f6cb74d2a7

Observation cb7fc0a8-3ee4-45b7-ba58-37f468155d54 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.447099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.447099Z digest=sha256:c9ffe72224f87ba119de1d4a431bc029b359e25506fd1ddb5478ea5142dbb4a4

Observation f18f105c-723e-4f69-bc5d-835bcd223f7d · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.064270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:24.628303Z digest=sha256:a9fb47ad68f28ae63b5246f171bead8f22f73d79c914988c82fdc34a89972e88

Observation 99fdb9fa-25a6-4dd7-960c-7f37b6558ffc · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DocVQA: A Dataset for VQA on Document Images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.790502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.790502Z digest=sha256:0d883d489b5888670badc82364ca70c7722dbc008e40f96678084efc49e37870

Observation c7fb0e1e-04da-47cc-bf80-2bb65dc0b676 · outbound

This paper cites Infographicvqa.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Infographicvqa

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.879645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.879645Z digest=sha256:0210d425cd1b54412780544abc019628664feb191fb072b0a5a50db1be5356f7

Observation d5a71263-a21a-4380-abef-d0154ff3ec05 · outbound

This paper cites Schema theory revisited.Review of educational research, 75(4):531–566, 2005.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Schema theory revisited.Review of educational research, 75(4):531–566, 2005

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.764860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:24.937787Z digest=sha256:2a2b6470aba12578132a7ffe5dfa8a16847ab3a7f4bec30996416be84fbfda4b

Observation f8731295-39b9-4461-a89c-a346ff2d16ed · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-vqa: Visual question answering by reading text in images

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.564102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.026922Z digest=sha256:ad4b612059af750f9966ca58f6047b688893d48e25498faaec1c86f42403015c

Observation 50c017f3-90c0-4df5-a43d-a94f52ab6a77 · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Compositional chain of thought prompting for large multimodal models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.342412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.120302Z digest=sha256:c764662018f2a725d524c9818d41cd69a9892c8b15cfb06f9fd677dbb989991f

Observation c4c16501-60db-4847-a6c4-4a5b7eb6b375 · outbound

This paper cites Modeling context between objects for referring expression understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context between objects for referring expression understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.230063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.230063Z digest=sha256:ca122d9cc5c945271ef44aa130ac96b38076deedd4ae89ab2064aea85aa3aecb

Observation 04158600-31a1-4870-a91c-ace12a753643 · outbound

This paper cites Gpt-4 technical report, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gpt-4 technical report, 2023

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.330170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.330170Z digest=sha256:42fb7c80f2dfef8dba11376118daa334e75d7a3c4d7dd008b8b2a3d486f488a7

Observation 644297db-8b94-4964-ad8d-d9d0c68fa519 · outbound

This paper cites Chatgpt: A large language model for natural language processing, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chatgpt: A large language model for natural language processing, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.097424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.475870Z digest=sha256:7d9566df86739f8d2b01b6b666f9f0a9cff72e234ec66af267427dd6eb8569bb

Observation 8742e0ac-8199-418e-b456-4aa0f5a1a862 · outbound

This paper cites Learning to predict visual attributes in the wild.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning to predict visual attributes in the wild

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.829599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.551507Z digest=sha256:cdbc6bd808ea00f4021e3b3330760afc6a8fa9bb3bb89f4dee32e98e6ce7a8ba

Observation e55e88bc-d4d8-4cd2-9c50-ea5b04d010a0 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Plummer, Liwei Wang, Chris M

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.502500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.619246Z digest=sha256:630a76e205d3ca5d5cf7b8f39530fd9a71f4af4a9488038b28b427f816a6851c

Observation 43f3024b-f2bc-4c52-91c7-2f7eddfb5c10 · outbound

This paper cites Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.286746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:25.770648Z digest=sha256:8b95896417ea7d35114abec61046f90d0e09be0153f8d857c7d786de4bfc01cd

Observation efde47e4-f32c-43f0-8ffa-cd59057f1d59 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.907620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.907620Z digest=sha256:26db1dff047e5cffa54c80409c400af96eeaad91fa95be71b9760402277bf7d5

Observation 444391bf-6e3d-43a6-988c-750be1226fc7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning transferable visual models from natural language supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.963697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.963697Z digest=sha256:745cbe62a7351df53d3240ba02a5695141836e749e6edfc1f6b5eb79f7226733

Observation fc710403-a39d-40f4-a2ec-e273c051da41 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.136393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.136393Z digest=sha256:f1aa05ab9835ea078a698a969acfd712450e6230b4c10a25016724fbfc29e0c9

Observation 2d14563a-fa7f-4563-b0d4-e07d0bba71c7 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.065652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:26.325068Z digest=sha256:4225e1682a36fd0f597a73e76321616d4e1df8cd8947e41051cecc638c535bcc

Observation 40e0e4a3-f065-4fdc-94c8-bfc9dde9bf10 · outbound

This paper cites A- okvqa: A benchmark for visual question answering using world knowledge.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A- okvqa: A benchmark for visual question answering using world knowledge

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.744111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:26.448121Z digest=sha256:90fca2ad6948cee99c8fee808b4f0f9f3baba6c38be349742d620886f6810f58

Observation f446b261-ed47-44a9-b1e3-e3c89900bad2 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.608996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.608996Z digest=sha256:06222ff248a7e20f2a83b839f8218c40c41444dde9f73c30a8bc3fb8b0ff49af

Observation a79b54ec-36a7-41a4-9f59-194a176e5c8c · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Objects365: A large-scale, high-quality dataset for object detection

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.721908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.721908Z digest=sha256:0accbdbd28a9ad99cae1aad0afe008dd2dd2d652d3d5cf0c18861922139ba09c

Observation d96e4623-5ba0-4a39-a2e4-78bb8261499a · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.829722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.829722Z digest=sha256:72e5376275c47566e584ebb016096ed322771a8e8891557e3d0805ddeaa4facd

Observation 33b36a20-3f15-494b-9b7e-85c21bb343e3 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Textcaps: a dataset for image captioning with reading comprehension

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.964879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.964879Z digest=sha256:476d28d51e2eaebb7af1d8acdebf902e81a7b1cf6ab614658fa911a3d072c6d0

Observation 85c83ac2-07e3-4182-ac75-6bc7c05496cb · outbound

This paper cites Towards vqa models that can read.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Towards vqa models that can read

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.138590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.138590Z digest=sha256:97c3187644932d6aa15398a9d34668ce74e4527b200e9d52bbfaabf45e0822d4

Observation 63324793-9ecd-451d-83e5-20c423547bf7 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vipergpt: Visual inference via python execution for reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.298358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.298358Z digest=sha256:85148562b3acefaa508c818a469df2568497495360a99c729e32db7fc4f235b6

Observation b6a94df1-76aa-44e5-b482-72557d258da4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.425640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.425640Z digest=sha256:96a1546f8b6e5a81bc28c993dcd32376c5dd3547008170728b306c61c9b5f90e

Observation f4576e48-1534-43e1-841f-2936f52b26de · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.674640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.674640Z digest=sha256:b0abecafe9aa051418dc9db700c7a3a9810cee6d95f009b36c10b2736a4e47bf

Observation a8556a4e-68e7-4b55-b6a2-e46ddf9619c1 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2.5: A party of foundation models, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.385503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:27.845136Z digest=sha256:c40aae0c1fb1073072f974ba7816f6c4220d9ccfcd5612aedcace8e1867cf26c

Observation 00b24c17-26cc-4e54-9912-f809d9449f77 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.020503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.020503Z digest=sha256:3eb2ed64709516b9a4c5ae7736c723f5558fd008cc0a9336078775968704574e

Observation a7227ddf-c6ae-4f25-a30b-dc91397c878b · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V3det: Vast vocabulary visual detection dataset

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.108520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:28.235980Z digest=sha256:485c9fde424b7b47b46531e79bd558eaf40d8c27e61f3664d8dc6b3ef76df28c

Observation 7ea5dd24-af13-40fd-b158-c0121e0842bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.436180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.436180Z digest=sha256:ab3fc3a83c7ce4002f23ed21f2d0e3715af3e60e6724d6e4d98757e0d2a9e515

Observation 6664e084-8326-4525-88ac-8aa5900c3b19 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.654535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.654535Z digest=sha256:f993f513c550b0db55a815614d24e2d5277e2c17d2e5e9a271a8b6e088270443

Observation 99fb6267-65f7-4ac8-bcbe-8a35e9da365b · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.800849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.800849Z digest=sha256:14e8f55f1fcb1b70f1fa0105c62df2b732ce7fac825006022ee21973f10269e7

Observation 64d076da-b523-4ae9-95ed-dfada6d68c35 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Universal instance perception as object discovery and retrieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.854933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:28.943646Z digest=sha256:7e066ed787c484e13c5994ed38f82154f68de2011905d0a5f20c581c12856edc

Observation ceb3ff64-846a-4a98-bd96-1e607ffe50f1 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:29.090452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:29.090452Z digest=sha256:b7c2fbff396926685288b5656181f0a0dbab239240ae3b97ea10db8c7bc83484

Observation 3bce61a9-aea9-4e95-9caf-c3fdd31c82ca · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems, 36, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:29.715909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:29.715909Z digest=sha256:0bf7dbf0a18d379ccea1d0731068aed54a46424c8c8046692a17b82513950e0d

Observation 47d9c578-d58f-4748-af17-7ea6ee6efd43 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret: Refer and ground anything anywhere at any granularity

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.509787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:31.986056Z digest=sha256:599f091575c1185f7c349d14cb81250753e254ceb80f321b808feee51c7a8de7

Observation 81556e13-6f8a-4852-8d99-066d8bf27378 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.111913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.111913Z digest=sha256:1a563c9bd9d643c7f0ffc17b1c37efd47f168a404472d7071108db63552557a9

Observation c74538d7-e8a4-4a2b-b73e-4170ddac37f4 · outbound

This paper cites Modeling context in referring expressions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context in referring expressions

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.216567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.216567Z digest=sha256:134aaddda1ee22e4cfe5a00853f9e668e481a9e9b2c2e57785cc6a065de09258

Observation 88ee7ea1-47a9-43a7-870b-bd16d1a7f6d9 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.326619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.326619Z digest=sha256:ce125b3344d3f759489bf1f0144843bcdc471817998065c69b12e542f8f84969

Observation aad44938-ec86-46b1-9972-491612c46e4b · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Osprey: Pixel understanding with visual instruction tuning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.463254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.463254Z digest=sha256:e1f1a55dc9bd1bee3143929502327be251b6dfa8e2c624fc466f520495e9caf5

Observation 1db65240-c2c3-444a-a8f7-6bbfac99a870 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.600372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.600372Z digest=sha256:791b0cdefdacdf608f00bd0dfbdfa092d93c37837cd6eef27732ed79dab311f1

Observation 60d458aa-0cd3-4836-96de-c49499754562 · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Griffon: Spelling out all object locations at any granularity with large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.294153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:32.713818Z digest=sha256:4010768efe7c5c2e8ee857417f3cde41d298650217ac7626e6a7e3ef8ba785fb

Observation 250fc1e9-7210-495a-9cb2-b8a494f739d5 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Automatic Chain of Thought Prompting in Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.834055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.834055Z digest=sha256:ea237c1c3acc9e040658423fcf63a77de6909d4727e28dbd78908585477dff89

Observation 43994f56-3434-4817-80de-4d69d73512ac · outbound

This paper cites yes” or “no.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models yes” or “no

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.033887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:32.967018Z digest=sha256:b28f9ad48946c47b9df4a98daad8b0d1f6fb1c21fcbcda6b3863b384e78ae6f5

Observation 420e162c-67b3-4f6f-9d11-b5510053330e · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.655327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:33.117627Z digest=sha256:2cba5ea7a8a1418c020d30d01c99e8436754afbd50ae9c42c67905f17b907e53

Observation f776522e-df36-4c05-9dbb-8041809823c6 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.188645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:33.381162Z digest=sha256:e20d60f7e3e44aa8c059b707521376754fe843e3ca34205c3fefa4a2c01aa57f

Observation 8fa7385a-6ddc-471a-ab2d-f46df1dbc8bb · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:35.804909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:33.586439Z digest=sha256:9db5acd94c7740fd863a2007a46777c53c79fbd6a0f6b596bb58bf0e5813527f

Observation af56e164-1237-4768-94ab-da81c32ec363 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:35.445068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:33.665448Z digest=sha256:6c0b066202853b41addcb261d9058b74e5e21d825c8e6ee837043ddc9c23e4aa

Observation 1d7846c4-193e-4223-9e09-2d8cbd0b1552 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 102

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.375321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:53:33.769358Z digest=sha256:6abb44ceaf4b6dc0887b0d69034646629e6fcc982f225f8213d13247cd31fe4b

Pith citing papers

Observation 02361d15-17e7-43df-9bbc-8057663b10f3 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.886378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.886378Z digest=sha256:ee4f1a5e9312cea53218e43c4d3c95850cc357d57fd2f7d8f36d56b04d600dc8

Observation 5d8824a4-823a-4543-b995-e55d68069cfd · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.006201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:a7bb28d04d4e84bb199f539f672dd658da458ccf54ab39d1370cc5c250fa22f9

Observation 5b38d76a-528f-4b16-b509-d2292a23dc31 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Reference 267

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:7299fba589becd18b3e342d93717fcf92deb629ce0debde8c09391170095126c