Pith. sign in

Paper Citation Record · LEDGER

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

As of 9 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 2 inbound Pith citation observations for arXiv:2505.20753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20753 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:33.769358Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T04:17:40.198357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.004072Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0c5ce36-7756-4d7f-99c4-076090780625 · outbound

This paper cites Tallyqa: Answering complex counting questions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tallyqa: Answering complex counting questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.496274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.496274Z digest=sha256:abab1534079375aa368f0159ce3791f27e9154a2de98353a39766ef12274921e

Observation 378366f4-92e8-4dfe-8de8-9cc157e94372 · outbound

This paper cites Macmillan, 2005.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Macmillan, 2005

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.548421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.548421Z digest=sha256:5e0c396a6edef246fe85847392dcb3bc1fa4fea72ec2dc379cccb1c51cc42697

Observation fdf90947-2c0e-48fd-b9a6-444821b05537 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.693636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.693636Z digest=sha256:827f43d454e0235ca3675ae5b76665fae77016f2c839ced0bc1ec36296d8e573

Observation 31ab5a2e-056d-4f9d-ab3c-b2824f777881 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Graph of thoughts: Solving elaborate problems with large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.758407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.758407Z digest=sha256:34dd17da1f7ceaa126700deaf30a50322cd50d7a623ec864340c5103b1fe0347

Observation c7367a42-a8d2-4aba-b715-d790901e698b · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vizwiz: nearly real-time answers to visual questions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.844787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.844787Z digest=sha256:537489eabae52680d782ba70feaf30c22ded3cf65ed63539618712a0676adf00

Observation 27d52929-5863-4763-8f58-0d862d6a24ce · outbound

This paper cites Due: End-to-end document understanding benchmark.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Due: End-to-end document understanding benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.894399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.894399Z digest=sha256:dc8a92c685beeccf17dd54434a7f76acb68ca1eb70d9759d240b80177fd04656

Observation a14884bd-d4e2-49a5-babe-6df5e63e82e7 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.976064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.976064Z digest=sha256:f822a25c2c4a3237569e7beebb7cea092e5f56cfd46ca864d187112ef47eaeb4

Observation 6f9a441d-3597-404b-865a-d2b17460588e · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.043444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.043444Z digest=sha256:21c7ff9c9edcdaa342455224b0406619560250e3eb2c85004a448afb470bde71

Observation bda90484-709f-472d-b9fa-185e123158c1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.114407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.114407Z digest=sha256:528dc68d0a2f7ba9369882e597a28e9069312a508ff8f9e990339653ac3f17cd

Observation 65924100-3232-4961-af83-da8e2ddea302 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.174756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.174756Z digest=sha256:b0f139ea0e704e348d9ae48e1fa97bc62826fe8610028177ec0c9fb50a5d05ab

Observation 4db8759c-387c-4251-be4e-56fc6039b0cb · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.255910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.255910Z digest=sha256:48ad153d6c97d39ed3ac22229286c3baa84d9cfeb6be96bd299e86efb8baf45e

Observation 52b4758f-07a6-4826-a16b-2257cfbfeaff · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.325250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.325250Z digest=sha256:3999c2d4e1c014a6ee8ececa7a17310bb911296a16f3ce3ae1cd8eb2fbc63954

Observation 7a573766-a33e-4c3b-a218-901eee60fb5d · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.400173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.400173Z digest=sha256:231e0768f4685aec8cba5715e31063935eccf2b66bed007cc62cd0d6f2ee93b6

Observation 417f4e4c-8ac4-44a4-ab4c-d8069a1cf6ed · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.469521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.469521Z digest=sha256:c65051353ed96ae010baef848855538ab40e06b6702429c10fc7c0ba565bbbcd

Observation 955f3862-cd6d-425b-9487-cacf88b4967d · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.529695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.529695Z digest=sha256:843a0520230a57e1541b067aec033ac9a6cb17cf4d747a87e3935b9376e3ff28

Observation df4562eb-e1eb-4298-ab9d-ce26fb8d5219 · outbound

This paper cites Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.586311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.586311Z digest=sha256:3ad54319762ca73e52fa453eef83ca496c6c65d9f51c8e4c08d0727d82502662

Observation 580d697d-f7a0-417f-9820-53c92980fb3e · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.636978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.636978Z digest=sha256:4898e85b1a5a8b30c11134a2ceeea1b15b497e1d4ece96dd3a0fe621af8d9216

Observation 9d33b65c-aaa8-43f7-ac12-8c5fefa27395 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.701415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.701415Z digest=sha256:431764c795c33fce3012910d15ec0cbb8beccd3f32c4d37e3eb8d9df663a0e8e

Observation 92a0534e-ba85-4e6b-94bb-03596443725c · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual programming: Compositional visual reasoning without training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.754461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.754461Z digest=sha256:63801050d35c5d041fbd28fabb495a2df50e89935b0d7fe6d422bb97fadfb7c2

Observation 64bfc413-9da2-4864-9bc9-332813dd0d7f · outbound

This paper cites GREC: Generalized Referring Expression Comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models GREC: Generalized Referring Expression Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.818723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.818723Z digest=sha256:c6ce4cde7a8072a3f0b889e89421a4e26f2df692c8e5e033210bd1b9dc10605e

Observation 608935c6-b7d3-4b0e-9d73-7fb1f6999ddf · outbound

This paper cites spaCy: Industrial-strength Natural Language Processing in Python.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models spaCy: Industrial-strength Natural Language Processing in Python

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:20.893626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:20.893626Z digest=sha256:a21923663d8b51a5f3aea7f5572f774d38913e83cf2f0f010221d7e1c0236a7a

Observation adcf352b-ff59-4a77-a590-39a1900d31a6 · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.001252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.001252Z digest=sha256:c1652c0ea836b389a9108de264dfecc8e9a0bb7512c3ea1ae84046b1b97684df

Observation 5ca6894f-45a7-4e25-b8b5-4e6c8779bbe1 · outbound

This paper cites Hudson and Christopher D.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Hudson and Christopher D

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.112689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.112689Z digest=sha256:2e470d8a09486b052b50690c0d912e3160c9d0901df2e981499e847f3b181c95

Observation e40686fc-98d3-4eb9-b2aa-989e442e7940 · outbound

This paper cites Vcoder: Versatile vision encoders for multimodal large lan- guage models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vcoder: Versatile vision encoders for multimodal large lan- guage models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.187118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.187118Z digest=sha256:2875c87aa65685ad8f784ad6c86af44d73a457f402b51389eca44abe2f3d6d42

Observation 5a3bc90a-502a-4109-a130-f4f4218e71e7 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.232788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.232788Z digest=sha256:85f887bf96bc63af524fded8c1454c7f326f42d78067b0cf7ca85716125112db

Observation fc28b69b-2bdd-4f01-92cb-1e75ae9a1463 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Dvqa: Understanding data visualizations via question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.291645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.291645Z digest=sha256:60dfc7f6c419bfbf82a9cc78481190a3a494da6adc8a124bfb3a1fa2962aaddc

Observation 7e8d1255-7b28-40ec-9aeb-c7560ea16bc5 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.375712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.375712Z digest=sha256:af066cd1eb40f9ec535dbe01c1a6ed46772882b10a571db0e91b76134a531bdc

Observation 6c46aa50-7e9e-497d-8cdd-58f9336c2bdf · outbound

This paper cites A diagram is worth a dozen images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A diagram is worth a dozen images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.469500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.469500Z digest=sha256:426874e52ec559ab0d577960defa096ec1325026686f55270370e8deecda6eef

Observation b88d2eda-89bf-441f-a912-3f178702cb78 · outbound

This paper cites Ocr-free document understanding transformer.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-free document understanding transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.575290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.575290Z digest=sha256:6e4c1e6d5f11ef888cbbbe83dc4147ea3b240a968a8aa49e116143f7cd11e8cc

Observation fde02abc-151c-4d07-8be9-3d5787e222d9 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.660540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.660540Z digest=sha256:e97cf21c13a7a47d15fc31786bdf518f97792e4d1c0f23fbf26cd79b9ff256e9

Observation 8733d8a4-d5f7-4e6d-8edb-c83824827eb1 · outbound

This paper cites Shamma, Michael S.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shamma, Michael S

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.781128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.781128Z digest=sha256:326200adcbe50c6a41221c0a17c3208a61aee35ce0d1fb0619178d012bedd5a7

Observation d264e598-dba1-4a0c-8182-39169e1c6db4 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.886226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.886226Z digest=sha256:6b0038eed128d0bc1447d66eb52e5393f85e4750489918280ebdbf59176a73ec

Observation 435ca4ef-4f62-4aeb-99fb-5769e0e1ef27 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:21.997699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:21.997699Z digest=sha256:9f4aa6292d39c63f5ddcb653038a5e53e2e8aab0f5147c4212ebfe2d8114fdfb

Observation aaa4e6ba-fc9f-48f5-9b0b-facb283dbcf0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.107442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.107442Z digest=sha256:89f6c7dc45fb7479a72249894bbf62153c2181256a37f4cc49c8b8cb53312c0a

Observation a55e395e-5397-4627-8430-f5d8ea9f42d7 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.204884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.204884Z digest=sha256:b674a4a8d2b7f0a8d0564800bacf27cadd4c54853124af082c779b7baeb01070

Observation 4f1502bd-9178-44d1-928a-fb4d474499b0 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.323235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.323235Z digest=sha256:d038d961224970c5d77a488019ab8e33cb0bf7c32a7f891307c0dc3eef766d92

Observation 8473446e-f3dc-4a6b-9da8-71901626e012 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:42.125743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:22.424312Z digest=sha256:69c5b75c8bd03b420a0aefcb1008918b5d8e798146406c163f9479515bcb01ad

Observation c11fef41-281f-476d-8165-e17a14aa3f49 · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Openorca: An open dataset of gpt augmented flan reasoning traces

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.848288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:22.545417Z digest=sha256:b6a49cc963b0acba2eedf4f725b2c2ee66f682b1ee2a0bc94c86f8b5d0db8b26

Observation a2e8eb45-df10-44a8-8d88-9f589bc5a246 · outbound

This paper cites Microsoft coco: Common objects in context.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft coco: Common objects in context

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.675306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.675306Z digest=sha256:88276142ebd1d5a997775d3be975be3ef946975e0218910782050e36902422a5

Observation 32adae86-aba9-41c4-98a9-81111160a3a0 · outbound

This paper cites Sphinx-x: Scaling data and parameters for a family of multi-modal large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Sphinx-x: Scaling data and parameters for a family of multi-modal large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.608715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:22.820414Z digest=sha256:ccfb8ccf0a964e7c330922d807afc1fa705f99d6ca95dbab9cc2869a7d23906c

Observation 82437ce4-a71e-489f-8774-159e3913f8a1 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:22.964209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:22.964209Z digest=sha256:3f9780f3d30bbbb49ef4dbe2a496852b4dcc0a64517bad91e179ed637dd78dba

Observation 2e9833d0-f52b-4065-8c01-1211b45fa3a0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Improved baselines with visual instruction tuning, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.130259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.130259Z digest=sha256:37a3cb93b2527b653dde6554c54a2110a037a88bb8e2237f3abac3e8c02a319d

Observation 63cea63f-afb5-4efd-bb9f-2b4d68737187 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.373071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.373071Z digest=sha256:9302e558158c7895fb39d45d598f156a12999e7c500c4dba8e4dd680ce1aeae2

Observation 9b47bb85-f4b1-4233-ab38-4ddb71c2176b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.500464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.500464Z digest=sha256:1ab84c711e8d593504180baaf83d1180e4fd8e1dc5a9452e33eccf3a84d74ba1

Observation 05240b21-1ac2-4421-a709-8474e902b52d · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.608745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.608745Z digest=sha256:261a7f76b26981da0d8266d94ff5bbaac31b734007c5f5d745fde6f202d19203

Observation ea90225d-36a0-4628-b4d1-b53b74070a7d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.686995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.686995Z digest=sha256:8a0a753efe45652f148689ca1f261556466b8e5fe28dbde4823ff16a4f7b2b93

Observation e2adaeff-cf3f-4379-9212-7a1a5c377f61 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.781434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.781434Z digest=sha256:c9e94a6a5d3e00acd18556654f1571a133cd05d604db0d902416f992c9405398

Observation 1c56508f-e269-44b5-96d9-07a2bec943ec · outbound

This paper cites Decoupled Weight Decay Regularization.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Decoupled Weight Decay Regularization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.884963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.884963Z digest=sha256:4e849e947947411954947f853b0d486ba05d39dbbab257ece15832790cd16911

Observation 2d5289b8-a2f0-4180-8ac6-af8bf1b6a8b7 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:23.995606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:23.995606Z digest=sha256:25875474a00ff823d079e552dbb61ce94b7890ec629752b0789a46e85b832b6f

Observation a68b5598-0cb6-4120-bbc7-4155901aeac1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.328177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:24.139378Z digest=sha256:9216e52467dd7f04ec6b2c39c663a4fb32f8b9f3a415db3c3efee248753cb338

Observation 93c279a6-7aaf-40bb-99c5-b7769f486fd0 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.296339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.296339Z digest=sha256:5c59bcde97c0790176bf3c682ccbf34f17ede25ec4b2be2eeecaeef36971ebc0

Observation cb7fc0a8-3ee4-45b7-ba58-37f468155d54 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.447099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.447099Z digest=sha256:37b69bd74adf12058fb8bf3efe12aa45450ad7e512ebf746563dcd69405b4c69

Observation f18f105c-723e-4f69-bc5d-835bcd223f7d · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:41.064270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:24.628303Z digest=sha256:fa6f6bfea561ca294b0ecd8eaecad82f4f069107cf14e14a48bf5843bef29b44

Observation 99fdb9fa-25a6-4dd7-960c-7f37b6558ffc · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DocVQA: A Dataset for VQA on Document Images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.790502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.790502Z digest=sha256:f52ca2aa534253f33cfa03e32e3f56a5ab359a4a28562314869f8ff249b0cba8

Observation c7fb0e1e-04da-47cc-bf80-2bb65dc0b676 · outbound

This paper cites Infographicvqa.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Infographicvqa

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:24.879645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:24.879645Z digest=sha256:6d9796b61c97f6f5352b0e38b4d5002e903b3390cd0cc93db7073043bf448c4b

Observation d5a71263-a21a-4380-abef-d0154ff3ec05 · outbound

This paper cites Schema theory revisited.Review of educational research, 75(4):531–566, 2005.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Schema theory revisited.Review of educational research, 75(4):531–566, 2005

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.764860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:24.937787Z digest=sha256:87b3cb14a00df1aca52116ca5579cdb6688890540758d439d48bd56b6efc2b42

Observation f8731295-39b9-4461-a89c-a346ff2d16ed · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-vqa: Visual question answering by reading text in images

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.564102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.026922Z digest=sha256:dcc6843fc448b527a654014d1f480368561f53a2d09e83108472650f2ba90df8

Observation 50c017f3-90c0-4df5-a43d-a94f52ab6a77 · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Compositional chain of thought prompting for large multimodal models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.342412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.120302Z digest=sha256:28837a4532a730909807f96cbe3f3ea8ef15e5025315963c8b486235719d3898

Observation c4c16501-60db-4847-a6c4-4a5b7eb6b375 · outbound

This paper cites Modeling context between objects for referring expression understanding.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context between objects for referring expression understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.230063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.230063Z digest=sha256:e946dda7bd69afa2d5337509d1894805960639708a1736530d63849531b39cd0

Observation 04158600-31a1-4870-a91c-ace12a753643 · outbound

This paper cites Gpt-4 technical report, 2023.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gpt-4 technical report, 2023

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.330170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.330170Z digest=sha256:debebb0c449d119fe64f1781b48cc91e597e888039f7317c13f6e87c3c15aacb

Observation 644297db-8b94-4964-ad8d-d9d0c68fa519 · outbound

This paper cites Chatgpt: A large language model for natural language processing, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chatgpt: A large language model for natural language processing, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:40.097424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.475870Z digest=sha256:170961a70cca10ee4b2999961af6eedf65ae1fc7d933ea6e01e58f3c120f9f12

Observation 8742e0ac-8199-418e-b456-4aa0f5a1a862 · outbound

This paper cites Learning to predict visual attributes in the wild.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning to predict visual attributes in the wild

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.829599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.551507Z digest=sha256:0773d3ce7ea68dba3a81b8723bff00857cb48a6a4e9c197f0c6185453fc5ba40

Observation e55e88bc-d4d8-4cd2-9c50-ea5b04d010a0 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Plummer, Liwei Wang, Chris M

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.502500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.619246Z digest=sha256:0b700d90b174c56b100bcc7c8c1c3bc9a895b1dcc88102eccec60dd6ee002b71

Observation 43f3024b-f2bc-4c52-91c7-2f7eddfb5c10 · outbound

This paper cites Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.286746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:25.770648Z digest=sha256:a7132b184d5abaae1044e2cb13197baa76df12b36ba1eb128460f13f93c2d6f2

Observation efde47e4-f32c-43f0-8ffa-cd59057f1d59 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.907620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.907620Z digest=sha256:71e48504a2b2ba096b9936c849d2f3460549eaa3af30823dae927a78380dacef

Observation 444391bf-6e3d-43a6-988c-750be1226fc7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning transferable visual models from natural language supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:25.963697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:25.963697Z digest=sha256:f7bdb3d1a7a12c7acb3c874014d38ca23bc07021eccf306510812948e5806e1e

Observation fc710403-a39d-40f4-a2ec-e273c051da41 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.136393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.136393Z digest=sha256:14d3c44d81d6ddc21b02dac445b9d9abdb317bf7a189ada41b40eb42c889aaff

Observation 2d14563a-fa7f-4563-b0d4-e07d0bba71c7 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:39.065652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:26.325068Z digest=sha256:f60a4a91e33f24bab3fb6691c6b08c85cbb979c7d3bb4ac355738b96e5319488

Observation 40e0e4a3-f065-4fdc-94c8-bfc9dde9bf10 · outbound

This paper cites A- okvqa: A benchmark for visual question answering using world knowledge.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A- okvqa: A benchmark for visual question answering using world knowledge

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.744111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:26.448121Z digest=sha256:8e60dd1a7aec44f2ebc9b6ddb3d664aa435f7987dd94f85a349ca19ad3ff7443

Observation f446b261-ed47-44a9-b1e3-e3c89900bad2 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.608996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.608996Z digest=sha256:872500efc46ca76294fc8c7024d101eb4da232454d79d98dea64ec94d2a8e2bb

Observation a79b54ec-36a7-41a4-9f59-194a176e5c8c · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Objects365: A large-scale, high-quality dataset for object detection

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.721908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.721908Z digest=sha256:df75dba54aab16c122914d51a3ecea4fbc57a0cb99bef1017451756590b12864

Observation d96e4623-5ba0-4a39-a2e4-78bb8261499a · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.829722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.829722Z digest=sha256:67c11944e3e86dcde473443096edcb05ff92e2c8954de1dd220dc679841edc1d

Observation 33b36a20-3f15-494b-9b7e-85c21bb343e3 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Textcaps: a dataset for image captioning with reading comprehension

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:26.964879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:26.964879Z digest=sha256:d9a46afcac7a67a72072832cafdf165f21fb45af312f719a4738a7698b4301f0

Observation 85c83ac2-07e3-4182-ac75-6bc7c05496cb · outbound

This paper cites Towards vqa models that can read.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Towards vqa models that can read

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.138590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.138590Z digest=sha256:5354c56ec70251dbe45a09fe66feadd7300e30c1fcf8477854b24f77abfb75d1

Observation 63324793-9ecd-451d-83e5-20c423547bf7 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vipergpt: Visual inference via python execution for reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.298358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.298358Z digest=sha256:f3ed41cd56490ae2ef6ed23162a9a33598f6a80beb92f1beaff1506d1377ec62

Observation b6a94df1-76aa-44e5-b482-72557d258da4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.425640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.425640Z digest=sha256:bef1f9a7168ddba441af773b5e356a4f2ead9346efb7b61c30f8d57186b4c09c

Observation f4576e48-1534-43e1-841f-2936f52b26de · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:27.674640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:27.674640Z digest=sha256:9132f9debba27eaac5c3ea2888fb8b1c336c6b6cd2de8c40bae391d9f683b7f4

Observation a8556a4e-68e7-4b55-b6a2-e46ddf9619c1 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2.5: A party of foundation models, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.385503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:27.845136Z digest=sha256:22104cf0f77b7cccc4c12b193a3421205739f7d93c90fee55754ec086fe3d8b7

Observation 00b24c17-26cc-4e54-9912-f809d9449f77 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.020503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.020503Z digest=sha256:e70691d825d7c8d7f49e8e92b2c5c47d7a29356ff064a14d45f3ee5b9b3faf83

Observation a7227ddf-c6ae-4f25-a30b-dc91397c878b · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V3det: Vast vocabulary visual detection dataset

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:38.108520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:28.235980Z digest=sha256:1ccb94f913fe21e30baefee42cc0e2a44171b161544cef07098ade5794b36a45

Observation 7ea5dd24-af13-40fd-b158-c0121e0842bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.436180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.436180Z digest=sha256:1dfd29f2dcaa66233fb80ec5b3396fbe20bd721676d9f0ac1ca5fd179b126045

Observation 6664e084-8326-4525-88ac-8aa5900c3b19 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.654535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.654535Z digest=sha256:046c7a3c093b78adfa253f05ad267bf857c06e3d0b4a025b6b1fbcee8fd5e993

Observation 99fb6267-65f7-4ac8-bcbe-8a35e9da365b · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:28.800849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:28.800849Z digest=sha256:d7c2cde3a28e3b8b6342c2b09719ee79a0b5137b6713fa959c71bec833e41617

Observation 64d076da-b523-4ae9-95ed-dfada6d68c35 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Universal instance perception as object discovery and retrieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.854933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:28.943646Z digest=sha256:fa10d7c8738123920ed513d90f95b052552c9e5c3a699252cd225e0db037d3d9

Observation ceb3ff64-846a-4a98-bd96-1e607ffe50f1 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:29.090452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:29.090452Z digest=sha256:37d76746e57fa8d3e6ed51ec843e844e570955c684162daddfb046df0c62fec9

Observation 3bce61a9-aea9-4e95-9caf-c3fdd31c82ca · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems, 36, 2024.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:29.715909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:29.715909Z digest=sha256:3f8e9bc0182c705fcbe02deb5d61362c883fd290d2c88146a39a801b856c0964

Observation 47d9c578-d58f-4748-af17-7ea6ee6efd43 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret: Refer and ground anything anywhere at any granularity

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.509787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:31.986056Z digest=sha256:6f462e9eedb9af22e4d977eef5197dafe32a1aa2b3683a93b547e7dc1e2f26ac

Observation 81556e13-6f8a-4852-8d99-066d8bf27378 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.111913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.111913Z digest=sha256:b747849d73085084b787911e18f3c1fab06f73a4482aa260c87f54c64a9f39d8

Observation c74538d7-e8a4-4a2b-b73e-4170ddac37f4 · outbound

This paper cites Modeling context in referring expressions.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context in referring expressions

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.216567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.216567Z digest=sha256:a873d5f0ac4fb0199d838142ee6eb44e6b52112b39acfe6f9eee2ffd9a8f0508

Observation 88ee7ea1-47a9-43a7-870b-bd16d1a7f6d9 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.326619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.326619Z digest=sha256:83fb90be3109d99b4515c712fc60d7b1c677f31b54bdd9313c14a33dc73ad0c5

Observation aad44938-ec86-46b1-9972-491612c46e4b · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Osprey: Pixel understanding with visual instruction tuning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.463254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.463254Z digest=sha256:8f24d3b40c04d38ff99e706b3d8bd09f6ba5ce7b27874c79d91caf83d8262b47

Observation 1db65240-c2c3-444a-a8f7-6bbfac99a870 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.600372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.600372Z digest=sha256:b1fdb350c890b3db07a09bac88f477dcde2896a3366fdfa08a8b39c3bbfc5a10

Observation 60d458aa-0cd3-4836-96de-c49499754562 · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Griffon: Spelling out all object locations at any granularity with large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.294153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:32.713818Z digest=sha256:7ba6163b206ae03ca32a35f5f7dbbe0b6f594c17933a24219c1c8c0aa470c5a2

Observation 250fc1e9-7210-495a-9cb2-b8a494f739d5 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Automatic Chain of Thought Prompting in Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:32.834055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:32.834055Z digest=sha256:fa4888d8930673b99175e42df110435fb1bafe0eec2226b06c0657c859e9914a

Observation 43994f56-3434-4817-80de-4d69d73512ac · outbound

This paper cites yes” or “no.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models yes” or “no

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:53:37.033887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:32.967018Z digest=sha256:12b95bc783a94dfbc62f8f0ba48e5aa1359a67ac752a7a40848671cb545087a5

Observation 420e162c-67b3-4f6f-9d11-b5510053330e · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.655327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:33.117627Z digest=sha256:62bb41d51617cc9dee809876e3f12e9e8f977a12357bb1f5cc131faa0d3cdb44

Observation f776522e-df36-4c05-9dbb-8041809823c6 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.188645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:33.381162Z digest=sha256:b9d36e047d7c01130028dd08c9653f0e5e0336d6eb59e7023e2d1e19a84bf910

Observation 8fa7385a-6ddc-471a-ab2d-f46df1dbc8bb · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:35.804909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:33.586439Z digest=sha256:d64212c3918bf59db0da41771a38d82dcb418fa0b7b7270a6a3ad4fe0d6cd968

Observation af56e164-1237-4768-94ab-da81c32ec363 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:35.445068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:33.665448Z digest=sha256:8079fd28d034ccd693e74dd734520267402cb2e3b8e6c428a0d079814d0e0a45

Observation 1d7846c4-193e-4223-9e09-2d8cbd0b1552 · outbound

This paper cites an unresolved cited work.

Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work

Reference 102

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:53:36.375321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:53:33.769358Z digest=sha256:5333b05a4c8febd744fef7382512072ac6c3d267c74b8cdb6b27d35bac542e56

Pith citing papers

Observation 5d8824a4-823a-4543-b995-e55d68069cfd · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.006201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:eb9560b8ef9e1b383e0e3931a13b2a371ecd583c5b4b877419ab998e7a4aac78

Observation 5b38d76a-528f-4b16-b509-d2292a23dc31 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

Reference 267

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:9a18973bcd83bc57710f88bfe59b8917221efaf059df93ccb16394f9bd4718ea