Pith. sign in

Paper Citation Record · LEDGER

Explain Before You Answer: A Survey on Compositional Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 100 of 266 outbound references and 18 inbound Pith citation observations for arXiv:2508.17298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17298 v3

Coverage vector

measured 100 of 266 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:17.882743Z

measured 118 of 118 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.785540Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:52.515977Z

Reference resolution

100 of 266 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf9a193c-382c-40e5-bdc0-23b094869826 · outbound

This paper cites A Benchmark for Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Benchmark for Compositional Visual Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.513735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.513735Z digest=sha256:27cb44088866953d584d3437691bde9ac6cb8008fefc2c6176b331e0d8f72206

Observation 83611022-30e1-43e8-9896-77fd7aa45eae · outbound

This paper cites Take A Step Back: Rethinking the Two Stages in Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Take A Step Back: Rethinking the Two Stages in Visual Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.518180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.518180Z digest=sha256:e3c6d3515805aaf1dfb5928317802447e2b6f9af73120b7bf51f4e94170bc4cc

Observation 924adcb7-85f5-44d2-b420-1bf109c57bac · outbound

This paper cites The Book of Why: The New Science of Cause and Effect.

Explain Before You Answer: A Survey on Compositional Visual Reasoning The Book of Why: The New Science of Cause and Effect

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.521694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.521694Z digest=sha256:47ee7c4617f9c4a041208cdd6f057eed6bd7c54e2574e8405a8206a2b826cfe3

Observation 5bd42100-d56e-4ac5-a5df-668a858ae1f0 · outbound

This paper cites Inferring and Executing Programs for Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Inferring and Executing Programs for Visual Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.526199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.526199Z digest=sha256:34b2beac42c8cb3a85dc934d029fe0c4fa546b70bb64d6e6bca12c461bc6afb3

Observation 70ba1066-30bf-407e-8f16-35b1c2b99abb · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.530233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.530233Z digest=sha256:12668595f18a1e8f56b7d874effa760d6337e1379598a3869a3a5d34adc29457

Observation f44fd22e-c38b-4cdb-a66b-86b7bf678bee · outbound

This paper cites RA VEN: A Dataset for Relational and Analogical Visual REasoNing.

Explain Before You Answer: A Survey on Compositional Visual Reasoning RA VEN: A Dataset for Relational and Analogical Visual REasoNing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.533928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.533928Z digest=sha256:77a841fd08f186b18547e79ce7842a563bccdd95a0696d7b0a09bfeb02f56a57

Observation bf4e9e29-bbb8-482e-ab01-1cdef4a157c6 · outbound

This paper cites Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.574350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:09:17.538323Z digest=sha256:71c7f19e0eec4dd361b3078e58594c287223da8921087dae0e55acc1b6fa4e67

Observation f9d54651-546f-40c1-be8f-cc8b0b4c4ba3 · outbound

This paper cites Maintaining Reasoning Consistency in Compositional Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Maintaining Reasoning Consistency in Compositional Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.542217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.542217Z digest=sha256:ef1355d4f26adaa53c563047bc11e9be540f695926d0ac36d11408a457155db4

Observation 5f3077c1-cef7-4d19-835a-8a9cc3205176 · outbound

This paper cites Visual Instruction Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.545552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.545552Z digest=sha256:9b238ebd1191ac0816dea94b2709053a269a8c35e8a85a480a4aa1e16e9a7b5d

Observation 0b3b3d39-4673-414f-84e2-c0fb42fc0b6d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.549166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.549166Z digest=sha256:04092d50d495a9a2e2ab9a23cbd33bca39c2438da3c034edd208491474185e8b

Observation 719d8d4f-3419-4e9f-ab68-3fc22040ea3e · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.553400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.553400Z digest=sha256:edfd3957382d06a9c2bd6b32b6d25336394ea0bd4f9dd99e7c4e2ac82f00bb5b

Observation 9372ac24-3045-4847-a1fa-c748255303ca · outbound

This paper cites InstructBLIP: towards general-purpose vision-language models with instruction tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning InstructBLIP: towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.557143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.557143Z digest=sha256:4392bf3c8fbe07ea9684ddbe0d4afb11dbee00aa70b566ca9a4ece74b35c5cd8

Observation e174f2f3-7221-4093-849e-c6e360e8dc2d · outbound

This paper cites A Survey of Multimodel Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Survey of Multimodel Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.560723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.560723Z digest=sha256:a90aad2c51727de0316594757623db0cdc4fa968ba951fe9007615eee77070d2

Observation 6e70e9c5-698a-4ac0-a3df-92b461415d20 · outbound

This paper cites Interpretable Visual Reasoning: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Interpretable Visual Reasoning: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.564558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.564558Z digest=sha256:c0e322bc0db29aa6d167297c47f7bcfe62b43b69c999c158ec1a8a5091b12a8b

Observation e0a07f5f-c192-4958-a691-081a514198a5 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.568659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.568659Z digest=sha256:38fdf11ed4c3708716edd98867ca12ed2a888877d6bc36c02afb76b9d69b82ea

Observation 0e2b5f02-331d-461e-aadf-b9937102a53b · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.572947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.572947Z digest=sha256:299645cc17197c9c9de01900fa3d257c1eeaa4593e223b366ff7bcd60e65054e

Observation 512ec9a5-7708-4f86-bf24-29a0b4fdf499 · outbound

This paper cites Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.576724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.576724Z digest=sha256:6cc099c32b8a7ea10112c7c14fd80ffc7a4fbeb8ca05f5eff84eb342a41e17e2

Observation 815e063b-a98a-4958-ac2d-0ddd06033f11 · outbound

This paper cites NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.580278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.580278Z digest=sha256:8dbcd147e44d95d8be6ec23b3ebcdb05e1796a7a9c951f304706baa2087e9171

Observation d43da455-eb95-4631-abf3-20648f47a5c2 · outbound

This paper cites UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.584057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.584057Z digest=sha256:e03d55de7f0226978e0e454b362d9bcc7f1fb68944f91e82ea262385e39f7a66

Observation 987e6fbc-1c5d-487a-92b8-82eb754d8265 · outbound

This paper cites A Study of Visual Reasoning in Medical Diagnosis.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Study of Visual Reasoning in Medical Diagnosis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.587663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.587663Z digest=sha256:0d1e8cceabf3aae0a3637dc236fd7a38debf15e5de175c2c814ac71bfb8f68c0

Observation 96a9982d-3315-4005-abf4-39a1116d614a · outbound

This paper cites Medical Visual Question Answering via Conditional Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Medical Visual Question Answering via Conditional Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.590944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.590944Z digest=sha256:40859333243734b9f253dfe0393a5c9ac031ce4f84850d92e698d397b10c26f7

Observation 79ab4d51-8731-415c-9a36-9f3079477b69 · outbound

This paper cites Programmatically Grounded, Compositionally Generalizable Robotic Manipulation.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.594301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.594301Z digest=sha256:107597eb0d85b5afe95c8c197ec063881bcd866728389c43ff661d9161ebd20c

Observation 33a7b91a-c907-4811-8035-29075a1d6879 · outbound

This paper cites Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.598754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.598754Z digest=sha256:e6db415325871ea04c3b710c619b01eb749adf2af62f01e4bea30c6220e54467

Observation 297a2486-68bd-44b4-8619-b25f63c90620 · outbound

This paper cites When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,.

Explain Before You Answer: A Survey on Compositional Visual Reasoning When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.602375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.602375Z digest=sha256:eca60a72f6892aa60988b13f8b30ee8070140ada9e24e4068c06df62ab331b25

Observation 99030bc0-26ea-4050-a685-58bd8e373122 · outbound

This paper cites Challenges and Prospects in Vision and Language Research.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Challenges and Prospects in Vision and Language Research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.605883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.605883Z digest=sha256:90eb84b843933642a0213c4917ad87c1ac5d87e1a11e02391590179b558e4546

Observation 196cf821-1c40-4821-b251-93dbff6c6ea7 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.609224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.609224Z digest=sha256:3162ec11239d3fcad853c0b20bbb85890bfc03ac67ee8f04ebaf0516da58b5ab

Observation 2fd5befd-2d65-4ff3-a00c-4e9608088e1a · outbound

This paper cites Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.612998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.612998Z digest=sha256:f83a46bbb77b7be7291c925f4a2775ccb68254210386a0abc86fd3bf92be6083

Observation 4d087fe7-1755-40e9-98ea-16743f97646e · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.616473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.616473Z digest=sha256:cff271290af922b7f3f2a87056c246f05fe782e585b4cbce85b634ca6a5ac574

Observation 3daeb3fe-c31d-429a-a273-1e2da67a7da6 · outbound

This paper cites HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.620594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.620594Z digest=sha256:bc80db70d3dced0b4a8f8f277cce36e7e54cea5f04a1346c37427e64bce0f073

Observation 53e0aaa8-566d-42d4-9045-52f301a33acb · outbound

This paper cites Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.624093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.624093Z digest=sha256:63ca9bc2f60d57e6db66395144d22e7b40a4c7f107a730240734b44f83748a29

Observation b1388ad6-63e2-47fc-baad-db756fcd9eea · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.627597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.627597Z digest=sha256:60306261c15243b345df2e896d41f76f25dba30bd031531edad1bf2792660a59

Observation 8d1ceb2c-3880-4c4f-b55b-82c3818c16cc · outbound

This paper cites Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.631786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.631786Z digest=sha256:ececeaa4e11e64a17af1cdde5d10ca77e2ba4c8f0cd51db72a16433d700ac8f6

Observation 85856c28-b570-4790-9b5d-f4e1b6ca1e72 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.635381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.635381Z digest=sha256:567bd603f55f306450dcbdb7fd7d821e36902fb4278f63dd0ed81930d1b55d6c

Observation 64d1eec0-1361-4820-ad73-17f8e09fad9b · outbound

This paper cites Cascaded Mutual Modulation for Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Cascaded Mutual Modulation for Visual Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.638828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.638828Z digest=sha256:8c5a7652517da4ab3a7e57a09efbaa46547962ba59eb0bf9790b0b64415288a2

Observation a2df30ad-5b05-4301-9b66-9afb134bfb95 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Mul- timodal Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Compositional Chain-of-Thought Prompting for Large Mul- timodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.642304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.642304Z digest=sha256:3684d33432a1471b77fb907ae2b3942714cdc2f3225720b8d9b2ea2451207dbe

Observation d784b18a-9cbd-4deb-90b5-9eda1f65b23a · outbound

This paper cites Building Machines That Learn and Think Like People.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Building Machines That Learn and Think Like People

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.646029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.646029Z digest=sha256:5b01615dbb83b2f8745dc51b1b7d0f4aa24cecef99007e38cc7f57ebf7e69206

Observation 918a7128-311b-4c4c-ad4e-de30c815e0a4 · outbound

This paper cites Meta Module Network for Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Meta Module Network for Compositional Visual Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.649583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.649583Z digest=sha256:e273889f929e573da57204caeec1e1d5f1eff31f1c3cc9c09a5e4cc49afbe676

Observation 3d198a3e-97a6-4c85-a11e-aa27679b8ef0 · outbound

This paper cites Improving Visual Reasoning Through Semantic Representation.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Improving Visual Reasoning Through Semantic Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.653229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.653229Z digest=sha256:d70146d0871176f70c87759eb163b61372b41b18e6d04b64708e17ff5bff0715

Observation b25d1405-e144-46fd-ab8e-b15b530a418f · outbound

This paper cites What’s Left? Concept Grounding with Logic-Enhanced Foundation Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning What’s Left? Concept Grounding with Logic-Enhanced Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.656988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.656988Z digest=sha256:6f53568652e911ffe5d448aa8bd6e5b89d4ee263a13ed8da54f2e2339de5fa5f

Observation 683be0f2-2665-4239-afec-bcb50526e037 · outbound

This paper cites Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.660571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.660571Z digest=sha256:a7f0fa1e9381e9c2a2c25b66b1746b4d055b231952620bf9a0835a8569dfe0fa

Observation fc3fc673-c443-42a2-b420-165676db93b1 · outbound

This paper cites Visual Question Answering: a State-of-the-Art Review.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Question Answering: a State-of-the-Art Review

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.664117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.664117Z digest=sha256:90447bddaa08c10121f9e7f933636019328e9f28f6a89c331c6e36cafe1ea634

Observation 2d033639-8414-435d-8263-eeebc82ec403 · outbound

This paper cites Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.667724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.667724Z digest=sha256:9673279fa24581de6dc3b80c6842ef0563ef749a869dc9d5fe4d0f964b2c8a15

Observation e4977987-e095-4c48-8059-9b2636c812a4 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.671253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.671253Z digest=sha256:e475450c5b9bf2627c1cb2d57b6a03ead92258a18815e75f6a0afbb3f77ee10b

Observation f4ae5747-05cf-4036-bf44-e2481bc8e361 · outbound

This paper cites A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.674787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.674787Z digest=sha256:b35a2c5f61c74f5fb04ccdbfae6ba36e8a124a018539aa4a52312a391488633b

Observation c960d985-fdbd-4ba9-abb0-3c3a02cd2d7c · outbound

This paper cites Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.678225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.678225Z digest=sha256:c5519f9d1c3e8d0b7d97b541bfb46bcf8f2a4172729b8b284c199fc7cbd2ede9

Observation 06ab9bbb-a02a-4267-a8bb-86804d9c61db · outbound

This paper cites Large Multimodal Agents: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Large Multimodal Agents: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.681651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.681651Z digest=sha256:bf2722e30f4e634142da85e49c6eca98b5e1b12422da1a23556b7eedbb9b146b

Observation bbef4951-f830-443c-94dd-8e38f18c843d · outbound

This paper cites From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.685501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.685501Z digest=sha256:d8e23f4c0660923eb6f7a51e30954c284e248eaef6e6c291cb8a7ed288b51b52

Observation 1c15a348-49a0-4fee-8ca8-1f237413e057 · outbound

This paper cites Robust Visual Question Answering: Datasets, Methods, and Future Challenges.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Robust Visual Question Answering: Datasets, Methods, and Future Challenges

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.689858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.689858Z digest=sha256:dd63c56e6594bdcd0f3e4103c79e8bf8adc7020f41b30531df5119403ba0b7c8

Observation ba5105d8-c291-475e-bc8e-29ac4f2cd677 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.693163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.693163Z digest=sha256:3c5044cdabd9918bb70c03739a6a1428658ba42787dd0de7c40b47a804d83590

Observation 94d521e2-9d73-4f3e-81d2-107b850daa29 · outbound

This paper cites How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model.

Explain Before You Answer: A Survey on Compositional Visual Reasoning How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.696947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.696947Z digest=sha256:a76bc84a8f3d500e8785154be605a0a1d1faad8998f50251af11eeb7b4f362ce

Observation deb20ad6-fd0f-4386-b1f4-d13d21265ad4 · outbound

This paper cites Tool learning with large language models: a survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Tool learning with large language models: a survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.700806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.700806Z digest=sha256:b85cc6d82cf14521471fbfbf6a9090ec99615c00401bf7b8f5f637b967dbd64f

Observation 77ab8f79-80a7-451e-8d54-432ec300cc0d · outbound

This paper cites Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.704220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.704220Z digest=sha256:f4d6b6351019f31aad1ec75bd43412598c5c474bd34f9fa61f9d48ecb58230c3

Observation 02920ede-5fbd-4de5-b835-f6a418c41248 · outbound

This paper cites Multimodal Intelligence: Representation Learning, Information Fusion, and Applications.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.707897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.707897Z digest=sha256:47fc91c0bee53f49e359a6e3119d0a8d0183637ae27ba06b24074ea30513be66

Observation 2f2659f7-bea5-4572-afa9-8c2b0efb2b45 · outbound

This paper cites Transformation Driven Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Transformation Driven Visual Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.712111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.712111Z digest=sha256:59767d1923c5c6d226b51841b4c81dda4d6132de9939704e1350e35399e19574

Observation 1b81e6d5-89b6-45c5-b512-b315b8302a38 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Residual Learning for Image Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.716819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.716819Z digest=sha256:b09f2879c49a1fc6cc28dd36ce4ed4a27a59788eff915905fbdbf561dac5609b

Observation b1010c49-1a83-43c7-bccc-18c96da62785 · outbound

This paper cites Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.720344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.720344Z digest=sha256:703fdf85b4a43ba3788188b1d72c3839ca2bfb2d9248fe37915251673db01d30

Observation ca4d0ffc-6df0-41e3-ab35-cf09a055772c · outbound

This paper cites Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.724028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.724028Z digest=sha256:438b82ee958c6f26bf7990093f1853922cca70be4fc54ba992172ed35ee4f72c

Observation 90671fdd-138e-4c0c-9b78-b6022419b97b · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Explain Before You Answer: A Survey on Compositional Visual Reasoning You Only Look Once: Unified, Real-Time Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.727468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.727468Z digest=sha256:a3761e2ec8839b19eb56057b67a18d5d52dbbc7d55eb184412bb32a2a86d313a

Observation 688bf382-7594-46c3-b2a0-26b482a034c4 · outbound

This paper cites Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.730761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.730761Z digest=sha256:79fa0656d0082e4bbdc33fbda7eb98beb44c84257a7f12196efb893c251dd2eb

Observation 8dde71b0-5067-4b54-af20-57d9a50dc076 · outbound

This paper cites Bilinear Attention Networks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Bilinear Attention Networks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.734056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.734056Z digest=sha256:fa644220d9b5b8ae31faba179c30a1b87a98bcc0a0e9ce6ebfa0ac8d793d957e

Observation 941cd9c1-6dfe-4474-9d73-c0242e7b12d4 · outbound

This paper cites Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.737492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.737492Z digest=sha256:735ed04603cfc01e2fb8fddeedd3a05f7c684f65723bef5bd0c3dd24a1a537b1

Observation 062c41e1-e85c-4bc4-bd55-c6aedc3761e8 · outbound

This paper cites Deep Modular Co-Attention Networks for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Modular Co-Attention Networks for Visual Question Answering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.741248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.741248Z digest=sha256:a42e339bd99c26b1dbf801b163ab39ef12b96b4fc8c54b71d88390d3d69953ed

Observation a0c7384f-0162-4842-a675-56c7f7180120 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.745210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.745210Z digest=sha256:73d9c903715d0c21e95883c39ecafcbbcaedfafef2fa6231eecfae8e10f4df12

Observation a2a5f4f7-200d-47c1-959f-5b217f210384 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.748896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.748896Z digest=sha256:a44e62b1c7dcdbc632620d04ccec277b6b4c9f1d4999446e5087260ac2154901

Observation b4ac0030-bd44-4619-9630-5854e783e3f8 · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning UNITER: UNiversal Image-TExt Representation Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.752421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.752421Z digest=sha256:1f5e2ae165bdf70afe128003cce3ebd2c67d50c2ad3002bfb727c49d7f0e0468

Observation 55445a2a-d4e0-4e52-b950-feee7d366620 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.756195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.756195Z digest=sha256:279f1b4f43353fad87a4d5b2f585f658f0f02c9e5acb0591bf0398da6841daca

Observation 57c32c0e-3d8e-406c-be76-bceed02b541c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.759948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.759948Z digest=sha256:0d5c3703540db98b0d2204640246cc5520d43b43358c75e08f8005602891b881

Observation 2b23ce4f-6699-4b99-8f80-0196f07275b9 · outbound

This paper cites Otter: A Multi-Modal Model With In-Context Instruction Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Otter: A Multi-Modal Model With In-Context Instruction Tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.763717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.763717Z digest=sha256:b8eb91e121c58ab4bd805f89d513f6f3c8a51d1f1323d19f3bd84e97fad4f2c4

Observation f7ed05ab-d852-437c-b3cd-3ff03a1c8f2a · outbound

This paper cites CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023.

Explain Before You Answer: A Survey on Compositional Visual Reasoning CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.767580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.767580Z digest=sha256:87bed41fd47f7e54bc7951e0608dd8effc5f682f3e965b76faad13a03448c331

Observation 84fb288c-bba1-47aa-91a0-ddd47bc24e45 · outbound

This paper cites Logics and Languages.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Logics and Languages

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.771070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.771070Z digest=sha256:051e1e0061c0e29add9708b2c9822baca40067785b792e4f8454fef6826f208f

Observation f4ad16d1-0bed-4eee-960b-6027a2807b91 · outbound

This paper cites From Recognition to Cognition: Visual Commonsense Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From Recognition to Cognition: Visual Commonsense Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.775107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.775107Z digest=sha256:029f5d77ee467756855b4fb4279fa6d1d55bab64847f5534496036e0ea0e7bd3

Observation baf19ec6-f9e9-43a4-bb37-7391120e77c8 · outbound

This paper cites Visual Programming: Compositional Visual Reasoning without Training.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Programming: Compositional Visual Reasoning without Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.778592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.778592Z digest=sha256:937f8dea9991398c4d482c806763aea8541e2c6b0bdef940e7f5399500c94c81

Observation 99f6c90f-0741-4e80-bb46-f36ffe1334c4 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.782737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.782737Z digest=sha256:0149c101faf0b6982a5528da23ed3bafb8e83f937e50bd39b01cbf57302a6e3c

Observation 2bc9b9a0-95bb-4356-8361-fd0d53e5ac2c · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.786176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.786176Z digest=sha256:09f70251d16e038828a5017ed4a4391654974f77ab1459792450abf307216802

Observation 42dac3b9-737c-4faa-9eaa-dd61249f1bd7 · outbound

This paper cites Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.789732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.789732Z digest=sha256:cd724f1a2cc222747c89674345926c2bd6db59377ecc9e40a27972b9ea0f31d0

Observation 1e1b648b-eb9d-4cd7-a375-5885712aef3d · outbound

This paper cites DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.793148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.793148Z digest=sha256:90682c8c05d55b4a090923efd6108a3d072c6de96c0c2118af64ea65ea444b4f

Observation 219d2711-261d-4c2a-8821-b4e16c67407e · outbound

This paper cites Iterated Learning Improves Compositionality in Large Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Iterated Learning Improves Compositionality in Large Vision-Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.796395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.796395Z digest=sha256:1a7262f40ee2425d1c861141ff6f14121732a5844f9b8896cfb738f98d5ba8b4

Observation 0be00509-5529-45a1-9744-ea2ffe4877ee · outbound

This paper cites Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.799892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.799892Z digest=sha256:904e941e41fc94e9bec208fe4f6e13b40c875da074a26a531557157cd4ebfea5

Observation 402ae171-e88e-405b-8586-0d8f42ab9c48 · outbound

This paper cites Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.804131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.804131Z digest=sha256:fd043c0155a143e92ad48e6958e13fa33ddb75d265de6560dc80c70a392a936d

Observation fd81ad9a-57cb-4db9-a2e5-1c4001c63617 · outbound

This paper cites an unresolved cited work.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.807533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.807533Z digest=sha256:e2eb9e9f1bb1af1d116ea42d2512d83d4631c924441d708b6aa832a8f4138b90

Observation 5003c3c9-2241-4807-bf99-040ceabd0e34 · outbound

This paper cites Visual Compositional Learning for Human-Object Interaction Detection.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Compositional Learning for Human-Object Interaction Detection

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.810619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.810619Z digest=sha256:1d2549a49365756374df4072b480c4cc715a39b79266e3df839a0d28f0c1cb20

Observation 768452f4-e668-425f-9073-9215c305e181 · outbound

This paper cites Learn- ing Visual Composition through Improved Semantic Guidance.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Learn- ing Visual Composition through Improved Semantic Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.814504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.814504Z digest=sha256:bcdecf7fce27d8f30e5fc9d40a4e47aef5d21401cacc91c8d62a2c8380210fd5

Observation 192f2e11-89cd-4f08-a990-e135ce90ee3a · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.817979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.817979Z digest=sha256:62281158ebd39d7c3af839216d50fdba79abddd525223e07d2d566f35070543a

Observation 8930522c-42f9-4904-9d26-9b80d0fa34ba · outbound

This paper cites Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.821769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.821769Z digest=sha256:6ca93b36a48b2589980c9b5f9ca0f4705258c97cb230bafabdb4432bfb2355f1

Observation 987437d3-de0a-4659-b914-f0228fce6e58 · outbound

This paper cites Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.825678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.825678Z digest=sha256:4f879fb8b217746ebfc880b924c041f4ff7c9197d8e046c656d628721561e150

Observation 278a00d6-4fee-4601-8e17-672634a531d7 · outbound

This paper cites DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage.

Explain Before You Answer: A Survey on Compositional Visual Reasoning DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.829651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.829651Z digest=sha256:fa9de2238b7d960808b0d84d91c745ef124aed75d97b7fd734f952eeaadce5ba

Observation 29a18e16-e797-45cb-9f18-3830955ce0c2 · outbound

This paper cites Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.833925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.833925Z digest=sha256:5489b660bc706e7f01da00e550642add88476b2964c78538c3505839c62192a2

Observation b9db1b69-797a-4db1-95d2-c85d21c96437 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.837525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.837525Z digest=sha256:7452a15edd8869b4a5f20cec5be4612e41096f10c4c4a14b36970f460069be5b

Observation f2878906-a8f7-4076-99eb-00848f53b154 · outbound

This paper cites Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.840955Z digest=sha256:5cd8a868cca647e7f9e9ff60cfad931278eba87f988d7a01a2aae45f40f9aafa

Observation 9da7e67f-463c-45f7-bad2-459093ead803 · outbound

This paper cites From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.844636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.844636Z digest=sha256:0892f447a477aaf4d207b3ec521cd66da88ed1c0596649fe1817dae0cab5d188

Observation 604da5ec-4610-4637-958d-0ea41dd807fb · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.848322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.848322Z digest=sha256:fb3a50666b7f717a37cad715c3679ba9de6465def3dd80c96eef181023d32ac1

Observation 61c5613e-7010-4d73-8542-7ce9f4cea9a7 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.852321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.852321Z digest=sha256:992dd86568cffac174801a16ef4ce103f0c32e2cbdd0c4e460acb4ec0d9d493e

Observation fb4e0f90-1e84-4ef2-953a-0144669a021b · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Explain Before You Answer: A Survey on Compositional Visual Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.857411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.857411Z digest=sha256:56f40c3290274bb616288ff199ec007a5e1790911f8a21b50049ef2d830c8efb

Observation a50aad1e-035b-4fc0-94c0-b621b2a5082a · outbound

This paper cites Grounded Chain-of-Thought for Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.861449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.861449Z digest=sha256:37bd37a2de92d9571d8be1c9f01f457eea4a31d66ab1aa3aba7b4c606ecc6ad6

Observation cd5e1b7d-bb3d-48b5-9f69-0480c2afb333 · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.865349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.865349Z digest=sha256:99304175c8ee013c7beb6caa18d779aa0bf257c66c48aeaeab21a66de2a7de19

Observation 4b9dacbf-46aa-4491-8e98-36ae4e4777fe · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Explain Before You Answer: A Survey on Compositional Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.868949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.868949Z digest=sha256:44adbb026ea60ff45a68b2647216d903b7ad2be1e0c9427e20220bd91fbf3f40

Observation e37c4b36-3475-43e2-af63-4f2ab3164986 · outbound

This paper cites ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty.

Explain Before You Answer: A Survey on Compositional Visual Reasoning ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.872314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.872314Z digest=sha256:bdf42bc56efe54ba27518adfaffaa793adfc7d6641ba917fc39f3ac62a20ac5f

Observation bb85a1dc-0e06-459c-8797-1cbb510a0bbb · outbound

This paper cites Visual Reasoning by Progressive Module Networks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Reasoning by Progressive Module Networks

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.875627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.875627Z digest=sha256:299548b9b0bd697d37acc11d7cecb5b2033e121df0fe9c5f5e31078c299cfb11

Observation b1c5f0c2-31fa-45fc-8e88-0d21b6eb56a9 · outbound

This paper cites Predicate Hierarchies Improve Few-Shot State Classification.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Predicate Hierarchies Improve Few-Shot State Classification

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.342455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T17:09:17.878832Z digest=sha256:29ba49545b5d0b15c33515d48e62951ffc4f00b5b97090f69393d37fcc338cc8

Observation bf64629d-8068-4973-8e84-c06fc4d03311 · outbound

This paper cites Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.882743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.882743Z digest=sha256:3f4313359b78921543528a7b1ae3264bb1a08c476f90f8fcc05d9ce62e354fa4

Pith citing papers

Observation 29394350-4395-427a-a682-5f18b5a73c2a · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:3081952bb57765296410414dc7b81a48be3a54f9d11883d2100f669d12c7430c

Observation ebf4352f-6fef-4a35-8e24-3dbbcb426783 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:29.860215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:29.860215Z digest=sha256:dd9b7e3c89171c09cf10503870d78c0d3b8a5cb353a83c8d86d35fcc15d67b5f

Observation bd14d4ff-deff-46a7-baa7-ea0fadf4891e · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:02da041b326c14dfb420859ab746a1941c46a6189ecc52c96805302437c03ff5

Observation 0c472c11-a78d-4401-866d-41d28e17055c · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:67c267f4bd3f8794a3bf526d0174764803b43915bb1435f6a71ede55ac95458d

Observation fb7f47e8-e6dc-4614-bcff-b50c7d0c8fd3 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:39e1add872eae1c2c483953bd062d8e39c1ad212096ef6f9827296df61adf5d0

Observation 58a0db0e-ef8d-42d6-9981-1ec18f092e7f · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:1128f128741105c40f9c80700f57e4aea1cc62cd8ef9a46f28dc7ab1b504578d

Observation ac0e57e9-6067-4f03-b674-14f2ab0d56d3 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:8cb217764a2e3b01e77a779c2864b74b341b766c1075ccd3fed00f2cded04d1a

Observation b1fb141f-8865-4542-a2d2-af25d41ca59a · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:63aedfbb97123a6e772f7159e41002887899690f32235806735edd3c7230485d

Observation 7b62a7a6-9669-4863-8b9d-7ec826435d21 · inbound

ARIS: Agentic and Relationship Intelligence System for Social Robots cites this paper.

ARIS: Agentic and Relationship Intelligence System for Social Robots Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T19:24:45.954058Z digest=sha256:9644fb3a452908ec9d189b8af86e510623af6441a91864b5fe3dfc34bf80be83

Observation a4274836-00e4-42d0-8f9e-846c186d3db9 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:45:35.282421Z digest=sha256:558cf1d395d9d1c58b30a0fdd3d5a60535199c55df97c9aafdba74c9f0743e05

Observation 5ee90f4c-549c-4719-886f-36445f090c3e · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T23:09:57.202867Z digest=sha256:8c5f5102e163518a4933d6d0b25522140265115442af066119b15149f1048392

Observation 8fd5c76c-67fe-4ad7-ad5a-cdd969354892 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:46:26.117927Z digest=sha256:02ed160ce534d7d6042c5f9d0705102aeafb043aef530805f9aaf83f5f52551a

Observation f249f835-12cb-48f8-ad7b-7876f7ebab44 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T02:22:11.160316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:22:11.160316Z digest=sha256:f3d9cfe548b01f3cbdb829c19ccacded0cfe811c842decea135fe648a12e3051

Observation 6f9d7c78-088b-4eae-938e-c9872e5efe55 · inbound

GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes cites this paper.

GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T04:59:57.934468Z digest=sha256:1c260be905d698badac35a52c4c40e1893c0c19e9285fa547c4398f45beb3014

Observation bcf55865-2ecb-4bf6-907e-250b4f2f899e · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:167c0f51c0b858167ffdf3340059d4d62fc013bef1b0938f34442c1faed6fa50

Observation 8a48c267-2168-4d6c-9355-5843f187446e · inbound

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs cites this paper.

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T05:40:03.047274Z digest=sha256:269813433750ae5e2f6dcbf1dec7eb6eb07ff024248cdfc0aa708ba5df57d10d

Observation aaa942f2-5d63-45c0-9d8e-f0ec27d554d8 · inbound

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs cites this paper.

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T09:39:49.120139Z digest=sha256:3b5febf52f3b059bd9f3a96485063753a718ec52976fa77d1bb98f7468c46063

Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · inbound

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction cites this paper.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.785540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.785540Z digest=sha256:aab0532014c413b0145f8ba7828fa5d1411556da6b3a2238d04b1c1a3aad0cd9