Pith. sign in

Paper Citation Record · LEDGER

Explain Before You Answer: A Survey on Compositional Visual Reasoning

As of 21 August 2026, this Paper Citation Record lists 100 of 266 outbound references and 18 inbound Pith citation observations for arXiv:2508.17298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17298 v3

Coverage vector

measured 100 of 266 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:17.882743Z

measured 118 of 118 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.785540Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:52.515977Z

Reference resolution

100 of 266 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf9a193c-382c-40e5-bdc0-23b094869826 · outbound

This paper cites A Benchmark for Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Benchmark for Compositional Visual Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.513735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.513735Z digest=sha256:177d2f3ba609d454bb869afbc2c76fd57839d572009e716f84f743a24050158c

Observation 83611022-30e1-43e8-9896-77fd7aa45eae · outbound

This paper cites Take A Step Back: Rethinking the Two Stages in Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Take A Step Back: Rethinking the Two Stages in Visual Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.518180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.518180Z digest=sha256:2a51864c92c2713c00008ecf96cae4c8f6c9fe6cceff210932a8565dc310612b

Observation 924adcb7-85f5-44d2-b420-1bf109c57bac · outbound

This paper cites The Book of Why: The New Science of Cause and Effect.

Explain Before You Answer: A Survey on Compositional Visual Reasoning The Book of Why: The New Science of Cause and Effect

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.521694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.521694Z digest=sha256:5f32ce2fff3c594503421d8ee8b43ce9b55ca04d82f14ce3de1389d9f80e4c2f

Observation 5bd42100-d56e-4ac5-a5df-668a858ae1f0 · outbound

This paper cites Inferring and Executing Programs for Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Inferring and Executing Programs for Visual Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.526199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.526199Z digest=sha256:95e5f24b7d26faa4c199f5342dc0f25497e0b71a85a7ff95c3f3d193b221b8df

Observation 70ba1066-30bf-407e-8f16-35b1c2b99abb · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.530233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.530233Z digest=sha256:c9b64321ac98663328da33af5541af503c03cf9a64cd2e99fe9120171e3a3780

Observation f44fd22e-c38b-4cdb-a66b-86b7bf678bee · outbound

This paper cites RA VEN: A Dataset for Relational and Analogical Visual REasoNing.

Explain Before You Answer: A Survey on Compositional Visual Reasoning RA VEN: A Dataset for Relational and Analogical Visual REasoNing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.533928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.533928Z digest=sha256:d4c39ec480ee45ff9f079547b0ee3458e871e26489237ec15422825355a4bf0b

Observation bf4e9e29-bbb8-482e-ab01-1cdef4a157c6 · outbound

This paper cites Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning about Moving Objects

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.574350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T17:09:17.538323Z digest=sha256:2afd48ede9e2ee08b2fcb280dc30e98c7c0170c0eab17e99ebdb0ef6e7e39324

Observation f9d54651-546f-40c1-be8f-cc8b0b4c4ba3 · outbound

This paper cites Maintaining Reasoning Consistency in Compositional Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Maintaining Reasoning Consistency in Compositional Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.542217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.542217Z digest=sha256:0dad9ab372d6e799f11e096c939bd861a570b4489acbd9fe76e636b348b118dc

Observation 5f3077c1-cef7-4d19-835a-8a9cc3205176 · outbound

This paper cites Visual Instruction Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.545552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.545552Z digest=sha256:e4d1fd11865fc74324b5bd1c58b061dd7bc15cadcdd0a52baee6dce19e6077e7

Observation 0b3b3d39-4673-414f-84e2-c0fb42fc0b6d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.549166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.549166Z digest=sha256:c2da8bd0caa70ce51d018e19bc0bdeed1971f519ab2bc0edcd81d323439bc1f5

Observation 719d8d4f-3419-4e9f-ab68-3fc22040ea3e · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.553400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.553400Z digest=sha256:3d45cae4cceb926af169b09e4f17285286a9f42ac5cd0a742a3253f9572157c8

Observation 9372ac24-3045-4847-a1fa-c748255303ca · outbound

This paper cites InstructBLIP: towards general-purpose vision-language models with instruction tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning InstructBLIP: towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.557143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.557143Z digest=sha256:3aabb608ef920b4f8a7aef1319d3b81bd54378a9f902260a020973a375c8af34

Observation e174f2f3-7221-4093-849e-c6e360e8dc2d · outbound

This paper cites A Survey of Multimodel Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Survey of Multimodel Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.560723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.560723Z digest=sha256:8b7c5aa7de997f2a612fb0cf6ad284e5951da078febaa0acae93993ab085cbd9

Observation 6e70e9c5-698a-4ac0-a3df-92b461415d20 · outbound

This paper cites Interpretable Visual Reasoning: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Interpretable Visual Reasoning: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.564558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.564558Z digest=sha256:65bc5789355961286386483688d19dd0011d54c25cd1f941d39c53651f513fe2

Observation e0a07f5f-c192-4958-a691-081a514198a5 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.568659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.568659Z digest=sha256:14f173ad7201afd10f493d5bee67c6bed011ba8820106ddd9ad82edee1093228

Observation 0e2b5f02-331d-461e-aadf-b9937102a53b · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.572947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.572947Z digest=sha256:797d383339465685c04d85f4f0bf28e2de20a4652d8644d9322204259eb33dd7

Observation 512ec9a5-7708-4f86-bf24-29a0b4fdf499 · outbound

This paper cites Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Position: Prospective of Autonomous Driving - Multimodal LLMs World Models Embodied Intelligence AI Alignment and Mamba

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.576724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.576724Z digest=sha256:f019a59b1e47318aab8e2098e7a4aede071a1440526dcfe7fa6162673cfc4b23

Observation 815e063b-a98a-4958-ac2d-0ddd06033f11 · outbound

This paper cites NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.580278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.580278Z digest=sha256:29bb8b39443d8a2ebc1be7403a4783588039608ec176286a20652e481de65587

Observation d43da455-eb95-4631-abf3-20648f47a5c2 · outbound

This paper cites UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning UA V-based Urban Structural Damage Assessment Using Object- based Image Analysis and Semantic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.584057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.584057Z digest=sha256:011b60e279885d769c7b02162c7bcfee548a1bbb87e4891f3c1a6a71719cc3fa

Observation 987e6fbc-1c5d-487a-92b8-82eb754d8265 · outbound

This paper cites A Study of Visual Reasoning in Medical Diagnosis.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A Study of Visual Reasoning in Medical Diagnosis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.587663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.587663Z digest=sha256:f1954e65d703aeb4073140cd63e1c28bb135b8e2de408ee5a4458608fe24c062

Observation 96a9982d-3315-4005-abf4-39a1116d614a · outbound

This paper cites Medical Visual Question Answering via Conditional Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Medical Visual Question Answering via Conditional Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.590944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.590944Z digest=sha256:bfc3ea1add2b4219395217d19d24d961b64a1a91edf179e7d83a1ead06aa1151

Observation 79ab4d51-8731-415c-9a36-9f3079477b69 · outbound

This paper cites Programmatically Grounded, Compositionally Generalizable Robotic Manipulation.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.594301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.594301Z digest=sha256:8b6ca93b7d924506828d17f7a8b96bae1cf10ecf2341a081274f7740eb4b7fcd

Observation 33a7b91a-c907-4811-8035-29075a1d6879 · outbound

This paper cites Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.598754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.598754Z digest=sha256:280f1f4116a5f3910b83036d4ac5b14b428095c3e974e31259c6b471adf5e4d6

Observation 297a2486-68bd-44b4-8619-b25f63c90620 · outbound

This paper cites When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,.

Explain Before You Answer: A Survey on Compositional Visual Reasoning When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It? In The Eleventh International Conference on Learning Representations ,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.602375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.602375Z digest=sha256:89df0266ad9990b0ccf022afbfb5dcb55d147fece3830b4a89614f6bde475367

Observation 99030bc0-26ea-4050-a685-58bd8e373122 · outbound

This paper cites Challenges and Prospects in Vision and Language Research.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Challenges and Prospects in Vision and Language Research

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.605883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.605883Z digest=sha256:0adac1c37cd1767343518eb5656c49b25eb0fae252919bdfa4fa34a0d74950cd

Observation 196cf821-1c40-4821-b251-93dbff6c6ea7 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.609224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.609224Z digest=sha256:bb5012496212d6c8d023862e1c4afd494540cd8b38666106b69d991fd848bca3

Observation 2fd5befd-2d65-4ff3-a00c-4e9608088e1a · outbound

This paper cites Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.612998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.612998Z digest=sha256:ca11acd2fad1bb53391ea77e948713e0dddc2a82ce6e9c1ebf022eb087f6756b

Observation 4d087fe7-1755-40e9-98ea-16743f97646e · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.616473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.616473Z digest=sha256:490ae04be3392594b32eca160702b314d78a3d41438334507358ac8d7f4d9b41

Observation 3daeb3fe-c31d-429a-a273-1e2da67a7da6 · outbound

This paper cites HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.620594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.620594Z digest=sha256:5dbc86291e197d74247167ad68d4e835f721256243a868173f59784330744682

Observation 53e0aaa8-566d-42d4-9045-52f301a33acb · outbound

This paper cites Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.624093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.624093Z digest=sha256:63fa4af390c8bbb15457b8e2697c3fbb2a0a1090c08dcafadcdfcb6f4006a29c

Observation b1388ad6-63e2-47fc-baad-db756fcd9eea · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.627597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.627597Z digest=sha256:56ed2dd6c53bbc26a2a6bbd7c142e56725655f1ee1b7170e54fad956b7a8e3da

Observation 8d1ceb2c-3880-4c4f-b55b-82c3818c16cc · outbound

This paper cites Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Enhancing Multimodal Compositional Reasoning of Visual Language Models With Generative Negative Mining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.631786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.631786Z digest=sha256:4e0dc65803898e4e1351522389db8e0d4208155b9e18d7b4f1dbb88589be3434

Observation 85856c28-b570-4790-9b5d-f4e1b6ca1e72 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.635381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.635381Z digest=sha256:961b01c0e6c3b605612ab6cf05f7d7fad82cc5739041dd5357bbe9029b035878

Observation 64d1eec0-1361-4820-ad73-17f8e09fad9b · outbound

This paper cites Cascaded Mutual Modulation for Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Cascaded Mutual Modulation for Visual Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.638828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.638828Z digest=sha256:cda944c3a4af8347c88bd85f0de8ad61ee0a18aa7d585ff50cdbe8f52acaca91

Observation a2df30ad-5b05-4301-9b66-9afb134bfb95 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Mul- timodal Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Compositional Chain-of-Thought Prompting for Large Mul- timodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.642304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.642304Z digest=sha256:a2affc56c81f8567286cdadcc51d5e91b06ff7ced33e96d43e71ea2173abc89e

Observation d784b18a-9cbd-4deb-90b5-9eda1f65b23a · outbound

This paper cites Building Machines That Learn and Think Like People.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Building Machines That Learn and Think Like People

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.646029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.646029Z digest=sha256:be673f705635d833313925aae0b26df7c9a9191ba613710f029b1ce624629b4b

Observation 918a7128-311b-4c4c-ad4e-de30c815e0a4 · outbound

This paper cites Meta Module Network for Compositional Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Meta Module Network for Compositional Visual Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.649583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.649583Z digest=sha256:39084a67c460f6b09b4a9e7a074581aaeac1aa73e9b69dc31621eb5e9e72da8f

Observation 3d198a3e-97a6-4c85-a11e-aa27679b8ef0 · outbound

This paper cites Improving Visual Reasoning Through Semantic Representation.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Improving Visual Reasoning Through Semantic Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.653229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.653229Z digest=sha256:a7261f9c06d39f407213950c938bc19bac2904fb1596aab0e473a3305eb4a29d

Observation b25d1405-e144-46fd-ab8e-b15b530a418f · outbound

This paper cites What’s Left? Concept Grounding with Logic-Enhanced Foundation Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning What’s Left? Concept Grounding with Logic-Enhanced Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.656988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.656988Z digest=sha256:a94489f8f224def05c0c55df3c6d46258c3132b594bb9670f3d415bf6962849a

Observation 683be0f2-2665-4239-afec-bcb50526e037 · outbound

This paper cites Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.660571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.660571Z digest=sha256:c1c2dfda0b5523444abd93f5696cb895f133e150985b71dffc92eb7be4a2c4e7

Observation fc3fc673-c443-42a2-b420-165676db93b1 · outbound

This paper cites Visual Question Answering: a State-of-the-Art Review.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Question Answering: a State-of-the-Art Review

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.664117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.664117Z digest=sha256:4785aeb65a96b09e81194079a4837672f0b8202794003090f4dd72aa2ae141c5

Observation 2d033639-8414-435d-8263-eeebc82ec403 · outbound

This paper cites Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.667724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.667724Z digest=sha256:b3b628091c40cfa80bcd919a85761604cc2c7867ba066f91fc9f663183c9a1b5

Observation e4977987-e095-4c48-8059-9b2636c812a4 · outbound

This paper cites Mind with Eyes: from Language Reasoning to Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.671253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.671253Z digest=sha256:ed6870b8abf0eddd396cf19d77ad681095c9eb2c0a90a2326879873c96e52b36

Observation f4ae5747-05cf-4036-bf44-e2481bc8e361 · outbound

This paper cites A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge.

Explain Before You Answer: A Survey on Compositional Visual Reasoning A survey of neurosymbolic visual reasoning with scene graphs and common sense knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.674787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.674787Z digest=sha256:db9355d00e56f0e93c8950384f4f44e901f85bf53b8eac1bdd06cee66cb18c90

Observation c960d985-fdbd-4ba9-abb0-3c3a02cd2d7c · outbound

This paper cites Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.678225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.678225Z digest=sha256:6d215afb3263a1440bf112eaba7170767e807fbca554f8e121e0b7f024464a6b

Observation 06ab9bbb-a02a-4267-a8bb-86804d9c61db · outbound

This paper cites Large Multimodal Agents: A Survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Large Multimodal Agents: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.681651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.681651Z digest=sha256:e107a86c203b57a13e7beafdecef115500f2dcc29f3c7ebe6b9b4d9919613e52

Observation bbef4951-f830-443c-94dd-8e38f18c843d · outbound

This paper cites From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From image to language: A critical analysis of Visual Question Answering (VQA) approaches, challenges, and opportunities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.685501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.685501Z digest=sha256:ea31c056e95d81cd30001626af41acea61dfca0a7cef0f9f13f1aaf24a49ab12

Observation 1c15a348-49a0-4fee-8ca8-1f237413e057 · outbound

This paper cites Robust Visual Question Answering: Datasets, Methods, and Future Challenges.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Robust Visual Question Answering: Datasets, Methods, and Future Challenges

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.689858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.689858Z digest=sha256:7246436e533206796ce6a2c44afb13e8f51d70a4f9bc3ded1331273ca1a363b2

Observation ba5105d8-c291-475e-bc8e-29ac4f2cd677 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.693163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.693163Z digest=sha256:ff6a509e78b1a4523a7172841af7d20abad0412cf6438c24defa1fbb12ea673f

Observation 94d521e2-9d73-4f3e-81d2-107b850daa29 · outbound

This paper cites How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model.

Explain Before You Answer: A Survey on Compositional Visual Reasoning How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.696947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.696947Z digest=sha256:8b7e07f74890062648d54bbe91b710378b6d4f05a674ab9553ccb808d4acbf2a

Observation deb20ad6-fd0f-4386-b1f4-d13d21265ad4 · outbound

This paper cites Tool learning with large language models: a survey.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Tool learning with large language models: a survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.700806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.700806Z digest=sha256:da89e4c8b0b10e214543189eb9aada8010fe748a221e7b9fd75611baea7f030e

Observation 77ab8f79-80a7-451e-8d54-432ec300cc0d · outbound

This paper cites Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.704220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.704220Z digest=sha256:390821ec614c7a41aed1f665e48927882c2da526ed3dde007e3cf70d51029318

Observation 02920ede-5fbd-4de5-b835-f6a418c41248 · outbound

This paper cites Multimodal Intelligence: Representation Learning, Information Fusion, and Applications.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.707897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.707897Z digest=sha256:20767d5eaaa2db8b15972d4b23b7d09f584fc6deb4249b5173754ddec3028098

Observation 2f2659f7-bea5-4572-afa9-8c2b0efb2b45 · outbound

This paper cites Transformation Driven Visual Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Transformation Driven Visual Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.712111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.712111Z digest=sha256:8c477fc69eb22383b008b5f6cc3e7b18f0c443114ce821717988a991454b7a43

Observation 1b81e6d5-89b6-45c5-b512-b315b8302a38 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Residual Learning for Image Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.716819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.716819Z digest=sha256:ec4236be2eab92f709dfd4a4c91861f9ca5fef38d0b582d143eead0760e30d00

Observation b1010c49-1a83-43c7-bccc-18c96da62785 · outbound

This paper cites Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Two-stage Rule-induction visual reasoning on RPMs with an application to video prediction

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.720344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.720344Z digest=sha256:b9f868ec02a5763e9fdf16466baa9766f0ff05ba1365fe95471053b0717952e4

Observation ca4d0ffc-6df0-41e3-ab35-cf09a055772c · outbound

This paper cites Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.724028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.724028Z digest=sha256:e1a53f4dd2a604193efce6883aeab91bc34fdba76464494c7fd77632d2036187

Observation 90671fdd-138e-4c0c-9b78-b6022419b97b · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Explain Before You Answer: A Survey on Compositional Visual Reasoning You Only Look Once: Unified, Real-Time Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.727468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.727468Z digest=sha256:c2eee62aae1599b975139d38ed16f59a022fdd2045dfab4b96be164250acffeb

Observation 688bf382-7594-46c3-b2a0-26b482a034c4 · outbound

This paper cites Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.730761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.730761Z digest=sha256:0e4ea6845ad4584c339983bb0db144d66645579336fd9d77c2be46eb4137bcdb

Observation 8dde71b0-5067-4b54-af20-57d9a50dc076 · outbound

This paper cites Bilinear Attention Networks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Bilinear Attention Networks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.734056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.734056Z digest=sha256:615754a8a0f3766e94f9b1167dbdc52160503c2fdc43c1000d92cbf0f8ea41c4

Observation 941cd9c1-6dfe-4474-9d73-c0242e7b12d4 · outbound

This paper cites Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Improved Fusion of Visual and Language Representations by Dense Symmetric Co- Attention for Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.737492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.737492Z digest=sha256:6a223ed39ea6f6d43cf245a2360aa0e17e8cdb72e0c17da95bd2e3730a8f65ac

Observation 062c41e1-e85c-4bc4-bd55-c6aedc3761e8 · outbound

This paper cites Deep Modular Co-Attention Networks for Visual Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Deep Modular Co-Attention Networks for Visual Question Answering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.741248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.741248Z digest=sha256:61f9b81de78293dd32b7e60776bfbb480d2d391edc6b7d465b9b9f7aad924a83

Observation a0c7384f-0162-4842-a675-56c7f7180120 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.745210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.745210Z digest=sha256:1bc24757bd7d7caf78bb22b0353eeb02ab251b40f7c79b390c2098e0c2dc63de

Observation a2a5f4f7-200d-47c1-959f-5b217f210384 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Learning Transferable Visual Models From Natural Language Supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.748896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.748896Z digest=sha256:cc59b6734e83c5aaf9b3410b13b333a59e1e09dca32bbe3a97fcecc0bb230166

Observation b4ac0030-bd44-4619-9630-5854e783e3f8 · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning UNITER: UNiversal Image-TExt Representation Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.752421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.752421Z digest=sha256:2ccc2d13ffbf58b6976d42c24deffbd6f025e78f4fdf6b95f9daeda6b834190f

Observation 55445a2a-d4e0-4e52-b950-feee7d366620 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.756195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.756195Z digest=sha256:54c95018a29e9305e1f6bced8cb2d724702a4ae3efd929b1535d6c4f9f81a331

Observation 57c32c0e-3d8e-406c-be76-bceed02b541c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.759948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.759948Z digest=sha256:aa8bacab34cc4c03828d87f787382102de58c0b74bcd0293312413625102d2e7

Observation 2b23ce4f-6699-4b99-8f80-0196f07275b9 · outbound

This paper cites Otter: A Multi-Modal Model With In-Context Instruction Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Otter: A Multi-Modal Model With In-Context Instruction Tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.763717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.763717Z digest=sha256:da250f89785a6335d5c8cbb14861e68cb8044bd9f76deb3076e75e99a6dc705c

Observation f7ed05ab-d852-437c-b3cd-3ff03a1c8f2a · outbound

This paper cites CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023.

Explain Before You Answer: A Survey on Compositional Visual Reasoning CREPE: Can Vision-Language Founda- tion Models Reason Compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.767580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.767580Z digest=sha256:34dfba2f35d4612dca221e2a4887604a78fd5213f3d2f842e62a3bfd692bbc93

Observation 84fb288c-bba1-47aa-91a0-ddd47bc24e45 · outbound

This paper cites Logics and Languages.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Logics and Languages

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.771070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.771070Z digest=sha256:f882ec4acd66259576e6ede9e914f76d6326caa5ccd872e35a33106efbf09f76

Observation f4ad16d1-0bed-4eee-960b-6027a2807b91 · outbound

This paper cites From Recognition to Cognition: Visual Commonsense Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From Recognition to Cognition: Visual Commonsense Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.775107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.775107Z digest=sha256:59c4125436929897e3ecea973bcf7bfb4d5d638c4071a374ec6776e10cb85eb9

Observation baf19ec6-f9e9-43a4-bb37-7391120e77c8 · outbound

This paper cites Visual Programming: Compositional Visual Reasoning without Training.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Programming: Compositional Visual Reasoning without Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.778592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.778592Z digest=sha256:6d4ca0b826128e5c72f5832c6a501ca29896271416317fbfaaf51f39d0f8b3c0

Observation 99f6c90f-0741-4e80-bb46-f36ffe1334c4 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Explain Before You Answer: A Survey on Compositional Visual Reasoning GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.782737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.782737Z digest=sha256:b2c1d8ca813599e92d5edd7c216434d3e30ad11ab782a293ad6e90ef22416fea

Observation 2bc9b9a0-95bb-4356-8361-fd0d53e5ac2c · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.786176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.786176Z digest=sha256:355cad1853ac4e53c84aa6c89bbf9e4e9e02d92cb51193ad49cd4da811eeac5f

Observation 42dac3b9-737c-4faa-9eaa-dd61249f1bd7 · outbound

This paper cites Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Bongard-HOI: Benchmarking Few-Shot Visual Reasoning for Human-Object Interactions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.789732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.789732Z digest=sha256:ac998009b0835d6388b5886dc67fe099fb616ad712f1f0bdd3b7821444cea421

Observation 1e1b648b-eb9d-4cd7-a375-5885712aef3d · outbound

This paper cites DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.793148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.793148Z digest=sha256:7df71e7b1d130a282654d1f50ad126f3fd54af30ef40b9e347d4e3fe7e2d8f6b

Observation 219d2711-261d-4c2a-8821-b4e16c67407e · outbound

This paper cites Iterated Learning Improves Compositionality in Large Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Iterated Learning Improves Compositionality in Large Vision-Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.796395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.796395Z digest=sha256:52198aa27b3a282b746cfcea0e5cbfb6e4737fb37136a1e6e2b90438e1277e59

Observation 0be00509-5529-45a1-9744-ea2ffe4877ee · outbound

This paper cites Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.799892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.799892Z digest=sha256:9b637c5bcabf33a742471ea3c64c2ccbec5099f069c97d7353812777b31b69ea

Observation 402ae171-e88e-405b-8586-0d8f42ab9c48 · outbound

This paper cites Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.804131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.804131Z digest=sha256:2c331c5eb460fd211d1524a2fb0722583e66b7ea58e85812b14cc700d7d6fe84

Observation fd81ad9a-57cb-4db9-a2e5-1c4001c63617 · outbound

This paper cites an unresolved cited work.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.807533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.807533Z digest=sha256:22383c0ac4c86915228d7e2cf1ae3eb85003ff941f2d2926bc3cd0442cad70d4

Observation 5003c3c9-2241-4807-bf99-040ceabd0e34 · outbound

This paper cites Visual Compositional Learning for Human-Object Interaction Detection.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Compositional Learning for Human-Object Interaction Detection

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.810619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.810619Z digest=sha256:5475395f017696b338ebc4e8d14900c8730d760b6731ab2fd94bce01eaffb375

Observation 768452f4-e668-425f-9073-9215c305e181 · outbound

This paper cites Learn- ing Visual Composition through Improved Semantic Guidance.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Learn- ing Visual Composition through Improved Semantic Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.814504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.814504Z digest=sha256:850153724f65b7376f0fcae75c629b71374b34f6b27773855b02b87cf08c3ff0

Observation 192f2e11-89cd-4f08-a990-e135ce90ee3a · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Socratic Models: Composing Zero-Shot Multimodal Reasoning with Lan- guage

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.817979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.817979Z digest=sha256:ded59edc80962ab70fc74241397dceb06be3493e60e2f472ec871bb99d905413

Observation 8930522c-42f9-4904-9d26-9b80d0fa34ba · outbound

This paper cites Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Synthetic Visual Genome: Dense Scene Graphs at Scale with Multimodal Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.821769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.821769Z digest=sha256:957056c13454be84c57f2a8ec27283deac38b5f78da1b8f7df8e2e46a5242782

Observation 987437d3-de0a-4659-b914-f0228fce6e58 · outbound

This paper cites Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.825678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.825678Z digest=sha256:b18f21d29b9bffd783de79789d062138c91689baa98f170d9a3cc16ddce01d15

Observation 278a00d6-4fee-4601-8e17-672634a531d7 · outbound

This paper cites DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage.

Explain Before You Answer: A Survey on Compositional Visual Reasoning DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.829651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.829651Z digest=sha256:0943068f4811c20eed47357a7da5241d611737caa8125fba558ad8db4d9bbef9

Observation 29a18e16-e797-45cb-9f18-3830955ce0c2 · outbound

This paper cites Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Divide and Conquer: Answering Questions With Object Factorization and Compositional Reasoning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.833925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.833925Z digest=sha256:d229fcf017e1709bbc01fd81a290c3c01c4a304324ea984a98fe6b3bf968173b

Observation b9db1b69-797a-4db1-95d2-c85d21c96437 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.837525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.837525Z digest=sha256:38add85e20ba6dd0981fcbb192d00d899057e98a2cedcc057449f17602a0aa54

Observation f2878906-a8f7-4076-99eb-00848f53b154 · outbound

This paper cites Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.840955Z digest=sha256:f074de611471d100b1bacfb02d5e2a94fbe789417d3cd1aedf90bb11d5455115

Observation 9da7e67f-463c-45f7-bad2-459093ead803 · outbound

This paper cites From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis.

Explain Before You Answer: A Survey on Compositional Visual Reasoning From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.844636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.844636Z digest=sha256:6dff808f3d478fb791b5f3f9d59e3e12100b7db025062c120d5ceb578b8e0f3f

Observation 604da5ec-4610-4637-958d-0ea41dd807fb · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.848322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.848322Z digest=sha256:5c086f66477f5afbdbc490c92813b5a36580222e4e21fb3608ee147fb46ab0ee

Observation 61c5613e-7010-4d73-8542-7ce9f4cea9a7 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.852321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.852321Z digest=sha256:a7e76efd88a13ce665619698f3a012e0e4e62b2cd86f9c51ca34557ba8916936

Observation fb4e0f90-1e84-4ef2-953a-0144669a021b · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Explain Before You Answer: A Survey on Compositional Visual Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.857411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.857411Z digest=sha256:f8864222fd0f3d78022c2eb9053f42f4573e69276a5a9ce97a05e2d7c73c5968

Observation a50aad1e-035b-4fc0-94c0-b621b2a5082a · outbound

This paper cites Grounded Chain-of-Thought for Multimodal Large Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.861449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.861449Z digest=sha256:6141961a0ae5589b14c2e38a26365bfc75c9c26fd7a25d60b1a93e25a24f3e55

Observation cd5e1b7d-bb3d-48b5-9f69-0480c2afb333 · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.865349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.865349Z digest=sha256:4211b0003a613b6dae0333e70ca0e20c268edf19c240117fb83e1e571c5b1982

Observation 4b9dacbf-46aa-4491-8e98-36ae4e4777fe · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Explain Before You Answer: A Survey on Compositional Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.868949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.868949Z digest=sha256:ada830818782ecf228679b4507a41012e6fc1f0e04722e54d7d08e99161e7b4d

Observation e37c4b36-3475-43e2-af63-4f2ab3164986 · outbound

This paper cites ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty.

Explain Before You Answer: A Survey on Compositional Visual Reasoning ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.872314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.872314Z digest=sha256:028ff239436c60ddbb0338dc3b47dff18a4839b859614a5e7c7bd50c33a60c01

Observation bb85a1dc-0e06-459c-8797-1cbb510a0bbb · outbound

This paper cites Visual Reasoning by Progressive Module Networks.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Reasoning by Progressive Module Networks

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.875627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.875627Z digest=sha256:3142aa1b751b126ce3863b93443d89f2963d7c8cfd41d725c8a0dbce5213de1e

Observation b1c5f0c2-31fa-45fc-8e88-0d21b6eb56a9 · outbound

This paper cites Predicate Hierarchies Improve Few-Shot State Classification.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Predicate Hierarchies Improve Few-Shot State Classification

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.342455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T17:09:17.878832Z digest=sha256:d467e6bc5a4cf473121db9ef1581115198b20333ef44a2202b461ecdcf1a3622

Observation bf64629d-8068-4973-8e84-c06fc4d03311 · outbound

This paper cites Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models.

Explain Before You Answer: A Survey on Compositional Visual Reasoning Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:17.882743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:17.882743Z digest=sha256:93b41e72bdac17703e4e5b0ab3cbff689b5a419b1e3063a9dce6d5f182037280

Pith citing papers

Observation 29394350-4395-427a-a682-5f18b5a73c2a · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:a297b0b132a000a02d3b9cb382a751a8ab87758009f12390e2bccb017aa410e5

Observation ebf4352f-6fef-4a35-8e24-3dbbcb426783 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:29.860215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:29.860215Z digest=sha256:8d3d38bcfbdb3a1921b6dbe994d0caafd8101f664eb6d57039836632b222a241

Observation bd14d4ff-deff-46a7-baa7-ea0fadf4891e · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:3239866aba7326bb3f684d77d55fb3255902212e5c1ac046d5da5bdd541afee9

Observation 0c472c11-a78d-4401-866d-41d28e17055c · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:5210d32f79df45d0ba15ce2d85cef0fb67c2f27c5c8067c487c0fe1ef2de2472

Observation fb7f47e8-e6dc-4614-bcff-b50c7d0c8fd3 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:6f328689da4de5cf498a076864cdcae9d056b8c464198473b1643451fd99e0e6

Observation 58a0db0e-ef8d-42d6-9981-1ec18f092e7f · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:a9c19611690aeda8afdf8551b997bf0a347f0455e3e0fadb1fc3bb0c95eadc14

Observation ac0e57e9-6067-4f03-b674-14f2ab0d56d3 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:4de32bd1aa5ffc9ce9d9df8f05a27d4219249f33e52571d75b7399f54ef1e60c

Observation b1fb141f-8865-4542-a2d2-af25d41ca59a · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:ac3e4e8a99bfd5d3a77df06d8a1d5cc3ecf20e9423581607f2026b1c9363a58a

Observation 7b62a7a6-9669-4863-8b9d-7ec826435d21 · inbound

ARIS: Agentic and Relationship Intelligence System for Social Robots cites this paper.

ARIS: Agentic and Relationship Intelligence System for Social Robots Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T19:24:45.954058Z digest=sha256:cb9c29cd34be2d30cd581faa453642121c6b02f22c57e3d51d9126b6737a817b

Observation a4274836-00e4-42d0-8f9e-846c186d3db9 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:45:35.282421Z digest=sha256:71672bcf3d5c21e82b497ad82383c93c4c6b191d23c55eea5076659b09ace156

Observation 5ee90f4c-549c-4719-886f-36445f090c3e · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:09:57.202867Z digest=sha256:9f1eb1a7b36d34dab572fbecba9e7b0f3b7f1b2acb04b09ace5243f5fca60f93

Observation 8fd5c76c-67fe-4ad7-ad5a-cdd969354892 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:46:26.117927Z digest=sha256:a713845185fcd36bcd098125593ffb7e45e2f708b7a4abda020ef53c12e0e0b8

Observation f249f835-12cb-48f8-ad7b-7876f7ebab44 · inbound

Dynamic Execution Commitment of Vision-Language-Action Models cites this paper.

Dynamic Execution Commitment of Vision-Language-Action Models Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T02:22:11.160316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:22:11.160316Z digest=sha256:1faf28dcf6508645abb3df3b035b43530c366d237ed661ea2e7242c5952e9bb5

Observation 6f9d7c78-088b-4eae-938e-c9872e5efe55 · inbound

GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes cites this paper.

GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:59:57.934468Z digest=sha256:90f5ec2b624384731c77474016fb5e58eb7d367313be01c1959e410f64f11c0f

Observation bcf55865-2ecb-4bf6-907e-250b4f2f899e · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:b0e2b106db32ab3ba90b632e3aa667402ebff7472f908ff786044ebc05f6eb5c

Observation 8a48c267-2168-4d6c-9355-5843f187446e · inbound

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs cites this paper.

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T05:40:03.047274Z digest=sha256:7a46849226f013ff4fce8bff680afaf4e60533e747a5f8a8f023d12875f52158

Observation aaa942f2-5d63-45c0-9d8e-f0ec27d554d8 · inbound

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs cites this paper.

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T01:19:37.179138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T09:39:49.120139Z digest=sha256:08b63d3863a165bbf72bde77cb7a636c7b0957364b4347d6502f76167fbdffbb

Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · inbound

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction cites this paper.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.785540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.785540Z digest=sha256:cae46eb1e25abbd1235579237d7e4e8778448bb13da7e3693b7eb6f3408b2852