Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.348354Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2502.00711.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.348354Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:44.595993Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T14:33:46.092736Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 649cd177-51ed-4f04-9e08-d8538d083313 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaliGemma 2: A Family of Versatile VLMs for Transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e22241-9f19-4b57-a981-2e8df316fd40 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual instruction tuning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb8f5b2-24eb-467b-9b82-dc9abbc70039 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Mdetr-modulated detection for end-to-end multi-modal understanding,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cad07e75-4fab-4b70-b696-a2df3173f6ad · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 78d15b40-0858-4e91-b723-e66af2381bdf · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Cola: A benchmark for compositional text-to-image retrieval,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd748ad8-7833-406c-b731-c2e9de31d20c · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual programming: Compositional visual reasoning without training,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13ca5d95-21ba-4f0c-83a7-a6ef98b275c2 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Interpretable visual reasoning: A survey,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d897cae-9bd3-4bce-925a-d4d70a53e570 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87d224cb-6bc4-47bf-90d1-cea4ba5904f5 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rephrase, augment, reason: Visual grounding of questions for vision-language models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c97c9b4-04b5-43ad-876c-31afd58a9186 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6cc5687-cb05-4ac4-89df-ad9c12f108cb · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ecd5353-0806-429f-bd7d-adc11a0573bb · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vision–language model for visual question answering in medical imagery,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 09bf68a5-9800-46fb-a189-f9a5aa0935ad · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Lingoqa: Visual question answering for autonomous driving,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 493c5b92-5199-4a6a-9220-f8dc3ab48917 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ac84f1-acd6-40fd-a085-f22403b49a64 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Dealing with Semantic Underspecification in Multimodal NLP
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201d67f5-58d4-40b1-87a2-6cb46d878e3b · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Open visual knowledge extraction via relation-oriented multimodality model prompting,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45b14339-5b4d-41df-aaf5-dcf88d0e2447 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PV2TEA: Patching Visual Modality to Textual-Established Information Extraction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c41710f0-af35-4fab-8584-779222491bc2 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Recurrent fusion network for image captioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9cc4f3b7-a6f7-44f6-a865-61ff0b598e07 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Boosting image captioning with attributes,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec9a6f63-0b01-4463-8b84-a56fea172bf6 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Promptcap: Prompt-guided image captioning for vqa with gpt-3,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b3de275-8ce0-4c20-9023-989fc64a8130 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Injecting semantic concepts into end-to-end image captioning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44a583f8-05d8-490e-9b87-d344108e3e9b · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual Commonsense based Heterogeneous Graph Contrastive Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 576ee7d7-bb2a-411f-bda0-bc7818ca20da · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Covlm: Composing visual entities and relationships in large language models via communicative decoding,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 859becd4-8893-49f5-8051-0a6a55c14891 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Bridging knowledge graphs to generate scene graphs,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17da6dd9-5f85-4145-b893-db9cff8abcfb · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Co-training improves prompt-based learning for large language models,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 352cb6f8-0e86-4031-9d49-c0d03e1ee2c6 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language model as attributed training data generator: A tale of diversity and bias,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f23dd74d-2d9e-4d04-965a-e696057e3844 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language models are zero-shot reasoners,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72faaa99-7ef1-4082-8c89-e4dd2dface47 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd099ee-9f1a-4352-b2f0-45ae8faa3e49 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37485df7-ccff-4907-a53b-d104118fabc5 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e32c95e-1594-4df8-a12c-5606734b071c · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Improving vision-and-language reasoning via spatial relations modeling,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 581bbdad-363a-4dde-a6dd-c3153a860408 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework GPT-4 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a49c73-976b-48ed-b03b-edf5ca7db438 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Chain of thought prompting elicits reasoning in large language models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6301d550-8086-41db-9af7-2e5291357021 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Re- flexion: Language agents with verbal reinforcement learning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48cb50bd-3dce-4d0a-b28a-5375e5202022 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Making the v in vqa matter: Elevating the role of image understanding in visual question answering,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc1b197-7c6f-4a7c-ab72-dbc64b4aa10b · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework A-okvqa: A benchmark for visual question answering using world knowledge,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c7d8b91-5454-4649-8acd-ffe132606bac · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vizwiz grand challenge: Answering visual questions from blind people,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b4a91c25-d32e-4a6f-a5d3-5e61ae026181 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework @ crepe: Can vision-language foundation models reason compositionally?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebb616e-16a6-4190-a4bd-660078d895be · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Qwen2.5-vl,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f87daaa-787a-4922-9028-3e10613dc958 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework 4o mini: Advancing cost-efficient intelligence, 2024,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 39ae7dd0-f211-4e54-874e-c68fc8233f56 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Hello gpt-4o,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b03738ee-646c-4502-8f5f-188b317569f9 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Introducing gpt-5,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51f45e83-d199-4222-bc23-0a51e053ed2b · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Learning to localize objects improves spatial reasoning in visual-llms,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d453f790-a4cd-49b5-b5db-c2bb756e6185 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Eliminating the language bias for visual question answering with fine- grained causal intervention,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bbc7420b-076e-4705-9814-f9e1961ef556 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e538e280-5cfb-46b7-a517-36b48a81d7a5 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfe634d1-97ed-4d42-a49c-e3f291b3d187 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 614997c3-fa05-4221-81c6-761128f6e041 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 886cfc9d-39ab-4256-aaa3-4f24b4803d38 · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0426853d-c62e-4887-8569-5ba81f7ec38d · outbound
VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Available: https://openreview.net/forum?id=L4nOxziGf9
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 798e4ab8-facb-4bdd-b162-b30819d9d17f · inbound
Augmented Vision-Language Models: A Systematic Review VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.