Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2302.11713.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:22:10.964965Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:59:46.832017Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b6bfed60-da5a-4134-b7a6-b87f524b9d9b · inbound
PaLI-X: On Scaling up a Multilingual Vision and Language Model Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a4f59d38-5b45-41f7-af56-f14df06e3ed3 · inbound
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3109636d-67a5-415a-baff-0e63cb21a129 · inbound
UniCoRN: Unified Commented Retrieval Network with LMMs Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a8baad-8059-40d3-af0d-6ecc6d78427f · inbound
Towards General Continuous Memory for Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1780cfb-b22e-4278-807b-4b7dcdd9248a · inbound
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b593f0-fa3a-43dd-82f9-2874a2ab121a · inbound
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa67e94-d0bc-4c11-8444-04a73de81995 · inbound
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 261c7fb3-9da8-48e0-95c6-3f92a1fafaf9 · inbound
Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eed1381-9216-4444-9d52-308e0adc1ca9 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b0c8a1-b73e-4bf0-bf58-3901c577b7e7 · inbound
MMSearch-R1: Incentivizing LMMs to Search Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · inbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d9639e-f645-4088-80e7-25eeee52864d · inbound
Augmented Vision-Language Models: A Systematic Review Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c4a3da-e9a7-4fa5-b398-b019731279ff · inbound
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ea7054b4-df8a-41ef-a8c0-4066d635aa22 · inbound
DeepEyesV2: Toward Agentic Multimodal Model Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e880ec27-3ea7-45b3-8290-d63e329b92e4 · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f8b9170-1c4c-4886-bf4a-bcaa4ed1c4c7 · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1c256f-8a67-45d1-bb0d-622d9aac99ed · inbound
Evaluating the Search Agent in a Parallel World Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4fd4f9ae-c567-4359-a94a-2355379060a8 · inbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df65d05-b063-4535-bd42-fb830a64fa75 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c589ea0-4333-4bed-9a40-68f66b2496c2 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d5088e-69fd-4c27-a7c0-95a147f5c88f · inbound
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f65682d5-c688-4328-a206-d66109f75ec6 · inbound
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17253701-54e0-494f-99cf-108285154a07 · inbound
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4dfc9e15-e7ef-4c77-9ef1-356eb94b06a6 · inbound
Delineating Knowledge Boundaries for Honest Large Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6fabff02-2f74-430d-995d-3bcac17718dc · inbound
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7219241-79d1-4bb5-9a7d-668821257ec8 · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cdc4ff57-90a9-4fd6-938f-fe148251fffa · inbound
Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fcf7514f-5791-4cb5-b6b6-aacdc5ce8d7c · inbound
SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a1c92dfb-b7d4-4e0d-8081-6a6b85515075 · inbound
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bfdd7a1d-5754-44bc-b4ee-de7fe08b7503 · inbound
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ee170d-edfb-4ad2-8703-e2033d4bd1f9 · inbound
UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10cad550-eafc-443e-b7b9-4b4853ee7ce6 · inbound
UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf954f1-ee33-40a1-9b60-d5801e39a2c0 · inbound
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.