Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:47:21.298721Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2509.25339.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:47:21.298721Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:06:09.166363Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T00:39:16.760441Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4576053d-b36e-44ee-8d3e-6b87bef647f0 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Analyzing the Behavior of Visual Question Answering Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fd9518-4c04-4cea-8dd2-b5fa5540390d · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Don't just assume; look and answer: Overcoming priors for visual question answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b88778ee-6288-4281-8351-953ba692dc25 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Lawrence Zitnick, and Devi Parikh
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d64e739f-4735-4809-b787-2a826ecbf900 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes TouchStone: Evaluating Vision-Language Models by Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0907f8-1167-4777-8fe4-5de53dae2734 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b10dda8-45ad-4bc0-8edf-8be54046ef0e · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f86f860-a92a-44e7-b326-bebd1325be26 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes An Introduction to Vision-Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee4b375-35a2-4ed8-951f-e29d1c51bba5 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Rubi: Reducing unimodal biases for visual question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eba9351e-66a9-47a2-ad2e-540a1cbf8742 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Frankland, Thomas L
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5e17b1aa-5508-4161-917a-a22e05ba2199 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37: 0 27056--27087, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eac2d1b9-92f1-472a-bcbc-4ca0ba736604 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b5919fa9-9530-42e9-ae14-9f2d8b8e5407 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Dense and aligned captions (dac) promote compositional reasoning in vl models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation acea20a7-fd49-4782-8089-bc3a54530cbe · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Teaching structured vision & language concepts to vision & language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 240e69e9-fc52-4a20-909d-48b0a6ede26c · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Datasheets for datasets
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd2025af-03d4-40ec-a42a-b14a45d18f16 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemini: A family of highly capable multimodal models, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2fa9098-eb69-4e5e-b187-95de0361c95c · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemini 2.0 Flash Model Card , April 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5e20987b-301c-43fc-ab17-94d3f2ad7b21 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemma 3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f355e9d-40b0-4cd3-b177-8eea9ed46ac1 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Announcing Gemma 3n preview: powerful, efficient, mobile-first AI , May 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7437e8e0-9bab-47b3-a9c9-376ea8443af4 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f577566f-61d2-4934-9b3d-35f29ed3da87 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Horizon Alpha - Advanced AI Language Model , August 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9f9fe8bf-3b2b-48e9-99bc-53a8a006fad1 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f251e5ee-0697-4eb8-a8c3-24fac17a7b2d · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Conme: Rethinking evaluation of compositional reasoning for modern vlms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f6ff9b14-0eab-41b6-a60f-77e3d8b0e79d · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Large language models are zero-shot reasoners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a861aeaa-a769-43b1-a18a-347c688b418f · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Dvoichnye kody s ispravleniem vypadenii, vstavok i zameshchenii simvolov
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 790b2c43-80eb-455f-98d9-5eb7cd5321a3 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes OtterHD: A High-Resolution Multi-modality Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eddabbf-1ee2-43b6-8ea9-a3153d23b222 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes LLaVA-OneVision: Easy Visual Task Transfer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1035144-e976-4713-ab9f-1a64ad82b23e · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f7edbe-9fa1-4a7c-967f-4bed859fc1cc · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation abe1a6c5-fdbc-4a82-8d77-db8e8afadbe3 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Omnibench: Towards the future of universal omni-language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be46e89-72bb-4df3-beef-1dfb9c7b4398 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Introducing LFM2: The Fastest On-Device Foundation Models on the Market Liquid AI , August 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88745ae0-5b33-4b9c-80e3-c104f21a1df4 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Improved Baselines with Visual Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1aa9b3-f745-4527-ae8d-3781a72d6564 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9131d634-98cb-4b2f-8364-d37164c47830 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4702157-2270-4dc0-9fcb-6d4f0173892a · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88151aa5-fe11-456a-a41f-07f3871f88af · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes SmolVLM: Redefining small and efficient multimodal models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c014dd17-f679-4637-b26c-a042e91c0753 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Umap: Uniform manifold approximation and projection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42490883-eacb-40aa-9618-25eb19f76c80 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation , August 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f18667f4-307f-4d7c-b4ea-65ca6c822495 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Lafter: Label-free tuning of zero-shot classifier using language and unlabeled image collections
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 62648a1b-cc44-408d-be5c-3143792037bd · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a51e2431-9cff-4337-bf93-b9962ae81c41 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gpt-4 technical report, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2626458-faec-4587-b9ce-05513ee4366b · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes OpenAI o3 and o4-mini System Card , August 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9bac3cc7-64ce-441a-89c8-b9b6b325e354 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Teaching clip to count to ten
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1cdcc5b-8858-4d7c-b2c6-f3db5c05edca · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Humanity's Last Exam
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4652664-a200-4bfb-8d52-05d83c12a132 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Scaling Vision Pre-Training to 4K Resolution
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0ec413-9a2b-4721-a793-212da32eeb2b · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes PaliGemma 2: A Family of Versatile VLMs for Transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5975e3-ade4-440b-9b3b-f59e1b89f7b2 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Winoground: Probing vision and language models for visio-linguistic compositionality
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a2fec2-73df-4032-b5fc-bf50c58878fd · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Chi, Quoc V Le, and Denny Zhou
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e666cc3-0a11-4a25-9bc3-288cd37cd3c5 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c13a608-bfde-4df0-a934-a2a56f5654b1 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes V?: Guided visual search as a core mechanism in multimodal llms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 12ca737c-8477-4d65-b0c4-96001c0d6c4b · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f1514b-204d-408f-af04-91d59147c537 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e05dc6-e450-4575-8263-99b7c967bfb4 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480d0587-444c-4c5a-a899-75d96eefc4fd · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23e0a24-5d2f-4bed-807e-4584ed11507a · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Yin and yang: Balancing and answering binary visual questions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb33b780-0800-415d-8170-c2bf6728fb1e · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63efc223-c776-498c-bac0-d3b890c6cbae · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62803457-e2fe-4e8e-8a31-9fc12368c114 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Why are Visually-Grounded Language Models Bad at Image Classification?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6054d7eb-f7fb-4438-a0a0-69574ec18ba8 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c7314233-8846-4750-9ec1-ab18580096ea · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36af6328-f6f0-4cdb-b958-695a745e5880 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes @esa (Ref
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e11a02-4912-44fe-8084-ecb42ece83a2 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64999212-2f4a-430a-a251-1ef55cd8d540 · outbound
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2269bc1-a5ce-4470-8187-3f2c773314d1 · inbound
REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.