Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:31.975509Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2506.08008.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:31.975509Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:44:45.178051Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:29:29.186840Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d1878cd5-cea2-4902-b674-fde8032d49ac · outbound
Hidden in plain sight: VLMs overlook their visual representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da569a7-e35e-43b7-831e-8a15fb0eeb92 · outbound
Hidden in plain sight: VLMs overlook their visual representations Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5f4e8e-52ec-447f-a42e-25e53966848f · outbound
Hidden in plain sight: VLMs overlook their visual representations OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29bfddc-1992-4023-be6f-1dcf41c3317c · outbound
Hidden in plain sight: VLMs overlook their visual representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf576ad-2880-4a83-927c-45829b2203b0 · outbound
Hidden in plain sight: VLMs overlook their visual representations Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 580c1030-0ba6-4645-b8a8-c33cea9c1479 · outbound
Hidden in plain sight: VLMs overlook their visual representations Probing the 3D Awareness of Visual Foundation Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26744c1-08d4-4e66-a479-43e0be18273b · outbound
Hidden in plain sight: VLMs overlook their visual representations PaliGemma: A versatile 3B VLM for transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cd8f14e-8f94-4bec-a68b-2d43b66e6b0c · outbound
Hidden in plain sight: VLMs overlook their visual representations Evaluating Multiview Object Consistency in Humans and Image Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c760c7e7-9e4f-4753-a33c-653eb22cef32 · outbound
Hidden in plain sight: VLMs overlook their visual representations Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde34bae-c1ab-4412-b835-e865f41fe766 · outbound
Hidden in plain sight: VLMs overlook their visual representations ShapeNet: An Information-Rich 3D Model Repository
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf1eaae-e437-4f15-9902-44650ed9bd71 · outbound
Hidden in plain sight: VLMs overlook their visual representations An Empirical Study of Training Self-Supervised Vision Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dab55f0-ed59-4e5a-8f40-64a1bd16929b · outbound
Hidden in plain sight: VLMs overlook their visual representations Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f89ad20a-b7b2-4269-8be2-b30b28854f2a · outbound
Hidden in plain sight: VLMs overlook their visual representations Gonzalez, Ion Stoica, and Eric P
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9705a3f-eb6c-4689-b927-f444ef4c5b14 · outbound
Hidden in plain sight: VLMs overlook their visual representations Imagenet: A large-scale hierarchical image database
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6324917d-df62-4403-96b1-d25a6accdb8f · outbound
Hidden in plain sight: VLMs overlook their visual representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37462020-3f3a-46e7-8157-55b8d09095c0 · outbound
Hidden in plain sight: VLMs overlook their visual representations MouSi: Poly-Visual-Expert Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1d3bfd-f903-4eff-8b91-cd78b2677d51 · outbound
Hidden in plain sight: VLMs overlook their visual representations BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535000f8-1b2d-4c6a-9863-30297af8d5d9 · outbound
Hidden in plain sight: VLMs overlook their visual representations A Neural Algorithm of Artistic Style
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b990cb-d837-4715-8c96-4b3604a8777f · outbound
Hidden in plain sight: VLMs overlook their visual representations Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9200f337-18a2-4632-9222-10fe1cd965fd · outbound
Hidden in plain sight: VLMs overlook their visual representations The functional correspondence problem
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b3675a38-0f23-4adb-a3f8-4399078d05e8 · outbound
Hidden in plain sight: VLMs overlook their visual representations What matters when building vision-language models?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa320425-5567-4e1b-a1b2-e289f5b01d27 · outbound
Hidden in plain sight: VLMs overlook their visual representations The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5150e7b4-13cd-457a-9a7f-9b7b2c84bded · outbound
Hidden in plain sight: VLMs overlook their visual representations Improved Baselines with Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 361ab3f8-5750-485e-9c03-72e1fa7965f0 · outbound
Hidden in plain sight: VLMs overlook their visual representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94772e6-f240-42d1-a7de-8f870d376c8f · outbound
Hidden in plain sight: VLMs overlook their visual representations Transformer-based neural texture synthesis and style transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b9114a3f-163d-495a-a67c-fd0091df70d4 · outbound
Hidden in plain sight: VLMs overlook their visual representations SPair-71k: A Large-scale Benchmark for Semantic Correspondence
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a2ae01-b915-4b82-bab3-8e94fa98192a · outbound
Hidden in plain sight: VLMs overlook their visual representations Indoor segmentation and support inference from rgbd images
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e001612a-9f8c-484f-94b2-ab6402190f12 · outbound
Hidden in plain sight: VLMs overlook their visual representations DINOv2: Learning Robust Visual Features without Supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89a713b-4577-4ba0-9b57-9d3534bdff1b · outbound
Hidden in plain sight: VLMs overlook their visual representations Learning Transferable Visual Models From Natural Language Supervision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61c7834-0bfd-4419-8d62-6bfe83bf8bad · outbound
Hidden in plain sight: VLMs overlook their visual representations Vision Transformers for Dense Prediction
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · outbound
Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bd96aa-b993-4134-9747-392bb262feb0 · outbound
Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944a45ef-63b3-4ef9-b4f7-15b93ad1ce04 · outbound
Hidden in plain sight: VLMs overlook their visual representations Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6972fb65-294c-42bc-b23c-9462b1bfbb11 · outbound
Hidden in plain sight: VLMs overlook their visual representations Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 237f0622-530e-4866-bb3b-e867106bf4de · outbound
Hidden in plain sight: VLMs overlook their visual representations Disn: Deep implicit surface network for high-quality single-view 3d reconstruction
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b48f5f0-46a2-4ef5-aaaf-6c5027b37b2d · outbound
Hidden in plain sight: VLMs overlook their visual representations MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec4e2d1-989f-45d6-b10f-c0998f888da2 · outbound
Hidden in plain sight: VLMs overlook their visual representations Sigmoid Loss for Language Image Pre-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66242d9-6b7d-4b1b-bf6c-4d27d675e400 · outbound
Hidden in plain sight: VLMs overlook their visual representations A General Protocol to Probe Large Vision Models for 3D Physical Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 097b52e8-2d6e-47ad-9065-505b1eeaa5d6 · outbound
Hidden in plain sight: VLMs overlook their visual representations write newline
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aeb6766-572d-484d-9d53-34305de2ff79 · outbound
Hidden in plain sight: VLMs overlook their visual representations @esa (Ref
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8edfdc6e-430b-43eb-b4c3-34f7e275f8d4 · outbound
Hidden in plain sight: VLMs overlook their visual representations Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f879e721-808e-4138-b4c2-9b7742b2229a · outbound
Hidden in plain sight: VLMs overlook their visual representations A, B, C, D
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4f2ae1-7184-47fd-8352-468f01ff091c · inbound
PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters Hidden in plain sight: VLMs overlook their visual representations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a901a8f-b164-4932-a899-d8b48b76f59b · inbound
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Hidden in plain sight: VLMs overlook their visual representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 906c870a-fb4f-47a1-bc78-34f6579d5ef1 · inbound
Mull-Tokens: Modality-Agnostic Latent Thinking Hidden in plain sight: VLMs overlook their visual representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0ece36ae-f997-4c5f-aced-1b408dbee82a · inbound
Egocentric Bias in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e02c4924-6ba0-4eec-9be2-0cddcca0ab73 · inbound
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Hidden in plain sight: VLMs overlook their visual representations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 843cc588-bcef-45b4-ad02-e53dfa76a3d0 · inbound
Vision Language Models Cannot Reason About Physical Transformation Hidden in plain sight: VLMs overlook their visual representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4825b433-659f-4773-8bee-187497140f64 · inbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e41c6df5-c7aa-4fd7-9275-eec59f7b92e8 · inbound
Watch Before You Answer: Learning from Visually Grounded Post-Training Hidden in plain sight: VLMs overlook their visual representations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation deee1951-098a-4113-85f5-857b9badc27d · inbound
Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification Hidden in plain sight: VLMs overlook their visual representations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 934e2904-d7e1-4c89-978d-3a8e55ac61cf · inbound
Do Vision Language Models Need to Process Image Tokens? Hidden in plain sight: VLMs overlook their visual representations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2c8aec45-5da4-4685-9d43-3c26b0e35eca · inbound
Boosting Visual Instruction Tuning with Self-Supervised Guidance Hidden in plain sight: VLMs overlook their visual representations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation df2215f2-cab2-497c-97a8-f12aad1d624d · inbound
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models Hidden in plain sight: VLMs overlook their visual representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 575432cf-dadc-494c-807f-1a53dc2b3433 · inbound
Do multimodal models imagine electric sheep? Hidden in plain sight: VLMs overlook their visual representations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9dc668a8-f105-4526-8b89-61ae92877cf9 · inbound
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc64514a-24f0-456a-b219-b2dbf321a2ba · inbound
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b957b92a-4bcf-4ae6-8c5e-30ce662dcb60 · inbound
Diagnosing Visual Ignorance in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43114add-9cbc-41f2-99ff-81fde306c970 · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM Hidden in plain sight: VLMs overlook their visual representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ab637ef-2e2c-4309-8082-436d893e5348 · inbound
Visual Access Boundaries in Vision-Language Model Reasoning Hidden in plain sight: VLMs overlook their visual representations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd8021e-4afd-4b44-89b7-8baeb11f729c · inbound
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Hidden in plain sight: VLMs overlook their visual representations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.