Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.06581.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.161806Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · inbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d172e97-d16d-4b23-96e8-07f47b5f7817 · inbound
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Vision language models are blind: Failing to translate detailed visual features into words
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90ab0e07-92d8-469e-a697-85fb8720ac3f · inbound
Seed1.5-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fea50c66-695f-4e0d-8199-6795462b794a · inbound
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Vision language models are blind: Failing to translate detailed visual features into words
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b7d359-8125-4274-9e8c-390a39552d4b · inbound
Grounded Reinforcement Learning for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · inbound
MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · inbound
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · inbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 54beb4b6-eab4-4c6a-a62f-4932d1ec6ffa · inbound
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production Vision language models are blind: Failing to translate detailed visual features into words
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f940b53a-8cb9-47ba-ba9b-7bc8d974cbef · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Vision language models are blind: Failing to translate detailed visual features into words
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c24374-7d38-43d7-8e26-b3ad50485149 · inbound
MiMo-Embodied: X-Embodied Foundation Model Technical Report Vision language models are blind: Failing to translate detailed visual features into words
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb0aee7d-01c0-4023-a49a-babad6dda0c5 · inbound
Vision Language Models Cannot Reason About Physical Transformation Vision language models are blind: Failing to translate detailed visual features into words
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f717ea5-af40-4dee-a8c1-6ba866d69223 · inbound
ReflectCAP: Detailed Image Captioning with Reflective Memory Vision language models are blind: Failing to translate detailed visual features into words
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 507e832f-3a9e-4d07-b76f-eac62eab3c70 · inbound
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53b4db02-75ec-40cf-9ebc-67c356f64060 · inbound
Context Unrolling in Omni Models Vision language models are blind: Failing to translate detailed visual features into words
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 74437cb0-9d1c-4e63-bed6-ff4a7af29f26 · inbound
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Vision language models are blind: Failing to translate detailed visual features into words
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fdb723b4-7ba4-4341-8e5a-1cc979e0f73d · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a0e4a19a-dfff-4ede-84aa-bc715d538aab · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5f8806d-6710-47e2-a050-161487d2d9e5 · inbound
Binding Visual Features Point by Point Vision language models are blind: Failing to translate detailed visual features into words
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · inbound
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e26043c4-e69a-4fba-bf18-736ff1b56518 · inbound
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Vision language models are blind: Failing to translate detailed visual features into words
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a38d71dd-185c-4860-8360-5c4924cca342 · inbound
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 58793a2d-0661-4d17-a12c-6c5c5216f1e3 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Vision language models are blind: Failing to translate detailed visual features into words
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 287748e6-e590-4333-866d-e815b816a8c8 · inbound
Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR Vision language models are blind: Failing to translate detailed visual features into words
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1626bef0-46b0-43d5-9dcb-86301cac682a · inbound
Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms Vision language models are blind: Failing to translate detailed visual features into words
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe5085fe-7473-43f1-869e-7b787f611cb6 · inbound
The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6dd3fd32-7674-4264-9f5e-f80b07eefd3f · inbound
The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a92488-f13b-4b92-a37b-504dac08ccda · inbound
The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c3781e-bf2f-49f8-8d26-6fbf92dacf31 · inbound
Information-Regularized Attention for Visual-Centric Reasoning Vision language models are blind: Failing to translate detailed visual features into words
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64fac733-b34b-4efc-8a60-fa541118694f · inbound
An Exam for Active Observers Vision language models are blind: Failing to translate detailed visual features into words
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.