Pith. sign in

Paper Citation Record · LEDGER

Vision language models are blind: Failing to translate detailed visual features into words

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.06581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06581 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.161806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.161806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.161806Z digest=sha256:0bfde9a81dbd1982ad52b4515eac28f6182356c0ab3c035d2d17a0f46dea39be

Observation 2d172e97-d16d-4b23-96e8-07f47b5f7817 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.399915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:10ed495440ebae73644d9b90fdc0f5c32e04b43f7da9fb05adf1347037c4f393

Observation 90ab0e07-92d8-469e-a697-85fb8720ac3f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.234599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d16722ab0cee4c2546790e7887ba9d7fe79d9c94dfdcff4bc259d74562f78cdf

Observation fea50c66-695f-4e0d-8199-6795462b794a · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Vision language models are blind: Failing to translate detailed visual features into words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.991717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.991717Z digest=sha256:13c351b47d974b9825a0b5d83cd04f9b31e8a30ebb4bba54c13fa7bf60b7cdf7

Observation 25b7d359-8125-4274-9e8c-390a39552d4b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.218019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:edd022595502c823aad92e970dfb2ddb531f2d9d18656cd52f290641b11d24b3

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:1d7bf715ab9ae885b58c9518e95af0c49928f08c88948eac4e4635031fae05c1

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:d5f18e026a6fd9b8415d8eea42ddfe3e56087be42f89dca66e265a860890eb1e

Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.991005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:74567a379d46f0534a83bfd04c17df891386ddf538615dde30c44ff6fdbaeb81

Observation 54beb4b6-eab4-4c6a-a62f-4932d1ec6ffa · inbound

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production cites this paper.

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:34.429111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:34.429111Z digest=sha256:618fdbedc224b40bdd0fe113c3e442f0f5ecb632a116307434980a7eb439f3ea

Observation f940b53a-8cb9-47ba-ba9b-7bc8d974cbef · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.881826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.881826Z digest=sha256:3569c9d36c7c46b3bc95faaa7e82f70fca85e0d15d201f61062dba8f25c0177b

Observation 50c24374-7d38-43d7-8e26-b3ad50485149 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.704948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:22adfc1ab3055d61e01d88734daa00a2c3002ee4e0dab542c93f2107840d4fc7

Observation fb0aee7d-01c0-4023-a49a-babad6dda0c5 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:d09431ff079be67373673e0df35804ae32f3c2e5f7353b616d8c4ec8830e512e

Observation 5f717ea5-af40-4dee-a8c1-6ba866d69223 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Vision language models are blind: Failing to translate detailed visual features into words

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:02.305599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:7b245fc9162a18028855cd4ebdfdfa4286865cb71f618b3e40f9256695a69593

Observation 507e832f-3a9e-4d07-b76f-eac62eab3c70 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:89f8510ca2d2846c6b7c1bb1c5a3b916f03e74b9e0a7781d4d9de8d9ac58acba

Observation 53b4db02-75ec-40cf-9ebc-67c356f64060 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.871090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:92b4a94f77ddd22512386deab93ffae4bc27d62afa5553d96846c50df6c99780

Observation 74437cb0-9d1c-4e63-bed6-ff4a7af29f26 · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Vision language models are blind: Failing to translate detailed visual features into words

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:18.598684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:ecb117c677978e5a58ae46604b6c5183a6029e74ef4432c1674fc19b02cf9fb9

Observation fdb723b4-7ba4-4341-8e5a-1cc979e0f73d · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:59:45.707454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:c42f06c63ee16fbada04312c1fd88296efde1625b4c4d11605b43bcf4673aa6e

Observation a0e4a19a-dfff-4ede-84aa-bc715d538aab · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:54:57.786081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:5172d91e24be9e2d0cfa7446041aaae0729211db63fde8cf3d0954fa5920cc96

Observation e5f8806d-6710-47e2-a050-161487d2d9e5 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.285291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:fc8296de0b8e5743736a3c4c6099e160629a23bad59d499d4191d473794f939e

Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · inbound

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes cites this paper.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.666857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:869ddec5b68f85a08af1e34a9eb4a3c0de0c2d78549c24af8637b68172701d2a

Observation e26043c4-e69a-4fba-bf18-736ff1b56518 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.051325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:a8f18fe1993c9f77800997232a4878fc320016d4646531d073be7ab4d1da85d1

Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.978721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:932e2e8e522ec6fc03ccab1b7e5e8c9e3b0ca2acc3e85bd0dbe1b8ee10e581e3

Observation a38d71dd-185c-4860-8360-5c4924cca342 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.439678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:d808bb0814c3d0b987aaa6f89a5e6b08fc0ab9d645537492e6d8f9f035333cb6

Observation 58793a2d-0661-4d17-a12c-6c5c5216f1e3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Vision language models are blind: Failing to translate detailed visual features into words

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.076380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:eeb7b5c02c3f688205f82457d7c109171b94eb93986d41f204f886f62bac1e8b

Observation 287748e6-e590-4333-866d-e815b816a8c8 · inbound

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR cites this paper.

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR Vision language models are blind: Failing to translate detailed visual features into words

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:31.073799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:32:39.976049Z digest=sha256:5b7ffdd88a6b191eb78a71ce3ee6844d0f7a7ccca8c42e01a6902ad2b593519a

Observation 1626bef0-46b0-43d5-9dcb-86301cac682a · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms Vision language models are blind: Failing to translate detailed visual features into words

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:10:02.337868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:53f845fd8711ddcd8287199edb1f01120093906b0dae17fa3477895008c4b334

Observation fe5085fe-7473-43f1-869e-7b787f611cb6 · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.238996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:24:11.228672Z digest=sha256:5cab74fbedf3f0960915de859613128ced7b1c012db1e72db0ea91052f6bb8fc

Observation 6dd3fd32-7674-4264-9f5e-f80b07eefd3f · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:58:52.700648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:58:52.700648Z digest=sha256:b2655a688fbdcc779dff20584fb8188831e3b628c1c97d489590990bb2442d7a

Observation 90a92488-f13b-4b92-a37b-504dac08ccda · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:42:46.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:42:46.146973Z digest=sha256:5de9eeb1b4ede18986917ed98ed404259efaf951d0d3dc141f5970f8ed52ed1a

Observation b9c3781e-bf2f-49f8-8d26-6fbf92dacf31 · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.256397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:2265817a9a0634c5969d3433b9ae0f0598329e45c086081d0c11dfc6a2f8047f

Observation 64fac733-b34b-4efc-8a60-fa541118694f · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers Vision language models are blind: Failing to translate detailed visual features into words

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:06.631736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:06.631736Z digest=sha256:65d88972a2f5dd8afe301d16475b0be44ab6fe373b5846f6e6eca17fcb6d81e5