Pith. sign in

Paper Citation Record · LEDGER

Vision language models are blind: Failing to translate detailed visual features into words

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.06581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06581 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.161806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.161806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.161806Z digest=sha256:85e12ca2dd7d7bd9e990e9b80f7e2b553136b08ccb4091c75f3b52ae90f23688

Observation 2d172e97-d16d-4b23-96e8-07f47b5f7817 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.399915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:66ad19ffad88b59a62fdf44f40b00522bd86559f2418fd3a5e9b0741c841caf8

Observation 90ab0e07-92d8-469e-a697-85fb8720ac3f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.234599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:497900219e9edc72e970b94f8998b79f849aa19df37ad335d86d94fa3e19155f

Observation fea50c66-695f-4e0d-8199-6795462b794a · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Vision language models are blind: Failing to translate detailed visual features into words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.991717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.991717Z digest=sha256:d5aa4c0f5652272a6472ae5b8425c001d5a6d30df47b43660919e366128bc10b

Observation 25b7d359-8125-4274-9e8c-390a39552d4b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.218019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:c75587b35fe1f6ece41c90890d32493a88d2e8a4784b29c27b46972bcba256e0

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:c892cd4ad56aa855a97c1b10db6c072df3a0ac2254b0ef80a0c2df5682b46ae6

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:c2ab7aa3b194dc47467afbace64e26469b8303f43140b11ee5a3fde2b173924d

Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.991005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:527941e2f556b6c3502a450cfcee89a86036b22dc1f643051d214de137214265

Observation 54beb4b6-eab4-4c6a-a62f-4932d1ec6ffa · inbound

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production cites this paper.

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:34.429111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:34.429111Z digest=sha256:946ff0cb0f14b96180663c45d7b6ca9c802323333a8d1ef6801fb04b9e0954dd

Observation f940b53a-8cb9-47ba-ba9b-7bc8d974cbef · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.881826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.881826Z digest=sha256:00c6771ac3699f28999a5dd41b3f8415554990e43bcd63a34ae6835bc4b6e283

Observation 50c24374-7d38-43d7-8e26-b3ad50485149 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.704948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:e62481b90c667998114cba075ee7c479cb9fc6a3f27e2eab739ff83f50f16fc6

Observation fb0aee7d-01c0-4023-a49a-babad6dda0c5 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:37c7c7ea7b363d58aa80f2b4c3d5f90f6c4ba61b00c40e7e680c316dc1dd9548

Observation 5f717ea5-af40-4dee-a8c1-6ba866d69223 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Vision language models are blind: Failing to translate detailed visual features into words

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:02.305599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:f580d3460740777b327eecfafcfe1431f1d9f6fd80dc0300923792b89a0253d5

Observation 507e832f-3a9e-4d07-b76f-eac62eab3c70 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:dc7bec772550b6d7b52f0503bcf3eb4923daa7d2e227d224fa29908cb7cde0c9

Observation 53b4db02-75ec-40cf-9ebc-67c356f64060 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.871090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:b744755615cdfac2b4cf0294f89b25b519d78ce72b5ea6dba0a22739c4b6a6aa

Observation 74437cb0-9d1c-4e63-bed6-ff4a7af29f26 · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Vision language models are blind: Failing to translate detailed visual features into words

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:18.598684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:d75f9ccb07da07ecaaf9d855d743534b8f8b67881ccd0fadc1e1b23aaa00ef73

Observation fdb723b4-7ba4-4341-8e5a-1cc979e0f73d · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:59:45.707454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:af87a44d4fbd7abcc511ea4efcbed13b221bff54a3eb415c0b21529e43dc5afd

Observation a0e4a19a-dfff-4ede-84aa-bc715d538aab · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:54:57.786081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:3075d0b67d2bc6aac0c7c28dd76814cd042ec39ca58fd7d5666f874ffe6e3119

Observation e5f8806d-6710-47e2-a050-161487d2d9e5 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.285291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:fdb0d29059ea5d04598b674e5bfcad0da5ffbafbc7034c66bc248f48feaba069

Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · inbound

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes cites this paper.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.666857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:b84c59124e291175d022ace2483368ef30091bc143a4711a10b7086188907cff

Observation e26043c4-e69a-4fba-bf18-736ff1b56518 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.051325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:686b9dda8a9e4f41ef77beacab14a6d8c435572eb213b22cb55848adf850c55b

Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.978721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:65e1badf529fca7b4144c7de5279417c3522635e4c969c65e04a5417599016a2

Observation a38d71dd-185c-4860-8360-5c4924cca342 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.439678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:3afbc4c6dcd8bd05bd1e055d32393d600ab1408d69cdd16283558ef7e2da96d7

Observation 58793a2d-0661-4d17-a12c-6c5c5216f1e3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Vision language models are blind: Failing to translate detailed visual features into words

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.076380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:0c34809f5e21be8527d8fa2d68d9f4c932e372414a62eb8ff3bef8b456a5196d

Observation 287748e6-e590-4333-866d-e815b816a8c8 · inbound

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR cites this paper.

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR Vision language models are blind: Failing to translate detailed visual features into words

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:31.073799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:32:39.976049Z digest=sha256:2e8d6fc284e46ac5fb6d9277220fff6812460c5f50a5e8d5326fd7e12ac6ecb5

Observation 1626bef0-46b0-43d5-9dcb-86301cac682a · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms Vision language models are blind: Failing to translate detailed visual features into words

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:10:02.337868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:15481c074cc510d33bd07c752cfef1507a8151196e28c7d32186c2fc910cac41

Observation fe5085fe-7473-43f1-869e-7b787f611cb6 · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.238996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:24:11.228672Z digest=sha256:2aa759e4f863f29557f5775bff01c085de39457162744b4d3f225604d9e4ad87

Observation 6dd3fd32-7674-4264-9f5e-f80b07eefd3f · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:58:52.700648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:58:52.700648Z digest=sha256:3eba495acf1cfa474bf6d9f6b7b060be66948d8f113b0ac6156054037411b9e6

Observation 90a92488-f13b-4b92-a37b-504dac08ccda · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:42:46.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:42:46.146973Z digest=sha256:444e73e43f8d51f74649adbd58e6b23b107022132bd8be1318c4106d83b3d6d1

Observation b9c3781e-bf2f-49f8-8d26-6fbf92dacf31 · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.256397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:78f7e79dcb81ef2a47ab9583d696a35ecbbe49bd0ecb271763a8e411c189f515

Observation 64fac733-b34b-4efc-8a60-fa541118694f · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers Vision language models are blind: Failing to translate detailed visual features into words

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:06.631736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:06.631736Z digest=sha256:4e22cbfd0feb16aeb0cd6de1870e78acacc32cbbf7ee0b6838c48553bf2b53fa