Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Really Understand Visual Language?

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.00193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00193 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:26:36.002606Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:13:21.512791Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6baffd48-76b9-4dd9-a2bc-015b9f46e169 · inbound

Explainability for Vision Foundation Models: A Survey cites this paper.

Explainability for Vision Foundation Models: A Survey Do Vision-Language Models Really Understand Visual Language?

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-10T17:26:36.002606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:26:36.002606Z digest=sha256:167278e25733a7de3c55d061cb19d83ead014933b75bcdac9199828e29f2b186

Observation 8bcc6c06-589e-4811-b439-600c4ccc9304 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Do Vision-Language Models Really Understand Visual Language?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:12.945758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:12.945758Z digest=sha256:8ee7d40b83406050863c9bd761019311b2dbc90fc2734158e75c7a686b2586db

Observation 3ef1a41b-ba4f-41a4-bf3f-c49dd05ac35d · inbound

WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis cites this paper.

WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis Do Vision-Language Models Really Understand Visual Language?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:00:32.189282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:00:32.189282Z digest=sha256:331273d850a7c5575887ef9acc85d6ce143a86787dfd256c48d4a094f09c08cb

Observation aee70df0-c41f-468b-ad14-146a75c38881 · inbound

ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues cites this paper.

ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues Do Vision-Language Models Really Understand Visual Language?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:06.467280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:06.467280Z digest=sha256:714954683e1a7e75e73a8a09edae2bd127a6e004a46570d6fb4fdc5d3a309bf3

Observation a9ba543e-5102-49ad-bd2c-1d24ad5fe75d · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models Do Vision-Language Models Really Understand Visual Language?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.125186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.125186Z digest=sha256:cf1e1a85556c5f564653efd9443ac67e7e0a8f5a5f7ea568d4d743a633e1906f

Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · inbound

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features cites this paper.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.649412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.649412Z digest=sha256:0b834e9481b3d93b70e5a2fa4744d2ca7be62b9577f73fd68482773c901da66b

Observation bb536c98-020c-4ac6-8c8b-7f2f360deeef · inbound

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions cites this paper.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Do Vision-Language Models Really Understand Visual Language?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.284722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:3e15595ad83bfea5c3952e0f89b804d4056d5b8387bb7b08b2a3d175c9ad91c3

Observation 26b2dbf2-6d71-4bae-9ccb-acd527940f04 · inbound

Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding cites this paper.

Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding Do Vision-Language Models Really Understand Visual Language?

Reference 26

Resolution
verified exact
orphan_title_repair, observed 2026-05-13T17:18:02.458267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T17:15:31.368264Z digest=sha256:aebac80aee19df6a696ab526174806381c4793305a2166e37b79da2b2ed51522

Observation 51ea26a9-8f75-4e2d-bd1e-165529af6f56 · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Do Vision-Language Models Really Understand Visual Language?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.514602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:c6f3f5c09be3d6979ccc9ae27615ea2fec326af4bb307baafc7d2719be755a5b

Observation 8b54b4b5-1c0a-46a6-b413-a2b356e5ee53 · inbound

Prior Bias in Vision Language Models on UML Diagram Interpretation cites this paper.

Prior Bias in Vision Language Models on UML Diagram Interpretation Do Vision-Language Models Really Understand Visual Language?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T06:36:12.761601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:36:12.761601Z digest=sha256:622fff4a50dc3703b99c5cc41e4007d809d591e2ba7dc97ee3d8732894763380

Observation 8e712cc3-8472-41e2-b623-e29e0c4283d1 · inbound

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams cites this paper.

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams Do Vision-Language Models Really Understand Visual Language?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T07:19:14.207385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:19:14.207385Z digest=sha256:02042c20592efae177182e1d66b335d077e177eab96cb3543ed6115e8352a33d