Pith. sign in

Paper Citation Record · LEDGER

VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2412.00947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00947 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:12.982888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:30.214152Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f7140633-96f6-4c00-9bee-71368dcfa5a6 · inbound

Efficient Few-Shot Continual Learning in Vision-Language Models cites this paper.

Efficient Few-Shot Continual Learning in Vision-Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T23:37:28.642453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:37:28.642453Z digest=sha256:6364517159e995d77fad2682d2564e28fc25293a5c0a611b0855438bc2faf5d8

Observation 99452ea1-9ee4-4289-9077-fe367dc07d56 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:12.982888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:12.982888Z digest=sha256:96ee486b0be67cc222b6340d9e89cf5436c8026b8651c85c6887cdf1bc7a84e2

Observation 6f47a08a-978f-4f72-b89c-b76d940c876a · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.336310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:53.336310Z digest=sha256:5bc4528873f774f03eb1931dd776ddbb4cdcb15b53241abcf71a141178229117

Observation f4afe626-f042-4eb8-948c-6478f66f7fef · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:08.012180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:08.012180Z digest=sha256:e5a29bc22485af8ea1d045ce4f533a628f4c21fa5a106f81168211125d79ca72

Observation 830a32d2-05da-4338-85dc-07fca2b14c57 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:51.700025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:51.700025Z digest=sha256:17be71f541ed5d59aeeed49568d034a1b3049e449be8f1c1c81808f572a06024

Observation c6fda004-17c1-49e6-a231-5e1f0f261a9f · inbound

Why MLLMs Struggle to Determine Object Orientations cites this paper.

Why MLLMs Struggle to Determine Object Orientations VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:04.546609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:19:44.979076Z digest=sha256:5e20bb032b1e92f204cd5ebf8ab2a56ae8afde85d0e2c92c9e00766f5df9b266

Observation 21f03969-e72a-430e-80f0-858351a178fa · inbound

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading cites this paper.

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:16:26.557200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T11:45:57.291112Z digest=sha256:24bd8a2c756eece7ccdf8c7dd565d246df2d05411206bfc0fe6217b7022ee38b

Observation 4a1f0191-67c1-401c-a645-e41a857a2b47 · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.501063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:b093c4b9ea48d56c19edd324bcc3a7303df56a1e6849dee8450c1c0ee6ab306c

Observation 7722e7a8-e3ed-4aec-be36-78b28d282a84 · inbound

A Dataset for Dynamic Human Preferences for Vision Language Models cites this paper.

A Dataset for Dynamic Human Preferences for Vision Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:30.233405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:19:12.597144Z digest=sha256:7de50e81a29c872b6780261fd42fccbe8db286aceda722f2b621d315c107be90