Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:26.760637Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2509.08266.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:26.760637Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8bbe823e-462c-40e3-b69f-d81d2a79d9d7 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vision Language Models are Biased
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 401146e2-0882-438c-8edd-f1ed24fb398a · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Open ai: Introducing openai o3 and o4-mini, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a50d699-9317-403e-b6f0-fa7c18a6eef9 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Google deepmind: Gemini 2.5 pro, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42696a3c-e460-4e02-b84b-5a5fa5eeee81 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features VLind-Bench: Measuring Language Priors in Large Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a831130e-7f07-4a94-b749-20bb5a609c37 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d34adfe-0899-4e2c-9216-88ccf8b798d7 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15927bdb-d807-466a-911a-3ca3587833fb · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b317174-849f-4293-b417-c7c08d6f86af · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Mitigating object hallucinations in large vision-language models with assembly of global and local attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088d296f-b968-48ad-bad2-d164b2ac3698 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Qwen2.5-vl, January 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d57f32d-4e42-4dd6-a132-2d26bd3076d7 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Kimi-VL Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dae0dc1-d113-46ce-93d8-f074cc831067 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82fc1942-6ff6-48f2-a64b-25d2d07aeddf · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df664d15-dfa3-41ba-8c56-15d0995bc94c · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 112c2379-fe21-4f38-97ed-77e6e4658966 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2bc3755-82a8-4d66-a4b7-cf976b1af856 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ed5b933-1c0a-4ca8-8453-642c384360be · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d64897c2-e6a6-488a-b583-2d0ead436151 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 5,6 and 7
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28fd5db4-1fd8-465e-bdae-e4042f821720 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dd5e933-aa7c-4eab-b042-7da77c559e15 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eee89944-913f-45a8-8018-bd2c559f5bac · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d70d92a-852e-4597-b085-c4614e156384 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 8 and 9
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 675d7e65-d541-4b06-a7be-441f0b0639a3 · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ffc17cc-26f5-4d94-80b4-55c3c9bb822f · outbound
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features [1], when using the same prompts and data as them
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.