Pith. sign in

Paper Citation Record · LEDGER

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2509.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08266 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:26.760637Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bbe823e-462c-40e3-b69f-d81d2a79d9d7 · outbound

This paper cites Vision Language Models are Biased.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vision Language Models are Biased

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.320688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.320688Z digest=sha256:66d9129a6a0c865559d4842dc417eb59c1cff393c1dfb6b956ab4fef2ae94071

Observation 401146e2-0882-438c-8edd-f1ed24fb398a · outbound

This paper cites Open ai: Introducing openai o3 and o4-mini, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Open ai: Introducing openai o3 and o4-mini, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.323971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:24.418011Z digest=sha256:cb3e83b1d2ecb560da2bf0f95c18fbc8ffbcf7cb94b153d1b84f0e8b971aa85b

Observation 5a50d699-9317-403e-b6f0-fa7c18a6eef9 · outbound

This paper cites Google deepmind: Gemini 2.5 pro, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Google deepmind: Gemini 2.5 pro, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.136766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:24.537822Z digest=sha256:981e11649cb64e155b75d588a81e68eb3c76b37f75f5463c70ea8884eecd7d4a

Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · outbound

This paper cites Do Vision-Language Models Really Understand Visual Language?.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.649412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.649412Z digest=sha256:ca2529d37604bc8a263a22014f1a6e0795a815671d3a1150b6be302b76065de3

Observation 42696a3c-e460-4e02-b84b-5a5fa5eeee81 · outbound

This paper cites VLind-Bench: Measuring Language Priors in Large Vision-Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features VLind-Bench: Measuring Language Priors in Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.757401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.757401Z digest=sha256:69635823988cf31e36b8f5bfa019f3ba35465c73fd8d8bf7eaa441e8fbed15a8

Observation a831130e-7f07-4a94-b749-20bb5a609c37 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.845483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:24.886900Z digest=sha256:b30350ce37c11d321802342f3fdd04a590dfed7314e7f3e6ff26337324f44d8d

Observation 6d34adfe-0899-4e2c-9216-88ccf8b798d7 · outbound

This paper cites Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.989862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.989862Z digest=sha256:c9353c232e12b27ace29ebc420984311f67f1b71a10e21616f9a67eb54c5283b

Observation 15927bdb-d807-466a-911a-3ca3587833fb · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.135570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.135570Z digest=sha256:bb353013948a1af51d0c051d40420fdcdfb534a821144e4c087a06b1be7e85bd

Observation 4b317174-849f-4293-b417-c7c08d6f86af · outbound

This paper cites Mitigating object hallucinations in large vision-language models with assembly of global and local attention.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Mitigating object hallucinations in large vision-language models with assembly of global and local attention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.632342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.210104Z digest=sha256:bbe181ab0cc4e7d52d2266b9500953f22035dff2d4868120db664c6de8b4c38c

Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.290792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.290792Z digest=sha256:c48d9b6432c7cb0c72652b056f8e333b4d54415b9c7073dce16914c08af02c6e

Observation 088d296f-b968-48ad-bad2-d164b2ac3698 · outbound

This paper cites Qwen2.5-vl, January 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Qwen2.5-vl, January 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.372227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.372227Z digest=sha256:0a795411f26734853412bff2ea49c665deb73a2ecfe87c8f6b4799cc157accb2

Observation 0d57f32d-4e42-4dd6-a132-2d26bd3076d7 · outbound

This paper cites Kimi-VL Technical Report.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Kimi-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.467092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.467092Z digest=sha256:f38480e07f62ad992d78806960ddaba5375c9187796e7fa82a1d1217f9028ce9

Observation 9dae0dc1-d113-46ce-93d8-f074cc831067 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.516499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.565607Z digest=sha256:734bb080329a2ea866ad4206a60ce3350868c1e1038a08def13c7893df209964

Observation 82fc1942-6ff6-48f2-a64b-25d2d07aeddf · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.272055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.660579Z digest=sha256:4d5a437f99f663719175a58db5aa79689253c26079d6168ec9ff07d389f0d55e

Observation df664d15-dfa3-41ba-8c56-15d0995bc94c · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.061393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.723912Z digest=sha256:d4aa49bdbb97f67d94c2fe13e22ed0cb0bf26e935c9c6e0c45d9d62938753722

Observation 112c2379-fe21-4f38-97ed-77e6e4658966 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:28.837798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.849248Z digest=sha256:4804f8957c2654088ff92a69b44a88b291333f2b051067168b2e6bec38dc16bc

Observation f2bc3755-82a8-4d66-a4b7-cf976b1af856 · outbound

This paper cites We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.617063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:25.949159Z digest=sha256:a424a6777f00f4971f1633f0ecd44180e4460885268ca2a380700c7d15a034ba

Observation 1ed5b933-1c0a-4ca8-8453-642c384360be · outbound

This paper cites In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9).

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.422583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.085382Z digest=sha256:cb7f06b3081368b0d41247ae88a467d32a228dde6331d9eb426b68d22f8b6833

Observation d64897c2-e6a6-488a-b583-2d0ead436151 · outbound

This paper cites Refer Figures 5,6 and 7.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 5,6 and 7

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.218182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.172634Z digest=sha256:cdd37877f1caf52d1082a705c81169530ad988268079ac022eca4d5558d94cee

Observation 28fd5db4-1fd8-465e-bdae-e4042f821720 · outbound

This paper cites We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.939333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.255914Z digest=sha256:7bb6d07c58c6601e93d601fdd21635d1eb2fb36512fc3824e1ad7c91cd7dd8f9

Observation 5dd5e933-aa7c-4eab-b042-7da77c559e15 · outbound

This paper cites When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.748866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.369271Z digest=sha256:37779ea43f348f30391e6e757e5fed1bf06d82382dec5554bfc4ec62f64d844a

Observation eee89944-913f-45a8-8018-bd2c559f5bac · outbound

This paper cites Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.549741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.458136Z digest=sha256:00106c1a0e8124e6c3d43be40a3eda40f7ed9bc4b065af62eae6230ad542cbe1

Observation 8d70d92a-852e-4597-b085-c4614e156384 · outbound

This paper cites Refer Figures 8 and 9.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 8 and 9

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.330933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.564356Z digest=sha256:0991d14ebabef84982866e6210b7f24be1d925a11e3161e642687a595854f23c

Observation 675d7e65-d541-4b06-a7be-441f0b0639a3 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:27.138585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.675080Z digest=sha256:7f05863cd405244f18597ff0fe748681b6b10fbb97edfa372be2b4d6e8bdba49

Observation 0ffc17cc-26f5-4d94-80b4-55c3c9bb822f · outbound

This paper cites [1], when using the same prompts and data as them.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features [1], when using the same prompts and data as them

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.003109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T20:58:26.760637Z digest=sha256:4f815ea8bfab8d8e7cd21dc77d74a38ca93e71246a28c2b2bd0362e7033fa8cf

Pith citing papers

No inbound Pith citation observations are available.