Pith. sign in

Paper Citation Record · LEDGER

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features

As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2509.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08266 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:26.760637Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bbe823e-462c-40e3-b69f-d81d2a79d9d7 · outbound

This paper cites Vision Language Models are Biased.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vision Language Models are Biased

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.320688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.320688Z digest=sha256:ee7bd1cda6e1fead97b82069b8ca897380b69744c350618a9373229672efb7e3

Observation 401146e2-0882-438c-8edd-f1ed24fb398a · outbound

This paper cites Open ai: Introducing openai o3 and o4-mini, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Open ai: Introducing openai o3 and o4-mini, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.323971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:24.418011Z digest=sha256:1f9de455eb1550c557a3968cdc314246d216a8932953f9eb1e8772b2ee0fc61b

Observation 5a50d699-9317-403e-b6f0-fa7c18a6eef9 · outbound

This paper cites Google deepmind: Gemini 2.5 pro, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Google deepmind: Gemini 2.5 pro, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.136766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:24.537822Z digest=sha256:247ae7ba3886bb25613a8113c616aba6c3451dc91ed5e5a09cb90faffae0ba28

Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · outbound

This paper cites Do Vision-Language Models Really Understand Visual Language?.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.649412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.649412Z digest=sha256:0b834e9481b3d93b70e5a2fa4744d2ca7be62b9577f73fd68482773c901da66b

Observation 42696a3c-e460-4e02-b84b-5a5fa5eeee81 · outbound

This paper cites VLind-Bench: Measuring Language Priors in Large Vision-Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features VLind-Bench: Measuring Language Priors in Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.757401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.757401Z digest=sha256:4d27e3789a8a4a6e354c627a284808274cff1d9e42ce8ec601f13885e160b2ab

Observation a831130e-7f07-4a94-b749-20bb5a609c37 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.845483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:24.886900Z digest=sha256:b222268ab8826d8677ec9144fdc631d2c0bbeb7073e22a3fe43880377a46c55b

Observation 6d34adfe-0899-4e2c-9216-88ccf8b798d7 · outbound

This paper cites Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.989862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.989862Z digest=sha256:5a7cfd69b645132071110fda53cf05eecd8611b78ca4cdbcdce2737c573be416

Observation 15927bdb-d807-466a-911a-3ca3587833fb · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.135570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.135570Z digest=sha256:548e25eb86703e423314bb9b0b5a899e927dbc1d8cf348c219a1ffb8aeee2ea6

Observation 4b317174-849f-4293-b417-c7c08d6f86af · outbound

This paper cites Mitigating object hallucinations in large vision-language models with assembly of global and local attention.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Mitigating object hallucinations in large vision-language models with assembly of global and local attention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.632342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.210104Z digest=sha256:9a5ecf263554599f53ce78035f78fe40a54e57dd124b9579bbf00d731a5a7c97

Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.290792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.290792Z digest=sha256:dd4f5b469031b860b1aee8c8455e733ea0537d95d72911a227389d4bebeaefdc

Observation 088d296f-b968-48ad-bad2-d164b2ac3698 · outbound

This paper cites Qwen2.5-vl, January 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Qwen2.5-vl, January 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.372227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.372227Z digest=sha256:f798fc0c0e4598dee1a9a10484ecc04d5bd288cd12877e38331ececdcb0aba2d

Observation 0d57f32d-4e42-4dd6-a132-2d26bd3076d7 · outbound

This paper cites Kimi-VL Technical Report.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Kimi-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.467092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.467092Z digest=sha256:95c53c8cb00a31b8b0e955c28a8abd81db478b7402903f8d12c79d2891ed3f82

Observation 9dae0dc1-d113-46ce-93d8-f074cc831067 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.516499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.565607Z digest=sha256:3fa444d19d9966f5b494684b2c8646c781de7cf90895c2201949da9fdf2b33db

Observation 82fc1942-6ff6-48f2-a64b-25d2d07aeddf · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.272055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.660579Z digest=sha256:457821a5580e75499a369bb2503c312ab0ede5f561214cb50314865e39f9ad7d

Observation df664d15-dfa3-41ba-8c56-15d0995bc94c · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.061393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.723912Z digest=sha256:e54742c125dd5ca0575982cb0d4ef313dcbac3ce1053b97b29e2225f4c49ccdb

Observation 112c2379-fe21-4f38-97ed-77e6e4658966 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:28.837798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.849248Z digest=sha256:d7e8c3b75dfd6130cf081e82c28673ba28c594604c224f6c176e45bf9e911e62

Observation f2bc3755-82a8-4d66-a4b7-cf976b1af856 · outbound

This paper cites We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.617063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:25.949159Z digest=sha256:73cd30e8b3eefd44413e815aa35157167d15748598fb2ef39ba97b5b3acc8634

Observation 1ed5b933-1c0a-4ca8-8453-642c384360be · outbound

This paper cites In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9).

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.422583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.085382Z digest=sha256:7528851f3c1f5087cc6005fed76af6973638d9cd1afb0345ce131e21c59f861a

Observation d64897c2-e6a6-488a-b583-2d0ead436151 · outbound

This paper cites Refer Figures 5,6 and 7.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 5,6 and 7

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.218182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.172634Z digest=sha256:96750e9533024c4d2209c3e103c3afeaa111f3301047e2eb13362ad12ea3c3f3

Observation 28fd5db4-1fd8-465e-bdae-e4042f821720 · outbound

This paper cites We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.939333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.255914Z digest=sha256:9bfda200ced3c2287c3a5f022f0e0b82fe9f0736511f09ffaf1f775f8f84c0f5

Observation 5dd5e933-aa7c-4eab-b042-7da77c559e15 · outbound

This paper cites When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.748866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.369271Z digest=sha256:83b79fca7a7b698bf6f5ca4aab516cf677f0c8ace86ed01ab5a964730a2bed64

Observation eee89944-913f-45a8-8018-bd2c559f5bac · outbound

This paper cites Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.549741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.458136Z digest=sha256:746f4f2d3c31878704a72b424f47d9306bd7fdf7463220830d9d4d377aab2db4

Observation 8d70d92a-852e-4597-b085-c4614e156384 · outbound

This paper cites Refer Figures 8 and 9.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 8 and 9

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.330933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.564356Z digest=sha256:a96ab6187e0ceaa11b166640cc7839de835c5f62a15277740e7af2c012ed7222

Observation 675d7e65-d541-4b06-a7be-441f0b0639a3 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:27.138585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.675080Z digest=sha256:fe4a2cc95d25d70be5630150a140b4d2c30f29569bffc2601d02391703158753

Observation 0ffc17cc-26f5-4d94-80b4-55c3c9bb822f · outbound

This paper cites [1], when using the same prompts and data as them.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features [1], when using the same prompts and data as them

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.003109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T20:58:26.760637Z digest=sha256:10b3231bee6cc6a4046fe7a6008a25f2e694cb388880c14d1f18663f0cb1f48f

Pith citing papers

No inbound Pith citation observations are available.