Pith. sign in

Paper Citation Record · LEDGER

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2509.25339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.25339 v3

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:47:21.298721Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:06:09.166363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:39:16.760441Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4576053d-b36e-44ee-8d3e-6b87bef647f0 · outbound

This paper cites Analyzing the Behavior of Visual Question Answering Models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Analyzing the Behavior of Visual Question Answering Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.040327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.040327Z digest=sha256:a4578eee1c32c0a9aca660d488fc576f8fd58192e0bcc6a910f345d065f7e32d

Observation 36fd9518-4c04-4cea-8dd2-b5fa5540390d · outbound

This paper cites Don't just assume; look and answer: Overcoming priors for visual question answering.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Don't just assume; look and answer: Overcoming priors for visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.424122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.044962Z digest=sha256:ef85e79a19fbe2020aa41d6346bbe977443ac70f22d3ffcfde4098ae88ec1cf4

Observation b88778ee-6288-4281-8351-953ba692dc25 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.401379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.049331Z digest=sha256:3f71f90c2b8685cdc334f7fbfdb20d5406f0576b0cfacee16cc5f6b47ffb818b

Observation d64e739f-4735-4809-b787-2a826ecbf900 · outbound

This paper cites TouchStone: Evaluating Vision-Language Models by Language Models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes TouchStone: Evaluating Vision-Language Models by Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.053486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.053486Z digest=sha256:9e0d044fcbe21274be949c85b2ef871ed2809d2bfcaba8c459b54fb414ef2fb0

Observation 0c0907f8-1167-4777-8fe4-5de53dae2734 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.058171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.058171Z digest=sha256:a76f4bb3d858b6f3a8fb119e341fbddc329c64940615f974e303dc9f6b2b3f82

Observation 9b10dda8-45ad-4bc0-8edf-8be54046ef0e · outbound

This paper cites VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.063096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.063096Z digest=sha256:1c125d30dae1fcb4922e5f04402b5338ef384633dcc61501a4e9cc449f19fbb7

Observation 4f86f860-a92a-44e7-b326-bebd1325be26 · outbound

This paper cites An Introduction to Vision-Language Modeling.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes An Introduction to Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.067942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.067942Z digest=sha256:4fa0d2dabfb13bc200f266864db29fcc5cd10916ab7b0ac0c9d81cad87528110

Observation 3ee4b375-35a2-4ed8-951f-e29d1c51bba5 · outbound

This paper cites Rubi: Reducing unimodal biases for visual question answering.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Rubi: Reducing unimodal biases for visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.380876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.072234Z digest=sha256:6b53529c170e070225e5f1e86c63816739e0fc91454f58a193e51215e0080abb

Observation eba9351e-66a9-47a2-ad2e-540a1cbf8742 · outbound

This paper cites Frankland, Thomas L.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Frankland, Thomas L

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.365038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.075941Z digest=sha256:927a225a950452cfa4ceb9f6dc335a6e512ecb0e104765be5d36dcb5f07481f1

Observation 5e17b1aa-5508-4161-917a-a22e05ba2199 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37: 0 27056--27087, 2024.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37: 0 27056--27087, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.348689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.080040Z digest=sha256:e305b516b235ba03cb4102686d475e4760dfc381dab91cf6f72d5044fb14b2a6

Observation eac2d1b9-92f1-472a-bcbc-4ca0ba736604 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.326194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.084154Z digest=sha256:608b312f791b0d9d18a087b4e6526ade48ab21973e077e1a97e972e62447dc2d

Observation b5919fa9-9530-42e9-ae14-9f2d8b8e5407 · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Dense and aligned captions (dac) promote compositional reasoning in vl models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.298236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.088177Z digest=sha256:9815a10da2db8ceccfcba9004bc17f70d5e5cfe4c30f6ff5b4a2133473cac690

Observation acea20a7-fd49-4782-8089-bc3a54530cbe · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.279557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.092076Z digest=sha256:6577bea96e60a671c7b544228a7fd233acb5ad021927c5742317246ddb6f459e

Observation 240e69e9-fc52-4a20-909d-48b0a6ede26c · outbound

This paper cites Datasheets for datasets.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Datasheets for datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.095843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.095843Z digest=sha256:1c5c1a66735460b5845fa7fd97606a2fdfa69eda0b009b63ede9893a0231455b

Observation bd2025af-03d4-40ec-a42a-b14a45d18f16 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemini: A family of highly capable multimodal models, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.099947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.099947Z digest=sha256:6ee3d443c215c6196fe61211bcd184097292aa82f848ef1ed640d8c8317fe5b3

Observation c2fa9098-eb69-4e5e-b187-95de0361c95c · outbound

This paper cites Gemini 2.0 Flash Model Card , April 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemini 2.0 Flash Model Card , April 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.250312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.103540Z digest=sha256:b7b6095467e2a44b26d9d4404231c880b5e9e219a734f9a02a7314f64bc08221

Observation 5e20987b-301c-43fc-ab17-94d3f2ad7b21 · outbound

This paper cites Gemma 3 Technical Report.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gemma 3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.107273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.107273Z digest=sha256:b66922f9c3011e2a782000802dcc650fb9d0c29ad8334ed63077c1150509a150

Observation 6f355e9d-40b0-4cd3-b177-8eea9ed46ac1 · outbound

This paper cites Announcing Gemma 3n preview: powerful, efficient, mobile-first AI , May 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Announcing Gemma 3n preview: powerful, efficient, mobile-first AI , May 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.230181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.111628Z digest=sha256:3a013ad9b43c96599f53efade2d41c86cdd8e0d26e47e420cdb8822bba55c802

Observation 7437e8e0-9bab-47b3-a9c9-376ea8443af4 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.213241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.118171Z digest=sha256:1c920f290f57c025b91a04a3df9967f7af4b6f0b86774121c2b8eb173b2849e9

Observation f577566f-61d2-4934-9b3d-35f29ed3da87 · outbound

This paper cites Horizon Alpha - Advanced AI Language Model , August 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Horizon Alpha - Advanced AI Language Model , August 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.194800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.122326Z digest=sha256:eb065bdee9ee5e7d8053bc1b052173ead7f996b2f5baaa526e04235fc3deea7b

Observation 9f9fe8bf-3b2b-48e9-99bc-53a8a006fad1 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.125964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.125964Z digest=sha256:22c8389be4881cc97f8822c8ab79057ab1a8c61cb14bfeb5c13b778b75aadc00

Observation f251e5ee-0697-4eb8-a8c3-24fac17a7b2d · outbound

This paper cites Conme: Rethinking evaluation of compositional reasoning for modern vlms.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Conme: Rethinking evaluation of compositional reasoning for modern vlms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.155519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.130094Z digest=sha256:8baabe76e8926cbc966f1f29d9e6d25ca5875690a2db6cf75506f4a3bf52718e

Observation f6ff9b14-0eab-41b6-a60f-77e3d8b0e79d · outbound

This paper cites Large language models are zero-shot reasoners.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Large language models are zero-shot reasoners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.136115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.135099Z digest=sha256:815700c3b03978232644bf8846ca8a63215f78f24523719c2caa1e1cb9541788

Observation a861aeaa-a769-43b1-a18a-347c688b418f · outbound

This paper cites Dvoichnye kody s ispravleniem vypadenii, vstavok i zameshchenii simvolov.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Dvoichnye kody s ispravleniem vypadenii, vstavok i zameshchenii simvolov

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.117058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.139304Z digest=sha256:3e56a588e1c44dd78fb1727cf8806397da0ce84adc9f50baee43fee0304bb7a1

Observation 790b2c43-80eb-455f-98d9-5eb7cd5321a3 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes OtterHD: A High-Resolution Multi-modality Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.143108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.143108Z digest=sha256:e91839cd7996d216d27feb0bd69d5e97a601e309b6229b64d1ec1e800a59cc11

Observation 3eddabbf-1ee2-43b6-8ea9-a3153d23b222 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.147071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.147071Z digest=sha256:7cdbd6d854b046abf000b7749578237aab5ad41046c5d931df4931cc93c5e421

Observation e1035144-e976-4713-ab9f-1a64ad82b23e · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.151496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.151496Z digest=sha256:3a5ed51a18834373204a6970faf10da75b2d1b512f5fa964c1d5947ef073ca6d

Observation 04f7edbe-9fa1-4a7c-967f-4bed859fc1cc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.099750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.156664Z digest=sha256:3f01ae4ed0240ea4a63e3940166aa302c7d894955c1d32a90ffbc888d35d185d

Observation abe1a6c5-fdbc-4a82-8d77-db8e8afadbe3 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Omnibench: Towards the future of universal omni-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.160513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.160513Z digest=sha256:3e70e8f02ec8f1ec70447026dfd108e873568a9b40db1bbbaa8cf18cd7d6daa2

Observation 7be46e89-72bb-4df3-beef-1dfb9c7b4398 · outbound

This paper cites Introducing LFM2: The Fastest On-Device Foundation Models on the Market Liquid AI , August 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Introducing LFM2: The Fastest On-Device Foundation Models on the Market Liquid AI , August 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.085503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.164491Z digest=sha256:80e4e96c5db34e4644aed169fed824af3b0474916613f193100c76e02feeabfc

Observation 88745ae0-5b33-4b9c-80e3-c104f21a1df4 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Improved Baselines with Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.169973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.169973Z digest=sha256:55171dd80d9fdb1230ab6a7181ad953119ef8f0323eb522d3777a0044138f82d

Observation 5f1aa9b3-f745-4527-ae8d-3781a72d6564 · outbound

This paper cites Visual instruction tuning.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.068206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.174736Z digest=sha256:189a8c8e31dc8b1a928d66028a034e065807ea026968dd14c591a69288d5c9d4

Observation 9131d634-98cb-4b2f-8364-d37164c47830 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.178919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.178919Z digest=sha256:4ec4b16f065bac4fe81e8fd38487ea6692f0ad0c9d68287af8112c8eadf883a1

Observation e4702157-2270-4dc0-9fcb-6d4f0173892a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.043783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.182824Z digest=sha256:c03bb7f34f3ef189c461be7a082244582b6354bc812a68a136b10fc4a8477a9c

Observation 88151aa5-fe11-456a-a41f-07f3871f88af · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes SmolVLM: Redefining small and efficient multimodal models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.186700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.186700Z digest=sha256:354c81ebcb7d71e3902f8b531e933cdde71d98e988c325cd70e0079b5d1e4629

Observation c014dd17-f679-4637-b26c-a042e91c0753 · outbound

This paper cites Umap: Uniform manifold approximation and projection.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Umap: Uniform manifold approximation and projection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.191082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.191082Z digest=sha256:714ccc90b39a6f74b805c4ba92dd4905d513bc7eee339f259217c247add1c1cd

Observation 42490883-eacb-40aa-9618-25eb19f76c80 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation , August 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation , August 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:22.022615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.195198Z digest=sha256:468d5e6bc1a704d77fbadc202b8f84b2d37161130f1acb4127d122608bd99513

Observation f18667f4-307f-4d7c-b4ea-65ca6c822495 · outbound

This paper cites Lafter: Label-free tuning of zero-shot classifier using language and unlabeled image collections.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Lafter: Label-free tuning of zero-shot classifier using language and unlabeled image collections

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.999220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.198933Z digest=sha256:4f5b0be582ca55c37f8ce07debd435af227e813e0e3bff56cfc8c2630f36aa81

Observation 62648a1b-cc44-408d-be5c-3143792037bd · outbound

This paper cites an unresolved cited work.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:47:21.976497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.202438Z digest=sha256:e48d1d99b78f65e6450609451c53401a55a9980640eb847f5f4353873e30bce9

Observation a51e2431-9cff-4337-bf93-b9962ae81c41 · outbound

This paper cites Gpt-4 technical report, 2024.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Gpt-4 technical report, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.959217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.207164Z digest=sha256:5a4d7a82cfe14898e8b3b35e00251ada9fb0a3792375ca1e841dafa8c882f5e4

Observation e2626458-faec-4587-b9ce-05513ee4366b · outbound

This paper cites OpenAI o3 and o4-mini System Card , August 2025.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes OpenAI o3 and o4-mini System Card , August 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.937484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.212289Z digest=sha256:f7197ad14275edc47399763809c202f95e9135acdf5bb594c03921ee0aa81b39

Observation 9bac3cc7-64ce-441a-89c8-b9b6b325e354 · outbound

This paper cites Teaching clip to count to ten.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Teaching clip to count to ten

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.921064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.217036Z digest=sha256:5e869a8903ce6e7c92f31794542a571b9f22ddd7bc5c1fd44d067c57d95d2c14

Observation f1cdcc5b-8858-4d7c-b2c6-f3db5c05edca · outbound

This paper cites Humanity's Last Exam.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Humanity's Last Exam

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.221916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.221916Z digest=sha256:ed0ac56cd7de212fd806f3038968708d91fef745436866fe875433636bb836b7

Observation d4652664-a200-4bfb-8d52-05d83c12a132 · outbound

This paper cites Scaling Vision Pre-Training to 4K Resolution.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Scaling Vision Pre-Training to 4K Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.226501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.226501Z digest=sha256:699749d729c846dc9c9d49281973e87cf67e7af99d57a3fe90a2d8624000949d

Observation cc0ec413-9a2b-4721-a793-212da32eeb2b · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.230285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.230285Z digest=sha256:a1489836667d0e2ee59d15440130f365d0e4e2c50539c993f5cb0310a2a8acf5

Observation 7c5975e3-ade4-440b-9b3b-f59e1b89f7b2 · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.234168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.234168Z digest=sha256:4b6955ab32b17aa83e63b3792e0e14e16a27533423b4d6dc91628aaa16a9190f

Observation 45a2fec2-73df-4032-b5fc-bf50c58878fd · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Chi, Quoc V Le, and Denny Zhou

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.238120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.238120Z digest=sha256:957dbc40efbcf6fdca6cfbf55c503c606a847129ffffef0ae1301db93bd58461

Observation 1e666cc3-0a11-4a25-9bc3-288cd37cd3c5 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.242391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.242391Z digest=sha256:934f5745a7c4e936683cf20c1715dd86750d2a257447540918846baa25c11983

Observation 1c13a608-bfde-4df0-a934-a2a56f5654b1 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes V?: Guided visual search as a core mechanism in multimodal llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.881817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.246432Z digest=sha256:e113e64902be4bbea0175c55db3b5028521cfa4230212f0a135650f0b4678ea5

Observation 12ca737c-8477-4d65-b0c4-96001c0d6c4b · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.249979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.249979Z digest=sha256:f063c3fb0aaa4ad709608998fdac335295f66a4388cc970f0e84e915fcbfc751

Observation 10f1514b-204d-408f-af04-91d59147c537 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.254143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.254143Z digest=sha256:8f8dbba0c0259077b76b1019e12ed53c9f8f4ca4116d0616704e4dfe4a5fd68d

Observation 31e05dc6-e450-4575-8263-99b7c967bfb4 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.258016Z digest=sha256:0b24ff8c290845d37390c77273c121f256f26c57f14e0ef9bcd9411cc8cbe5d8

Observation 480d0587-444c-4c5a-a899-75d96eefc4fd · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.261811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.261811Z digest=sha256:a4bc58b48f09f24bccf6242278279011fccdb6d0bfdb8422ea556198475ac909

Observation b23e0a24-5d2f-4bed-807e-4584ed11507a · outbound

This paper cites Yin and yang: Balancing and answering binary visual questions.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Yin and yang: Balancing and answering binary visual questions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.864345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.265911Z digest=sha256:34b4a5ab1ce5a27e2cdf71e4dea3b3f80fe05148b08508a4b1a2ab83f6aff11c

Observation cb33b780-0800-415d-8170-c2bf6728fb1e · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.269398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.269398Z digest=sha256:b800a1826021881a4010c284afff1d72a4f06315fa4f3bcc9fe067ec2b9e0322

Observation 63efc223-c776-498c-bac0-d3b890c6cbae · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.273387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.273387Z digest=sha256:9650a0a2b764223d76a17a8e677cfb6c25f51f5401ac59f61f4547d88e236f6b

Observation 62803457-e2fe-4e8e-8a31-9fc12368c114 · outbound

This paper cites Why are Visually-Grounded Language Models Bad at Image Classification?.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.277306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.277306Z digest=sha256:2e8d5c75c2c4bdb023201a486f2b7c76a326a414058645a2a4e22c204a4289ed

Observation 6054d7eb-f7fb-4438-a0a0-69574ec18ba8 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:47:21.844294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T15:47:21.281527Z digest=sha256:e3357d46aff8d28de6ea57fbd3d6deb695dfc795a6c488386a10f0b0094ae077

Observation c7314233-8846-4750-9ec1-ab18580096ea · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.285423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.285423Z digest=sha256:546ac04f5d460c8d7849c9bd0d71f89a05eb04014acb28259bbb2bdc1ed7ddba

Observation 36af6328-f6f0-4cdb-b958-695a745e5880 · outbound

This paper cites @esa (Ref.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes @esa (Ref

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.289600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.289600Z digest=sha256:4068282e4a7fbc8659825d851f66998725100f37b3e7866a706fe4136ee2d700

Observation 37e11a02-4912-44fe-8084-ecb42ece83a2 · outbound

This paper cites an unresolved cited work.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.294663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.294663Z digest=sha256:e5e1bf92fc149fed63c3dccb9a0a2f2576e48f2e24aa28269201d88b0896aec7

Observation 64999212-2f4a-430a-a251-1ef55cd8d540 · outbound

This paper cites an unresolved cited work.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.298721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.298721Z digest=sha256:5d10f78e871aa5da05fde1d7bf568666529de0d3b02044cfdc3a1afba73005ce

Pith citing papers

Observation f2269bc1-a5ce-4470-8187-3f2c773314d1 · inbound

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation cites this paper.

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:16.762070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T21:06:09.166363Z digest=sha256:9db258e9356c56d7c905c93a40eb3d8baa87098341be210e2f3bf2c89f4ca3bd