Pith. sign in

Paper Citation Record · LEDGER

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2605.22903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22903 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:58:35.466105Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:04:54.866011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbe1f486-96d7-424c-91f0-7054527bb109 · outbound

This paper cites Qwen3-vl technical report.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Qwen3-vl technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.414805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6024408a7c838064149fea43a34fcbe61191dac5e2032a13fb507168c3b09b33

Observation 0abe4d4c-c03b-4302-85cd-2b32521e79bb · outbound

This paper cites Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.417944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6a46d6ee903e1309b2e140d35db1454535ba86000d2ac62aa0f9fe864afd8499

Observation 552c21fc-48ec-4058-9730-eab872861c57 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.404362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6f6e7482d5fa7cdd15b0ac5da3d3bb85e534b37f6705f2d8e3ce28c628452d0b

Observation 82c008e2-0908-4703-a7c5-a2ea5468e827 · outbound

This paper cites On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.408744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:fdcc7672e7592c281d9432325e78e1ac5bc2c5fd7e3fd04d25a4c56e55bd2309

Observation 507598b1-d0f3-448f-af5a-2f5b33506d7a · outbound

This paper cites Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.430218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:d2e13ec8edfd83ac374877d4a74c8868f6764e632f229be7517a68eb5783bf07

Observation 2667e7d8-a612-4032-a7b8-f24cdd80279c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.433280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:c7b07638dd272b74f08d3ef9ef98dc6e167b0e2447750a99f2ba62327f9f15b1

Observation 44229f4b-d8e8-4844-84ad-7a99734b9ae4 · outbound

This paper cites Does ob- ject grounding really reduce hallucination of large vision- language models?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Does ob- ject grounding really reduce hallucination of large vision- language models?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.436022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:98b000347da9fb3ce304b1af60aa0eaf423df192b1e24149cf9fc069f3a4b9f1

Observation 947ae1ed-30ae-483b-b353-83d142c2bb61 · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.455640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:731ef8a44d103472b041fe662b0d111ad66644e50f5157d92bf1cc84c0c069fa

Observation 61bd3256-36dc-4fbd-9ac0-82fbdcb23f19 · outbound

This paper cites Do vision-language models really understand visual lan- guage?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do vision-language models really understand visual lan- guage?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.366989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:b5f1f40e6d44f7189fa9583c8ed0ac320748178746b45cc3e9f7bf55f372828f

Observation 4681251b-5cfc-4e35-ad9a-11c63022dfb1 · outbound

This paper cites A survey on evaluation of multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A survey on evaluation of multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.360769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:3ef3e94816699769334c9b471b8d6fa563db851ee0f162bec818de5d3d61dde2

Observation cfba36bb-ffba-4f15-8cf3-b8625cdd2523 · outbound

This paper cites Robustifying vision-language models via dynamic token reweighting.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Robustifying vision-language models via dynamic token reweighting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.370424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:8daa7ed017857db432a5497edd49f14bd0d6528ec2d6257a4c175fffbdaa5cfc

Observation cffb5a42-1fe7-4f34-a153-2fab6f2df224 · outbound

This paper cites A comprehensive analysis for visual object hallucination in large vision-language mod- els.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A comprehensive analysis for visual object hallucination in large vision-language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.351708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:68326d90aeb36039d85b446a4e9b4c7cbc843025f0df83b2c42ad7e9e5dd55ca

Observation 4c60f163-65b7-4e8e-bc91-9aaae62c3554 · outbound

This paper cites Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.354421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:dac5f37f5123387ef048d50609cc92d0df4a0f36657ae59ef65a76e0c6ed5dc7

Observation 66a0dbf9-2dc0-4d41-9339-4c7251f2b2f7 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.370018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:d9cc1c75d8a33714e97c7f0e26290f3d33948191fa0d6651066c40f14d569e75

Observation 2e54ad26-7c3b-4709-a535-267b303a00a7 · outbound

This paper cites Halp: Detecting hallucinations in vision- language models without generating a single token.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halp: Detecting hallucinations in vision- language models without generating a single token

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.372796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:4b182b69b624e70df27f7a0a0369b07649c9b8fabe42fa944f4750b4103027fd

Observation c5d31bb2-1eea-4429-b775-cb854953af8c · outbound

This paper cites VLind-bench: Measuring language priors in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? VLind-bench: Measuring language priors in large vision- language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.396618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f7b5325a37095ed82f97cb1eb26dc9660385565edd8bed52b0281c78d9a59a70

Observation 61a0e2ee-6037-406d-91cf-cec25196026f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Evaluating object hallucination in large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.441466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:9460b2c73404c31251d5ef017d72480160d663e47f76e83177ed22c940c67ff7

Observation d9b41419-e78a-47a7-973a-0a00a3ca8605 · outbound

This paper cites Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.450173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:2589d10ed395853be08b3c41683911857da5c837a440a3cb32550cd0fdb6a6b5

Observation 9a45dd42-6685-4f8a-971f-8f440030f551 · outbound

This paper cites On the predictive power of representation dispersion in lan- guage models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On the predictive power of representation dispersion in lan- guage models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.438835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7dbeeb65420a02a6ce46c74f21c188d07e424e5eec9a9c3f877e61e8a6f4951a

Observation 27a0c21f-cfe2-4c96-a0ca-7d1894b3c780 · outbound

This paper cites Visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.452744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:310825d9e1c17c20726f91b7126ca57b3b28fe461c0b9d7ba43573824df07e20

Observation 45fec4a6-30ed-4fec-bacb-23411fda4561 · outbound

This paper cites Improved baselines with visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.399786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:83714873cd7f6f1eed8bd7c7c7969014a66f7718c516faeee3b33c1d0d027c96

Observation d5a0d606-df2b-4ff0-8f2f-633a89bf4bf0 · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.444251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:707ebe06c595fed5a38d35c11efeaa0878f2a3dc2c367365cf5305ce9b47fcc0

Observation ef4fa151-afe6-440d-b380-f383da9658f4 · outbound

This paper cites H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.378422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:62779dbea10ee35522884894fd3a63ca9cf37d10ea1329b95895fd71128cd73b

Observation 7e12274c-f23f-4824-8eb8-f4eb33aa2341 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Learning transferable visual models from natural language supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.381260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:3832f0fd030017bdb611ce70e653cebca32a06ec006c7e0166f09764a9c9fb80

Observation 260af8aa-7554-4818-8223-def6d7ceeac8 · outbound

This paper cites Sam 2: Segment anything in images and videos.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Sam 2: Segment anything in images and videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.384382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:622bfe4affb0184d56e150d7c210a5802c8a0ed103e5e117c30d450bdce44a38

Observation d93c712f-4fe3-4800-b463-ff00bc216cc2 · outbound

This paper cites Object hallucination in image cap- tioning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Object hallucination in image cap- tioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.387523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:bae7b74904c3e8df3f9f8ada90aca5e43a4d1d1967185ce05243cecca5a9e9d2

Observation 0daaaf25-dd19-4fb2-a546-397cc0b490e5 · outbound

This paper cites The effective rank: A mea- sure of effective dimensionality.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? The effective rank: A mea- sure of effective dimensionality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.390543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:108029e410846abf6d7666fd527acd6246b78fc6ccd84fec1967235defff710d

Observation bb71e7c2-0b06-4a11-9331-f8655048e47c · outbound

This paper cites ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.374343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:d9be5a380cc25a374de6ea1e795b14b2ad7177322093e14b7db8142efec27f04

Observation 02ef412c-421d-4547-8520-da619b63df75 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.363858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7ad842f89bd70fee84ed28c86e841e719f6d342fe8735f575fa630f0beab3d5a

Observation 120be32e-0543-4118-9022-460690a2f274 · outbound

This paper cites From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.357415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:ec16d1fe6f8e30150daa166b663e5b4dab6efd221cc5e1837ffb20e1d2b8ec00

Observation b0323300-e024-4c4f-87bf-9b40bbf91ebc · outbound

This paper cites Openai gpt-5 system card.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Openai gpt-5 system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.375826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6a61d69627edb3d2bec533cb505eea3690491ea8c2dff19f5e05a92c289fbfc2

Observation b7f055d5-0f5d-48f1-b2d3-c4e6dc8bd7a1 · outbound

This paper cites From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.341090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:595618f3f5852db69860a7848a3fe94be17102a7fe2bb407a20797dc36aae93c

Observation 0ad782f6-701b-4e5e-8fdb-606186261e0d · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.343505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:0ae4be73413ceff7ed39b7f74817d225283c694433c76c91ebecda15c5083d77

Observation d8014817-b633-43a5-8298-515acb0150f5 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.346183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:fdb3daa90ffaf4838a3df8864a5dad135db30233176d783bb3304ea69f94781e

Observation b23ff460-f89e-400d-9b19-7971c18df27d · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.348881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:dbe82f9446cfca545b3380d2593ea1855aad1863fdbe7f6ef8d8bd5d7aa23190

Observation 146fb5d1-6a65-41a6-915a-ab71d202d2a0 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.393477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:29845eb2236f35f391412ffcba06460372684661cfe1a325c9aac34ba271d056

Observation 7c74344a-82d7-4d5a-afdc-65411650281d · outbound

This paper cites Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.458707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f92712f25795db1e5aacd234764e7351392cdd88bd36b6018916a328e43e7b48

Observation fcd32957-bc16-4b8f-9ce6-1d80b3451caa · outbound

This paper cites no images.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? no images

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T06:00:24.447119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7422db7cdc7cb7f1bb29034e78b0e75cde01e0ef73afc09c2a75edf90b5b7e34

Pith citing papers

Observation 0cfc13e8-9803-4225-9a17-91e9ffaaf842 · inbound

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs cites this paper.

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:54.866011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:54.866011Z digest=sha256:60e5a9891306ef9347c0e4b36a8a39636db2c60fdf2efd8157ff81acddc66244

Observation eb984bca-97a2-4abf-9791-12f2555fddcd · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T12:08:02.370883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:08:02.370883Z digest=sha256:87142295c3dd906f88594711ff723f6deb239d3964caa9adc3bb83dbb150daa4

Observation 7f2c230e-7d58-4dee-82b8-89759e04746b · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.125422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.125422Z digest=sha256:f6903c5a223a6f31f649a766af63bac87aa42bcaaf1d9ec233d71afb1c37fbac