Pith. sign in

Paper Citation Record · LEDGER

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2605.22903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22903 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:58:35.466105Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:04:54.866011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbe1f486-96d7-424c-91f0-7054527bb109 · outbound

This paper cites Qwen3-vl technical report.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Qwen3-vl technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.414805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:fa8ba03c7d844318625149c3d8688fcb5e875844948c93a84c099a18a89c2701

Observation 0abe4d4c-c03b-4302-85cd-2b32521e79bb · outbound

This paper cites Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.417944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:5d870a1c4af7f96ae0eb33c1ce40c8be5ee2c895a28e37cac077d4add88d72dd

Observation 552c21fc-48ec-4058-9730-eab872861c57 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.404362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:9d782bba8a531bfb67a571404bdab475118372395aa7ebf95fe54e59f166cf5c

Observation 82c008e2-0908-4703-a7c5-a2ea5468e827 · outbound

This paper cites On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.408744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:5c804e627e3c00b3c35c9c5e080b2d30afce7ec1001c883f77c622f8a235cd77

Observation 507598b1-d0f3-448f-af5a-2f5b33506d7a · outbound

This paper cites Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.430218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:1bca72c06b0cea5c31391a30fdcb04402dbc1dcbab4512cd55e89cc37f849716

Observation 2667e7d8-a612-4032-a7b8-f24cdd80279c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.433280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:c9f48e3008326567d28a35f084f5fb2f59224f75cc5f3dce0716c2b4fbcdaf6b

Observation 44229f4b-d8e8-4844-84ad-7a99734b9ae4 · outbound

This paper cites Does ob- ject grounding really reduce hallucination of large vision- language models?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Does ob- ject grounding really reduce hallucination of large vision- language models?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.436022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:cac25d171e1e6a69338271e7b1fca8e422a29f853ec8e3c967cf0b809f153914

Observation 947ae1ed-30ae-483b-b353-83d142c2bb61 · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.455640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:343f625cf8cfd292265f049ed05ebfe2a40ed3ba0545c86feaeb591c48651d39

Observation 61bd3256-36dc-4fbd-9ac0-82fbdcb23f19 · outbound

This paper cites Do vision-language models really understand visual lan- guage?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do vision-language models really understand visual lan- guage?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.366989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:1b8af12b183c4ef2ecb74521a8f8305ad778774a165711fd21bbd83a25da3fda

Observation 4681251b-5cfc-4e35-ad9a-11c63022dfb1 · outbound

This paper cites A survey on evaluation of multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A survey on evaluation of multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.360769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:2bd7d6c299af20205e4bfa844a19562a0192fcf7bb07ab59ff42757fad5c9424

Observation cfba36bb-ffba-4f15-8cf3-b8625cdd2523 · outbound

This paper cites Robustifying vision-language models via dynamic token reweighting.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Robustifying vision-language models via dynamic token reweighting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.370424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:20f2298ff970e2b267c07a17e838b8d0c7a2a4c8eb2170ca38a2905244f738f0

Observation cffb5a42-1fe7-4f34-a153-2fab6f2df224 · outbound

This paper cites A comprehensive analysis for visual object hallucination in large vision-language mod- els.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A comprehensive analysis for visual object hallucination in large vision-language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.351708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f6dce3ca29611cbc3b0dd68c77fb3516766ab475a81fff523488695352282de2

Observation 4c60f163-65b7-4e8e-bc91-9aaae62c3554 · outbound

This paper cites Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.354421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f7f345a73bff34ee5b74ffa0461b175cf7c6bbe9212a10e3e0ead2adb04b1a0b

Observation 66a0dbf9-2dc0-4d41-9339-4c7251f2b2f7 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.370018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7f7c5eae2dd000193d549170efff6140cbb18c642cb79cf8154c0d1522f0885a

Observation 2e54ad26-7c3b-4709-a535-267b303a00a7 · outbound

This paper cites Halp: Detecting hallucinations in vision- language models without generating a single token.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halp: Detecting hallucinations in vision- language models without generating a single token

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.372796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:1b0c24e10b734305f0d8b46c5e39af36de42417826a7d3f6f05b490c92a7b72d

Observation c5d31bb2-1eea-4429-b775-cb854953af8c · outbound

This paper cites VLind-bench: Measuring language priors in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? VLind-bench: Measuring language priors in large vision- language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.396618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:4cddd9e9e578496a5422ca8135f956ce92fc0bf01afd22ae00e369de428dd38e

Observation 61a0e2ee-6037-406d-91cf-cec25196026f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Evaluating object hallucination in large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.441466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:3020b972b1b413462aa264b4d53d65b0c1ed809b1c63204a371133e005168cf5

Observation d9b41419-e78a-47a7-973a-0a00a3ca8605 · outbound

This paper cites Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.450173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:4278c8bf7097b38e8a06b51c96cc34a04de04dd9dffe84e3ba50fa8a0bd55f35

Observation 9a45dd42-6685-4f8a-971f-8f440030f551 · outbound

This paper cites On the predictive power of representation dispersion in lan- guage models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On the predictive power of representation dispersion in lan- guage models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.438835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:2ce045b741ce1fdd7abcc0990997f1429903e7282630ea254e0666499213ba84

Observation 27a0c21f-cfe2-4c96-a0ca-7d1894b3c780 · outbound

This paper cites Visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.452744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:295249b350b557205892816a5ad8fe996d1cb91480da9113440dfbb3fef65448

Observation 45fec4a6-30ed-4fec-bacb-23411fda4561 · outbound

This paper cites Improved baselines with visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.399786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:16d3d990a032330ab72afc238253e4614b1eafe11bc28279e71edf4116be8eed

Observation d5a0d606-df2b-4ff0-8f2f-633a89bf4bf0 · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.444251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:543bc26bb656ff4a4e3eb8b1c85a2fdc950a72c7519287ec1f8bf707e7df6da1

Observation ef4fa151-afe6-440d-b380-f383da9658f4 · outbound

This paper cites H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.378422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:c9d9d34abc1ddf0d8226071198ca13a4ea6de135bec05a8f29c753ceb3aad793

Observation 7e12274c-f23f-4824-8eb8-f4eb33aa2341 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Learning transferable visual models from natural language supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.381260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:a3e9d5e59020286f6008757a2b3b76b1ee489533f1ec52a1dc311d7877b24e68

Observation 260af8aa-7554-4818-8223-def6d7ceeac8 · outbound

This paper cites Sam 2: Segment anything in images and videos.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Sam 2: Segment anything in images and videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.384382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:e53047f2727a2a8219e6372d986c0b402a4bedd6ecc9e313aa74170083c9999a

Observation d93c712f-4fe3-4800-b463-ff00bc216cc2 · outbound

This paper cites Object hallucination in image cap- tioning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Object hallucination in image cap- tioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.387523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:cd7ddcdcc5e178fe9cb0ecefa9797f6d2ab9d90019f866934c0f50feae9612ff

Observation 0daaaf25-dd19-4fb2-a546-397cc0b490e5 · outbound

This paper cites The effective rank: A mea- sure of effective dimensionality.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? The effective rank: A mea- sure of effective dimensionality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.390543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:a52d8f2aa14018d2735d2693f5c93eadb382a3bacb16c07418e6643b8d0635a9

Observation bb71e7c2-0b06-4a11-9331-f8655048e47c · outbound

This paper cites ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.374343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:9e08df66f6c2a94681aeebd9bc8e694f017388f2126811726fd278b37850c776

Observation 02ef412c-421d-4547-8520-da619b63df75 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.363858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:d433b0e9cf1e1d1b0b8d0fa96230dd922bad447351e66eeb0a28d3f6a65da248

Observation 120be32e-0543-4118-9022-460690a2f274 · outbound

This paper cites From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.357415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:e88a1381b22efae0729ca6e29a908d7ddcd210bf7f424a95c72077779f284b35

Observation b0323300-e024-4c4f-87bf-9b40bbf91ebc · outbound

This paper cites Openai gpt-5 system card.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Openai gpt-5 system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.375826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f57802d0ecf9038b6c505d5bbeba5118c6582487868301f802f37c576470402c

Observation b7f055d5-0f5d-48f1-b2d3-c4e6dc8bd7a1 · outbound

This paper cites From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.341090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:ebcf7bf78d05e5914b4324dc517b8a14b4973e0bc16586437dba3ae517587b1e

Observation 0ad782f6-701b-4e5e-8fdb-606186261e0d · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.343505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:0d336c6ff42bbc59d831e62f92c6df336f2acede6d76daf2dd3639a0b10ba629

Observation d8014817-b633-43a5-8298-515acb0150f5 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.346183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:31b3bf81211b8b5ee1cbf83073a8a46ffbe874d0fabfcd34410dc359b072f86d

Observation b23ff460-f89e-400d-9b19-7971c18df27d · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.348881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:eff066778f92994cf330a39ce0087fedede59850a99a5eca746b1744d01a42ae

Observation 146fb5d1-6a65-41a6-915a-ab71d202d2a0 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.393477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:bd3d700102355af8504c1d28a8451b68efe7d6c047db8a9c279fd54d61b97789

Observation 7c74344a-82d7-4d5a-afdc-65411650281d · outbound

This paper cites Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.458707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6d8fe63cf50dbe9b602b57eb742a3044b39ba291508b10b0fe915755f522be16

Observation fcd32957-bc16-4b8f-9ce6-1d80b3451caa · outbound

This paper cites no images.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? no images

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T06:00:24.447119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:34ccf0c2c474ad18ef33ec1286bd60473e3a741016c142849aba273c59980f0b

Pith citing papers

Observation 0cfc13e8-9803-4225-9a17-91e9ffaaf842 · inbound

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs cites this paper.

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:54.866011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:54.866011Z digest=sha256:7ebc78f18fc0e6fa1da7fc30aa441850969943f812cff33f1bf28e862de35d1d

Observation eb984bca-97a2-4abf-9791-12f2555fddcd · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T12:08:02.370883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:08:02.370883Z digest=sha256:87142295c3dd906f88594711ff723f6deb239d3964caa9adc3bb83dbb150daa4

Observation 7f2c230e-7d58-4dee-82b8-89759e04746b · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.125422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.125422Z digest=sha256:f6903c5a223a6f31f649a766af63bac87aa42bcaaf1d9ec233d71afb1c37fbac