Pith. sign in

Paper Citation Record · LEDGER

Hidden in plain sight: VLMs overlook their visual representations

As of 15 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2506.08008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08008 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:31.975509Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:44:45.178051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:29.186840Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1878cd5-cea2-4902-b674-fde8032d49ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Hidden in plain sight: VLMs overlook their visual representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.827278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.827278Z digest=sha256:34c13a60f450e3d0f32452e842ba878a506b79e784afedc6358f4a75072f640e

Observation 3da569a7-e35e-43b7-831e-8a15fb0eeb92 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Hidden in plain sight: VLMs overlook their visual representations Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.832536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.832536Z digest=sha256:433ea1733fcb6e8f0d3f49297d063eccf8ad6b9147a597e0f6978c325667a678

Observation 4f5f4e8e-52ec-447f-a42e-25e53966848f · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.836549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.836549Z digest=sha256:f0a7ec5ab301ac18242f54ad52019b7b7b741f23dbfb025243482b883bae60aa

Observation d29bfddc-1992-4023-be6f-1dcf41c3317c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Hidden in plain sight: VLMs overlook their visual representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.840203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.840203Z digest=sha256:dd3901543bcc77afbacc0a6cd95299840c1c8ce1150be33b538312c457a203f9

Observation faf576ad-2880-4a83-927c-45829b2203b0 · outbound

This paper cites Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors.

Hidden in plain sight: VLMs overlook their visual representations Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.658742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.843707Z digest=sha256:830e0b37b4a321992cbc85a465fa77f8f01075aa53bc2f9a3c0e0f7406f3a5cd

Observation 580c1030-0ba6-4645-b8a8-c33cea9c1479 · outbound

This paper cites Probing the 3D Awareness of Visual Foundation Models.

Hidden in plain sight: VLMs overlook their visual representations Probing the 3D Awareness of Visual Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.847196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.847196Z digest=sha256:d10c9f094d3d6d5e3035a66bfc5d07d11aa73160aa5b781fcd428885fd11009a

Observation e26744c1-08d4-4e66-a479-43e0be18273b · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Hidden in plain sight: VLMs overlook their visual representations PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.851720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.851720Z digest=sha256:56a90d942fdcdf74471ecc1b8485ce9c3c9b49ee6b4867c068fef326c01f3303

Observation 7cd8f14e-8f94-4bec-a68b-2d43b66e6b0c · outbound

This paper cites Evaluating Multiview Object Consistency in Humans and Image Models.

Hidden in plain sight: VLMs overlook their visual representations Evaluating Multiview Object Consistency in Humans and Image Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.855332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.855332Z digest=sha256:99ee316ae06780f5d135599ebf3f998f49938040ed3ae71a5c6df230dad1eb79

Observation c760c7e7-9e4f-4753-a33c-653eb22cef32 · outbound

This paper cites Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild.

Hidden in plain sight: VLMs overlook their visual representations Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.858960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.858960Z digest=sha256:e2798e5ba2e43c554792452156bca9794e20c527e48887398d9d753b447e6735

Observation bde34bae-c1ab-4412-b835-e865f41fe766 · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Hidden in plain sight: VLMs overlook their visual representations ShapeNet: An Information-Rich 3D Model Repository

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.862693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.862693Z digest=sha256:f001535b11bbd9afdac126713679b741f6aac203cf271dbd864173596afb258a

Observation 6bf1eaae-e437-4f15-9902-44650ed9bd71 · outbound

This paper cites An Empirical Study of Training Self-Supervised Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.866030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.866030Z digest=sha256:64c4975f49fc3da53866598c73132a17cc8c81f71236c069d70681eee5dd4d20

Observation 5dab55f0-ed59-4e5a-8f40-64a1bd16929b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Hidden in plain sight: VLMs overlook their visual representations Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.648354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.870323Z digest=sha256:44e6e96c53fcf42dfc3c042a41546c17474f5bf4130c7490847cc52e79f5e6c9

Observation f89ad20a-b7b2-4269-8be2-b30b28854f2a · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Hidden in plain sight: VLMs overlook their visual representations Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.873622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.873622Z digest=sha256:c1939d2218073be81d3431e833adffd5a1757177b281efbd8d39725a9abbf4a2

Observation b9705a3f-eb6c-4689-b927-f444ef4c5b14 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Hidden in plain sight: VLMs overlook their visual representations Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.877003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.877003Z digest=sha256:c0dafc145e696c6130bdcdb3157927062e45ef9eebaf91446f509949b65fd2c3

Observation 6324917d-df62-4403-96b1-d25a6accdb8f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hidden in plain sight: VLMs overlook their visual representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.880937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.880937Z digest=sha256:5116debc43929a603bb073204a2638b709582d71c62df8989e793f63146835f2

Observation 37462020-3f3a-46e7-8157-55b8d09095c0 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations MouSi: Poly-Visual-Expert Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.884425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.884425Z digest=sha256:ff80ddb990a145dbcb905d695805312516b91956332a2fdb609bc7da2fc2e111

Observation fb1d3bfd-f903-4eff-8b91-cd78b2677d51 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Hidden in plain sight: VLMs overlook their visual representations BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.888488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.888488Z digest=sha256:2962ea3cbfd57c983be173483acad2c8400340faf19c879707a1f69d1f3afe52

Observation 535000f8-1b2d-4c6a-9863-30297af8d5d9 · outbound

This paper cites A Neural Algorithm of Artistic Style.

Hidden in plain sight: VLMs overlook their visual representations A Neural Algorithm of Artistic Style

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.892084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.892084Z digest=sha256:1b32f275a3628223aef5dd588309a99aafc43354b7edfbac240291304d9fc0fd

Observation 28b990cb-d837-4715-8c96-4b3604a8777f · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Hidden in plain sight: VLMs overlook their visual representations Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.895470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.895470Z digest=sha256:da03842b9dbb73711097c1b8b491cdf444d29b83c878224ff22c90df5e50cb51

Observation 9200f337-18a2-4632-9222-10fe1cd965fd · outbound

This paper cites The functional correspondence problem.

Hidden in plain sight: VLMs overlook their visual representations The functional correspondence problem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.632004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.899045Z digest=sha256:ff9dc3352e3e595410f90c4e8c7a928cefd523d035a492b29536b63088937592

Observation b3675a38-0f23-4adb-a3f8-4399078d05e8 · outbound

This paper cites What matters when building vision-language models?.

Hidden in plain sight: VLMs overlook their visual representations What matters when building vision-language models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.902130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.902130Z digest=sha256:35c5c88b68fdfda1d40ddabb72428c14d91c0c8ad106aa25016101bdcd0facea

Observation aa320425-5567-4e1b-a1b2-e289f5b01d27 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Hidden in plain sight: VLMs overlook their visual representations The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.905842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.905842Z digest=sha256:96df94e53a8a77bf1a6c2115da88874ebeb1e96a018683f474e61a0a5b2e530d

Observation 5150e7b4-13cd-457a-9a7f-9b7b2c84bded · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Hidden in plain sight: VLMs overlook their visual representations Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.909053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.909053Z digest=sha256:c15ceb4a69371fd371cda1fe5b75b5a5d888699b68fb2e640da4e6d520347a72

Observation 361ab3f8-5750-485e-9c03-72e1fa7965f0 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Hidden in plain sight: VLMs overlook their visual representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.912439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.912439Z digest=sha256:a72c8d73964c5925f6b2409f0f6c6a739b9c706225645fcee4173ec5b03efa80

Observation e94772e6-f240-42d1-a7de-8f870d376c8f · outbound

This paper cites Transformer-based neural texture synthesis and style transfer.

Hidden in plain sight: VLMs overlook their visual representations Transformer-based neural texture synthesis and style transfer

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:25:32.344244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.915449Z digest=sha256:f3510049b6b9967bfeff4466991c1190ea2d5a10e8571433c811f75d5068b11d

Observation b9114a3f-163d-495a-a67c-fd0091df70d4 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

Hidden in plain sight: VLMs overlook their visual representations SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.919369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.919369Z digest=sha256:17f20252afc019ec30c9e89158e09078c72f9912f342e93317cbccb77065e391

Observation 44a2ae01-b915-4b82-bab3-8e94fa98192a · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Hidden in plain sight: VLMs overlook their visual representations Indoor segmentation and support inference from rgbd images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.922941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.922941Z digest=sha256:e8c1bfe8d28005bf31012300b4e9fae259763eccc308349b059299c5b49562cf

Observation e001612a-9f8c-484f-94b2-ab6402190f12 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Hidden in plain sight: VLMs overlook their visual representations DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.926174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.926174Z digest=sha256:89ba426ae53f9bcb713d8d6f04595c598c716bc348c34ec270e4a755d411b86e

Observation e89a713b-4577-4ba0-9b57-9d3534bdff1b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Hidden in plain sight: VLMs overlook their visual representations Learning Transferable Visual Models From Natural Language Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.930340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.930340Z digest=sha256:54a347f4d86743b5c8b3f097f5eea32b0373a4af1601e3c9454ad0cc3287e3da

Observation a61c7834-0bfd-4419-8d62-6bfe83bf8bad · outbound

This paper cites Vision Transformers for Dense Prediction.

Hidden in plain sight: VLMs overlook their visual representations Vision Transformers for Dense Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.933811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.933811Z digest=sha256:39aa4bca417d147623ee3349d6fd97244f69e46cc10c71834c5a5d85847fe3a5

Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.936899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.936899Z digest=sha256:269d97d2cd2ef7940cf7b1216890317c2b4f7513b3a6979d07ddea3c1daa6e56

Observation 96bd96aa-b993-4134-9747-392bb262feb0 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.940685Z digest=sha256:70c9f7c494095411a3ca186c6fe2bab4443726d603536d7f8df6d4b5d908d049

Observation 944a45ef-63b3-4ef9-b4f7-15b93ad1ce04 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Hidden in plain sight: VLMs overlook their visual representations Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.944023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.944023Z digest=sha256:9d3c93707037be9c9def128273381cb7b691a33b721517393a09a5cd54a7804e

Observation 6972fb65-294c-42bc-b23c-9462b1bfbb11 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hidden in plain sight: VLMs overlook their visual representations Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.947615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.947615Z digest=sha256:a5f83337815439c76b46dc11086e000543e5ccec8c7d57900231107c8f3f0c16

Observation 237f0622-530e-4866-bb3b-e867106bf4de · outbound

This paper cites Disn: Deep implicit surface network for high-quality single-view 3d reconstruction.

Hidden in plain sight: VLMs overlook their visual representations Disn: Deep implicit surface network for high-quality single-view 3d reconstruction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.604495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.950731Z digest=sha256:16d1873ea35475272a4d794cd9a4f3ee02d58c2ba5dbcf1e54e303d3b79474b1

Observation 8b48f5f0-46a2-4ef5-aaaf-6c5027b37b2d · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Hidden in plain sight: VLMs overlook their visual representations MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.954005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.954005Z digest=sha256:807e2f7132d6eeb92e3a01b203c7bbb28e5803b8267d720b84896462ee9a4cb7

Observation aec4e2d1-989f-45d6-b10f-c0998f888da2 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Hidden in plain sight: VLMs overlook their visual representations Sigmoid Loss for Language Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.957341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.957341Z digest=sha256:f32962f6bba9e4f694ba8ccbfbac22909f6132f27487f0dc946955d91ef9ca47

Observation f66242d9-6b7d-4b1b-bf6c-4d27d675e400 · outbound

This paper cites A General Protocol to Probe Large Vision Models for 3D Physical Understanding.

Hidden in plain sight: VLMs overlook their visual representations A General Protocol to Probe Large Vision Models for 3D Physical Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.960678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.960678Z digest=sha256:46a8ec157079440429f2d82b8e4434c2c6be65b730b60c555a6f13944c9fe311

Observation 097b52e8-2d6e-47ad-9065-505b1eeaa5d6 · outbound

This paper cites write newline.

Hidden in plain sight: VLMs overlook their visual representations write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.964053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.964053Z digest=sha256:736eee5bf45e5ac4cdfd38482465a122ba418bd770dc4aa34d823db3a5d403f1

Observation 5aeb6766-572d-484d-9d53-34305de2ff79 · outbound

This paper cites @esa (Ref.

Hidden in plain sight: VLMs overlook their visual representations @esa (Ref

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.967916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.967916Z digest=sha256:71e73430747d707f805722917f1b9cc5286d78c6da5820c5e855f3f5ba638c25

Observation 8edfdc6e-430b-43eb-b4c3-34f7e275f8d4 · outbound

This paper cites an unresolved cited work.

Hidden in plain sight: VLMs overlook their visual representations Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.971733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.971733Z digest=sha256:319987e072b0f9021a940a09da1abf3b8154152717c4e997b7657ab646c826c6

Observation f879e721-808e-4138-b4c2-9b7742b2229a · outbound

This paper cites A, B, C, D.

Hidden in plain sight: VLMs overlook their visual representations A, B, C, D

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.975509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.975509Z digest=sha256:6c5d66e25623384749116787fa84b8f47834a906deaa952bfe03a612506e0ad7

Pith citing papers

Observation db4f2ae1-7184-47fd-8352-468f01ff091c · inbound

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters cites this paper.

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters Hidden in plain sight: VLMs overlook their visual representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:23.812923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:02:23.812923Z digest=sha256:0580db9965a58c94ec3b75a5824948fd078239fbeeee3f317e447ea8dc19a872

Observation 3a901a8f-b164-4932-a899-d8b48b76f59b · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:24:21.380035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:9ad709bfca7c4013954685c2325494a5f5f1be6dda50fd3929e747961e58f2a7

Observation 906c870a-fb4f-47a1-bc78-34f6579d5ef1 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.582363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:d9a8d18be044b803f0e8738e606e5643d081db25a554abad74047345f06fd03a

Observation 0ece36ae-f997-4c5f-aced-1b408dbee82a · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:42.981300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:42.981300Z digest=sha256:65cc7ef85d6733b1780b3461479918f54b546794d554d880161b6d989c13ff75

Observation e02c4924-6ba0-4eec-9be2-0cddcca0ab73 · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Hidden in plain sight: VLMs overlook their visual representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:07.008117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:110596527aee25c0d65c1469ed27f2d9811bfb57fb919ce099fdb2503c87d96e

Observation 843cc588-bcef-45b4-ad02-e53dfa76a3d0 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:976ae7c564c005905dcc4770eb20e1d6de3cd698bced7d22092ae33de4074c79

Observation 4825b433-659f-4773-8bee-187497140f64 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.843334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:860e9a2817d3981d925b02eaa775efc8c7113e7877a44e69bd6e83aa4a5b5870

Observation e41c6df5-c7aa-4fd7-9275-eec59f7b92e8 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training Hidden in plain sight: VLMs overlook their visual representations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:0053e76e1427619ede5eeb4dec2cd6a3185fb9c5f8bcad4c8061e38686bf1269

Observation deee1951-098a-4113-85f5-857b9badc27d · inbound

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification cites this paper.

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification Hidden in plain sight: VLMs overlook their visual representations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:00.893146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:11:13.523021Z digest=sha256:963cd773efd04844b1913be92a2ad9cedf46b3829126956d21821847aea69675

Observation 934e2904-d7e1-4c89-978d-3a8e55ac61cf · inbound

Do Vision Language Models Need to Process Image Tokens? cites this paper.

Do Vision Language Models Need to Process Image Tokens? Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.692738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:26:37.371488Z digest=sha256:56c5fc08a98e9e0afaa062eecde9ee48c9475a44e11c6ad121d30d7d72b95c98

Observation 2c8aec45-5da4-4685-9d43-3c26b0e35eca · inbound

Boosting Visual Instruction Tuning with Self-Supervised Guidance cites this paper.

Boosting Visual Instruction Tuning with Self-Supervised Guidance Hidden in plain sight: VLMs overlook their visual representations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:50:58.634642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:27:52.208827Z digest=sha256:ae213ca16b334bf05ee41c9012d16c9ef44446ba05ad51d210c7facd6dbe6768

Observation df2215f2-cab2-497c-97a8-f12aad1d624d · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:35:26.811566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:502ab22a7c86f48beea540f28bcef3f99e4f43394247f594323a411adf5a7a2f

Observation 575432cf-dadc-494c-807f-1a53dc2b3433 · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Hidden in plain sight: VLMs overlook their visual representations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.524703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:696cf14c68bdcfee368fedd032a8c41833a3a9c0e9637644be0438e2d3f06e5a

Observation 9dc668a8-f105-4526-8b89-61ae92877cf9 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:51.063386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T19:14:03.809826Z digest=sha256:7de3c1a0de6ce1c229b6a96819be5560f8e457e50b1c1f5e2499d83584f90abd

Observation dc64514a-24f0-456a-b219-b2dbf321a2ba · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.103332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T22:07:35.021986Z digest=sha256:b14c8f87b62f7d67c101fd4c5d295a73bf275218e3d770b580ba25bf17f04f7b

Observation b957b92a-4bcf-4ae6-8c5e-30ce662dcb60 · inbound

Diagnosing Visual Ignorance in Vision-Language Models cites this paper.

Diagnosing Visual Ignorance in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.193506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:51:25.269986Z digest=sha256:cf06194e9a6de3adb75e39319466609a1196993929d7a46b63cf9af31735a946

Observation 43114add-9cbc-41f2-99ff-81fde306c970 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM Hidden in plain sight: VLMs overlook their visual representations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.190447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:2aaea7d64230c78ac132d1e3702c2b650fc7e51f5489ebf97ba0bd0f0dde94db

Observation 4ab637ef-2e2c-4309-8082-436d893e5348 · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:27.325903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:27.325903Z digest=sha256:2a4c0a544586ca756536430730487bb9236683d6450e1ee2a1be19399aebb29f

Observation 3fd8021e-4afd-4b44-89b7-8baeb11f729c · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Hidden in plain sight: VLMs overlook their visual representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:45.178051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:45.178051Z digest=sha256:fa6ec2289782e8218302e28ef3adec35ea83cb4163011ce1ff77d20348ed7da5