Pith. sign in

Paper Citation Record · LEDGER

BLINK: Multimodal Large Language Models Can See but Not Perceive

As of 14 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 91 inbound Pith citation observations for arXiv:2404.12390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.12390 v4

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:18:15.439163Z

measured 181 of 181 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 91 of 91 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:12:25.253133Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact4
  • verified fuzzy38
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch29

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fcf71cd7-bd5a-48a6-857f-95277b0fa783 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.762219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c5b48efb3172a9d1098cc10366b5d23e10cc9c94fbdcadcb2c48a0db6647762a

Observation 93d55819-5eef-47df-922c-cecf99ea015b · outbound

This paper cites In: AAAI (2019) 10.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: AAAI (2019) 10

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.768026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c80abcf9ae13393b0e94702d2877e723e43a0147104f2ba03adcd8c57f041277

Observation dfe95755-cdf2-487f-b0d0-d6ae0c098bae · outbound

This paper cites Advances in Neural Information Processing Systems35, 23716–23736 (2022) 2, 4, 22.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in Neural Information Processing Systems35, 23716–23736 (2022) 2, 4, 22

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.771138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a6c13b4f21ccd6d7e2c5fc742dc0c0573bc5d5d2a1e778388d1a5ee135713a43

Observation f55ea2a4-f1c8-482a-b894-d35b1898c965 · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE international conference on computer vision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.773873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f1ef45f7ea2ea9e5aece1504c7c91e93341fa937a7cbedff25f451c119385d7b

Observation d50abe5a-8d1e-41da-a66e-4df0232e8cf2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.666166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e3d201b73aaa6c674d9d08ba37ee4e1ae6e75ca89e2f32df920f98bde209d18b

Observation b698a4e8-b64d-4f81-ae13-8a1661b32373 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.776403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6818b64e43b9e4e413516bb9bcc66f7c1f5f1c246569d4e78f4d93f725371a8e

Observation 52168342-21d1-4b73-8c7e-b36acd46bc7c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

BLINK: Multimodal Large Language Models Can See but Not Perceive Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.619179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2bc6ae6a04a364c07145294715f29b22a65092c5ce36509b4817ee2996f8f099

Observation a4b4fa38-fef0-4f95-bd78-9c2c825e419d · outbound

This paper cites In: CVPR (2017) 3, 7.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2017) 3, 7

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.778804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:03818688a88faca82999b46a5cbc6fc13249961954abee2bf79d756558c79320

Observation 5a7b5cd3-a7f7-4fac-84c2-11fe68dee4bc · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.781036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e8d91b82724481f1e3854415725f35ed5aa7de0b646ba326db8c0906886fa951

Observation 7254d039-d974-4a4c-bc93-e12cd269cf20 · outbound

This paper cites ACM Trans.

BLINK: Multimodal Large Language Models Can See but Not Perceive ACM Trans

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.783620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d8eba31051f49dd136c9a9adb39c96a53ce839d5e87996b3f12c887ba8da4fdb

Observation b8bad57a-08cb-4d1f-a0c4-6a097a15bc29 · outbound

This paper cites Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language.

BLINK: Multimodal Large Language Models Can See but Not Perceive Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.599633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6af3596a25b8d33e10af62a8cb965f3d41ec2f7f71df893c6e5e7ab8c89fb999

Observation b72a520a-b593-4a50-a263-629843d8af92 · outbound

This paper cites In: 1993 (4th) International Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: 1993 (4th) International Conference on Computer Vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.786158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c0daefe54488e36d36944f60acef2abdf7e329b687aa27a132ccc18c71b158dc

Observation fed26eb1-3b7d-43be-a90e-39e4331dfbbb · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020) 4, 9, 11, 21, 22.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems33, 1877–1901 (2020) 4, 9, 11, 21, 22

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.788689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:13183caa77a42f4f805092457fce73c25522651b4bbe832f44c3fe7060bfe537

Observation a9fe8895-9705-4417-a3fb-5a19979de761 · outbound

This paper cites ACM Trans.

BLINK: Multimodal Large Language Models Can See but Not Perceive ACM Trans

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.791003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:bc478d423125c615164e6c72f52e873ccfe0252b4c0885c527d983ba7902570f

Observation 2c66a042-f3bd-4b04-9676-7d13b331b5d2 · outbound

This paper cites In: CVPR (2021) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2021) 4

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.793107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:98e152ff3f1ca9e19ab5630ec97a1cd484b8f3dc925e8068e03da2a9ea5061cf

Observation ac9a7f0f-0c40-48eb-83fd-5225b857aa86 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.795278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:0790adda6847f861fc6f74f52ed8c511003254e0ddef706e0dcf95a34ad746d2

Observation 996ca3fc-1ad2-4bbe-8cf0-7c41edf64524 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

BLINK: Multimodal Large Language Models Can See but Not Perceive MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:13:09.112362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9804ff03aadd18abb970dbad339dd8cac9dfa55e54ef167dc14ecca36c297eef

Observation 146fed86-e719-4157-a572-2704ba86e6f1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

BLINK: Multimodal Large Language Models Can See but Not Perceive ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.593758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:bb7535ca3528f35d585121fa2ccdf7527a8af498018ad350b376484ee649febb

Observation 46b2e885-baad-45ce-85c9-8704c2eed1ce · outbound

This paper cites Advances in neural information processing systems29 (2016) 3, 7, 8.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems29 (2016) 3, 7, 8

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.798544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1dff1091e25b4a490fba74bed2d225f5035b69e1624551fba8b39a618104f986

Observation 8cb85d00-298c-49aa-a09e-3e103837ffbe · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.800932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:af84139996693b2848bbb1298456928c29b49a44737f4e8e606ef3911efbe22a

Observation 21afd6bf-0673-4a84-9c75-5b750877a0df · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.803360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d55a5e840bf0bdafe3448ced28a0ce56d7fb682664e392d619ff8cef7975fac7

Observation 998556e5-a872-49a4-9278-fa654d61a263 · outbound

This paper cites https://github.com/open-compass/opencompass (2023) 11.

BLINK: Multimodal Large Language Models Can See but Not Perceive https://github.com/open-compass/opencompass (2023) 11

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.805845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:046db43d9e32f036a6c395f7d2a307e0d945040ad7111cc297cb16f72e35ba07

Observation 0d8d7a4d-c62a-491b-9f6b-a5d314d42318 · outbound

This paper cites com/InternLM/xtuner (2023) 11, 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive com/InternLM/xtuner (2023) 11, 12, 23, 24

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.808247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8005ae138e47eb42c57a1f93fe49d506f4eb2c9fda9a33584a883bf5ec406672

Observation 267c7f17-6c3d-449f-9439-fc0dd8361675 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.810690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f7f6ab0ce5ddd37223ebb30a9178082acb47b0641d8314e693b1d288383002ee

Observation b890ceed-b7cc-4a19-999e-3d6fd1dc3926 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.813663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:4b9fe026c9061ea0958919dee05fca6414c2f5bec2f87d22262e10a0af218c63

Observation d4b6a04b-d8b2-46a4-8867-e734fd207a15 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

BLINK: Multimodal Large Language Models Can See but Not Perceive InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:30:28.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:230f5df7dfeb3924ebd8e2f806d97e3cea002af1a046cd243725c2c6288e26c3

Observation c210f13f-37f6-4e5f-abb1-8a52e32644d6 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.816068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7051dd17410875840efd5d329fd6b29519e39101854021155415fbe1808f6677

Observation 881999c7-fa79-45d0-b917-f348e6369c1d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.509047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c65dcc24cc5c15b9fe0f980d141959d2f3c1c4164571ccf71c5878222def17e6

Observation 5001d30e-24ca-480d-828c-8cacd8c79a23 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.818380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:61dfbb7c9013c3588d5c2232a7fd808f5b560b09afb102c4716ca259e3221d39

Observation 5d08287a-8053-4e97-8d0d-dfe69ce9caae · outbound

This paper cites In: Findings of the Association for Computational Linguistics: ACL 2023.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Findings of the Association for Computational Linguistics: ACL 2023

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-05-15T20:18:15.490957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:de233f52c6f4775567ce3bac82712c723dc83717a5871a569a87728863d5d895

Observation 5a80d92d-f532-42ca-83d1-312f60f4cf4d · outbound

This paper cites 37 Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang.

BLINK: Multimodal Large Language Models Can See but Not Perceive 37 Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang

Reference 31

Resolution
metadata mismatch
doi, observed 2026-05-15T20:18:15.485601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d8bc1743accbe6bbe43d71bc01e2d7357649c9628acf7bf9f0c6f9595cce3f13

Observation 95663a59-111c-4a1c-9e73-ba265c6dcc19 · outbound

This paper cites Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering.

BLINK: Multimodal Large Language Models Can See but Not Perceive Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.582600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d2fbf665a98733d7f27dfdced6a78b7685786a1c6dcd587362b239f8aad82f1d

Observation d7c261a0-478f-475b-b63f-e59ab1eb3003 · outbound

This paper cites In: Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 4

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.821099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:78a306ee81a328aa51eb34e69d8ff16031f6ff74be961cb5f0ce43ad73712d43

Observation 2d3701d7-a80e-4cd3-a109-1f6dd3636a23 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:57246efd0a8117df9f986e13254ffdc0550c275a3cf2859f534e9a9154073c9d

Observation 9cdd2190-f815-4841-9bcf-ced8e62c9fe2 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.825953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:361e1b46fe36edc4b8d4cb59442a7fee6aa3af7ac079452affc04d588e7e4051

Observation 0bfe30f8-de75-4de1-a072-3d0165bfd8df · outbound

This paper cites In: Alvey vision conference.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Alvey vision conference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.828507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a140392f0a73e4863ceaf1e6ffebd37b008760319d3f2e10c3c6b90a60bbd0ed

Observation a030bb59-3043-41cb-89f5-d0b525733949 · outbound

This paper cites Cambridge university press (2003) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Cambridge university press (2003) 2

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.831015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8f0340f9082ecd4f8fdc56b6198a3cd227aaff1b855e10f0ef0e63e0734065f2

Observation 9a16ecd1-667b-4023-b391-602484d52b8b · outbound

This paper cites PromptCap: Prompt-Guided Task-Aware Image Captioning.

BLINK: Multimodal Large Language Models Can See but Not Perceive PromptCap: Prompt-Guided Task-Aware Image Captioning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.672292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7f7e46b6962732b04e562c78acd8f0d8cd3e8adc3822a3d0ebaa6fac5172511c

Observation 532dd131-098b-434e-a28d-6f80c7074153 · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

BLINK: Multimodal Large Language Models Can See but Not Perceive TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.689615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6b764556b8ab7916f4e371872e3c98c01744553bfa05dcc40e3bd12c8a8b1604

Observation 3b0a7b7e-1c97-493d-b1ad-5a7ae2346f39 · outbound

This paper cites Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.496968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f03c9a4290cc92b6a9fdb2bc9276250547e0a0326281b1e8e19cb880beed1f95

Observation 0b5a1051-327b-4ebf-8c93-0762dfdced0b · outbound

This paper cites International journal of computer vision123, 32–73 (2017) 3, 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive International journal of computer vision123, 32–73 (2017) 3, 4

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.833614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6ab09083f21d975680c32e4bd925a9cdb2b81da11514255bc797eea3a7cb37bf

Observation 7e6c7465-7eff-4c68-ae11-60ebfed00596 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.836030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f38467690e2d3da690bc1b0c38029f7d5af8c1d6316c04ea893945db3a81972f

Observation 1d5c8533-1aa0-446e-9863-d49c8d67f1cb · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.530071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:353a4a590930a1a5d3fc4815f0b291679b52669ac006bd311c0dc9c1d7b84eb8

Observation 68ea6d96-69d5-4450-8375-3014c5517409 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

BLINK: Multimodal Large Language Models Can See but Not Perceive SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.545662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d1f42ee08dfed8d878b19b948cb183b40ae8783ce3b15cc6101200d4c6e8fa9b

Observation c1412af2-b079-464d-bc52-db91f758aca3 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.558098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1e1bf4f76e19aecb93d9b1fb68cc58593f942e73a9383461dcfd90e640bc2b58

Observation a6cadab9-e8d2-42c2-9ed3-29cf7110eab7 · outbound

This paper cites In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.839038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:ce784393e19758701ba896ded5cda4a0f9d36a42fc1e615acb8599b8afce9661

Observation a2d66600-7c1c-4a2f-9eaf-5685180dfec8 · outbound

This paper cites Transactions of the Association for Computational Linguistics11, 635–651 (2023) 2, 9, 21.

BLINK: Multimodal Large Language Models Can See but Not Perceive Transactions of the Association for Computational Linguistics11, 635–651 (2023) 2, 9, 21

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.841877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b21b9d14e20b72aaea765c0cc4f6562f6eab3b63569ebd0c1fe82d2225316af7

Observation 2da219b6-9c93-4953-8dcc-6bc7fbcd1f54 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.844279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:fc800ee092848a86e783958bffd888a59bcd832df8a93848ed634e05c37f5bb0

Observation 14f5665c-0157-4308-98c8-275c83c14c88 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.846689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:de3d7b5bcfd924902790441ec2d95c1fff584abf3d752a30ba9395fc1660a355

Observation 509e345e-cd95-4443-b391-bee8bc86f682 · outbound

This paper cites io/blog/2024-01-30-llava-next/ 2, 4, 8, 11, 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive io/blog/2024-01-30-llava-next/ 2, 4, 8, 11, 12, 23, 24

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.849407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:4aabbe513f640c55c0d648ee1303cf18259799c37a55d9d4cf08c6a2a0877828

Observation 1dd3f07d-6473-444b-97a2-667afbffde6c · outbound

This paper cites Advances in neural information processing systems36 (2024) 2, 11.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems36 (2024) 2, 11

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.693679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:ef95fa253667e84f5134763b6e6ee3aa4b229fbcfbcba9487d5f0eaa59923d19

Observation cf854146-2f12-47d1-acae-ff5f845c4577 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.697244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6c851691354b1ad01024c07f2a025f36b4d626cc0f2f14346926f9f7d0eff779

Observation 84113f97-4b89-46be-8fbb-382d41d8eb83 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:591b1a2a1d2fb23a81db82a48b61205106fcbce81a49f376d2633a0ab7889922

Observation 434afe2a-894c-47b0-ab3b-eaacd2eeb993 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.700983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:851a4818ffe9543e6d131a5ceb486a5120501941060b8e1a9ea940a406e4380e

Observation a66033a4-545f-455c-9f27-947c3b512e4d · outbound

This paper cites In: Proceedings of the seventh IEEE international conference on computer vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the seventh IEEE international conference on computer vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.704616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:5947f6606dad97079c3f1c3995dbf01c442c922ef24c806c09ec1f56ecff1040

Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.684545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f588d6ae07afd1ceca96db9b6b430c2dc0426ac825baa8fd1a8179aa37ab5f72

Observation 827638ea-6485-465e-b8cb-f6601dc6fd4e · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

BLINK: Multimodal Large Language Models Can See but Not Perceive MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.678753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:4c177b94636452dd348c4c99d2f2edb67349d50f086b66af4bd363a37f19cad0

Observation cc83e521-21b9-41d8-9c3e-aa999faf7126 · outbound

This paper cites MIT press (2010) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive MIT press (2010) 2

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.709099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c8561988ce45d52ba40f4f3c8b200c03affe52685c2462eee4941028ecfb3e6e

Observation 0eeac0e1-ac71-45ef-9a00-80fac5c18bb4 · outbound

This paper cites Science 194(4262), 283–287 (1976) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Science 194(4262), 283–287 (1976) 2

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.713373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:39c158d6166246e18aa1eaa9ba68b0ad42a7f34240eef1c85b85b405c0af17f9

Observation f49f3f48-d0ac-4389-8315-31b5ccd71c56 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

BLINK: Multimodal Large Language Models Can See but Not Perceive SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.503362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:56df2db2eb83cc46b5cb2f9855cea9aff30a11f655166fab2b0d55fcd54e41e8

Observation bfb6c897-1a2e-4c19-be5e-fd2aee2e9fd9 · outbound

This paper cites Cambridge tiass., HIT479(480), 104 (1969) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Cambridge tiass., HIT479(480), 104 (1969) 2

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.717033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9e33f4127a4090bd5102de6af3598c5104f0de9de693749f1ca504ed869cde82

Observation 1ed26a87-e17c-4207-8559-e3d004ac42b5 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.720719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:efd1c5df9efcd1f1ebe35cf01b45e1d4092f01d8316ab07fa50ad708a34e0df2

Observation f91b33e7-a376-4ba9-ac0e-49845d32d7ce · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

BLINK: Multimodal Large Language Models Can See but Not Perceive SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.522936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:ec40438fa1d83ae099b4b2e2ed4a9b70145d0e6fb6c708bb9d723e94d16f0165

Observation 7ff96374-5266-4210-a5a0-706d7dcf6277 · outbound

This paper cites In: International conference on machine learning.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: International conference on machine learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.724244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:48df6ead8acc9d7e4a661a5c60cc7ee87dd4da562a48f975645e186d904189cd

Observation 410795db-2cbf-4622-a1a2-34c8ce65d5ff · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

BLINK: Multimodal Large Language Models Can See but Not Perceive LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.537757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a1fcffa8733861718efb15c4e9f25e82ca74d28458adbeaf70c8e153dead7f0a

Observation 33c1b918-6657-4f1a-8e24-48fcad2e8f2f · outbound

This paper cites In: European Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: European Conference on Computer Vision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.728266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6f03037d299e46bf5b5876cd6467775e3fe62c9cd1bc5d7174ecca3d31d93b01

Observation dcb66970-d321-4245-98b2-a09e0c80b1e4 · outbound

This paper cites What does CLIP know about a red circle? Visual prompt engineering for VLMs.

BLINK: Multimodal Large Language Models Can See but Not Perceive What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.551709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b9c76064a9a297fa6ec71b1ecd05c695299207f7001afea3d2e0636942b0cdc3

Observation 5eb3dfad-7ab7-465a-ac49-c3ca3ec61d1f · outbound

This paper cites CVPR (2021) 14.

BLINK: Multimodal Large Language Models Can See but Not Perceive CVPR (2021) 14

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.731452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:3844196bc7396785c41fb23aa6c5de9a19f150dad2432b84ece0ab68ec0fd6c3

Observation b8d17ed1-6576-4d50-ad15-e05a10d25bd1 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

BLINK: Multimodal Large Language Models Can See but Not Perceive EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.564212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b7a7434aab00c441fc346132549a2c531a08aba0520e7fa6425d7ee2a6360ad1

Observation b4aef595-17b8-4569-bcd5-562b37ff4ffe · outbound

This paper cites Emergent Correspondence from Image Diffusion.

BLINK: Multimodal Large Language Models Can See but Not Perceive Emergent Correspondence from Image Diffusion

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.570881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7394ae49c7b30cbd45b31dcfa64c9d2978307a4de7a559cfce8aa4dddc7aa128

Observation c17df47c-94db-4fdc-82ff-2839578afde7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Gemini: A Family of Highly Capable Multimodal Models

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.576419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d0460d599b34a0afa89379f57f82c2184797102737dab5727ffc49972b5e11ae

Observation a7075c6b-8591-482f-83b5-e0ec6c073911 · outbound

This paper cites https://github.com/InternLM/InternLM (2023) 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive https://github.com/InternLM/InternLM (2023) 12, 23, 24

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.734581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2282ced85729cc04c5fcab65466f9649c61c8207bba2bf8f2a1e6925d5349ff6

Observation f4cae9cd-991d-4d5b-901a-c596728f644a · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.737730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b348c8033032fcfaf1cfba7feda5882a397d4d429450c9b4eff28e75b51c60c6

Observation 2a3aae24-6ba9-4f12-96a2-661cbac603c4 · outbound

This paper cites IEEE Transactions on pattern analysis and machine intelligence24(9), 1226–1238 (2002) 2 20 Fu et al.

BLINK: Multimodal Large Language Models Can See but Not Perceive IEEE Transactions on pattern analysis and machine intelligence24(9), 1226–1238 (2002) 2 20 Fu et al

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.740549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b209d716ea42bab0a3099073e28a9c537249f303f53700fc72eaa981fc9aca00

Observation c06fee5a-2d65-4152-a868-6783f875cfff · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.588230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8786b38cfac9a2cdc22240103785a6bc1df9df0d9620522e6d5b07f46ab7592e

Observation 56a1c8a3-f519-442d-a6bb-eb1cd0f04b5a · outbound

This paper cites In: Pro- ceedings of IEEE Conference on Computer Vision and Pattern Recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Pro- ceedings of IEEE Conference on Computer Vision and Pattern Recognition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.743471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1da68fd45f6185c4b434459e3156c105f03f8d4cfb6868bac0db7275a9e1ae41

Observation 96aa267c-ca6c-47f5-8249-15ff9908c10e · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.745998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:37b21bfa0bcce574b588e1019b7f40e78f067f509796de2aba67b2ed8d1591e8

Observation 21050174-c0c5-400c-b14a-73891c3f4531 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.606264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:0b1efcfbdbda6d89763f37fe8c7ac1484c8edf86871b8726a604793728ae779b

Observation 8a04abfe-161d-49f3-96c3-2728f4fad943 · outbound

This paper cites DIRE for Diffusion-Generated Image Detection.

BLINK: Multimodal Large Language Models Can See but Not Perceive DIRE for Diffusion-Generated Image Detection

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.612995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b5fbba504d8bf8a9153ab67441965b4319bdc184a9c4ff33dbd5347818bc7509

Observation fc6a34db-744a-4677-a780-d0a488b920e3 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.748558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:eb1af0b4d64021f97c87d84aa2b94e3a32c3270d64fbb82ceaad5a4ce41214b8

Observation 3be8907e-eadd-4f79-ab9f-c308f9a3827c · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

BLINK: Multimodal Large Language Models Can See but Not Perceive List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.625753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b76dab30a06b619a4160bbae49c21c2e1517b76b6c723faf48ec93a9b2a2204c

Observation f3fee5fc-46c8-4076-b40f-91e0f1a6489a · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

BLINK: Multimodal Large Language Models Can See but Not Perceive Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.632677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1eb3026fa5fa6a2b90fc9300ab260cb5428a2adab5cb89b1757e92f7391b7e23

Observation 7fced311-c5dd-45cc-9e03-b3a3fb4f3f82 · outbound

This paper cites In: CVPR (2024) 14.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2024) 14

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.751646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1e32c15a04d42369de47652366af68bc4935b2cdd80b59c264252639683e3a62

Observation e8d55ce6-fd0d-41eb-80e4-30352dbd407a · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.754381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a6bdd924cbe27e3ad2cda97958658544b473caa5549f4ab266e08e415296d48f

Observation 84fa579c-1abd-48d2-b560-b96da5f3a348 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

BLINK: Multimodal Large Language Models Can See but Not Perceive The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:07.065720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:08b15d024bcbda81508ed744abc643bcd14a50d72325f759dcf3395ea8c6db55

Observation 8daddfca-2328-4e99-867c-dfe394bb0e23 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

BLINK: Multimodal Large Language Models Can See but Not Perceive MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.655809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1004e257b6a4401fd1966269fdaa66537ea9af92749a4b8310d212923028e9ad

Observation 566ba0cc-e3c3-4f53-9bf6-6bb74f6bc283 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

BLINK: Multimodal Large Language Models Can See but Not Perceive MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.661112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:4a1822d2b2b97d2eba3e9d64bd1967c073e7faa7fc15c6e6221d939f5daddd19

Observation d8059fec-5cab-4dab-b8b9-1830dc7cbd2b · outbound

This paper cites Advances in Neural Information Processing Systems35, 27469–27483 (2022) 3, 9.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in Neural Information Processing Systems35, 27469–27483 (2022) 3, 9

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.757131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c707dc3c77298d363d948c516b2a66f444ee22c3921f7bf869a9495f63ecd2c6

Observation fa975452-116c-4d8e-860a-9083bc9de309 · outbound

This paper cites In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019) 4

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.759785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a440050914ecbf784d8aa5211138ec5cd37e3eb65ed7466e86c319a6b862f78b

Observation f777c19d-b5ac-4ae3-8fec-a230e670e8f9 · outbound

This paper cites You are an AI assistant who will help me to match an answer with several options of a single-choice question.

BLINK: Multimodal Large Language Models Can See but Not Perceive You are an AI assistant who will help me to match an answer with several options of a single-choice question

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.765110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e7dc1aa2665df6d3132b546268a7837663c173c9f6762728ce7f07a00aced3c9

Pith citing papers

Observation 449cc6a2-7af6-44e2-8b8c-60d35ee072a1 · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:28a4754b8076591b67420c6c7eded7b670130c5a248bfea18fe40f67e4dd5ba7

Observation db3428af-e77a-4f8a-995c-184368d595ba · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:09:30.442631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:19553f8a117bef3a585134a0d5d02504860fccf7d751f3aee3b082d0e77a654a

Observation f2da540d-4c8c-4aa0-8482-62bf96050759 · inbound

Depth Anything V2 cites this paper.

Depth Anything V2 BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T14:56:33.945280Z digest=sha256:d6b89c9b3542ae88bed7075f2d9a24a8cb2d5c5f4720b8c6d717291a0bed0784

Observation 35fbd366-aacc-42ef-8c5b-3de03abd4298 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:05:03.712019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:44e653e4a1d971c834a83d083c69fb53f397bc90fa18dfa8e3fb034cc11facd1

Observation 2490bbe3-8e77-4a4e-ae5a-95e5937f0c81 · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:c055a4f14ac23e1f566cc0c7dde2d7f4ecc8c5cfa134c99122113549815d00c2

Observation d256730f-a2e3-4983-90d8-0043a8ba6f24 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:b5aacef204692917f1c5bacb8999cb610c3f84e3c617f66c6c5757c89f8b47c8

Observation a910a12c-9f03-4125-b335-711b47e51e93 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:59:32.699685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:740ea0d1f78b8112b1dca65ba255fdaed5f8839c8fe06b10607603f478ca5ed8

Observation 02f44911-d49a-470d-b4a6-7515bb7319a4 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:135a6230c5b5e5d6d56e7e703bc3916e5d4302e2bf8afd938a26e0c78fa306b0

Observation 28a73a8d-5055-4bab-a60d-3fca6d663680 · inbound

Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs cites this paper.

Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:25.253133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:25.253133Z digest=sha256:e1d5ad4aefbd0bfda9a89ae1456951ea54e5b2d9f137c75f3df5714423bad915

Observation deb76522-dca0-477e-99ba-391bc10795e0 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.795142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.795142Z digest=sha256:35012604063fb3fccdc9e17508bd6d37323d07e1390d6bfdaf48cf6c545939dc

Observation d144e7b3-8901-49ed-b6ce-5bb0da669ad6 · inbound

SketchAgent: Language-Driven Sequential Sketch Generation cites this paper.

SketchAgent: Language-Driven Sequential Sketch Generation BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:54:48.212530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:54:48.212530Z digest=sha256:5c85a859246b403904c508dacc5d97a587b2f0e81e1b33624697bdc3a81ff42e

Observation e878cf0e-87c1-4f84-935a-a03c82514ac9 · inbound

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features cites this paper.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.565408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.565408Z digest=sha256:22be0bd8cf37196d7a22a563670e3fe99f0744abe74e1e85383025b06c323e2a

Observation 87569496-d69f-4c50-87a4-35694e820fe0 · inbound

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models cites this paper.

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:55.601768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:55.601768Z digest=sha256:518580551286c5d3a504c25c40b0813052527a4aeff682941a6ed965365743d3

Observation 7921fe95-972e-483f-b52d-a833ce1a2b35 · inbound

MageBench: Bridging Large Multimodal Models to Agents cites this paper.

MageBench: Bridging Large Multimodal Models to Agents BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.059138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.059138Z digest=sha256:531b2e16abfd1a38601c0cb506965ae95bd7fcec070410f6ba8d18835bacc22e

Observation b97e81f7-0bde-406c-b183-6b7e09787a8c · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f882b9864460fac82ffd41834aad9917145bc71d4423e809d8ac9982d409f196

Observation 372766fd-7225-41f3-82ff-0c4daa529ec4 · inbound

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents cites this paper.

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:49.515784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:24:49.515784Z digest=sha256:51f13fe4a5f11970565a441f523a8aef893df493be420eda838cfc1b55950f5c

Observation 71b020b1-6cf7-446e-9f1e-99d3d252a245 · inbound

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model cites this paper.

Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:03:22.525123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:03:22.525123Z digest=sha256:555d00f86e06d932cbaf23671a4cb71b92b71788571cd9bd96e84c7c5150fbae

Observation 248a84da-d421-42b5-b827-57a691d103b1 · inbound

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics cites this paper.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.559022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.559022Z digest=sha256:8afb4fae3b3e9c9e144da9081f5cfa4dc9362084a8ef8ea0f2ac77426c2604fa

Observation b98382ba-30bc-42b8-915d-d86fbd198ffc · inbound

Do large language vision models understand 3D shapes? cites this paper.

Do large language vision models understand 3D shapes? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.729389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.729389Z digest=sha256:2ea61690b4c3639408fcf877d35f3c9dbd0ff6af0326819b8d2600f93b4d05ef

Observation f743acbe-c706-4107-a2f2-b81bf92ef702 · inbound

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models cites this paper.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.065731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.065731Z digest=sha256:27e952f349f72b86ef8119cb3cf7bf08874ef7b1da29f39b66966cd4487e94f8

Observation d9b79575-df07-451c-ab2a-4b28228f0063 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 258

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.207957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:68d8c80e4fa78658c8dd06220e591e2492b81ae3c374c23cbe57873ec0ecac93

Observation 2ca317ef-07e8-4b03-8e4d-b73c4f167674 · inbound

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks cites this paper.

MMFactory: A Universal Solution Search Engine for Vision-Language Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:08:17.696188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:08:17.696188Z digest=sha256:8a40fae348878a060fa7e9fc3722d15101c3ba3d24d6333a5dfd4425efdc6448

Observation fd405925-dcba-4488-a58c-c3c3c3f77105 · inbound

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding cites this paper.

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:02.983278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:02.983278Z digest=sha256:54634f117b3c1cfdae7eaa2b3c45172049b8df1876f3c5531f4dace3eec8b005

Observation b4ec5e97-8b24-4e94-838b-2075b0fcb267 · inbound

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction cites this paper.

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:17:54.034110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:17:54.034110Z digest=sha256:8b9ffd3a2aaecb6061b270408986007287874611ab86524a20dde09995ff4dbe

Observation 7d442750-fa52-491e-a38f-d1a8e0582475 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.886875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.886875Z digest=sha256:2ca2462ad3272ae7f03977438dd984ccf34dad5a63edfad80789f26edf4ea734

Observation cb91d072-e5e7-429c-b9d4-f256aecc44af · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:16db5a3a87de0c15ba189f270e8576c0ea317068cd1e5f9949deecde1dff6027

Observation 2637342f-824d-4ee1-8cec-4d9dbe651456 · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:22:12.223017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:5dc836ca74e1d40b6763575a6edc0c4234f9e73053ecabf1f448eef1f5854b7d

Observation 2a242639-a304-471e-8c5f-5b5ad78f44e1 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7094bb60dd070ccfa518a631c6047a71bdbd31f4b273e12b99b8834ac5981eb7

Observation 3024c293-8945-492a-9bb2-c065a21ddb9c · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.174278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:b88beacca7e4110e6461ae691b4cb2606847816548d4d66a74231c123aebf739

Observation 0ecb08d9-d824-4b1f-abdb-58bfcb12287a · inbound

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations cites this paper.

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:23.368760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:23.368760Z digest=sha256:eee67edbde50edea59fd002390a00f1eaa21d6cf297ed9db098ab9a82cf34ecd

Observation 2d189f2b-a1f3-4e59-a6ec-335fb4838dc6 · inbound

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study cites this paper.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.901934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.901934Z digest=sha256:44b81b587c84619ce90b4a2e671e0ada0cc6861361c40b2801e94ce822d785e2

Observation 395a5a37-0c82-48f3-a1a5-3d79e44675ee · inbound

Synthetic Visual Genome cites this paper.

Synthetic Visual Genome BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.197159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.197159Z digest=sha256:ad3b6e8a83287e6285da01d54e48e24b7fa83cb6cfd400f42afd74511bda731c

Observation fb1d3bfd-f903-4eff-8b91-cd78b2677d51 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.888488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.888488Z digest=sha256:2962ea3cbfd57c983be173483acad2c8400340faf19c879707a1f69d1f3afe52

Observation 6840f7bc-b91c-4ab6-8362-746e39cd4238 · inbound

MANBench: Is Your Multimodal Model Smarter than Human? cites this paper.

MANBench: Is Your Multimodal Model Smarter than Human? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:33.576525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:33.576525Z digest=sha256:780a33d5763520f7922ce99fb5470a606b584fea6d0dbaec4084f9ccfe44704c

Observation 0c53e5fd-1426-421a-b279-e7f525546c99 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.659375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.659375Z digest=sha256:0eb0582126500580612d53324aba1d0f5e1142de20304276d8cd70cabf1020d5

Observation 7095c77d-5117-459b-ac10-881acf5f2746 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:65b1a92aa4c97a40f4ac5ca7ef133842b9668122bb1647c7d1e423d46c29307a

Observation 84230edf-4209-4308-9141-436bcd19977b · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.011869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:c6f81a011341ef351d58ca1e3502e25b0640b9dbb4d3fb3e7b707c9c31630392

Observation e1a4f98a-1838-4266-8391-2cfa32afdb38 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:28.730737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:28.730737Z digest=sha256:5611081eb9584bcb249aab28eed80a8ec397c3596f76192ab2c968fd1fca1ed3

Observation ab5b5abd-9b09-4f0f-bccc-629838e11b29 · inbound

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor cites this paper.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.095890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.095890Z digest=sha256:aa321d80a803faf0797d15b78ea6000ae4c1df0afa09189e0dedad1ca52db793

Observation 4e757f10-5bb8-4a03-bdcd-822a6fc97c6c · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.096889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.096889Z digest=sha256:6f2e844488117d45db618f4dadff1b0af4f9b8763f4c195a4138354d0886c62d

Observation d517cdae-32fd-4521-a495-ce6109d7355a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:ac4185a4fc8f52fca90715a8db5a5dd0c0efb18017ea866552be7795074e7bd6

Observation b3588781-8ee6-4566-963c-9e74225587a6 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.631783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.631783Z digest=sha256:14e25b3ba274b6072c87d2fc98d6f5f44f30c7e37a404386fdc55019b0ae0b48

Observation e48ae2b1-7886-4639-8cb2-d9c322917be1 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.924950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.924950Z digest=sha256:6e4e05a898e9b5fd75dd040b4a6b4dde9e29103ec66cc8979a177149f6633455

Observation 3e8d8889-933c-49ba-96c7-2b7824e177ec · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:08.613463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:08.613463Z digest=sha256:123c62d2b05fec1b659e11e19056439c3958305d3ed2c4b76ae1327aa5e556bb

Observation 80da89b1-98fb-43e9-8400-1d37f0472ea6 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:49.532620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:49.532620Z digest=sha256:a9fb65e9d44275c6a69dc65be0ea468a3487bece07e54735e070ec309666bb0b

Observation f0d3fd3c-a303-42f7-b729-f26b80345f4b · inbound

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models cites this paper.

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:45:24.419680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T22:43:16.761970Z digest=sha256:182590c4e37899d49ceeedc4f249afbf6772b9436d1f9114072aaa86008bde26

Observation c8e4af68-f378-4907-8e8f-4673229054ca · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:11.109093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:11.109093Z digest=sha256:b84d62c3f0720100c7f85a16a6ab67d77444871a6d4f6d49e9430dfcb510aa91

Observation ef67d5fd-d15f-4fda-b4b0-c5ab2d75330a · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:37:59.940145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:d1ea3332002bcc9d78d83d20c1e40c82f149d3a0986619e17eb076c21d850d3e

Observation ac112fd3-56c9-4fba-b4f9-128456022928 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:ad61899620a6300a977b366cc69dbcfb8bc9e4ea9113c159eccc5fe6199d3339

Observation fcab69be-4460-4126-8fe5-a1838387d2b4 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:3aa6b553301ebc5d826badf28d0004f6b6a5bbbbca4aaa3a031fdabac44276e8

Observation cb7da1c8-5250-4d3a-88a7-3395600d6638 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:6b012d9eaa789d858eb7c83d75d5ee08e3d0ddd121377214c0352164fe61309a

Observation 5c7952db-71b9-4556-b907-5c59a7f6b83a · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:53b98fbb2f411c0ea9ac7b0e5b2b1d272d66e1e09d7679460674a14b2d0a7da3

Observation ed1cadea-6207-4b5c-8aa3-40ecee404002 · inbound

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs cites this paper.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:3fa458f69aeb02f01294eff56f065fb607663ac8f4462aa0a22929d5e4a2479f

Observation ac8a1ce3-7d20-450c-90f2-fe7fab2e5efc · inbound

RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction cites this paper.

RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T15:29:17.567557Z digest=sha256:d9a300c4296e3ccfd510b4e5c4152854378d6371aaac5a4375117257dd57afe0

Observation d57cdf37-8b01-47fa-bdff-def804182c4c · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:64c191c888838cca302acdd97ead3f047096abf2ca4d56b6cc016cadb093993b

Observation 5a4ce0a0-ebc7-45a1-9d80-6b86ae9c5d7e · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.604914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:5c89ec49ca7cf54c0ed89e58e176117a6e0390cb2d3aa8ac2738224763cc2389

Observation 653ca010-be3a-4fd0-8c2c-ccb7b67d391b · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:159b1852f02c47926e6a58b8a60e03a891c73a980325c28d53ea0348274d96e0

Observation 5de67ce0-c44b-4ec6-8abc-43c140d464b1 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:4128f0ee8718ebef6bcc3bd7736eba8991cd77b39a0a5a3dd18ce87c914f068a

Observation 0680746e-0946-49b5-982f-d46c5df6dcb0 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:13:46.799961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:c0e52bb8f56b3633a3e1499ac1e45fda3122c1d2427e9e7c905ef8f79db359e5

Observation 32f15706-6d81-4527-a732-c653269f5fe0 · inbound

What's Holding Back Latent Visual Reasoning? cites this paper.

What's Holding Back Latent Visual Reasoning? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:03:15.479718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:59:14.134917Z digest=sha256:4c528703905fef469530760f1c576dc7a30bea2a73ba79dcd79ab4c9e428da14

Observation 45972e6c-4f3a-49ce-ad25-59e5e90a968d · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:33:14.185570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:5f32c5f8aa5e60ef50047f11a535c2935f7d4941e35dd298a3a7f8cf172a4642

Observation 41f4ffe9-b6b8-4ff8-819f-eb9b5000c5c0 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.371150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f94841be218df01458368c6ec73531a3a07271a7a33ad7beb1d47f7af893a583

Observation 7ab9d2bd-131f-4b57-bd9f-16c354b70c5e · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.208797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:31630120c5198071e6efc8af1e917bf21a6cffa7eba3de2f2c8d4486d3915ae9

Observation c2a322ea-d23f-496d-8dcc-18c43c501de5 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:47.176193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:1410fcca74c78b4c370b04c9912e787ce2ae6b7fb77ce969a9951d887ec72793

Observation 7ad6f110-2913-4142-bc38-ab555372727b · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:39:53.709641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:a21d86668ed138baa301010cdeecb1464b5d0e628a78456800fa7e4e4e1f232a

Observation b280e90a-091a-4308-a785-3ce2c8c263d9 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:04:58.341972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:b71beac1f7c34b7e45fd8fab91106c78aab95eb0a9cc56d21b5d8898d9cc4692

Observation 870f1a3d-4d80-4c90-9713-611b8411423a · inbound

PInVerify: An Offline Embodied Benchmark for Active Instance Verification cites this paper.

PInVerify: An Offline Embodied Benchmark for Active Instance Verification BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:33:13.552711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:30:59.306097Z digest=sha256:77d779a83fe0872ca3ed901370e73bed428d9bc7705b02d1779293e5e2613d56

Observation 508d172c-b047-4f6a-9dad-28b2652407c4 · inbound

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? cites this paper.

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T17:37:14.727108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T21:56:25.350827Z digest=sha256:a4cb980d55319c9c8dbecc83b9ac9ae336caf96a645f8bf80ba2c9ba8a9b7eed

Observation b148aae3-f0e4-4bae-aa81-fc50c8314a12 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:aa36025c0ae7a602fffe722dd5d09d60b8f7d564ca3e5ae46292a2fa027eaf4b

Observation 6351c0a6-442f-44bf-bf28-b82b86efef77 · inbound

Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers cites this paper.

Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:50:49.101338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:08554a0da65a791ed89a1359df54fd8657cf69418685d655ce86dd5188df1e99

Observation 10e1f930-285b-421e-babd-a0e88f519f27 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 126

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:08:55.668706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:a2add0705e67c59b9044a5b282d3e8a70e68045cd603852bef8c28b14f685557

Observation fdc46151-68b7-4d87-a1d1-123208f6c79c · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:49:02.264842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:a62097319e1bd0a50ac9214ea86e5ded996cc096444093fd4714df502eb4f4f0

Observation 13ac3131-27ae-4e06-b3cf-71889a166f29 · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.569062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T17:56:07.864580Z digest=sha256:ca25e279dd9e9460a083d20698c73a07b3b598b31823e2f9278ed8b18f346ddc

Observation d41550c2-7965-45d5-9bed-ea092705becc · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:39.058350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T10:22:08.535453Z digest=sha256:a21144aa9136f3208e81a475e23a793dce36410cede07022943d48ae3986bfa9

Observation a7bf10db-3965-469f-9212-8b0376aa8f1e · inbound

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception cites this paper.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.076334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:9633163afc3e4c501d381e8300ec996379485136449d2a1f3ccd10f73c94c836

Observation 9b6bcb5c-80c8-492c-b2eb-08e16619da25 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.287801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:caf28eaca6a1f460850f9b89d5675d4a3e35ddcd9dbde8b1d5dc6e1833e20f29

Observation 65ca5b1c-22f4-4574-a1e6-af195f0604d6 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.200965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:634dbde6d513870d1f1155d2f35d5bbbfcce95505f1868f843ef931bcfeb96b3

Observation f0acdfa3-2518-438c-9d35-bf02dd304185 · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:10:02.344679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:054c1b302b8d42b3fbd08ca2375f44c4c0b8a2c79c93308d6b1d21c4b0b7c6d0

Observation f79a77b4-41dc-4519-8def-0f5e1e5622ac · inbound

C3-Bench: A Context-Aware Change Captioning Benchmark cites this paper.

C3-Bench: A Context-Aware Change Captioning Benchmark BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:50:10.197607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T21:02:52.529391Z digest=sha256:4ff5bd8692158573ba96a848034cfa2337ff422731cb9001c46689849d41c2ae

Observation 5ae8d44e-6985-4875-b221-1261d8c47ec4 · inbound

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues cites this paper.

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:19:50.970106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T05:18:01.931929Z digest=sha256:fd343dcd93eccf708ae0fdd5004f858d65906a12e8cff18f9d3ccc45b7f5d8d1

Observation 6265e2a3-e2ac-48c7-afb9-a4c84b86d857 · inbound

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine cites this paper.

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:26:58.578997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T13:19:29.775959Z digest=sha256:2c90fe11053809670930a6c0974e1ceea7a17406ad03076d1565d021dd14920c

Observation ae5cf0f9-eb9a-47b2-9552-76e22787a979 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:7246ee06832f46233af1331a9aecf5c26fe0adfe81dca2d8460b245e94699dd4

Observation 21908eb4-c7c1-4ead-a2f4-43266c6052d7 · inbound

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning cites this paper.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T07:59:56.441098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T07:59:56.441098Z digest=sha256:789b194dbd73f51a83994b469e7a59eb6730181b388c997924057f2021b2c2d3

Observation 3e4cd782-5ba3-4906-b57b-b3866b366861 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:03.439176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:03.439176Z digest=sha256:b6f91c0b3e4d46038c995fe0b3c5c0bfe095f7c61847bede1af3c8418b0fb957

Observation 968911c8-70be-4eee-8090-d09f98d0d120 · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.400943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.400943Z digest=sha256:d495c2c4e1ca4854eca724b1a49b50e7bfb48ba0ff76c9e6fc47ab26e255ae0d

Observation fb24cba7-7b31-4c49-be5a-c6f5f1953e27 · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.095760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.095760Z digest=sha256:2914acb8421467791e5fc1c32a98acaf0cd17fb1ec2f4f3a25f71b22908c458d

Observation 933e3f93-953b-4aa7-b85e-f48e1221ae1d · inbound

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications cites this paper.

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T12:20:51.527995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:20:51.527995Z digest=sha256:50559e34838ff04c3b37fc0b1056c3453b80bac6425d8feda275ad3c281fc8c8

Observation e91dee9a-f065-41c2-b5cc-1b61574b181e · inbound

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents cites this paper.

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.939147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:55.939147Z digest=sha256:e018049f330cb622b388f6c3cfcd71255e5923e320133ae120c357f05e51e834

Observation 2f4fecf4-dedb-4797-aab3-cfc5ecd2ec38 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.296769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.296769Z digest=sha256:0104fc924cf11d478a30d711b8877ab058592c7c3a2eb3556c71e7bed21e3e45

Observation 195d31a5-11c3-492e-af18-b14ebf648ec6 · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.008737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.008737Z digest=sha256:1ad0f413e447f493fd1cbd823b97282cda090621573ecf173824610c87ed5b6a

Observation e048748a-49a0-49a5-85bb-66e7701d78db · inbound

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus cites this paper.

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:56:01.753162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:56:01.753162Z digest=sha256:664c71e8ef96a9198b116f73ab53bc58d0c4012630767991e7c9a791fa30850a