Pith. sign in

Paper Citation Record · LEDGER

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2505.10541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10541 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:13:38.669736Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:53:15.920563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:55:15.290281Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e65b50f-6a00-4907-aeca-e697b5d9b09d · outbound

This paper cites write newline.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.127029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.127029Z digest=sha256:b5f9af0d47f471f2efed35a5782583479776af26bef00847d1829599e900dd59

Observation 9e63ff9a-cf2b-4e2f-91ef-9f8f5f1bdbf2 · outbound

This paper cites Claude 3 haiku: our fastest model yet.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Claude 3 haiku: our fastest model yet

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.973460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:36.149728Z digest=sha256:7e06d86fea79bfe808866d226071b7411b52d4d17df90a97ba4e90d93652b810

Observation c0766cfa-f7b5-4a79-b752-4cc7ea13fd01 · outbound

This paper cites Y., Bhiwandiwalla, A., Tseng, S.-Y., Olson, M.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Y., Bhiwandiwalla, A., Tseng, S.-Y., Olson, M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.255024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.255024Z digest=sha256:46c2926fb2c027ac1bb0179bafa0cfec54b4e8121e79aea7b76ad00b6a517506

Observation 3d037445-bc1f-4ad0-95b6-eb6646212a53 · outbound

This paper cites F., G \'o mez, L., and Karatzas, D.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis F., G \'o mez, L., and Karatzas, D

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.277213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.277213Z digest=sha256:c237c9bc15211c66cf3b1022012635818aaf5832438fb35dd7dc10c2dbb3235d

Observation dcfe01a4-b34e-4c45-8b1f-319cbbd9dea6 · outbound

This paper cites H., Vora, S., Liong, V.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis H., Vora, S., Liong, V

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.303024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.303024Z digest=sha256:094d04671633f3b13d1a60c96ec37031acd404e3f951299e9280ec29dc063dbc

Observation 439fd75e-fe80-46f7-8b96-def93d14fa28 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.403553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.403553Z digest=sha256:2ece333ce063b58bf308c9a15d34ba173d1a23292b4f063974d015b15abb2e23

Observation e6d834d3-9ba7-4424-8e7b-6145d80a4a3a · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Dress: Instructing large vision-language models to align and interact with humans via natural language feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.824259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:36.508976Z digest=sha256:de58675b91d45ff28e2617b5e9b469c215637ea54cf4e1c5c0d24c6b0b2269a4

Observation 95434e8c-a688-45e2-ba2d-08060f68889f · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.523424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.523424Z digest=sha256:d26361ca172cfa4afc7aeec2735c063ad2b352248fa2e85539ce195256d0cbfa

Observation 037abd1d-c667-4fd2-9824-58e20bdac74c · outbound

This paper cites H., Yu, F., Wan, X., and Wang, B.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis H., Yu, F., Wan, X., and Wang, B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.751673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:36.535151Z digest=sha256:2027922c56eafe2b694c3595c15a1da07a079fbb3055c2f988f0c4070914f156

Observation b6de1c68-313a-4035-a964-4ac65c790d2c · outbound

This paper cites A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.556970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.556970Z digest=sha256:9e87acc22a36a7576f8faeccd6362b3c6f38a121d588427d89f6837cb0841e38

Observation 13945699-3fac-49b0-8aaa-ca25d471b533 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.663456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.663456Z digest=sha256:56bcadee9d90daf8f04fc051734a9197bd1fb8e7a77171dafeacf1f1672283b7

Observation ddefc0fa-4660-4846-9098-4672dd91a1b8 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.802693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.802693Z digest=sha256:322ce0890789c5a7e856a59d809dddcf252a3e696ffc04a19e68a8f79decef84

Observation 623384f0-0c02-49c4-889a-0dac13b88e5e · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.853423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.853423Z digest=sha256:4f0eaa81cd3ebd5a2c07bee6a7f6fd123ca33840ce30e7da5c904044bb2eda62

Observation 8059e091-9fd4-45f9-a8cb-c97e28baca0f · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.873714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.873714Z digest=sha256:2d6f74f0275973e0704fe7d2054fbccf9596569da5d621b4f3febabb0a136158

Observation e05b8189-4235-45af-b516-9166d7fce7c3 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis CogVLM2: Visual Language Models for Image and Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.894005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.894005Z digest=sha256:8aaff5f165c25edde1ba2308a4b3b3dd9118841887a7bd16dd36c14c838cc61f

Observation 4162649e-6b6a-4df6-a251-80f50038dad4 · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:36.988347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:36.988347Z digest=sha256:a31d4abc047130b855b097d537df9dbeec579d243ca41920cafcd2c519a5b6d4

Observation 4418e467-ff88-4fec-9f4a-7231c02a0469 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.013746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.013746Z digest=sha256:3c784fe68feb6db61846a5b11ac06f1107728970f979a87bf5907fe87f4568c2

Observation b42bf34f-0c98-41eb-9012-0ec11b9c3f63 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.590512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.036793Z digest=sha256:9f002c3a994eb052a8562dce640b2cb0cac6e3a2785af4544b5898ba92f0ab07

Observation e5144ad0-7b10-472d-8470-5b5cb7e2a691 · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.394832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.097040Z digest=sha256:96da2506f43f3ebf15aac7aaf2ac14da1ba32ec41441b25cd84dd7a4828d3f0b

Observation f4188452-f5df-4d9d-939a-372a2929b963 · outbound

This paper cites The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.161934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.161934Z digest=sha256:343bfb85580fc7f1d6daabed962b53c5e3f942983088984ccc4d69cce6d03e92

Observation 3f50f6ce-c668-4786-bf7a-dfe77276435a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.184098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.184098Z digest=sha256:7bc214cedce38f42c9f3c68bb61e4eada1ea8ae661b6f1d5197ecbbf86b240d1

Observation 5415a960-083b-48b7-8ea6-a74e4d64db68 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.251991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.251991Z digest=sha256:1b52cb8b6c702fe16d13d3c62390f955e98652fadc96dc82c7ca482df0619ad5

Observation e174b456-dba6-4805-9865-ac6a25267585 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.333684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.333684Z digest=sha256:1574813871bae6d99d666a3445484fbebd8b2937654f19db3643f9f64406a4fc

Observation 15a68c54-70e4-428d-beab-f690768d86f4 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Evaluating object hallucination in large vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.386927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.386927Z digest=sha256:f4293ae12dff78a2001e2a2792a0cdf38ea553fdae5076eb70cf619907dc80ba

Observation 4d3d3625-3f5b-4bdf-84d0-70b160d07d33 · outbound

This paper cites X., Tian, P., Yin, C.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis X., Tian, P., Yin, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.418823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.418823Z digest=sha256:4e9ec59bcb615fb3bed1b1ca02e42e5cdb0bb867056d4d4de5af62b19e4caccd

Observation 4e7cd78a-e8f8-4978-b5c4-17ebfe5758b7 · outbound

This paper cites an unresolved cited work.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.498864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.498864Z digest=sha256:8df5485ba2db4d1d932ad45a951aa6ee84e664589b9e7b98215404a19b26cbb8

Observation a838b0bf-022e-4f1e-a280-44925a1c28ed · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.587409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.587409Z digest=sha256:c89057fdb7a7384d4b81d11ba616e62cf2f61915ff5557ad50c37d3b3d91fa5f

Observation 56f88c9f-0cc8-4441-92be-cc0b506b312b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.182714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.661696Z digest=sha256:14a61ff1372105f99007ffa4d1cad32a2b82f61d5792b0b0045b01b220b8d731

Observation 48c02c35-1732-4f56-938d-96686c26fb47 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Docvqa: A dataset for vqa on document images

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:42.031472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.689152Z digest=sha256:ce24eae72b8adb7a0f6ea78781a0d86cd7af6fc4190f0fcb21d432645d4e8cc2

Observation ec068e0e-6881-40ee-a553-be31aee913a8 · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.699635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.699635Z digest=sha256:827355c28fcf02995a8f227441eb60e686f652da892c614d2b56825a331137fb

Observation 39528761-a855-4a0a-b23f-2cfde322fefa · outbound

This paper cites Gpt-4o system card, 2024.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Gpt-4o system card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.881845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.708140Z digest=sha256:8a134923835b6e189e8c059cb843a4b5d9e61ed963f067b8e7b37403a83be3fa

Observation 64bdf4d0-5e00-49c9-a0c5-9e3fbce4c7da · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Perception test: A diagnostic benchmark for multimodal video models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.835948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.717062Z digest=sha256:dc1940ca2a52aa061ebf285bd071ef8c669d5c04c63f99c4d006f97fc7679e37

Observation 9e4ba0fa-2975-4356-a414-f8cd75e51486 · outbound

This paper cites U., Hezel, N., and Jung, K.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis U., Hezel, N., and Jung, K

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.793389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.788172Z digest=sha256:45e5234b3dc33ef5ac1ea4a760ac967d1510921752db92d059eba29f0ce2f4e7

Observation 814fce51-8e6c-4722-8abc-abe0fbc87f9d · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.804909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.804909Z digest=sha256:959f5453addbff257cc36f8810b5ab0f4f660096f044d30650879adc28563fd3

Observation 778cad13-ab66-4424-84f3-0101bbc5fa4e · outbound

This paper cites Aligning large multimodal models with factually augmented RLHF.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Aligning large multimodal models with factually augmented RLHF

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.819200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.819200Z digest=sha256:b3a0305499aee972fc0f689ff725c3730a1fc7a4d4cb369e9387692483ce8d48

Observation 1cd8ac92-a3bd-4306-9da4-355dff4b3e83 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Slidevqa: A dataset for document visual question answering on multiple images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.692449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:37.907391Z digest=sha256:330d15bacc56c2e1f884eed14c1e4c97acb6dbcadc6b76da71a435d5cd5fcdd3

Observation e3284111-3a0a-4561-b56b-95fd3705422a · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.958941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.958941Z digest=sha256:10260bc49d4fabdc1deb8318daacf4f61fefe2b5c86f5d07075503e6ca4d996c

Observation 09fd7815-4961-4df3-a3dd-66e6ed87c57e · outbound

This paper cites Attention is all you need.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Attention is all you need

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.051212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.051212Z digest=sha256:80fd10173d52f92ffe616221e9be79fa61f41974809fdc975ad3c185029196d8

Observation 740686f6-9c71-45a0-bb11-87c78d0b8654 · outbound

This paper cites Analyzing the Structure of Attention in a Transformer Language Model.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Analyzing the Structure of Attention in a Transformer Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.071899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.071899Z digest=sha256:f1ed23087b78b411bbf4f4ff807c48c193fac06dfdf7f6160490765eb7821f1b

Observation 6808bb5d-b3d9-4bc9-a85d-98138846febf · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.088801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.088801Z digest=sha256:12ba2f4ae5c6b49752df2c127baa6d841797462ce870831f7ee5f7fd9ebf72e0

Observation c105421a-acd4-4a3d-a25b-ea67ef4f8f3f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.101688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.101688Z digest=sha256:5ebf6988ab0799e5c6ccf3722c345a33e7b2e1d63210e9316e8be8af4ced61d0

Observation 69ca77c9-7cfb-49ff-b405-fd6937c4d8f9 · outbound

This paper cites Needle in a multimodal haystack.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Needle in a multimodal haystack

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.482125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:38.124510Z digest=sha256:4f8b7731232cfc7d754cd6ca9b4c7510ee86365dcad919b7623d2bb42bf8573c

Observation 66a3c391-a451-4848-b51d-fb017986a3a7 · outbound

This paper cites H., Le, Q.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis H., Le, Q

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:13:41.428374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:13:38.212476Z digest=sha256:25b9f6b1ccccc226c87759ec566af3b216980c17072056b8a25e8ebb78bff23b

Observation 0fbb95f1-2556-4a89-a1a0-5148abc81d3f · outbound

This paper cites MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.283855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.283855Z digest=sha256:f59015cf6ccf637d69307e63e6540b76036ffe65c3fdb52130904a25f4d748fb

Observation 66753a6f-05d7-4180-9997-5c9abc452796 · outbound

This paper cites Qwen2 Technical Report.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Qwen2 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.332827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.332827Z digest=sha256:259512929594957f6ccc61e57e4d4139e71308d5b263a469063e6b6001f3f95f

Observation 64d88026-b91b-4d5e-947d-0caff1f3308d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.380837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.380837Z digest=sha256:59f304b212fb8e367f2df24bc226af8a9123356d0a2ec78040d96e89586f25f1

Observation ef60c97a-fed6-4657-8c78-8620ef2887a4 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.425106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.425106Z digest=sha256:fbc13cda247eef9e9a0410efcc3e375361de897e4d980679f7dd5b69154940dd

Observation 6574eee5-562a-4fbb-b353-15b4f62a7736 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.518278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.518278Z digest=sha256:7aad4ac7dedec1b06e3e2470eb5d4402549a338b4d7655e5c00691c93d8ad67d

Observation 0aaaecd9-e590-4823-902f-79a487135dfb · outbound

This paper cites Sigmoid loss for language image pre-training.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Sigmoid loss for language image pre-training

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.618356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.618356Z digest=sha256:f2fe8337d361f1133a8bcd2a617c79dd5390bf577041aaa9f5aab2bab5d6e434

Observation b1ab4e11-b315-46b1-9795-1d51107345e7 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:38.669736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:38.669736Z digest=sha256:7d899a53aaddd1bd44182f6a43f6b1f93cf4b2cfaf445b4b6a6efcd46b22f866

Pith citing papers

Observation 5cfd14b8-5ae3-4165-9053-3f6bcb53be59 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.293168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:f8db554b0978f759f827d370ca69dd2e246b4f42d069016cdad08dda0c4b5acc