Pith. sign in

Paper Citation Record · LEDGER

On the robustness of multimodal language model towards distractions

As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2502.09818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09818 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:25:58.596793Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T01:00:27.572834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:11:17.919527Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved13
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d37f96c1-4c41-4b8c-919a-0a402e321173 · outbound

This paper cites Vqa: Visual question answering.

On the robustness of multimodal language model towards distractions Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.350990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.384098Z digest=sha256:ec19b95766ec755cbedbd6b7dc621fa8fc11a996369d423cdf47ea314da0a4b7

Observation c5f01393-06f9-4c4a-8a5f-d819ccc290c6 · outbound

This paper cites Choquette- Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt.

On the robustness of multimodal language model towards distractions Choquette- Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.331941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.391159Z digest=sha256:ba6f5930292a40423ae5e70120d8d881d3bdb40a1b73f6a2b83238f842870c9e

Observation b75241f2-2493-4802-9562-c21873821f9d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

On the robustness of multimodal language model towards distractions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.396754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.396754Z digest=sha256:555dd646f39d4c84c1d6a09c215448334af270f79b8a76f47f0d8d31accbe7f6

Observation dffe2634-4015-4213-af49-7a35f2e7c4c4 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

On the robustness of multimodal language model towards distractions HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.403744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.403744Z digest=sha256:d32eee818f6b3b9a06bfc58244f4a1b68bfa7e3171631967461db4f4ef4c1392

Observation 1e68e2c4-f4ec-469a-9f33-44081099c8c6 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

On the robustness of multimodal language model towards distractions Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.410363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.410363Z digest=sha256:cfd2d1dec7fec77c6e15b862183726ba633ba7a9bfec533117ea0098cc105c89

Observation 0cede9a2-b671-4098-82ae-3f68818a795a · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

On the robustness of multimodal language model towards distractions How Robust is Google's Bard to Adversarial Image Attacks?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.416124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.416124Z digest=sha256:9797214103fed3129c5e1f1b01e334199e8b8c0f8bdf01e628766ce73ee60423

Observation 0fb5e3ff-151b-4fce-97c8-024e0321cc11 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

On the robustness of multimodal language model towards distractions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.422649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.422649Z digest=sha256:c4dfa167e086d78c138737e1619458cf0b2e07b7da56f886e027bb6c67596d48

Observation 32aa141f-b92e-476b-97fd-c2eb6c6fdb1d · outbound

This paper cites Phi3v-finetuning, 2023.

On the robustness of multimodal language model towards distractions Phi3v-finetuning, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.300025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.429141Z digest=sha256:737c565efb3bb90769813380d199dacb6bb46f4fa04fa8607b73093535063a36

Observation 19b751bb-1b39-44f4-8568-0f24ab94da08 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

On the robustness of multimodal language model towards distractions HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.434448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.434448Z digest=sha256:8b49ab7e495a8d53563fc78c3e260cba76d6927056833d7a1949242aecc235ca

Observation a8e5ac55-2ec6-4eed-b9b3-ba4cff84de8e · outbound

This paper cites Cogvlm2: Visual language models for image and video understanding, 2024.

On the robustness of multimodal language model towards distractions Cogvlm2: Visual language models for image and video understanding, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.281755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.440591Z digest=sha256:d1d1ee51f73f7fd87c7904d294b9acaf62f66ad516587f8099f8eb5c8df1dd05

Observation 16072865-e658-44f3-ad2a-1c9ca4557ae9 · outbound

This paper cites Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages.

On the robustness of multimodal language model towards distractions Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.446123Z digest=sha256:d8f9fdae9aa50f0b68605458d4d136f788094746d83e79160bed13d7f2e35ede

Observation bbc59deb-19da-45ee-865c-468d971daa7f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

On the robustness of multimodal language model towards distractions Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.262535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.451812Z digest=sha256:30e41ac5cd7047dc38399428c60ea64127b94a03a343138608fc8cc8f8d23b4a

Observation 4a7358af-0757-499b-bb12-dfbab7ebdd5b · outbound

This paper cites Open- clip, 2021.

On the robustness of multimodal language model towards distractions Open- clip, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.458116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.458116Z digest=sha256:e41b7fe2fd1f1b793a69f6090c938aa051edde087b089aff31a136cb7bc6ebf5

Observation a3ac1863-e170-4d35-8487-75b3139b806d · outbound

This paper cites Adversarial examples for evaluating math word problem solvers.

On the robustness of multimodal language model towards distractions Adversarial examples for evaluating math word problem solvers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.230192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.464365Z digest=sha256:dc4d34f139e7e4940e78dbbf3ff64779a40d93217ea8b794d1c8d1c634d305a6

Observation 69848a13-51bb-4a10-a771-a3e89c5162d9 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

On the robustness of multimodal language model towards distractions LISA: Reasoning Segmentation via Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.469758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.469758Z digest=sha256:bfbfe1b0092b6095078c188f4afb8bf60b1c00771aac219c3e700c8558ed0208

Observation b327ad1f-67ed-4043-8270-8c0e367b89bb · outbound

This paper cites Evaluating object hallucination in large vision-language models.

On the robustness of multimodal language model towards distractions Evaluating object hallucination in large vision-language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.476188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.476188Z digest=sha256:9a4ba79ad393ce07cdcb3e3195015aab5b9abf9a915043d7bf6010ce2b80b7ed

Observation 010bf48e-4ce5-44fe-a456-0a2d80a7dd25 · outbound

This paper cites Visual instruction tuning.

On the robustness of multimodal language model towards distractions Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.200014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.481824Z digest=sha256:fde60c6701bd97e20bd62730c47dbaeb059d4925369a673dbafc21ad2f6a1335

Observation 2394962e-aa9d-4613-86c3-49bfbe126f3e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

On the robustness of multimodal language model towards distractions MMBench: Is Your Multi-modal Model an All-around Player?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.486806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.486806Z digest=sha256:ba5c5c8302429941943f63a13db438b15360097b51b96c5c84e48ab7675e806d

Observation 38ffbaa5-1e8d-4562-aeaa-3b5fac7cd4cd · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

On the robustness of multimodal language model towards distractions Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.177301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.493078Z digest=sha256:c934cb8079a246dc998a01b6989b3f2f0e3b131815da2d7a6a94acdb65371a12

Observation cd7b49c2-a6e2-43bb-9766-1022caeef332 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts, 2024.

On the robustness of multimodal language model towards distractions Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.157478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.498077Z digest=sha256:01a141bd0eaf42efb3448c42addfeff404f45770162c6de205df413d2a4ff119

Observation 06812f5e-2460-458d-828b-3204904a853d · outbound

This paper cites Understanding zero-shot adversarial robust- ness for large-scale models.

On the robustness of multimodal language model towards distractions Understanding zero-shot adversarial robust- ness for large-scale models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.136750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.503335Z digest=sha256:7f3492d9596fbaac5377f306a6e62352de7a133c5ce4d0d27222c2c3a502ea9c

Observation 3b5c2d96-6aef-4ffc-a56e-fd306bc2d6c2 · outbound

This paper cites Gpt-4v(ision) system card.

On the robustness of multimodal language model towards distractions Gpt-4v(ision) system card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.508674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.508674Z digest=sha256:9e1bebe4534f481bda19def0201b0a9e9657358076ff05ba07f34630d5265d67

Observation cc9d91b8-b302-4cf8-aba5-62fb66f13ecf · outbound

This paper cites Gpt-3.5 turbo.

On the robustness of multimodal language model towards distractions Gpt-3.5 turbo

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.102629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.514303Z digest=sha256:4a8bf65af3a9cacfb2e7eaf7c651852aec6455af903eb98355c21f609c877511

Observation 403be801-f927-4acd-b089-efe41a4ba526 · outbound

This paper cites Are nlp models really able to solve simple math word problems? In NAACL-HLT, 2021.

On the robustness of multimodal language model towards distractions Are nlp models really able to solve simple math word problems? In NAACL-HLT, 2021

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.081273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.519582Z digest=sha256:c3ef22c5582118f7c6ded1b6559d43cc4be702f2fe9332f184ec84bfd607113d

Observation b32d299b-6207-43d8-96f0-8fead312d049 · outbound

This paper cites Homepage.

On the robustness of multimodal language model towards distractions Homepage

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.061065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.525855Z digest=sha256:8375dc6b30ea22ecc7c8ad2eb315e71c4d285f60acfe051c65a0a41a5995ade9

Observation 3439da96-30d7-4e0b-b9f1-8b278ed8d560 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

On the robustness of multimodal language model towards distractions Visual adversarial examples jailbreak aligned large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.008222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.539490Z digest=sha256:51948c8feade0d93b071437920ea241c4a806a79f21e895b433f9fe587ff306b

Observation f6299b28-b9e6-4049-91a2-2ea929020757 · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer.

On the robustness of multimodal language model towards distractions Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.986614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.545301Z digest=sha256:d7923a9374602a114af03a335e4685290b5649afc49e2c8c5a9e3010730c9b4f

Observation a5ce974a-f315-43a0-a3e6-d0e34f3137da · outbound

This paper cites On the adversarial robustness of multi-modal foundation models, 2023.

On the robustness of multimodal language model towards distractions On the adversarial robustness of multi-modal foundation models, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.967524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.550995Z digest=sha256:f718b97cc0532eda1d63903aac9306ee3212ce2a3ce710e659789ab587312f4d

Observation f81d975d-aed2-4629-8323-c292deccde64 · outbound

This paper cites Robust clip: Unsupervised ad- versarial fine-tuning of vision embeddings for robust large vision-language models.

On the robustness of multimodal language model towards distractions Robust clip: Unsupervised ad- versarial fine-tuning of vision embeddings for robust large vision-language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.948610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.556555Z digest=sha256:3eaaf640a0f7130aa90487a056d7574ce57d0e3c7b383222e862477019865ea4

Observation 87b13a6c-1f33-46f3-b4c0-dc1caf2df972 · outbound

This paper cites Large language models can be easily distracted by irrelevant context, 2023.

On the robustness of multimodal language model towards distractions Large language models can be easily distracted by irrelevant context, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.928046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.562026Z digest=sha256:ef6f18b56af2959628cf72a6c6acbe1f747b08fd5aa7e244da3422072e8ba071

Observation 4a016203-6941-4686-b224-43f566834c56 · outbound

This paper cites Towards vqa models that can read.

On the robustness of multimodal language model towards distractions Towards vqa models that can read

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.908306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.568109Z digest=sha256:4ace34fb5953b2c3143a7d58aa7591f6d1c86ac6c02a4a48c2fa931af8da7915

Observation a31188e5-e4eb-4d69-bb43-306c9759323f · outbound

This paper cites Unsplash api.

On the robustness of multimodal language model towards distractions Unsplash api

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.889018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.573558Z digest=sha256:7e3b6a0cc9873befbc8da17729145eaa203ea68226abe3cc7db86680d635fb21

Observation 88d427a4-54cc-4ea6-92a4-42de4e6d29f4 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

On the robustness of multimodal language model towards distractions Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.870325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.580371Z digest=sha256:1c239098323e3df67362ae58d400ccb2c943eebc431beec561ceb17b440f6ac8

Observation 511b0599-cc63-4731-94ce-6bdcecf8d2ce · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

On the robustness of multimodal language model towards distractions MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.585610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.585610Z digest=sha256:aea6d30f8c9a7c94957b0b5d9f59df0ba734039d1e8c9a1e3d3167f011cd7fe1

Observation c0365983-3f1b-44a0-9ff2-b663f50f370d · outbound

This paper cites On evaluating ad- versarial robustness of large vision-language models, 2023.

On the robustness of multimodal language model towards distractions On evaluating ad- versarial robustness of large vision-language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.851745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.591386Z digest=sha256:e523e0f5a403dd34c82d529b9e7bd0720797bff9ce7c90022fce5f110fe77235

Observation 02225cbf-15c6-4914-b4b6-8fdb52c6be58 · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

On the robustness of multimodal language model towards distractions Analyzing and mitigating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.827053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.596793Z digest=sha256:4062e63c6ecf34a83a205e41e25fcb67cae5d373a402bede3c391cb736ca9fdb

Observation f4d9cd83-44e4-40e9-bdf3-23e12f2c026c · outbound

This paper cites an unresolved cited work.

On the robustness of multimodal language model towards distractions Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T20:25:59.041438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T20:25:58.533546Z digest=sha256:622bdeae749f1bd454a67c00be5b223cf96c111577aaf970261bff4923f549a1

Pith citing papers

Observation 1d1532fb-9e7f-4bea-bd60-4c445ac19072 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability On the robustness of multimodal language model towards distractions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:27.572834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:27.572834Z digest=sha256:69abe123b5c3a44324dcfaf246b01fc83478b25d5a0205c72585ffc020b92196

Observation a79769f4-ea2d-41a4-b959-8c00f88e3b89 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models On the robustness of multimodal language model towards distractions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:17.922848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:8db244239fabfe86611ea13b8f940bccdf7edc67d9dbcbdfc2b5c819e2862d1a