Pith. sign in

Paper Citation Record · LEDGER

Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2307.10490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.10490 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:38:15.346908Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8bcc565e-2fd1-4a78-95d7-86c2094b28be · inbound

Whispers in the Machine: Confidentiality in Agentic Systems cites this paper.

Whispers in the Machine: Confidentiality in Agentic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:03:53.924075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T03:59:03.972043Z digest=sha256:da5493ebee33f6bda9d0639d39dba36a09424a30b25d6aa063a6beac92530b43

Observation ba6f02e6-5b5e-4765-a2d5-7ce7ab6c3470 · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.756346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:48adcad9ecdd116319e48a586cfee9b2797813fa20e2b25143fb1507d5cd6a67

Observation 2586924b-d8a2-40a5-8ac9-2d56e349ac96 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 293

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.252950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:40b64c29521b3a95d7ae823d39801cfe157340d5b1df35d653a1dd1ef96e4548

Observation 5fbed452-07e8-4d5e-bcb8-94d749fcab45 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.262296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.262296Z digest=sha256:79e2557690b8207ec68942bb519aed2a0fb32e3e9b7cb4debc1236cff3a9250e

Observation 72c7d2c3-637d-49f7-9744-6d1eec08e422 · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.829234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:7ae91e65dc20fc903cbba18857a1e36445d77e3a14834b2b8ba41cefdfddb416

Observation e177162d-e68a-4e9f-a1e1-382d99dbc1e8 · inbound

Seven Security Challenges in Cross-domain Multi-agent LLM Systems cites this paper.

Seven Security Challenges in Cross-domain Multi-agent LLM Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:34.852642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:34.852642Z digest=sha256:1df3eef5726248ce283642401936d6aaeb1ecc4513ce2d96a2f36eee8d94fbe0

Observation e45d494d-1e18-483f-80c2-3892c392dc02 · inbound

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities cites this paper.

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:58.079163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:58.079163Z digest=sha256:251bd763d212d0101c816784256a534cf6682b883594f8800f1534f7eeec7304

Observation 75da2dfd-2a96-4e97-9eda-f29973741e30 · inbound

Normative Conflicts and Shallow AI Alignment cites this paper.

Normative Conflicts and Shallow AI Alignment Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:54.417585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:54.417585Z digest=sha256:34f7dd0059343106c596ac412f26f5680a717e694f6aba1b0a5d23605fffa910

Observation 80107157-42c9-4510-a17b-5ec0990ce669 · inbound

Prompt Injection 2.0: Hybrid AI Threats cites this paper.

Prompt Injection 2.0: Hybrid AI Threats Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:05.101373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:05.101373Z digest=sha256:0f4c0f0ce72738e95294083793265b2b6ac9fdf0fb48198c878a702acf331d4e

Observation 6cd493d6-d804-492f-b72c-8cde5d1c5bb4 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.603513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.603513Z digest=sha256:cfc0da3a7b5e883a1e80c9c19e18f15bc03375a24c2c8c40f4e98e49e2670d8b

Observation f12926ff-74ab-4146-9cad-112b3fb9769a · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.258454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.258454Z digest=sha256:3f3dbc57f72fff5cc6f638eee3e057269a28245b487d3f5ccb4ea87f1ec454ed

Observation b451794c-2a09-460f-9695-11d6ad1a2e58 · inbound

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous cites this paper.

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:09.526632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:09.526632Z digest=sha256:4dd5473292c05c7df473f281303388b0590ff9e7344ad0bff523a558074fd5f8

Observation d2734992-083f-4ae0-b78f-89aaeff0a021 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.045980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:39d659d4bd6f4415310eb2426c4f4906d4194a08969bbf757c13d73938c80d98

Observation f42a9e4d-5ec1-476b-9dd8-94feee168b2e · inbound

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents cites this paper.

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:14.163930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:14.163930Z digest=sha256:9e1d2ae52d8b02036ef4ae4baf83b1892eed14cfb07158e6fbf49aa0c9be5941

Observation 8463fab9-571f-4e8e-b7dd-dabc3bfac939 · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.422099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:5a3aab89f962441cea2fe62b077f27b7ab29ecb80ff6ae3efa02ed42bc03cd9b

Observation 34480220-1d5d-4fee-98a5-820723d20d35 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:bdba15920a70330f454ba94ee046c2e7c8ced07b78573d8a2022908dce2caca6

Observation ba77251d-9df2-4d85-8ca2-bdeb9251ed60 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.867051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:f3b4f74d53b8baa1f96bc17d36015bebb0b3d1bc94b5a3419e98c7b3565b0c13

Observation e7e78990-4ca0-4ece-913f-9385457e017b · inbound

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents cites this paper.

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:04.276860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:51:47.820649Z digest=sha256:3c242f15c31818be5d92b7b0eaad7aaacf8eef313eb5b3d9b2895f4133576295

Observation f7c90804-d766-4002-a5b8-6af4d28cd976 · inbound

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks cites this paper.

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:07.944342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:37:28.671785Z digest=sha256:1578f818e3ea5a38503913a60b18a3930298e45f98c9a1b4a16f7ef9fe0bf68c

Observation 3d59fd30-10a5-4c33-a868-42015cdb1a8f · inbound

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents cites this paper.

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-08T22:39:20.664808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:08:24.524671Z digest=sha256:f65926e111bcd703a5e2c6495d2f3c873abe291900cb1de3d69424bef8395841

Observation 86b07d54-192f-42a8-bf58-7f0780903955 · inbound

Semantic Denial of Service in LLM-controlled robots cites this paper.

Semantic Denial of Service in LLM-controlled robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:14.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T07:59:42.478294Z digest=sha256:0159749c430645c515517a8903dc842e4b0edc5ec2cd5d68bda9744e3095cf29

Observation 2bf34282-1d7d-4985-99ff-d4341cef654e · inbound

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems cites this paper.

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:26.904379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T09:44:14.993126Z digest=sha256:a63647d130df0269f1c13047ce8f4925de9045a6cfd88aed143b42208bb09ee3

Observation 8b73f468-d843-4c22-a309-b934f46880ed · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.007157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:c555c19733447f2907eecb37312d57c8ac9f663fd2bba33021745e0249d8575f

Observation dd5b62e4-8956-45c1-8185-8601b0578196 · inbound

Laundering AI Authority with Adversarial Examples cites this paper.

Laundering AI Authority with Adversarial Examples Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.017954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:19:38.662062Z digest=sha256:2ed7a8811766fe227f2302347219f2e23bc6003339774b5caf8b4d76be44fe42

Observation b3141f92-532f-4220-9c54-27138fe399e7 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.541381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:d7484ff7600203fb3200ad6b4784a9ed27ecc9b47019a6c7e945ae99e1765b27

Observation a11d1780-b5a2-4e01-8618-3e765bf64265 · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:48:11.621232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T09:46:42.413501Z digest=sha256:dbead7e7c0868604d89e875cc76f2159d5564e226bac84b9e26b904ea822028d

Observation 37c91339-e6ee-4f0d-a953-64b7410774ae · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.902347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:57:29.491043Z digest=sha256:ed03c094968faaee7d9250f3a4f078ebc21654ee5f7befb7671f48d3ff3dfad2

Observation b2cc4d12-f2ca-4f45-ae83-c5cd23ffa2c7 · inbound

The Surface You Test Is Not the Surface That Breaks cites this paper.

The Surface You Test Is Not the Surface That Breaks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.206657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T06:37:19.674012Z digest=sha256:6590f178562fd7bbf8a2f2d425bd9fbb04df3cd63d63fcbc9ad801f8e6cc45ef

Observation f1131dda-ded6-43fc-b755-4e2713447359 · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.607755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:58921303957de8db72ce507e3ba23dda4b20e36a95d189c4f962dcdf9a478993

Observation 10cc6aee-624b-410a-a368-a9a0c102c966 · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.719751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:2fea3b4755c53bf707187eaafddeb59843de1becfde3fa73e3e3fd9092e0f42c

Observation acb5163e-3814-4a3b-85b2-f480b497ffa0 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.301820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:90c75f8be6fbe6a2e9bfedf60e88247728ed5c27fad5ec8019c274c9d6b2c676

Observation bfa57969-40fc-4e2e-8e8c-8cc1183fea77 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T02:14:12.243970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:14:12.243970Z digest=sha256:b0f3557eaee67d2319858bd2f8fef86d42132571e61b8642651886c618332ae2

Observation e9c648c7-be23-4f06-805c-ed9d7d789a84 · inbound

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices cites this paper.

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T13:02:42.673767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:02:42.673767Z digest=sha256:26018915f6dc1b7e55c853362f76e3f7dc5f09ad7ffdb212b1d8fe65dda71fdc

Observation 3ca2b105-d4f9-4efc-a76a-fcd377754eae · inbound

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents cites this paper.

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:43:50.536410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:43:50.536410Z digest=sha256:31f25f63807e35ae95113bb6de0746c0d8d62b6b6174ed079511fc40ac89ec51

Observation 9b7750ee-0b58-4992-bd03-9af838bcdeb8 · inbound

Agent Security Needs Redefinition through a Holistic Framework cites this paper.

Agent Security Needs Redefinition through a Holistic Framework Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 295

Resolution
unresolved
no resolver link, observed 2026-08-01T06:04:46.731247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:04:46.731247Z digest=sha256:0b79b328f9fe5be767b15cda137cc257d9a0f82198a38ae5f0b91b901bcf1361

Observation 293a4c02-1cd1-473b-8c3d-18967550e9d9 · inbound

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails cites this paper.

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:20:22.213214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:20:22.213214Z digest=sha256:9ef93b2e3c44e04ebd5446f00518b4b0703208699607ecd8df4637ba07a86cfe

Observation 6e5903d9-8340-4ce3-87d2-e302054779fd · inbound

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots cites this paper.

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:15.346908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:15.346908Z digest=sha256:968152a6554a9b42cd149dfa67384fe7ab246ac03321117dd8d8033416d4ffbc