Pith. sign in

Paper Citation Record · LEDGER

Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2307.10490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.10490 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:07:35.853723Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8bcc565e-2fd1-4a78-95d7-86c2094b28be · inbound

Whispers in the Machine: Confidentiality in Agentic Systems cites this paper.

Whispers in the Machine: Confidentiality in Agentic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:03:53.924075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T03:59:03.972043Z digest=sha256:60b1ad1e6e12e44767a863f4471a914cc420ca58153bb4ac81b3df635c1289a8

Observation ba6f02e6-5b5e-4765-a2d5-7ce7ab6c3470 · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.756346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:fc036f578285d1f39f7ea93e9c07d4538031d557b15f8ac42ff03cd33f477785

Observation 2586924b-d8a2-40a5-8ac9-2d56e349ac96 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 293

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.252950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:5df89b59222dc15d3e0dfe14bc8108945fb1d70322e17511ee08e5fe2712856b

Observation 5fbed452-07e8-4d5e-bcb8-94d749fcab45 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.262296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.262296Z digest=sha256:597dbfea2913ceaf6a5e38a13d0f5b5aa4d2f174a35033f748928f7aa042678a

Observation 72c7d2c3-637d-49f7-9744-6d1eec08e422 · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.829234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:0da5d5d7347fd51086750139fd906829b7dd6bda4bb378f1387adae1547d870a

Observation e177162d-e68a-4e9f-a1e1-382d99dbc1e8 · inbound

Seven Security Challenges in Cross-domain Multi-agent LLM Systems cites this paper.

Seven Security Challenges in Cross-domain Multi-agent LLM Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:34.852642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:34.852642Z digest=sha256:edba8bba522f2a721a03330c78603168e0b350d1267ac2afe9cce71e738bc0f5

Observation e45d494d-1e18-483f-80c2-3892c392dc02 · inbound

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities cites this paper.

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:58.079163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:58.079163Z digest=sha256:2a55e478d9c7973d8683be5e2b5dd6e01a37588696e43a020e055701a1e8f111

Observation 75da2dfd-2a96-4e97-9eda-f29973741e30 · inbound

Normative Conflicts and Shallow AI Alignment cites this paper.

Normative Conflicts and Shallow AI Alignment Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:54.417585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:40:54.417585Z digest=sha256:f09d66d25653ccebb185709f38731f3d1e746c2c3cdd4b00017c6e71ee992d5d

Observation 80107157-42c9-4510-a17b-5ec0990ce669 · inbound

Prompt Injection 2.0: Hybrid AI Threats cites this paper.

Prompt Injection 2.0: Hybrid AI Threats Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:05.101373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:05.101373Z digest=sha256:858f51eb6e6144ba03cd5084249750d6d20410a2a1793e66885b561346b77d9e

Observation 6cd493d6-d804-492f-b72c-8cde5d1c5bb4 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.603513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.603513Z digest=sha256:133783126f0bc4bb096862f63cbb4d9747859605e7c7f7f02c46d3dfc00e02cb

Observation f12926ff-74ab-4146-9cad-112b3fb9769a · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.258454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.258454Z digest=sha256:9ea7c0d6c9fbd8eba1fec1dbbd5375039a1e1dc391d98df46203a68857dcf642

Observation b451794c-2a09-460f-9695-11d6ad1a2e58 · inbound

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous cites this paper.

Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:09.526632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:09.526632Z digest=sha256:3a087fc8895873d383d061c737ce0bf73d4caafe37ff54d4b919895ace47ead8

Observation d2734992-083f-4ae0-b78f-89aaeff0a021 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.045980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:328837e50df7f99ced2e81eaba3c2c1772387065f5b1de03f330aa52547d45f5

Observation f42a9e4d-5ec1-476b-9dd8-94feee168b2e · inbound

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents cites this paper.

Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:14.163930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:14.163930Z digest=sha256:a84768c9ae5c2959b53e3b84af939f546723c735ae4dbd887d9508eb850cf0bf

Observation 8463fab9-571f-4e8e-b7dd-dabc3bfac939 · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.422099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:0e164e09067742ad16ffbef3438496b7fa85f3ad170fc35cbfc25a23bddab177

Observation 34480220-1d5d-4fee-98a5-820723d20d35 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:353cfb7cf14aeb792396b226095001ef743b40e3b756b603409c70f9009a3c63

Observation ba77251d-9df2-4d85-8ca2-bdeb9251ed60 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.867051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:20f4989df4ba976be57d6ebecb708df0f81fa8e104df8b464b9d5fbb4a061a4a

Observation e7e78990-4ca0-4ece-913f-9385457e017b · inbound

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents cites this paper.

Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:04.276860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:51:47.820649Z digest=sha256:b40319401ace897d86f5581cb4ab1a73b96ac32549dba21acccede8aaf3f6bb2

Observation f7c90804-d766-4002-a5b8-6af4d28cd976 · inbound

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks cites this paper.

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:07.944342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:37:28.671785Z digest=sha256:190339ca408a9b4d397a9a8fb056292dd000c64790db5ca4fb8f527144eaa1e3

Observation 3d59fd30-10a5-4c33-a868-42015cdb1a8f · inbound

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents cites this paper.

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-08T22:39:20.664808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:08:24.524671Z digest=sha256:014512aa767ad32a072241e40d260723bcb6532f195e658f67e816b9adaacce6

Observation 86b07d54-192f-42a8-bf58-7f0780903955 · inbound

Semantic Denial of Service in LLM-controlled robots cites this paper.

Semantic Denial of Service in LLM-controlled robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:14.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T07:59:42.478294Z digest=sha256:0a4e2b0985a00551b6f48c80cfe51bfaec090f832fd8adb314746b02fdd7fcb2

Observation 2bf34282-1d7d-4985-99ff-d4341cef654e · inbound

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems cites this paper.

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:26.904379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T09:44:14.993126Z digest=sha256:d2fea83259825d1c6df2b187ddbfbfa4cab18263631cf2a3868e03a41fd60f65

Observation 8b73f468-d843-4c22-a309-b934f46880ed · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.007157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:2cd95cc2e4f9c649b1eaab0a9866a8d9e2afdb8128ba4fb1a6e9ff5c7e33b6b0

Observation dd5b62e4-8956-45c1-8185-8601b0578196 · inbound

Laundering AI Authority with Adversarial Examples cites this paper.

Laundering AI Authority with Adversarial Examples Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.017954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:19:38.662062Z digest=sha256:0aa2c6408836bb0d3077e81baea1a481b42eb214362aaf761427e8cda53ea9f9

Observation b3141f92-532f-4220-9c54-27138fe399e7 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.541381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:8876fa99a58067ee71c9acad0489aab7bedcd918b010392f8a8e1beb11f3fd41

Observation a11d1780-b5a2-4e01-8618-3e765bf64265 · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:48:11.621232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T09:46:42.413501Z digest=sha256:16ca190e11cf643538157ec2f5a0685c2258ff563e2d8db2b786071603105713

Observation 37c91339-e6ee-4f0d-a953-64b7410774ae · inbound

Hallucination as Exploit: Evidence-Carrying Multimodal Agents cites this paper.

Hallucination as Exploit: Evidence-Carrying Multimodal Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.902347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T08:57:29.491043Z digest=sha256:bafae1aefd749625e0e2cecb970edf8b8ea7f8436e50a3cab6e6dd06cdb94d4a

Observation b2cc4d12-f2ca-4f45-ae83-c5cd23ffa2c7 · inbound

The Surface You Test Is Not the Surface That Breaks cites this paper.

The Surface You Test Is Not the Surface That Breaks Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.206657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T06:37:19.674012Z digest=sha256:0f9946087f025b595d1e3e76c8196f399d17d00e56f7438c69665c4fd5ab8d5b

Observation f1131dda-ded6-43fc-b755-4e2713447359 · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.607755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:026a68ddb9bfc48012d5f48ead052f980cfb15e09a7c13e5283ed3f76e207dea

Observation 10cc6aee-624b-410a-a368-a9a0c102c966 · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.719751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:cee3948e2b2494767917c0ab98564120729a8205ffcded7a5d6e40c54fbccf23

Observation acb5163e-3814-4a3b-85b2-f480b497ffa0 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.301820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:7d252828dc25a9dbc5d0f3af5a099c9a2d729c388a66c59bd54ee2de28f21aff

Observation bfa57969-40fc-4e2e-8e8c-8cc1183fea77 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T02:14:12.243970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:14:12.243970Z digest=sha256:1edf5e0d323c8a6e3651faadcaf2a59872abf31385ac22b72a4e52d8a40ae1f1

Observation e9c648c7-be23-4f06-805c-ed9d7d789a84 · inbound

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices cites this paper.

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T13:02:42.673767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:02:42.673767Z digest=sha256:76e79f96cdea4448876da3a107e88a1005e72a40990285567d7b4aa4f07637d4

Observation 3ca2b105-d4f9-4efc-a76a-fcd377754eae · inbound

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents cites this paper.

Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:43:50.536410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:43:50.536410Z digest=sha256:adf23d3938d8b4d6ce3ced45f13b51c022421a2d74920f9fac3646cb21bfd648

Observation 9b7750ee-0b58-4992-bd03-9af838bcdeb8 · inbound

Agent Security Needs Redefinition through a Holistic Framework cites this paper.

Agent Security Needs Redefinition through a Holistic Framework Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 295

Resolution
unresolved
no resolver link, observed 2026-08-01T06:04:46.731247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:04:46.731247Z digest=sha256:25a68f533cd496f0fbaabb16f90e3732c0797197759ab6c61f87c1059bbaac26

Observation 293a4c02-1cd1-473b-8c3d-18967550e9d9 · inbound

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails cites this paper.

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:20:22.213214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:20:22.213214Z digest=sha256:8ca5c02291eb4a91a972bd696f14854595d49ea9e5c481c7a76db8a28a6328e2

Observation 646d89a3-18e4-4702-a6e0-1f7169807a60 · inbound

PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents cites this paper.

PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:07:35.853723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:07:35.853723Z digest=sha256:7de0e55918fc9405217e0127871e59325910f36bbcc4ef75b005a76ea5c2a2ac

Observation 6e5903d9-8340-4ce3-87d2-e302054779fd · inbound

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots cites this paper.

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:38:15.346908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:38:15.346908Z digest=sha256:87f80e422da8bc136c0d75bad7fbfda31f0d404953d40c13dc335d2ddd944738