REVIEW 11 cited by
Imprompter: Tricking LLM Agents into Improper Tool Use
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Model (LLM) Agents are an emerging computing paradigm that blends generative machine learning with tools such as code interpreters, web browsing, email, and more generally, external resources. These agent-based systems represent an emerging shift in personal computing. We contribute to the security foundations of agent-based systems and surface a new class of automatically computed obfuscated adversarial prompt attacks that violate the confidentiality and integrity of user resources connected to an LLM agent. We show how prompt optimization techniques can find such prompts automatically given the weights of a model. We demonstrate that such attacks transfer to production-level agents. For example, we show an information exfiltration attack on Mistral's LeChat agent that analyzes a user's conversation, picks out personally identifiable information, and formats it into a valid markdown command that results in leaking that data to the attacker's server. This attack shows a nearly 80% success rate in an end-to-end evaluation. We conduct a range of experiments to characterize the efficacy of these attacks and find that they reliably work on emerging agent-based systems like Mistral's LeChat, ChatGLM, and Meta's Llama. These attacks are multimodal, and we show variants in the text-only and image domains.
Forward citations
Cited by 11 Pith papers
-
Agent Data Injection Attacks are Realistic Threats to AI Agents
Agent data injection (ADI) forges trusted agent metadata via probabilistic delimiter injection and bypasses defenses built only for instruction injection.
-
DualView: Preventing Indirect Prompt Injection in Personal AI Agents
DualView extends Dual-LLM symbol isolation into the shared user environment via dual Agent/Human views, blocking both immediate and stored IPI at 0% ASR while preserving near-baseline utility.
-
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
Memory can be exploited to hijack LLM agents' tool selection and induce persistent behavioral deviations even against explicit instructions and safety constraints.
-
Agent Security Needs Redefinition through a Holistic Framework
Agent security should be redefined around four contextual authorization properties instead of the content of the action performed.
-
AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
AgentSentinel combines system-level tracing with LLM-based auditing to block 79.6% of attacks in the authors' 60-scenario computer-use agent benchmark.
-
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
GCG-optimized trigger strings embedded in HTML can command LLM web agents to perform attacker-chosen actions, including credential exfiltration.
-
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
MCP-powered LLM agents are vulnerable to prompt injection from third-party services, and simple detection or filtering defenses do not reliably stop these attacks.
-
Lessons from Defending Gemini Against Indirect Prompt Injections
Google DeepMind reports that adversarially fine-tuning Gemini 2.5 cut indirect prompt-injection success by roughly half on average in tested settings, but adaptive attackers still found gaps.
-
On the Future of Software Reuse in the Era of AI Native Software Engineering
The paper frames AI-assisted generative code reuse as a cargo-cult-like practice and lays out a research agenda for its quality, security, licensing, and maintainability challenges.
-
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.
-
Data-Centric Safety and Ethical Measures for Data and AI Governance
A conceptual framework that maps dataset safety practices to six stages of the AI lifecycle, synthesizing existing documentation and red-teaming recommendations.
Discussion (0). Continue with ORCID to comment.