Pith. sign in

REVIEW 11 cited by

Imprompter: Tricking LLM Agents into Improper Tool Use

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14923 v2 pith:EU2IPK3G submitted 2024-10-19 cs.CR

classification cs.CR
keywords attacksagent-basedagentsemergingsystemsagentattackautomatically
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Model (LLM) Agents are an emerging computing paradigm that blends generative machine learning with tools such as code interpreters, web browsing, email, and more generally, external resources. These agent-based systems represent an emerging shift in personal computing. We contribute to the security foundations of agent-based systems and surface a new class of automatically computed obfuscated adversarial prompt attacks that violate the confidentiality and integrity of user resources connected to an LLM agent. We show how prompt optimization techniques can find such prompts automatically given the weights of a model. We demonstrate that such attacks transfer to production-level agents. For example, we show an information exfiltration attack on Mistral's LeChat agent that analyzes a user's conversation, picks out personally identifiable information, and formats it into a valid markdown command that results in leaking that data to the attacker's server. This attack shows a nearly 80% success rate in an end-to-end evaluation. We conduct a range of experiments to characterize the efficacy of these attacks and find that they reliably work on emerging agent-based systems like Mistral's LeChat, ChatGLM, and Meta's Llama. These attacks are multimodal, and we show variants in the text-only and image domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agent Data Injection Attacks are Realistic Threats to AI Agents

    cs.CR 2026-07 accept novelty 7.0 of 10

    Agent data injection (ADI) forges trusted agent metadata via probabilistic delimiter injection and bypasses defenses built only for instruction injection.

  2. DualView: Preventing Indirect Prompt Injection in Personal AI Agents

    cs.CR 2026-07 conditional novelty 7.0 of 10

    DualView extends Dual-LLM symbol isolation into the shared user environment via dual Agent/Human views, blocking both immediate and stored IPI at 0% ASR while preserving near-baseline utility.

  3. From Storage to Steering: Memory Control Flow Attacks on LLM Agents

    cs.CR 2026-03 unverdicted novelty 7.0 of 10

    Memory can be exploited to hijack LLM agents' tool selection and induce persistent behavioral deviations even against explicit instructions and safety constraints.

  4. Agent Security Needs Redefinition through a Holistic Framework

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Agent security should be redefined around four contextual authorization properties instead of the content of the action performed.

  5. AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents

    cs.CR 2025-09 conditional novelty 6.0 of 10

    AgentSentinel combines system-level tracing with LLM-based auditing to block 79.6% of attacks in the authors' 60-scenario computer-use agent benchmark.

  6. Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree

    cs.CR 2025-07 conditional novelty 5.0 of 10

    GCG-optimized trigger strings embedded in HTML can command LLM web agents to perform attacker-chosen actions, including credential exfiltration.

  7. We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems

    cs.LG 2025-06 conditional novelty 5.0 of 10

    MCP-powered LLM agents are vulnerable to prompt injection from third-party services, and simple detection or filtering defenses do not reliably stop these attacks.

  8. Lessons from Defending Gemini Against Indirect Prompt Injections

    cs.CR 2025-05 conditional novelty 5.0 of 10

    Google DeepMind reports that adversarially fine-tuning Gemini 2.5 cut indirect prompt-injection success by roughly half on average in tested settings, but adaptive attackers still found gaps.

  9. On the Future of Software Reuse in the Era of AI Native Software Engineering

    cs.SE 2025-08 conditional novelty 4.0 of 10

    The paper frames AI-assisted generative code reuse as a cargo-cult-like practice and lays out a research agenda for its quality, security, licensing, and maintainability challenges.

  10. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.

  11. Data-Centric Safety and Ethical Measures for Data and AI Governance

    cs.CY 2025-06 conditional novelty 3.0 of 10

    A conceptual framework that maps dataset safety practices to six stages of the AI lifecycle, synthesizing existing documentation and red-teaming recommendations.

Pith tools