Pith. sign in

REVIEW 6 cited by

Effective Prompt Extraction from Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.06865 v3 pith:M6MWZVPK submitted 2023-07-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords promptpromptsmodelsattacksextractionlanguageframeworkhigh
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The text generated by large language models is commonly controlled by prompting, where a prompt prepended to a user's query guides the model's output. The prompts used by companies to guide their models are often treated as secrets, to be hidden from the user making the query. They have even been treated as commodities to be bought and sold on marketplaces. However, anecdotal reports have shown adversarial users employing prompt extraction attacks to recover these prompts. In this paper, we present a framework for systematically measuring the effectiveness of these attacks. In experiments with 3 different sources of prompts and 11 underlying large language models, we find that simple text-based attacks can in fact reveal prompts with high probability. Our framework determines with high precision whether an extracted prompt is the actual secret prompt, rather than a model hallucination. Prompt extraction from real systems such as Claude 3 and ChatGPT further suggest that system prompts can be revealed by an adversary despite existing defenses in place.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agent Data Injection Attacks are Realistic Threats to AI Agents

    cs.CR 2026-07 accept novelty 7.0 of 10

    Agent data injection (ADI) forges trusted agent metadata via probabilistic delimiter injection and bypasses defenses built only for instruction injection.

  2. Privacy and Security Threat for OpenAI GPTs

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A large-scale study finds that over 98.8% of sampled OpenAI custom GPTs leak their system instructions to crafted adversarial prompts, and hundreds of GPTs transmit user conversation data to third parties.

  3. A Critical Evaluation of Defenses against Prompt Injection Attacks

    cs.CR 2025-05 conditional novelty 6.0 of 10

    StruQ, SecAlign, Instruction Hierarchy, PromptGuard, and Attention Tracker are substantially less effective and utility-preserving than claimed when evaluated with diverse prompts and adaptive attacks.

  4. Federated In-Context Learning: Iterative Refinement for Improved Answer Quality

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fed-ICL iteratively refines QA answers via federated in-context learning with only label transmission, showing convergence on a linear attention model and gains on MMLU and TruthfulQA.

  5. System Prompt Extraction Attacks and Defenses in Large Language Models

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A benchmarking study shows that chain-of-thought, few-shot, and modified sandwich queries can recover LLM system prompts with high similarity-based success, and output filtering is the most reliable tested defense.

  6. Securing AI Systems: A Guide to Known Attacks and Impacts

    cs.CR 2025-06 conditional novelty 3.0 of 10

    A practitioner-oriented review that organizes known adversarial attacks on predictive and generative AI systems into eleven types mapped to confidentiality, integrity, and availability impacts.

Pith tools