Pith. sign in

REVIEW 10 cited by

Prompt Stealing Attacks Against Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12959 v1 pith:5TKSJTDP submitted 2024-02-20 cs.CR cs.CL

classification cs.CRcs.CL
keywords promptpromptsattacksextractorparameterstealinganswersattack
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing reliance on large language models (LLMs) such as ChatGPT in various fields emphasizes the importance of ``prompt engineering,'' a technology to improve the quality of model outputs. With companies investing significantly in expert prompt engineers and educational resources rising to meet market demand, designing high-quality prompts has become an intriguing challenge. In this paper, we propose a novel attack against LLMs, named prompt stealing attacks. Our proposed prompt stealing attack aims to steal these well-designed prompts based on the generated answers. The prompt stealing attack contains two primary modules: the parameter extractor and the prompt reconstruction. The goal of the parameter extractor is to figure out the properties of the original prompts. We first observe that most prompts fall into one of three categories: direct prompt, role-based prompt, and in-context prompt. Our parameter extractor first tries to distinguish the type of prompts based on the generated answers. Then, it can further predict which role or how many contexts are used based on the types of prompts. Following the parameter extractor, the prompt reconstructor can be used to reconstruct the original prompts based on the generated answers and the extracted features. The final goal of the prompt reconstructor is to generate the reversed prompts, which are similar to the original prompts. Our experimental results show the remarkable performance of our proposed attacks. Our proposed attacks add a new dimension to the study of prompt engineering and call for more attention to the security issues on LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Federation over Text: Insight Sharing for Multi-Agent Reasoning

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    FoT lets multiple LLM agents federate text-based reasoning traces into a cross-task insight library, raising average task accuracy by 24% and cutting reasoning tokens by 28%.

  2. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  3. Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Malicious calendar invites and emails can poison Gemini's context, enabling data exfiltration, app control, and physical-world actions.

  4. Has My System Prompt Been Used? Large Language Model Prompt Membership Inference

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A permutation test on BERT embeddings of LLM outputs can detect, with statistical significance, when response distributions differ because a chat service uses a different system prompt than a candidate prompt.

  5. Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks

    cs.CR 2026-04 conditional novelty 5.0 of 10

    Refusal-aligned LLMs leak system instructions under encoding/serialization prompts at high rates, and one-shot CoT instruction reshaping substantially reduces that leakage without retraining.

  6. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  7. A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives

    cs.CR 2025-08 conditional novelty 4.0 of 10

    The paper classifies model extraction attacks and defenses into attack, defense, and computing environment categories and surveys their current state.

  8. A Survey on Model Extraction Attacks and Defenses for Large Language Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A taxonomy of model extraction attacks and defenses for large language models, with proposed evaluation metrics and future research directions.

  9. System Prompt Extraction Attacks and Defenses in Large Language Models

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A benchmarking study shows that chain-of-thought, few-shot, and modified sandwich queries can recover LLM system prompts with high similarity-based success, and output filtering is the most reliable tested defense.

  10. Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

    cs.NE 2025-08 reject novelty 3.0 of 10

    The claimed MFO-DBO algorithm and its CEC2017/PV results are absent from the manuscript, which instead contains an unrelated prompt-stealing attack paper.

Pith tools