REVIEW 14 cited by
Stealing Part of a Production Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under \$20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.
Forward citations
Cited by 14 Pith papers
-
Can Watermarking Techniques Help Prevent LLM Model Stealing?
Softplus-then-perturb with embedding-seeded Gaussian noise defeats PCA/averaging/RPCA dimension-extraction attacks on Mistral-7B and GPT-2 with only modest quality loss.
-
Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
A trained transformer's next-token distribution often matches the empirical next-token distribution of its pretraining corpus, with agreement improving as models grow, while a persistent tail of mismatches remains.
-
Characterizing Linear Alignment Across Language Models
Linear (affine) maps between final hidden states of independent LLMs preserve downstream performance and can enable text generation when tokenizers and scale align, enabling a practical HE-based privacy protocol.
-
CrypTorch: PyTorch-based Auto-tuning Compiler for Machine Learning with Multi-party Computation
An MPC-ML compiler that modularizes and auto-tunes operator approximations, delivering 1.2–1.8x speedups over an optimized baseline under user-set accuracy bounds.
-
Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
Malicious calendar invites and emails can poison Gemini's context, enabling data exfiltration, app control, and physical-world actions.
-
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
WikiMem, a Wikidata-derived canary dataset and a calibrated NLL-ranking metric, identifies which human-fact associations an LLM has memorized, with higher rates for famous people and larger models.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
Approximating Language Model Training Data from Weights
A gradient-based greedy selection method (SELECT) recovers effective substitute fine-tuning data from two language model checkpoints, approaching the original model's performance on classification and SFT tasks.
-
How Well Do AI Systems Solve AP Physics? A Comparative Evaluation of Large Language Models on Algebra-Based Free Response Questions
ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1 average 82–92% on AP Physics 1/2 free-response questions but systematically fail spatial, visual, and conceptual tasks.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives
The paper classifies model extraction attacks and defenses into attack, defense, and computing environment categories and surveys their current state.
-
A Survey on Model Extraction Attacks and Defenses for Large Language Models
A taxonomy of model extraction attacks and defenses for large language models, with proposed evaluation metrics and future research directions.
-
Differentiating hype from practical applications of large language models in medicine -- a primer for healthcare professionals
LLMs are statistical text predictors that cannot reason about truth, so in medicine they should be used only with human oversight for tasks like summarization, data extraction, and scheduling.
-
Report on NSF Workshop on Science of Safe AI
An NSF workshop report articulating a cross-disciplinary research agenda for designing and verifying safe, trustworthy AI systems.
Discussion (0). Sign in to comment.