Pith. sign in

REVIEW 4 cited by

Language Model Inversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13647 v1 pith:VE4WIBH2 submitted 2023-11-22 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelinversionlanguagepromptsrecoverconsiderdistributioninformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Language models produce a distribution over the next token; can we use this information to recover the prompt tokens? We consider the problem of language model inversion and show that next-token probabilities contain a surprising amount of information about the preceding text. Often we can recover the text in cases where it is hidden from the user, motivating a method for recovering unknown prompts given only the model's current distribution output. We consider a variety of model access scenarios, and show how even without predictions for every token in the vocabulary we can recover the probability vector through search. On Llama-2 7b, our inversion method reconstructs prompts with a BLEU of $59$ and token-level F1 of $78$ and recovers $27\%$ of prompts exactly. Code for reproducing all experiments is available at http://github.com/jxmorris12/vec2text.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can Watermarking Techniques Help Prevent LLM Model Stealing?

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Softplus-then-perturb with embedding-seeded Gaussian noise defeats PCA/averaging/RPCA dimension-extraction attacks on Mistral-7B and GPT-2 with only modest quality loss.

  2. Black-Box Inference of LLM Architectural Properties with Restrictive API Access

    cs.LG 2026-07 unverdicted novelty 6.0 of 10

    NightVision recovers LLM hidden dimension to 23% average relative error (9% on MoE) and depth/parameter count to 53% on models >3B parameters using common-set prompting, spectral analysis, and TTFT under single-logit ...

  3. inversedMixup: Data Augmentation via Inverting Mixed Embeddings

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Mixing BERT embeddings and inverting them into text with LLaMA produces interpretable augmented sentences, improves few-shot classification on some datasets, and exposes 'manifold intrusion' in text Mixup.

  4. Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare

    cs.CR 2026-06 unverdicted novelty 4.0 of 10

    TRiSM-guided agentic workflows reduced RAG poisoning attack success from 31% to 10%, data-field injection from 42% to 25%, eliminated network injection, and raised report accuracy from 72.5% to 86.5% across five LLMs ...

Pith tools