Pith. sign in

REVIEW 5 cited by

Submix: Practical Private Prediction for Large-Scale Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.00971 v1 pith:MICA4MAN submitted 2022-01-04 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords privatesubmixlanguagemodelsprivacycorpuspredictionattacks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent data-extraction attacks have exposed that language models can memorize some training samples verbatim. This is a vulnerability that can compromise the privacy of the model's training data. In this work, we introduce SubMix: a practical protocol for private next-token prediction designed to prevent privacy violations by language models that were fine-tuned on a private corpus after pre-training on a public corpus. We show that SubMix limits the leakage of information that is unique to any individual user in the private corpus via a relaxation of group differentially private prediction. Importantly, SubMix admits a tight, data-dependent privacy accounting mechanism, which allows it to thwart existing data-extraction attacks while maintaining the utility of the language model. SubMix is the first protocol that maintains privacy even when publicly releasing tens of thousands of next-token predictions made by large transformer-based models such as GPT-2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation

    cs.IT 2026-07 conditional novelty 6.0 of 10

    For histogram-based DP synthetic data, mixing the private histogram with a floor-raised version of a same-domain public distribution is asymptotically the best linear privacy mechanism.

  2. Lower Bounds for Public-Private Learning under Distribution Shift

    cs.LG 2025-07 reject novelty 6.0 of 10

    For Gaussian mean estimation and linear regression with distribution shift, the paper claims that public data never provides complementary value: either public data alone suffices, or (for large shifts) private data a...

  3. Differentially Private In-context Learning via Sampling Few-shot Mixed with Zero-shot Outputs

    cs.LG 2025-01 reject novelty 5.0 of 10

    DPS-MOZO samples each generated token from the product of per-example distributions mixed with the zero-shot distribution to make in-context learning differentially private without additive noise.

  4. Public Data Assisted Differentially Private In-Context Learning

    cs.AI 2025-09 conditional novelty 4.0 of 10

    A private ICL algorithm that aggregates LLM responses with DPM clustering and uses public data representatives achieves near-non-private utility at epsilon=1.

  5. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

    cs.CR 2025-12 conditional novelty 2.0 of 10

    A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.

Pith tools