Pith. sign in

REVIEW 2 cited by

Private prediction for large-scale synthetic text generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12108 v2 pith:IZZLG4AP submitted 2024-07-16 cs.LG cs.CLcs.CR

classification cs.LGcs.CLcs.CR
keywords dataprivacyprivatesyntheticpredictioncontrastdifferentialensure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an approach for generating differentially private synthetic text using large language models (LLMs), via private prediction. In the private prediction framework, we only require the output synthetic data to satisfy differential privacy guarantees. This is in contrast to approaches that train a generative model on potentially sensitive user-supplied source data and seek to ensure the model itself is safe to release. We prompt a pretrained LLM with source data, but ensure that next-token predictions are made with differential privacy guarantees. Previous work in this paradigm reported generating a small number of examples (<10) at reasonable privacy levels, an amount of data that is useful only for downstream in-context learning or prompting. In contrast, we make changes that allow us to generate thousands of high-quality synthetic data points, greatly expanding the set of potential applications. Our improvements come from an improved privacy analysis and a better private selection mechanism, which makes use of the equivalence between the softmax layer for sampling tokens in LLMs and the exponential mechanism. Furthermore, we introduce a novel use of public predictions via the sparse vector technique, in which we do not pay privacy costs for tokens that are predictable without sensitive data; we find this to be particularly effective for structured data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Differentially Private Generation of Domain-Specific Text

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Applying a new benchmark to five specialized domains, the paper shows current privacy-preserving text generators lose much of their utility and fidelity, especially at strict privacy levels and on gated datasets.

  2. Differentially-private text generation degrades output language quality

    cs.CL 2025-09 conditional novelty 5.0 of 10

    DP fine-tuning systematically degrades LLM output length, grammatical correctness, and lexical diversity, and this degradation grows as the privacy budget shrinks.

Pith tools