Pith. sign in

REVIEW 2 cited by

Targeted Attack on GPT-Neo for the SATML Language Model Data Extraction Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07735 v1 pith:IEH4EG6C submitted 2023-02-13 cs.CL cs.AIcs.CR

classification cs.CLcs.AIcs.CR
keywords dataextractionattacksmodelattacklanguagetargetedtraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Previous work has shown that Large Language Models are susceptible to so-called data extraction attacks. This allows an attacker to extract a sample that was contained in the training data, which has massive privacy implications. The construction of data extraction attacks is challenging, current attacks are quite inefficient, and there exists a significant gap in the extraction capabilities of untargeted attacks and memorization. Thus, targeted attacks are proposed, which identify if a given sample from the training data, is extractable from a model. In this work, we apply a targeted data extraction attack to the SATML2023 Language Model Training Data Extraction Challenge. We apply a two-step approach. In the first step, we maximise the recall of the model and are able to extract the suffix for 69% of the samples. In the second step, we use a classifier-based Membership Inference Attack on the generations. Our AutoSklearn classifier achieves a precision of 0.841. The full approach reaches a score of 0.405 recall at a 10% false positive rate, which is an improvement of 34% over the baseline of 0.301.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Poisoned Chalice of LLM Evaluation Report

    cs.SE 2026-07 conditional novelty 4.0 of 10

    A competition report showing that white-box membership inference attacks on code LLMs mostly fail (AUC ~0.56–0.61) except for one structure-aware method (SERSEM, AUC ~0.77) that generalizes to a held-out model.

  2. LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures

    cs.CR 2025-05 conditional novelty 3.0 of 10

    This survey categorizes attacks on large language models by lifecycle phase and maps them to prevention and detection defenses, concluding that only a few defenses are highly effective.

Pith tools