Pith. sign in

REVIEW 1 cited by

Unlearning Traces the Influential Training Data of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15241 v2 pith:CL6CAYDN submitted 2024-01-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords traininginfluencemodeldatasetdatasetsmethodsunlearninguntrac
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Identifying the training datasets that influence a language model's outputs is essential for minimizing the generation of harmful content and enhancing its performance. Ideally, we can measure the influence of each dataset by removing it from training; however, it is prohibitively expensive to retrain a model multiple times. This paper presents UnTrac: unlearning traces the influence of a training dataset on the model's performance. UnTrac is extremely simple; each training dataset is unlearned by gradient ascent, and we evaluate how much the model's predictions change after unlearning. Furthermore, we propose a more scalable approach, UnTrac-Inv, which unlearns a test dataset and evaluates the unlearned model on training datasets. UnTrac-Inv resembles UnTrac, while being efficient for massive training datasets. In the experiments, we examine if our methods can assess the influence of pretraining datasets on generating toxic, biased, and untruthful content. Our methods estimate their influence much more accurately than existing methods while requiring neither excessive memory space nor multiple checkpoints.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairshare Data Pricing via Data Valuation for Large Language Models

    cs.GT 2025-01 conditional novelty 5.0 of 10

    A game-theoretic model and simulations claim that pricing LLM training data at each buyer's maximum willingness to pay, computed from data-valuation scores, is optimal for buyers and sellers over the long run.

Pith tools