Pith. sign in

REVIEW 1 cited by

On the Feasibility of In-Context Probing for Data Attribution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12259 v2 pith:PYS3C5I6 submitted 2024-07-17 cs.CL

classification cs.CL
keywords dataattributiontasksconnectiongradient-basedmethodsmodeltraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data attribution methods are used to measure the contribution of training data towards model outputs, and have several important applications in areas such as dataset curation and model interpretability. However, many standard data attribution methods, such as influence functions, utilize model gradients and are computationally expensive. In our paper, we show in-context probing (ICP) -- prompting a LLM -- can serve as a fast proxy for gradient-based data attribution for data selection under conditions contingent on data similarity. We study this connection empirically on standard NLP tasks, and show that ICP and gradient-based data attribution are well-correlated in identifying influential training data for tasks that share similar task type and content as the training data. Additionally, fine-tuning models on influential data selected by both methods achieves comparable downstream performance, further emphasizing their similarities. We also examine the connection between ICP and gradient-based data attribution using synthetic data on linear regression tasks. Our synthetic data experiments show similar results with those from NLP tasks, suggesting that this connection can be isolated in simpler settings, which offers a pathway to bridging their differences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairshare Data Pricing via Data Valuation for Large Language Models

    cs.GT 2025-01 conditional novelty 5.0 of 10

    A game-theoretic model and simulations claim that pricing LLM training data at each buyer's maximum willingness to pay, computed from data-valuation scores, is optimal for buyers and sellers over the long run.

Pith tools