Pith. sign in

REVIEW 6 cited by

Data Shapley in One Training Run

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11011 v3 pith:OHW2QTGL submitted 2024-06-16 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords dataattributionmodelmodelspretrainingshapleyalgorithmcontribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts. However, existing approaches require re-training models on different data subsets, which is computationally intensive, foreclosing their application to large-scale models. Furthermore, they produce the same attribution score for any models produced by running the learning algorithm, meaning they cannot perform targeted attribution towards a specific model obtained from a single run of the algorithm. This paper introduces In-Run Data Shapley, which addresses these limitations by offering scalable data attribution for a target model of interest. In its most efficient implementation, our technique incurs negligible additional runtime compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage for the first time. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Asymptotic Analysis of the Shapley Value for Dataset Valuation

    cs.GT 2026-07 conditional novelty 7.0 of 10

    Under smooth RKHS embedding utilities, a fixed owner's Shapley value is O(1/I)-close in L1 to an explicit leading term of scale (log I)/I driven by a first-order population signal.

  2. Newfluence: Boosting Model interpretability and Understanding in High Dimensions

    stat.ML 2025-07 conditional novelty 6.0 of 10

    In high-dimensional regression, classical influence functions underestimate true leave-one-out influence by a per-point factor, and the proposed Newfluence estimator corrects this bias.

  3. DICE: Data Influence Cascade in Decentralized Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    DICE defines and approximates multi-hop data influence in decentralized learning, showing that influence is shaped by data, topology, and loss curvature.

  4. Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation

    cs.LG 2026-03 reject novelty 5.0 of 10

    Local Shapley restricts data valuation to per-test support sets and reuses subset trainings, but the claimed exactness and concentration bounds are flawed.

  5. KAIROS: Scalable Model-Agnostic Data Valuation

    cs.LG 2025-06 conditional novelty 5.0 of 10

    KAIROS derives a closed-form Maximum Mean Discrepancy influence score that approximates leave-one-out data rankings and detects noise, mislabels, and backdoors without retraining.

  6. Collective Bargaining in the Information Economy Can Address AI-Driven Power Concentration

    cs.CY 2025-06 conditional novelty 4.0 of 10

    A policy agenda proposes collective bargaining by information producers as the principal fix for AI-driven market concentration and collapse of the information commons.

Pith tools