Pith. sign in

REVIEW 2 cited by

Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18639 v4 pith:7Y4HQVT3 submitted 2024-10-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords diffusionmodelsattributiontrainingscorecontributiondatadistributions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As diffusion models become increasingly popular, the misuse of copyrighted and private images has emerged as a major concern. One promising solution to mitigate this issue is identifying the contribution of specific training samples in generative models, a process known as data attribution. Existing data attribution methods for diffusion models typically quantify the contribution of a training sample by evaluating the change in diffusion loss when the sample is included or excluded from the training process. However, we argue that the direct usage of diffusion loss cannot represent such a contribution accurately due to the calculation of diffusion loss. Specifically, these approaches measure the divergence between predicted and ground truth distributions, which leads to an indirect comparison between the predicted distributions and cannot represent the variances between model behaviors. To address these issues, we aim to measure the direct comparison between predicted distributions with an attribution score to analyse the training sample importance, which is achieved by Diffusion Attribution Score (\textit{DAS}). Underpinned by rigorous theoretical analysis, we elucidate the effectiveness of DAS. Additionally, we explore strategies to accelerate DAS calculations, facilitating its application to large-scale diffusion models. Our extensive experiments across various datasets and diffusion models demonstrate that DAS significantly surpasses previous benchmarks in terms of the linear data-modelling score, establishing new state-of-the-art performance. Code is available at \hyperlink{here}{https://github.com/Jinxu-Lin/DAS}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Directional Influence Function: Estimating Training Data Influence in Constrained Learning

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Directional Influence Function estimates training-point impact on constrained learners by linearizing the variational inequality of optimality and solving a small QP.

  2. A LoRA is Worth a Thousand Pictures

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LoRA weight vectors, projected with PCA and a per-PC calibration, cluster and retrieve artistic styles more accurately than CLIP, DINO, and style-specialized image features.

Pith tools