Pith. sign in

REVIEW 3 cited by

The Journey, Not the Destination: How Data Guides Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06205 v1 pith:E4Z4YTHR submitted 2023-12-11 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusionmodelsattributionstraineddataimagesmethodtraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models trained on large datasets can synthesize photo-realistic images of remarkable quality and diversity. However, attributing these images back to the training data-that is, identifying specific training examples which caused an image to be generated-remains a challenge. In this paper, we propose a framework that: (i) provides a formal notion of data attribution in the context of diffusion models, and (ii) allows us to counterfactually validate such attributions. Then, we provide a method for computing these attributions efficiently. Finally, we apply our method to find (and evaluate) such attributions for denoising diffusion probabilistic models trained on CIFAR-10 and latent diffusion models trained on MS COCO. We provide code at https://github.com/MadryLab/journey-TRAK .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

    cs.CY 2026-06 accept novelty 6.5 of 10

    Legal and ethical bans on CSAM access and generation break standard AI safety techniques, creating 15 open problems that demand new methods for dataset cleaning, concept fusion prevention, fine-tuning resilience, dete...

  2. Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning

    cs.LG 2025-10 conditional novelty 5.0 of 10

    Restricting diffusion unlearning to a tuned time window and to low-pass-filtered images improves image quality and unlearning speed relative to uniform unlearning in both face-image and text-to-image settings.

  3. Blink of an eye: a simple theory for feature localization in generative models

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Critical windows in language and diffusion models are characterized, under a shared degradation process, as the interval where a target sub-population remains separable while a smaller sub-population becomes indisting...

Pith tools