Pith. sign in

REVIEW 8 cited by

Synthetic Data from Diffusion Models Improves ImageNet Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08466 v1 pith:UYFC2VQR submitted 2023-04-17 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords modelssamplesclassificationgenerativeimagenetx256accuracydata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep generative models are becoming increasingly powerful, now generating diverse high fidelity photo-realistic samples given text prompts. Have they reached the point where models of natural images can be used for generative data augmentation, helping to improve challenging discriminative tasks? We show that large-scale text-to image diffusion models can be fine-tuned to produce class conditional models with SOTA FID (1.76 at 256x256 resolution) and Inception Score (239 at 256x256). The model also yields a new SOTA in Classification Accuracy Scores (64.96 for 256x256 generative samples, improving to 69.24 for 1024x1024 samples). Augmenting the ImageNet training set with samples from the resulting models yields significant improvements in ImageNet classification accuracy over strong ResNet and Vision Transformer baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 79 citations worldwide. Full citation record

  1. S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

    cs.CV 2025-10 conditional novelty 7.0 of 10

    A 139k-image synthetic dataset with diffusion- and DINO-derived masks, trained with a multi-mask decoder, improves cross-dataset salient-object detection and reaches state-of-the-art after fine-tuning.

  2. SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

    cs.CR 2025-06 conditional novelty 7.0 of 10

    A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...

  3. MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Saliency-guided diffusion mixing that preserves Grad-CAM-highlighted diagnostic regions improves medical image classification accuracy and AUC across four public datasets.

  4. Understanding Trade offs When Conditioning Synthetic Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Diverse layout-plus-prompt conditioning of diffusion models generates synthetic data that improves few-shot object detection mAP by up to 177% over real-data-only training, while prompt-only conditioning wins when con...

  5. Taming Diffusion for Dataset Distillation with High Representativeness

    cs.CV 2025-05 conditional novelty 6.0 of 10

    D3HR maps VAE latents to a near-Gaussian space via DDIM inversion, fits per-class Gaussian statistics, and selects the most moment-matching subset to generate distilled images, reporting state-of-the-art accuracy.

  6. Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective

    cs.CV 2025-07 reject novelty 5.0 of 10

    Edge-case synthesis with a fine-tuned text-to-image model improves fisheye object detection, but the gain is not isolated from simply adding more data.

  7. Generative Data Augmentation for Object Point Cloud Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A mask-conditioned diffusion model generates labeled point cloud variants and filtered pseudo-labels, improving object part segmentation with only 10% hand labels.

  8. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

Pith tools