Pith. sign in

REVIEW 4 cited by

Training on Thin Air: Improve Image Classification with Generated Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15316 v1 pith:NYGVGRJE submitted 2023-05-24 cs.CV cs.LG

classification cs.CVcs.LG
keywords datatrainingapproachdiffusiongeneratedimagesclassificationdiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Acquiring high-quality data for training discriminative models is a crucial yet challenging aspect of building effective predictive systems. In this paper, we present Diffusion Inversion, a simple yet effective method that leverages the pre-trained generative model, Stable Diffusion, to generate diverse, high-quality training data for image classification. Our approach captures the original data distribution and ensures data coverage by inverting images to the latent space of Stable Diffusion, and generates diverse novel training images by conditioning the generative model on noisy versions of these vectors. We identify three key components that allow our generated images to successfully supplant the original dataset, leading to a 2-3x enhancement in sample complexity and a 6.5x decrease in sampling time. Moreover, our approach consistently outperforms generic prompt-based steering methods and KNN retrieval baseline across a wide range of datasets. Additionally, we demonstrate the compatibility of our approach with widely-used data augmentation techniques, as well as the reliability of the generated data in supporting various neural architectures and enhancing few-shot learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

    cs.CR 2025-06 conditional novelty 7.0 of 10

    A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...

  2. Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

    cs.LG 2026-07 accept novelty 6.5 of 10

    Post-generation selection via Homogeneous-Heterogeneous real-data splits and a fidelity-diversity score raises synthetic-image utility for classification and segmentation without retraining generators.

  3. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Conditioning a fine-tuned text-to-image model on per-image background/pose captions and then randomly recombining those contexts across classes improves few-shot fine-grained classifier accuracy.

  4. Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DiffGRE generates synthetic images through cross-image interpolation in diffusion and CLIP latent spaces, filters them for diversity, and uses pseudo-labels to improve on-the-fly fine-grained category discovery.

Pith tools