REVIEW 13 cited by
Effective Data Augmentation With Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Data augmentation is one of the most prevalent tools in deep learning, underpinning many recent advances, including those from classification, generative models, and representation learning. The standard approach to data augmentation combines simple transformations like rotations and flips to generate new images from existing ones. However, these new images lack diversity along key semantic axes present in the data. Current augmentations cannot alter the high-level semantic attributes, such as animal species present in a scene, to enhance the diversity of data. We address the lack of diversity in data augmentation with image-to-image transformations parameterized by pre-trained text-to-image diffusion models. Our method edits images to change their semantics using an off-the-shelf diffusion model, and generalizes to novel visual concepts from a few labelled examples. We evaluate our approach on few-shot image classification tasks, and on a real-world weed recognition task, and observe an improvement in accuracy in tested domains.
Forward citations
Cited by 13 Pith papers
-
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...
-
Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
A leak-free stacking protocol and mask-conditioned synthetic generation improve sand-boil segmentation to 0.707 IoU, but stacking underperforms the best single model and synthetic gains come only from label-fidelity f...
-
ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.
-
Reading a Ruler in the Wild
RulerNet detects centimeter marks on rulers in natural images and estimates pixel-per-centimeter scale via a learned geometric progression model, outperforming prior OCR-based and frequency-based baselines.
-
Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery
DiffGRE generates synthetic images through cross-image interpolation in diffusion and CLIP latent spaces, filters them for diversity, and uses pseudo-labels to improve on-the-fly fine-grained category discovery.
-
Inpainting is All You Need: A Diffusion-based Augmentation Method for Semi-supervised Medical Image Segmentation
A diffusion-based inpainting pipeline generates synthetic image-label pairs for medical segmentation and improves dice scores under limited labels on four datasets.
-
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
SITN improves cross-domain few-shot detection and segmentation by using weakened-noise diffusion and background inpainting to synthesize helpful training images, outperforming prior methods on all reported benchmarks.
-
QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis
Diffusion-generated video/audio samples weighted by a learned quality scorer are claimed to improve multimodal sentiment analysis on CH-SIMS, CMU-MOSI, and MUStARD.
-
Multimodal Forecasting of Sparse Intraoperative Hypotension Events Powered by Language Model
IOHFuseLM combines patient attributes and MAP waveforms through token-level alignment and diffusion-augmented pretraining, improving IOH event detection on two datasets over six baselines.
-
VisAlgae 2023: A Dataset and Challenge for Algae Detection in Microscopy Images
The VisAlgae 2023 dataset and challenge provide a new public benchmark for detecting six microalgae species in microscopy images, with baseline and top-10 leaderboard results.
-
Generative Data Augmentation for Object Point Cloud Segmentation
A mask-conditioned diffusion model generates labeled point cloud variants and filtered pseudo-labels, improving object part segmentation with only 10% hand labels.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models
Diffusion-based synthetic expansion of pedestrian attribute training sets yields 1.1 to 3.3 point mA gains over the base PAR model on PA100k, PETAzs, and RAPzs.
Discussion (0). Continue with ORCID to comment.