Pith. sign in

REVIEW 19 cited by

AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02781 v2 pith:LLQHMVSO submitted 2019-12-05 stat.ML cs.CVcs.LG

classification stat.MLcs.CVcs.LG
keywords robustnessaugmixdataimproveuncertaintyaccuracydistributionimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern deep neural networks can achieve high accuracy when the training distribution and test distribution are identically distributed, but this assumption is frequently violated in practice. When the train and test distributions are mismatched, accuracy can plummet. Currently there are few techniques that improve robustness to unforeseen data shifts encountered during deployment. In this work, we propose a technique to improve the robustness and uncertainty estimates of image classifiers. We propose AugMix, a data processing technique that is simple to implement, adds limited computational overhead, and helps models withstand unforeseen corruptions. AugMix significantly improves robustness and uncertainty measures on challenging image classification benchmarks, closing the gap between previous methods and the best possible performance in some cases by more than half.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 567 citations worldwide. Full citation record

  1. ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.

  2. On the Reliability of Cue Conflict and Beyond

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.

  3. Frequency Prior Guided Matching: A Data Augmentation Approach for Generalizable Semi-Supervised Polyp Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FPGM learns a frequency prior from labeled polyp edges and aligns unlabeled image spectra to it, improving semi-supervised polyp segmentation and zero-shot generalization on six datasets.

  4. Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.

  5. Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SynOOD generates synthetic near-boundary OOD images with MLLM-guided inpainting and energy-score gradients, then fine-tunes CLIP image and text features, reporting state-of-the-art OOD detection on ImageNet benchmarks.

  6. MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MUG combines manually corrected pseudo-labels, cross-modal random track recombination, and a Mamba-Transformer network to reach new state-of-the-art F1 scores on the LLP audio-visual video parsing benchmark.

  7. Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.

  8. ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Reorganizing facial video into four channel-concatenated quadrants before tokenization yields 56.00% test accuracy on AI4Pain video-only pain classification, the highest reported under that benchmark protocol.

  9. Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Interleaving clean and noisy training epochs improves clean, corrupted, and out-of-distribution accuracy on CIFAR-100 and ImageNet for CNNs and ViTs, with impulse noise best for ResNets and Gaussian noise best for ViTs.

  10. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

  11. Data-Augmented Quantization-Aware Knowledge Distillation

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A teacher-only metric, M=DEV-CMI, ranks data augmentations for low-bit quantized knowledge distillation and improves accuracy on CIFAR and Tiny ImageNet.

  12. DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The abstract and body describe different papers; the body proposes TriReWeight, a re-weighting wrapper for generative data augmentation claimed to add 2.9 to 7.9 accuracy points in small-dataset classification.

  13. Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    KR-NFT tunes CLIP text features with image-conditioned scaling and shifting plus a knowledge regularization loss, improving OOD detection on base and unseen classes without forgetting pre-trained knowledge.

  14. Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided Optimization

    cs.CV 2025-07 reject novelty 5.0 of 10

    LEAwareSGD modulates the learning rate with a Lyapunov exponent estimate and claims state-of-the-art accuracy on three single-domain generalization benchmarks, but the method is underspecified and its hyperparameters ...

  15. A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A unified tokenizer maps facial video and fNIRS into one token space; the segment-latent transformer hits 57.33% test accuracy on AI4Pain pain recognition.

  16. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  17. Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models

    cs.CV 2025-09 conditional novelty 4.0 of 10

    Diffusion-based synthetic expansion of pedestrian attribute training sets yields 1.1 to 3.3 point mA gains over the base PAR model on PA100k, PETAzs, and RAPzs.

  18. PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

    cs.CV 2025-08 reject novelty 4.0 of 10

    PQ-DAF uses pose-conditioned diffusion generation plus CogVLM filtering to augment few-shot driver distraction training data, and reports large accuracy gains that are compromised by a non-standard train/test protocol.

  19. Spatial RoboGrasp: Generalized Robotic Grasping Control Policy

    cs.RO 2025-05 conditional novelty 4.0 of 10

    Spatial RoboGrasp combines AugFusion, monocular depth, and grasp prompts in a diffusion policy, claiming large gains under exposure change, without released artifacts or error bars.

Pith tools