REVIEW 19 cited by
AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern deep neural networks can achieve high accuracy when the training distribution and test distribution are identically distributed, but this assumption is frequently violated in practice. When the train and test distributions are mismatched, accuracy can plummet. Currently there are few techniques that improve robustness to unforeseen data shifts encountered during deployment. In this work, we propose a technique to improve the robustness and uncertainty estimates of image classifiers. We propose AugMix, a data processing technique that is simple to implement, adds limited computational overhead, and helps models withstand unforeseen corruptions. AugMix significantly improves robustness and uncertainty measures on challenging image classification benchmarks, closing the gap between previous methods and the best possible performance in some cases by more than half.
Forward citations
Cited by 19 Pith papers
-
ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.
-
On the Reliability of Cue Conflict and Beyond
Stylized cue-conflict bias scores are confounded by impure cues, imbalance, ratio metrics and restricted labels; REFINED-BIAS supplies pure balanced stimuli and full-label MRR sensitivity for reliable diagnosis.
-
Frequency Prior Guided Matching: A Data Augmentation Approach for Generalizable Semi-Supervised Polyp Segmentation
FPGM learns a frequency prior from labeled polyp edges and aligns unlabeled image spectra to it, improving semi-supervised polyp segmentation and zero-shot generalization on six datasets.
-
Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.
-
Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
SynOOD generates synthetic near-boundary OOD images with MLLM-guided inpainting and energy-score gradients, then fine-tunes CLIP image and text features, reporting state-of-the-art OOD detection on ImageNet benchmarks.
-
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
MUG combines manually corrected pseudo-labels, cross-modal random track recombination, and a Mamba-Transformer network to reach new state-of-the-art F1 scores on the LLP audio-visual video parsing benchmark.
-
Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.
-
ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment
Reorganizing facial video into four channel-concatenated quadrants before tokenization yields 56.00% test accuracy on AI4Pain video-only pain classification, the highest reported under that benchmark protocol.
-
Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance
Interleaving clean and noisy training epochs improves clean, corrupted, and out-of-distribution accuracy on CIFAR-100 and ImageNet for CNNs and ViTs, with impulse noise best for ResNets and Gaussian noise best for ViTs.
-
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.
-
Data-Augmented Quantization-Aware Knowledge Distillation
A teacher-only metric, M=DEV-CMI, ranks data augmentations for low-bit quantized knowledge distillation and improves accuracy on CIFAR and Tiny ImageNet.
-
DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
The abstract and body describe different papers; the body proposes TriReWeight, a re-weighting wrapper for generative data augmentation claimed to add 2.9 to 7.9 accuracy points in small-dataset classification.
-
Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
KR-NFT tunes CLIP text features with image-conditioned scaling and shifting plus a knowledge regularization loss, improving OOD detection on base and unseen classes without forgetting pre-trained knowledge.
-
Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided Optimization
LEAwareSGD modulates the learning rate with a Lyapunov exponent estimate and claims state-of-the-art accuracy on three single-domain generalization benchmarks, but the method is underspecified and its hyperparameters ...
-
A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities
A unified tokenizer maps facial video and fNIRS into one token space; the segment-latent transformer hits 57.33% test accuracy on AI4Pain pain recognition.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
-
Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models
Diffusion-based synthetic expansion of pedestrian attribute training sets yields 1.1 to 3.3 point mA gains over the base PAR model on PA100k, PETAzs, and RAPzs.
-
PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
PQ-DAF uses pose-conditioned diffusion generation plus CogVLM filtering to augment few-shot driver distraction training data, and reports large accuracy gains that are compromised by a non-standard train/test protocol.
-
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
Spatial RoboGrasp combines AugFusion, monocular depth, and grasp prompts in a diffusion policy, claiming large gains under exposure change, without released artifacts or error bars.
Discussion (0). Sign in to comment.