REVIEW 8 cited by
How Does Mixup Help With Robustness and Generalization?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mixup is a popular data augmentation technique based on taking convex combinations of pairs of examples and their labels. This simple technique has been shown to substantially improve both the robustness and the generalization of the trained model. However, it is not well-understood why such improvement occurs. In this paper, we provide theoretical analysis to demonstrate how using Mixup in training helps model robustness and generalization. For robustness, we show that minimizing the Mixup loss corresponds to approximately minimizing an upper bound of the adversarial loss. This explains why models obtained by Mixup training exhibits robustness to several kinds of adversarial attacks such as Fast Gradient Sign Method (FGSM). For generalization, we prove that Mixup augmentation corresponds to a specific type of data-adaptive regularization which reduces overfitting. Our analysis provides new insights and a framework to understand Mixup.
Forward citations
Cited by 8 Pith papers
-
Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation
The test error of random-feature ridge regression with arbitrary data augmentation admits a closed-form asymptotic characterization in the proportional regime that depends only on population covariances and augmentati...
-
CrackForward: Context-Aware Severity Stage Crack Synthesis for Data Augmentation
A context-aware generative model synthesizes crack growth patterns with directional propagation and learned morphology to augment training data and improve crack segmentation performance.
-
Bringing Balance to Hand Shape Classification: Mitigating Data Imbalance Through Generative Models
Pre-training an EfficientNet classifier on GAN-generated balanced hand images, then fine-tuning on real data, raises accuracy on the imbalanced RWTH handshape benchmark from 80.6% to 85.3%.
-
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Introduces a margin-adaptive confidence ranking method that learns an estimator from simulated diversity and derives margin-dependent generalization bounds for use in fixed-sequence testing of LLM-human agreement.
-
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Learning a margin-based confidence ranker for LLM judges improves agreement-target success in cascaded selective evaluation compared to heuristic confidence scores.
-
Medical Model Synthesis Architectures: A Case Study
MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.
-
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
Develops a margin-adaptive learned confidence estimator for LLMs with generalization guarantees to improve agreement rates with human judgments over heuristic baselines.
-
Why Pool When You Can Flow? Active Learning with GFlowNets
Training a GFlowNet to generate high-BALD molecules gives pool-size-independent acquisition and near-BALD classification quality on JAK2 virtual screening.
Discussion (0). Sign in to comment.