Pith. sign in

REVIEW 7 cited by

How Does Mixup Help With Robustness and Generalization?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.04819 v4 pith:VMKXUTU2 submitted 2020-10-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords mixuprobustnessgeneralizationadversarialanalysisaugmentationcorrespondsloss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mixup is a popular data augmentation technique based on taking convex combinations of pairs of examples and their labels. This simple technique has been shown to substantially improve both the robustness and the generalization of the trained model. However, it is not well-understood why such improvement occurs. In this paper, we provide theoretical analysis to demonstrate how using Mixup in training helps model robustness and generalization. For robustness, we show that minimizing the Mixup loss corresponds to approximately minimizing an upper bound of the adversarial loss. This explains why models obtained by Mixup training exhibits robustness to several kinds of adversarial attacks such as Fast Gradient Sign Method (FGSM). For generalization, we prove that Mixup augmentation corresponds to a specific type of data-adaptive regularization which reduces overfitting. Our analysis provides new insights and a framework to understand Mixup.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation

    stat.ML 2026-05 conditional novelty 7.0 of 10

    The test error of random-feature ridge regression with arbitrary data augmentation admits a closed-form asymptotic characterization in the proportional regime that depends only on population covariances and augmentati...

  2. CrackForward: Context-Aware Severity Stage Crack Synthesis for Data Augmentation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    A context-aware generative model synthesizes crack growth patterns with directional propagation and learned morphology to augment training data and improve crack segmentation performance.

  3. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Introduces a margin-adaptive confidence ranking method that learns an estimator from simulated diversity and derives margin-dependent generalization bounds for use in fixed-sequence testing of LLM-human agreement.

  4. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

    cs.LG 2026-05 conditional novelty 5.0 of 10

    Learning a margin-based confidence ranker for LLM judges improves agreement-target success in cascaded selective evaluation compared to heuristic confidence scores.

  5. Medical Model Synthesis Architectures: A Case Study

    cs.AI 2026-05 unverdicted novelty 5.0 of 10

    MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.

  6. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    Develops a margin-adaptive learned confidence estimator for LLMs with generalization guarantees to improve agreement rates with human judgments over heuristic baselines.

  7. Why Pool When You Can Flow? Active Learning with GFlowNets

    cs.LG 2025-08 conditional novelty 3.0 of 10

    Training a GFlowNet to generate high-BALD molecules gives pool-size-independent acquisition and near-BALD classification quality on JAK2 virtual screening.

Pith tools