Pith. sign in

REVIEW 1 cited by

Reweighting Augmented Samples by Minimizing the Maximal Expected Loss

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.08933 v1 pith:ZYRURGMC submitted 2021-03-16 cs.LG

classification cs.LG
keywords lossaugmentationaugmentedsamplesdataexpectedmaximalmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data augmentation is an effective technique to improve the generalization of deep neural networks. However, previous data augmentation methods usually treat the augmented samples equally without considering their individual impacts on the model. To address this, for the augmented samples from the same training example, we propose to assign different weights to them. We construct the maximal expected loss which is the supremum over any reweighted loss on augmented samples. Inspired by adversarial training, we minimize this maximal expected loss (MMEL) and obtain a simple and interpretable closed-form solution: more attention should be paid to augmented samples with large loss values (i.e., harder examples). Minimizing this maximal expected loss enables the model to perform well under any reweighting strategy. The proposed method can generally be applied on top of any data augmentation methods. Experiments are conducted on both natural language understanding tasks with token-level data augmentation, and image classification tasks with commonly-used image augmentation techniques like random crop and horizontal flip. Empirical results show that the proposed method improves the generalization performance of the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A fully online, loss-based reweighting scheme that down-weights low-loss samples during LLM pretraining yields small average benchmark gains at 1.4B and 7B scale, together with a convergence bound under convexity and ...

Pith tools