Pith. sign in

REVIEW 2 cited by

Sample adaptive data augmentation with progressive scheduling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00415 v1 pith:X44ULU5T submitted 2024-11-30 cs.SD eess.AS

classification cs.SDeess.AS
keywords augmentationdatatrainingfixedmethodstrategyaishell-1approach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data augmentation is a widely adopted technique utilized to improve the robustness of automatic speech recognition (ASR). Employing a fixed data augmentation strategy for all training data is a common practice. However, it is important to note that there can be variations in factors such as background noise, speech rate, etc. among different samples within a single training batch. By using a fixed augmentation strategy, there is a risk that the model may reach a suboptimal state. In addition to the risks of employing a fixed augmentation strategy, the model's capabilities may differ across various training stages. To address these issues, this paper proposes the method of sample-adaptive data augmentation with progressive scheduling(PS-SapAug). The proposed method applies dynamic data augmentation in a two-stage training approach. It employs hybrid normalization to compute sample-specific augmentation parameters based on each sample's loss. Additionally, the probability of augmentation gradually increases throughout the training progression. Our method is evaluated on popular ASR benchmark datasets, including Aishell-1 and Librispeech-100h, achieving up to 8.13% WER reduction on LibriSpeech-100h test-clean, 6.23% on test-other, and 5.26% on AISHELL-1 test set, which demonstrate the efficacy of our approach enhancing performance and minimizing errors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Replacing square spectrogram patches with full-frequency temporal patches plus patch-aligned masking improves audio classification accuracy and reduces compute for Transformer and Mamba models.

  2. Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Acoustic-focused augmentation of a 960-hour dataset is reported to reduce out-of-distribution word error rates by up to 19.24 percent, suggesting acoustic diversity, not linguistic diversity, drives ASR robustness.

Pith tools