Pith. sign in

REVIEW 1 cited by

Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08568 v1 pith:XD36XI4V submitted 2024-06-12 cs.SD eess.AS

classification cs.SDeess.AS
keywords datadysarthricspeechaugmentationdasrmodelssynthesistraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic speech recognition (ASR) research has achieved impressive performance in recent years and has significant potential for enabling access for people with dysarthria (PwD) in augmentative and alternative communication (AAC) and home environment systems. However, progress in dysarthric ASR (DASR) has been limited by high variability in dysarthric speech and limited public availability of dysarthric training data. This paper demonstrates that data augmentation using text-to-dysarthic-speech (TTDS) synthesis for finetuning large ASR models is effective for DASR. Specifically, diffusion-based text-to-speech (TTS) models can produce speech samples similar to dysarthric speech that can be used as additional training data for fine-tuning ASR foundation models, in this case Whisper. Results show improved synthesis metrics and ASR performance for the proposed multi-speaker diffusion-based TTDS data augmentation for ASR fine-tuning compared to current DASR baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Acoustic-focused augmentation of a 960-hour dataset is reported to reduce out-of-distribution word error rates by up to 19.24 percent, suggesting acoustic diversity, not linguistic diversity, drives ASR robustness.

Pith tools