Pith. sign in

REVIEW 1 cited by

Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10680 v1 pith:OJ7IC6KM submitted 2024-08-20 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords languagesmodeloriginalwhisperforgettinglora-basedmodelsmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained multilingual speech foundation models, like Whisper, have shown impressive performance across different languages. However, adapting these models to new or specific languages is computationally extensive and faces catastrophic forgetting problems. Addressing these issues, our study investigates strategies to enhance the model on new languages in the absence of original training data, while also preserving the established performance on the original languages. Specifically, we first compare various LoRA-based methods to find out their vulnerability to forgetting. To mitigate this issue, we propose to leverage the LoRA parameters from the original model for approximate orthogonal gradient descent on the new samples. Additionally, we also introduce a learnable rank coefficient to allocate trainable parameters for more efficient training. Our experiments with a Chinese Whisper model (for Uyghur and Tibetan) yield better results with a more compact parameter set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...

Pith tools