Pith. sign in

REVIEW 3 cited by

Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17382 v1 pith:YABA3ISH submitted 2023-11-29 eess.AS

classification eess.AS
keywords whisperadaptationadaptingdataresultscode-switchhoursperformance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This paper details the experimental results of adapting the OpenAI's Whisper model for Code-Switch Mandarin-English Speech Recognition (ASR) on the SEAME and ASRU2019 corpora. We conducted 2 experiments: a) using adaptation data from 1 to 100/200 hours to demonstrate effectiveness of adaptation, b) examining different language ID setup on Whisper prompt. The Mixed Error Rate results show that the amount of adaptation data may be as low as $1\sim10$ hours to achieve saturation in performance gain (SEAME) while the ASRU task continued to show performance with more adaptation data ($>$100 hours). For the language prompt, the results show that although various prompting strategies initially produce different outcomes, adapting the Whisper model with code-switch data uniformly improves its performance. These results may be relevant also to the community when applying Whisper for related tasks of adapting to new target domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR

    cs.CL 2025-06 conditional novelty 6.0 of 10

    AsyncSwitch improves code-switched ASR on Whisper by adapting the decoder on text before speech-text alignment and full fine-tuning.

  2. CAFE A Novel Code switching Dataset for Algerian Dialect French and English

    cs.SD 2024-11 conditional novelty 6.0 of 10

    CAFE is a new spontaneous speech corpus for Algerian dialect, French, and English code-switching, with 2.6 hours manually annotated and a Whisper benchmark reaching MER 0.310.

  3. Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

    cs.CL 2024-12 conditional novelty 5.0 of 10

    An LSTM-based encoder refiner plus language-aware dual adapters with a fusion module cuts Mandarin-English code-switching ASR errors on SEAME.

Pith tools