Pith. sign in

REVIEW 3 cited by

OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16658 v3 pith:FI5VOSIF submitted 2024-01-30 cs.CL eess.AS

classification cs.CLeess.AS
keywords owsmmodelsspeechdataopenmodelperformancewhisper-style
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits. However, previous versions of OWSM (v1 to v3) are still based on standard Transformer, which might lead to inferior performance compared to state-of-the-art speech encoder architectures. This work aims to improve the performance and efficiency of OWSM without additional data. We present a series of E-Branchformer-based models named OWSM v3.1, ranging from 100M to 1B parameters. OWSM v3.1 outperforms its predecessor, OWSM v3, in most evaluation benchmarks, while showing an improved inference speed of up to 25%. We further reveal the emergent ability of OWSM v3.1 in zero-shot contextual biasing speech recognition. We also provide a model trained on a subset of data with low license restrictions. We will publicly release the code, pre-trained models, and training logs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.

  2. ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...

  3. Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.

Pith tools