REVIEW 3 cited by
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits. However, previous versions of OWSM (v1 to v3) are still based on standard Transformer, which might lead to inferior performance compared to state-of-the-art speech encoder architectures. This work aims to improve the performance and efficiency of OWSM without additional data. We present a series of E-Branchformer-based models named OWSM v3.1, ranging from 100M to 1B parameters. OWSM v3.1 outperforms its predecessor, OWSM v3, in most evaluation benchmarks, while showing an improved inference speed of up to 25%. We further reveal the emergent ability of OWSM v3.1 in zero-shot contextual biasing speech recognition. We also provide a model trained on a subset of data with low license restrictions. We will publicly release the code, pre-trained models, and training logs.
Forward citations
Cited by 3 Pith papers
-
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
A new open suite of 13 multilingual speech models, up to 18B parameters, yields empirical scaling laws for ASR and speech translation performance.
-
ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition
A three-stage iterative LoRA training recipe (Focus, Feed Back, Fix) is applied to Whisper-large-v3 and Qwen2-Audio, reporting WER reductions on a multilingual ASR benchmark, with the gains attributed to the iterative...
-
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.
Discussion (0). Continue with ORCID to comment.