Pith. sign in

REVIEW 3 cited by

Exploring Transformers for Large-Scale Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.09684 v2 pith:V4SKWMOI submitted 2020-05-19 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords transformersdatarecognitionspeecharoundbeencomparedcondition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While recurrent neural networks still largely define state-of-the-art speech recognition systems, the Transformer network has been proven to be a competitive alternative, especially in the offline condition. Most studies with Transformers have been constrained in a relatively small scale setting, and some forms of data argumentation approaches are usually applied to combat the data sparsity issue. In this paper, we aim at understanding the behaviors of Transformers in the large-scale speech recognition setting, where we have used around 65,000 hours of training data. We investigated various aspects on scaling up Transformers, including model initialization, warmup training as well as different Layer Normalization strategies. In the streaming condition, we compared the widely used attention mask based future context lookahead approach to the Transformer-XL network. From our experiments, we show that Transformers can achieve around 6% relative word error rate (WER) reduction compared to the BLSTM baseline in the offline fashion, while in the streaming fashion, Transformer-XL is comparable to LC-BLSTM with 800 millisecond latency constraint.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A multilevel approach to accelerate the training of Transformers

    cs.LG 2025-04 conditional novelty 5.0 of 10

    A multilevel scheme that alternates fine transformer training with two half-depth coarse models reaches the single-level training loss with 44 percent fewer FLOPs on one small language-model setup.

  2. Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning

    cs.AI 2025-04 conditional novelty 5.0 of 10

    A Conformer-based Arabic ASR trained from scratch on 15,000 hours of weak labels outperforms several open and closed-source models on standard Arabic benchmarks.

  3. Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    A new Badini Kurdish speech corpus and a comparison of Wav2Vec2 versus Whisper show Wav2Vec2 achieves 82.67% accuracy versus Whisper's 53.17%.

Pith tools