Pith. sign in

REVIEW 5 cited by

WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.01547 v5 pith:EPOWYISK submitted 2021-02-02 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords wenetmodelnon-streamingproductionrecognitionspeechattentiontoolkit
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify streaming and non-streaming end-to-end (E2E) speech recognition in a single model. The main motivation of WeNet is to close the gap between the research and the production of E2E speechrecognition models. WeNet provides an efficient way to ship ASR applications in several real-world scenarios, which is the main difference and advantage to other open source E2E speech recognition toolkits. In our toolkit, a new two-pass method is implemented. Our method propose a dynamic chunk-based attention strategy of the the transformer layers to allow arbitrary right context length modifies in hybrid CTC/attention architecture. The inference latency could be easily controlled by only changing the chunk size. The CTC hypotheses are then rescored by the attention decoder to get the final result. Our experiments on the AISHELL-1 dataset using WeNet show that, our model achieves 5.03\% relative character error rate (CER) reduction in non-streaming ASR compared to a standard non-streaming transformer. After model quantification, our model perform reasonable RTF and latency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition

    cs.SD 2025-06 conditional novelty 6.0 of 10

    An iterative segmentation-based self-training method for Whisper improved long dysarthric speech recognition and achieved second place in both WER and SemScore at the SAP Challenge.

  2. TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

    eess.AS 2024-12 conditional novelty 6.0 of 10

    TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.

  3. When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A cascaded speech-to-text translation model that feeds five aligned ASR candidates and self-supervised speech units to a translation model matches end-to-end performance on GigaST, with an English-to-Chinese BLEU of 38.1.

  4. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

  5. Complexity boosted adaptive training for better low resource ASR performance

    cs.SD 2024-12 conditional novelty 4.0 of 10

    CBA training, a two-stage adaptive ASR training scheme using a MinMax-IBF sample-complexity policy, improves WER/CER over WeNet Conformer with SpecAugment on LibriSpeech 100h and AISHELL-1.

Pith tools