Pith. sign in

REVIEW 3 cited by

Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00842 v5 pith:DKVJK264 submitted 2022-02-02 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords modelmulti-talkeroutputstreamingt-sotcostmodelsmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a token-level serialized output training (t-SOT), a novel framework for streaming multi-talker automatic speech recognition (ASR). Unlike existing streaming multi-talker ASR models using multiple output branches, the t-SOT model has only a single output branch that generates recognition tokens (e.g., words, subwords) of multiple speakers in chronological order based on their emission times. A special token that indicates the change of ``virtual'' output channels is introduced to keep track of the overlapping utterances. Compared to the prior streaming multi-talker ASR models, the t-SOT model has the advantages of less inference cost and a simpler model architecture. Moreover, in our experiments with LibriSpeechMix and LibriCSS datasets, the t-SOT-based transformer transducer model achieves the state-of-the-art word error rates by a significant margin to the prior results. For non-overlapping speech, the t-SOT model is on par with a single-talker ASR model in terms of both accuracy and computational cost, opening the door for deploying one model for both single- and multi-talker scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Error Analysis in a Modular Meeting Transcription System

    eess.AS 2025-09 conditional novelty 5.0 of 10

    In a modular meeting transcription pipeline, missing speech segments, not primary-to-cross-channel leakage, cause most of the gap to oracle segmentation.

  2. SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.

  3. Joint ASR and Speaker Role Tagging with Serialized Output Training

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Fine-tuning Whisper with serialized output training and role-specific tokens produces role-aware transcripts in one pass, cutting multi-talker word error rate by 10 to 40 percent versus a WavLM CTC baseline.

Pith tools