Pith. sign in

REVIEW 1 cited by

End-to-End Speaker-Attributed ASR with Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.02128 v1 pith:FT4NLQD5 submitted 2021-04-05 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speakermodelspeaker-attributedspeecharchitecturecountingdatasetend-to-end
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identification for monaural multi-talker audio. Firstly, we thoroughly update the model architecture that was previously designed based on a long short-term memory (LSTM)-based attention encoder decoder by applying transformer architectures. Secondly, we propose a speaker deduplication mechanism to reduce speaker identification errors in highly overlapped regions. Experimental results on the LibriSpeechMix dataset shows that the transformer-based architecture is especially good at counting the speakers and that the proposed model reduces the speaker-attributed word error rate by 47% over the LSTM-based baseline. Furthermore, for the LibriCSS dataset, which consists of real recordings of overlapped speech, the proposed model achieves concatenated minimum-permutation word error rates of 11.9% and 16.3% with and without target speaker profiles, respectively, both of which are the state-of-the-art results for LibriCSS with the monaural setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR

    eess.AS 2025-06 conditional novelty 6.0 of 10

    A speaker-activity mask injected into an ASR encoder lets one model instance transcribe each talker in overlapped speech without speaker embeddings.

Pith tools