REVIEW 2 cited by
End-to-End Neural Speaker Diarization with Permutation-Free Objectives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations. Instead, our model has a single neural network that directly outputs speaker diarization results. To realize such a model, we formulate the speaker diarization problem as a multi-label classification problem, and introduces a permutation-free objective function to directly minimize diarization errors without being suffered from the speaker-label permutation problem. Besides its end-to-end simplicity, the proposed method also benefits from being able to explicitly handle overlapping speech during training and inference. Because of the benefit, our model can be easily trained/adapted with real-recorded multi-speaker conversations just by feeding the corresponding multi-speaker segment labels. We evaluated the proposed method on simulated speech mixtures. The proposed method achieved diarization error rate of 12.28%, while a conventional clustering-based system produced diarization error rate of 28.77%. Furthermore, the domain adaptation with real-recorded speech provided 25.6% relative improvement on the CALLHOME dataset. Our source code is available online at https://github.com/hitachi-speech/EEND.
Forward citations
Cited by 2 Pith papers
-
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
A lip-sync speaker-ID plus distillation and hierarchical-fusion model reports state-of-the-art weighted F1 on MELD (67.75%) and IEMOCAP (72.44%).
-
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Cosine similarity between t-vector speaker embeddings in a streaming transducer speech translation model detects speaker changes (F1 up to 0.68) and classifies gender (0.989 accuracy).
Discussion (0). Continue with ORCID to comment.