REVIEW 3 cited by
TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel attention based statistics pooling layer to map variable-length utterances to a fixed-length embedding (t-vector). TitaNet is a scalable architecture and achieves state-of-the-art performance on speaker verification task with an equal error rate (EER) of 0.68% on the VoxCeleb1 trial file and also on speaker diarization tasks with diarization error rate (DER) of 1.73% on AMI-MixHeadset, 1.99% on AMI-Lapel and 1.11% on CH109. Furthermore, we investigate various sizes of TitaNet and present a light TitaNet-S model with only 6M parameters that achieve near state-of-the-art results in diarization tasks.
Forward citations
Cited by 3 Pith papers
-
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
Introduces CoMuMDR, a code-mixed, multi-modal, multi-domain discourse corpus for conversations, with benchmarks showing poor relation classification by current models.
-
CASPER: A Large Scale Spontaneous Speech Dataset
CASPER presents a 102-hour spontaneous English conversation dataset with per-speaker channels, speaker metadata, and baseline ASR and diarization results.
-
Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models
A T5-based text-only diarization model reports WDER 4.9 on short dialogues, beating audio baselines, but the result hinges on an uncontrolled in-domain versus out-of-domain comparison.
Discussion (0). Continue with ORCID to comment.