Pith. sign in

REVIEW 4 cited by

TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.04410 v1 pith:TWTBL3GC submitted 2021-10-08 eess.AS cs.SD

classification eess.AScs.SD
keywords speakertitanetdiarizationarchitecturecontextconvolutionsdepth-wiseerror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel attention based statistics pooling layer to map variable-length utterances to a fixed-length embedding (t-vector). TitaNet is a scalable architecture and achieves state-of-the-art performance on speaker verification task with an equal error rate (EER) of 0.68% on the VoxCeleb1 trial file and also on speaker diarization tasks with diarization error rate (DER) of 1.73% on AMI-MixHeadset, 1.99% on AMI-Lapel and 1.11% on CH109. Furthermore, we investigate various sizes of TitaNet and present a light TitaNet-S model with only 6M parameters that achieve near state-of-the-art results in diarization tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Introduces CoMuMDR, a code-mixed, multi-modal, multi-domain discourse corpus for conversations, with benchmarks showing poor relation classification by current models.

  2. CASPER: A Large Scale Spontaneous Speech Dataset

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CASPER presents a 102-hour spontaneous English conversation dataset with per-speaker channels, speaker metadata, and baseline ASR and diarization results.

  3. Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

    cs.CL 2025-06 reject novelty 4.0 of 10

    A T5-based text-only diarization model reports WDER 4.9 on short dialogues, beating audio baselines, but the result hinges on an uncontrolled in-domain versus out-of-domain comparison.

  4. Sparse Autoencoder Insights on Voice Embeddings

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Sparse autoencoders trained on Titanet speaker embeddings yield latent units that identify and steer language and music features, replicating LLM-style feature splitting and steering in audio data.

Pith tools