REVIEW 3 cited by
BUT System Description to VoxCeleb Speaker Recognition Challenge 2019
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this report, we describe the submission of Brno University of Technology (BUT) team to the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2019. We also provide a brief analysis of different systems on VoxCeleb-1 test sets. Submitted systems for both Fixed and Open conditions are a fusion of 4 Convolutional Neural Network (CNN) topologies. The first and second networks have ResNet34 topology and use two-dimensional CNNs. The last two networks are one-dimensional CNN and are based on the x-vector extraction topology. Some of the networks are fine-tuned using additive margin angular softmax. Kaldi FBanks and Kaldi PLPs were used as features. The difference between Fixed and Open systems lies in the used training data and fusion strategy. The best systems for Fixed and Open conditions achieved 1.42% and 1.26% ERR on the challenge evaluation set respectively.
Forward citations
Cited by 3 Pith papers
-
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.
-
Learning Emotion-Invariant Speaker Representations for Speaker Verification
CopyPaste-based parallel data, cosine similarity loss, and energy-based emotion masking reduce speaker verification EER by 19.29% relative on the Dusha emotional speech corpus.
-
Towards Robust Uncertainty-Aware Speaker Modeling
Inter- and intra-speaker hardness in an uncertainty-aware softmax plus source-prior uncertainty calibration improves speaker verification reliability under domain shift.
Discussion (0). Continue with ORCID to comment.