Pith. sign in

REVIEW 3 cited by

BUT System Description to VoxCeleb Speaker Recognition Challenge 2019

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12592 v1 pith:VH6QY2IS submitted 2019-10-16 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords systemschallengefixednetworksopenconditionsfusionkaldi
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we describe the submission of Brno University of Technology (BUT) team to the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2019. We also provide a brief analysis of different systems on VoxCeleb-1 test sets. Submitted systems for both Fixed and Open conditions are a fusion of 4 Convolutional Neural Network (CNN) topologies. The first and second networks have ResNet34 topology and use two-dimensional CNNs. The last two networks are one-dimensional CNN and are based on the x-vector extraction topology. Some of the networks are fine-tuned using additive margin angular softmax. Kaldi FBanks and Kaldi PLPs were used as features. The difference between Fixed and Open systems lies in the used training data and fusion strategy. The best systems for Fixed and Open conditions achieved 1.42% and 1.26% ERR on the challenge evaluation set respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.

  2. Learning Emotion-Invariant Speaker Representations for Speaker Verification

    cs.SD 2025-05 conditional novelty 5.0 of 10

    CopyPaste-based parallel data, cosine similarity loss, and energy-based emotion masking reduce speaker verification EER by 19.29% relative on the Dusha emotional speech corpus.

  3. Towards Robust Uncertainty-Aware Speaker Modeling

    cs.SD 2026-07 conditional novelty 4.0 of 10

    Inter- and intra-speaker hardness in an uncertainty-aware softmax plus source-prior uncertainty calibration improves speaker verification reliability under domain shift.

Pith tools