Pith. sign in

REVIEW 5 cited by

Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.14153 v1 pith:YRSD6JX2 submitted 2020-09-29 eess.AS cs.SD

classification eess.AScs.SD
keywords challengerecognitionspeakermodelsvoxcelebanalysisarchitecturebaseline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report describes our submission to the VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020. We perform a careful analysis of speaker recognition models based on the popular ResNet architecture, and train a number of variants using a range of loss functions. Our results show significant improvements over most existing works without the use of model ensemble or post-processing. We release the training code and pre-trained models as unofficial baselines for this year's challenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech

    eess.AS 2025-02 accept novelty 7.0 of 10

    ASVspoof 5 provides a large, speaker-diverse, crowdsourced benchmark for speech spoofing and deepfake detection, including 32 attacks and a post-processing pipeline to reduce shortcut artifacts.

  2. NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

    eess.AS 2025-08 conditional novelty 5.0 of 10

    NanoCodec achieves competitive speech quality at 12.5 frames per second and 0.6-1.78 kbps, with a causal decoder for low-latency speech LLM inference.

  3. SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

    eess.AS 2025-07 conditional novelty 5.0 of 10

    SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.

  4. An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS

    eess.AS 2025-06 conditional novelty 5.0 of 10

    In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.

  5. RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.

Pith tools