REVIEW 5 cited by
Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report describes our submission to the VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020. We perform a careful analysis of speaker recognition models based on the popular ResNet architecture, and train a number of variants using a range of loss functions. Our results show significant improvements over most existing works without the use of model ensemble or post-processing. We release the training code and pre-trained models as unofficial baselines for this year's challenge.
Forward citations
Cited by 5 Pith papers
-
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
ASVspoof 5 provides a large, speaker-diverse, crowdsourced benchmark for speech spoofing and deepfake detection, including 32 attacks and a post-processing pipeline to reduce shortcut artifacts.
-
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
NanoCodec achieves competitive speech quality at 12.5 frames per second and 0.6-1.78 kbps, with a causal decoder for low-latency speech LLM inference.
-
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.
-
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.
Discussion (0). Continue with ORCID to comment.