Pith. sign in

REVIEW 2 cited by

Pushing the limits of raw waveform speaker recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.08488 v2 pith:4YWPLQHM submitted 2022-03-16 eess.AS cs.AI

classification eess.AScs.AI
keywords modelspeakerwaveforminputsrecognitionself-supervisedbestequal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention. However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-based counterparts, which demonstrate equal error rates under 1% on the popular VoxCeleb1 test set. This paper proposes a novel speaker recognition model based on raw waveform inputs. The model incorporates recent advances in machine learning and speaker verification, including the Res2Net backbone module and multi-layer feature aggregation. Our best model achieves an equal error rate of 0.89%, which is competitive with the state-of-the-art models based on handcrafted features, and outperforms the best model based on raw waveform inputs by a large margin. We also explore the application of the proposed model in the context of self-supervised learning framework. Our self-supervised model outperforms single phase-based existing works in this line of research. Finally, we show that self-supervised pre-training is effective for the semi-supervised scenario where we only have a small set of labelled training data, along with a larger set of unlabelled examples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RED: Robust Environmental Design

    cs.CV 2024-11 conditional novelty 6.0 of 10

    RED learns per-class colorful grid patterns for road sign backgrounds that make any small patch class-discriminative, sharply reducing vulnerability to patch attacks.

  2. Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A large-scale benchmark shows twelve audio deepfake detectors are all substantially vulnerable to statistical and optimization-based adversarial attacks.

Pith tools