Pith. sign in

REVIEW 2 cited by

RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.08104 v2 pith:A2US2WVC submitted 2019-04-17 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords speakerdeepneuralwaveformsback-endclassificationembeddingend-to-end
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains. In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation. In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification. Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance. Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis. Effective back-end classification models that suit the proposed speaker embedding are also explored. We propose an end-to-end system that comprises two deep neural networks, one front-end for utterance-level speaker embedding extraction and the other for back-end classification. Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation. The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InterGridNet: An Electric Network Frequency Approach for Audio Source Location Classification Using Convolutional Neural Networks

    cs.SD 2025-02 conditional novelty 6.0 of 10

    A shallow RawNet-style network with neural architecture search classifies ENF-containing recordings into nine power grids plus None, reaching 92 percent on SP Cup 2016 but trailing the authors' own prior 96 percent fu...

  2. Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

    eess.AS 2025-07 reject novelty 5.0 of 10

    A joint training loop makes an F5-TTS model produce audio that a paired wav2vec 2.0/LCNN discriminator can recognize, enabling watermark-free attribution; however, the reported generalization gain is not isolated from...

Pith tools