Pith. sign in

REVIEW 1 cited by

Raw Differentiable Architecture Search for Speech Deepfake and Spoofing Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.12212 v2 pith:UEQXILX2 submitted 2021-07-26 eess.AS

classification eess.AS
keywords architecturenetworkdetectionapproachesdeepfakedifferentiableparameterssearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end approaches to anti-spoofing, especially those which operate directly upon the raw signal, are starting to be competitive with their more traditional counterparts. Until recently, all such approaches consider only the learning of network parameters; the network architecture is still hand crafted. This too, however, can also be learned. Described in this paper is our attempt to learn automatically the network architecture of a speech deepfake and spoofing detection solution, while jointly optimising other network components and parameters, such as the first convolutional layer which operates on raw signal inputs. The resulting raw differentiable architecture search system delivers a tandem detection cost function score of 0.0517 for the ASVspoof 2019 logical access database, a result which is among the best single-system results reported to date.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Fusing CQCC spectral features with Wav2Vec2.0 embeddings via cross-attention lowers average equal error rate from 10.87% to 6.80% across four speech deepfake benchmarks.

Pith tools