Pith. sign in

REVIEW 9 cited by

CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00332 v3 pith:XM2MVCYJ submitted 2023-03-01 cs.SD eess.AS

classification cs.SDeess.AS
keywords networkcomputationalefficientinferencespeakerverificationarchitecturecontext-aware
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Time delay neural network (TDNN) has been proven to be efficient for speaker verification. One of its successful variants, ECAPA-TDNN, achieved state-of-the-art performance at the cost of much higher computational complexity and slower inference speed. This makes it inadequate for scenarios with demanding inference rate and limited computational resources. We are thus interested in finding an architecture that can achieve the performance of ECAPA-TDNN and the efficiency of vanilla TDNN. In this paper, we propose an efficient network based on context-aware masking, namely CAM++, which uses densely connected time delay neural network (D-TDNN) as backbone and adopts a novel multi-granularity pooling to capture contextual information at different levels. Extensive experiments on two public benchmarks, VoxCeleb and CN-Celeb, demonstrate that the proposed architecture outperforms other mainstream speaker verification systems with lower computational cost and faster inference speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  2. HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage audio-visual deepfake detector, HOLA, uses 1.81M pre-training samples and hierarchical cross-modal fusion modules to achieve first place and near-perfect AUC on AV-Deepfake1M++ video-level detection.

  3. Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A multi-stage pipeline combining curriculum learning, chain-of-thought data, and RL with verifiable rewards achieves 11.57% WER on the MLC-SLM multilingual ASR test set, versus a 20.17% baseline.

  4. Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A speaker-aware progressive OSD model using WavLM, Campplus, and VAD-gated masking reports 82.76% F1 on AMI, above the listed prior best of 79.21%.

  5. VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A 293-speaker longitudinal dataset with weekly samples over up to 17 years is introduced and used to show that speaker verification error grows with age, especially for female and middle-aged speakers.

  6. NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

    eess.AS 2026-07 conditional novelty 5.5 of 10

    A hierarchical NVAE plug-in generates diverse pseudo-speaker embeddings that raise ASV EER above 38% on FACodec/CosyVoice2 with a controllable privacy-utility trade-off.

  7. X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

    eess.AS 2026-07 conditional novelty 5.0 of 10

    An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.

  8. Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification

    eess.AS 2025-09 conditional novelty 5.0 of 10

    A Bi-LSTM variant of ECAPA-TDNN's Res2Block cuts speaker-verification EER by 23% on VoxCeleb1-O at nearly the same parameter count.

  9. DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A dual-branch pooling method, DRASP, combines global statistics with segment-level attention and improves MOS prediction correlation with human ratings.

Pith tools