Pith. sign in

REVIEW 3 cited by

An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.12838 v2 pith:EQWBHBHO submitted 2023-05-22 eess.AS cs.SD

classification eess.AScs.SD
keywords fusionfeaturefeaturesgloballocaleres2netspeakerverification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective fusion of multi-scale features is crucial for improving speaker verification performance. While most existing methods aggregate multi-scale features in a layer-wise manner via simple operations, such as summation or concatenation. This paper proposes a novel architecture called Enhanced Res2Net (ERes2Net), which incorporates both local and global feature fusion techniques to improve the performance. The local feature fusion (LFF) fuses the features within one single residual block to extract the local signal. The global feature fusion (GFF) takes acoustic features of different scales as input to aggregate global signal. To facilitate effective feature fusion in both LFF and GFF, an attentional feature fusion module is employed in the ERes2Net architecture, replacing summation or concatenation operations. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the ERes2Net in speaker verification. Code has been made publicly available at https://github.com/alibaba-damo-academy/3D-Speaker.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    cs.MM 2025-06 conditional novelty 6.0 of 10

    StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.

  2. Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Qwen-Audio-3.0-TTS claims state-of-the-art controllable multilingual text-to-speech across 16 languages and 20 Chinese dialects, using a 12.5 Hz tokenizer and multi-stage RL.

  3. Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

    cs.SD 2025-09 conditional novelty 4.0 of 10

    Neural audio codecs match or beat Opus for speaker verification on VoxCeleb1 below 12 kbps and stay within about 1.5 percentage points EER above it.

Pith tools