Pith. sign in

REVIEW 5 cited by

SpEx+: A Complete Time Domain Speaker Extraction Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.04686 v2 pith:MQ55BJAU submitted 2020-05-10 eess.AS cs.SD

classification eess.AScs.SD
keywords speakerspextime-domainextractionspeechfrequency-domainsolutioncomplete
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech. We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches. Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference. The size of the analysis window for time-domain and the size for frequency-domain input are also different. Such mismatch has an adverse effect on the system performance. To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+. Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder. Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FlowTSE: Target Speaker Extraction with Flow Matching

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Conditional flow matching on mel-spectrograms with a phase-conditioned vocoder matches or beats published TSE baselines on Libri2Mix.

  2. Single-Channel Target Speech Extraction Utilizing Distance and Room Clues

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Adding room dimensions and reverberation time as inputs to a distance-based target speech extraction model improves SDR by about 1.2 dB on unseen simulated and real rooms.

  3. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0 of 10

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.

  4. M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

    eess.AS 2025-05 conditional novelty 5.0 of 10

    M3ANet aligns EEG and speech representations with InfoNCE contrastive learning and encodes speech with multi-scale convolutions plus GroupMamba, improving brain-assisted target speaker extraction on three datasets.

  5. DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions

    cs.SD 2025-02 conditional novelty 4.0 of 10

    DCF-Net, a time-frequency target speaker extraction model with a DualStream Fusion Block, reports 21.6 dB SI-SDRi on WSJ0-2Mix and a 0.4% target confusion rate.

Pith tools