REVIEW 5 cited by
SpEx+: A Complete Time Domain Speaker Extraction Network
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Speaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech. We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches. Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference. The size of the analysis window for time-domain and the size for frequency-domain input are also different. Such mismatch has an adverse effect on the system performance. To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+. Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder. Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively.
Forward citations
Cited by 5 Pith papers
-
FlowTSE: Target Speaker Extraction with Flow Matching
Conditional flow matching on mel-spectrograms with a phase-conditioned vocoder matches or beats published TSE baselines on Libri2Mix.
-
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
Adding room dimensions and reverberation time as inputs to a distance-based target speech extraction model improves SDR by about 1.2 dB on unseen simulated and real rooms.
-
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.
-
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
M3ANet aligns EEG and speech representations with InfoNCE contrastive learning and encodes speech with multi-scale convolutions plus GroupMamba, improving brain-assisted target speaker extraction on three datasets.
-
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
DCF-Net, a time-frequency target speaker extraction model with a DualStream Fusion Block, reports 21.6 dB SI-SDRi on WSJ0-2Mix and a 0.4% target confusion rate.
Discussion (0). Continue with ORCID to comment.