Pith. sign in

REVIEW 1 cited by

On The Compensation Between Magnitude and Phase in Speech Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.05470 v2 pith:55P4EPEW submitted 2021-08-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechlossmagnitudedomainmetricsseparationcompensationcomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural network (DNN) based end-to-end optimization in the complex time-frequency (T-F) domain or time domain has shown considerable potential in monaural speech separation. Many recent studies optimize loss functions defined solely in the time or complex domain, without including a loss on magnitude. Although such loss functions typically produce better scores if the evaluation metrics are objective time-domain metrics, they however produce worse scores on speech quality and intelligibility metrics and usually lead to worse speech recognition performance, compared with including a loss on magnitude. While this phenomenon has been experimentally observed by many studies, it is often not accurately explained and there lacks a thorough understanding on its fundamental cause. This paper provides a novel view from the perspective of the implicit compensation between estimated magnitude and phase. Analytical results based on monaural speech separation and robust automatic speech recognition (ASR) tasks in noisy-reverberant conditions support the validity of our view.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concatenating speech-recognition and active-speaker visual embeddings improves audio-visual speech enhancement in low-SNR multi-speaker settings, and the CPU real-time system is released open source.

Pith tools