{"id":"ef6cedbe-10aa-49a5-8134-c5980fbd9fe3","arxiv_id":"2412.08306","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A generative diffusion speech enhancement model preserves syllable stress better than discriminative enhancers for non-native English speech, and human perception matches automatic stress detection.","lead":"This paper tests whether speech enhancement models, which remove background noise, preserve the syllable stress patterns that signal word meaning in English. It finds that a generative diffusion-based enhancer (CDiffuSE) keeps stress detection accuracy higher than two discriminative enhancers, and human listeners agree.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central CDiffuSE-vs-discriminative ranking rests on single-run accuracy differences as small as 0.15 percentage points, with no confidence intervals or significance tests reported.","rationale":"The reader correctly identifies missing statistical support as part of the rationale, but the designated weakest assumption is the use of clean-speech syllable boundaries on noisy and enhanced audio. That alignment concern is real and worth testing, yet it is not the single most load-bearing issue: even if the boundaries were perfectly valid, the headline claim would still be unsupported because the accuracy gaps are tiny and unreplicated. The statistical concern directly targets the central assertion that CDiffuSE outperforms the discriminative models, and a concrete paired-significance test can settle whether the observed ranking is credible. I therefore partially agree with the reader's weakest assumption while shifting the focus to the absence of uncertainty quantification, which is the condition most essential to the central claim.","tokens_in":9628,"tokens_out":8450,"duration_ms":106348,"concrete_test":"Use the saved fold-level predictions to run paired McNemar tests, or bootstrap 95% confidence intervals on per-syllable accuracy differences, between CDiffuSE and each discriminative model for every SNR and feature set; repeat the 5-fold cross-validation with at least five random seeds and report mean plus/minus standard deviation. If the CDiffuSE advantage is not significant at p < 0.05 in most conditions, the conclusion should be revised to state that no reliable advantage was established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that CDiffuSE's higher stress-detection accuracy over DTLN and Denoiser reflects a true model-level effect rather than run-to-run variation. The paper reports only point estimates from a single 5-fold cross-validation run, with no standard deviations, confidence intervals, significance tests, or repeated seeds. In Table I with heuristic features on GER, CDiffuSE beats Denoiser by 0.45 pp at 0 dB, 0.15 pp at 5 dB, 0.00 pp at 10 dB, and 0.15 pp at 20 dB; with postprocessing the gaps are 0.5, 0.25, 0.4, and 0.55 pp. These differences are well within the variability expected from a VAE+DNN trained with random initialization on roughly 3.7k utterances. The ITA gaps are larger, around 1.5-2.1 pp, but they are still unreplicated. Furthermore, Table II with wav2vec 2.0 features shows Denoiser beating CDiffuSE at 0 dB on both languages (91.42 vs 89.40 for GER; 91.40 vs 90.30 for ITA), so the unqualified conclusion that CDiffuSE outperforms the discriminative models overstates the full set of results. Table III adds no statistical support: the perceptual percentages are aggregates over 25 items per language with no inter-rater reliability or significance measure, and the accuracy values there differ markedly from the corresponding Table I values, indicating that the perceptual subset is noisier and not reconciled with the main evaluation. Without a demonstration that the observed margins are statistically reliable, the paper's central ranking is not yet supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates three end-to-end speech enhancement (SE) models—DTLN and Denoiser (discriminative) and CDiffuSE (generative)—as preprocessors for automatic syllable stress detection on the ISLE corpus (German and Italian learners of English). Noisy speech at 0, 5, 10, and 20 dB SNR is enhanced by each model, and a jointly optimized VAE+DNN classifier is trained on either 38-dimensional heuristic acoustic/context features or 768-dimensional wav2vec 2.0 features, with and without a one-stressed-syllable postprocessing step. A perceptual study with 25 listeners compares how similar enhanced audio is to clean audio in terms of stress placement. The central claims are that (1) heuristic features are more noise-robust than wav2vec 2.0 features, (2) the generative model CDiffuSE outperforms the discriminative models for stress detection, and (3) perceptual results align with automatic detection results.","tokens_in":9985,"tokens_out":2676,"duration_ms":28516,"significance":"If the central ranking were statistically well supported, the paper would provide a useful practical recommendation: use CDiffuSE as an SE front-end for syllable stress detection in CALL systems and prefer heuristic features over wav2vec 2.0 under low SNR. The use of a real L2 corpus (ISLE), two language groups, three SE models, two feature families, and a listener study are strengths, as is the focus on a relatively understudied downstream task (prosodic stress preservation rather than ASR or intelligibility). However, the paper's main conclusion rests on point estimates from a single 5-fold run with no confidence intervals or significance tests, and the perceptual component lacks inter-rater reliability. Because the reported margins are often below 0.5 percentage points and one feature condition reverses the ranking, the central claims are not yet established to the standard needed for a journal publication.","major_comments":[{"comment":"The central claim that CDiffuSE outperforms the discriminative models is supported only by point estimates from a single 5-fold cross-validation run, with no standard deviations, confidence intervals, or paired significance tests. Several GER differences are very small: at 0 dB CDiffuSE beats Denoiser by 0.45 pp, at 5 dB by 0.15 pp, at 10 dB by 0.00 pp, and at 20 dB by 0.15 pp (without postprocessing); with postprocessing the gaps are 0.5, 0.25, 0.4, and 0.55 pp. These margins are well within the run-to-run variability expected for a VAE+DNN trained on roughly 3.7k utterances. The paper must report variability across seeds or folds and a paired test (e.g., McNemar or bootstrap over utterances) before the ranking can be accepted.","section":"§V-A, Table I"},{"comment":"The conclusion that 'audios enhanced with the SE model belonging to generative modeling category (CDiffuSE) outperforms the discriminative SE models' is stated without qualification, but Table II shows that with wav2vec 2.0 features at 0 dB, Denoiser beats CDiffuSE in both languages (GER: 91.42 vs 89.40; ITA: 91.40 vs 90.30). The paper's own abstract limits the robustness claim to 'when heuristic features are used,' but the conclusion does not. The conclusion should be restricted to the heuristic-feature condition or should provide an explanation for the reversal under self-supervised features.","section":"§V-A, Table II and §VI"},{"comment":"The perceptual study is presented as confirming the automatic results, but no inter-rater reliability, no confidence intervals, and no significance test are reported for the listener preference percentages (45.91% vs 30.34% vs 23.75% for ITA; 37.60% vs 34.88% vs 27.52% for GER). With only 25 items per language and one vote per item per listener, the margins between Denoiser and DTLN in GER (34.88% vs 27.52%) may not be meaningful. In addition, the automatic accuracies in Table III (e.g., ITA CDiffuSE 82.65%) differ substantially from the corresponding Table I values (around 92.4–93.0%), so the perceptual subset is not representative of the main evaluation and the alignment between the two tables is not explained.","section":"§V-B, Table III"},{"comment":"The evaluation assumes that syllable boundaries computed on clean speech remain valid for noisy and enhanced audio. The dataset section states that phoneme alignments are computed on clean recordings and corrected by linguists, and Section III-A says syllable segmentation is performed 'on each speech signal using syllable timestamps' without re-alignment. If enhancement shifts timing or introduces artifacts, the features are computed on misaligned segments, and accuracy differences may reflect alignment error rather than SE quality. The paper should either re-align enhanced audio or provide evidence that the clean timestamps remain accurate after enhancement.","section":"§II and §III-A"}],"minor_comments":[{"comment":"The cross-references in the text appear as 'Table ??' and 'Table ??' in both subsections of Section V-A; these placeholders should be replaced with the actual table numbers.","section":"§V-A"},{"comment":"There are several typographical errors, including 'discirminative' in the introduction, 'Wave2vec 2.0' in the Table II header, 'on on the V oiceBank-DEMAND dataset' in Section III-B3, and 'Adressing' in Section III-B. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The perceptual study section does not specify how the 50 word stimuli were chosen, how many times each subject listened to each sample, or whether the order of presentation was randomized; those details should be added for reproducibility.","section":"§III-C"},{"comment":"The architecture description states that 'all the DNN and VAE parameters ... are optimal and we choose them by maximizing the performance on the validation set,' but no search grid or selection criterion is given; reporting the hyperparameter range would improve reproducibility.","section":"§IV"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practical question, and the authors have assembled a substantial experimental setup. However, the lack of any statistical validation for the headline ranking is a load-bearing gap; without it, the recommendation to prefer CDiffuSE and heuristic features is not justified. I would also encourage the editor to consider whether the fixed clean-speech alignment is a known limitation that the authors can address with a re-alignment experiment or a sensitivity analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is the first paper I know of that compares discriminative and generative end-to-end speech enhancement models specifically for syllable stress preservation, and it includes a human perceptual check. That question is worth asking. Second, the headline claim — generative CDiffuSE consistently beats discriminative DTLN and Denoiser — is not actually supported by the paper's own tables.\n\nWhat's good: the experimental setup is sensible. The authors take the ISLE corpus (non-native German and Italian speakers of English), corrupt with Gaussian noise at 0–20 dB, enhance with three standard models (DTLN, Denoiser, CDiffuSE), run a VAE+DNN stress detector on two feature sets (heuristic acoustic/context and wav2vec 2.0), apply the one-stressed-syllable postprocessing, and compare. The perceptual study with 25 listeners, asking which enhanced version best matches clean stress placement, is a reasonable idea and its winner (CDiffuSE) matches the automatic ranking on the subset.\n\nThe soft spot is statistical, and it is load-bearing. Every accuracy is a single-run point estimate from 5-fold CV, with no standard deviations, confidence intervals, or significance tests. In Table I, CDiffuSE beats Denoiser by 0.15–0.55 percentage points on GER; those margins are within random-initialization noise. The ITA margins (about 1.5–2.1 pp) are larger but still unreplicated. More telling, Table II shows Denoiser beating CDiffuSE at 0 dB on both languages (91.42 vs 89.40 GER; 91.40 vs 90.30 ITA), so the conclusion that CDiffuSE outperforms the discriminative models across the board is simply an overstatement. Table III adds no statistical support: the perceptual percentages are aggregates over 25 items per language, no inter-rater reliability is given, and the subset accuracies (e.g., 82.65% for ITA CDiffuSE) are far below the Table I values, suggesting the perceptual subset is noisier and not reconciled with the main evaluation.\n\nOne more issue: syllable boundaries come from clean-speech alignments and are applied to enhanced audio without re-alignment. If an SE model shifts timing or introduces artifacts, the features are computed on misaligned segments. The authors never audit this.\n\nWho is this for? Speech technologists building CALL front-ends and anyone interested in prosodic feature robustness. The observation that heuristic features hold up better than wav2vec 2.0 under noise, and that diffusion-based enhancement may be gentler on stress cues, is a useful pointer. But as it stands, the evidence is suggestive, not conclusive.\n\nMy recommendation: send it to a serious referee. The question is legitimate, the experiments are cheap to extend, and a good referee can ask for repeated seeds, error bars, paired tests, a disclosed perceptual protocol with reliability, reconciliation of Table III with Table I, and a softened conclusion. I would not cite it as a definitive ranking, but it deserves referee time.","headline":"Useful first comparison of SE models for stress preservation, but the central CDiffuSE ranking is unsupported by unreplicated point estimates and contradicted by Table II at 0 dB.","tokens_in":10533,"tokens_out":3221,"would_cite":false,"duration_ms":29409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative speech enhancement preserves syllable stress better than discriminative models.","keywords":["syllable stress detection","speech enhancement","diffusion models","discriminative vs generative models","heuristic acoustic features","wav2vec 2.0 features","ISLE corpus","perceptual study"],"falsifier":"Re-run stress detection after re-aligning syllable boundaries on each enhanced waveform and compare model rankings; if CDiffuSE's advantage over Denoiser and DTLN shrinks or disappears, the headline ranking comes from clean-audio alignment timings rather than from genuine stress preservation.","tokens_in":9448,"feed_emoji":"🔊","tokens_out":4405,"duration_ms":46366,"temperature":0.7,"pith_summary":"The paper asks whether speech enhancement, used to clean noisy audio before automatic syllable stress detection, preserves the stress patterns that listeners and language-learning systems depend on. It compares three enhancement models—two discriminative and one generative—on English speech from non-native German and Italian speakers at SNRs from 0 to 20 dB. The central claim is that the generative diffusion-based model (CDiffuSE) keeps stress detection accuracy highest, that heuristic acoustic and context features stay robust under noise while wav2vec 2.0 features degrade sharply at low SNR, and that human listeners rank the same model best. If this holds, generative enhancement is the better front-end for stress-focused CALL systems, and heuristic features are safer than self-supervised ones in noisy conditions.","feed_headline":"Diffusion-based enhancer best preserves syllable stress","feed_subtitle":"Generative CDiffuSE beats discriminative models in stress detection, and human listeners agree.","key_machinery":"The pipeline computes syllable boundaries once on clean audio using linguist-corrected phoneme alignments and a syllabification tool, then applies those timestamps to noisy and enhanced audio to cut syllable segments. For each syllable it extracts either a 38-dimensional heuristic feature set (sonority-based prominence contours plus binary context features) or 768-dimensional wav2vec 2.0 frame averages, feeds them to a jointly optimized VAE+DNN classifier, and applies a postprocessing rule that keeps exactly one stressed syllable per word. The three speech enhancement models—CDiffuSE (generative diffusion), Denoiser (waveform-domain discriminative), and DTLN (real-time discriminative)—are the interventions being compared inside this fixed pipeline.","core_discovery":"On the paper's own terms, the discovery is a ranking: the generative CDiffuSE model outperforms the discriminative DTLN and Denoiser models for preserving syllable stress across almost all tested SNR levels, and this ranking matches human perception. The authors also find that heuristic acoustic and context features are robust to noise and enhancement, while wav2vec 2.0 representations suffer large accuracy drops at 0 and 5 dB SNR, both on noisy and enhanced audio. The perceptual study reinforces the automatic results: listeners judged CDiffuSE-enhanced audio as most similar to clean speech, followed by Denoiser and then DTLN, and the same order appears in automatic stress detection accuracy on the subset used for listening.","pith_inferences":["Because syllable boundaries come from clean audio and are never re-estimated on enhanced signals, the reported ranking could partly reflect how well each enhancer preserves timing rather than stress cues; a boundary re-alignment test would separate these effects.","wav2vec 2.0 features are averaged over entire syllables, so additive Gaussian noise is averaged into every feature dimension, which may explain their severe low-SNR degradation; frame-level classifiers or noise-augmented training could change that comparison.","CDiffuSE's iterative denoising may act as a prosody-preserving prior rather than just a noise remover; a direct extension would use its learned representations for stress detection without explicit syllable segmentation.","The perceptual study used only 50 words and 25 listeners, so the alignment between human and automatic judgment is promising but would need a larger stimulus set to generalize across speakers and word types."],"forward_implications":["CALL systems operating in noisy environments can use CDiffuSE as a preprocessing front-end to keep syllable stress detection accuracy close to clean-speech levels.","Heuristic sonority-based features are the safer choice at low SNRs, contrary to the common expectation that self-supervised representations like wav2vec 2.0 dominate traditional features.","Human perception tracks automatic stress detection, so automatic accuracy on enhanced speech is a reliable proxy for what learners hear.","Discriminative enhancers can actively hurt stress detection compared to leaving speech noisy, especially on the Italian dataset, meaning enhancement is not always beneficial for this task.","The one-stressed-syllable postprocessing rule becomes more important when wav2vec 2.0 features are used at 0 dB, where raw classification is least reliable."],"supporting_citations":[{"why":"Supplies the jointly optimized VAE+DNN method used as the stress detection classifier throughout the experiments.","marker":"[12]"},{"why":"Provides the ISLE corpus of non-native English from German and Italian speakers, the data source for all experiments.","marker":"[13]"},{"why":"Defines the sonority-based prominence acoustic features that form the heuristic feature set.","marker":"[14]"},{"why":"Defines the binary context feature vector that captures syllable nucleus type, neighboring phoneme categories, and word position.","marker":"[15]"},{"why":"Introduces wav2vec 2.0, the self-supervised representation model whose features are compared against heuristic features.","marker":"[16]"},{"why":"Prior work showing wav2vec 2.0 features can capture stress patterns, motivating their use in this study.","marker":"[20]"},{"why":"Introduces CDiffuSE, the conditional diffusion generative speech enhancement model being evaluated.","marker":"[25]"},{"why":"Introduces DTLN, the real-time discriminative speech enhancement model being evaluated.","marker":"[26]"},{"why":"Introduces the real-time DEMUCS-based Denoiser, the other discriminative speech enhancement model being evaluated.","marker":"[31]"}],"fun_headline_variants":["Generative SE model best preserves stress in noisy speech","CDiffuSE beats DTLN and Denoiser on stress preservation","Human listeners confirm: generative enhancer keeps stress","Heuristic features robust to noise with generative SE","Diffusion-based enhancement wins on stress detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that syllable boundaries and stress labels derived from clean audio remain correct after noise addition and enhancement, so the model never re-aligns to the processed signal and any timing shift would count as a stress-preservation failure instead of an alignment artifact.","fun_headline_variants_meta":{"raw":{"variants":["Generative SE model best preserves stress in noisy speech","CDiffuSE beats DTLN and Denoiser on stress preservation","Human listeners confirm: generative enhancer keeps stress","Heuristic features robust to noise with generative SE","Diffusion-based enhancement wins on stress detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000535,"raw_usage":{"total_tokens":2567,"prompt_tokens":935,"completion_tokens":1632,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":1555}},"tokens_in":551,"tokens_out":1632,"duration_ms":12167,"temperature":1.0,"reasoning_tokens":1555,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:57:30.668462+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run stress detection after re-aligning syllable boundaries on each enhanced waveform and compare model rankings; if CDiffuSE's advantage over Denoiser and DTLN shrinks or disappears, the headline ranking comes from clean-audio alignment timings rather than from genuine stress preservation.","supporting_citations":[{"cited_title":"A compar- ison of learned representations with jointly optimized vae and dnn for syllable stress detection,","cited_arxiv_id":null,"evidence_quote":"Supplies the jointly optimized VAE+DNN method used as the stress detection classifier throughout the experiments."},{"cited_title":"The ISLE corpus of non-native spoken English,","cited_arxiv_id":null,"evidence_quote":"Provides the ISLE corpus of non-native English from German and Italian speakers, the data source for all experiments."},{"cited_title":"Automatic detection of syllable stress using sonority based prominence features for pronunciation evaluation,","cited_arxiv_id":null,"evidence_quote":"Defines the sonority-based prominence acoustic features that form the heuristic feature set."},{"cited_title":"Comparison of automatic syllable stress detection quality with time-aligned boundaries and context dependencies,","cited_arxiv_id":null,"evidence_quote":"Defines the binary context feature vector that captures syllable nucleus type, neighboring phoneme categories, and word position."},{"cited_title":"Exploring the use of self-supervised representations for automatic syllable stress detection,","cited_arxiv_id":null,"evidence_quote":"Prior work showing wav2vec 2.0 features can capture stress patterns, motivating their use in this study."},{"cited_title":"Conditional diffusion probabilistic model for speech enhancement,","cited_arxiv_id":null,"evidence_quote":"Introduces CDiffuSE, the conditional diffusion generative speech enhancement model being evaluated."},{"cited_title":"Dual-Signal Transformation LSTM Network for Real-Time Noise Suppression","cited_arxiv_id":"2005.07551","evidence_quote":"Introduces DTLN, the real-time discriminative speech enhancement model being evaluated."}],"review_version":1}