REVIEW 3 major objections 5 minor 51 references
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A two-stage speech enhancer that first cleans continuous features, then generates discrete tokens, claims to beat single-paradigm models.
desk verdict Two-stage continuous-then-discrete GSE with a RootLM/BranchLM codec hierarchy is a real architectural contribution, but private fine-tuning data and missing significance tests keep the 'surpasses' claim from being proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the hierarchical language model built on the codec's residual vector quantization: a RootLM models shared acoustic content, while per-level BranchLMs predict each codebook level conditioned on the RootLM and on the previous level, which is what lets the system regenerate missing spectral or temporal content while keeping acoustic consistency. The supporting mechanism is the channel-split NAC-RoFormer, which lowers the computational cost of attending to 1024-dimensional codec features by grouping channels and alternating temporal and cross-group attention.
What would settle it
Train OmniGSE on public data only and evaluate on a held-out compound-distortion set whose distortion recipe differs from the training recipe; if its advantage over single-paradigm baselines disappears, the central claim of general superiority is not supported.
Extended reading notes
Core claim
Stage I uses a channel-split network with dual-path rotary-position attention to map the codec encoder's features for distorted speech toward the clean-speech features produced by a teacher codec encoder, and the codec encoder itself is fine-tuned on distorted input at the same time. Stage II takes those enhanced pre-quantized features as conditioning and autoregressively predicts the residual vector quantization tokens of clean speech. The hierarchical language model consists of one RootLM, which captures acoustic features shared across codebook levels, and a separate BranchLM per level, which captures the progressive relationship from one codebook level to the next; the levels are trained with teacher forcing under a cross-entropy loss. The discovery, stated on the paper's terms, is that this continuous-to-discrete collaboration resolves the precision-versus-flexibility tradeoff, and that the hierarchical LM reduces inter-level prediction conflicts that a single shared LM would introduce.
Load-bearing premise
The central claim assumes the performance gains come from the two-stage architecture itself, not from the private high-fidelity fine-tuning data or from the hand-chosen distortion mix aligning with the test sets.
Editorial extensions
If this is right
- A single OmniGSE model could replace separate denoising, dereverberation, bandwidth extension, declipping, and packet-loss concealment modules.
- Because the first stage improves the conditioning features, the second stage can focus on regenerating missing content instead of suppressing noise, so the two stages compound.
- The paper's ablation says using separate BranchLMs avoids the pattern conflicts that a single multi-level LM produces.
- The framework reports its restoration quality without extra self-supervised semantic features, which lowers conditioning cost.
Reading between the lines
- Editorial inference: the RootLM/BranchLM split is a general recipe for any residual-vector-quantization codec, so it could transfer to other neural codecs and to low-bitrate speech synthesis.
- Editorial inference: the architecture implies that discriminative feature cleaning can act as a universal conditioning front-end for generative audio models, not just for speech enhancement.
- Editorial inference: the paper does not report how sensitive the results are to the hand-chosen distortion probabilities; varying them per deployment is a direct test of the framework's real-world generality.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes OmniGSE, a two-stage general speech enhancement framework that combines a continuous feature enhancement stage (channel-split NAC-RoFormer) with a discrete token generation stage based on a hierarchical language model (RootLM plus BranchLMs). The authors evaluate on DNS2020 denoising/dereverberation, Voicefixer SR/GSR, and Interspeech 2022 PLC benchmarks, and they report ablations of the two stages, the NAC-RoFormer, the hierarchical LM, encoder fine-tuning, and teacher forcing. The central claim is that this cross-domain collaborative design surpasses existing single-paradigm discriminative and generative methods, with particular strength in compound distortion scenarios.
Significance. The architectural idea is coherent: a discriminative pre-enhancement stage feeds high-SNR continuous features to a generative LM-based token predictor, and the hierarchical RootLM/BranchLM design explicitly models RVQ inter-codebook dependencies. The ablations in Table 5 provide useful evidence that each component contributes to the final performance. The breadth of the benchmark coverage (denoising, dereverberation, super-resolution, restoration, packet loss) is a strength. However, the central comparative claim is currently not fully supported because the evaluation uses a private fine-tuning dataset unavailable to the baselines, reports only point estimates without confidence intervals, and the listening tests omit the most relevant generative baselines. If these evaluation gaps are addressed, the framework would be a solid contribution to general speech enhancement.
major comments (3)
- [Sec. 4.1, Tables 1-4] The comparison is confounded by the private fine-tuning data. Section 4.1 states that the authors 'fine-tuned our model on a private high-fidelity speech dataset,' but the baselines in Tables 1-4 were not given this data. Additionally, the training distortion recipe in Sec. 3.1 (noise 100%, reverb 50%, other distortions equally likely) may be aligned with the test-set conditions, and the baselines were not trained with this exact recipe. Therefore the reported gains cannot be attributed solely to the proposed architecture. I request a controlled experiment: (i) train OmniGSE without the private data; (ii) fine-tune at least one representative generative baseline (e.g., MaskSR or AnyEnhance) on the same private data and distortion pipeline; and (iii) report the performance difference to isolate the contribution of the private data.
- [Tables 1-3] No confidence intervals or significance tests are reported, and the claimed superiority is not consistent across all metrics. For example, in Table 2 OmniGSEfb has lower SBS (0.930 vs 0.941) and SIM (0.935 vs 0.943) than AnyEnhance, and in Table 3 OmniGSEfb has lower NISQA (4.293 vs 4.335) than MaskSR. Some OVRL differences in Table 1 are as small as 0.026 (3.444 vs 3.418). Without variance estimates or statistical tests, the statement in the abstract that OmniGSE 'surpasses existing models across multiple benchmarks' is overstated, and the specific claim of excelling in compound distortions is not uniformly supported by the GSR results.
- [Figures 3-4] The subjective listening tests compare OmniGSE only against discriminative baselines (FullSubNet, VoiceFixer, TF-GridNet) and omit the strongest generative competitors (MaskSR, AnyEnhance, LLaSE-G1) that are the direct point of comparison for the generative stage. Figure 3 and Figure 4 also lack details on the number of listeners, the number of stimuli, and any statistical analysis. Please extend the listening test to include at least the strongest generative baselines and report listener counts and significance testing.
minor comments (5)
- [Sec. 2.2] The word 'knowlege' should be 'knowledge' in the sentence about discrete codebooks encapsulating rich prior knowledge.
- [Sec. 4.1] The description of the private high-fidelity dataset is too vague; please report its size, duration, and whether it was used only for fine-tuning after the main training or as part of the main training mixture, since this materially affects the comparison.
- [Sec. 4.2] The full-band and wideband models contain roughly 0.97B and 1.23B parameters, respectively, yet Sec. 2.2 claims reduced computational cost. Please report inference time (e.g., RTF) and memory usage to substantiate the efficiency claim.
- [Figure 5] The x-axis label 'SNR' is ambiguous; please define whether this is signal-domain SNR or a feature-space measure, and additionally report the final objective scores for each conditioning feature rather than only the feature SNR.
- [Table 1] Some baseline entries have missing values (e.g., VoiceFixer SBS and SIM, SELM and GenSE NISQA/SBS/SIM); please indicate whether these are not reported or not applicable, and consider supplementing the missing metrics for a complete comparison.
Circularity Check
No circularity: the paper is an empirical system built on external benchmarks, with no derivation step that reduces to its own inputs.
full rationale
The central claim, that OmniGSE surpasses existing models on general speech enhancement benchmarks, is an empirical claim supported by comparisons on public test sets (Interspeech 2020 DNS, Voicefixer SR/GSR, Interspeech 2022 PLC). The training objectives are Lemb = MSE(Fenh, Ftea) in Eq. (6) and the code cross-entropy loss in Eq. (7), both of which supervise the model toward teacher-NAC targets for clean speech. These are standard training losses, not quantities used to derive benchmark scores. No parameter is fitted to the test metrics, and no metric used in Tables 1-4 appears as an input to the model. The paper does not invoke a uniqueness theorem or an ansatz from the authors' prior work; references to DAC, WavLM, DNSMOS, and NISQA are external, and no cited result is loaded with the paper's own conclusion. The disclosed use of a private high-fidelity speech dataset in Sec. 4.1 is a legitimate evaluation-fairness concern: the baselines were not retrained on that data, so the reported advantages could in principle reflect data rather than architecture. However, that is a controlled-comparison issue, not circularity under the defined patterns; it does not make the benchmark claim true by construction. The distortion simulation recipe (noise at 100%, reverb at 50%, other distortions equally likely) likewise matches common evaluation practice and is not a fitted parameter renamed as a prediction. Overall, no step of the paper's argument reduces, by definition or by self-citation, to its own inputs, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The pre-trained DAC codebooks provide a sufficient discrete representation of clean speech such that codec reconstruction error is small relative to enhancement gains.
- domain assumption The teacher NAC's clean-speech continuous embeddings are a valid regression target for stage I.
- domain assumption DNSMOS, NISQA, PLCMOS, SBS, and SIM, along with the collected MOS, are reliable proxies for perceptual quality and speaker fidelity.
- domain assumption The synthetic distortion pipeline (sequential noise, reverb, then one of clipping, super-resolution, or packet loss) is representative of real-world compound distortions.
Cite this review
Pith. "Pith review of From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models." pith.science (2026). https://pith.science/paper/73UJPIK3
@misc{pith2026250719062,
author = {Pith},
title = {Pith review of: From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/73UJPIK3}},
note = {Machine review of arXiv:2507.19062}
}
read the original abstract
This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang. 2021. Hi-Fi Multi-Speaker English TTS Dataset. In Interspeech. ISCA, 2776–2780
work page 2021
-
[2]
Sebastian Braun and Ivan Tashev. 2020. Data Augmentation and Loss Normaliza- tion for Deep Noise Suppression. In SPECOM (Lecture Notes in Computer Science, Vol. 12335). Springer, 79–86
work page 2020
-
[3]
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022. WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing. IEEE J. Sel. Top. Signal Process. 16,...
work page 2022
-
[4]
Sanyuan Chen, Chengyi Wang, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al . 2025. Neural codec language models are zero-shot text to speech synthesizers. IEEE Transactions on Audio, Speech and Language Processing (2025)
work page 2025
-
[5]
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023. High Fidelity Neural Audio Compression. Trans. Mach. Learn. Res. 2023 (2023)
work page 2023
-
[6]
Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi. 2020. Real Time Speech Enhancement in the Waveform Domain. In INTERSPEECH. ISCA, 3291–3295
work page 2020
-
[7]
Lorenz Diener, Marju Purin, Sten Sootla, Ando Saabas, Robert Aichner, and Ross Cutler. 2023. PLCMOS - A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms. In INTERSPEECH. ISCA, 2533–2537
work page 2023
-
[8]
Lorenz Diener, Sten Sootla, Solomiya Branets, Ando Saabas, Robert Aichner, and Ross Cutler. 2022. INTERSPEECH 2022 Audio Deep Packet Loss Concealment Challenge. In INTERSPEECH. ISCA, 580–584
work page 2022
Show all 51 references
-
[9]
Harishchandra Dubey, Ashkan Aazami, Vishak Gopal, Babak Naderi, Sebastian Braun, Ross Cutler, Alex Ju, Mehdi Zohourian, Min Tang, Mehrsa Golestaneh, et al. 2024. Icassp 2023 deep noise suppression challenge. IEEE Open Journal of Signal Processing (2024)
2024
-
[10]
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra
-
[11]
Xiang Hao, Xiangdong Su, Radu Horaud, and Xiaofei Li. 2021. Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement. In ICASSP. IEEE, 6633–6637
2021
-
[12]
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE ACM Trans. Audio Speech Lang. Process. 29 (2021), 3451–3460
2021
-
[13]
Shengpeng Ji, Ziyue Jiang, Xize Cheng, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Ruiqi Li, Ziang Zhang, Xiaoda Yang, Rongjie Huang, Yidi Jiang, Qian Chen, Siqi Zheng, Wen Wang, and Zhou Zhao. 2024. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio L...
2024 arXiv
-
[14]
Boyi Kang, Xinfa Zhu, Zihan Zhang, Zhen Ye, Mingshuai Liu, Ziqian Wang, Yike Zhu, Guobin Ma, Jun Chen, Longshuai Xiao, et al. 2025. LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement. arXiv preprint arXiv:2503.00493 (2025)
2025 arXiv
-
[15]
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2023. High-Fidelity Audio Compression with Improved RVQGAN. In NeurIPS
2023
-
[16]
Haoyang Li, Jia Qi Yip, Tianyu Fan, and Eng Siong Chng. 2025. Speech En- hancement Using Continuous Embeddings of Neural Audio Codec. CoRR abs/2502.16240 (2025)
2025 arXiv
-
[17]
Nan Li, Xiguang Zheng, Chen Zhang, Liang Guo, and Bing Yu. 2022. End-to-End Multi-Loss Training for Low Delay Packet Loss Concealment. In INTERSPEECH. ISCA, 585–589
2022
-
[18]
Xu Li, Qirui Wang, and Xiaoyu Liu. 2024. MaskSR: Masked Language Model for Full-band Speech Restoration. CoRR abs/2406.02092 (2024)
2024 arXiv
-
[19]
Baiyun Liu, Qi Song, Mingxue Yang, Wuwen Yuan, and Tianbao Wang. 2022. PLCNet: Real-time Packet Loss Concealment with Semi-supervised Generative Adversarial Network. In INTERSPEECH. ISCA, 575–579
2022
-
[20]
Plumbley
Haohe Liu, Ke Chen, Qiao Tian, Wenwu Wang, and Mark D. Plumbley. 2024. Audiosr: Versatile Audio Super-Resolution at Scale. In ICASSP. IEEE, 1076–1080
2024
-
[21]
Haohe Liu, Xubo Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, DeLiang Wang, Chuanzeng Huang, and Yuxuan Wang. 2022. VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration. In INTERSPEECH. ISCA, 4232–4236
2022
-
[22]
Ziyin Liu, Tilman Hartwig, and Masahito Ueda. 2020. Neural Networks Fail to Learn Periodic Functions and How to Fix It. In NeurIPS
2020
-
[23]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In ICLR (Poster). OpenReview.net
2019
-
[24]
Wei Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, and Yun-Ning Hung. 2024. Music Source Separation With Band-Split Rope Transformer. In ICASSP. IEEE, 481–485
2024
-
[25]
Ye-Xin Lu, Yang Ai, and Zhen-Hua Ling. 2023. MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. InINTERSPEECH. ISCA, 3834–3838
2023
-
[26]
Gabriel Mittag, Babak Naderi, Assmaa Chehadi, and Sebastian Möller. 2021. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets. In Interspeech. ISCA, 2127–2131
2021
-
[27]
Chandan K. A. Reddy, Vishak Gopal, and Ross Cutler. 2022. Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors. In ICASSP. IEEE, 886–890
2022
-
[28]
Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan, and Johannes Gehrke. 2020. The INTERSPEECH 2020 Deep Noise Suppression Challeng...
2020
-
[29]
Takaaki Saeki, Soumi Maiti, Shinnosuke Takamichi, Shinji Watanabe, and Hiroshi Saruwatari. 2024. SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics. CoRR abs/2401.16812 (2024)
2024 arXiv
-
[30]
Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. RoFormer: Enhanced transformer with Rotary Position Embedding. Neurocomputing 568 (2024), 127063
2024
-
[31]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[32]
Ter- riberry, Michael Klingbeil, Paris Smaragdis, and Arvindh Krishnaswamy
Jean-Marc Valin, Ahmed Mustafa, Christopher Montgomery, Timothy B. Ter- riberry, Michael Klingbeil, Paris Smaragdis, and Arvindh Krishnaswamy. 2022. Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model. In INTERSPEECH. ISCA, 570–574
2022
-
[33]
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. 2017. CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit. University of Edinburgh. The Centre for Speech Technology Research (CSTR) 6 (2017), 15
2017
-
[34]
Zhong-Qiu Wang, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, and Shinji Watanabe. 2023. TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech Separation. IEEE ACM Trans. Audio Speech Lang. Process. 31 (2023), 3221–3236
2023
-
[35]
Ziqian Wang, Xinfa Zhu, Zihan Zhang, Yuanjun Lv, Ning Jiang, Guoqing Zhao, and Lei Xie. 2024. SELM: Speech Enhancement using Discrete Tokens and Language Models. In ICASSP. IEEE, 11561–11565
2024
-
[36]
Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux. 2019. WHAM!: Extending Speech Separation to Noisy Environments. In INTERSPEECH. ISCA, 1368–1372
2019
-
[37]
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. 2022. Vision Transformer with Deformable Attention. In CVPR. IEEE, 4784–4793
2022
-
[38]
Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2024. BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec. CoRR abs/2409.05377 (2024)
2024 arXiv
-
[39]
Haici Yang, Jiaqi Su, Minje Kim, and Zeyu Jin. 2024. Genhancer: High-fidelity speech enhancement via generative modeling on discrete codec tokens. In Proc. Interspeech 2024. 1170–1174
2024
-
[40]
Jixun Yao, Hexin Liu, Chen Chen, Yuchen Hu, Chng Eng Siong, and Lei Xie. 2025. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling. CoRR abs/2502.02942 (2025)
2025 arXiv
-
[41]
Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi Dai, Hongzhan Lin, Jianyi Chen, Xingjian Du, Liu- meng Xue, Yunlin Chen, Zhifei Li, Lei Xie, Qiuqiang Kong, Yike Guo, and Wei Xue
-
[42]
Jia Qi Yip, Shengkui Zhao, Dianwen Ng, Eng Siong Chng, and Bin Ma. 2024. Towards audio codec-based speech separation. arXiv preprint arXiv:2406.12434 (2024)
2024 arXiv
-
[43]
Jianwei Yu and Yi Luo. 2023. Efficient Monaural Speech Enhancement with Universal Sample Rate Band-Split RNN. In ICASSP. IEEE, 1–5
2023
-
[44]
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for Longer Sequences. InNeurIPS
2020
-
[45]
Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019. LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech. In INTERSPEECH. ISCA, 1526–1530
2019
-
[46]
Junan Zhang, Jing Yang, Zihao Fang, Yuancheng Wang, Zehua Zhang, Zhuo Wang, Fan Fan, and Zhizheng Wu. 2025. AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement. CoRR abs/2501.15417 (2025)
2025
-
[47]
Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Anurag Kumar, Jan Pirklbauer, Marvin Sach, Shinji Watanabe, Tim Fingscheidt, and Yanmin Qian. 2024. URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement. Co...
2024 arXiv
-
[48]
Zihan Zhang, Jiayao Sun, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, and Lei Xie. 2024. Bs-Plcnet: Band-Split Packet Loss Concealment Network with Multi- Task Learning Framework and Multi-Discriminators. In ICASSP Workshops. IEEE, 23–24
2024
-
[49]
Watcharasupat, and Woon-Seng Gan
Shengkui Zhao, Bin Ma, Karn N. Watcharasupat, and Woon-Seng Gan. 2022. FR- CRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement. In ICASSP. IEEE, 9281–9285
2022
-
[2022]
IEEE ACM Trans
FSD50K: An Open Dataset of Human-Labeled Sound Events. IEEE ACM Trans. Audio Speech Lang. Process. 30 (2022), 829–852
2022
-
[2025]
CoRR abs/2502.04128 (2025)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis. CoRR abs/2502.04128 (2025)
2025 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.