REVIEW 4 major objections 6 minor 1 cited by
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Phone-purity guided tokens cut dysarthric speech word error by up to 1.77%.
desk verdict A credible, well-controlled empirical study of phone-purity-guided discrete tokens for dysarthric ASR; the gains are real but modest, and the main weak spot is the unvalidated use of forced-alignment phone labels. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the phone-purity regularization term added to the two quantization objectives. For PPG K-means, the cluster-mean update becomes a weighted average of the standard data mean and the 'purest centroid'—the mean of all frames in the cluster that share the top-1 most frequent phone label—with weight λ. For PPG VAE-VQ, the training loss adds α times a phone-entropy term that measures how evenly the quantized features are spread across phone classes, using a Gaussian phone posterior. Both terms pull codebook entries toward phone-homogeneous regions of the feature space, so the discrete tokens carry more phonetic discrimination than purely unsupervised clusters.
What would settle it
Re-run the PPG K-means and PPG VAE-VQ pipelines on UASpeech using phone labels from an independent phone recognizer (or manually corrected alignments) instead of the GMM-HMM forced alignments; if the WER gains over non-PPG tokens disappear or reverse, the benefit is an artifact of the label source rather than of phone-purity guidance itself.
Extended reading notes
Core claim
The central claim is that phonetic label supervision during discrete token extraction improves downstream ASR on dysarthric speech. Two quantization methods are modified: K-means cluster centroids are updated toward the mean of the frames in the cluster that share the most frequent phone label, and VAE-VQ's reconstruction loss is augmented with an entropy term that penalizes uncertainty of phone labels conditioned on the quantized features. On UASpeech, the resulting PPG tokens outperform their unsupervised counterparts in every codebook-size condition tested, yielding up to 0.99% and 1.77% absolute WER gains on hybrid TDNN and E2E Conformer systems, and the best system combination reaches 23.25% WER. The authors also report consistent gains on a phone-purity metric and sharper T-SNE cluster boundaries, linking the WER reductions to improved phonetic separation in the token space.
Load-bearing premise
The gains rest on the assumption that the frame-level phone labels obtained by GMM-HMM forced alignment on dysarthric speech are accurate enough to define useful phone-purity targets; if articulation imprecision makes these alignments systematically wrong, the purity loss would reinforce alignment errors and the WER gains could shrink or reverse.
Editorial extensions
If this is right
- The PPG tokens can be dropped into existing hybrid and end-to-end ASR pipelines as direct replacements for the input features, with no change to the back-end model architecture.
- Because the gains hold at both codebook sizes tested (100 and 500), phone-purity guidance offers a way to use smaller codebooks without sacrificing recognition accuracy.
- The combination of systems built on PPG and non-PPG tokens yields the best result, suggesting that the supervised and unsupervised tokens encode complementary information.
- Since the quantization operates on any continuous frame-level features, the phone-purity guidance should extend to other speech foundation models such as WavLM or Whisper.
Reading between the lines
- The phone-purity metric could serve as a cheap validation criterion for choosing codebook size or the regularization weight, removing the need to run full ASR for every configuration.
- If the WER gains track phone purity as the paper reports, then a natural testable extension is to apply PPG tokens to other low-resource disordered speech domains, such as children's speech, where unsupervised tokens are known to underperform.
- A caveat worth testing: the method depends on the quality of the GMM-HMM forced-alignment phone labels, so gains might change if those labels are replaced by a different phone-label source.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes phone-purity guided (PPG) discrete token extraction for dysarthric speech recognition. Frame-level phonetic labels obtained by GMM-HMM forced alignment are used to regularize two quantization procedures: K-means, by pulling each cluster centroid toward the mean of its most frequent phone label (Eq. 2), and VAE-VQ, by adding an entropy-based phone-purity loss computed from a Gaussian phone-class posterior (Eqs. 3-6). The discrete tokens are extracted from a fine-tuned HuBERT model and evaluated with hybrid TDNN and end-to-end Conformer ASR systems on the UASpeech corpus at codebook sizes 100 and 500. The paper reports statistically significant WER reductions for PPG tokens over non-PPG K-means and VAE-VQ tokens, with the largest gains of 0.99% and 1.77% absolute, and a combined system achieving 23.25% WER. The authors claim this is the first study of discrete tokens for dysarthric speech recognition.
Significance. If the central claim holds, the paper introduces a simple and potentially general way to inject phonetic supervision into discrete token extraction, improving ASR for atypical speech where unsupervised tokenization is known to lose phonetic discrimination. The controlled experimental design is a genuine strength: PPG and non-PPG tokens are trained and evaluated under the same recipes, with held-out WER evaluation and MAPSSWE paired significance tests. The use of two different back-end ASR architectures and multiple codebook sizes, plus system combination, strengthens the empirical case. However, the load-bearing reliance on forced-alignment labels whose accuracy is never measured, the circularity of the phone-purity metric, and missing key hyperparameters currently limit the reproducibility and the strength of the causal claim. The contribution is a useful empirical result, but additional validation is needed before the claimed mechanism can be regarded as established.
major comments (4)
- [Sec. III.A, Eq. (2) and Sec. III.B, Eq. (5)] The phone-purity supervision depends entirely on the frame-level phonetic labels Y obtained by GMM-HMM forced alignment, but the paper never reports any measure of alignment accuracy on dysarthric speech. Since articulation imprecision can systematically bias forced alignments, the regularizers in Eq. (2) and Eq. (5) may be optimizing against corrupted targets. The WER gains are small in absolute terms (0.40-1.77%), so an aligner-specific artifact could plausibly produce the observed pattern. The authors should validate the label source, for example by reporting alignment accuracy on a manually transcribed subset, comparing with an alternative aligner (e.g., a different GMM-HMM recipe or a neural aligner), or ablating with labels from a different source. Without such evidence, the central claim that phone purity, rather than aligner-specific noise, drives the gains is not fully established.
- [Table II and Sec. IV.C.3] The phone-purity metric in Table II is not independent evidence for the proposed method. The metric is computed against the same reference label set that defines the PPG objective; the objective explicitly encourages label-homogeneous clusters by construction (the K-means target p_k is the mean of top-1-label features, and the VAE-VQ loss is the entropy of the label posterior). Higher phone purity for PPG tokens is therefore partly guaranteed by construction, and Table II cannot validate label quality or independently confirm the WER gains. The authors should either evaluate purity against an independent label source (e.g., manual phone transcriptions or a different aligner) or explicitly reframe Table II as a sanity check rather than supporting evidence.
- [Sec. III.A, Eq. (2)] The K-means phone-purity regularization weight λ is never reported, although the corresponding VAE-VQ weight α is given (α=1.2 and 1.05 for K=100 and K=500 in Sec. III.B). Since the modified centroid update in Eq. (2) is the core mechanism of the proposed K-means method, the paper should report λ for each codebook size and ideally provide a sensitivity analysis over λ. In addition, the definition of p_k ('the mean average over all the frame level un-quantized features ... that share the top-1 most frequent reference phonetic label') is ambiguous in degenerate cases, such as a cluster with no unique most frequent label or with all labels distinct; the paper should clarify how such cases are handled.
- [Table I and Sec. IV.C] The claim that PPG tokens 'consistently outperform' non-PPG tokens is stronger than the reported significance pattern. For example, Sys. 11 vs Sys. 10 on TDNN at K=500 is not significant overall (30.24 vs 30.64, no dagger), and several PPG entries (e.g., Sys. 7, 21, 25, 29) are not significant against their direct non-PPG counterparts in some intelligibility subgroups. The paper should qualify the consistency claim by explicitly listing which pairwise comparisons are significant and discussing whether non-significant results correlate with codebook size or intelligibility subgroup. Additionally, Table I contains many pairwise comparisons tested at α=0.05 without any multiple-comparison correction; the authors should either apply a correction or state explicitly why the uncorrected MAPSSWE tests are considered appropriate.
minor comments (6)
- [Abstract] The first sentence, 'Discrete tokens extracted provide efficient and domain adaptable speech features,' is missing an object; it should read 'Discrete tokens extracted from speech foundation models provide ...'.
- [Sec. III.B, Eq. (5)] The entropy definition should specify the base of the logarithm and how zero posterior probabilities are handled, since Eq. (6) can in principle produce numerically zero probabilities for some phonetic classes.
- [Sec. IV.A] The paper states that all dysarthric training utterances are used for phone purity analysis, but Table II does not specify whether purity is computed on the same data used for clustering or on a held-out set; this should be clarified.
- [Table III] The paper says the final system is 'further contrasted against SOTA,' but the reported 23.25% WER is higher than several prior published results (e.g., 16.53% in CUHK-2024). Since the contribution is about discrete tokens rather than achieving SOTA, the comparison should be framed explicitly as not a direct SOTA claim, with a note on the different training conditions (augmentation, adaptation, and system combination).
- [References] Reference [30] is a duplicate of reference [11]; the duplicate should be removed and the citation renumbered.
- [Throughout] The spelling 'V AE-VQ' appears in several places with an extra space; this should be unified to 'VAE-VQ' for consistency.
Circularity Check
Held-out WER claim is self-contained; the phone-purity confirmation is partly by construction but not load-bearing.
-
self definitional
[Sec. III.A Eq. (2), Sec. III.B Eq. (5), Sec. IV.C Table II]
"The phone-purity regularization term, LPur, in the above Equation (3) is the entropy loss computed using the V AE-VQ quantized features against the reference phonetic labels."
The PPG regularizers directly optimize the same construct that Table II reports: Eq. (5) minimizes label entropy against the reference phonetic labels, and Eq. (2) pulls each K-means centroid toward the mean of features sharing the top-1 reference label. Thus the reported 'consistent improvements on the phone purity metric' are in part guaranteed by construction and do not independently validate phonetic discrimination or label quality. This is not the central claim: the WER gains are measured on the held-out UASpeech B2 test set, so they do not reduce to the training objective.
full rationale
The paper's central claim -- statistically significant WER reductions from PPG tokens -- is evaluated on the held-out UASpeech B2 test set with a significance test, and the baselines are standard non-PPG K-means/VAE-VQ tokens under matched codebook sizes. That claim is therefore not circular: the test labels and decoding are external to the token-extraction objective. The only self-referential element is the phone-purity metric: the method is explicitly defined to increase purity/minimize label entropy against the forced-alignment labels, and Table II's phone purity is the same construct, so those metric gains are partly by construction rather than independent evidence. The unvalidated GMM-HMM forced-alignment labels are a robustness risk (if the alignments are wrong, the regularizer reinforces their errors), but that is an empirical validity concern, not a circularity in the derivation. No load-bearing argument reduces to a self-citation or a fitted parameter renamed as a prediction. Overall the central derivation is self-contained; score 2 reflects the one minor by-construction confirmation metric.
Assumptions & free parameters
free parameters (4)
- lambda (K-means phone-purity weight) =
not reported
- alpha (VAE-VQ phone-purity weight) =
1.2 for K=100, 1.05 for K=500
- codebook size K =
100 and 500
- Gaussian phone class parameters (mu_j and Sigma_j) =
estimated from data, values not reported
assumptions (4)
- domain assumption Frame-level phonetic reference labels from GMM-HMM forced alignment are accurate enough to define phone purity.
- domain assumption HuBERT fine-tuned with a bottleneck and CTC provides continuous features suitable for quantization and downstream ASR.
- standard math Standard K-means and VAE-VQ training assumptions hold, including convergence and SGD behavior.
- standard math The MAPSSWE matched-pairs test is a valid basis for comparing WERs.
Cite this review
Pith. "Pith review of Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition." pith.science (2026). https://pith.science/paper/H2ECDZMD
@misc{pith2026250104379,
author = {Pith},
title = {Pith review of: Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/H2ECDZMD}},
note = {Machine review of arXiv:2501.04379}
}
read the original abstract
Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch against normal voice remains unexplored. To improve their phonetic discrimination that is weakened during unsupervised K-means or vector quantization of continuous features, this paper proposes novel phone-purity guided (PPG) discrete tokens for dysarthric speech recognition. Phonetic label supervision is used to regularize maximum likelihood and reconstruction error costs used in standard K-means and VAE-VQ based discrete token extraction. Experiments conducted on the UASpeech corpus suggest that the proposed PPG discrete token features extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG based K-means or VAE-VQ tokens across varying codebook sizes by statistically significant word error rate (WER) reductions up to 0.99\% and 1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech test set of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained by combining systems using different token features. Consistent improvements on the phone purity metric were also achieved. T-SNE visualization further demonstrates sharper decision boundaries were produced between K-means/VAE-VQ clusters after introducing phone-purity guidance.
Figures
Forward citations
Cited by 1 Pith paper
-
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
Regularized federated learning (parameter, embedding, and KL-loss based) consistently outperforms FedAvg for dysarthric and elderly speech recognition by up to 0.55% absolute WER, and per-batch communication approache...
Reference graph
Works this paper leans on
-
[1]
Combining in-domain and out-of-domain speech data for automatic recognition of disordered speech,
H. Christensen, M. Aniol, P. Bell, P. Green, T. Hain, and P. Swietojanski, “Combining in-domain and out-of-domain speech data for automatic recognition of disordered speech,” in INTERSPEECH, 2013
work page 2013
-
[2]
Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition,
F. Xiong, J. Barker, Z. Yue, and H. Christensen, “Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition,” in ICASSP, 2020
work page 2020
-
[3]
Recent Progress in the CUHK Dysarthric Speech Recognition System,
S. Liu, M. Geng, S. Hu, X. Xie, M. Cui, J. Yu, X. Liu, and H. Meng, “Recent Progress in the CUHK Dysarthric Speech Recognition System,” IEEE/ACM TASLP, 2021
work page 2021
-
[4]
M. Geng, X. Xie, Z. Ye, T. Wang, G. Li, S. Hu, X. Liu, and H. Meng, “Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2022
work page 2022
-
[5]
Hyper-parameter adaptation of conformer asr systems for elderly and dysarthric speech recognition,
T. Wang, S. Hu, J. Deng, Z. Jin, M. Geng, Y . Wang, H. Meng, and X. Liu, “Hyper-parameter adaptation of conformer asr systems for elderly and dysarthric speech recognition,” in INTERSPEECH, 2023
work page 2023
-
[6]
Self-supervised asr models and features for dysarthric and elderly speech recognition,
S. Hu, X. Xie, M. Geng, Z. Jin, J. Deng, G. Li, Y . Wang, M. Cui, T. Wang, H. Meng, and X. Liu, “Self-supervised asr models and features for dysarthric and elderly speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
work page 2024
-
[7]
H. Wang, Z. Jin, M. Geng, S. Hu, G. Li, T. Wang, H. Xu, and X. Liu, “Enhancing pre-trained asr system fine-tuning for dysarthric speech recognition using adversarial data augmentation,” in ICASSP, 2024
work page 2024
-
[8]
F. Xiong, J. Barker, and H. Christensen, “Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition,” in ICASSP, 2019
work page 2019
Show all 44 references
-
[9]
Investigation of Data Augmentation Techniques for Disordered Speech Recognition,
M. Geng, X. Xie, S. Liu et al. , “Investigation of Data Augmentation Techniques for Disordered Speech Recognition,” INTERSPEECH, 2020
2020
-
[10]
Exploring Self-supervised Pre-trained ASR Models for Dysarthric and Elderly Speech Recognition,
S. Hu, X. Xie, Z. Jin, M. Geng, Y . Wang, M. Cui, J. Deng, X. Liu, and H. Meng, “Exploring Self-supervised Pre-trained ASR Models for Dysarthric and Elderly Speech Recognition,” in ICASSP, 2023
2023
-
[11]
Benefits of Pre-Trained Mono- and Cross- Lingual Speech Representations for Spoken Language Understanding of Dutch Dysarthric Speech,
P. Wang and H. Van hamme, “Benefits of Pre-Trained Mono- and Cross- Lingual Speech Representations for Spoken Language Understanding of Dutch Dysarthric Speech,” EURASIP J. Audio Speech Music Process. , 2023
2023
-
[12]
Adversarial Data Augmentation for Disordered Speech Recognition,
Z. Jin, M. Geng, X. Xie, J. Yu, S. Liu, X. Liu, and H. Meng, “Adversarial Data Augmentation for Disordered Speech Recognition,” in INTERSPEECH, 2021
2021
-
[13]
Towards automatic data augmentation for disordered speech recognition,
Z. Jin, X. Xie, T. Wang, M. Geng, J. Deng, G. Li, S. Hu, and X. Liu, “Towards automatic data augmentation for disordered speech recognition,” in ICASSP, 2024
2024
-
[14]
vq-wav2vec: Self-supervised learning of discrete speech representations,
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” Proc. ICLR, 2020
2020
-
[15]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , 2020
2020
-
[16]
Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
-
[17]
Wavlm: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Select...
2022
-
[18]
Explo- ration of Efficient End-to-End ASR using Discretized Input from Self- Supervised Learning,
X. Chang, B. Yan, Y . Fujita, T. Maekaku, and S. Watanabe, “Explo- ration of Efficient End-to-End ASR using Discretized Input from Self- Supervised Learning,” in Proc. INTERSPEECH 2023 , 2023
2023
-
[19]
Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,
X. Chang, B. Yan, K. Choi, J.-W. Jung, Y . Lu, S. Maiti, R. Sharma, J. Shi, J. Tian, S. Watanabe et al. , “Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,” in ICASSP, 2024
2024
-
[20]
Towards universal speech discrete tokens: A case study for asr and tts,
Y . Yang, F. Shen, C. Du, Z. Ma, K. Yu, D. Povey, and X. Chen, “Towards universal speech discrete tokens: A case study for asr and tts,” in ICASSP, 2024
2024
-
[21]
Children’s speech recognition through discrete token enhancement,
V . N. Sukhadia and S. A. Chowdhury, “Children’s speech recognition through discrete token enhancement,” in INTERSPEECH, 2024
2024
-
[22]
Vqtts: High-fidelity text-to-speech synthesis with self-supervised vq acoustic feature,
C. Du, Y . Guo, X. Chen, and K. Yu, “Vqtts: High-fidelity text-to-speech synthesis with self-supervised vq acoustic feature,” in Interspeech, 2022
2022
-
[23]
Unicats: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding,
C. Du, Y . Guo, F. Shen, Z. Liu, Z. Liang, X. Chen, S. Wang, H. Zhang, and K. Yu, “Unicats: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding,” Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
-
[24]
Expresso: A benchmark and analysis of discrete expressive speech resynthesis,
T. A. Nguyen, W.-N. Hsu, A. D’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, F. Kreuk, Y . Adi, and E. Dupoux, “Expresso: A benchmark and analysis of discrete expressive speech resynthesis,” in INTERSPEECH, 2023
2023
-
[25]
What dysarthrias can tell us about the neural control of speech,
R. D. Kent, J. F. Kent, G. Weismer, and J. R. Duffy, “What dysarthrias can tell us about the neural control of speech,” Journal of Phonetics , 2000
2000
-
[26]
Dysarthric Speech Database for Universal Access Research,
H. Kim, M. Hasegawa-Johnson, A. Perlman, J. Gunderson, K. Watkin, and S. Frame, “Dysarthric Speech Database for Universal Access Research,” in INTERSPEECH, 2008
2008
-
[27]
Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech Recognition,
A. Hernandez, P. A. P ´erez-Toro, E. Noeth, J. R. Orozco-Arroyave, A. Maier, and S. H. Yang, “Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech Recognition,” in IN- TERSPEECH, 2022
2022
-
[28]
Adversarial data augmentation using vae-gan for disordered speech recognition,
Z. Jin, X. Xie, M. Geng, T. Wang, S. Hu, J. Deng, G. Li, and X. Liu, “Adversarial data augmentation using vae-gan for disordered speech recognition,” in ICASSP, 2023
2023
-
[29]
Multi-Stage Audio-Visual Fusion for Dysarthric Speech Recognition With Pre-Trained Models,
C. Yu, X. Su, and Z. Qian, “Multi-Stage Audio-Visual Fusion for Dysarthric Speech Recognition With Pre-Trained Models,” IEEE Trans- actions on Neural Systems and Rehabilitation Engineering , 2023
2023
-
[30]
Benefits of pre-trained mono- and cross- lingual speech representations for spoken language understanding of dutch dysarthric speech,
P. Wang and H. Van hamme, “Benefits of pre-trained mono- and cross- lingual speech representations for spoken language understanding of dutch dysarthric speech,” EURASIP J. Audio Speech Music Process. , 2023
2023
-
[31]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunsk...
2023
-
[32]
k-means++: the advantages of careful seeding,
D. Arthur and S. Vassilvitskii, “k-means++: the advantages of careful seeding,” in Proceedings of the Eighteenth Annual ACM-SIAM Sympo- sium on Discrete Algorithms , 2007
2007
-
[33]
An overview of gradient descent optimization algorithms,
S. Ruder, “An overview of gradient descent optimization algorithms,” arXiv preprint arXiv:1609.04747 , 2016
2016 arXiv
-
[34]
Some statistical issues in the comparison of speech recognition algorithms,
L. Gillick and S. Cox, “Some statistical issues in the comparison of speech recognition algorithms,” in ICASSP, 1989
1989
-
[35]
Two-pass decoding and cross-adaptation based system combination of end-to-end conformer and hybrid tdnn asr systems,
M. Cui, J. Deng, S. Hu, X. Xie, T. Wang, S. Hu, M. Geng, B. Xue, X. Liu, and H. Meng, “Two-pass decoding and cross-adaptation based system combination of end-to-end conformer and hybrid tdnn asr systems,” in Interspeech 2022, 2022, pp. 3158–3162
2022
-
[36]
The HTK Book,
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey et al. , “The HTK Book,” Cambridge University Engineering Department , 2002
2002
-
[37]
A time delay neural net- work architecture for efficient modeling of long temporal contexts,
V . Peddinti, D. Povey, and S. Khudanpur, “A time delay neural net- work architecture for efficient modeling of long temporal contexts,” in Interspeech 2015, 2015, pp. 3214–3218
2015
-
[38]
Purely sequence-trained neural networks for asr based on lattice-free mmi,
D. Povey, V . Peddinti, D. Galvez, P. Ghahremani, V . Manohar, X. Na, Y . Wang, and S. Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi,” in Interspeech 2016, 2016, pp. 2751– 2755
2016
-
[39]
Espnet: End-to-end speech processing toolkit,
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y . Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduch- intala, and T. Ochiai, “Espnet: End-to-end speech processing toolkit,” in INTERSPEECH, 2018
2018
-
[40]
Speaker adaptation for wav2vec2 based dysarthric asr,
M. K. Baskar, T. Herzig, D. Nguyen, M. Diez, T. Polzehl, L. Burget, and J. ˇCernock´y, “Speaker adaptation for wav2vec2 based dysarthric asr,” in INTERSPEECH, 2022
2022
-
[41]
Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition,
M. Geng, X. Xie, Z. Ye, T. Wang, G. Li, S. Hu, X. Liu, and H. Meng, “Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition,” IEEE/ACM TASLP, 2022
2022
-
[42]
Ad- versarial Data Augmentation Using V AE-GAN for Disordered Speech Recognition,
Z. Jin, X. Xie, M. Geng, T. Wang, S. Hu, J. Deng, G. Li, and X. Liu, “Ad- versarial Data Augmentation Using V AE-GAN for Disordered Speech Recognition,” in ICASSP, 2023
2023
-
[43]
DuTa-VC: A Duration-aware Typical- to-atypical V oice Conversion Approach with Diffusion Probabilistic Model,
H. Wang, T. Thebaud, J. Villalba, M. Sydnor, B. Lammers, N. De- hak, and L. Moro-Velazquez, “DuTa-VC: A Duration-aware Typical- to-atypical V oice Conversion Approach with Diffusion Probabilistic Model,” in INTERSPEECH, 2023
2023
-
[44]
Use of Speech Impairment Severity for Dysarthric Speech Recognition,
M. Geng, Z. Jin, T. Wang, S. Hu, J. Deng, M. Cui, G. Li, J. Yu, X. Xie, and X. Liu, “Use of Speech Impairment Severity for Dysarthric Speech Recognition,” in INTERSPEECH, 2023
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.