REVIEW 2 major objections 6 minor 65 references
Phone Segmentation and Recognition through Phonological Activation Mapping
T0 review · 2 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Self-supervised speech models already encode phonetic structure; two simple heads recover phone labels and boundaries from under a minute of transcriptions.
desk verdict Clean engineering result: SPAM plus two gradient-free heads turns latent S3M phonetics into joint segmentation+recognition from under a minute of labels, with solid OOD evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
S3M-based Phonological Activation Mapping (SPAM): each frame representation is projected onto difference-of-means phonological vectors (one per binary feature channel) and affinely normalized, yielding a time-aligned matrix of feature activations that both the recognition and segmentation heads read directly.
What would settle it
Estimate the phonological vectors only on TIMIT, then measure recognition error and boundary R-value on a held-out language or atypical-speech set that contains many phones and feature combinations never seen in TIMIT; if performance collapses relative to a model that sees even a few minutes of that language’s data, the claim that the directions transfer fails.
Extended reading notes
Core claim
Phonetic structure is already latent in the representation geometry of self-supervised speech models; once that structure is isolated as linear phonological directions and turned into a frame-wise activation map (SPAM), two simple non-gradient heads recover both phone boundaries and phone labels with under a minute of labeled speech and generalize to unseen phones and out-of-domain data.
Load-bearing premise
The linear phonological directions estimated from center-pooled frames of a single English corpus stay aligned enough with true phonetic features across languages, accents, and atypical speech that nearest-neighbor matching and peak detection on the resulting activations still work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that phonetic structure is already latent in self-supervised speech model (S3M) representations and can be steered for joint phone segmentation and recognition without gradient-based training of the heads. It constructs S3M-based Phonological Activation Mapping (SPAM) by estimating linear phonological vectors as differences of means over center-pooled frames (Eq. 1), projecting each frame onto those vectors with affine normalization (Eqs. 2–3), and stacking the resulting activations. Two gradient-descent-free heads then operate on SPAM: a recognition head that nearest-neighbor matches the center-frame activations to PanPhon canonical feature vectors (Eq. 4), and a segmentation head that ensembles multi-scale cosine differences, backward-contrast signals, and a mel-spectrogram difference via a product of non-negative signals followed by peak detection (Eqs. 5–11). With phonological vectors estimated from TIMIT (or fractions thereof), the method is evaluated for R-value segmentation and PFER recognition across English, accented, atypical, and multilingual corpora, showing competitive or superior out-of-domain performance relative to TIMIT-trained CTC/FCE/BCE baselines and usable results with under a minute of labeled data.
Significance. If the empirical claims hold, the work offers a practical, sample-efficient route to time-aligned phonetic transcription that is especially relevant for low-resource languages, atypical speech, and fieldwork settings where large transcribed corpora and heavy fine-tuning are unavailable. Strengths include the explicit design of parameter-free (or closed-form) heads, the sample-efficiency ablation down to ~18 utterances (Fig. 3), the unseen-phone analysis under oracle segmentation, the broad multi-domain evaluation (Tables I–II), and the public release of modeling and evaluation code. The approach also supplies an interpretable intermediate representation (SPAM) that unifies segmentation and recognition under a single phonological geometry, which is a useful conceptual contribution even if absolute topline numbers remain higher for large supervised systems.
major comments (2)
- [Table II / §IV-B] Table II: SPAM’s in-domain PFER on PR-tmt (22.9) is substantially worse than the CTC baseline (7.2) and the SotA toplines, while remaining competitive only on the multilingual average. The abstract and §IV-B claim “strong” recognition performance; the manuscript should either qualify this claim more carefully (e.g., “strong OOD generalization relative to TIMIT-trained baselines”) or provide additional analysis of when the nearest-neighbor head fails (phonotactics, inventory size, silence/closure handling).
- [§V-B] §V-B (oracle segmentation) correctly identifies the segmenter as the primary bottleneck (PFER drops to 11.1 on TIMIT and 8.4 on VoxAngeles with GT boundaries). Given that the central claim is joint segmentation-and-recognition, the paper should either strengthen the segmentation head (e.g., by reporting precision/recall or boundary-error distributions) or more explicitly frame recognition results under predicted vs. oracle boundaries so readers can assess the joint system’s practical utility.
minor comments (6)
- [§III-A] §III-A: Center pooling is asserted to be preferable for phonological arithmetic, citing prior work; a one-sentence quantitative comparison (center vs. average pooling) on the same TIMIT vectors would make the design choice self-contained.
- [Eq. (2)] Eq. (2) and footnote 1: γ is fixed to 4 for all experiments. A brief sensitivity check (or statement that results are insensitive within a range) would reassure readers that the constant is not a hidden free parameter.
- [Table I / Fig. 3] Table I / §IV-A: R-value is reported without error bars or multiple random seeds for the data-subsampling experiments in Fig. 3. Even a single standard deviation over a few seeds would strengthen the sample-efficiency claim.
- [§III-D] §III-D: The product ensemble (Eq. 11) and the seven-signal list are clear, but the theoretical minima φ_k are not tabulated; listing them (or stating they are the analytic lower bounds of each cosine-based term) would aid reproducibility.
- [§III-B] Fig. 1 caption and §III-B: The silence, closure, and release channels are important engineering details; a short note on how they affect the PanPhon nearest-neighbor match (Eq. 4) would clarify the recognition pipeline.
- [Throughout] Minor typography: “V oxangeles” / “V oxAngeles” spacing is inconsistent; “SotA” is used without expansion on first occurrence in §IV.
Circularity Check
Mild self-citation of linear phonological directions plus definitional reuse of PanPhon for both vector estimation and recognition readout; empirical OOD metrics remain independent of construction.
-
self citation load bearing
[Section I; Section III-A]
"Recent work [27], [28] shows that these phonological features can be accurately modeled as linear directions, i.e., phonological vectors, in the representation space of some S3Ms. For instance, adding the voicing vector to a representation of [s] moves it toward [z] (Figure 1, left). Estimating these vectors is known to be highly sample-efficient [27]"
The claim that S3M representations contain accurate, sample-efficient linear phonological directions recoverable by difference-of-means is the foundation of SPAM and is supported only by citations whose author lists substantially overlap with the present paper. Downstream task numbers still stand independently, but the geometric premise itself is not re-derived here.
-
self definitional
[Section III-A Eq. (1); Section III-C Eq. (4)]
"We use PanPhon [34] to assign each phone its phonological features. ... The phonological vector for channel i is a difference of means: vi = µi − µ∁i ... Because the SPAM channels are PanPhon features, recognition requires no trained classifier: a segment is labeled with the phone whose phonological feature vector best matches its SPAM activations. ... ˆv = arg max_v σ(mc(s))⊤ pv. This amounts to a nearest-neighbor lookup in PanPhon"
Phonological channels and their vectors are defined from PanPhon feature assignments on the training phones; the recognition head then classifies by matching the resulting activations against the identical PanPhon canonical vectors. For phones whose features appear in training, successful recovery is a direct readout of the same feature system used to construct the directions (quality still depends on S3M geometry).
full rationale
The paper's load-bearing empirical results (R-value on OOD/atypical/multilingual sets in Table I, PFER on PRiSM in Table II, sample-efficiency curves in Fig. 3, and seen/unseen phone PFER under oracle segmentation) are obtained by applying closed-form heads to held-out ground-truth annotations that are independent of the TIMIT-fitted means. The recognition head is a parameter-free nearest-neighbor match (Eq. 4) and the segmenter is an ensemble of cosine differences plus a closed-form least-squares backward contrast (Eqs. 5–11); neither quantity is forced by the fit. Circularity is limited to (a) the premise that difference-of-means recovers accurate linear phonological directions, which rests on self-citations [27],[28], and (b) the fact that both vector construction and phone readout use the same PanPhon feature inventory. These do not make the reported numbers tautological. Score 2 is therefore appropriate; no fitted-input-as-prediction, uniqueness-import, or ansatz-smuggling reductions exist.
Assumptions & free parameters
free parameters (4)
- gamma (SPAM scaling constant)
- ensemble of seven segmentation signals and their theoretical minima phi_k
- S3M layer and model choice (final layer of WavLM-large)
- 20 ms boundary hit threshold and strict R-value mode
assumptions (4)
- domain assumption Phonological features appear as approximately linear directions recoverable by difference-of-means in S3M representation space.
- domain assumption PanPhon’s 21 ternary articulatory features (plus silence/closure/release channels) are a sufficient and language-universal basis for phone identity.
- domain assumption Center-pooling of phone spans yields more reliable phonological vectors than average-pooling.
- ad hoc to paper Cosine distance peaks (and product ensemble thereof) on SPAM activations correspond to phone boundaries.
invented entities (2)
-
S3M-based Phonological Activation Mapping (SPAM)
independent evidence
-
Backward-contrast segmentation signals beta_ell
Cite this review
Pith. "Pith review of Phone Segmentation and Recognition through Phonological Activation Mapping." pith.science (2026). https://pith.science/paper/Q6WBZERT
@misc{pith2026260709020,
author = {Pith},
title = {Pith review of: Phone Segmentation and Recognition through Phonological Activation Mapping},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q6WBZERT}},
note = {Machine review of arXiv:2607.09020}
}
read the original abstract
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.
Figures
Reference graph
Works this paper leans on
-
[1]
TIMIT acoustic-phonetic continuous speech corpus,
J. S. Garofolo, L. F. Lamel, W. M. Fisher, D. S. Pallett, N. L. Dahlgren, V . Zue, and J. G. Fiscus, “TIMIT acoustic-phonetic continuous speech corpus,” 1993
1993
-
[2]
Phonetic segmentation of the UCLA phonetics lab archive,
E. Chodroff, B. Pa ˇzon, A. Baker, and S. Moran, “Phonetic segmentation of the UCLA phonetics lab archive,” inLREC-COLING, 2024
2024
-
[3]
L. D. Shriberg, R. D. Kent, T. McAllister, J. L. Preston, and M. L. Speights,Clinical phonetics. Plural Publishing, 2025
2025
-
[4]
EduSpeak®: A speech recognition and pronunciation scoring toolkit for computer-aided language learning applications,
H. Franco, H. Bratt, R. Rossier, V . R. Gadde, E. Shriberg, V . Abrash, and K. Precoda, “EduSpeak®: A speech recognition and pronunciation scoring toolkit for computer-aided language learning applications,” Language Testing, vol. 27, pp. 401 – 418, 2010. [Online]. Available: https://api.semanticscholar.org/CorpusID:143273296
2010
-
[5]
PRiSM: Benchmarking phone realization in speech models,
S. Bharadwaj, C.-J. Li, Y . Kim, K. Choi, E. Yeo, R. S.-E. Shim, H. Zhou, B. Boldt, K. R. Jacome, K. Chang, D. Agrawal, K. Xu, C.-H. H. Yang, J. Zhu, S. Watanabe, and D. R. Mortensen1, “PRiSM: Benchmarking phone realization in speech models,” inACL, 2026
2026
-
[6]
Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments,
D. R. Mortensen, J. Picone, X. Li, and K. Siminyu, “Tusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments,” inProc. Interspeech, 2021, pp. 3660–3664
2021
-
[7]
Prosodic abx: A language-agnostic method for measuring prosodic contrast in speech representations,
H. Sun, S. McIntosh, K. Choi, E. Yeo, D. Saito, and N. Minematsu, “Prosodic abx: A language-agnostic method for measuring prosodic contrast in speech representations,”Interspeech, 2026
2026
-
[8]
Speech playground: An interactive tool for speech analysis and comparison,
S. McIntosh, D. Saito, and N. Minematsu, “Speech playground: An interactive tool for speech analysis and comparison,”arXiv preprint arXiv:2607.00418, 2026
arXiv 2026
Show all 65 references
-
[9]
Towards language-agnostic stipa: Universal phonetic transcription to support language documentation at scale,
J. L. Suchardt, H. El-Shazli, and P. Cassotti, “Towards language-agnostic stipa: Universal phonetic transcription to support language documentation at scale,” inEMNLP, 2025
2025
-
[10]
Language documentation twenty-five years on,
F. Seifart, N. Evans, H. Hammarstr ¨om, and S. C. Levinson, “Language documentation twenty-five years on,”Language, vol. 94, no. 4, pp. e324– e345, 2018
2018
-
[11]
The buckeye corpus of conversational speech: labeling conventions and a test of transcriber reliability,
M. A. Pitt, K. Johnson, E. Hume, S. Kiesling, and W. Raymond, “The buckeye corpus of conversational speech: labeling conventions and a test of transcriber reliability,”Speech Communication, vol. 45, pp. 89–95, 2005
2005
-
[12]
Reliability studies in broad and narrow phonetic transcription,
L. D. Shriberg and G. L. Lof, “Reliability studies in broad and narrow phonetic transcription,”Clinical Linguistics & Phonetics, vol. 5, no. 3, pp. 225–279, 1991
1991
-
[13]
Simple and Effective Zero-shot Cross- lingual Phoneme Recognition,
Q. Xu, A. Baevski, and M. Auli, “Simple and Effective Zero-shot Cross- lingual Phoneme Recognition,” inInterspeech, 2022
2022
-
[14]
ZIPA: A family of efficient models for multilingual phone recognition,
J. Zhu, F. Samir, E. Chodroff, and D. R. Mortensen, “ZIPA: A family of efficient models for multilingual phone recognition,” inACL, 2025
2025
-
[15]
POWSM: A phonetic open whisper-style speech foundation model,
C.-J. Li, K. Chang, S. Bharadwaj, E. Yeo, K. Choi, J. Zhu, D. Mortensen, and S. Watanabe, “POWSM: A phonetic open whisper-style speech foundation model,” inACL, 2026
2026
-
[16]
An Empirical Recipe for Universal Phone Recognition,
S. Bharadwaj, C.-J. Li, K. Choi, E. Yeo, W. Chen, S. Watanabe, and D. R. Mortensen, “An Empirical Recipe for Universal Phone Recognition,” in Interspeech, 2026
2026
-
[17]
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,
A. Graves, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” inICML, 2006
2006
-
[18]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” inNeurIPS, 2017
2017
-
[19]
Attention-based models for speech recognition,
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y . Bengio, “Attention-based models for speech recognition,”Advances in neural information processing systems, vol. 28, 2015
2015
-
[20]
Phoneme Segmentation Using Self-Supervised Speech Models,
L. Strgar and D. Harwath, “Phoneme Segmentation Using Self-Supervised Speech Models,” inSLT, 2023
2023
-
[21]
Phone-to-audio alignment without text: A semi-supervised approach,
J. Zhu, C. Zhang, and D. Jurgens, “Phone-to-audio alignment without text: A semi-supervised approach,” inICASSP, 2022
2022
-
[22]
Explore wav2vec 2.0 for mispronunciation detection
X. Xu, Y . Kang, S. Cao, B. Lin, and L. Ma, “Explore wav2vec 2.0 for mispronunciation detection.” inInterspeech, 2021
2021
-
[23]
Speech intelligibility assessment of dysarthric speech by using goodness of pronunciation with uncertainty quantification,
E. J. Yeo, K. Choi, S. Kim, and M. Chung, “Speech intelligibility assessment of dysarthric speech by using goodness of pronunciation with uncertainty quantification,” inInterspeech, 2023
2023
-
[24]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS, 2020
2020
-
[25]
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”IEEE/ACM TASLP, 2021
2021
-
[26]
WavLM: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., “WavLM: Large-scale self-supervised pre- training for full stack speech processing,”J-STSP, 2022
2022
-
[27]
[b]=[d]- [t]+[p]: Self-supervised speech models discover phonological vector arithmetic,
K. Choi, E. Yeo, C. J. Cho, D. Harwath, and D. R. Mortensen, “[b]=[d]- [t]+[p]: Self-supervised speech models discover phonological vector arithmetic,” inACL Findings, 2026
2026
-
[28]
Self- supervised speech models encode phonetic context via position-dependent orthogonal subspaces,
K. Choi, E. Yeo, C. J. Cho, D. R. Mortensen, and D. Harwath, “Self- supervised speech models encode phonetic context via position-dependent orthogonal subspaces,”arXiv preprint arXiv:2603.12642, 2026
2026
-
[29]
Opening the black box of wav2vec feature encoder,
K. Choi and E. J. Yeo, “Opening the black box of wav2vec feature encoder,”arXiv preprint arXiv:2210.15386, 2022
2022 arXiv
-
[30]
Analysing discrete self supervised speech representation for spoken language modeling,
A. Sicherman and Y . Adi, “Analysing discrete self supervised speech representation for spoken language modeling,” inICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[31]
Leveraging allophony in self-supervised speech models for atypical pronunciation assessment,
K. Choi, E. Yeo, K. Chang, S. Watanabe, and D. R. Mortensen, “Leveraging allophony in self-supervised speech models for atypical pronunciation assessment,” inNAACL, 2025
2025
-
[32]
Layer-wise analysis of a self- supervised speech representation model,
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self- supervised speech representation model,” inASRU, 2021
2021
-
[33]
What do Speech Foundation Models Learn? Analysis and Applications,
A. Pasad, “What do Speech Foundation Models Learn? Analysis and Applications,” Ph.D. dissertation, Toyota Technical Institute at Chicago, 2025, accessed on https://arxiv.org/abs/2508.12255
2025 arXiv
-
[34]
Panphon: A resource for mapping ipa segments to articulatory feature vectors,
D. R. Mortensen, P. Littell, A. Bharadwaj, K. Goyal, C. Dyer, and L. Levin, “Panphon: A resource for mapping ipa segments to articulatory feature vectors,” inCOLING, 2016
2016
-
[35]
Multi-level acoustic segmentation of continuous speech,
J. Glass and V . Zue, “Multi-level acoustic segmentation of continuous speech,” inICASSP, 1988
1988
-
[36]
Segmentation and modeling in segment- based recognition,
J. W. Chang and J. R. Glass, “Segmentation and modeling in segment- based recognition,” inEurospeech, 1997
1997
-
[37]
Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation,
F. Kreuk, J. Keshet, and Y . Adi, “Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation,” inInterspeech, 2020
2020
-
[38]
A simple hmm with self-supervised represen- tations for phone segmentation,
G.-P. Yang and H. Tang, “A simple hmm with self-supervised represen- tations for phone segmentation,” inSLT, 2024
2024
-
[39]
Unsupervised Speech Segmentation and Variable Rate Representation Learning Using Segmental Contrastive Predictive Coding,
S. Bhati, J. Villalba, P. ˙Zelasko, L. Moro-Velazquez, and N. Dehak, “Unsupervised Speech Segmentation and Variable Rate Representation Learning Using Segmental Contrastive Predictive Coding,”IEEE/ACM TASLP, 2022
2022
-
[40]
Universal phone recognition with a multilingual allophone system,
X. Li, S. Dalmia, J. Li, M. Lee, P. Littell, J. Yao, A. Anastasopoulos, D. R. Mortensen, G. Neubig, A. W. Blacket al., “Universal phone recognition with a multilingual allophone system,” inICASSP, 2020
2020
-
[41]
Allophant: Cross-lingual Phoneme Recognition with Articulatory Attributes,
K. Glocker, A. Herygers, and M. Georges, “Allophant: Cross-lingual Phoneme Recognition with Articulatory Attributes,” inInterspeech, 2023
2023
-
[42]
Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer,
N. Tits, P. Bhatnagar, and T. Dutoit, “Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer,”Acoustics, vol. 6, no. 3, pp. 772–781, Sep. 2024
2024
-
[43]
Self-Supervised Speech Representations are More Phonetic than Semantic,
K. Choi, A. Pasad, T. Nakamura, S. Fukayama, K. Livescu, and S. Watanabe, “Self-Supervised Speech Representations are More Phonetic than Semantic,” inInterspeech, 2024
2024
-
[44]
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, ˙I. Polat, Y . Fen...
2020
-
[45]
A new text-independent method for phoneme segmentation,
G. Aversano, A. Esposito, and M. Marinaro, “A new text-independent method for phoneme segmentation,” inMidwest Symposium on Circuits and Systems (MWSCAS). IEEE, 2001
2001
-
[46]
Montreal forced aligner: Trainable text-speech alignment using kaldi,
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” inInterspeech, vol. 2017, 2017, pp. 498–502
2017
-
[47]
XLSR Inclusive English Speech-to-IPA,
K. Labs, “XLSR Inclusive English Speech-to-IPA,” 2025. [Online]. Available: https://huggingface.co/collections/KoelLabs/ xlsr-inclusive-english-speech-to-ipa
2025
-
[48]
L2-ARCTIC: A Non-native English Speech Corpus,
G. Zhao, S. Sonsaat, A. Silpachai, I. Lucic, E. Chukharev-Hudilainen, J. M. Levis, and R. Gutierrez-Osuna, “L2-ARCTIC: A Non-native English Speech Corpus,” inInterspeech, 2018
2018
-
[49]
Speech Accent Archive,
S. Weinberger, “Speech Accent Archive,” 2015, retrieved from https: //accent.gmu.edu
2015
-
[50]
Building a time-aligned cross-linguistic reference corpus from language documentation data (DoReCo),
L. Paschen, F. Delafontaine, C. Draxler, S. Fuchs, M. Stave, and F. Seifart, “Building a time-aligned cross-linguistic reference corpus from language documentation data (DoReCo),” inProc. LREC. European Language Resources Association, 2020
2020
-
[51]
Scaling human and g2p supervision for robust phonetic transcription,
A. Metzger, A. Srivastava, and R. Mukhamedvaleev, “Scaling human and g2p supervision for robust phonetic transcription,” inInterspeech, 2026
2026
-
[52]
Global TIMIT learner simple english,
H. Ding, S. Liao, Y . Zhan, H. Feng, W. He, X. Hu, Y . Wu, J. Yuan, and M. Liberman, “Global TIMIT learner simple english,” Web Download. LDC2020S11, Philadelphia, 2020. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2020S11
2020
-
[53]
Global TIMIT learner treebank english,
H. Luan, Y . Wang, H. Feng, W. He, X. Hu, Y . Wu, J. Yuan, and M. Liberman, “Global TIMIT learner treebank english,” Web Download. LDC2020S09, Philadelphia, 2020. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2020S09
2020
-
[54]
Global TIMIT Thai,
M. Liberman, J. Yuan, C. Cieri, J. Wright, and N. Chanchaochai, “Global TIMIT Thai,” Web Download. LDC2022S13, Philadelphia, 2022. [Online]. Available: https://catalog.ldc.upenn.edu/LDC2022S13
2022
-
[55]
Dysarthric speech corpus in Tamil for rehabilitation research,
T. A. Mariya Celin, T. Nagarajan, and P. Vijayalakshmi, “Dysarthric speech corpus in Tamil for rehabilitation research,” in2016 IEEE Region 10 Conference (TENCON), 2016, pp. 2610–2613
2016
-
[56]
A weighted speaker-specific confusion transducer-based augmentative and alternative speech communication aid for dysarthric speakers,
T. A. Mariya Celin, G. Anushiya Rachel, T. Nagarajan, and P. Vi- jayalakshmi, “A weighted speaker-specific confusion transducer-based augmentative and alternative speech communication aid for dysarthric speakers,”IEEE Transactions on Neural Systems and Rehabilitation Engineeri...
2019
-
[57]
The TORGO database of acoustic and articulatory speech from speakers with dysarthria,
F. Rudzicz, A. K. Namasivayam, and T. Wolff, “The TORGO database of acoustic and articulatory speech from speakers with dysarthria,”Language Resources and Evaluation, vol. 46, no. 4, pp. 523–541, 2012
2012
-
[58]
Dysarthria detection and severity assessment using rhythm-based metrics
A. Hernandez, E. J. Yeo, S. Kim, and M. Chung, “Dysarthria detection and severity assessment using rhythm-based metrics.” inInterspeech, 2020
2020
-
[59]
An improved speech segmentation quality measure: the r-value,
O. J. R ¨as¨anen, U. K. Laine, and T. Altosaar, “An improved speech segmentation quality measure: the r-value,” inInterspeech, 2009
2009
-
[60]
English mfa acoustic model v3.1.0,
M. McAuliffe and M. Sonderegger, “English mfa acoustic model v3.1.0,” https://mfa-models.readthedocs.io/acoustic/English/ EnglishMFAacousticmodelv3 1 0.html, Tech. Rep., Jun 2024
2024
-
[61]
V oxcommunis corpus,
E. Ahn and E. Chodroff, “V oxcommunis corpus,” https://osf.io/t957v, Jan 2022
2022
-
[62]
Thai mfa acoustic model v3.0.0,
M. McAuliffe and M. Sonderegger, “Thai mfa acoustic model v3.0.0,” https://mfa-models.readthedocs.io/acoustic/Thai/ ThaiMFAacousticmodelv3 0 0.html, Tech. Rep., Feb 2024
2024
-
[63]
Wav2Gloss: Generating Interlinear Glossed Text from Speech,
T. He, K. Choi, L. Tjuatja, N. Robinson, J. Shi, S. Watanabe, G. Neubig, D. Mortensen, and L. Levin, “Wav2Gloss: Generating Interlinear Glossed Text from Speech,” inACL, 2024
2024
-
[64]
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,” inInterspeech, 2022
2022
-
[65]
Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks,
H. Kamper and B. v. Niekerk, “Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks,” in Interspeech, 2021
2021
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.