REVIEW 3 major objections 5 minor 1 cited by
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Multilingual speech datasets carry serious quality problems concentrated in less-institutionalized languages, making many low-resource benchmark results untrustworthy without per-language inspection.
desk verdict A genuinely useful dataset audit with concrete new findings, whose headline correlation claim outruns its evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The audit protocol is the load-bearing instrument. On the quantitative side, the paper computes signal-to-noise ratio, voice activity detection, median utterance duration, median word count, and average hours per speaker for each language subset, which exposes micro-level issues like short clips, silence-heavy recordings, and single-speaker subsets. On the qualitative side, volunteer native speakers reviewed 100 randomly sampled text-plus-audio sentences per language across about 40 languages for coherence, audio-text alignment, dialect, topic domain, and language identification, which exposes macro-level issues. Targeted classifier scripts — a marker-word algorithm for Norwegian Bokmål versus Nynorsk, a marker-word algorithm for Arabic Fusha versus dialect, and the canto-filter package for Cantonese versus Standard Written Chinese — convert sociolinguistic distinctions into testable proportions. The micro/macro taxonomy does the argument's work: micro-level issues can be fixed programmatically, while macro-level issues require language planning decisions about orthography, register, and dialect scope.
What would settle it
Re-run the audit on the full validated splits of the flagged languages (for example FLEURS yue_hk, FLEURS ar_eg, MCV17 nn_no, and MCV17 nan_tw) with two independent native-speaker annotation teams and reported agreement per language, then compare macro-level issue rates and the WER change after cleaning. If full-split classification finds the issue rates close to zero or the cleaning has no measurable effect on WER, the central prevalence and institutionalization-correlation claims would collapse.
Extended reading notes
Core claim
The paper's central claim is that the quality of multilingual speech datasets systematically tracks a language's institutionalization. Micro-level defects — extremely short utterances, low proportions of actual speech, imbalanced topic domains, and lack of speaker diversity — appear across languages but are detectable programmatically and can be mitigated by filtering; macro-level defects — unspecified writing systems in digraphic languages, ambiguous registers in diglossic languages, and underspecified dialect boundaries — are concentrated in less-institutionalized languages and require human linguistic expertise to diagnose. The evidence includes Norwegian subsets that mix Bokmål and Nynorsk despite their locale labels, a FLEURS Arabic subset labeled Egyptian that is almost entirely Modern Standard Arabic, a FLEURS Cantonese subset that contains no Cantonese at all, and a Common Voice Taiwanese Southern Min subset that is dictionary-like, dual-script, and poorly aligned. The paper further shows the evaluation cost: a Norwegian Bokmål ASR model's substitution error rate rises by roughly 25 percentage points on the mixed Nynorsk-labeled subset. What follows is that published WER numbers for many low-resource languages may reflect dataset artifacts rather than model capability.
Load-bearing premise
The prevalence and correlation claims rest on the assumption that 100 randomly sampled sentences per language, judged by volunteer native speakers with no reported inter-rater agreement, represent the quality of each full language subset, even though the paper itself notes in its limitations that many languages remain uninspected.
Editorial extensions
If this is right
- Norwegian subsets mix the two written standards: MCV17's nn_no contains 8.1% Bokmål sentences and FLEURS's nb_no contains 8.8% Nynorsk, and a Bokmål-trained ASR model's substitution error rate is about 25 percentage points higher on the mixed nn_no subset, so WER comparisons across these datasets are not clean measures of the same language.
- FLEURS's yue_hk subset is almost entirely Standard Written Chinese (89.8% of sampled prompts), not Cantonese, meaning model evaluations on that subset can appear successful while producing outputs in the wrong register.
- FLEURS's ar_eg subset is 98.6% Modern Standard Arabic rather than Egyptian Arabic, showing that locale labels in diglossic settings can encode the wrong variety.
- MCV17's nan_tw subset has 21 hours of raw audio but only 48.3% actual speech, and its text prompts are mostly dictionary-style single words or phrases written redundantly in two scripts with mismatched audio, making it nearly unusable without substantial restructuring.
- Dataset creators should conduct sociolinguistic assessment, prescribe orthography and register choices, enforce multi-level quality checks, and release detailed metadata as a standard part of building speech datasets for less-institutionalized languages.
Reading between the lines
- If the institutionalization-quality correlation extends beyond the audited languages, raw WER comparisons across low-resource languages should be treated as provisional until each subset passes a sociolinguistic audit; this is stricter than the paper's own recommendation to add human evaluation.
- The paper's framing implies a feedback loop it only gestures at: dataset creation is itself a language-planning intervention, so tracking how contributors' script and register choices shift across Common Voice versions for a digraphic language would provide a measurable way to watch community orthographic norms form.
- A testable extension of the paper's argument is that prescribing orthography and register rules before collection would lower downstream WER for a digraphic or diglossic language even when audio quantity is held constant, which would separate macro-level contamination from data scarcity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript audits three widely used multilingual speech datasets (Mozilla Common Voice 17.0, FLEURS, and VoxPopuli) for data quality issues, dividing them into micro-level issues (short durations, low speech proportion, topic imbalance, limited speaker diversity) and macro-level issues (unspecified writing systems, register/variety confusion, ambiguous dialect scope). It presents case studies for Norwegian Bokmål/Nynorsk, Arabic, Cantonese/Hong Kong Chinese, Fula, Kabuverdianu, and Taiwanese Southern Min, and it proposes a language-planning checklist for future dataset creation. The paper's headline claim is that macro-level issues are more prevalent in less-institutionalized, often under-resourced languages and that there is a 'strong positive correlation' between a language's institutionalization status and dataset quality.
Significance. The concrete case studies are valuable and, for the most part, well supported: the Norwegian WER experiment, the canto-filter analysis of Cantonese subsets, the nan_tw structure analysis, and the references to public issue reports give the audit a credible empirical core. The proposed guidelines for sociolinguistic assessment and proactive language planning are a useful contribution to dataset-creation practice, and the paper explicitly connects dataset construction to language planning in a way that is likely to stimulate debate. If the headline correlation were supported, the paper would have a strong systematic claim with clear consequences for low-resource ASR evaluation. As written, however, that central claim is asserted rather than demonstrated, which both overstates the evidence and obscures the genuinely solid case-study findings.
major comments (3)
- [Abstract and §1] The claim of a 'strong positive correlation between a language's institutionalization status and its dataset quality' is never measured or operationalized. No institutionalization variable is defined or scored, no per-language prevalence table is provided, and no correlation statistic is reported. The support offered in §4 is qualitative: selected case studies plus the observation that VoxPopuli, whose 16 European Parliament languages are mostly highly institutionalized, shows no macro-level issues. This is not a correlation. The claim appears in the abstract and introduction as a main result, so it is load-bearing; please either define and measure institutionalization and quality across a defined set of languages, or restrict the abstract and introduction to the observed case-study pattern rather than asserting a statistical relationship.
- [§1, Table 16, §8] The prevalence claim that macro-level issues are 'more prevalent' in less-institutionalized languages is not supported by the described qualitative protocol. The paper reports that around 40 languages were reviewed by volunteer native speakers at 100 randomly sampled sentences each, but it does not report inter-rater agreement, confidence intervals, how the roughly 40 languages were selected, or a uniform macro-issue verdict for every reviewed language. Macro-level issues such as dialect scope (ff_sn, kea_cv) are difficult to diagnose reliably from 100 sentences, and the appendix table lists languages without recording outcomes. Section 8's coverage limitation is honest, but it does not fix the unsupported prevalence statement, which is used to justify the paper's main claim. Please either provide the missing protocol details and per-language results, or weaken the prevalence statements to 'the reviewed languages exhibit...'.
- [§4.1–4.2] The classification results in Tables 2, 4, and 5 rest on marker-based scripts whose precision and recall are not reported. This matters because the Norwegian, Arabic, and Cantonese percentages are presented as quantitative evidence for the macro-level claims. The Norwegian WER experiment is a strong independent confirmation for that particular case, but the Arabic and Cantonese prevalence numbers would be more convincing with a validation set, a precision/recall estimate, or a release of the annotated samples used for manual verification. Please add such validation or explicitly label these numbers as approximate heuristic classifications.
minor comments (5)
- [§2 and Table 6] The paper says VoxPopuli has 16 European languages in §2 and in Table 6, but §3.2 and Figure 1 refer to '14 languages' in VoxPopuli; this inconsistency should be corrected, and the speech-proportion analysis should state which languages were excluded.
- [§3.2] The sentence 'While all 14 languages in V oxPopuli have at least 89% speech' is followed by 'the da_dk training set in FLEURS' with broken spacing; please copyedit the section for formatting errors such as 'theda_dktraining set'.
- [§4.3] The sentence 'This is significant because the the Guinean variant of Fula has the most speakers' contains a duplicated article; please fix the typo.
- [Appendix B] The text 'around 12 hours in each language' is misspelled as 'aroudn 12 hours'; please correct.
- [§4.2.2] The canto-filter package used in the Cantonese analysis is authored by one of the paper's co-authors; the text cites the package but should also state this authorship and affiliation explicitly for transparency.
Circularity Check
No significant circularity: the audit is empirical, and the sole self-citation is an independently published, open-source tool corroborated by external reports.
full rationale
The paper reports an empirical quality audit of three public speech datasets; it contains no fitted parameters, no derivation chain, and no prediction that is computed from its own inputs. The central generalization that macro-level issues are more prevalent in less-institutionalized languages is a qualitative synthesis of case studies and volunteer reviews. It may be under-supported, since no institutionalization variable is scored, no inter-rater agreement is reported, and the 100-sentence sample per language is small, but under-support is an evidentiary weakness, not circularity: the claim is not defined in terms of the dataset labels, and the paper does not fit a parameter and then rename that fit as a finding. The only self-citation is the canto-filter package of Lau et al. (2024), used in Section 4.2.2 to classify FLEURS yue_hk as Standard Written Chinese rather than Cantonese. That package is a previously published, open-source tool developed by one co-author, and its output is independently corroborated in the paper by a public Hugging Face discussion, the FLoRes GitHub issue #61, and the independent experiments of Xie and Chen (2025). The citation therefore provides external, checkable evidence rather than an unverified self-referential premise, and the paper's conclusions do not reduce to that citation by construction. Accordingly, no circular step can be exhibited under the required standard; the appropriate finding is a low score reflecting only a minor self-citation.
Assumptions & free parameters
assumptions (4)
- domain assumption The sociolinguistic categories digraphia and diglossia, as defined by Dale (1980) and Ferguson (1959), are applicable to the dataset languages under review.
- domain assumption ISO 639-3 and BCP 47 locale codes are intended to reflect actual linguistic content, so mismatches are defects.
- ad hoc to paper The classification marker lists for Norwegian, Arabic, and Cantonese are sufficient to distinguish written varieties.
- ad hoc to paper Community-driven dataset creation should incorporate language planning principles.
Cite this review
Pith. "Pith review of Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning." pith.science (2026). https://pith.science/paper/YQHFISXU
@misc{pith2026250617525,
author = {Pith},
title = {Pith review of: Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning},
year = {2026},
howpublished = {\url{https://pith.science/paper/YQHFISXU}},
note = {Machine review of arXiv:2506.17525}
}
read the original abstract
Our quality audit for three widely used public multilingual speech datasets - Mozilla Common Voice 17.0, FLEURS, and Vox Populi - shows that in some languages, these datasets suffer from significant quality issues, which may obfuscate downstream evaluation results while creating an illusion of success. We divide these quality issues into two categories: micro-level and macro-level. We find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages. We provide a case analysis of Taiwanese Southern Min (nan_tw) that highlights the need for proactive language planning (e.g. orthography prescriptions, dialect boundary definition) and enhanced data quality control in the dataset creation process. We conclude by proposing guidelines and recommendations to mitigate these issues in future dataset development, emphasizing the importance of sociolinguistic awareness and language planning principles. Furthermore, we encourage research into how this creation process itself can be leveraged as a tool for community-led language planning and revitalization.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
Balalaika is a data-centric annotation pipeline for Russian speech that combines semantic VAD, ASR ensembling, and prosody enrichment to build a 5.1k-hour corpus showing gains in denoising and TTS.
Reference graph
Works this paper leans on
-
[1]
Lin Alivin. 1999. https://sino-platonic.org/complete/spp089_taiwanese.pdf Writing Taiwanese : The development of modern written Taiwanese . Department of Asian and Middle Eastern Studies, University of Pennsylvania
work page 1999
-
[2]
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2020. https://aclanthology.org/2020.lrec-1.520/ Common Voice : A massively-multilingual speech corpus . In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 4218--4...
work page 2020
-
[3]
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick Von Platen, Yatharth Saraf, Juan Pino, et al. 2022. https://arxiv.org/abs/2111.09296 XLS-R : Self-supervised cross-lingual speech representation learning at scale . In Interspeech, Incheon, Korea
arXiv 2022
-
[4]
Lo \" c Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023. https://arxiv.org/abs/2308.11596 SeamlessM4T : Massively multilingual & multimodal machine translation . arXiv preprint arXiv:2308.11596
arXiv 2023
-
[5]
Martijn Bartelds, Nay San, Bradley McDonnell, Dan Jurafsky, and Martijn Wieling. 2023. https://doi.org/10.18653/v1/2023.acl-long.42 Making more of little data: Improving low-resource automatic speech recognition using data augmentation . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), volume 1, pages 715--...
-
[6]
Emily M. Bender and Batya Friedman. 2018. https://doi.org/10.1162/tacl_a_00041 Data statements for natural language processing: Toward mitigating system bias and enabling better science . Transactions of the Association for Computational Linguistics, 6:587--604
-
[7]
Joseph Lo Bianco. 2018. https://www.taylorfrancis.com/chapters/edit/10.4324/9781315561271-5/reinvigorating-language-policy-planning-intergenerational-language-revitalization-joseph-lo-bianco Reinvigorating language policy and planning for intergenerational language revitalization . In The Routledge handbook of language revitalization, pages 36--48. Routledge
work page doi:10.4324/9781315561271-5/reinvigorating-language-policy-planning-intergenerational-language-revitalization-joseph-lo-bianco 2018
-
[8]
Kristen Brustad. 2017. https://www.jstor.org/stable/10.1163/j.ctt1w76vkk.7 Diglossia as ideology . In The politics of written language in the Arab world, pages 41--67. Brill
Show all 63 references
-
[9]
Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Candido Junior, Anderson da Silva Soares, Sandra Alu \' sio, and Moacir Antonelli Ponti. 2023. https://doi.org/10.21437/Interspeech.2023-496 ASR data augmentation in low-resource settings using cross-lingual mul...
2023 doi
-
[10]
Chung-Cheng Chiu, Arun Narayanan, Wei Han, Rohit Prabhavalkar, Yu Zhang, Navdeep Jaitly, Ruoming Pang, Tara N Sainath, Patrick Nguyen, Liangliang Cao, et al. 2021. https://doi.org/10.1109/SLT48900.2021.9383518 RNN-T models fail to generalize to out-of-domain audio: Causes and ...
2021
-
[11]
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli. 2021. https://arxiv.org/abs/2006.13979 Unsupervised cross-lingual representation learning for speech recognition . In Interspeech, Brno, Czechia
2021 arXiv
-
[12]
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. 2023. https://doi.org/10.1109/SLT54892.2023.10023141 FLEURS : Few-shot learning evaluation of universal representations of speech . In IEEE Spoken Lang...
2023
-
[13]
Robert L. Cooper. 1989. https://doi.org/10.1017/CBO9780511620812 Language planning and social change . Cambridge UniversityPress
1989 doi
-
[14]
Marta R Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. https://arxiv.org/abs/2207.04672 No language left behind: Scaling human-centered machine translation . arX...
2022 arXiv
-
[15]
Ian R.H. Dale. 1980. https://doi.org/doi:10.1515/ijsl.1980.26.5 Digraphia . International Journal of the Sociology of Language, 1980(26):5--14
1980 doi
-
[16]
Afia Dil. 1986. https://doi.org/10.1515/9783110873641-076 Diglossia in Bangla : A study of shifts in the verbal repertoire of the educated classes in Dhaka, Bangladesh . The Fergusonian Impact, 2:451--65
1986 doi
-
[17]
Siyuan Feng, Bence Mark Halpern, Olya Kudina, and Odette Scharenborg. 2024. https://doi.org/10.1016/j.csl.2023.101567 Towards inclusive automatic speech recognition . Computer Speech & Language, 84
2024
-
[18]
Ferguson
Charles A. Ferguson. 1959. https://doi.org/10.1080/00437956.1959.11659702 Diglossia . Word, 15(2):325--340
1959
-
[19]
Ferguson
Charles A. Ferguson. 1996. Epilogue: diglossia revisited. Understanding Arabic: Essays in contemporary Arabic linguistics in honor of El-Said Badawi , pages 49--67
1996
-
[20]
Mahault Garnerin, Solange Rossato, and Laurent Besacier. 2021. https://doi.org/10.18653/v1/2021.gebnlp-1.10 Investigating the impact of gender representation in ASR training data: A case study on Librispeech . In 3rd Workshop on Gender Bias in Natural Language Processing, page...
2021 doi
- [21]
-
[22]
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingual ...
2022 doi
-
[23]
Jacob Høigilt and Gunvor Mejdell. 2017. https://doi.org/10.1163/9789004346178 The Politics of Written Language in the Arab World: Writing Change . Brill, Leiden, The Netherlands
2017 doi
-
[24]
Youngjoo Jung and Bora Kim. 2023. https://doi.org/10.1177/18793665231188380 Coexistence of multiple writing systems: Classifying digraphia in post-socialist countries . Journal of Eurasian Studies, 14(2):139--150
2023 doi
-
[25]
Shigeki Karita, Richard Sproat, and Haruko Ishikawa. 2023. https://doi.org/10.18653/v1/2023.cawl-1.8 Lenient evaluation of J apanese speech recognition: Modeling naturally occurring spelling inconsistency . In Proceedings of the Workshop on Computation and Written Language (CA...
2023 doi
-
[26]
Rickford, Dan Jurafsky, and Sharad Goel
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R. Rickford, Dan Jurafsky, and Sharad Goel. 2020. https://doi.org/10.1073/pnas.1915768117 Racial disparities in automated speech recognition . Proceedings of the National Ac...
2020 doi
-
[27]
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2022. https://doi.org/10.1162/tacl_a_00447 Quality at a glance: An audit of web-crawled multilingual data...
2022 doi
-
[28]
Anuj Kumar, Pooja Reddy, Anuj Tewari, Rajat Agrawal, and Matthew Kam. 2012. https://doi.org/10.1145/2207676.2208564 Improving literacy in developing countries using speech recognition-supported games on mobile devices . In Proceedings of the SIGCHI Conference on Human Factors ...
2012
-
[29]
Chaak-ming Lau, Mingfei Lau, and Ann Wai Huen To. 2024. https://aclanthology.org/2024.eurali-1.4 The extraction and fine-grained classification of written C antonese materials through linguistic feature detection . In Proceedings of the 2nd Workshop on Resources and Technologi...
2024
-
[30]
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/2d8911db9ecedf866015091b28946e15-Paper-Conference.pdf Voice...
2023
-
[31]
Viktor Leggio and Yaron Matras
D. Viktor Leggio and Yaron Matras. 2017. https://doi.org/10.1017/9781316562949 Orthography development on the internet: Romani on YouTube . Creating orthographies for endangered languages, page 254–275
2017 doi
-
[32]
Yuanchao Li, Peter Bell, and Catherine Lai. 2022. https://doi.org/10.1109/ICASSP43922.2022.9746289 Fusing ASR outputs in joint training for speech emotion recognition . In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7362--7366, Singapore
2022
-
[33]
Yuanchao Li, Zeyu Zhao, Ondrej Klejch, Peter Bell, and Catherine Lai. 2023. https://doi.org/10.21437/Interspeech.2023-2078 ASR and emotional speech: A word-level investigation of the mutual impact of speech and emotion recognition . In Interspeech, pages 1244--1248, Dublin, Ireland
2023 doi
-
[34]
Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Paden Tomasello, Jacob Kahn, Gilad Avidov, Ronan Collobert, and Gabriel Synnaeve. 2021. https://doi.org/10.21437/Interspeech.2021-1758 Rethinking evaluation in ASR : Are our models robust enough? In Interspeech, pages 311--315,...
2021 doi
-
[35]
Behrooz Mahmoodi-Bakhtiari. 2018. https://doi.org/doi:10.1515/9783110455793-011 Spoken vs. written Persian : Is Persian diglossic? , pages 183--212. De Gruyter Mouton, Berlin, Boston
2018 doi
-
[36]
Nina Markl and Stephen Joseph McNulty. 2022. https://aclanthology.org/2022.lrec-1.680/ Language technology practitioners as language managers: arbitrating data bias and predictive bias in ASR . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, page...
2022
-
[37]
Gleb Mazovetskiy and Taku Kudo. 2024. https://jedworkshop.github.io/JLR2024/materials/b-2.pdf Data processing for Japanese text-to-pronunciation models. JLR workshop at the 30th annual conference of the association for natural language processing
2024
-
[38]
Teresa L. McCarty. 2018. https://doi.org/10.4324/9781315561271 Community-based language planning: Perspectives from indigenous language revitalization . In The Routledge handbook of language revitalization, pages 22--35. Routledge
2018 doi
-
[39]
Dalai Mengke, Yan Meng, and Peter Mihajlik. 2024. https://aclanthology.org/2024.sigul-1.40/ Tandem long-short duration-based modeling for automatic speech recognition . In Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages (SIGUL),...
2024
-
[40]
Sainath, and Trevor Strohman
Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach, Tara N. Sainath, and Trevor Strohman. 2019. https://doi.org/10.1109/ASRU46091.2019.9003913 Recognizing long-form speech using streaming end-to-end models . In Automatic Speech Recognition and Understanding Wor...
2019
-
[41]
Iuliia Nigmatulina, Tannon Kew, and Tanja Samardzic. 2020. https://aclanthology.org/2020.vardial-1.2 ASR for non-standardised languages with dialectal variation: the case of S wiss G erman . In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialect...
2020
-
[42]
Katsuhiro J. Ota. 2005. http://hdl.handle.net/10125/11520 An investigation of written Taiwanese . Master's thesis, University of Hawaii, August
2005
-
[43]
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2024. https://dl.acm.org/doi/10.5555/3666122.3669586 The RefinedWeb dataset for Falcon LLM : outperforming cura...
2024
-
[44]
Matthew Perez, Mimansa Jaiswal, Minxue Niu, Cristina Gorrostieta, Matthew Roddy, Kye Taylor, Reza Lotfian, John Kane, and Emily Mower Provost. 2022. https://doi.org/10.21437/Interspeech.2022-10943 Mind the gap: On the value of silence representations to lexical-based speech em...
2022 doi
-
[45]
Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2024. http://jmlr.org/papers/v25/23-1318.html Scaling speech technology to 1,000+ languages . Journal of Machine Learning Res...
2024
-
[46]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. https://dl.acm.org/doi/10.5555/3618408.3619590 Robust speech recognition via large-scale weak supervision . In International Conference on Machine Learning (ICML), pages 28492--28...
2023
-
[47]
Jaiswal, Shubham Prakash, Bendi Pragnya Sree, and Animesh Mukherjee
Anand Kumar Rai, Siddharth D. Jaiswal, Shubham Prakash, Bendi Pragnya Sree, and Animesh Mukherjee. 2024. https://doi.org/10.48550/arXiv.2410.16712 DENOASR: Debiasing ASRs through Selective Denoising . In IEEE International Conference on Knowledge Graphs (ICKG), Abu Dhabi, UAE
-
[48]
Thomas Reitmaier, Dani Kalarikalayil Raju, Ondrej Klejch, Electra Wallington, Nina Markl, Jennifer Pearson, Matt Jones, Peter Bell, and Simon Robinson. 2024. https://doi.org/10.1145/3613904.3642026 Cultivating spoken language technologies for unwritten languages . In Proceedin...
2024
-
[49]
Gerald Roche. 2017. https://doi.org/doi:10.1515/ijsl-2017-0001 Introduction: The transformation of Tibet's language ecology in the twenty-first century . International Journal of the Sociology of Language, 2017(245):1--35
2017 doi
-
[50]
Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zal \'a n Borsos, F \'e lix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zal \'a n Borsos, F \'e lix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023. https://doi.org/doi:10.48550/arXiv.2306.12925 AudioPaLM : A large language model th...
-
[51]
Kathleen D Sackett. 2017. Community-driven, goal-centred orthography development: A tsakhur case study. Mari C. Jones and Damien Mooney, Creating Orthographies for Endangered Languages, pages 88--108
2017
-
[52]
Don Snow. 2010. https://doi.org/doi:10.1515/ijsl.2010.052 Hong Kong and modern diglossia . International Journal of the Sociology of Language, 2010(206):155--179
2010 doi
-
[53]
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, and Team. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1211 C asablanca: Data and models for multidialectal A rabic speech recognition . In Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2...
2024 doi
-
[54]
Zhiyuan Tang, Dong Wang, Yanguang Xu, Jianwei Sun, Xiaoning Lei, Shuaijiang Zhao, cheng wen, Xingjun Tan, Chuandong Xie, Shuran Zhou, Rui Yan, Chenjia Lv, Yang Han, Wei Zou, and Xiangang Li. 2021. https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/0...
2021
-
[55]
Chara Tsoukala, Kosmas Kritsis, Ioannis Douros, Athanasios Katsamanis, Nikolaos Kokkas, Vasileios Arampatzakis, Vasileios Sevetlidis, Stella Markantonatou, and George Pavlidis. 2023. https://doi.org/10.18653/v1/2023.fieldmatters-1.5 ASR pipeline for low-resourced languages: A ...
2023 doi
-
[56]
Ehsan Variani, David Rybach, Cyril Allauzen, and Michael Riley. 2020. https://doi.org/10.1109/ICASSP40776.2020.9053600 Hybrid autoregressive transducer ( HAT ) . In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6139--6143
2020
-
[57]
Dusan Varis and Ond r ej Bojar. 2021. https://aclanthology.org/2021.emnlp-main.650/ Sequence length is a domain: Length-based overfitting in transformer models . In Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8246--8257
2021
-
[58]
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021. https://doi.org/10.18653/v1/2021.acl-long.80 VoxPopuli : A large-scale multilingual speech corpus for representation learning, semi-super...
2021 doi
-
[59]
Williams, Don Nix, and Peter Fairweather
Susan M. Williams, Don Nix, and Peter Fairweather. 2000. https://archive.isls.org/conferences/icls/2000/proceedings/abstracts/ab115.html Using speech recognition technology to enhance literacy instruction for emerging readers . In International Conference of the Learning Scien...
2000
-
[60]
Peng Xie and Kani Chen. 2025. https://arxiv.org/abs/2310.17953 Developing a multilingual dataset and evaluation metrics for code-switching: A focus on Hong Kong's polylingual dynamics . In International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1--5
2025 arXiv
- [61]
-
[62]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[63]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.