Pith. sign in

REVIEW 3 major objections 5 minor 43 references

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that NaijaVoices, a newly collected 1,800-hour, 5,455-speaker corpus for Igbo, Hausa, and Yoruba, is the first dataset at this scale for these languages, and that finetuning on it cuts word error rates by 42–76% relative…

desk verdict A genuinely large new speech corpus for Igbo, Hausa, and Yoruba, with credible external FLEURS gains, but the headline WER numbers are inflated by a likely sentence-level train/test leakage in the internal test set. read the letter →

arxiv 2505.20564 v3 pith:GZFSGUH2 submitted 2025-05-26 cs.CL

classification cs.CL
keywords NaijaVoicesautomaticspeechrecognitionIgboHausaYorubadatasetlow-resourcelanguagesdatafarming
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces NaijaVoices, a speech-text dataset of 1,838 hours from 5,455 speakers across Igbo, Hausa, and Yoruba, built from culturally grounded sentences that were manually composed rather than scraped from the web. The authors claim that finetuning standard ASR models on this corpus produces large improvements: average relative word error rate reductions of 75.86% for Whisper small, 52.06% for MMS base, and 42.33% for XLSR base. They also find that small finetuned models can outperform a much larger unfinetuned state-of-the-art model, suggesting that dataset scale and quality matter more than model size in this setting. The paper argues that its 'data farming' collection approach, in which trained facilitators supervise community voice donors, solves the usual tension between scalability and recording quality.

What carries the argument

The load-bearing mechanism is the 'data farming' pipeline: 144 sentence generators (48 per language) manually compose culturally aligned sentences guided by 100 themes; a double-review system checks the generated texts for correctness, including diacritics; then trained Facilitators supervise Voice Donors in their communities as they record each sentence twice using a recording app, with a second review of random recordings for quality. This pipeline produces 1,917,686 recordings with 645,138 unique sentences, high speaker diversity, and controlled audio quality, and it is what the paper credits for the WER gains.

What would settle it

Take a random sample of about 500 recordings from the released dataset, have independent native-speaker linguists transcribe them from scratch with full diacritics, and measure the agreement rate with the published transcripts. A low agreement rate, or a finetuned model whose error rate does not improve on an independently recorded corpus collected through a different protocol, would undercut the claim that NaijaVoices is both large and high-quality.

Watch

Extended reading notes

Core claim

The central discovery is that a large, culturally diverse, community-collected speech corpus can dramatically improve automatic speech recognition for three under-resourced African languages. Concretely, finetuning Whisper small for five epochs on NaijaVoices lowers its average WER from 225.47% to 54.43%, with similar large gains for MMS and XLSR; the finetuned small models also beat the 2.3B-parameter Seamless M4T v2 unfinetuned baseline. The paper further shows that the corpus covers a wider acoustic space than Common Voice for Hausa and Yoruba, and that most recordings fall within a clean-audio signal-to-noise threshold.

Load-bearing premise

The quality argument rests on the assumption that the 1.9 million recordings are paired with accurate, diacritic-correct transcripts, a guarantee the paper bases on its double review of generated sentences and a spot-check of a random audio subset rather than on independent verification.

Editorial extensions

If this is right

  • If the WER reductions hold, NaijaVoices could serve as a foundation dataset for building production speech systems in Igbo, Hausa, and Yoruba, which currently lack support in major voice-enabled technologies.
  • The Phase 2 subset, with 6,000+ sentences per speaker, provides the long-form recordings needed for text-to-speech synthesis as well as ASR.
  • The finding that monolingual finetuning works better for Igbo while multilingual finetuning helps Yoruba and Hausa suggests practical guidance for how to allocate training data across related languages.
  • Because small finetuned models outperform a much larger unfinetuned model, the dataset could enable accurate ASR on devices with limited compute.
  • The CC BY-NC-SA license makes the corpus publicly available, so future work can reproduce and build on these results directly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the 100 culturally themed prompts and the sentence-generation workflow could be reused as a template for creating similar corpora in other under-resourced languages, potentially accelerating speech data cultivation across Africa and beyond.
  • Beyond the paper: the internal NV Test set is drawn from the same pipeline as the training data, so the reported internal WER gains may partly reflect distribution overlap; the FLEURS results, where the models still improve but by smaller margins, are the more conservative evidence of generalization.
  • Beyond the paper: the dataset's demographic spread (65.9% aged 18–29, 25.7% aged 6–17) means it may be less representative of older speakers, so testing on older-voice recordings would reveal whether the corpus supports age-balanced ASR.
  • Beyond the paper: the audio-quality analysis relies on a standard SNR estimator, but the paper does not report how the highest-SNR fraction correlates with transcript accuracy, so a joint analysis of audio quality and transcription correctness would be a natural next check.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces NaijaVoices, a 1,838-hour, 5,455-speaker speech-text corpus for Igbo, Hausa, and Yoruba, built through a 'data farming' pipeline of manual sentence generation, double review, facilitator-mediated recording, and community engagement. The authors report acoustic diversity and SNR analyses, and they fine-tune Whisper, XLSR, MMS, and SeamlessM4T models, claiming large ASR improvements, including average relative WER reductions of 75.86% for Whisper, 52.06% for MMS, and 42.33% for XLSR in the Abstract and Conclusion. The paper argues that this is the first corpus of this scale for the three target languages and that it demonstrates a scalable, culturally grounded data collection model.

Significance. If the reported results hold, NaijaVoices is a major community resource: it is an order of magnitude larger in hours and speakers than existing corpora for Igbo, Hausa, and Yoruba, and its detailed account of sentence generation, review, and community-based recording is a valuable methodological template. The authors deserve credit for releasing the dataset, for evaluating on the independent FLEURS benchmark, and for reporting concrete per-language WERs rather than only aggregate claims. The FLEURS results provide genuine evidence that the dataset improves ASR on an externally created test set. However, the headline generalization claims rest on an internal test set whose construction appears to allow sentence-level leakage, and the experimental section provides no variance information. These issues affect the central quantitative claim and require revision before the paper's main conclusions can be accepted at face value.

major comments (3)
  1. [4.1 (with 2.2 and 3.1)] The internal NV Test is speaker-disjoint but not sentence-disjoint. Section 2.2 states that each sentence was recorded twice and that no voice donor recorded the same sentence more than once, and Section 3.1 reports 1,917,686 recordings over 645K unique sentences, so many written sentences have multiple distinct audio recordings by different speakers. The stratified split in Section 4.1 (99.7% train, 0.3% dev/test, plus all samples from 20 unseen speakers) removes speaker overlap but does not prevent identical transcripts from appearing in both training and test. This allows the models to memorize sentence-level vocabulary and language-model patterns, artificially lowering the internal NV Test WERs in Table 3 and inflating the headline relative reductions (75.86% Whisper, 52.06% MMS, 42.33% XLSR) reported in the Abstract and Conclusion. The FLEURS results are not affected by this leakage, but they do not validate the specific internal WER magnitudes. I request a sentence-disjoint split (e.g., group by sentence ID before splitting) and a re-reporting of all headline numbers, or at minimum a FLEURS-only headline together with a clear caveat that the NV Test measures matched-transcript conditions.
  2. [Table 3 / Section 4.2] All results are single runs with no variance, confidence intervals, or significance tests. Several comparisons that drive the discussion, such as Igbo NV Test monolingual XLSR 41.54 versus multilingual 52.44, or Yoruba FLEURS 76.43 versus 91.28, may be within run-to-run noise, especially with five epochs of fine-tuning and per-language test sets of roughly one hour. Please report multiple seeds or error bars, or soften the comparative claims about which fine-tuning configuration is best.
  3. [2.1, 2.2, 3.3] The 'high-quality' claim and the internal WER numbers depend on transcript accuracy, but no quantitative transcript-quality metrics are reported. The double-review process is described, yet there is no inter-reviewer agreement rate, post-hoc error rate, or analysis of diacritic consistency, and the second review covers only a random subset of recordings. If systematic transcription or diacritic errors exist, internal NV Test WERs will be artificially low. Please report a measured transcription error rate on a held-out sample and, if possible, a diacritic-consistency study.
minor comments (5)
  1. [3.3 / Figure 3] The WADA-SNR interpretation appears contradictory: the caption says a value of 100 means that the audio has a great proportion of silence with respect to speech, while the text says 65% of samples achieve the highest possible SNR score and the remaining samples are within the clean audio threshold (SNR 20-100). Please clarify whether high WADA-SNR scores indicate clean speech or silence, and report the actual SNR distribution.
  2. [Table 2] The maximum Hausa duration of 1,321.08 seconds (roughly 22 minutes) is implausible for a sentence recording and should be verified or excluded; it may also affect the reported total hours.
  3. [4.1] The construction of the NV Test from both the stratified 0.3% split and the additional 20 unseen speakers needs a precise description; the current text suggests the two procedures overlap or double-count, and a small diagram or pseudocode would remove ambiguity.
  4. [Abstract / Index Terms] The index terms contain the typo 'datatset' and the terms are run together without separators; please fix the spelling and format the list properly.
  5. [1] The claim that no voice-enabled technology supports any of the 2,000+ African languages is very broad for a citation to a 2020 Mozilla blog post; please qualify the claim by date, scope, and the set of technologies considered.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: ASR gains are empirical, evaluated on external FLEURS, and not derived from fitted parameters or self-citation.

full rationale

The paper's central claims are an empirical dataset contribution, not a derivation from fitted parameters. The ASR experiments fine-tune standard pretrained models (Whisper, MMS, XLSR) on NaijaVoices and evaluate on both an internal NV Test split and the externally constructed FLEURS benchmark. Relative WER reductions are measured outcomes, not quantities fitted to the data being predicted, so the 'fitted input called prediction' and 'self-definitional' patterns do not apply. No load-bearing self-citation is present: the 'data farming' blog footnote is framing, and the MMS feature extractor and WADA-SNR are external tools used for analysis. The one substantive caveat is that the NV Test split is speaker-disjoint but not explicitly sentence-disjoint: Section 2.2 says 'each sentence recorded twice' and Section 4.1 describes stratified sampling plus 20 unseen speakers, so a test utterance's transcript may also occur in training, which could inflate the internal NV Test WERs and the averages that include them. This is an evaluation-leakage/correctness concern rather than circularity under the strict definition, because the test audio comes from unseen speakers and the FLEURS results are independent of the NaijaVoices training text. The dataset contribution is therefore self-contained against an external benchmark, and the circularity score is low.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper does not posit new physical entities or fitted constants. Its central claims rest on dataset quality assumptions (transcript accuracy, audio cleanliness, test representativeness) and on external tooling (WADA-SNR, MMS embeddings, FLEURS) whose validity for these languages is taken as given.

assumptions (5)
  • domain assumption WADA-SNR value 100 indicates clean, high-quality audio
    Section 3.3 uses WADA-SNR and treats 100 as the highest possible SNR and 20-100 as a clean audio threshold, quoting [30]. This interpretation is imported from external tooling and is not validated on this data.
  • domain assumption MMS encoder embeddings capture meaningful acoustic diversity
    Section 3.2 uses averaged MMS 1B embeddings as a proxy for acoustic diversity without validating that this proxy correlates with human-perceived diversity of accents and styles.
  • domain assumption FLEURS is a valid external benchmark for Igbo, Hausa, and Yoruba
    Section 4.1 evaluates on FLEURS and assumes its transcripts, pronunciation, and test splits are correct for these three languages.
  • domain assumption NV Test set from 20 unseen speakers with the fewest samples is representative
    Section 4.1, Data Split constructs the internal test set from a non-random selection of speakers (those with fewest samples); this assumes the resulting split measures generalization to the broader population.
  • domain assumption Using the English Whisper tokenizer for Igbo is adequate for evaluation
    Section 4.1 states the English tokenizer was used for Igbo because Whisper's tokenizer lacks Igbo; this may reduce measured performance and is unverified as a fair evaluation choice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages." pith.science (2026). https://pith.science/paper/GZFSGUH2

@misc{pith2026250520564,
  author       = {Pith},
  title        = {Pith review of: The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GZFSGUH2}},
  note         = {Machine review of arXiv:2505.20564}
}
read the original abstract

The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient data. Popular voice-enabled technologies do not support any of the 2000+ African languages, limiting accessibility for circa one billion people. While previous dataset efforts exist for the target languages, they lack the scale and diversity needed for robust speech models. To bridge this gap, we introduce the NaijaVoices dataset, a 1,800-hour speech-text dataset with 5,000+ speakers. We outline our unique data collection approach, analyze its acoustic diversity, and demonstrate its impact through finetuning experiments on automatic speech recognition, averagely achieving 75.86% (Whisper), 52.06% (MMS), and 42.33% (XLSR) WER improvements. These results highlight NaijaVoices' potential to advance multilingual speech processing for African languages.

Figures

Figures reproduced from arXiv: 2505.20564 by the authors.

Figure 1
Figure 1. Illustration of the creation of the NaijaVoices dataset. Beginning with the sentence generation phase where a group of language experts, called sentence generators, composed sentences based on themes/prompts provided. This generation phase was followed by a double review system, leading to over 1.9M texts which were recorded with our unique approach in the recording phase. Here, the ‘Facilitator’ finds and closely g… view at source ↗
Figure 2
Figure 2. Acoustic diversity analysis of NaijaVoices and Common Voice datasets for Hausa and Yoruba languages. Using the unfinetuned MMS 1B model [2], we extract 1,280- dimensional feature vectors (from the last layer of the en￾coder), reduce them to two dimensions with PCA [28], and visualize them [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Distribution of SNR values from audio samples in the NaijaVoices dataset, showing that the majority of audio samples lie within 100. According to [29], a value of 100, the highest possible value, means that the audio has a great proportion of silence w.r.t speech. 4. Automatic Speech Recognition We perform automatic speech recognition (ASR) experi￾ments to demonstrate the potential of the NaijaVoices dataset for mul… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 36 canonical work pages

  1. [1]

    While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8]

    Introduction & Prior Work The importance of large, high-quality, and diverse training data in the performance of speech technologies cannot be overstated, as substantial datasets are essential for building high-performing speech processing models [1, 2, 3, 4, 5, 6]. While notable progress has been made in speech processing, African languages – including o...

  2. [2]

    Figure 1 illustrates our pipeline for creating the NaijaV oices dataset

    How NaijaVoices was created Our dataset creation process operates on our coined ethos of ‘data farming’1 which, in contrast to data mining which extracts and depletes data from providing communities, em- ploys a reciprocal relationship (akin to farming) which en- sures that providing communities are engaged in, empow- ered by, and mutually benefit from th...

  3. [3]

    It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences

    The NaijaVoices Dataset The NaijaV oices dataset captures the essence of Nigerian culture through authentic, expert-generated, and culturally rich sentences, offering a level of originality rarely found in online texts [21]. It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to forma...

  4. [4]

    Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20]

    Automatic Speech Recognition We perform automatic speech recognition (ASR) experi- ments to demonstrate the potential of the NaijaV oices dataset for multilingual speech research. Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20]. 4.1. Experimental Setup Model.For...

  5. [5]

    Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities

    Conclusion With the NaijaV oices dataset, we demonstrated the possi- bility of cultivating speech data for African languages at an unprecedented large scale (in terms of number of hours and speakers). Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities. Our dataset demonstrates unparalleled ac...

  6. [6]

    Acknowledgement We express our profound gratitude to the entire NaijaV oices community for making this dataset possible, and the Lacuna Fund for funding the creation of the NaijaV oices dataset. Finally we acknowledge the support of the IV ADO and the Canada First Research Excellence Fund (CFREF) / Apog´ee Funds, the Canada CIFAR AI Chairs Program, as wel...

  7. [7]

    Robust speech recognition via large-scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,”ICML, 2022

  8. [8]

    Scaling speech technology to 1,000+ languages,

    V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babuet al., “Scaling speech technology to 1,000+ languages,”JMLR, vol. 25, no. 97, pp. 1–52, 2024

Show all 43 references
  1. [9]

    The data provenance initiative: A large scale audit of dataset licensing & attribution in AI,

    S. Longpre, R. Mahari, A. Chen, N. Obeng-Marnu, D. Sileo et al., “The data provenance initiative: A large scale audit of dataset licensing & attribution in AI,”arXiv preprint arXiv: 2310.16787, 2023

  2. [10]

    IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,

    T. Javed, J. Nawale, E. George, S. Joshiet al., “IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,” inFindings of ACL 2024, aug 2024, pp. 10 740–10 782

  3. [11]

    IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,

    A. Sankar, S. Anand, P. S. Varadhan, S. Thomaset al., “IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,”NeurIPS DB, 2024

  4. [12]

    BASE TTS: Lessons from building a billion-parameter text-to- speech model on 100k hours of data,

    M. Łajszczak, G. C´ambara, Y . Li, F. Beyhanet al., “BASE TTS: Lessons from building a billion-parameter text-to- speech model on 100k hours of data,”arXiv preprint arXiv: 2402.08093, 2024

  5. [13]

    Replication data for Igbo Natural Language Processing Tasks I,

    G. O. Nweya, A. S. Oluwole, E. F. Onwuegbuzia, S. O. Ejinwaet al., “Replication data for Igbo Natural Language Processing Tasks I,” 2022. [Online]. Available: https://doi.org/10.7910/DVN/RXBNCZ

  6. [14]

    Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,

    T. `Og´unr`em´ı, C. D. Manning, and D. Jurafsky, “Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,”CALCS, 2023

  7. [15]

    Masakhane - machine translation for Africa,

    I. Orife, J. Kreutzer, B. Sibanda, D. Whitenacket al., “Masakhane - machine translation for Africa,”arXiv preprint arXiv: 2003.11529, 2020

  8. [16]

    Partici- patory research for low-resourced machine translation: A case study in African languages,

    W. Nekoto, V . Marivate, T. Matsila, T. Fasubaaet al., “Partici- patory research for low-resourced machine translation: A case study in African languages,” inFindings of EMNLP 2020, Online, nov 2020, pp. 2144–2160

  9. [17]

    The state and fate of linguistic diversity and inclusion in the NLP world,

    P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, “The state and fate of linguistic diversity and inclusion in the NLP world,” inProceedings of ACL, Online, jul 2020, pp. 6282–6293

  10. [18]

    A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,

    D. I. Adelani, J. O. Alabi, A. Fan, J. Kreutzer, X. Shen, M. Reid, D. Ruiter, D. Klakowet al., “A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,”NAACL, 2022

  11. [19]

    R. Muhire. (2020) How Rwanda is making voice tech more open. [Online]. Available: https://foundation.mozilla.org/en/ blog/how-rwanda-making-voice-tech-more-open/

  12. [20]

    GlobalPhone: A multi- lingual text & speech database in 20 languages,

    T. Schultz, N. T. Vu, and T. Schlippe, “GlobalPhone: A multi- lingual text & speech database in 20 languages,” inICASSP. IEEE, 2013, pp. 8126–8130

  13. [21]

    YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,

    K. Olaleye, D. Oneat ¸˘a, and H. Kamper, “YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,” inSLT 2022. IEEE, 2023, pp. 731–738

  14. [22]

    `Ir`oy`ınspeech: A multi-purpose Yor`ub´a speech corpus,

    T. `Og´unr`em´ı, K. T ´ubos´un, A. Anuoluwapo, I. Orife, and D. I. Adelani, “`Ir`oy`ınspeech: A multi-purpose Yor`ub´a speech corpus,”Proceedings of LREC, 2023. [Online]. Available: https://arxiv.org/abs/2307.16071v2

  15. [23]

    V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,

    O. Ahia, A. Aremu, D. Abagyan, H. Gonen, D. I. Adelani et al., “V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,”Findings of EMNLP, 2024

  16. [24]

    BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,

    J. Meyer, D. Adelani, E. Casanova, A. ¨Oktemet al., “BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,” inINTERSPEECH, 2022, pp. 2383– 2387

  17. [25]

    Common V oice: A massively-multilingual speech corpus,

    R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyeret al., “Common V oice: A massively-multilingual speech corpus,” in Proceedings of LREC 2020. European Language Resources Association, 2020, pp. 4218–4222. [Online]. Available: https://aclanthology.org/2020.lrec-1.520/

  18. [26]

    FLEURS: Few-shot learning evaluation of universal representations of speech,

    A. Conneau, M. Ma, S. Khanuja, Y . Zhanget al., “FLEURS: Few-shot learning evaluation of universal representations of speech,”SLT, 2022

  19. [27]

    Quality at a glance: An audit of web-crawled multilingual datasets,

    J. Kreutzer, I. Caswell, L. Wang, A. Wahabet al., “Quality at a glance: An audit of web-crawled multilingual datasets,” TACL, 2021

  20. [28]

    Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,

    I. Abdulmumin, M. Beukman, J. O. Alabi, C. C. Emezue, E. Asiko, T. P. Adewumi, S. H. Muhammad, M. Adeyemi, O. Yousuf, S. Singh, and T. Gwadabe, “Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,”WM ˆE, 2022

  21. [29]

    The FineWeb Datasets: Decanting the web for the finest text data at scale,

    G. Penedo, H. Kydl´ıˇcek, L. B. Allal, A. Lozhkov, M. Mitchell et al., “The FineWeb Datasets: Decanting the web for the finest text data at scale,”arXiv preprint arXiv: 2406.17557, 2024

  22. [30]

    JW300: A wide-coverage parallel corpus for low-resource languages,

    ˇZ. Agi ´c and I. Vuli ´c, “JW300: A wide-coverage parallel corpus for low-resource languages,” inProceedings of ACL, A. Korhonen, D. Traum, and L. M `arquez, Eds. Florence, Italy: Association for Computational Linguistics, jul 2019, pp. 3204–3210

  23. [31]

    `Ir`oy`ınspeech: Yor`ub´a speech corpus,

    T. `Og´unr`em´ı, K. T´ubos´un, A. Anuoluwapo, I. Orife, and D. I. Adelani, “`Ir`oy`ınspeech: Yor`ub´a speech corpus,”Proceedings of LREC, 2023

  24. [32]

    Hausa speech corpus,

    U. A. Ibrahim, “Hausa speech corpus,” 2021. [Online]. Available: https://doi.org/10.17632/j6kjmfrbby.2

  25. [33]

    Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,

    B. Wanjawa, L. D. A. Wanzare, F. Indede, O. McOnyango, E. Ombui, and L. Muchemi, “Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,”Journal for Language Technology and Com- putational Linguistics, 2022

  26. [34]

    LIII. On lines and planes of closest fit to systems of points in space,

    K. P. F.R.S., “LIII. On lines and planes of closest fit to systems of points in space,”Philosophical Magazine Series 1, vol. 2, pp. 559–572, 1901

  27. [35]

    Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,

    C. Kim and R. M. Stern, “Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,” in INTERSPEECH, 2008, pp. 2598–2601

  28. [36]

    Audio quality feature,

    Soapbox Labs, “Audio quality feature,” 2025, documentation. Accessed: 2025-02-06. [Online]. Available: https://docs. soapboxlabs.com/guides-&-tutorials/audio-quality-feature/

  29. [37]

    (n.d.) Evaluation

    Hugging Face. (n.d.) Evaluation. Chapter 5 of the Audio Course. [Online]. Available: https://huggingface.co/learn/ audio-course/en/chapter5/evaluation

  30. [38]

    Unsupervised cross-lingual representation learning for speech recognition,

    A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,” inINTERSPEECH, 2021, pp. 2426– 2430

  31. [39]

    Seamless: Multilingual expressive and streaming speech translation,

    Seamless Communication, L. Barrault, Y .-A. Chung, M. C. Meglioli, D. Dale, N. Donget al., “Seamless: Multilingual expressive and streaming speech translation,”ArXiv, vol. abs/2312.05187, 2023

  32. [40]

    Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,

    K. Ogueji, Y . Zhu, and J. Lin, “Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,” inProceedings of the 1st Workshop on Multilingual Representation Learning, D. Ata- man, A. Birch, A. Conneau, O. Firat,...

  33. [41]

    Data collection and quality challenges in deep learning: A data-centric ai perspective,

    S. E. Whang, Y . Roh, H. Song, and J.-G. Lee, “Data collection and quality challenges in deep learning: A data-centric ai perspective,”arXiv preprint arXiv: 2112.06409, 2021

  34. [44]

    What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,

    X. Yu, Z. Zhang, F. Niu, X. Hu, X. Xia, and J. Grundy, “What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,”39th International Confer- ence on Automated Software Engineering (ASE), pp. 656–668, 2024

  35. [2022]

    Available: https://arxiv.org/abs/2203.06404v1

    [Online]. Available: https://arxiv.org/abs/2203.06404v1

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.