Pith. sign in

REVIEW 3 major objections 6 minor 17 references

Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A year of grassroots collection delivered 30,000 sentences per language and 268 hours of speech for three Kenyan languages, the paper reports.

desk verdict The corpora are a real contribution; the manuscript just doesn't show them, and that gap is fixable before publication. read the letter →

arxiv 2501.11003 v1 pith:AJANWMQG submitted 2025-01-19 cs.CL

classification cs.CL
keywords NaturallanguageprocessingLow-resourcelanguagesAfricanCorpusbuildingCrowdsourcingKidaw'idaKalenjinDholuo
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports a one-year, community-based effort to build the first NLP corpora for Kidaw'ida and Kalenjin, and an additional parallel corpus for Dholuo, three under-resourced Kenyan languages. The authors claim they collected 30,000 sentences per language, translated each into Kiswahili, and made the resulting three parallel corpora freely downloadable on Zenodo. They also report 56, 92, and 120 hours of speech data for the three languages on Mozilla Common Voice, with voice collection still ongoing. If the repositories contain what the paper describes, these resources lower the barrier to machine translation, speech recognition, and other NLP applications for languages that currently have almost no digital data.

What carries the argument

The load-bearing mechanism is a two-track collection pipeline: contributors write or transcribe sentences in the target language, bilingual contributors translate them into Kiswahili, and Data Collection Leads—native speakers with high language proficiency—check spelling, grammar, fluency, and translation quality. The text is then uploaded to Mozilla Common Voice, which provides the infrastructure for recording and validating the same sentences across many voices, while the parallel sentences are stored in spreadsheets mirrored on GitHub and Zenodo. This pipeline converts native-speaker availability and local language knowledge into structured parallel text and speech data without relying on web crawling or existing digital sources.

What would settle it

Open the Zenodo record 13355021 and the Mozilla Common Voice language pages for Kidaw'ida, Kalenjin, and Dholuo, and count the sentence pairs, unique sentences, recording hours, and validated speakers; if any language has substantially fewer than 30,000 sentence pairs or fewer recording hours than the table reports, the paper's central claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that a small team can create substantial parallel text and speech resources for languages that lack them by recruiting native speakers known to the team, paying small stipends, recording conversations, transcribing them, and translating the results into Kiswahili. The reported result is 30,000 Kidaw'ida–Kiswahili, 30,000 Kalenjin–Kiswahili, and 30,000 Dholuo–Kiswahili sentence pairs, together with 56, 92, and 120 hours of speech recordings on Mozilla Common Voice, with speaker counts of 24, 41, and 44 respectively. The paper states that these are, to its knowledge, the first NLP corpora for Kidaw'ida and Kalenjin, and it makes all of the resources freely available under open licenses so that baseline models can be trained and community expansion can continue.

Load-bearing premise

The central claim rests on the accuracy of the reported counts: that the Zenodo record and Mozilla Common Voice pages actually contain the stated 30,000 sentences per language and the stated speech hours, even though the paper shows no sample sentences, file counts, or quality checks.

Editorial extensions

If this is right

  • Baseline machine translation and speech recognition models can now be trained for Kidaw'ida, Kalenjin, and Dholuo using freely downloadable data, giving developers a starting point to improve on.
  • The corpora give Kidaw'ida and Kalenjin a digital presence they previously lacked, making it possible to build NLP tools for health, agriculture, education, and commerce for their speakers.
  • The reported gender balance among contributors and Data Collection Leads makes the speech data more likely to represent female voices, which is often missing in low-resource speech datasets.
  • As the open repositories grow, community members can add more sentences and recordings, which should improve model accuracy over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the reported counts survive independent checks, the paper's main practical lesson is that the bottleneck for low-resource African language NLP is not technical but organizational: trusted bilingual community members, paid modestly, can produce useful corpora in about a year.
  • The decision to strip code-switched words during transcription, while sensible for clean parallel corpora, may make the data less representative of everyday mixed-language speech, so downstream systems may need separate treatment of code-switching.
  • Because the paper provides no sample sentences or quality metrics, an immediate extension would be to publish a held-out test set with human-validated references, allowing future work to measure translation and speech-recognition quality consistently.
  • The selective-crowdsourcing method is unlikely to scale to the millions of sentences required by large language models without layering in automated validation, so the corpora may be more immediately useful for baselines and community tools than for foundation-model pretraining.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a one-year, grant-funded project to build three parallel text corpora (Kidaw'ida–Kiswahili, Kalenjin–Kiswahili, Dholuo–Kiswahili) of 30,000 sentences each, together with speech datasets contributed to Mozilla Common Voice. The authors describe a 'selective crowdsourcing' methodology, a quality-assurance process based on Data Collection Leads, and community-engagement challenges. The principal contribution is the claimed public release of these resources for low-resource African language NLP.

Significance. If the resources exist as described, they constitute the first NLP corpora for Kidaw'ida and Kalenjin and an additional Dholuo–Kiswahili parallel corpus, with a non-trivial speech component. The paper's participatory approach, attention to gender balance, and emphasis on open licensing are strengths, and the work addresses a genuine resource gap. However, the paper does not currently substantiate the central empirical claims: the text corpus sizes and speech statistics are asserted without in-paper evidence, making the significance conditional on external verification.

major comments (3)
  1. [Section 5] The claim 'We collected 30,000 text sentences for each of the three languages' is not supported within the paper. No file inventory, per-language sentence-pair counts, token or character statistics, duplication checks, or sample entries are provided, and the Zenodo URL alone does not allow a reader to verify the count or the contents of the deposit. Please include a dataset summary table with per-language counts, at least one sample aligned sentence pair per language, and confirmation of the license and file format of the deposited corpus.
  2. [Table 1] The speech data table reports hours and speaker counts for each language on Mozilla Common Voice, but it does not state the dataset version, the language page URLs, or whether the hours are total recorded hours or validated hours. Common Voice distinguishes these two metrics, and without specifying which is reported the numbers cannot be audited. Please state the dataset snapshot and report both total and validated hours and speaker counts.
  3. [Section 4.1] The quality-assurance process is described only qualitatively: Data Collection Leads 'checked the contributors' data for correct spelling, grammar, fluency, and proper translation.' No numbers are given, such as the number of DCLs per language, the fraction of sentences reviewed, the correction rate, or how disagreements were resolved. Because the paper presents 'selective crowdsourcing' as a methodological contribution, these quantitative details are needed to support the quality claims.
minor comments (6)
  1. [Section 3] The sentence describing Nakatumba-Nabende et al. reads 'A total of sentences of monolingual data for five languages was collected.' The number is missing; please complete the sentence.
  2. [Section 3] In the description of Ogayo et al., 'Dhouo' appears to be a typo for 'Dholuo'; please correct it.
  3. [Figure 1] The methodology behind the 'Distribution of NLP activity in Africa' figure is not described. Please specify the data source, the search terms used, and the time period covered.
  4. [Section 5] The text states 'The same repository is on Github' but no GitHub URL is provided; please add it so readers can access the version-controlled source.
  5. [Table 1 caption] The caption contains a stray space: 'T able 1' should be 'Table 1'.
  6. [Section 4.1] The sentence 'The issue of cultural appropriateness of data is often cited as militating against using texts of foreign origin, but we propose that such text can catalyse ideas' is presented without a supporting reference; please either cite relevant literature or reframe it as an observation from the project.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports empirical corpus construction and contains no derivation, fitted parameter, or prediction that reduces to its own inputs.

full rationale

The paper is a data-collection report rather than a derivational or predictive study: it contains no equations, no fitted parameters, no model, and no quantity that is predicted from another quantity. The central claim in Section 5, 'We collected 30,000 text sentences for each of the three languages and had contributors proficient in both their mother tongue and Kiswahili provide translations,' is an empirical self-report whose truth depends on the external Zenodo and Mozilla Common Voice repositories, not on a derivation chain internal to the manuscript. That evidential gap is a verifiability concern, not circularity. The only apparent self-citation is reference [13] (Kencorpus), which lists author Lilian Wanzare as a co-author; that citation appears in Related Work as an example of prior Kenyan corpus building and is not used to justify the present paper's central claim, its methodology, or any uniqueness assertion. No ansatz is smuggled in via citation, no result is renamed, and no fitted input is later called a prediction. Accordingly, no circular step can be exhibited, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No free parameters were fitted and no explicit mathematical axioms were introduced. The project relies on standard assumptions about language proficiency, consent, and data quality, which are described qualitatively but not formalized. The paper introduces no new theoretical entities. The contributions are empirical artifacts (corpora) whose existence must be verified externally.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo." pith.science (2026). https://pith.science/paper/AJANWMQG

@misc{pith2026250111003,
  author       = {Pith},
  title        = {Pith review of: Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AJANWMQG}},
  note         = {Machine review of arXiv:2501.11003}
}
read the original abstract

Natural Language Processing is a crucial frontier in artificial intelligence, with broad applications in many areas, including public health, agriculture, education, and commerce. However, due to the lack of substantial linguistic resources, many African languages remain underrepresented in this digital transformation. This paper presents a case study on the development of linguistic corpora for three under-resourced Kenyan languages, Kidaw'ida, Kalenjin, and Dholuo, with the aim of advancing natural language processing and linguistic research in African communities. Our project, which lasted one year, employed a selective crowd-sourcing methodology to collect text and speech data from native speakers of these languages. Data collection involved (1) recording conversations and translation of the resulting text into Kiswahili, thereby creating parallel corpora, and (2) reading and recording written texts to generate speech corpora. We made these resources freely accessible via open-research platforms, namely Zenodo for the parallel text corpora and Mozilla Common Voice for the speech datasets, thus facilitating ongoing contributions and access for developers to train models and develop Natural Language Processing applications. The project demonstrates how grassroots efforts in corpus building can support the inclusion of African languages in artificial intelligence innovations. In addition to filling resource gaps, these corpora are vital in promoting linguistic diversity and empowering local communities by enabling Natural Language Processing applications tailored to their needs. As African countries like Kenya increasingly embrace digital transformation, developing indigenous language resources becomes essential for inclusive growth. We encourage continued collaboration from native speakers and developers to expand and utilize these corpora.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

17 extracted references · 15 canonical work pages

  1. [1]

    610–623 (2021) 11

    Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S.: On the dangers of stochastic parrots: Can language models be too big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 610–623 (2021) 11

  2. [2]

    arXiv preprint arXiv:2010.02353 (2020)

    Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Kolawole, T., Fagbohungbe, T., Akinola, S.O., Muhammad, S.H., Kabongo, S., Osei, S., et al.: Participatory research for low-resourced machine translation: A case study in african languages. arXiv preprint arXiv:2010.02353 (2020)

  3. [3]

    In: Selected Proceedings of the 40th Annual Conference on African Linguistics, pp

    Bamgbose, A.: African languages today: The challenge of and prospects for empowerment under globalization. In: Selected Proceedings of the 40th Annual Conference on African Linguistics, pp. 1–14 (2011). Cascadilla Proceedings Project Somerville

  4. [4]

    PhD thesis, Syracuse University (2016)

    Nduati, R.N.: The post-colonial language and identity experiences of transna- tional kenyan teachers in us universities. PhD thesis, Syracuse University (2016)

  5. [5]

    Peter Bostock, Kenya (1995)

    Mcharo, A.F.: Mizango na Maza Ra Kidawida Ra Kufuma Kokala. Peter Bostock, Kenya (1995). https://books.google.co.ke/books?id=0zOBHAAACAAJ

  6. [6]

    Petit trait´ e de glottophagie, Paris, Payot (1974)

    Calvet, L.-J.: Linguistique et colonialisme. Petit trait´ e de glottophagie, Paris, Payot (1974)

  7. [7]

    Journal of African Studies 1986(28), 27–47 (1986)

    Sakamoto, K.: Social organization and ritual among the taita of kenya the process of the social change from kichuku type to muzi type. Journal of African Studies 1986(28), 27–47 (1986)

  8. [8]

    International Journal of Academic Research in Business and Social Sciences 8(8), 476–503 (2018)

    Naibei Faith, K., Lwangale, D.: A comparative study of the kalenjin dialects. International Journal of Academic Research in Business and Social Sciences 8(8), 476–503 (2018)

Show all 17 references
  1. [9]

    International journal of innovative research and development 5 (2016)

    Chelimo, F.J., Chelelgo, K.: Pre-colonial political organization of the kalenjin of kenya: An overview. International journal of innovative research and development 5 (2016)

  2. [10]

    Ogot, B.A.: Historical portrait of western kenya up to 1895 bethwell a. ogot. Historical Studies and Social Change in Western Kenya: Essays in Memory of Professor Gideon S. Were 13(4), 99–120 (2002)

  3. [11]

    African Identities 16(1), 87–102 (2018)

    Omulo, A.G., Williams, J.J.: A survey of the influence of ‘ethnicity’, in african governance, with special reference to its impact in kenya vis-` a-vis its luo community. African Identities 16(1), 87–102 (2018)

  4. [12]

    Cambridge University Press, United Kingdom (2000)

    Heine, B., Nurse, D.: African Languages: An Introduction. Cambridge University Press, United Kingdom (2000)

  5. [13]

    In: Wartena, C

    Wanjawa, B., Wanzare, L., Indede, F., McOnyango, O., Ombui, E., Muchemi, L.: Kencorpus: A kenyan language corpus of Swahili, dholuo and luhya for natural language processing tasks. In: Wartena, C. (ed.) Journal for Language Technology and Computational Linguistics, Vol. 36 No....

  6. [14]

    Transactions of the Association for Computational Linguistics 9, 1116–1131 (2021)

    Adelani, D.I., Abbott, J., Neubig, G., D’souza, D., Kreutzer, J., Lignos, C., Palen-Michel, C., Buzaaba, H., Rijhwani, S., Ruder, S., et al.: Masakhaner: Named entity recognition for african languages. Transactions of the Association for Computational Linguistics 9, 1116–1131 (2021)

  7. [15]

    Applied AI Letters 5(2), 92 (2024)

    Nakatumba-Nabende, J., Babirye, C., Nabende, P., Tusubira, J.F., Mukiibi, J., Wairagala, E.P., Mutebi, C., Bateesa, T.S., Nahabwe, A., Tusiime, H., et al.: Building text and speech benchmark datasets and models for low-resourced east african languages: Experiences and lessons....

  8. [16]

    In: Proc

    Ogayo, P., Neubig, G., Black, A.W.: Building tts systems for low resource languages under resource constraints. In: Proc. S4SG 2022 (2022)

  9. [17]

    ArXiv abs/2307.09288 (2023) 13

    Touvron, H., Martin, L., Stone, K.R., Albert, P., Almahairi, A., Babaei, Y., Bash- lykov, N., Batra, S., Bhargava, P., et al.: Llama 2: Open foundation and fine-tuned chat models. ArXiv abs/2307.09288 (2023) 13

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.