Pith. sign in

REVIEW 4 major objections 5 minor 85 references

HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The HausaNLP Catalogue gathers Hausa datasets, tools, and papers into one open repository that the paper argues will accelerate Hausa NLP.

desk verdict Useful community survey with a genuinely new catalog artifact, but the catalog's curation and completeness are not yet documented; fixable, not fatal. read the letter →

arxiv 2505.14311 v3 pith:Z65XP4U6 submitted 2025-05-20 cs.CL

classification cs.CL
keywords HausaNLPlow-resourcelanguageslanguageresourcesdatasetcatalognaturalprocessingAfricanlargemodelsmachinetranslation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Hausa is spoken by over 120 million first-language and 80 million second-language speakers, yet it remains a low-resource language in NLP because its datasets and tools are scarce and dispersed. This paper argues that a single open, community-driven catalogue it introduces, the HausaNLP Catalogue, can serve as the foundation for accelerating Hausa NLP by aggregating datasets, tools, and research works in one discoverable place. It supports this claim with a review of progress and gaps across text classification, machine translation, named entity recognition, speech recognition, and question answering, and with an analysis of why Hausa is hard to fit into large language models. If the catalogue works as advertised, researchers and practitioners get a central access point that lowers the cost of finding and reusing Hausa resources.

What carries the argument

The central object is the HausaNLP Catalogue itself, a curated online repository of Hausa NLP datasets, tools, and research papers. It functions as the mechanism that converts a scattered set of resources into a single discoverable access point; the paper argues that this aggregation is what lowers barriers for new researchers and accelerates progress. The accompanying task taxonomy and the discussion of tokenization and dialectal variation provide the organizational scheme for what the catalogue lists and what the field should build next.

What would settle it

Run an independent audit: take the datasets named in the paper's own review plus a sample of Hausa NLP datasets released in recent major venues, and check whether each is listed in the HausaNLP Catalogue with a working link and correct metadata. If a substantial share is missing or broken, the claim that the catalogue provides a foundation for Hausa NLP progress is not supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that the HausaNLP Catalogue provides both a foundation for accelerating Hausa NLP progress and a template for multilingual NLP research in other low-resource languages. The catalogue is a curated aggregation of the field's existing datasets, tools, and publications, maintained as a living, community-driven resource. In the same work, the authors survey the current state of Hausa NLP across five core tasks, document recurring obstacles such as limited open corpora, suboptimal tokenization, and dialectal variation, and propose strategic directions including dataset expansion, improved language modeling, and stronger community collaboration.

Load-bearing premise

The catalogue is sufficiently complete, accurate, and current to serve as the field's foundation; the paper gives no inclusion criteria, search protocol, update schedule, or completeness audit to back that up.

Editorial extensions

If this is right

  • Newcomers to Hausa NLP can use the catalogue as a single entry point to locate public datasets, reducing duplicated data-collection effort.
  • The task-by-task review makes visible where Hausa resources are thinnest, such as toxicity detection and question answering, guiding where annotation efforts should go.
  • The identified challenges of suboptimal tokenization and dialectal variation give model builders a concrete agenda for improving Hausa representation in large language models.
  • If the proposed community-driven model is adopted, resource updates can flow continuously from researchers rather than waiting for a single central team.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the catalogue's long-run value depends on documented curation practices; without stated inclusion criteria, an update schedule, and a completeness audit, it risks becoming stale and cannot guarantee the 'foundation' role claimed.
  • My inference: the same aggregation model could be applied to other low-resource languages, turning scattered datasets into comparable, cross-lingual resource hubs.
  • My inference: a practical near-term test of the catalogue is whether a researcher who knows only the paper can assemble a working Hausa sentiment or NER pipeline from the linked resources alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper surveys the current state of Hausa natural language processing, covering text classification, machine translation, part-of-speech tagging, summarization, question answering, named entity recognition, and speech recognition. It introduces the HausaNLP Catalogue (catalog.hausanlp.org), described as a curated, community-driven repository of datasets, tools, and research works. The paper also discusses challenges for Hausa in large language models, including tokenization and dialectal variation, and proposes future research directions. The central claim, stated in the Abstract and Section 1, is that the catalogue provides a foundation for accelerating Hausa NLP progress and that the paper constitutes a comprehensive review.

Significance. If the HausaNLP Catalogue is complete and trustworthy, it would be a valuable centralized access point for a low-resource language whose resources are currently dispersed. The survey itself is useful in assembling the literature across several tasks and in identifying gaps, and the authors are transparent about weaknesses in individual datasets. The community-oriented framing and the explicit connection to broader multilingual NLP questions are strengths. However, the foundation claim is not yet substantiated: the catalogue methodology is undocumented, no completeness or accuracy audit is provided, and the in-paper table of datasets contains demonstrable curation errors. These issues bear directly on the paper's main contribution rather than on its peripheral discussion.

major comments (4)
  1. [Section 1 and Abstract] The paper's central claim is that the HausaNLP Catalogue is a curated catalog that 'provides a foundation for accelerating Hausa NLP progress.' Nowhere does the paper describe how catalog entries were selected, what inclusion/exclusion criteria were used, how errors are reported or corrected, or how completeness is maintained. Without this methodology, the catalog's coverage cannot be verified, and the 'comprehensive review' claim becomes unfalsifiable. Please add a curation protocol, including search sources, inclusion criteria, and an audit procedure, or revise the claim to describe the catalog as an ongoing community effort rather than an established foundation.
  2. [Appendix Table 1] Table 1 contains several concrete curation errors that undermine the catalog's reliability as presented: two rows are numbered '3' (Aliyu et al., 2022 and Adelani et al., 2023); row 4 (Inuwa-Dutse, 2023) lacks a Size field; and the table omits datasets that the review itself discusses as major contributions, including HaVG (Section 3.2.2), HaVQA (Section 3.5), and AfriSenti (Section 1 cites Muhammad et al., 2023). These are not cosmetic issues; they show that the aggregation effort is not yet accurate enough to support the paper's foundational claim.
  3. [Section 1, 'living resource'] The catalog is described as a 'living resource,' but the paper provides no version identifier, snapshot date, or archival link. Because the paper invites readers to build on the catalog as a foundation, the published record must be reproducible. Please include a persistent snapshot or version number for catalog.hausanlp.org as of the paper's submission, or state how readers can obtain the exact version used in the paper.
  4. [Sections 3.1-3.7] The survey sections do not systematically link the datasets and tools they discuss to their catalog entries. For a resource intended as a centralized repository, each resource mentioned in the review should be traceable to a catalog identifier or URL. Without such cross-linking, the reader cannot verify that the catalog actually contains the resources claimed in the survey, and the 'comprehensive review' claim cannot be checked against the catalog.
minor comments (5)
  1. [Section 7] The section heading 'Appendex' is misspelled; it should be 'Appendix.'
  2. [Section 3.3] The phrase 'sourced from from radio stations' contains a duplicated 'from'; please correct the typo.
  3. [Section 3.4] The sentence 'perhaps conducted one the the earliest works' contains a duplicated 'the' and a small grammatical error; please revise.
  4. [Abstract] The abstract uses 'HAUSA NLP' in one place and 'HausaNLP' elsewhere; please standardize the naming.
  5. [Section 3.2.1] The reference to 'The Hausa Khamenei 2 corpus' is unclear and likely a typo; please clarify which corpus is meant and cite the source properly.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a survey and resource announcement whose claims are externally checkable, not derived from their own inputs.

full rationale

This paper is a survey plus resource announcement, not a derivation chain. Its central claim—that the HausaNLP Catalogue aggregates datasets, tools, and papers and thereby provides a foundation for accelerating Hausa NLP—is an externally checkable statement about a web resource, not a conclusion derived from fitted parameters or from a self-citation. There are no equations whose left-hand side is defined by its right-hand side, no parameter fitted to a subset and then reported as a prediction, and no uniqueness theorem imported from the authors' prior work to force a modeling choice. Many cited datasets and tools originate from the HausaNLP community, including the authors themselves, but in a field survey this is expected, and none of these citations carries a load-bearing argument that reduces to itself; the cited datasets (NaijaSenti, MasakhaNER, AfriSenti, and others) are independently published and inspectable. The Table 1 curation errors and missing inclusion criteria noted in the skeptical read concern completeness and auditability of the catalog, which is a correctness or evidence-quality issue, not circularity. No circular step meets the standard of exhibiting Eq. X = Eq. Y or a fitted input renamed as a prediction, so the score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented scientific entities. Its claims rest on background demographic facts, on the unstated assumption that the catalog is representative and current, and on the field-level belief that open corpora accelerate NLP.

assumptions (3)
  • domain assumption Hausa has roughly 120 million L1 and 80 million L2 speakers.
    Section 2 states these counts as background for the paper's significance argument, citing 'Hegazy et al.' without a usable year.
  • ad hoc to paper The catalogue is a complete and reliable index of Hausa NLP resources.
    The paper's central contribution depends on this premise, but no curation methodology, versioning, or completeness metric is provided.
  • domain assumption Open-source corpus availability is a primary driver of NLP progress.
    Section 1 motivates the catalog with this claim; it is a reasonable domain assumption but is not tested in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing." pith.science (2026). https://pith.science/paper/Z65XP4U6

@misc{pith2026250514311,
  author       = {Pith},
  title        = {Pith review of: HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z65XP4U6}},
  note         = {Machine review of arXiv:2505.14311}
}
read the original abstract

Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-language (L1) and 80 million second-language (L2) speakers worldwide. While significant advances have been made in high-resource languages, Hausa NLP faces persistent challenges, including limited open-source datasets and inadequate model representation. This paper presents an overview of the current state of Hausa NLP, systematically examining existing resources, research contributions, and gaps across fundamental NLP tasks: text classification, machine translation, named entity recognition, speech recognition, and question answering. We introduce HausaNLP (https://catalog.hausanlp.org), a curated catalog that aggregates datasets, tools, and research works to enhance accessibility and drive further development. Furthermore, we discuss challenges in integrating Hausa into large language models (LLMs), addressing issues of suboptimal tokenization and dialectal variation. Finally, we propose strategic research directions emphasizing dataset expansion, improved language modeling approaches, and strengthened community collaboration to advance Hausa NLP. Our work provides both a foundation for accelerating Hausa NLP progress and valuable insights for broader multilingual NLP research.

Figures

Figures reproduced from arXiv: 2505.14311 by the authors.

Figure 1
Figure 1. HausaNLP Catalogue: A repository of datasets, tools, and research papers on Hausa NLP, de￾veloped to improve access to and discovery of Hausa language resources and White, 2014). One of the recent advances in NLP is emergence of large language models (LLMs) such as ChatGPT, which demonstrated im￾pressive performance in various NLP tasks, such as dialogue generation and arithmetic reasoning (Qin et al., 2023). Howeve… view at source ↗
Figure 2
Figure 2. Taxonomy of Hausa NLP Research Progress: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 59 canonical work pages

  1. [1]

    Habeeba Ibraheem Abdullahi, Muhammad Aminu Ahmad, and Khalid Haruna. 2024. Twitter sentiment analysis for hausa abbreviations and acronyms. Science World Journal, 19(1):101--104

  2. [2]

    Idris Abdulmumin, Michael Beukman, Jesujoba Alabi, Chris Chinenye Emezue, Everlyn Chimoto, Tosin Adewumi, Shamsuddeen Muhammad, Mofetoluwa Adeyemi, Oreen Yousuf, Sahib Singh, and Tajuddeen Gwadabe. 2022 a . https://aclanthology.org/2022.wmt-1.98 Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced A ...

  3. [3]

    Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Muhammad, Ibrahim Sa ' id Ahmad, Subhadarshi Panda, Ond r ej Bojar, Bashir Shehu Galadanci, and Bello Shehu Bello. 2022 b . https://aclanthology.org/2022.lrec-1.694 H ausa visual genome: A dataset for multi-modal E nglish to H ausa machine translation . In Proceedin...

  4. [4]

    Idris Abdulmumin, Auwal Abubakar Khalid, Shamsuddeen Hassan Muhammad, Ibrahim Said Ahmad, Lukman Jibril Aliyu, Babangida Sani, Bala Mairiga Abduljalil, and Sani Ahmad Hassan. 2023. http://arxiv.org/abs/2311.12179 Leveraging closed-access multilingual embedding for automatic sentence alignment in low resource languages

  5. [5]

    Abdulqahar Mukhtar Abubakar, Deepa Gupta, and Susmitha Vekkot. 2024. Development of a diacritic-aware large vocabulary automatic speech recognition for hausa language. International Journal of Speech Technology, 27(3):687--700

  6. [6]

    Amina Imam Abubakar, Abubakar Roko, Aminu Muhammad Bui, and Ibrahim Saidu. 2021. https://doi.org/10.14569/IJACSA.2021.0120913 An enhanced feature acquisition for sentiment analysis of english and hausa tweets . International Journal of Advanced Computer Science and Applications, 12(9)

  7. [7]

    Emre Can Acikgoz, Mete Erdogan, and Deniz Yuret. 2024. https://www.scopus.com/inward/record.uri?eid=2-s2.0-85216584206&partnerID=40&md5=90ee9073dd5b9fcf7bdc8cfc5440bed3 Bridging the bosphorus: Advancing turkish large language models through strategies for low-resource language adaptation and benchmarking . page 242 – 268

  8. [8]

    David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala...

Show all 85 references
  1. [9]

    David Adelani, Md Mahfuz Ibn Alam, Antonios Anastasopoulos, Akshita Bhagia, Marta R. Costa-juss \`a , Jesse Dodge, Fahim Faisal, Christian Federmann, Natalia Fedorova, Francisco Guzm \'a n, Sergey Koshelev, Jean Maillard, Vukosi Marivate, Jonathan Mbuya, Alexandre Mourachko, S...

  2. [10]

    Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F

    David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda,...

  3. [11]

    David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D ' souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-...

  4. [12]

    David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime, Jesujoba Alabi, Atnafu Lambebo Tonja, Christine Mwase, Odunayo Ogundepo, Bonaventure F. P. Dossou, Akintunde Oladipo, Doreen Nixdorf, Chris Chinenye Emezue, sana al azzawi, Blessing Sibanda, Davis David, Lolwethu Ndolel...

  5. [13]

    Z eljko Agi \'c and Ivan Vuli \'c . 2019. https://doi.org/10.18653/v1/P19-1310 JW 300: A wide-coverage parallel corpus for low-resource languages . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3204--3210, Florence, Italy. As...

  6. [14]

    Ibrahim Said Ahmad, Shiran Dudy, Resmi Ramachandranpillai, and Kenneth Church. 2024. https://www.scopus.com/inward/record.uri?eid=2-s2.0-85204391034&partnerID=40&md5=0f6f3a02b415783b5f532f19c14af969 Are generative language models multicultural? a study on hausa culture and emo...

  7. [15]

    Ahmed and Dauda B

    U. Ahmed and Dauda B. 1970. An introduction to classical hausa and major dialects. Norther Nigeria Publishing Company

  8. [16]

    Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ond r ej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina Espa \ n a-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter,...

  9. [17]

    Adewale Akinfaderin. 2020. Hausamt v1. 0: Towards english--hausa neural machine translation. In Proceedings of the The Fourth Widening Natural Language Processing Workshop, pages 144--147

  10. [18]

    Saminu Mohammad Aliyu, Gregory Maksha Wajiga, Muhammad Murtala, Shamsuddeen Hassan Muhammad, Idris Abdulmumin, and Ibrahim Said Ahmad. 2022. Herdphobia: A dataset for hate speech against fulani in nigeria. In Seventh Widening Natural Language Processing Workshop (WiNLP)

  11. [19]

    Jamilu Awwalu, Saleh Elyakub Abdullahi, and Abraham Eseoghene Evwiekpaefe. 2021. A corpus based transformation-based learning for hausa text parts of speech tagging. International Journal of Computing and Digital Systems, 10:473--490

  12. [20]

    Muazzam Bashir, Azilawati Rozaimee, and Wan Malini Wan Isa. 2017. Automatic hausa languagetext summarization based on feature extraction using na \" ve bayes model. World Applied Science Journal, 35(9):2074--2080

  13. [21]

    A. Bello. 2015. The dialects of hausa. Ahmadu Bello University Press

  14. [22]

    Abdulkadir Abubakar Bichi, Ruhaidah Samsudin, Rohayanti Hassan, Layla Rasheed Abdallah Hasan, and Abubakar Ado Rogo. 2023. Graph-based extractive text summarization method for hausa text. Plos one, 18(5):e0285376

  15. [23]

    Erik Cambria and Bebo White. 2014. Jumping nlp curves: A review of natural language processing research. IEEE Computational intelligence magazine, 9(2):48--57

  16. [24]

    Bernard Caron. 2012. Hausa grammatical sketch

  17. [25]

    Jireh Yi-Le Chan, Khean Thye Bea, Steven Mun Hong Leow, Seuk Wai Phoong, and Wai Khuen Cheng. 2023. State of the art: a review of sentiment analysis based on sequential transfer learning. Artificial Intelligence Review, 56(1):749--780

  18. [26]

    Pinzhen Chen, Jind r ich Helcl, Ulrich Germann, Laurie Burchell, Nikolay Bogoychev, Antonio Valerio Miceli Barone, Jonas Waldendorf, Alexandra Birch, and Kenneth Heafield. 2021. https://aclanthology.org/2021.wmt-1.4 The U niversity of E dinburgh ' s E nglish- G erman and E ngl...

  19. [27]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...

  20. [28]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. http://arxiv.org/abs/1810.04805 Bert: Pre-training of deep bidirectional transformers for language understanding

  21. [29]

    Cheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing Sibanda, Bonaventure F. ...

  22. [30]

    Arwa Diwali, Kawther Saeedi, Kia Dashtipour, Mandar Gogate, Erik Cambria, and Amir Hussain. 2023. Sentiment analysis meets explainable artificial intelligence: A survey on explainable sentiment analysis. IEEE Transactions on Affective Computing

  23. [31]

    El-Kassas, Cherif R

    Wafaa S. El-Kassas, Cherif R. Salama, Ahmed A. Rafea, and Hoda K. Mohamed. 2021. https://doi.org/https://doi.org/10.1016/j.eswa.2020.113679 Automatic text summarization: A comprehensive survey . Expert Systems with Applications, 165:113679

  24. [32]

    Mohamed Helal Ahmed Sheref El-Shazly. 1987. The provenance of Arabic loan-words in Hausa: a phonological and semantic study. University of London, School of Oriental and African Studies (United Kingdom)

  25. [33]

    Goodwill Erasmo Ndomba, Medard Edmund Mswahili, and Young-Seob Jeong. 2025. https://doi.org/10.1109/ACCESS.2024.3522285 Tokenizers for african languages . IEEE Access, 13:1046--1054

  26. [34]

    Ankita Gandhi, Kinjal Adhvaryu, Soujanya Poria, Erik Cambria, and Amir Hussain. 2023. Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Information Fusion, 91:424--444

  27. [35]

    Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingua...

  28. [36]

    Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, and Philipp Zimmer. 2023. Simplistic collection and labeling practices limit the utility of benchmark datasets for twitter bot detection. In Proceedings of the ACM web conference 2023, pages 3660--3669

  29. [37]

    Hedderich, David Adelani, Dawei Zhu, Jesujoba Alabi, Udia Markus, and Dietrich Klakow

    Michael A. Hedderich, David Adelani, Dawei Zhu, Jesujoba Alabi, Udia Markus, and Dietrich Klakow. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.204 Transfer learning and distant supervision for multilingual transformer models: A study on A frican languages . In Proceedings...

  30. [38]

    A lexical semantic error analysis of arabic-speaking hausa language learners

    Mahmoud Fahmi Hegazy, Mohammad Ali Nofal, and MA Mahmoud Sayed. A lexical semantic error analysis of arabic-speaking hausa language learners

  31. [39]

    Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox

    Michael Y. Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox. 2024. https://aclanthology.org/2024.conll-babylm.1/ Findings of the second B aby LM challenge: Sample-efficient pret...

  32. [40]

    Umar Ibrahim, Abubakar Yakubu Zandam, Fatima Muhammad Adam, and Aminu Musa. 2024. A deep convolutional neural network-based model for aspect and polarity classification in hausa movie reviews. arXiv preprint arXiv:2405.19575

  33. [41]

    Umar Adam Ibrahim, Moussa Boukar Mahatma, and Muhammed Aliyu Suleiman. 2022. https://doi.org/10.1109/ITED56637.2022.10051610 Framework for hausa speech recognition

  34. [42]

    Sukairaj Hafiz Imam, Abubakar Ahmad Musa, and Ankur Choudhary. 2022. The first corpus for detecting fake news in hausa language. In Emerging Technologies for Computing, Communication and Smart Cities, pages 563--576, Singapore. Springer Nature Singapore

  35. [43]

    Isa Inuwa-Dutse. 2023. https://openreview.net/forum?id=K6FUc5TE-nY The first large scale collection of diverse hausa language datasets . In 4th Workshop on African Natural Language Processing

  36. [44]

    P.J. Jaggar. 2006. https://doi.org/10.1016/B0-08-044854-2/02071-X Hausa

  37. [45]

    Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2022. https://doi.org/10.1109/TKDE.2020.2981314 A survey on deep learning for named entity recognition . IEEE Transactions on Knowledge and Data Engineering, 34(1):50--70

  38. [46]

    Jiabei Liu, Keqin Li, Armando Zhu, Bo Hong, Peng Zhao, Shuying Dai, Changsong Wei, Wenqian Huang, and Honghua Su. 2024. Application of deep learning-based natural language processing in multilingual sentiment analysis. Mediterranean Journal of Basic and Applied Sciences (MJBAS...

  39. [47]

    Abeer Mahgoub, Ghada Khoriba, and Elhassan Anas Elsabry. 2024. https://doi.org/10.1016/j.procs.2024.10.181 Mathematical problem solving in arabic: Assessing large language models . volume 244, page 86 – 95

  40. [48]

    Martinez

    Angel R. Martinez. 2012. https://doi.org/https://doi.org/10.1002/wics.195 Part-of-speech tagging . WIREs Computational Statistics, 4(1):107--113

  41. [49]

    Idi Mohammed and Rajesh Prasad. 2024. Lexicon dataset for the hausa language. Data in Brief, 53:110124

  42. [50]

    Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu...

  43. [51]

    Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Nelson Odhiambo Onyango, Lilian D. A. Wanzare, Samuel Rutunda, Lukman Jibril Aliyu, Esubalew Alemneh, Oumaima Hourrane, Hagos Tesfahun Gebremicha...

  44. [52]

    Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Sa ' id Ahmad, Idris Abdulmumin, Bello Shehu Bello, Monojit Choudhury, Chris Chinenye Emezue, Saheed Salahudeen Abdullahi, Anuoluwapo Aremu, Al \' pio Jorge, and Pavel Brazdil. 2022. https://aclantho...

  45. [54]

    Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, et al. 2025 c . Brighter: Bridging the gap in human-annotated textual emotion recognition ...

  46. [55]

    Paul Newman. 2022. https://doi.org/10.1017/9781009128070.006 Loanwords , page 205–211. Cambridge University Press

  47. [56]

    Artur Nowakowski and Tomasz Dwojak. 2021. https://aclanthology.org/2021.wmt-1.14 A dam M ickiewicz U niversity ' s E nglish- H ausa submissions to the WMT 2021 news translation task . In Proceedings of the Sixth Conference on Machine Translation, pages 167--171, Online. Associ...

  48. [57]

    Ruba Obiedat, Duha Al-Darras, Esra Alzaghoul, and Osama Harfoushi. 2021. Arabic aspect-based sentiment analysis: A systematic literature review. IEEE Access, 9:152628--152645

  49. [58]

    Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.mrl-1.11 Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages . In Proceedings of the 1st Workshop on Multilingual Representation ...

  50. [59]

    Odunayo Ogundepo, Tajuddeen Gwadabe, Clara Rivera, Jonathan Clark, Sebastian Ruder, David Adelani, Bonaventure Dossou, Abdou Diop, Claytone Sikasote, Gilles Hacheme, Happy Buzaaba, Ignatius Ezeani, Rooweither Mabuya, Salomey Osei, Chris Emezue, Albert Kahira, Shamsuddeen Muham...

  51. [60]

    Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad, Aneesh Bose, Guneet Singh Kohli, Ibrahim Said Ahmad, Ketan Kotwal, Sayan Deb Sarkar, Ond r ej Bojar, and Habeebah Kakudi. 2023. https://doi.org/10.18653/v1/2023.findings-acl.646 H a VQA : A dataset for visual q...

  52. [61]

    Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is chatgpt a general-purpose natural language processing task solver? arXiv preprint arXiv:2302.06476

  53. [62]

    Ochilbek Rakhmanov and Tim Schlippe. 2022 a . https://aclanthology.org/2022.sigul-1.13 Sentiment analysis for H ausa: Classifying students ' comments . In Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages, pages 98--105,...

  54. [63]

    Ochilbek Rakhmanov and Tim Schlippe. 2022 b . Sentiment analysis for hausa: Classifying students’ comments. In Proceedings of the 1st Annual Meeting of the ELRA/ISCA Special Interest Group on Under-Resourced Languages, pages 98--105

  55. [64]

    Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2023. https://doi.org/10.1145/3560260 Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension . ACM Comput. Surv., 55(10)

  56. [65]

    Babangida Sani, Aakansha Soy, Sukairaj Hafiz Imam, Ahmad Mustapha, Lukman Jibril Aliyu, Idris Abdulmumin, Ibrahim Said Ahmad, and Shamsuddeen Hassan Muhammad. 2025 a . Who wrote this? identifying machine vs human-generated text in hausa. arXiv preprint arXiv:2503.13101

  57. [66]

    Muhammad Sani, Abubakar Ahmad, and Hadiza S Abdulazeez. 2022. Sentiment analysis of hausa language tweet using machine learning approach. Journal of Research in Applied Mathematics, 8(9):07--16

  58. [67]

    Sani Abdullahi Sani, Shamsuddeen Hassan Muhammad, and Devon Jarvis. 2025 b . https://aclanthology.org/2025.loreslm-1.7/ Investigating the impact of language-adaptive fine-tuning on sentiment analysis in H ausa language using A fri BERT a . In Proceedings of the First Workshop ...

  59. [68]

    Tim Schlippe, Edy Guevara Komgang Djomgang, Ngoc Thang Vu, Sebastian Ochs, and Tanja Schultz. 2012. Hausa large vocabulary continuous speech recognition. In Spoken Language Technologies for Under-Resourced Languages

  60. [69]

    Ayesha Shakith and L Arockiam. 2024. Enhancing classification accuracy on code-mixed and imbalanced data using an adaptive deep autoencoder and xgboost. The Scientific Temper, 15(03):2598--2608

  61. [70]

    Harisu Abdullahi Shehu, Kaloma Usman Majikumna, Aminu Bashir Suleiman, Stephen Luka, Md Haidar Sharif, Rabie A Ramadan, and Huseyin Kusetogullari. 2024. Unveiling sentiments: A deep dive into sentiment analysis for low-resource languages--a case study on hausa texts. IEEE Access

  62. [71]

    Peeyush Singhal, Rahee Walambe, Sheela Ramanna, and Ketan Kotecha. 2023. Domain adaptation: challenges, methods, datasets, and applications. IEEE access, 11:6973--7020

  63. [72]

    Atnafu Lambebo Tonja, Bonaventure F. P. Dossou, Jessica Ojo, Jenalea Rajab, Fadel Thior, Eric Peter Wairagala, Anuoluwapo Aremu, Pelonomi Moiloa, Jade Abbott, Vukosi Marivate, and Benjamin Rosman. 2024. http://arxiv.org/abs/2408.17024 Inkubalm: A small language model for low-r...

  64. [73]

    Aminu Tukur, Kabir Umar, and Anas Sa’idu Muhammad. 2020. Parts-of-speech tagging of hausa-based texts using hidden markov model. Dutse Journal of Pure and Applied Sciences (DUJOPAS), 6:303--313

  65. [74]

    Pavanpankaj Vegi, Sivabhavani J, Biswajit Paul, Abhinav Mishra, Prashant Banjare, Prasanna K R, and Chitra Viswanathan. 2022. https://aclanthology.org/2022.wmt-1.105 W eb C rawl A frican : A multilingual parallel corpora for A frican languages . In Proceedings of the Seventh C...

  66. [75]

    Ludwig Wittgenstein. 1994. Tractatus logico-philosophicus. Edusp

  67. [76]

    Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, and Tamar Regev. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.606 Quantifying the redundancy between prosody and text . In Proceedings of the 2023 Conference on Empirical Methods i...

  68. [77]

    BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanch...

  69. [78]

    S.A. Yakasai. 2025. Tauraruwa harshen hausa jiya da yau: Kalubale da madosa. Tauraruwa Journal of Hausa Studies, 1(1):1--9

  70. [79]

    Yizhe Yang, Huashan Sun, Jiawei Li, Runheng Liu, Yinghao Li, Yuhang Liu, Yang Gao, and Heyan Huang. 2024. https://doi.org/10.1016/j.aiopen.2024.08.001 Mindllm: Lightweight large language model pre-training, evaluation and domain application . AI Open, 5:1 – 26

  71. [80]

    Dong Yu and Lin Deng. 2016. Automatic speech recognition, volume 1. Springer

  72. [81]

    Aliyu Yusuf, Aliza Sarlan, Kamaluddeen Usman Danyaro, and Abdullahi Sani BA Rahman. 2023. Fine-tuning multilingual transformers for hausa-english sentiment analysis. In 2023 13th International Conference on Information Technology in Asia (CITA), pages 13--18. IEEE

  73. [82]

    Aliyu Yusuf, Aliza Sarlan, Kamaluddeen Usman Danyaro, Abdullahi Sani BA Rahman, and Mujaheed Abdullahi. 2024. Sentiment analysis in low-resource settings: A comprehensive review of approaches, languages, and data sources. IEEE Access

  74. [83]

    Rufai Yusuf Zakari, Zaharaddeen Karami Lawal, and Idris Abdulmumin. 2021. https://doi.org/10.24203/ijcit.v10i4.86 A systematic literature review of hausa natural language processing . International Journal of Computer and Information Technology (2279-0764), 10(4)

  75. [84]

    Abubakar Yakubu Zandam, Fatima Adam Muhammad, and Isa Inuwa-Dutse. 2023. Online threats detection in hausa language. In 4th Workshop on African Natural Language Processing

  76. [85]

    Wanru Zhao, Yihong Chen, Royson Lee, Xinchi Qiu, Yan Gao, Hongxiang Fan, and Nicholas D. Lane. 2024. https://www.scopus.com/inward/record.uri?eid=2-s2.0-85200573076&partnerID=40&md5=994f545e0b90dd866db79d0d1a71067e Breaking physical and linguistic borders: Multilingual federat...

  77. [86]

    Linan Zhu, Zhechao Zhu, Chenwei Zhang, Yifei Xu, and Xiangjie Kong. 2023. Multimodal sentiment analysis based on fusion methods: A survey. Information Fusion, 95:306--325

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.