REVIEW 4 cited by
AfroBench: How Good are Large Language Models on African Languages?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale multilingual evaluations, such as MEGA, often include only a handful of African languages due to the scarcity of high-quality evaluation data and the limited discoverability of existing African datasets. This lack of representation hinders comprehensive LLM evaluation across a diverse range of languages and tasks. To address these challenges, we introduce AfroBench -- a multi-task benchmark for evaluating the performance of LLMs across 64 African languages, 15 tasks and 22 datasets. AfroBench consists of nine natural language understanding datasets, six text generation datasets, six knowledge and question answering tasks, and one mathematical reasoning task. We present results comparing the performance of prompting LLMs to fine-tuned baselines based on BERT and T5-style models. Our results suggest large gaps in performance between high-resource languages, such as English, and African languages across most tasks; but performance also varies based on the availability of monolingual data resources. Our findings confirm that performance on African languages continues to remain a hurdle for current LLMs, underscoring the need for additional efforts to close this gap. https://mcgill-nlp.github.io/AfroBench/
Forward citations
Cited by 4 Pith papers
-
mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages
mRAKL reformulates multilingual knowledge graph completion as question answering and shows that retrieving context from Wikipedia improves tail-entity prediction for Tigrinya and Amharic, with gains up to 8.79 points ...
-
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
mSTEB is a new 200+ language speech and text benchmark showing that LLMs perform substantially worse on low-resource African and Americas/Oceania languages, especially in speech tasks.
-
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.
-
LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
LLMs and agentic AI are presented as a transformative opportunity for African insurance, with a call for African-led, equitable AI strategies.
Discussion (0). Sign in to comment.