REVIEW 8 cited by
IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench -- a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference~(AfriXNLI), mathematical reasoning~(AfriMGSM), and multi-choice knowledge-based question answering~(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings~(where test sets are translated into English) across 10 open and six proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages~(such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63\% of the best-performing proprietary model GPT-4o performance. In addition, machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, such as Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.
Forward citations
Cited by 8 Pith papers
-
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
Training-data mix—not base-model language coverage—was the main driver of improved African-language performance, with Qwen-3 bases gaining the most after continued pre-training.
-
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models
A new benchmark shows state-of-the-art LLMs perform poorly on three Taiwanese indigenous languages across MT, ASR, and summarization.
-
Improving Multilingual Math Reasoning for African Languages
SFT on translated OpenMathInstruct data outperforms directly generated synthetic data for math in African languages, and combining both yields the best AfriMGSM scores.
-
Voice of a Continent: Mapping Africa's Speech Technology Frontier
A new benchmark and fine-tuned Simba models improve speech recognition, synthesis, and language identification across 61 African languages, but the claimed state of the art lacks comparisons to prior task-specific systems.
-
Against 'softmaxing' culture
A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.
-
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
mSTEB is a new 200+ language speech and text benchmark showing that LLMs perform substantially worse on low-resource African and Americas/Oceania languages, especially in speech tasks.
-
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.
- Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
Discussion (0). Continue with ORCID to comment.