Pith. sign in

REVIEW 2 cited by

Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.00948 v1 pith:AV6U3SAF submitted 2024-12-01 cs.CL

classification cs.CL
keywords languagesmodelsafricanbenchmarkuhuralow-resourcelrlsanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluations of Large Language Models (LLMs) on knowledge-intensive tasks and factual accuracy often focus on high-resource languages primarily because datasets for low-resource languages (LRLs) are scarce. In this paper, we present Uhura -- a new benchmark that focuses on two tasks in six typologically-diverse African languages, created via human translation of existing English benchmarks. The first dataset, Uhura-ARC-Easy, is composed of multiple-choice science questions. The second, Uhura-TruthfulQA, is a safety benchmark testing the truthfulness of models on topics including health, law, finance, and politics. We highlight the challenges creating benchmarks with highly technical content for LRLs and outline mitigation strategies. Our evaluation reveals a significant performance gap between proprietary models such as GPT-4o and o1-preview, and Claude models, and open-source models like Meta's LLaMA and Google's Gemma. Additionally, all models perform better in English than in African languages. These results indicate that LMs struggle with answering scientific questions and are more prone to generating false claims in low-resource African languages. Our findings underscore the necessity for continuous improvement of multilingual LM capabilities in LRL settings to ensure safe and reliable use in real-world contexts. We open-source the Uhura Benchmark and Uhura Platform to foster further research and development in NLP for LRLs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KatotohananQA is a Filipino translation of TruthfulQA; seven LLMs scored 94.72% in English versus 83.87% in Filipino, with GPT-5 and GPT-5 mini showing the smallest gap.

  2. The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.

Pith tools