Pith. sign in

REVIEW 2 cited by

AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06897 v1 pith:IUPLWJBT submitted 2023-05-11 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords languagesafricanafriqacross-lingualcontentansweransweringquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

African languages have far less in-language content available digitally, making it challenging for question answering systems to satisfy the information needs of users. Cross-lingual open-retrieval question answering (XOR QA) systems -- those that retrieve answer content from other languages while serving people in their native language -- offer a means of filling this gap. To this end, we create AfriQA, the first cross-lingual QA dataset with a focus on African languages. AfriQA includes 12,000+ XOR QA examples across 10 African languages. While previous datasets have focused primarily on languages where cross-lingual QA augments coverage from the target language, AfriQA focuses on languages where cross-lingual answer content is the only high-coverage source of answer content. Because of this, we argue that African languages are one of the most important and realistic use cases for XOR QA. Our experiments demonstrate the poor performance of automatic translation and multilingual retrieval methods. Overall, AfriQA proves challenging for state-of-the-art QA models. We hope that the dataset enables the development of more equitable QA technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

    cs.CL 2025-07 conditional novelty 5.0 of 10

    KinyaColBERT, a morphology-aware two-tier ColBERT retriever, reports large MRR gains over multilingual baselines and commercial APIs on a new Kinyarwanda agricultural retrieval benchmark.

  2. The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes

    cs.CY 2025-02 conditional novelty 4.0 of 10

    A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.

Pith tools