Pith. sign in

REVIEW 1 cited by

AmQA: Amharic Question Answering Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03290 v2 pith:KQPRZPC2 submitted 2023-03-06 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords amharicdatasetlanguageamqaansweringbaselinedatasetsquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A translated Amharic version of SQuAD 2.0 is created with a proximity-aware alignment method, and fine-tuning XLM-R on it improves Amharic extractive QA scores on synthetic and human-curated dev sets.

Pith tools