Pith. sign in

REVIEW 2 cited by

AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09828 v3 pith:FYVKLUW7 submitted 2023-11-16 cs.CL

classification cs.CL
keywords languagesafricanevaluationmetricshumancometcorrelationdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the recent progress on scaling multilingual machine translation (MT) to several under-resourced African languages, accurately measuring this progress remains challenging, since evaluation is often performed on n-gram matching metrics such as BLEU, which typically show a weaker correlation with human judgments. Learned metrics such as COMET have higher correlation; however, the lack of evaluation data with human ratings for under-resourced languages, complexity of annotation guidelines like Multidimensional Quality Metrics (MQM), and limited language coverage of multilingual encoders have hampered their applicability to African languages. In this paper, we address these challenges by creating high-quality human evaluation data with simplified MQM guidelines for error detection and direct assessment (DA) scoring for 13 typologically diverse African languages. Furthermore, we develop AfriCOMET: COMET evaluation metrics for African languages by leveraging DA data from well-resourced languages and an African-centric multilingual encoder (AfroXLM-R) to create the state-of-the-art MT evaluation metrics for African languages with respect to Spearman-rank correlation with human judgments (0.441).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring In-context Example Generation for Machine Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DAT generates query-specific in-context translation examples using only an LLM, improving English-to-low-resource translation over zero-shot in most tested languages.

  2. Voice of a Continent: Mapping Africa's Speech Technology Frontier

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned Simba models improve speech recognition, synthesis, and language identification across 61 African languages, but the claimed state of the art lacks comparisons to prior task-specific systems.

Pith tools