{"id":"4589e09a-e911-48f1-8b2d-cd1c0e7b60db","arxiv_id":"2412.12417","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Human-translated benchmarks in eight African languages show GPT-4o accuracy is 12 to 20 percentage points below English, and fine-tuning on translated data recovers part of the gap.","lead":"The authors translated two established English benchmarks into eight low-resource African languages, measured how much worse large language models perform in those languages, and tested fine-tuning strategies to shrink the gap. This provides public evaluation tools for languages spoken by more than 160 million people and quantifies what actually helps improve model performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MMLU translations have no human quality check, and ROUGE-1 similarity to Google Translate is extreme (up to 79% of Bambara rows >0.95), so the headline 12.0-19.9 pp gap may partly reflect translation artifacts.","rationale":"The reader's weakest assumption is that the unvalidated MMLU translations could make the central numeric claims measure artifacts; the strongest claim indeed depends on translated benchmarks faithfully reflecting reasoning ability. My read confirms this as the most load-bearing concern, with concrete supporting evidence in the paper itself: no human MMLU quality review, ROUGE-1 distributions that flag either heavy MT similarity or low similarity, noisy Winogrande annotator agreement, and human-translated scores below machine-translated scores for some language-benchmark cells. This is not a fatal flaw: the paper releases substantial public resources, includes human evaluation for Winogrande (however noisy), reports reproducibility details, and the largest fine-tuning gains are large enough that they probably survive moderate translation noise. But the specific magnitude of the 12.0%-19.9% gap, and the fine-tuning improvements measured on the same translations, should be made robust to translation quality before the numbers are treated as precise capability measurements. The proposed check (independent expert review of a stratified MMLU sample and recomputation on the clean subset) would settle whether the concern lands. Since the reader already assigned a CONDITIONAL verdict for essentially this reason, no verdict adjustment is needed.","tokens_in":74299,"tokens_out":5999,"duration_ms":58849,"concrete_test":"Commission two independent medical-domain translators, blind to the vendor output and to Google Translate, to rate a stratified random sample of 100 MMLU rows per language (college medicine, clinical knowledge, virology) for translation adequacy and for whether the correct answer choice is preserved. Then recompute GPT-4o's English-vs-African gap and the mono-lingual fine-tuning gains using only rows rated adequate and answer-preserving by both reviewers. If the gap shrinks by more than about 5 percentage points or the fine-tuning gains largely disappear on the clean subset, the headline numbers are substantially translation artifacts; if the gap and gains persist, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim is that GPT-4o shows a 12.0%-19.9% absolute performance gap between English and the average of 11 African languages, supported by translated Winogrande and MMLU benchmarks. For this gap to measure capability rather than translation artifact, the translated prompts must preserve question semantics and answerability. That condition is not established for the MMLU sections: Appendix A reports no human quality review of the Translated.com output, and the only fidelity check is ROUGE-1 against Google Translate (Figure A.16). The scores are extreme and bimodal: up to 79% of Bambara MMLU rows have ROUGE-1 >0.95 (near-identical to MT), while up to 60% of Sesotho rows have ROUGE-1 <0.5 (highly dissimilar to MT). Both patterns are consistent with systematic translation problems, but in opposite directions, and the paper provides no validator to adjudicate correctness. For Winogrande, human evaluation exists, but inter-annotator agreement is very low (Cohen's kappa max 0.211 for quality and near 0 for appropriateness, Table A.16), so even the 'good translation' labels are noisy. The machine-translation comparison in Tables A.11-A.15 shows that human-translated MMLU can be substantially worse than machine-translated MMLU (e.g., GPT-4o Igbo college medicine: human 58.4 vs machine 68.2, difference -9.8 pp), so translation inadequacy can depress human-translated scores and inflate the measured gap. If MMLU translation errors are systematic, both the headline gap and the fine-tuning gains computed on the same translations partly measure artifacts rather than model capability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces human-translated versions of Winogrande and three clinical MMLU sections into eight low-resource African languages (plus three prior languages), evaluates a range of LLMs on these benchmarks, and reports a consistent 12.0%-19.9% absolute performance gap between English and the average of eleven African languages. It then fine-tunes Llama 3 70B over 400 configurations, reporting mono-lingual gains of 5.6%, cross-lingual gains of 2.9%, and a 3.0% out-of-the-box lift on culturally appropriate Winogrande items. The benchmarks, translations, and code are released publicly.","tokens_in":74685,"tokens_out":3937,"duration_ms":38086,"significance":"If the measurements are valid, this is a useful contribution: it expands reasoning-benchmark coverage to eight underrepresented languages, provides a large public translation resource, and gives practical evidence about fine-tuning with limited data. The paper is unusually transparent about reproducibility (seeds, prompts, hyperparameters, and full appendix tables) and about its own limitations, including the absence of statistical tests and variability in translation quality. However, the headline gap and the fine-tuning conclusions rest on translation fidelity and annotation reliability, both of which are not yet adequately established.","major_comments":[{"comment":"The MMLU translations received no human quality review; the only fidelity check is ROUGE-1 similarity to Google Translate, which is extreme and bimodal (up to 79% of Bambara rows have ROUGE-1 > 0.95, while up to 60% of Sesotho rows have ROUGE-1 < 0.5). ROUGE-1 similarity to a machine translation cannot distinguish a correct translation from one that merely resembles MT output, so the medical knowledge scores in Table 2 may be depressed by systematic translation errors. This is not hypothetical: Tables A.12-A.13 show that GPT-4o on Igbo college medicine scores 58.4 on the human translation but 68.2 on the machine translation, a 9.8-point gap in the opposite direction of the headline claim. Because these MMLU scores feed directly into the 12.0-19.9% gap and the fine-tuning experiments, the authors need to add a human validation sample of the MMLU translations (or an equivalent diagnostic) before the central claim can be taken at face value.","section":"Appendix Section A and Tables A.12-A.13"},{"comment":"The Winogrande quality and cultural-appropriateness annotations have very low inter-annotator agreement: the maximum Cohen's kappa for translation quality is 0.211 (Shona), and for appropriateness it is near zero for all languages (e.g., 0.074 for Bambara). The paper nonetheless uses these labels to define \"good/understandable\" translations and to split data into \"culturally appropriate\" vs \"culturally inappropriate\" subsets, from which it concludes a 3.0% average performance boost. With near-zero agreement, the appropriateness split is largely noise, and the measured lift could reflect translation quality or other confounds rather than a cultural construct. The authors acknowledge the low kappa but do not provide an adjudicated or consensus-based set of labels to show that the split is reliable. Please report results on an adjudicated subset or demonstrate that the effect is robust to alternative annotation aggregation schemes.","section":"Section C and Table A.16"},{"comment":"The headline performance gap is reported as point estimates without confidence intervals or significance tests. For MMLU virology, the test set has only 166 questions per language; a 12 percentage-point gap on that sample has a standard error of roughly 3 points, so the lower end of the claimed 12.0-19.9% range is not statistically distinguishable from chance-level differences for some language-benchmark combinations. The reproducibility checklist explicitly states that no statistical tests were used. This does not invalidate the direction of the result, but the central quantitative claim should be accompanied by uncertainty estimates (e.g., bootstrap confidence intervals) or at least by the per-benchmark sample sizes needed to assess the precision of the gap.","section":"Results, Tables 1 and 2"},{"comment":"The high/low-quality split is based on GPT-4o LLM-as-an-Annotator scores, with no human validation that these scores correlate with actual fine-tuning usefulness. The 5.4% average advantage of the high-quality tertile could be driven by surface features (e.g., question length, topic balance, or translationese patterns) rather than by the construct of \"quality.\" Since the abstract and discussion present high-quality dataset fine-tuning as a key finding, the authors should either validate the GPT-4o ratings against a human-rated sample or show that the effect persists after controlling for simple confounds such as length and lexical diversity.","section":"Fine-tuning with Varying Data Quality and Quantity"}],"minor_comments":[{"comment":"There is a typo in the phrase \"Lllama 3 70B\" (extra 'l'); please correct it.","section":"Fine-tuning with Varying Languages and Domains"},{"comment":"The color bar for Cohen's kappa ranges from 0.0 to 1.0, but several cells contain negative kappa values; this makes the visualization misleading. Please extend the color scale to negative values or recode the affected cells explicitly.","section":"Figure A.14"},{"comment":"The abstract states \"approximately 1 million human-translated words,\" but summing the per-language counts in the Methods (73,742 words for Winogrande plus 27,107 words for the three MMLU sections, times 8 languages) gives approximately 807,000 words, which is closer to 0.8 million. Please either recompute the total or revise the wording to be accurate.","section":"Abstract and Methods"}],"recommendation":"major_revision","confidential_remarks":"The paper's resource contribution is genuinely valuable, and the experimental design is much more transparent than typical for this area. The main risk is that the headline gap and the quality-based fine-tuning result depend on translation fidelity and annotation reliability that are not yet demonstrated. I would be open to accepting after the authors add a human validation sample for the MMLU translations, report uncertainty measures for the headline numbers, and either adjudicate or re-analyze the appropriateness annotations. The current version is not yet ready."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, the paper ships the goods: roughly a million words of translated Winogrande and clinical MMLU into Amharic, Bambara, Igbo, Sepedi, Shona, Sesotho, Setswana, and Tsonga, with cultural annotations, plus code and data on GitHub and HuggingFace. That is a real resource for multilingual evaluation. Second, the headline numbers—GPT-4o showing a 12.0–19.9 point absolute gap between English and the average of 11 African languages—are plausible but not bulletproof, because the MMLU half of that measurement rests on translations that received no human quality review.\n\nWhat the paper does well: the translation pipeline for Winogrande is documented in detail—translators, validators, evaluators, costs, inter-annotator agreement figures—and the data is released with original and corrected versions. That level of transparency is rare and valuable. The fine-tuning study is also broad: over 400 LoRA-tuned Llama 3 70B models, mono- and cross-lingual transfer, and a data-quality/quantity comparison. The finding that domain-matched fine-tuning (college medicine → clinical knowledge) gives the largest gains, and that high-quality data beats low-quality data of the same size, is credible and useful for practitioners.\n\nWhere I'd push back, and here the stress-test note is right: the MMLU translations are checked only with ROUGE-1 against Google Translate, which is not a validity check. The paper admits this in Appendix A, but it still uses those numbers to compute the headline gap. The bimodal ROUGE-1 pattern—Bambara very close to MT, Sesotho very far from it—suggests inconsistent translation quality, and the machine-translation comparison tables show cases where human-translated scores are worse than machine-translated ones (e.g., Igbo college medicine, GPT-4o: 58.4 vs 68.2). So part of the measured gap could be an artifact of translation errors, not just model capability. That doesn't sink the paper, but the claims should be tempered and ideally the MMLU translations should be spot-checked by humans.\n\nThe Winogrande cultural-appropriateness analysis is shakier. The inter-annotator kappa for appropriateness is near zero for most languages (max 0.09 for Zulu/Xhosa), and quality kappa tops out at 0.211 for Shona. The authors define \"inappropriate\" as \"either annotator flagged it\", which is a defensible choice given the noise, but the 3.0% lift figure is built on labels that are not reliably reproducible.\n\nMinor issues: the \"first known translations\" claim in the ethical statement is overreached, since IrokoBench covers at least some of these languages and the paper doesn't demonstrate absence elsewhere. The baseline tables (Table 1 and 2) have no error bars; fine-tuning tables do. The data-quality tertile split is post-hoc, and the exact split is a free parameter.\n\nBottom line: this deserves a serious referee. It's a resource paper with a lot of value, and the authors are upfront about many of the limitations. The reviewer should push for a human spot-check of MMLU translations and for more cautious wording on the gap and cultural-appropriateness claims, but the core resource is solid and the fine-tuning results are a real contribution. I'd take it to peer review and I'd cite it for the datasets.","headline":"A genuinely useful benchmark and fine-tuning resource for eight low-resource African languages, with translation-quality caveats that need to be stated more carefully.","tokens_in":75202,"tokens_out":2895,"would_cite":true,"duration_ms":26219,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Newly translated Winogrande and clinical MMLU benchmarks in eight African languages reveal a 12.0–19.9 percentage-point performance gap between English and the average African language for the best LLM (GPT-4o), and fine-tuning on the…","keywords":["LLM performance gap","low-resource African languages","benchmark translation","Winogrande","MMLU clinical sections","fine-tuning","cross-lingual transfer","cultural appropriateness"],"falsifier":"Take a random sample of 100 MMLU clinical-knowledge questions, have two independent professional medical translators translate them into Bambara and Igbo, then have a third expert resolve disagreements; if GPT-4o's accuracy on these adjudicated translations exceeds that on the paper's translations by more than a few points, translation quality is a substantial confound of the reported gap.","tokens_in":74140,"feed_emoji":"🗣️","tokens_out":9528,"duration_ms":75960,"temperature":0.7,"pith_summary":"The paper attempts to make LLM performance measurable in eight low-resource African languages by creating about one million human-translated words of established reasoning and medical benchmarks, and then uses those benchmarks to quantify how much worse state-of-the-art models perform compared with English. The headline finding is that the best model, GPT-4o, shows a performance gap of 12.0 to 19.9 percentage points between English and the average of the 11 African languages covered. The paper also argues that the gap is partly closable: fine-tuning on the translated data yields average mono-lingual gains of 5.6%, high-quality data beats low-quality data by 5.4%, cross-lingual transfer adds 2.9%, and GPT-4o scores about 3.0% higher out-of-the-box on questions that native speakers judge culturally appropriate. If true, the work provides reusable evaluation resources and evidence for data-centric strategies to reduce language inequity in LLMs.","feed_headline":"New tests expose 12–20 point LLM gap in African languages","feed_subtitle":"Human-translated benchmarks for 8 African languages quantify the gap and show fine-tuning recovers ~5.6%.","key_machinery":"The central object is the newly released parallel benchmark: Winogrande and three clinical sections of MMLU (college medicine, clinical knowledge, virology) translated into eight languages — Amharic, Bambara, Igbo, Sepedi, Shona, Sesotho, Setswana, Tsonga — and used together with existing translations for Afrikaans, Zulu, and Xhosa. These translated QA sets carry the argument because they make possible the first direct English-versus-African-language comparison on the same multiple-choice questions; the fine-tuning machinery (mono- vs cross-lingual data, GPT-4o quality scoring, and data-volume sampling) then tests whether the gap can be reduced and which data characteristics matter.","core_discovery":"The central discovery, on the paper's own terms, is that the performance gap between LLMs in English and in native African languages is large and measurable, not a side effect of missing benchmarks: GPT-4o drops 12.0% to 19.9% absolute in accuracy when evaluated on human-translated MMLU clinical sections and Winogrande in 11 African languages, with the largest single-language deficits for Bambara. The paper's second discovery is that fine-tuning on modest amounts of translated data narrows that gap, with average mono-lingual improvements of 5.6%, and that data quality and cultural suitability matter: using LLM-selected high-quality training samples improves performance by 5.4% over low-quality samples, and GPT-4o performs 3.0% better out-of-the-box on questions labeled culturally appropriate by native speakers.","pith_inferences":["Because model performance tracks language family and resource level (Afrikaans outscores Bambara by 22.9–56.1 points on every benchmark), the dominant driver of the gap is likely pre-training data exposure rather than task difficulty, which argues for directing translation and fine-tuning resources to Mande and Semitic family languages first.","The paper's machine-translation comparisons show that for several models, machine-translated queries can match or beat human-translated ones on MMLU, hinting that models have learned a 'translationese' distribution; a direct test would be to fine-tune the same model separately on human vs machine translations and compare generalization to new human text.","Because the cultural-appropriateness annotations show very low inter-annotator agreement (the highest kappa value is 0.211), the reported 3.0% cultural boost is likely an underestimate of the true effect; a replication using more annotators per item or a culturally grounded taxonomy could reveal a larger and more consistent effect."],"forward_implications":["If the gap is real, state-of-the-art LLMs are substantially less trustworthy in these languages on medical and reasoning questions, which matters for any attempt to deploy AI-assisted health information for the over 160 million speakers covered.","Fine-tuning on hundreds of translated examples can recover a meaningful part of the gap (5.6% average mono-lingual gain), so collecting small, high-quality translated datasets is a viable improvement path.","Aligning fine-tuning domain with target domain gives the largest boosts (up to 17.4% for college-medicine tuning evaluated on clinical knowledge), so resource creators should prioritize domain-matched data.","Quality filtering with an LLM annotator is worth about 5.4% over unfiltered data, meaning data selection can substitute for larger data volume in low-resource settings.","Questions that native speakers judge culturally appropriate are easier for GPT-4o (3.0% boost), so cultural adaptation of evaluation content affects measured capability."],"supporting_citations":[{"why":"Provides the Winogrande QA pairs and binary coreference task that were translated into the 8 African languages and used as a reasoning benchmark.","marker":"Sakaguchi et al. 2021"},{"why":"Provides the three clinical MMLU sections (college medicine, clinical knowledge, virology) that were translated and used as medical knowledge benchmarks.","marker":"Hendrycks et al. 2021b"},{"why":"Supplies the existing Afrikaans, Zulu, and Xhosa translations of Winogrande and clinical MMLU that the paper uses to extend the evaluation to 11 African languages.","marker":"BMGF 2024"},{"why":"Provides the Belebele reading-comprehension benchmark in 122 languages, used as an independent cross-check that corroborates the gap across all 11 African languages.","marker":"Bandarkar et al. 2024"}],"fun_headline_variants":["African LLM gap: 12–20 points, fine-tuning recovers 5.6%","Human-translated tests expose 12–20 point gap in African LLMs","Fine-tuning narrows LLM performance gap in 8 African languages","Cultural appropriateness adds 3% to LLM scores in African tests","New benchmarks quantify LLM deficit in African languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the translated benchmarks being accurate measures of meaning; the MMLU medical translations were never reviewed by human medical experts, and in Bambara up to 79% of MMLU rows are near-duplicates of automatic machine-translation output, so systematic translation errors could be inflating (or masking) part of the 12–20 point gap.","fun_headline_variants_meta":{"raw":{"variants":["African LLM gap: 12–20 points, fine-tuning recovers 5.6%","Human-translated tests expose 12–20 point gap in African LLMs","Fine-tuning narrows LLM performance gap in 8 African languages","Cultural appropriateness adds 3% to LLM scores in African tests","New benchmarks quantify LLM deficit in African languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000467,"raw_usage":{"total_tokens":2364,"prompt_tokens":1018,"completion_tokens":1346,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":1250}},"tokens_in":634,"tokens_out":1346,"duration_ms":9503,"temperature":1.0,"reasoning_tokens":1250,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:07:04.690958+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 100 MMLU clinical-knowledge questions, have two independent professional medical translators translate them into Bambara and Igbo, then have a third expert resolve disagreements; if GPT-4o's accuracy on these adjudicated translations exceeds that on the paper's translations by more than a few points, translation quality is a substantial confound of the reported gap.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the existing Afrikaans, Zulu, and Xhosa translations of Winogrande and clinical MMLU that the paper uses to extend the evaluation to 11 African languages."},{"cited_title":"N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; and Khabsa, M","cited_arxiv_id":null,"evidence_quote":"Provides the Belebele reading-comprehension benchmark in 122 languages, used as an independent cross-check that corroborates the gap across all 11 African languages."}],"review_version":1}