{"id":"dd532494-ad78-4691-8c82-ddcb4112b835","arxiv_id":"2504.14105","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Amplify's pilot with 155 experts produced an annotated dataset of 8,091 adversarial queries in seven languages, intended to evaluate AI safety and cultural relevance in Sub-Saharan Africa.","lead":"The paper describes Amplify, a platform through which African experts co-create tricky questions for AI chatbots, yielding over 8,000 examples in seven languages. It matters because locally written examples like these could help make AI models safer and more culturally accurate for communities that remain underrepresented in training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The pilot demonstrates a process for collecting culturally grounded queries, but the central claim requires the queries to actually be adversarial for model safety evaluation.","rationale":"The reader's weakest assumption focuses on the validation process in Appendix E and the absence of measured quality metrics; I agree that this is the central vulnerability. I would sharpen it further: the key missing quantity is not generic 'quality' but the adversarial effectiveness of the queries. An evaluation dataset must be shown to elicit unsafe or policy-violating outputs from representative models, or at least to exhibit high agreement with expert judgments that it would do so. The paper reports neither. Section 6.2.2's admission that Luganda and Igbo were hard to validate is especially telling because those languages are present in substantial numbers and are exactly the languages where the dataset would add the most value. This is an addressable gap rather than a fatal flaw: the methodology is documented in unusual detail, the local partnership structure is credible, and the data statistics are internally consistent. I therefore do not move the verdict; the existing CONDITIONAL verdict already captures the need for artifact availability and validation evidence. My concrete test would provide the missing outcome evidence and would also expose any per-language quality collapse that the current prose-level validation summary would hide.","tokens_in":17514,"tokens_out":2603,"duration_ms":25740,"concrete_test":"Take a random stratified sample of about 400 queries (approximately 50 per language, with English split by country), run them through two or three openly available LLMs under fixed decoding settings, and score the outputs with the same theme rubric used in the paper (hate speech, stereotypes, specialized advice, public interest, misinformation). Report the unsafe-response rate overall and per language, and compare it with a matched set of benign control queries. Also compute the human validation rejection rate from Appendix E on a separate subsample and, if feasible, inter-annotator agreement on validation labels for 50 queries. If the unsafe-response rate is close to the benign-control rate, or if it collapses in a specific language, the adversarial-evaluation claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the 8,091 annotated queries form an adversarial dataset usable for evaluating model safety and cultural relevance. The load-bearing condition is that a large fraction of the queries actually have the claimed property: they should elicit unsafe or policy-violating model responses, or at least reliably probe known failure modes, with correct annotations. The manuscript supplies process evidence but not outcome evidence. Section 4, step 7, says queries are translated, evaluated, and validated for coherence, semantic uniqueness, groundedness, and relevancy per Appendix E, but no results of that validation are reported: there are no pass/fail counts, no per-language or per-annotator quality rates, and no inter-annotator agreement. Section 6.2.2 states that validating Luganda and Igbo was difficult due to limited digital dictionaries and NLP tools, yet Table 11 counts 566 Luganda and 338 Igbo queries in the final dataset, so the least-validated languages constitute a material share of the data. Section 5 asserts that the queries 'are adversarial in nature and have a high likelihood of producing unsafe responses,' but no benchmark run, no safety classifier measurement, and no comparison with ordinary prompts is provided. Without such a measurement, the dataset could be coherent, culturally relevant, and well-annotated yet not adversarially effective; the claimed evaluation use case would then be unsupported. The descriptive statistics and thematic analyses may all be accurate, but the safety-evaluation claim requires evidence about model behavior, not only about input text.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Amplify Initiative, a data platform and methodology for co-creating localized, culturally grounded datasets with domain experts. It reports on a pilot conducted in Ghana, Kenya, Malawi, Nigeria, and Uganda, in which 155 experts produced 8,091 annotated adversarial queries in seven languages, targeting sensitive domains such as health, education, finance, and legal rights. The paper describes the seven-step methodology, the annotation taxonomy, the Android app used for collection, and provides descriptive statistics on language, domain, theme, and sensitive-characteristic distributions, along with qualitative examples of queries. The authors claim the resulting dataset can be used to evaluate LLM safety and cultural relevance in African contexts. The paper also candidly discusses practical challenges and limitations encountered during recruitment, training, validation, and scaling.","tokens_in":17751,"tokens_out":2585,"duration_ms":23674,"significance":"If the central claim can be substantiated, the paper would make a valuable contribution: it addresses a real gap in localized adversarial evaluation data for African languages, and it offers a participatory, community-centered methodology that is a useful template for other regions. The authors are transparent about limitations, and the paper includes concrete examples that give the reader a sense of the data's cultural specificity. The main strength is the process and the dataset's potential; the main weakness is that the paper provides no empirical outcome evidence that the queries actually function as adversarial safety-evaluation items. The significance of the contribution therefore hinges on whether the validation and quality assurance claims can be demonstrated, which the current manuscript does not do.","major_comments":[{"comment":"The validation step described in Section 4 (Step 7) and the validation principles in Appendix E (coherence, semantic uniqueness, groundedness, relevancy) are defined, but no results of the validation are reported. There are no pass/fail counts, no per-language or per-annotator quality rates, no inter-annotator agreement, and no rejection rate. Without these numbers, the claim that the 8,091 queries form a high-quality adversarial dataset is unsupported. This is load-bearing for the paper's central claim because the dataset's value as an evaluation benchmark depends on the queries being coherent, non-duplicative, grounded, and correctly annotated.","section":"Section 4, Step 7; Appendix E"},{"comment":"Section 6.2.2 states that validating data in languages like Luganda and Igbo was challenging due to the limited availability of digital dictionaries and NLP tools. Table 11 shows these languages account for 566 (Luganda) and 338 (Igbo) queries, over 11% of the dataset combined. The paper does not explain how the validation difficulties for these languages were resolved, nor does it provide any quality assurance evidence for them. This is concerning because the least-validated languages constitute a material portion of the data, and the paper gives no indication that the final dataset for these languages meets the same quality bar as the rest.","section":"Section 6.2.2; Table 11"},{"comment":"The paper asserts in Section 5 that the queries 'are adversarial in nature and have a high likelihood of producing unsafe responses from a large language model.' No empirical evidence is provided for this assertion: there is no benchmark run against any LLM, no safety classifier measurement, no comparison with non-adversarial or non-contextual prompts, and no measurement of the rate at which the queries actually elicit unsafe or policy-violating responses. Section 6.5.3 mentions future work to expand the dataset into a benchmarking suite, but the current claim of usability for safety evaluation is unsupported. The dataset could be coherent, culturally relevant, and well-annotated yet still not adversarially effective, so this missing evidence undermines the central claim.","section":"Section 5, first paragraph; Section 6.5.3"}],"minor_comments":[{"comment":"The affiliation for the Kenya-based author is listed as 'Jommo Kenyatta University Agriculture & Technology'; the standard spelling is 'Jomo Kenyatta University of Agriculture and Technology.'","section":"Author affiliation list"},{"comment":"In the sentence about Niger-Congo branches, 'V olta' contains a stray space and should be 'Volta.'","section":"Section 5.2"},{"comment":"The Igbo example query contains unusual dot separations (e.g., 'u. fo. du. ndi.'). If this reflects the tokenization or orthography, the presentation should be clarified; if it is a typographical artifact, it should be corrected.","section":"Table 4"},{"comment":"The paper mentions social desirability bias as a concern during the pilot but does not discuss any mitigation strategies or how it might affect the diversity of the collected queries.","section":"Section 6.3.1"},{"comment":"The abstract and Section 4 mention that the dataset and code are accessible via a link, but in the preprint no URL or DOI is actually visible, which hinders reproducibility and clarity about the dataset's availability.","section":"Abstract and Dataset Link"},{"comment":"The list of data authors is not formatted consistently (mixed capitalization, some entries with middle names, others with first name only). A consistent alphabetical list with full names would improve readability and credit attribution.","section":"Appendix H"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a dataset and methodology description rather than a full evaluation study. The missing outcome evidence for the adversarial quality of the queries is the key barrier to acceptance. Given the paper's stated purpose is to enable safety evaluation, the authors should either provide the missing quality metrics or substantially temper the claims about the dataset's immediate usability for evaluation. Additionally, the relationship between the authors and Google Research may warrant a check on the journal's disclosure norms, though the paper is transparent about this."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one if you work on multilingual evaluation or participatory data collection. The paper describes a pilot that produced 8,091 annotated queries in seven languages (English plus Luganda, Swahili, Chichewa, Igbo, Akan, and Nigerian Pidgin) with 155 domain experts across five Sub-Saharan African countries. The methodology is genuinely participatory: local researchers, hands-on training, a purpose-built app, validation principles in Appendix E, and a thoughtful discussion of incentives and ethical constraints. The descriptive statistics and thematic findings (e.g., mental health often framed as gendered, disability concerns in education, specific ethnic and cultural references) are plausible and give a real sense of local context. Credit where due: a new dataset in under-resourced languages with a documented end-to-end process is a useful contribution, and the paper is transparent about many limitations.\n\nThe soft spot is the one the abstract and conclusion rely on: the queries are described as adversarial and 'have a high likelihood of producing unsafe responses,' but the paper supplies no outcome evidence. There are no pass/fail counts from the validation step, no inter-annotator agreement, no rejection rates, and no runs of any model against the dataset. So we cannot tell whether a large share of these queries actually elicit unsafe or policy-violating responses, or whether the annotations are reliable. The paper even says validating Luganda and Igbo was hard due to missing NLP tools, yet those languages account for over 900 queries combined—a material fraction of the data. These are addressable gaps, but they are load-bearing for the safety-evaluation claim.\n\nI agree with the stress-test note: the process evidence is strong, but the central claim requires behavior-level evidence. This is not a fatal flaw if the paper is positioned as a methodology-and-pilot description rather than as a validated benchmark. But as written, the safety-evaluation use case is unsupported.\n\nThe dataset link mentioned in the abstract was not visible in the text I reviewed; I could not confirm public access.\n\nFor peer review: I would send this out. The topic matters, the dataset fills a documented gap, and the methodology is worth scrutiny. Serious referees should be able to help the authors either add validation measurements and benchmark results or reframe the claims. The paper is a solid first report, not a finished benchmark.","headline":"A credible participatory dataset paper whose core 'adversarial' claim is asserted, not yet demonstrated; deserving of peer review with room to strengthen.","tokens_in":18372,"tokens_out":2724,"would_cite":false,"duration_ms":24880,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pilot with 155 local experts produced 8,091 annotated adversarial queries in seven languages, giving model developers a way to test AI safety and cultural relevance beyond English.","keywords":["responsible AI","global AI","human-centered AI","socially centered data","adversarial queries","multilingual evaluation","African languages","data co-creation"],"falsifier":"Take a random sample of, say, 300 queries from the released dataset and have independent fluent speakers of each of the seven languages rate each query on the paper's own four validation principles—coherence, semantic uniqueness, groundedness, and relevance—and also judge whether it is genuinely adversarial; if a substantial share (for example, over 20 percent) fails, the claim that the dataset is ready to serve as an evaluation benchmark is undermined.","tokens_in":17348,"feed_emoji":"🌍","tokens_out":8030,"duration_ms":65117,"temperature":0.7,"pith_summary":"This paper reports on the Amplify Initiative pilot, which set out to show that localized AI evaluation data can be co-created with people who live in the contexts the data describes. A dataset co-authored by 155 experts across five countries in Sub-Saharan Africa produced 8,091 annotated adversarial queries—prompts likely to elicit unsafe model responses—in seven languages, including Luganda, Swahili, Chichewa, Igbo, Akan, and Nigerian Pidgin. The authors argue that this end-to-end approach, from partnerships and expert training through app-based collection, annotation, validation, and reward, is a reusable template for building datasets that test whether large language models are safe and culturally relevant outside English-centric settings. If that claim holds, model developers and local researchers would have a concrete path to evaluation data for languages and topics current benchmarks largely ignore.","feed_headline":"155 experts wrote 8,091 adversarial queries in seven languages","feed_subtitle":"The co-created dataset lets developers test AI safety and cultural fit in Luganda, Swahili, Chichewa, and beyond.","key_machinery":"The mechanism that carries the argument is the Amplify methodology, a seven-step pipeline for co-creating structured local-language datasets with domain experts. Its load-bearing parts are the expert selection process, the adversarial-query writing guidelines, and the annotation taxonomy of domains, themes (hate speech, stereotypes, specialized advice, public interest, misinformation or disinformation), and sensitive characteristics (age, gender, tribe, religion, disability, and others). These pieces turn open-ended local knowledge into structured records that can be retrieved, analyzed, and used as evaluation items. An Android app operationalizes the pipeline by training experts, unlocking query creation, blocking duplicates, and collecting annotations in a privacy-preserving way.","core_discovery":"The pilot's central discovery is that a structured, expert-led co-creation process can produce a substantial corpus of adversarial queries that capture local concerns rather than generic English internet content. The methodology proceeds through seven steps: forming partnerships with local researchers; choosing sensitive domains, topics, and experts; defining themes and sensitive characteristics; training experts; having them write and annotate queries in a purpose-built Android app; rewarding and recognizing contributors; and validating the collected data. The resulting dataset contains 8,091 queries annotated by domain, theme, and sensitive characteristic, so that a query about HIV misinformation in Uganda can be retrieved alongside its context. The paper shows, through five qualitative findings, that the data carry country-specific texture: health misinformation dominates, mental-health queries are gendered, disability concerns cluster in education, specific tribes appear as social groups, and cultural practices such as widow inheritance and healing dances surface across languages. The intended use is to evaluate large language models for safety and cultural relevance in these seven languages.","pith_inferences":["One implication the paper leaves implicit: if the dataset holds up under independent review, adversarial evaluation for low-resource languages can be produced without first building NLP tools for those languages, because the expert authors supply the linguistic judgment the tools would otherwise provide.","A testable next step the paper does not run is an audio-first version of the same pipeline, since the pilot itself notes many participants would rather speak than write, which might yield more naturalistic adversarial phrasing.","The annotation grid could serve as a cross-cultural diagnostic: rerunning the same domain, theme, and sensitive-characteristic structure in a different region would allow direct comparison of how harms like health misinformation are expressed across cultures."],"forward_implications":["The 8,091-query dataset can be used directly to evaluate large language models for safety and cultural relevance in Luganda, Swahili, Chichewa, Igbo, Akan, and Nigerian Pidgin.","The seven-step methodology provides a template for collecting similar localized evaluation data in other regions with minimal adaptation.","Because queries are annotated by domain, theme, and sensitive characteristic, the dataset supports targeted probing of specific harms, such as gendered mental-health stereotypes or health misinformation.","The structured, ontologically organized annotations can serve as a foundation for generating and validating synthetic data in these languages.","The pilot's training and recognition model—certificates, compensation, and data authorship—offers a pattern for sustainably engaging expert communities in data work."],"supporting_citations":[{"why":"Supplies the statistic that most LLM training text is over 90 percent English, motivating the data imbalance the pilot addresses.","marker":"Brinkmann et al., 2025"},{"why":"Establishes potential harms of large language models in Africa and shapes the sensitive domains selected for query collection.","marker":"Baguma et al., 2024"},{"why":"Documents the NLP diversity crisis that the pilot's localized dataset is meant to help address.","marker":"Joshi et al., 2020"},{"why":"Shows how few African languages translation tools cover, motivating localized collection rather than reliance on existing resources.","marker":"Dewitt Prat et al., 2024"},{"why":"Provides an existing multilingual reading-comprehension benchmark that the Amplify dataset complements in the evaluation-data space.","marker":"Bandarkar et al., 2024"},{"why":"Supplies an African-language benchmark as a comparison point for evaluation resources in the same languages and region.","marker":"Adelani et al., 2024"},{"why":"Evaluates large language model performance across African languages, setting the context for needing better localized data.","marker":"Ojo et al., 2024"},{"why":"Introduces a truthfulness benchmark for low-resource African languages, a direct predecessor in the evaluation-data space.","marker":"Bayes et al., 2024"}],"fun_headline_variants":["Local experts pen 8K adversarial queries for safer AI","8,091 co-created queries: making AI aware of local reality","Adversarial queries from 155 experts to test AI's cultural fit","8,091 questions written by local experts to harden AI for Africa","How 155 experts made 8,091 queries to keep AI locally aware"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value as safety-evaluation material rests on the assumption that the validation step actually caught the incoherent, duplicate, mislabeled, or non-adversarial queries, leaving a corpus that is genuinely fit for evaluating models.","fun_headline_variants_meta":{"raw":{"variants":["Local experts pen 8K adversarial queries for safer AI","8,091 co-created queries: making AI aware of local reality","Adversarial queries from 155 experts to test AI's cultural fit","8,091 questions written by local experts to harden AI for Africa","How 155 experts made 8,091 queries to keep AI locally aware"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000502,"raw_usage":{"total_tokens":2477,"prompt_tokens":993,"completion_tokens":1484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":1390}},"tokens_in":609,"tokens_out":1484,"duration_ms":10074,"temperature":1.0,"reasoning_tokens":1390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:55:25.228716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 300 queries from the released dataset and have independent fluent speakers of each of the seven languages rate each query on the paper's own four validation principles—coherence, semantic uniqueness, groundedness, and relevance—and also judge whether it is genuinely adversarial; if a substantial share (for example, over 20 percent) fails, the claim that the dataset is ready to serve as an evaluation benchmark is undermined.","supporting_citations":[],"review_version":1}