{"id":"375aa917-986d-40ba-a1c4-689f4e053d36","arxiv_id":"2412.11763","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"QUENCH introduces a generation-based quiz benchmark with masked entities and rationales, and documents a consistent Indic-versus-non-Indic performance gap across seven LLMs.","lead":"QUENCH is a 400-question English quiz benchmark built from YouTube quiz videos, with missing names and hand-written explanations. It shows large language models answer Indian-context questions substantially worse than non-Indian ones, revealing a measurable cultural knowledge gap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline Indic/non-Indic gap rests on an unvalidated LLM jury: BERTScore gaps are 0.6-2.4 points while GEval gaps are 12-32 points, so the reported gap could be judge bias about Indic content rather than real model knowledge.","rationale":"The paper is a useful and reasonably careful benchmark contribution: the dataset is novel, contamination checks are reported, and the evaluation setup is described in enough detail to be reproduced. However, the load-bearing empirical statement about a systematic Indic gap depends on the LLM jury being unbiased. The reader's weakest assumption identifies exactly this vulnerability, and my reading of Tables 4, 7, and 8 confirms it: standard semantic metrics show only small or inconsistent gaps, so the large GEval deltas are doing essentially all the work. The lack of any human validation of the jury, together with the internal count inconsistency in the human benchmarking section (20 sampled questions versus 10 Non-Indic plus 13 Indic), means the paper should not yet be accepted as establishing the magnitude, or even the existence, of a true Indic-specific reasoning deficit. A targeted human-validation experiment would settle the question directly. I therefore endorse the CONDITIONAL verdict without changing it: the concern is real but fixable, and the benchmark itself appears valuable.","tokens_in":21538,"tokens_out":3054,"duration_ms":39676,"concrete_test":"Stratified human validation of the jury: sample about 50 Indic and 50 Non-Indic entity predictions (and a matched set of rationale generations) across models, have two or three human annotators with Indic cultural competence score them using the exact same binary/Likert prompts and gold answers from Appendix C, then compare human scores with jury scores per subset and compute Δ(N-I) under human labels with bootstrap confidence intervals. If the human-judged gap stays within roughly 5 points of the GEval gap for most models, the Indic gap is real; if the human-judged gap shrinks substantially, the headline result is an LLM-judge artifact. Report inter-annotator agreement and also break out each judge's scores separately to detect self-preference effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that all benchmarked LLMs perform poorly on Indic-context questions, is supported almost entirely by LLM-jury (GEval) scores computed with judges that include GPT-4-Turbo, Mixtral-8x7B, and Meta-Llama-3-70B, three of the seven evaluated models. The paper does not validate this jury against human judgments, and the only human benchmarking is too small and internally inconsistent to serve that role: Section 6 says 20 questions were sampled with equal numbers from both subsets, then reports 10 Non-Indic and 13 Indic questions, and the authors explicitly refrain from broad conclusions. Because BERTScore differences between Indic and Non-Indic are only 0.6 to 2.4 points and BLEU/ROUGE differences are mixed or even negative, the 12 to 32 point GEval gaps in Table 4 carry nearly the entire evidentiary weight. The evaluation prompt in Appendix C gives judges the true answer and asks for a binary or Likert judgment, but if a judge consistently fails to accept valid Indic aliases, transliterations, or less canonical phrasings as correct, the measured gap would be inflated. This is plausible a priori because the judges themselves are LLMs likely trained on predominantly Western-centric data, yet no agreement analysis between judges and human annotators is reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces QUENCH, a manually curated English quiz benchmark of 400 open-domain questions with masked entities and human-written rationales, tagged as Indic (70 questions) or Non-Indic (330 questions). Seven LLMs (GPT-4-Turbo, GPT-3.5-Turbo, Gemini-1.5-Flash, Gemma-1.1-7B, Mixtral-8x7B, Llama-3-8B, Llama-3-70B) are evaluated zero-shot with and without chain-of-thought prompting using BLEU, ROUGE-L, BERTScore, and an LLM-as-judge jury (GPT-4-Turbo, Mixtral-8x7B, Llama-3-70B). The central empirical claim is that all benchmarked models perform substantially worse on Indic-context questions than on Non-Indic ones, with GEval gaps of 12 to 32 points for entity prediction and 8 to 20 points for rationale prediction. The authors also report that CoT has little effect, that larger models tend to outperform smaller ones, that gold-labeled rationales improve rationale generation, and that even the best model commits characteristic entity-recognition errors. A data-contamination check using WIMBD and Infinigram finds no significant leakage of the source material into common pretraining corpora.","tokens_in":21785,"tokens_out":4444,"duration_ms":43439,"significance":"QUENCH is a potentially valuable contribution: it is an open-domain, non-MCQ quiz benchmark with multi-entity masking and gold rationales, it is released with code and data, and the contamination check is careful and reassuring. If the reported Indic/non-Indic gap is taken at face value, the benchmark provides a reusable instrument for quantifying the Western-centric bias of LLM world knowledge. The strengths of the paper are the dataset construction, the multi-metric evaluation protocol, and the explicit leakage analysis. However, the headline gap is carried almost entirely by LLM-jury scores that are not validated against human judgments, and the standard lexical/semantic metrics show only small or mixed differences; the paper's central claim therefore needs additional evidence before it can be considered established.","major_comments":[{"comment":"The central Indic/non-Indic gap is computed with an LLM jury that is never validated against human judgments. The gap magnitudes differ dramatically across metrics: BERTScore deltas are only 0.6–2.4 points and BLEU/ROUGE deltas are often negative or mixed, while the GEval deltas are 12–32 points. Because the judges are themselves LLMs and one of them (GPT-4-Turbo) is also a benchmarked model, the possibility that the jury is biased against acceptable Indic paraphrases or transliterations is a real threat to the main claim. The paper should report per-judge scores, inter-judge agreement, and, crucially, a human-annotated sample scored with the same binary/Likert rubric. Without such validation, the headline gap may reflect judge bias rather than model knowledge.","section":"§4, Evaluation Metrics; Table 4; Appendix C"},{"comment":"The human benchmarking section is internally inconsistent and too small to serve as a jury validation. It states that 20 questions were sampled with equal numbers from both subsets, but then reports 10 Non-Indic and 13 Indic questions (totaling 23). The authors explicitly refrain from broad conclusions, and the human scores are not compared to GEval scores on the same questions. As written, this section cannot substantiate the claim in §5.3 that \"all the benchmarking LLMs perform poorly at questions with an Indic context for both entity and rationale generation tasks.\" The authors should either provide a proper human-evaluation study on LLM outputs or temper the claim to acknowledge that the jury metric is not externally validated.","section":"§6, Human Benchmarking; §5.3"},{"comment":"The claim that the Indic gap is observable \"across all metrics\" is not supported by the full tables. For rationale generation with predicted entities, BLEU deltas are negative for several models (e.g., Gemini 1.5 Flash without CoT: −11.8; GPT-4-Turbo: −1.7), BERTScore deltas are near zero (0.0 to 0.4), and only the GEval deltas are large. The manuscript does not provide confidence intervals, significance tests, or effect sizes for any of the reported differences. Given the small Indic subset (70 questions), the authors should report uncertainty estimates and base the headline conclusion on metrics that are robust to the judge-bias concern, or at least clearly separate the jury-based result from the standard-metric result.","section":"§5.3 and Tables 7–8"},{"comment":"The jury description is ambiguous about self-scoring. The text says each judge scores \"every other benchmarked LLM,\" which would give 3×6=18 judge-model combinations, yet the paper reports 21. If the judges in fact scored themselves (3×7=21), then the self-preference bias documented by Panickssery et al. (2024), which the authors cite, directly affects the reported scores and the model ranking. The authors should clarify whether self-evaluation was included, and if so, analyze its effect on the results.","section":"§4, Evaluation Metrics; footnote on jury composition"}],"minor_comments":[{"comment":"There are several typos and word-choice errors: \"B enchmark\" in the abstract, \"access\" should be \"assess\" in multiple places (e.g., Section 1 and Related Work), and \"explainations\" in Section 6.","section":"Abstract and throughout"},{"comment":"The wording \"we randomly sample 20 questions ... We sample equal numbers from both subsets and across all themes\" is contradicted by the reported 10 Non-Indic and 13 Indic questions; this arithmetic inconsistency should be fixed.","section":"§6, Human Benchmarking"},{"comment":"The prompt listed for \"generating rationale using gold labels\" appears to be identical to the entity-prediction prompt and does not actually provide the gold entity or ask for a rationale. This appears to be a copy-paste error, but as written it makes the gold-label rationale experiment non-reproducible from the paper alone.","section":"Appendix B, gold-label rationale prompt"},{"comment":"Figure 4(c)–(d) refer to \"Indic and Non-Indic languages,\" but the dataset is entirely in English; these should be labeled as \"contexts\" or \"subsets\" to avoid confusion.","section":"Figure 4 and Table 4 captions"},{"comment":"The error taxonomy lists \"Correct Answer\" as an error type, which is conceptually odd; the authors should rename this category (e.g., \"Correct answer despite failed reasoning\") or clarify why a correct answer is counted as an error.","section":"§7, Error Analysis, Table 5"},{"comment":"The conclusion says \"by accessing CoT,\" which should read \"by assessing CoT\" or \"by comparing CoT.\" Also, the statement that the Indic gap is \"significant\" should be qualified in light of the lack of statistical significance testing noted in the major comments.","section":"§8, Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a useful and carefully curated dataset, and the contamination analysis is a genuine strength. The main concern is that the paper's headline claim rests on an unvalidated LLM jury, with the human benchmarking section being too small and internally inconsistent to serve as validation. I would encourage the editor to require the authors to either add a human validation study of the jury or substantially weaken the claim to a reported observation about the LLM-jury metric. The ambiguity about self-scoring in the jury protocol should also be resolved, as it bears on the reported model rankings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: QUENCH is a real resource—manually transcribed quiz questions with masked entities and rationales, a careful contamination check, and a clean Indic/non-Indic split. That part is solid and worth building on. The headline claim—that all seven LLMs do much worse on Indic context—is plausible but not yet proven, because the effect is measured almost entirely by an LLM jury that includes one of the evaluated models and is never validated against human judgments.\n\nWhat's genuinely new: the dataset itself. Prior work has cultural benchmarks and open-domain QA, but I don't know of a generation-based, clue-driven quiz set with paired rationales and an Indic/non-Indic split. The annotation is careful, the contamination check is thorough (zero exact matches in C4, Dolma, etc.), and the error analysis is informative. The observation that CoT has little effect on this task is also interesting.\n\nWhere it gets soft: the 12–32 point gap between Indic and non-Indic on GEval is the paper's central evidence, but BERTScore differences are only 0.6–2.4 points, and BLEU/ROUGE are mixed. The jury includes GPT-4-Turbo, Mixtral, and Llama-3-70B—three of the seven evaluated models—with no agreement analysis against humans. The evaluation prompt tells the judge the true answer and asks for binary judgment, so any judge tendency to reject valid Indic aliases or transliterations would inflate the gap. That is a real possibility, given the judges are themselves LLMs trained on largely Western-centric data. Without a human validation sample, the size of the gap is uncertain. The human benchmarking section is too small and internally inconsistent to fill that role: it says 20 questions with equal numbers from both subsets, then reports 10 Non-Indic and 13 Indic. That needs correction.\n\nThe authors are honest about limitations, and the paper is clearly written. The central trend is consistent across models, which suggests something real is happening, but the magnitude is not credible as reported. A referee should ask for (1) a human-validated jury sample, (2) error bars or significance tests, and (3) an explanation of the 20 vs 23 discrepancy.\n\nWho this is for: anyone working on cultural bias in LLMs or on open-domain reasoning benchmarks. The dataset is likely to be reused even if the evaluation claim is revised. I'd send it to review, with the expectation of heavy revision.","headline":"A genuinely new dataset and a plausible but under-validated headline gap; the Indic/non-Indic result rests almost entirely on an LLM jury that is never checked against humans.","tokens_in":22346,"tokens_out":2092,"would_cite":true,"duration_ms":19721,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces QUENCH, a 400-question English quiz benchmark, and reports that all seven LLMs it tests score 12 to 32 points lower on Indic-context questions than on non-Indic ones, as judged by an LLM jury.","keywords":["QUENCH benchmark","LLM evaluation","Indic knowledge gap","quiz-based reasoning","entity prediction","rationale generation","chain-of-thought prompting","LLM-as-judge"],"falsifier":"Have a small panel of human quiz experts, blind to the Indic label, score a random sample of roughly 30 Indic and 30 matched non-Indic QUENCH answers for entity correctness and rationale quality; if the human-measured gap is close to zero while the LLM jury still reports 12 to 32 points, the paper's central claim of a large knowledge gap would not survive.","tokens_in":21322,"feed_emoji":"🧠","tokens_out":9190,"duration_ms":77532,"temperature":0.7,"pith_summary":"QUENCH is a benchmark of 400 English quiz questions, transcribed from public quizzing videos, in which answer entities are masked as X, Y, and Z and each question carries a gold rationale. The paper's central claim is that all seven LLMs it evaluates do consistently worse on the 70 Indic-context questions than on the 330 non-Indic questions, for both predicting the masked entity and generating the justifying rationale. On the paper's LLM-jury metric the gap runs from about 12 points for GPT-4-Turbo to about 32 points for Gemini 1.5 Flash. Because the questions are all in English and differ only in cultural context, the paper reads the gap as evidence of a real shortfall in non-Western knowledge and deduction, and presents QUENCH as a reusable instrument for measuring that gap.","feed_headline":"LLMs score 12–32 points worse on Indian-context quiz questions","feed_subtitle":"New benchmark: all seven LLMs tested do worse on Indic questions, by 12 to 32 points on jury scores.","key_machinery":"The central object is the QUENCH dataset: 400 English quiz questions with entities masked as X, Y, or Z and manually written free-text rationales, each tagged Indic or non-Indic based on whether the answer hinges on Indian geography, history, or culture. The evaluation machinery is a two-stage zero-shot protocol: the model first predicts the masked entities under a strict answer format, then generates a rationale from either its own predicted entities or the gold entities, and the whole pipeline is run both with and without chain-of-thought prompting. Scores come from BLEU, ROUGE-L, BERTScore, and an LLM-as-judge jury made of three models that each score every other model's outputs; the jury's binary entity verdict and 5-point rationale score are what produce the paper's headline Indic versus non-Indic gap.","core_discovery":"The paper establishes QUENCH as a valid open-domain, zero-shot quiz benchmark and reports the stable finding that every benchmarked LLM performs worse on Indic-context questions than non-Indic ones: an average gap of about 21 points for entity prediction and 14.7 points for rationale generation as judged by a three-model LLM jury. GPT-4-Turbo shows the smallest gap, around 12 points, while Gemini 1.5 Flash shows the largest at 32 points. The authors attribute the gap to pretraining corpora with a predominantly North American context, and they verify that the benchmark questions themselves are not present in major pretraining corpora, arguing that the gap is not a memorization artifact.","pith_inferences":["Because the gap is computed from jury scores while BERTScore differences are only 0.6 to 2.4 points, a human-scored validation study on the Indic subset would reveal whether the true knowledge gap is as large as the jury reports or partly an artifact of judge preferences.","The same masked-entity-plus-rationale format could be applied to quiz content from other non-Western regions to test whether the gap is India-specific or a general property of LLM pretraining distributions.","A natural next experiment the paper does not run is to give the models retrieval access or web search and measure whether the Indic gap narrows, which would separate memory failures from reasoning failures.","The finding that chain-of-thought does not help on QUENCH suggests the benchmark could serve as a stress test for prompt-engineering claims that are otherwise validated on Western-centric tasks."],"forward_implications":["Model rankings shift when cultural context is part of the test: GPT-4-Turbo leads on every metric, but the open-weight Meta-Llama-3-70B matches GPT-3.5-Turbo overall and shows a smaller performance swing across subsets.","English-only benchmarks that omit non-Western context overstate the general world-knowledge and deduction abilities of LLMs.","Chain-of-thought prompting does not reliably improve QUENCH scores, indicating the bottleneck is entity recall and cultural knowledge rather than the reasoning format.","Supplying gold entities instead of predicted ones raises rationale quality sharply (up to roughly 32 points for one model), showing LLMs can justify a known answer much better than they can retrieve it on their own."],"supporting_citations":[{"why":"Supplies the G-Eval LLM-as-judge method that underlies the jury-based GEval scores, the metric on which the headline Indic gap is measured.","marker":"Liu et al., 2023"},{"why":"Provides the jury-of-LLMs design used to reduce self-preference bias when scoring every benchmarked model's answers and rationales.","marker":"Verga et al., 2024"},{"why":"Establishes that LLM evaluators favor their own generations, motivating why the authors use multiple judges rather than a single model.","marker":"Panickssery et al., 2024"},{"why":"Defines BLEU, one of the standard n-gram metrics used alongside BERTScore and ROUGE-L.","marker":"Papineni et al., 2002"},{"why":"Defines BERTScore, the semantic-similarity metric that shows only small Indic/non-Indic differences and provides the contrast to the jury scores.","marker":"Zhang* et al., 2020"},{"why":"Supplies the North-American-context hypothesis the paper uses to explain why models underperform on Indic questions.","marker":"Zhou et al., 2022"},{"why":"Provides the contamination-check method used to argue that QUENCH questions are absent from major pretraining corpora, ruling out memorization as the explanation for performance.","marker":"Elazar et al., 2024"}],"fun_headline_variants":["QUENCH: Every LLM stumbles on India-specific quiz items","Indic context costs LLMs 12–32 points on new quiz","LLM bias: Indian quiz questions trip up all models","All 7 LLMs lag 12–32 points on Indian quiz"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline gap assumes the LLM jury grades Indic-context answers fairly, yet no human validation of the jury is reported and one of the jury members is itself an evaluated model, so a systematic jury bias against Indian-context answers would shrink the measured gap toward the near-zero BERTScore differences of 0.6 to 2.4 points.","fun_headline_variants_meta":{"raw":{"variants":["QUENCH: Every LLM stumbles on India-specific quiz items","Indic context costs LLMs 12–32 points on new quiz","LLM bias: Indian quiz questions trip up all models","All 7 LLMs lag 12–32 points on Indian quiz"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001306,"raw_usage":{"total_tokens":5265,"prompt_tokens":824,"completion_tokens":4441,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":4365}},"tokens_in":440,"tokens_out":4441,"duration_ms":26958,"temperature":1.0,"reasoning_tokens":4365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:36:33.690996+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a small panel of human quiz experts, blind to the Indic label, score a random sample of roughly 30 Indic and 30 matched non-Indic QUENCH answers for entity correctness and rationale quality; if the human-measured gap is close to zero while the LLM jury still reports 12 to 32 points, the paper's central claim of a large knowledge gap would not survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines BLEU, one of the standard n-gram metrics used alongside BERTScore and ROUGE-L."}],"review_version":1}