{"id":"9336e0a4-ee59-44f5-a8a7-154f4afd8555","arxiv_id":"2411.09492","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"MM-Eval is a new Modern Mongolian benchmark showing LLMs perform best at syntax, worse at semantics and knowledge, and worst at reasoning.","lead":"The authors built MM-Eval, a four-part benchmark with about 1,840 questions for testing how large language models handle Modern Mongolian. In tests on five models, all did better on grammar than on meaning, and all scored below 30 percent on math reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The semantics labels are the load-bearing point: §3.5 never verifies distractors are definitively wrong, so the syntax>semantics gap may be a labeling artifact; §4.3 even contradicts Table 1 on the best semantic model.","rationale":"MM-Eval is a useful and mostly transparent benchmark contribution, and the overall performance hierarchy may well survive closer inspection. The most load-bearing condition for the paper's headline result is not the choice of models or prompts but the correctness of the semantic labels. The central observation, that all models do materially better on syntax than on semantics, is an aggregate over 677 items whose distractors, per §3.5, are selected only by part-of-speech matching and are not described as manually verified. In contrast, the knowledge and reasoning sections explicitly mention manual verification and correction. A cloze item built by removing a vocabulary word can easily contain distractors that are grammatical and contextually plausible; without independent human judgment, the 'definitively incorrect' criterion is unsecured. If even 15-20% of semantics items are bad, model ordering and the syntax/semantics gap could shift materially, and the paper's central claim would be an artifact.\n\nThe reported results also contain a concrete inconsistency: Table 1 gives ChatGPT-4-Turbo 72.53% on Semantics and Qwen2-7B-Instruct 54.8%, while §4.3 says Qwen2-7B-Instruct performs well in Semantics (72.53%). The reader's strongest claim correctly follows Table 1, yet the prose says the opposite. This does not change the qualitative syntax>semantics ordering, but it weakens confidence in the result-reporting process and makes a manual audit urgent.\n\nI agree with the reader that label quality is the weakest assumption, but only partially: the reader's phrasing says distractors were manually verified, whereas §3.5 does not claim that for semantics. For these items the problem is not merely lack of inter-annotator agreement but absence of any reported verification. The proposed check, native-speaker adjudication of a sample of semantics items followed by recomputation on valid items, would settle whether the central claim is real. The dataset remains a valuable resource, and conditional acceptance is still appropriate, with the additional validation and correction of the prose/table mismatch as conditions. I therefore keep the verdict unchanged.","tokens_in":6343,"tokens_out":5396,"duration_ms":51163,"concrete_test":"Sample 100 Semantics items uniformly from MM-Eval and have three independent native Mongolian speakers judge each item: does the intended answer complete the sentence correctly, and is each distractor semantically impossible? Compute per-judge valid-item rate and pairwise agreement. Then recompute the Semantics column of Table 1 on the subset of items that all three judges mark valid. If the valid-item rate is below ~90%, or if the syntax>semantics ordering changes on the valid subset, the central claim is not supported. As a secondary check, run a random-choice baseline (25.0%) and a most-frequent-answer baseline to confirm the semantics scores are above chance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that every tested model scores higher on syntax than on semantics, and that semantics exposes a gap in deeper language understanding. For that comparison to mean anything, each of the 677 Semantics items must have exactly one correct answer and three unambiguously incorrect distractors. Section 3.5 shows how items are generated: a word matched from the textbook vocabulary list is removed, and three same-part-of-speech distractors are drawn from the same list. No manual check or correction step is reported for these distractors, in contrast with §3.6 and §3.7, where manual proofreading is explicitly described. Same-POS distractors are not guaranteed to be 'definitively incorrect': in a cloze item, several nouns or adjectives can be grammatically and pragmatically compatible, and whether they are semantically impossible is exactly the subjective judgment that was never independently validated. If a substantial share of Semantics items admit more than one plausible answer, the reported scores, and the uniform syntax-over-semantics ordering, are not a reliable measurement of deeper language understanding. There is also an internal sign of unvalidated results: Table 1 lists ChatGPT-4-Turbo as best at Semantics with 72.53% and Qwen2-7B-Instruct at 54.8%, while §4.3 states 'Qwen2-7B-Instruct performs well in semantics (72.53%)', attributing the table's best value to the wrong model. This makes it hard to audit which numbers the authors actually relied on.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MM-Eval, a benchmark for evaluating LLMs on Modern Mongolian (Cyrillic script), organized into a Dual Capability Framework with four hierarchical levels: syntax, semantics, knowledge, and reasoning. The dataset comprises 569 syntax MCQs drawn from a Mongolian textbook with shuffled-word distractors, 677 semantics MCQs constructed by cloze deletion with same-part-of-speech distractors from textbook vocabulary, 344 knowledge MCQs derived from WebQSP and ChatGPT-generated content, and 250 reasoning math problems translated from MGSM. The authors evaluate five models (Qwen2-7B-Instruct, GLM4-9b-chat, Llama3.1-8B-Instruct, GPT-4-Turbo, and DeepseekV2.5) and report that all models score higher on syntax than on semantics, that knowledge performance shows a moderate decline, and that all models perform poorly on reasoning (best 29.6%). The dataset is publicly released.","tokens_in":6673,"tokens_out":3137,"duration_ms":28083,"significance":"If the benchmark's labels and claims are sound, MM-Eval would be a useful resource for a genuinely under-served language: it provides a structured, multi-level evaluation covering both language proficiency and cognitive transfer, and the finding that semantic understanding lags syntactic competence in low-resource settings is potentially important. The paper also makes a plausible conceptual contribution in separating language abilities from cognitive abilities. However, the empirical value of the contribution currently rests on unverified assumptions about label quality and on an experimental design that involves the tested model in the construction of part of the data. The release of the dataset and the explicit descriptions of the construction pipeline are positive features, but they do not by themselves establish the reliability of the reported accuracies.","major_comments":[{"comment":"The semantics labels are the load-bearing component of the paper's central claim that syntax accuracy exceeds semantics accuracy for all models. The construction procedure selects same-part-of-speech distractors from the vocabulary list, but the manuscript reports no manual verification or correction step for these distractors, in contrast with §3.6 and §3.7 where manual proofreading is explicitly described. A cloze item with same-part-of-speech distractors can admit multiple plausible completions: a noun or adjective that is grammatical may still be semantically compatible with the sentence context, and whether it is 'definitively incorrect' is precisely the judgment that is never independently validated. If a substantial fraction of the 677 Semantics items have more than one correct answer, the reported semantics accuracies and the uniform syntax-over-semantics ordering are not a reliable measure of deeper language understanding. The authors must add a manual verification and correction step, report inter-annotator agreement, or otherwise demonstrate that each item has exactly one correct answer and three unambiguously incorrect distractors, and then re-release the dataset and recompute Table 1 if any items change.","section":"§3.5 (Semantics Eval)"},{"comment":"There is a direct inconsistency between Table 1 and §4.3 regarding the best-performing semantic model. Table 1 lists chatgpt4-turbo as achieving 72.53% on Semantics and qwen2-7b-instruct as 54.8%, while §4.3 states that 'Qwen2-7B-Instruct performs well in semantics (72.53%)', attributing the table's best value to the wrong model. Figure 2 is also described in a duplicated paragraph that swaps the roles of Table 1 and Figure 2. These inconsistencies make it impossible to know which numbers the authors actually stand behind and undermine confidence in the reported experimental results. The authors must correct the text, reconcile Table 1, Figure 2, and §4.3, and re-audit all reported numbers.","section":"§4.3 and Table 1"},{"comment":"The experimental comparison lacks a chance baseline and any measure of statistical uncertainty. All multiple-choice items have four options, so a random model would be expected to score 25% on Syntax, Semantics, and Knowledge by chance; several reported scores (e.g., Llama-3.1-8B on Semantics at 28.06%, and Qwen2 on Reasoning at 6%) are only slightly above or even below this floor, and without confidence intervals or significance tests the claim that 'all models performed better on syntactic tasks than semantic tasks' is not statistically supported. The authors should report binomial confidence intervals or standard errors for each accuracy, provide a random baseline, and ideally run multiple inference seeds (the current temperature=0 setting gives one deterministic outcome).","section":"§4.1 and Table 1"},{"comment":"The evaluation is partially circular for the closed-source models. The knowledge distractors in §3.6 are generated with the ChatGPT API, and the Mongolian translations of reasoning problems in §3.7 are produced and verified by the ChatGPT API; GPT-4-Turbo is then one of the evaluated models in §4.2. This means the tested model directly contributed content to the benchmark on which its own knowledge and reasoning scores are computed, so the reported 80.52% knowledge accuracy and 26.8% reasoning accuracy for GPT-4-Turbo are not independent measurements. The authors should either remove GPT-4-Turbo from the evaluated models, construct the knowledge and reasoning sections without using the evaluated model family, or demonstrate through a contamination analysis that the model's scores are not inflated by its own generated content. They should also report whether any contamination check was performed against model training corpora for the textbook- and translated-based items.","section":"§3.6, §3.7, and §4.2"}],"minor_comments":[{"comment":"The Discussion states that MM-Eval is 'limited by its single content source', but the benchmark actually uses three sources: the textbook, WebQSP, and MGSM, not to mention ChatGPT-generated knowledge items. This wording should be clarified to indicate that the language-ability section relies on a single textbook, while the cognitive-ability section uses multiple sources.","section":"§5 (Discussion)"},{"comment":"Model names are inconsistent: Table 1 uses lowercase 'deepseekv2.5' and 'chatgpt4-turbo', while §4.2 refers to 'GPT-4-Turbo-04-09' and 'DeepseekV2.5'. Please standardize model names, including for GLM4-9b-chat and Llama3.1-8B-Instruct.","section":"Table 1 and §4.2"},{"comment":"The two consecutive paragraphs beginning 'Figure 2 presents the corresponding results...' and 'Table 1 presents...' contain redundant and conflicting descriptions; one of the paragraphs should be removed and the remaining text should accurately describe which figure shows what.","section":"§4.3"},{"comment":"The paper cites a reference for GLM-130B but evaluates GLM4-9b-chat; a reference or model card for the actual evaluated GLM4 checkpoint should be provided.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a dataset/benchmark contribution rather than a methodological advance, and its publishability hinges on the trustworthiness of the dataset labels. The inconsistencies in the results section and the lack of statistical baselines suggest that the authors may have rushed the experimental reporting. Given that the dataset is released, it would be feasible to add the required verification, re-run the evaluation, and correct the narrative. I would not reject outright, but the revision needs to be substantive, not just editorial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is the resource, not the results. MM-Eval gives the first hierarchical benchmark for Modern Mongolian — 1,840 MCQ/math items across syntax, semantics, knowledge, and reasoning, with baselines for five LLMs. That is a real contribution for a language that barely appears in LLM evaluation. The Dual Capability Framework is a simple taxonomy, not deep theory, but it is a reasonable way to organize the evaluation. The choice to source syntax and semantics from a textbook and to translate WebQSP and MGSM is sensible, and the manual checks described for knowledge and reasoning are a point in the paper's favor.\n\nThe soft spots are mostly around the Semantics section, and one is load-bearing. Section 3.5 says distractor words were chosen to be 'plausible yet definitively incorrect,' but unlike sections 3.6 and 3.7 it reports no manual verification or quality check for those distractors. Same-part-of-speech distractors drawn from a vocabulary list can easily be grammatically and pragmatically compatible with the cloze sentence. If a non-trivial share of the 677 semantics items have more than one defensible answer, the paper's main empirical finding — every model scores higher on syntax than on semantics — becomes a labeling artifact rather than a measurement of deeper understanding. The paper needs an audit: at least a human agreement study or a clearly reported manual pass with numbers.\n\nThere is also a plain internal inconsistency that undermines confidence in the tables. Table 1 lists ChatGPT-4-Turbo as best on Semantics with 72.53% and Qwen2-7B at 54.8%, but Section 4.3 says 'Qwen2-7B-Instruct performs well in semantics (72.53%)'. That is not a subtle error; it attributes the table's best to the wrong model. The authors should fix it and double-check the rest of the reported numbers.\n\nMissing baselines and statistics are minor compared with the label problem, but still worth listing: no random baseline (25% for four-option items), no confidence intervals, no contamination analysis. The circularity (GPT-4 generated part of the data it is scored on) is real but partial, since manual verification is claimed for knowledge and reasoning. The transfer claim about knowledge relies on a comparison with high-resource languages that is not in the paper.\n\nBottom line: this paper is for anyone working on low-resource language evaluation, especially Mongolian NLP. The benchmark is a genuine resource, and the paper deserves peer review. I would condition acceptance on fixing the attribution error, validating the semantics labels with an independent manual pass, and adding simple baselines and confidence intervals.","headline":"A useful new benchmark for Mongolian LLM evaluation whose main finding rests on the least-verified section of the dataset; deserves review but needs a label audit.","tokens_in":7163,"tokens_out":2976,"would_cite":true,"duration_ms":24987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MM-Eval, a four-level benchmark for modern Mongolian in Cyrillic script, and reports that every LLM it tests handles syntax better than semantics, with reasoning scores far lower than either.","keywords":["MM-Eval","Modern Mongolian","LLM evaluation","low-resource language","syntax","semantics","knowledge","reasoning"],"falsifier":"Have professional Mongolian teachers independently re-annotate all 677 semantic items and a random sample of the syntax, knowledge, and reasoning sets. If agreement on the gold answers is low, or if many shuffled-order syntax options are judged acceptable in colloquial Mongolian, then the reported syntax-over-semantics gap and the reasoning ceiling of 29.6% would not survive.","tokens_in":6154,"feed_emoji":"🇲🇳","tokens_out":4797,"duration_ms":40513,"temperature":0.7,"pith_summary":"MM-Eval is a new benchmark for evaluating large language models on modern Mongolian, a low-resource language written in Cyrillic. The dataset is organized into four levels: 569 syntax questions, 677 semantics questions, 344 knowledge questions, and 250 reasoning problems. The paper tests five open- and closed-source models and finds a consistent pattern: all models score higher on syntax than on semantics, and all struggle with reasoning, where the best accuracy is 29.6%. The authors argue this reveals a gap between surface grammatical ability and deeper meaning understanding, while knowledge performance suggests that general knowledge transfers from high-resource languages to Mongolian.","feed_headline":"LLMs top 90% on Mongolian syntax but only 29% on reasoning","feed_subtitle":"A new four-level benchmark shows grammar beats meaning for five current LLMs in a low-resource language.","key_machinery":"The Dual Capability Framework organizes the benchmark into language abilities (syntax, semantics) and cognitive abilities (knowledge, reasoning). Syntax items are built by shuffling word order in textbook sentences, semantics items are fill-in-the-blank questions with same-part-of-speech distractors, knowledge items come from filtered WebQSP facts plus ChatGPT-generated and manually verified common-knowledge pairs, and reasoning items are MGSM math word problems translated into Mongolian. This framework lets the paper attribute low scores to either a missing language-specific ability or a missing general cognitive capacity.","core_discovery":"The central claim is that current LLMs display a clear capability hierarchy in modern Mongolian: syntactic competence exceeds semantic competence, and reasoning is the weakest area. GPT-4-Turbo reaches 90.69% on syntax and 80.52% on knowledge, Qwen2-7B-Instruct leads semantics at 54.8%, and DeepseekV2.5 achieves the highest reasoning score at 29.6%. The paper interprets this as evidence that models partially master Mongolian grammar but lack deeper language understanding and complex reasoning in this low-resource language.","pith_inferences":["If the syntax–semantics gap is genuine, a plausible follow-up is to test whether it narrows when models receive longer context or prompts in traditional Mongolian script, since the current design uses isolated sentences and may underestimate semantic ability.","The reasoning ceiling may partly reflect translation quality rather than pure reasoning capacity; running the same MGSM problems in English with the same models would separate translation failure from reasoning failure.","Because part of the knowledge section was generated with ChatGPT and then translated, models trained on similar outputs may score artificially high; a contamination-controlled version would be needed for stable conclusions."],"forward_implications":["Model rankings change across the four levels, so a single aggregate score for Mongolian would hide that GPT-4-Turbo leads syntax and knowledge while Qwen2-7B-Instruct leads semantics.","The near-ordering syntax > knowledge > semantics > reasoning across most tested models suggests that language ability and cognitive ability should be evaluated separately for low-resource languages.","Knowledge accuracy of 59–80% alongside weaker semantics implies that general knowledge transfers across languages better than language-specific semantic skill.","The reasoning floor of 5–29.6% identifies Mongolian math word problems as a clear target for future training data and model improvement.","The dataset provides a reusable test set for tracking whether future models improve on Mongolian without relying on machine-translation benchmarks."],"supporting_citations":[{"why":"Supplies the Modern Mongolian Language Textbook I sentences used to build the syntax and semantics questions.","marker":"(Hou et al., 2017)"},{"why":"Provides the WebQSP dataset filtered for country-related common-knowledge questions.","marker":"(Yih et al., 2016)"},{"why":"Provides the MGSM dataset that the reasoning section translates from English and Chinese into Mongolian.","marker":"(Shi et al., 2023)"},{"why":"Documents GPT-4-Turbo, the closed-source model that posts the best syntax and knowledge scores.","marker":"(OpenAI, 2023)"},{"why":"Documents DeepSeek, the model that achieves the highest reasoning accuracy in the evaluation.","marker":"(Dai et al., 2024)"},{"why":"Documents Qwen2, the open-source model that leads the semantic section.","marker":"(Yang et al., 2024)"},{"why":"Documents GLM, the open-source model family that includes GLM4-9b-chat evaluated in the experiments.","marker":"(Zeng et al., 2023)"}],"fun_headline_variants":["MM-Eval: LLMs hit 90% on Mongolian syntax, 29% on reasoning","Mongolian benchmark: Grammar beats meaning in LLMs","LLMs ace Mongolian grammar, fail at reasoning: MM-Eval","New benchmark: LLMs strong on syntax, weak on reasoning in Mongolian","MM-Eval shows LLMs parse Mongolian but can't reason"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's gold labels and distractor options are accurate: every multiple-choice and numeric answer was manually verified, but no inter-annotator agreement or independent quality metric is reported, so a substantial share of mislabeled items would change every reported accuracy.","fun_headline_variants_meta":{"raw":{"variants":["MM-Eval: LLMs hit 90% on Mongolian syntax, 29% on reasoning","Mongolian benchmark: Grammar beats meaning in LLMs","LLMs ace Mongolian grammar, fail at reasoning: MM-Eval","New benchmark: LLMs strong on syntax, weak on reasoning in Mongolian","MM-Eval shows LLMs parse Mongolian but can't reason"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1530,"prompt_tokens":876,"completion_tokens":654,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":559}},"tokens_in":492,"tokens_out":654,"duration_ms":5736,"temperature":1.0,"reasoning_tokens":559,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:34:54.634887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have professional Mongolian teachers independently re-annotate all 677 semantic items and a random sample of the syntax, knowledge, and reasoning sets. If agreement on the gold answers is low, or if many shuffled-order syntax options are judged acceptable in colloquial Mongolian, then the reported syntax-over-semantics gap and the reasoning ceiling of 29.6% would not survive.","supporting_citations":[],"review_version":1}