{"id":"fe0fe541-e332-4bdd-884d-c6d664bb914f","arxiv_id":"2504.19759","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Low-resource languages, not high-resource ones, drive multilingual moral reasoning performance when a model is fine-tuned, and models give inconsistent moral answers across languages.","lead":"A new benchmark, MMRB, tests moral reasoning in English, Chinese, Russian, Vietnamese, and Indonesian, and shows models answer inconsistently across languages. Fine-tuning experiments suggest that low-resource languages such as Vietnamese influence multilingual behavior more strongly than English does.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-lingual 'low-resource impact' in Tables 3 and 4 is an overgeneralization: low-resource training dominates only on low-resource evaluation targets, not on high-resource targets.","rationale":"Good-faith reading: the paper's contribution 3 and abstract assert a general cross-lingual asymmetry. The load-bearing condition is that low-resource training data produce larger changes across all evaluation languages, not only within low-resource languages. The printed Table 3 contradicts this condition when evaluation targets are EN/ZH/RU: on each such target, a high-resource training language yields the largest improvement (EN target: EN +5.9 vs ID +3.2/VI +1.8; ZH target: ZH +8.8 vs ID +4.0/VI +3.2; RU target: RU +9.4 vs ID +3.7/VI +5.0). Low-resource training only dominates on VI/ID targets. Thus the central claim is an overgeneralization of a real but narrower target-language-specific effect. This differs from the Reader's translation-quality concern, but it is more immediately checkable and directly impeaches the stated result. The test is a re-aggregation of the published deltas stratified by target language; if the pattern holds, the paper should be revised to claim that low-resource languages strongly affect low-resource evaluation performance, not multilingual reasoning generally. This reinforces the CONDITIONAL verdict: the benchmark and raw observations are still useful, but the headline generalization requires either a new analysis or a substantially weaker formulation.","tokens_in":7125,"tokens_out":13001,"duration_ms":121572,"concrete_test":"Re-aggregate Tables 3 and 4 by target-language stratum: for each training language, compute the mean cross-lingual delta on high-resource targets (EN/ZH/RU) and on low-resource targets (VI/ID), excluding the same-language cell. If low-resource training columns do not exceed high-resource training columns on the high-resource-target average, the headline claim should be narrowed to low-resource target effects; this check can be done from the printed values and should be accompanied by item-level bootstrap confidence intervals if reported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (abstract; §3.2.2) is that low-resource languages (VI, ID) have a stronger impact on multilingual moral reasoning than high-resource languages. The support in Table 3, however, is confined to low-resource evaluation targets. On the high-resource targets in ETHICSBASE, a high-resource training language produces the largest improvement every time: EN target deltas are +5.9 (EN), +3.6 (ZH), +2.9 (RU), +1.8 (VI), +3.2 (ID); ZH target deltas are +1.8/+8.8/+2.1/+3.2/+4.0; RU target deltas are +4.2/+6.8/+9.4/+5.0/+3.7. Low-resource training columns dominate only when the target is VI or ID (e.g., ETHICSBASE-VI: +19.9 for VI training, +27.1 for ID training; ETHICSBASE-ID: +16.2 for VI training, +21.2 for ID training). The same concentration appears in ETHICSPRO. Thus the data support a much narrower statement—fine-tuning a low-resource language strongly affects low-resource evaluation languages—not the general claim that low-resource languages are the critical driver of multilingual moral reasoning. Because that general claim is the paper's headline contribution, the conclusion is an overgeneralization from target-language-specific effects.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MMRB, a multilingual moral reasoning benchmark with 2,170 scenarios in English, Chinese, Russian, Vietnamese, and Indonesian across three context complexity levels (sentence, paragraph, document). The authors evaluate five LLMs on MMRB, report cross-lingual inconsistencies and a degradation with context complexity, and then fine-tune LLaMA-3-8B on monolingual alignment and poisoning data. The headline claim is that low-resource languages, particularly Vietnamese and Indonesian, have a disproportionately strong positive and negative impact on multilingual moral reasoning, challenging the assumption that high-resource languages are always the most effective sources of cross-lingual transfer.","tokens_in":7382,"tokens_out":3400,"duration_ms":35578,"significance":"The dataset and the fine-tuning experiments are potentially useful contributions to multilingual evaluation of moral reasoning, and the finding that low-resource-language fine-tuning can produce large effects on low-resource evaluation targets is interesting and actionable. However, the central claim as stated in the abstract and Section 3.2.2 is broader than what the data support: the alignment results show low-resource training languages dominating only on low-resource evaluation targets, not on high-resource targets. If the claim is narrowed to the observed target-language-specific pattern, the contribution is still worthwhile but substantially more modest. The paper also ships no code or data in the reviewed version, and the experiments lack repeated runs or confidence intervals, limiting the strength of the comparative conclusions.","major_comments":[{"comment":"The fine-tuning experiments in Tables 3 and 4 report single point estimates with no confidence intervals, no repeated runs, and no seed variation. Differences of a few accuracy points, which are used to rank languages by their cross-lingual impact, may be within run-to-run noise. At minimum, the authors should report multiple seeds with standard deviations or a statistical test over runs. This is load-bearing because the central comparative claim depends on the ordering of these deltas.","section":"Abstract and Section 3.2.2, Tables 3 and 4"},{"comment":"The interpretation of the Mann-Whitney U tests is inverted. The text states that 'most comparisons across languages fail to reach significance, indicating inconsistency in multilingual moral reasoning.' Failure to reject the null hypothesis means there is no detectable difference, not that the languages are inconsistent. The authors should rephrase this observation or report the direction and effect sizes of the tests that do reach significance.","section":"Section 3.2.1"},{"comment":"The same translated MMRB corpus is used both for fine-tuning and for evaluation, and no external benchmark is used to validate the cross-lingual transfer findings. Because the translation pipeline (DeepSeek plus manual verification plus Google Translate cross-validation) is the only bridge between languages, the observed disproportionate effects of Vietnamese and Indonesian could reflect translation artifacts or dataset-specific properties rather than a general property of low-resource languages. The Limitations section acknowledges translation inaccuracies but provides no analysis of translation quality, item difficulty equivalence, or consistency across languages. Adding a small external validation set or a per-language difficulty analysis would substantially strengthen the claim.","section":"Section 3.1 and Limitations"}],"minor_comments":[{"comment":"The ETHICSPRO rows in Table 4 are malformed: values such as '62.90↓9.2268.30↓3.8268.04↓4.0866.63↓5.4971.18↓0.94' run together without separators, making the table unreadable.","section":"Table 4"},{"comment":"The reference list contains two very similar entries for Wei et al. (2023) and Wei et al. (2022) for chain-of-thought prompting; these should be consolidated or disambiguated clearly.","section":"References"},{"comment":"The label 'ETHICSBASE-A VG' in Table 1 and the spacing in Figure 1 ('ETHICS BASE', 'ETHICS PRO', 'ETHICS MAX') are inconsistent and should be normalized.","section":"Table 1 and Figure 1"},{"comment":"The text refers to 'ETHICS BASE' with a space, while elsewhere it is 'ETHICSBASE'; please unify the naming.","section":"Section 3.2.1"},{"comment":"The win-rate diagram in Figure 2 is difficult to parse because the language labels are densely packed and the ordering is not explained; a clearer layout or a table of pairwise win rates would help.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe benchmark MMRB is a solid, usable contribution, and the question of how low-resource languages influence multilingual moral reasoning is genuinely worth asking. But the strong version of the claim — that low-resource languages have a disproportionately strong impact on multilingual moral reasoning — doesn't survive a close look at their own Tables 3 and 4. The low-resource advantage is mostly confined to low-resource evaluation targets. On high-resource targets (EN, ZH, RU), the largest alignment gains come from training in that same high-resource language, not from VI or ID. So what the data actually show is that fine-tuning on VI or ID strongly moves VI and ID evaluation, and less so the others. That's much narrower and less surprising than the abstract's claim.\n\nOther problems: the Mann-Whitney U interpretation in Section 3.2.1 is inverted. They say most cross-language comparisons fail to reach significance, which indicates inconsistency. Non-significant differences indicate no detectable difference, not inconsistency. Either they misapplied the test or misread it. Also, no confidence intervals or repeated runs anywhere; Tables 1-4 report single accuracies. Fine-tuning details are thin — no hyperparameters, data splits, or step counts. Table 4 has formatting typos that make it hard to parse. Translation quality is asserted but not quantified; the Limitations section admits potential translation inaccuracies, which is fine, but it is a core premise for cross-language comparisons, so some inter-annotator agreement or back-translation validation is needed.\n\nNone of this sinks the benchmark itself. It reuses existing English datasets translated into five languages, so novelty is incremental, but the resource is still usable. The evaluation of five LLMs across three complexity levels is useful, and the observation that performance degrades on paragraph-level but recovers on document-level when explicit moral principles are present is worth reporting.\n\nWho's this for? People working on multilingual evaluation and safety might want the benchmark. The central conclusion, though, should be treated as unresolved. I'd recommend the editor send it out for serious review, because the benchmark and the critique of the claim are both useful. The paper needs major revision before acceptance: redo the significance analysis, add error bars, report fine-tuning details, and reframe the conclusion to match the data.","headline":"A usable multilingual moral benchmark, but the headline low-resource finding is an overgeneralization of target-specific effects.","tokens_in":7912,"tokens_out":2426,"would_cite":false,"duration_ms":23202,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Low-resource languages are the strongest levers on an LLM's multilingual moral reasoning, for better and for worse.","keywords":["multilingual moral reasoning","low-resource languages","cross-lingual transfer","alignment fine-tuning","data poisoning","moral reasoning benchmark","LLM evaluation"],"falsifier":"Build an independent second translation of MMRB with item-level difficulty matched across languages, then rerun the monolingual alignment and poisoning experiments; if Indonesian and Vietnamese no longer dominate the cross-lingual deltas, the resource-level claim is an artifact of translation difficulty.","tokens_in":6923,"feed_emoji":"⚖️","tokens_out":6339,"duration_ms":57179,"temperature":0.7,"pith_summary":"This paper builds a five-language benchmark of moral scenarios at three levels of context and asks whether LLMs reason about morality consistently across languages, and which languages most affect cross-lingual behaviour. It finds that models are inconsistent across languages and that accuracy drops as context grows longer, with Vietnamese, the lowest-resource language in the set, hit hardest. The central result comes from fine-tuning a single open model: clean fine-tuning on Indonesian and Vietnamese improves other languages more than clean fine-tuning on English, while corrupted fine-tuning on Vietnamese damages other languages most. The paper argues that low-resource languages are high-leverage precisely because they are underrepresented in pretraining corpora, so fine-tuning has more room to change the model's moral associations in either direction.","feed_headline":"Low-resource languages shift LLM moral reasoning most","feed_subtitle":"Underrepresented languages are the strongest levers, for better or worse, in multilingual moral alignment.","key_machinery":"The load-bearing instrument is MMRB, a benchmark of 2,170 moral scenarios built from a sentence-level binary set, a paragraph-level reasoning set, and a document-level dilemma set with Virtue, Deontological, and Consequentialist answer branches, all translated into English, Chinese, Russian, Vietnamese, and Indonesian. The experimental machinery is monolingual fine-tuning: the same base model is trained on one language at a time, with correct labels for alignment and flipped labels for poisoning, and the cross-lingual accuracy deltas in Tables 3 and 4 provide the evidence for the resource-level claim.","core_discovery":"The paper claims that a language's resource level, not its typological distance or the model's familiarity, predicts how much that language's data moves multilingual moral reasoning. Fine-tuning LLaMA-3-8B on correctly labeled Indonesian data raises accuracy on every other language in MMRB, with Vietnamese clean data the next strongest; flipping labels in Vietnamese causes the sharpest cross-lingual drops, while English poisoning barely moves the scores. These cross-lingual transfer effects support a mechanism in which underrepresented languages have more headroom for fine-tuning to fill knowledge gaps, so their data quality matters disproportionately.","pith_inferences":["If the proposed mechanism is pretraining underrepresentation, the same pattern should appear for other safety-sensitive abilities, such as refusing harmful requests or avoiding social bias, when fine-tuned on an equally low-resource language.","The resource-level effect predicts even larger transfer from languages with less pretraining representation than Vietnamese; adding Swahili or a similar language to the same benchmark would test that gradient.","Since the paper reports that document-level scenarios restore accuracy with explicit moral principles, an extension could test whether explicit structure also suppresses poisoning transfer, which would clarify the interaction between context length and language effects."],"forward_implications":["Multilingual safety evaluation should treat low-resource languages as the most sensitive probes, since they show the largest alignment gains and the largest poisoning losses.","A relatively small amount of clean low-resource data can substitute for much larger English datasets in cross-lingual alignment.","A corrupted low-resource dataset can silently degrade moral behaviour in higher-resource languages, so curation effort should shift toward under-represented languages.","English-only benchmark scores overstate a model's cross-lingual moral consistency."],"supporting_citations":[{"why":"Supplies ETHICSBASE, the sentence-level moral scenarios that form the base layer of MMRB.","marker":"Hendrycks et al., 2020"},{"why":"Supplies ETHICSPRO, the paragraph-level moral reasoning data used for the middle complexity tier.","marker":"Ma et al., 2023"},{"why":"Provides high- and low-ambiguity moral scenarios used in curating the document-level ETHICSMAX tier.","marker":"Scherrer et al., 2024"},{"why":"Provides the moral dilemma framework with Virtue, Deontological, and Consequentialist branches that defines ETHICSMAX structure.","marker":"Rao et al., 2023b"},{"why":"Documents LLaMA-3's architecture and per-language training data distribution, motivating the low-resource hypothesis and supplying the base model for fine-tuning.","marker":"AI@Meta, 2024"},{"why":"Supplies the fine-tuning toolchain used for the monolingual alignment and poisoning experiments.","marker":"Zheng et al., 2024"},{"why":"Earlier multilingual moral-judgment study showing weaker performance in lower-resource languages; this paper's result extends it to cross-lingual transfer.","marker":"Khandelwal et al., 2024"}],"fun_headline_variants":["Low-resource languages sway LLM morality most","For LLM ethics, low-resource languages matter most","English barely moves LLM ethics; Vietnamese does","Language resource level predicts moral transfer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The five language versions of the translated scenarios are equally valid and equally difficult, so the larger transfer effects observed for Vietnamese and Indonesian reflect language resource levels rather than translation artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Low-resource languages sway LLM morality most","For LLM ethics, low-resource languages matter most","English barely moves LLM ethics; Vietnamese does","Language resource level predicts moral transfer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000151,"raw_usage":{"total_tokens":1108,"prompt_tokens":760,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":376,"completion_tokens_details":{"reasoning_tokens":290}},"tokens_in":376,"tokens_out":348,"duration_ms":3752,"temperature":1.0,"reasoning_tokens":290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:43:26.994742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build an independent second translation of MMRB with item-level difficulty matched across languages, then rerun the monolingual alignment and poisoning experiments; if Indonesian and Vietnamese no longer dominate the cross-lingual deltas, the resource-level claim is an artifact of translation difficulty.","supporting_citations":[],"review_version":1}