REVIEW 3 major objections 5 minor 1 cited by
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Low-resource languages are the strongest levers on an LLM's multilingual moral reasoning, for better and for worse.
desk verdict A usable multilingual moral benchmark, but the headline low-resource finding is an overgeneralization of target-specific effects. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is MMRB, a benchmark of 2,170 moral scenarios built from a sentence-level binary set, a paragraph-level reasoning set, and a document-level dilemma set with Virtue, Deontological, and Consequentialist answer branches, all translated into English, Chinese, Russian, Vietnamese, and Indonesian. The experimental machinery is monolingual fine-tuning: the same base model is trained on one language at a time, with correct labels for alignment and flipped labels for poisoning, and the cross-lingual accuracy deltas in Tables 3 and 4 provide the evidence for the resource-level claim.
What would settle it
Build an independent second translation of MMRB with item-level difficulty matched across languages, then rerun the monolingual alignment and poisoning experiments; if Indonesian and Vietnamese no longer dominate the cross-lingual deltas, the resource-level claim is an artifact of translation difficulty.
Extended reading notes
Core claim
The paper claims that a language's resource level, not its typological distance or the model's familiarity, predicts how much that language's data moves multilingual moral reasoning. Fine-tuning LLaMA-3-8B on correctly labeled Indonesian data raises accuracy on every other language in MMRB, with Vietnamese clean data the next strongest; flipping labels in Vietnamese causes the sharpest cross-lingual drops, while English poisoning barely moves the scores. These cross-lingual transfer effects support a mechanism in which underrepresented languages have more headroom for fine-tuning to fill knowledge gaps, so their data quality matters disproportionately.
Load-bearing premise
The five language versions of the translated scenarios are equally valid and equally difficult, so the larger transfer effects observed for Vietnamese and Indonesian reflect language resource levels rather than translation artifacts.
Editorial extensions
If this is right
- Multilingual safety evaluation should treat low-resource languages as the most sensitive probes, since they show the largest alignment gains and the largest poisoning losses.
- A relatively small amount of clean low-resource data can substitute for much larger English datasets in cross-lingual alignment.
- A corrupted low-resource dataset can silently degrade moral behaviour in higher-resource languages, so curation effort should shift toward under-represented languages.
- English-only benchmark scores overstate a model's cross-lingual moral consistency.
Reading between the lines
- If the proposed mechanism is pretraining underrepresentation, the same pattern should appear for other safety-sensitive abilities, such as refusing harmful requests or avoiding social bias, when fine-tuned on an equally low-resource language.
- The resource-level effect predicts even larger transfer from languages with less pretraining representation than Vietnamese; adding Swahili or a similar language to the same benchmark would test that gradient.
- Since the paper reports that document-level scenarios restore accuracy with explicit moral principles, an extension could test whether explicit structure also suppresses poisoning transfer, which would clarify the interaction between context length and language effects.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MMRB, a multilingual moral reasoning benchmark with 2,170 scenarios in English, Chinese, Russian, Vietnamese, and Indonesian across three context complexity levels (sentence, paragraph, document). The authors evaluate five LLMs on MMRB, report cross-lingual inconsistencies and a degradation with context complexity, and then fine-tune LLaMA-3-8B on monolingual alignment and poisoning data. The headline claim is that low-resource languages, particularly Vietnamese and Indonesian, have a disproportionately strong positive and negative impact on multilingual moral reasoning, challenging the assumption that high-resource languages are always the most effective sources of cross-lingual transfer.
Significance. The dataset and the fine-tuning experiments are potentially useful contributions to multilingual evaluation of moral reasoning, and the finding that low-resource-language fine-tuning can produce large effects on low-resource evaluation targets is interesting and actionable. However, the central claim as stated in the abstract and Section 3.2.2 is broader than what the data support: the alignment results show low-resource training languages dominating only on low-resource evaluation targets, not on high-resource targets. If the claim is narrowed to the observed target-language-specific pattern, the contribution is still worthwhile but substantially more modest. The paper also ships no code or data in the reviewed version, and the experiments lack repeated runs or confidence intervals, limiting the strength of the comparative conclusions.
major comments (3)
- [Abstract and Section 3.2.2, Tables 3 and 4] The fine-tuning experiments in Tables 3 and 4 report single point estimates with no confidence intervals, no repeated runs, and no seed variation. Differences of a few accuracy points, which are used to rank languages by their cross-lingual impact, may be within run-to-run noise. At minimum, the authors should report multiple seeds with standard deviations or a statistical test over runs. This is load-bearing because the central comparative claim depends on the ordering of these deltas.
- [Section 3.2.1] The interpretation of the Mann-Whitney U tests is inverted. The text states that 'most comparisons across languages fail to reach significance, indicating inconsistency in multilingual moral reasoning.' Failure to reject the null hypothesis means there is no detectable difference, not that the languages are inconsistent. The authors should rephrase this observation or report the direction and effect sizes of the tests that do reach significance.
- [Section 3.1 and Limitations] The same translated MMRB corpus is used both for fine-tuning and for evaluation, and no external benchmark is used to validate the cross-lingual transfer findings. Because the translation pipeline (DeepSeek plus manual verification plus Google Translate cross-validation) is the only bridge between languages, the observed disproportionate effects of Vietnamese and Indonesian could reflect translation artifacts or dataset-specific properties rather than a general property of low-resource languages. The Limitations section acknowledges translation inaccuracies but provides no analysis of translation quality, item difficulty equivalence, or consistency across languages. Adding a small external validation set or a per-language difficulty analysis would substantially strengthen the claim.
minor comments (5)
- [Table 4] The ETHICSPRO rows in Table 4 are malformed: values such as '62.90↓9.2268.30↓3.8268.04↓4.0866.63↓5.4971.18↓0.94' run together without separators, making the table unreadable.
- [References] The reference list contains two very similar entries for Wei et al. (2023) and Wei et al. (2022) for chain-of-thought prompting; these should be consolidated or disambiguated clearly.
- [Table 1 and Figure 1] The label 'ETHICSBASE-A VG' in Table 1 and the spacing in Figure 1 ('ETHICS BASE', 'ETHICS PRO', 'ETHICS MAX') are inconsistent and should be normalized.
- [Section 3.2.1] The text refers to 'ETHICS BASE' with a space, while elsewhere it is 'ETHICSBASE'; please unify the naming.
- [Figure 2] The win-rate diagram in Figure 2 is difficult to parse because the language labels are densely packed and the ordering is not explained; a clearer layout or a table of pairwise win rates would help.
Circularity Check
No significant circularity: the paper's claims are empirical readings of benchmark and fine-tuning results, not derived from fitted inputs, self-citation chains, or definitional equivalences.
full rationale
This is an empirical evaluation and fine-tuning study, not a derivation chain with equations or fitted parameters. The central claim that low-resource languages have a stronger positive and negative impact on multilingual moral reasoning is a direct reading of the delta tables (Tables 3 and 4), and the paper does not define 'low-resource impact' in terms of the quantities it then reports as findings. No parameter is fitted to a subset of data and then 'predicted' on a closely related quantity. No uniqueness theorem or load-bearing self-citation appears: the references include prior work on moral reasoning, but none of it is by the present authors, and none is invoked to forbid alternative explanations. The use of the same translated benchmark for fine-tuning and evaluation is a possible validity or leakage concern, and the skeptic's observation that the low-resource advantage is concentrated on low-resource evaluation targets is an overgeneralization critique, not a circularity critique. The limitations section explicitly acknowledges translation inaccuracies and limited generalizability, which further confirms that the authors do not claim to derive the result from prior assumptions. Thus, under the stated rules, there is no circular step to report and the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Accuracy on MMRB is a valid measure of moral reasoning ability.
- domain assumption Machine translation plus manual verification preserves scenario difficulty and moral content across languages.
- domain assumption Fine-tuning effects generalize beyond LLaMA-3-8B.
Cite this review
Pith. "Pith review of Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs." pith.science (2026). https://pith.science/paper/D736RBCN
@misc{pith2026250419759,
author = {Pith},
title = {Pith review of: Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/D736RBCN}},
note = {Machine review of arXiv:2504.19759}
}
read the original abstract
In this paper, we introduce the Multilingual Moral Reasoning Benchmark (MMRB) to evaluate the moral reasoning abilities of large language models (LLMs) across five typologically diverse languages and three levels of contextual complexity: sentence, paragraph, and document. Our results show moral reasoning performance degrades with increasing context complexity, particularly for low-resource languages such as Vietnamese. We further fine-tune the open-source LLaMA-3-8B model using curated monolingual data for alignment and poisoning. Surprisingly, low-resource languages have a stronger impact on multilingual reasoning than high-resource ones, highlighting their critical role in multilingual NLP.
Figures
Forward citations
Cited by 1 Pith paper
-
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
MET-D self-distills theory-selected moral grounds into native-language reasoning, lifting macro-F1 by ~3.7–4.2 points on MCLASH and MMoralExceptQA while raising native-language chains by ~62 points.
Reference graph
Works this paper leans on
-
[1]
AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card
2024
-
[2]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long a...
-
[3]
Hwang, Maxwell Forbes, and Yejin Choi
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.54 Moral stories: Situated reasoning about norms, intents, actions, and their consequences . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 698--718, Online and Punta Cana, Dominica...
-
[4]
Katharina Haemmerl, Bjoern Deiseroth, Patrick Schramowski, Jind r ich Libovick \'y , Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023. https://doi.org/10.18653/v1/2023.findings-acl.134 Speaking multiple languages affects the moral bias of language models . In Findings of the Association for Computational Linguistics: ACL 2023, pages 2137...
-
[5]
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020. Aligning ai with shared human values. arXiv preprint arXiv:2008.02275
arXiv 2020
-
[6]
Aditi Khandelwal, Utkarsh Agarwal, Kumar Tanmay, and Monojit Choudhury. 2024. Do moral judgment and reasoning capability of llms change with language? a study using the multilingual defining issues test. arXiv preprint arXiv:2402.02135
arXiv 2024
-
[7]
Pingchuan Ma, Zongjie Li, Ao Sun, and Shuai Wang. 2023. https://arxiv.org/abs/2305.02626 "oops, did i just say that?" testing and repairing unethical suggestions of large language models with suggest-critique-reflect process . Preprint, arXiv:2305.02626
work page Pith review arXiv 2023
-
[8]
OpenAI. 2023. Gpt-4. https://openai.com/gpt-4
work page 2023
Show all 19 references
-
[9]
Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. 2023 a . Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in llms. arXiv preprint arXiv:2310.07251
2023 arXiv
-
[10]
Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. 2023 b . https://doi.org/10.18653/v1/2023.findings-emnlp.892 Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLM s . In Findings of the Associat...
2023 doi
-
[11]
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2024. Evaluating the moral beliefs encoded in llms. Advances in Neural Information Processing Systems, 36
2024
-
[12]
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022. On the machine learning of ethical judgments from natural language. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational L...
2022
-
[13]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. https://arxiv.org/abs/2201.11903 Chain-of-thought prompting elicits reasoning in large language models . Preprint, arXiv:2201.11903
2023 arXiv
-
[14]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[15]
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. Llamafactory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372
2024 arXiv
-
[16]
Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng. 2023. Rethinking machine ethics--can llms perform moral reasoning through the lens of moral theories? arXiv preprint arXiv:2308.15399
2023 arXiv
-
[17]
Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. The moral integrity corpus: A benchmark for ethical dialogue systems. arXiv preprint arXiv:2204.03021
2022 arXiv
-
[18]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[19]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.