{"id":"2cabcd8d-c290-4229-8607-7d943b1bf260","arxiv_id":"2412.18863","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Multilingual LLMs produce different moral foundation scores across languages and models, and GPT models align more closely with human moral judgments than smaller open-source models.","lead":"Researchers asked four large language models to fill out a 36-item moral values questionnaire in eight languages and compared their answers with human responses. They found that the models' moral scores vary with language and model, and that some models align more closely with human answers than others.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The artificial word/negation bans in §3.3 are an untested control that likely distorts Likert responses, so the reported cross-language differences may be prompt artifacts rather than evidence of moral adaptation; a no-constraint control condition is needed before the central claim is accepted.","rationale":"The reader identified the same weakest assumption: the prompt-construction design may not preserve the intended meaning and response scale across all eight languages. My stress-test sharpens this into a concrete mechanism: the word bans and the negative-sentence ban in Section 3.3 are not neutral formatting choices. They target high-frequency vocabulary and negation that appear in the MFQ-2 items and in the low end of the Likert scale, so compliance pressure can shift the distribution of ratings in language-dependent ways. Because the same constraints are machine-translated into eight languages, subtle translation asymmetries can be confounded with cultural differences. The paper provides no control condition to separate these effects. This concern is load-bearing because the central claim is about what the models' language-conditioned moral scores mean. If the score differences are partly caused by instruction-following difficulty, then the cross-language variability is not evidence of moral adaptation. I do not recommend rejection: the issue is fixable by adding a control condition, and the reader's CONDITIONAL verdict already captures the need for revision. I also considered the statistical concern that 100 repeated outputs are treated as independent samples, but the prompt-artifact concern is more fundamental because it threatens the interpretation of every reported difference, not just the error bars. Thus the verdict should remain CONDITIONAL, with no change from the reader's recommendation.","tokens_in":19653,"tokens_out":4463,"duration_ms":47554,"concrete_test":"Run a control condition in each of the eight languages with the same four models and the same MFQ-2 items, but with a minimal instruction that only requests a Likert rating, e.g., 'Please choose the option that best describes how well the statement describes you', and no word/negation bans. Recompute the six foundation means, the language and model ANOVA in Table 1, and the Tukey HSD comparisons involving English. If the cross-language pattern and the English-vs-other differences disappear or change sign under the control, the Section 4.1 conclusion is an artifact of the Section 3.3 constraints; if the pattern is reproduced, the prompt-construction concern is settled and the conditional verdict can be upgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 4.1 is that multilingual LLMs adapt their moral reasoning to language-specific nuances rather than imposing English moral norms universally. The evidence for this claim rests entirely on language-condition differences in MFQ-2 scores. However, Section 3.3 injects a non-standard instruction set that forbids the words 'cannot', 'unable', 'instead', 'as', 'however', 'it', 'unfortunately', and 'important', and forbids any negative sentence about the prompt subject. These constraints conflict with the MFQ-2 items themselves, e.g., item 5 ('I think it is important for societies to cherish their traditional values') and item 1 ('Caring for people who have suffered is an important virtue'), and with the Likert anchor 'Does not describe me at all'. If a model treats rule 6 as prohibiting negative statements about the item, it may systematically avoid low response categories; if the translated instructions introduce asymmetric difficulty across languages, then the language main effects and pairwise differences in Tables 1 and 4 can be produced or suppressed by instruction-following behavior rather than by genuine moral content. No control condition without these constraints is reported, and no validation shows that the constrained prompt yields response distributions equivalent to a standard MFQ-2 administration. Under this confound, the ANOVA significance in Table 1 and the Tukey HSD results are not interpretable as evidence for moral adaptation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates whether multilingual LLMs impose English moral norms when prompted in eight languages. Using the 36-item MFQ-2, the authors collect Likert-scale responses from GPT-3.5-Turbo, GPT-4o-mini, Llama 3.1, and MistralNeMo in Arabic, Farsi, English, Spanish, Japanese, Chinese, French, and Russian, repeating each questionnaire 100 times per language and model. They report two-way ANOVAs showing significant language, model, and interaction effects, t-tests comparing WEIRD versus non-WEIRD language groups, and ANOVAs comparing model responses to human MFQ-2 data from Atari et al. (2023). The paper concludes that multilingual LLMs adapt their moral reasoning to language-specific nuances rather than universally imposing English moral norms, that GPT models align more closely with human moral judgments, and that smaller open-source models show stronger WEIRD biases.","tokens_in":19834,"tokens_out":6974,"duration_ms":62313,"significance":"If the findings are valid, the paper is among the first to apply MFQ-2 to multilingual LLMs and would provide a useful challenge to the simple 'English-norm imposition' narrative. The study has clear strengths: it uses a validated external instrument, adopts official MFQ-2 translations for eight languages, benchmarks against human data from Atari et al. (2023), and repeatedly samples model outputs to quantify variability. These features make the empirical core reproducible in principle. However, the central claim rests on a prompt-engineering design whose artificial word and negation constraints are unvalidated and plausibly confound the language-condition differences, so the contribution is conditional on the outcome of additional control experiments and sharper statistical interpretation.","major_comments":[{"comment":"The instruction rules directly conflict with the measurement instrument. Rule 5 forbids the words 'cannot', 'unable', 'instead', 'as', 'however', 'it', 'unfortunately', and 'important', yet MFQ-2 items and the task itself contain such words (e.g., item 5: 'I think it is important for societies to cherish their traditional values'; item 7: 'I believe that compassion ... is one of the most crucial virtues'). More seriously, Rule 6 forbids 'any negative sentences about the subject of the prompt', while the lowest Likert anchor is 'Does not describe me at all'—a negative sentence about the item. A model that follows Rule 6 will systematically avoid the lower end of the response scale, and if the translated instructions differ in how strongly they impose this rule across the eight languages, the language main effects in Table 1 and the descriptive differences in Table 4 can be generated by instruction-following artifacts rather than by moral content. The paper reports no control condition without these constraints and no validation that the constrained prompt yields response distributions comparable to a standard MFQ-2 administration. This confound directly undermines the central claim in Section 4.1 that models adapt to language-specific moral nuances.","section":"Section 3.3 and Appendix C"},{"comment":"There is an internal contradiction about the evidence for RQ1. The text first states, 'In each model, English shows significant differences from other languages across several moral foundations,' and then, two paragraphs later, states that 'Tukey's HSD post-hoc tests revealed that the differences involving English were not statistically significant for most moral foundations.' These statements cannot both be true without clarification of which tests and which foundations are being summarized. If the Tukey results are the correct reading, then the ANOVA language main effect is not specifically about English, and the conclusion that English norms are not imposed universally is not directly supported by the reported pairwise comparisons. The paper needs a consistent reporting of the post-hoc results, including the specific pairs that differ and effect sizes for the differences.","section":"Section 4.1"},{"comment":"The conclusion that models 'adapt their moral reasoning to reflect language-specific nuances' requires evidence that language-conditioned model responses align with human moral differences in the corresponding cultures. The paper compares models and humans via ANOVAs per foundation (Table 3), but this does not show per-language convergence; a model could have a small overall ANOVA difference from humans while still being misordered across languages. I recommend reporting a per-language agreement metric (e.g., mean absolute error or correlation across the six foundations) between model and human scores, and testing whether those metrics are better than a language-independent baseline. Without such an analysis, the 'adaptation' wording overstates what the data show.","section":"Sections 4.1 and 4.3"}],"minor_comments":[{"comment":"The paper is inconsistent about the number of moral foundations: Section 1 and Section 2.1 mention six foundations including liberty/oppression, while Section 2.2 says MFT introduced five dimensions and lists the original five. Please clearly distinguish the theoretical foundations of MFT from the six dimensions measured by MFQ-2 (care, equality, proportionality, loyalty, authority, purity).","section":"Section 2.2"},{"comment":"The instruction rule list differs between the main text and the appendix. In Section 3.3, Rule 1 is 'Do not elaborate on your reasoning' and the list has six rules; in Appendix C, the list starts with 'Do not say any other things instead of options' and has only five rules. The two descriptions should be aligned.","section":"Section 3.3 and Appendix C"},{"comment":"Classifying Russian as a WEIRD language is nonstandard for the psychology literature in which WEIRD refers to populations (Western, Educated, Industrialized, Rich, Democratic). Please justify this grouping or rename the contrast (e.g., 'European/Western' versus 'non-Western') to avoid a mischaracterization that affects the interpretation of Table 2.","section":"Section 3.2"},{"comment":"No decoding parameters (temperature, top-p, max tokens) are reported for any of the four models. These should be specified because the 100 repetitions per language and model may be degenerate if the models are sampled with temperature zero or a fixed seed.","section":"Section 3"},{"comment":"The ANOVA results do not report degrees of freedom, and the p-values are extremely small (many below 1e-18). The paper should report effect sizes (e.g., partial eta-squared) to give a sense of practical significance in addition to the F-tests, and should verify that the scientific notation in the table is accurate and reproducible.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate new application of MFQ-2 to multilingual LLMs, and the descriptive results are worth a look. But the headliner claim—that LLMs adapt to language-specific moral norms—rests on a prompt that artificially constrains the models' responses, and the statistical analysis treats repeated draws as independent samples. So I would not trust the specific numbers until the authors run a control condition and report sampling parameters.\n\nWhat's new: they're the first to use the updated MFQ-2 across eight languages with four models, and they compare against the human norms from Atari et al. (2023). The descriptive tables and figures show genuine variability across languages and models. That's useful, especially for people deciding which model to deploy in non-English settings. The paper also honestly lists limitations, including the newness of MFQ-2 and the gap between LLM responses and human psychological assessment.\n\nWhere it gets soft: Section 3.3 injects instructions that forbid the words 'cannot', 'unable', 'instead', 'as', 'however', 'it', 'unfortunately', and 'important', and forbid any negative sentence about the prompt subject. That conflicts with the Likert anchor 'Does not describe me at all' and with the questionnaire items themselves. If the model interprets rule 5 as prohibiting low ratings, or if the translated instructions create asymmetric difficulty across languages, the language main effects and pairwise differences in Tables 1 and 4 could be prompt artifacts. No control condition without these constraints is reported. This is the load-bearing confound, and it's fixable: run the same questionnaire with a neutral instruction prompt and compare.\n\nThe stats worry me too. Each cell is 100 model outputs, but the paper doesn't report temperature or other sampling parameters. Several cells show zero variance (e.g., Llama 3.1 French Care, Russian Care), which suggests deterministic generation or collapse. If the repetitions aren't independent, the ANOVA F-stats and Tukey results are not meaningful. The human-alignment analysis also pools model repetitions against individual human responses, which is a unit-of-analysis problem. Russia being grouped as WEIRD is a minor judgment call; I'd note it but not fight over it. And there's no code or data, which makes it hard to reproduce.\n\nWho it's for: people working on cultural bias and multilingual LLM evaluation, and anyone interested in the limitations of questionnaire-based probing. The paper deserves a serious referee, but the referee should push for the control condition and a proper statistical treatment. If those come in a revision, this could be a solid contribution. As it stands, I wouldn't use the numbers for policy or model selection.","headline":"Useful new MFQ-2 application to multilingual LLMs, but the untested prompt constraints and weak independence assumptions undermine the central cultural-adaptation claim.","tokens_in":20438,"tokens_out":3755,"would_cite":true,"duration_ms":33002,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multilingual language models shift their moral foundation scores with the language they are prompted in, rather than imposing a single English moral profile.","keywords":["moral foundations theory","multilingual LLMs","cultural bias","MFQ-2","moral alignment","cross-lingual evaluation","GPT-3.5-Turbo","Llama 3.1"],"falsifier":"Re-run the protocol in all eight languages without the forbidden-word and negative-sentence rules, using back-translation to match prompt meaning across languages; if the language-by-model interaction and the WEIRD/non-WEIRD gaps shrink or disappear, the observed cultural adaptation is an artifact of instruction constraints. Additionally, if the claim is genuine cultural knowledge, the model scores for a language should correlate with that language's human MFQ-2 scores; a language where model and human scores move in opposite directions would count against the claim.","tokens_in":19352,"feed_emoji":"🌐","tokens_out":8384,"duration_ms":66610,"temperature":0.7,"pith_summary":"The paper asks whether multilingual large language models impose English-centered moral norms when asked about morality in eight languages, or whether their moral reasoning shifts with the language. Using the updated 36-item Moral Foundations Questionnaire (MFQ-2), the author measures six moral foundations in four models: GPT-3.5-Turbo, GPT-4o-mini, Llama 3.1, and MistralNeMo. The central finding is that moral foundation scores vary significantly across languages and across models, with a language-by-model interaction, so no single English moral profile is imposed universally. The paper also finds that the two GPT models align more closely with human MFQ-2 responses than the smaller open-source models, especially in well-represented languages. A sympathetic reader would care because this challenges the default assumption that an LLM has one moral identity and points to language-specific evaluation and development as the path to fairer multilingual AI.","feed_headline":"LLMs adapt moral answers to each language, not just English","feed_subtitle":"Moral scores shift across eight languages and four models; GPT models sit closest to human responses.","key_machinery":"The central instrument is MFQ-2, a 36-item questionnaire covering six moral foundations—care, equality, proportionality, loyalty, authority, and purity—on a 5-point Likert scale. The carrying mechanism is a constrained prompting protocol: each item is presented in the official MFQ-2 translation for one of eight languages, with system rules requiring the model to answer only with a scale option and forbidding words such as 'cannot', 'instead', 'however', and 'it', plus any negative sentences. Each questionnaire is repeated 100 times per language per model. Statistical inference rests on two-way ANOVA with Tukey HSD post-hoc tests for language and model effects, t-tests for WEIRD versus non-WEIRD language groups, and ANOVA comparing model scores with human MFQ-2 responses from the validation study.","core_discovery":"The author's central claim is that multilingual LLMs adapt their moral reasoning to language-specific nuances rather than imposing English moral norms universally. Evidence comes from four models answering the same 36-item MFQ-2 in Arabic, Farsi, English, Spanish, Japanese, Chinese, French, and Russian: two-way ANOVA finds significant effects of language, model, and their interaction for all six foundations, while Tukey HSD tests show English differs significantly from most other languages only on scattered foundations. WEIRD and non-WEIRD language groups split systematically, with GPT models more balanced and smaller open-source models leaning toward WEIRD-language norms. Compared with human MFQ-2 responses in six of the eight languages, GPT-4o-mini and GPT-3.5-Turbo show the closest overall alignment, Llama 3.1 comes closest on Care, and MistralNeMo lags on every foundation; no model reproduces human moral judgments in all languages.","pith_inferences":["Editorial inference: the prompt rules banning words like 'it' and 'however' may bind more tightly in some languages than others, so part of the measured cross-language variation could be a translation artifact; a control condition without those rules would test this.","Editorial inference: if language-specific moral scores are genuine, then alignment benchmarks should treat each language as its own population, with human norms collected per language, rather than averaging moral scores across languages.","Editorial inference: the exclusion of Chinese and Farsi from the human-alignment analysis leaves two of the eight languages unanchored; adding human MFQ-2 norms for them could reorder the model rankings.","Editorial inference: the design cannot separate cultural adaptation from training-data mimicry; a follow-up that regresses model foundation scores on human cultural survey values per language would tell which explanation holds."],"forward_implications":["Moral evaluations of an LLM cannot be read off from English-only results; each language needs its own baseline.","Deployment choices should be model-specific: the larger GPT models are closer to human moral scores, while smaller open-source models lean more strongly toward WEIRD-language norms.","The language-by-model interaction means that moral output is shaped jointly by training data and architecture, so claims about a model's morality should name the language and model.","Because Arabic and Japanese show the largest deviations from human responses, LLM-based moral or social judgment in those languages carries real risk and warrants language-specific calibration."],"supporting_citations":[{"why":"Provides MFQ-2 itself, the human response data used for alignment, and the language selection rationale.","marker":"Atari et al. (2023)"},{"why":"Supplies the official MFQ-2 translations used to prompt the models in the eight languages.","marker":"Atari et al. (2022)"},{"why":"Defines MFQ-1 and the original five foundations that MFQ-2 extends and revises.","marker":"Graham et al. (2009, 2011)"},{"why":"Introduces Moral Foundations Theory, the psychological framework the questionnaire operationalizes.","marker":"Haidt and Joseph (2004)"},{"why":"Prior multilingual MFQ-1 study whose problems with negation and longer sentences motivate the paper's restrictive prompt rules.","marker":"Hämmerl et al. (2022)"}],"fun_headline_variants":["AI morals shift with language, not just English rules","Multilingual AI reveals cultural moral bias","GPT models mirror human morals across languages","Language changes LLM moral compass: study","LLMs adopt each language's moral norms uniquely"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the translated MFQ-2 items and the restrictive response rules mean the same thing in all eight languages; if the banned words or the Likert labels read differently in translation, the cross-language differences could be prompt artifacts rather than genuine moral variation.","fun_headline_variants_meta":{"raw":{"variants":["AI morals shift with language, not just English rules","Multilingual AI reveals cultural moral bias","GPT models mirror human morals across languages","Language changes LLM moral compass: study","LLMs adopt each language's moral norms uniquely"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":1167,"prompt_tokens":918,"completion_tokens":249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":182}},"tokens_in":534,"tokens_out":249,"duration_ms":2997,"temperature":1.0,"reasoning_tokens":182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:23:32.596680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the protocol in all eight languages without the forbidden-word and negative-sentence rules, using back-translation to match prompt meaning across languages; if the language-by-model interaction and the WEIRD/non-WEIRD gaps shrink or disappear, the observed cultural adaptation is an artifact of instruction constraints. Additionally, if the claim is genuine cultural knowledge, the model scores for a language should correlate with that language's human MFQ-2 scores; a language where model and human scores move in opposite directions would count against the claim.","supporting_citations":[],"review_version":1}