{"id":"856243e1-08bf-4b77-a992-210a3a225db3","arxiv_id":"2412.19926","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 1,730 LLM-generated ethical dilemmas, models reliably prefer truth over loyalty, community over individual, and long-term over short-term benefits, while showing sensitivity to question phrasing.","lead":"This paper tests how 20 large language models handle ethical dilemmas where both options are morally defensible. It finds that LLMs consistently favor truth, community, and long-term outcomes, but their answers change with prompt wording.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset-construction fidelity is the load-bearing assumption: without human validation of value-action alignment and consequence valence, the reported preference rates may measure prompt artifacts rather than LLM moral preferences.","rationale":"The paper's strongest claim is quantitative, with precise preference rates, and every one of those rates is computed over a corpus whose only quality checks are format compliance and the presence of the two value names in the explanation. The generation prompt asks an LLM to produce an action that follows one value and another action that follows the opposing value, but no independent check verifies that the generated text matches these labels. This is not an external or speculative objection; it is a measurement-instrument problem at the core of the study. If actions are mislabeled or scenarios are not genuine dilemmas, the aggregate numbers in Figure 2 do not describe LLM moral preferences; they describe properties of the generated corpus. The reader identified exactly this assumption as the weakest, and the conditional verdict is appropriate because the concern is addressable: human validation and data release would settle it. Other issues, such as prompt sensitivity, the unanimity selection filter, and the absence of statistical significance tests, are real but secondary; fixing them would not rescue a corpus that fails to instantiate the intended value conflicts. I therefore agree with the reader's framing and see no reason to change the verdict.","tokens_in":17538,"tokens_out":2577,"duration_ms":27645,"concrete_test":"Take a random stratified sample of 200 dilemmas, balanced across the four value pairs and both generator models, and have two independent human annotators label each item on three questions: (1) Does Action A align with Value 1 and Action B with Value 2? (2) Are both actions defensible, so the scenario is a genuine right-vs-right dilemma? (3) Are the attached consequences correctly valenced as positive and negative? Then recompute the four headline preference rates using only scenarios where both annotators agree the labels are correct. If the recomputed rates differ from the reported 93.48%, 83.69%, 72.37%, and 68.49% by more than five percentage points, or if a substantial fraction of sampled dilemmas fail validation, the central claim is not robust to dataset-generation bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that LLMs exhibit pronounced, stable preferences in right-vs-right dilemmas, e.g., truth over loyalty in 93.48% of unanimous scenarios. That claim inherits the correctness of the generated dataset. Section 3.2 constructs the 1,730 dilemmas with gpt-3.5-turbo and claude-3-sonnet, and then filters only on format compliance and whether the explanation mentions the two value names. There is no human validation that Action 1 actually instantiates Value 1, that Action 2 instantiates Value 2, that the scenario is a genuine right-vs-right dilemma rather than a temptation scenario, or that the generated consequences are correctly valenced as positive or negative. The generation prompt instructs the model to label actions by value, but that label is precisely the quantity under test. If the generator systematically pairs one value with more sympathetic or more socially normative actions, or if consequences are mis-valenced, the aggregate preference rates in Figure 2 and the deontology/consequentialism result in Section 5.3 would be artifacts of dataset construction rather than properties of LLMs. The dataset is not released, so this cannot be checked post hoc. The unanimity filter in Section 5.2 further restricts analysis to scenarios where models already agree, which can only amplify any generation bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a descriptive study of how large language models behave when confronted with ethical dilemmas in which both options are presented as morally defensible. The authors construct a corpus of 1,730 dilemmas based on Kidder's four value-conflict pairs (Truth vs. Loyalty, Individual vs. Community, Short-Term vs. Long-Term, Justice vs. Mercy) using two LLM generators, then evaluate 20 LLMs across six families under multiple prompt formulations. The main reported findings are that LLMs prefer truth over loyalty, long-term over short-term benefits, and community over individual, with a weaker preference in the Justice/Mercy pair; that larger models are less likely to change their choices when consequences are added; and that explicitly stated value preferences steer models more effectively than in-context examples. The paper also reports that LLMs vary substantially in their sensitivity to prompt formulation. The study is broad in scope and addresses an underexplored question, but the central empirical claims rest on an unvalidated, unreleased LLM-generated dataset, and several reporting gaps make the robustness of the headline percentages difficult to assess.","tokens_in":70,"tokens_out":11629,"duration_ms":173172,"significance":"If the dataset-construction assumptions hold, this would be a valuable descriptive contribution to machine ethics and value alignment: it provides a systematic map of LLM preferences across a well-known ethical taxonomy, covers a much larger scenario space than prior dilemma-based studies, and includes useful comparisons of prompt families with reversal controls. The finding that explicit value statements outperform few-shot demonstrations is practically relevant for prompt-design and alignment work. The paper is honest about some limitations, including prompt sensitivity and the simplified nature of its ethical framework. However, the study does not ship code or data, and its empirical claims currently depend on an LLM-generated corpus with no human validation and no reported inter-annotator agreement, which is a serious correctness risk for every headline percentage. The lack of statistical inference and the inconsistent reporting of the Justice/Mercy direction further weaken the current presentation.","major_comments":[{"comment":"The dataset-construction step is load-bearing for every headline result, and it is not validated. The 1,730 scenarios, the two actions, the assignment of each action to one side of the value pair, and the positive/negative consequences are all generated by gpt-3.5-turbo and claude-3-sonnet, and the only filters applied are format compliance and whether the explanation mentions the two value names. There is no human validation that Action 1 actually instantiates value 1, that Action 2 instantiates value 2, that the scenario is a genuine right-vs-right dilemma rather than a temptation scenario, or that the generated consequences have the intended valence. Because the generation prompt asks the model to label actions by the very values under test, the aggregate preference rates in Figure 2 and the consequence-flipping rates in Figure 3 can reflect the generators' biases or prompt artifacts rather than properties of the evaluated LLMs. The circularity is compounded by the fact that GPT-3.5 and Claude-3-Sonnet are themselves among the 20 evaluated models. The dataset is not released, so these concerns cannot be checked post hoc. Please report a human validation study on a representative sample, state the resulting agreement rates, and release the dataset or provide a clear availability statement.","section":"Section 3.2"},{"comment":"The headline preference rates (e.g., 'Truth is overwhelmingly preferred over Loyalty, with an average selection rate of 93.48%') are computed only on scenarios in which a model gives unanimous answers across all four prompt variants, after excluding all models whose all-four-prompt agreement is below that of Llama-3-8B (47.4%). This removes Claude-3-Sonnet, Mixtral-8x22B, Llama-2 (7B, 13B), Qwen2 (0.5B, 1.5B, 7B), and Yi-1.5 (6B, 9B) from Figure 2, so the cross-model comparison covers only 12 of the 20 evaluated models. The manuscript does not report how many scenarios survive the unanimity filter per model and per value pair, nor does it quantify the resulting selection bias. Restricting to unanimously answered scenarios can amplify any systematic bias in the generated dataset, because a generator that makes one action consistently more attractive will also produce more unanimous answers. Please report the number of retained scenarios per model/pair and show that the main conclusions are robust when the analysis is repeated on all scenarios, or justify the unanimity restriction as the primary analysis.","section":"Section 5.2, Figure 2"},{"comment":"The direction of the Justice-vs-Mercy preference is reported inconsistently. Section 5.2 states 'Mercy is chosen over Justice 68.49% of the time,' and the first block of numbers in Figure 2 is consistent with that reading. However, Section 6 (Summary of Findings) states 'an average of 68.49% favoring Justice.' The abstract omits the Justice/Mercy pair entirely. Since this is one of the four central descriptive findings, the discrepancy must be corrected, and the abstract should state the direction if the finding is included.","section":"Sections 5.2 and 6; Figure 2"},{"comment":"The claim that 'larger and more advanced models have a greater inclination to adopt the deontological principle' is stronger than the data shown. The flip rates in Figure 3 are not monotonic in model size or recency: Yi-1.5-34B flips 16.5%, Mixtral-8x7B 20.7%, and Llama-2-70B 12.3%, while Claude-3-Haiku flips 18.7%, so the 'larger' generalization is supported only by a loose trend among a subset of models. No confidence intervals or significance tests are reported for these flip rates or for the other headline percentages. Please either temper the claim to 'some of the largest models in this sample' or provide a statistical analysis that supports the size/recency generalization.","section":"Section 5.3, Figure 3"},{"comment":"The handling of invalid responses is not reported. Section 5 states that Mistral-7B-Instruct-v0.2 and Mixtral-8x7B-Instruct-v0.1 have average invalid response rates of about 20% and 10%, respectively, in some prompt conditions, but the paper does not say whether these responses are discarded, treated as missing, or imputed, nor whether the reported percentages are computed over valid responses only. Because these models are included in the preference-rate and consequence-flipping analyses, the per-model sample sizes differ, and the comparisons in Figures 2-5 may be based on different numbers of valid responses. Please specify the exclusion rule and report per-model, per-condition sample sizes.","section":"Section 5, Table 9"}],"minor_comments":[{"comment":"There are several typos in model names and technical terms: 'speciﬁced' in the abstract, 'Mistrial-7b' and 'Qwne2' in Section 4.2, and 'GPT-40' in Section 5.4. These should read 'specified', 'Mistral-7B', 'Qwen2', and 'GPT-4o' respectively.","section":"Abstract, Section 4.2, Section 5.4"},{"comment":"The figures are hard to read in the current rendering: the pair labels in Figure 2 (e.g., 'T ruth Loyalty') are ambiguous about which percentage belongs to which value, and Figure 6 has no color scale. Please add explicit value labels and legends or a color bar.","section":"Figures 2, 4, and 6"},{"comment":"The example in Section 3.2 and the instances in Table 4 include an 'Explanation' field, but the evaluation prompts in Table 6 do not contain this field. Please state explicitly that the Explanation is not shown to the evaluated models, so that the intended values are not leaked into the test prompts.","section":"Section 3.2 and Table 6"},{"comment":"The phrase 'scenarios where LLMs achieve unanimous agreement across all four prompts' should be stated more precisely: the unanimity is a property of a single model across the four prompt variants, not a property of the scenario independent of the model. Please clarify this in the text.","section":"Section 5.2"},{"comment":"Open models were run in 4-bit quantization while API models were not; this is a potential confound for cross-family and cross-size comparisons. Please discuss the likely impact of quantization, or provide a sensitivity analysis on at least one open model run at full precision.","section":"Section 4.2"},{"comment":"The comparison between explicit and few-shot preference handling uses different instruction wording: the explicit prompt states the preference directly, while the few-shot prompt requires the model to infer it from examples. The observed gap may therefore partly reflect instruction clarity rather than preference recognition alone. Please acknowledge this confound in the discussion.","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.CL and would be of interest to the ethics and alignment community. My recommendation is driven primarily by the unvalidated, unreleased LLM-generated dataset, which is load-bearing for all headline percentages, and by the inconsistent reporting of the Justice/Mercy direction in Section 6. If the authors provide a human validation study, release the dataset, and correct the reporting issues, I would support acceptance. I would also ask the editor to ensure that the authors cross-check the central numbers against the figures before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely useful empirical study of how LLMs navigate right-vs-right dilemmas, and it deserves a serious referee. But the central numbers—93.48% truth-over-loyalty, 83.69% long-term-over-short-term, etc.—should be treated as provisional until the dataset is human-validated and released. What the paper does well: it scales up a previously thin line of work. Instead of a few hand-picked dilemmas, it gives you 1,730 scenarios across four Kidder value pairs, evaluates 20 models from six families, and probes sensitivity, consistency, consequence-sensitivity, and preference following with multiple prompt variations. The finding that explicit value statements beat in-context exemplars is solid and practically relevant. The authors are also transparent about their construction process and limitations, which makes the paper easy to interrogate. The soft spots are real, though not fatal in my view. The dataset is generated by gpt-3.5-turbo and claude-3-sonnet with only format checks and lexical mention of the two values; there is no human check that Action 1 truly instantiates Value 1, or that the positive/negative consequences are correctly valenced. Since those labels are exactly what is being tested, a systematic generation bias could inflate the reported preferences. That is the load-bearing assumption. The unanimity filter in Section 5.2—keeping only scenarios where all four prompt variants agree—can only amplify any generation bias, and it also excludes several smaller models, so the aggregate percentages describe a subset of scenarios and models. The paper provides no statistical significance testing, so we cannot tell whether the between-pair differences are meaningful or noise. And the dataset is not released, so none of this can be checked post hoc. That said, I don't think the concern is that the authors fabricated anything. The example scenarios in the appendix look genuinely like right-vs-right dilemmas, and the aggregate preferences align with prior qualitative work on LLM moral tendencies. The bias, if present, is likely one of degree rather than kind. Who benefits: anyone working on LLM alignment, value-aware deployment, or machine ethics benchmarks. The paper is a solid contribution to an empirical literature that is still thin on scale. I would send it to review, but I would push for the authors to provide a human-validated subsample and release the data. Do not desk-reject; this is worth referee time.","headline":"A useful, well-scoped empirical map of LLM moral preferences on right-vs-right dilemmas, but the headline percentages rest on an LLM-generated dataset that has not been human-validated or released, so the results should be read as provisional.","tokens_in":780,"tokens_out":1029,"would_cite":true,"duration_ms":29770,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs show consistent, pronounced value preferences in right-vs-right dilemmas, favoring truth over loyalty in 93% of cases and community over the individual in 72%.","keywords":["ethical dilemmas","large language models","moral value preferences","value alignment","Kidder's framework","deontology","consequentialism","prompt sensitivity"],"falsifier":"Have independent human raters classify a random sample of the 1,730 dilemmas: which value each action instantiates, and whether the scenario truly poses a right-vs-right conflict. If raters frequently disagree with the generator's intended assignments, or if the measured preference rates (truth over loyalty at 93.48%, for example) fail to replicate on the subset of scenarios that pass human validation, then the reported percentages are artifacts of the generation pipeline rather than stable properties of LLM moral reasoning.","tokens_in":17317,"feed_emoji":"⚖️","tokens_out":15860,"duration_ms":134808,"temperature":0.7,"pith_summary":"This paper asks what large language models do when every option is morally defensible — dilemmas of right versus right, not right versus wrong. Drawing on Kidder's four recurring value conflicts (truth vs. loyalty, individual vs. community, short-term vs. long-term, justice vs. mercy), the authors build a corpus of 1,730 dilemmas and test 20 models from six families. On scenarios where a model consistently agrees with its own answers, it favors truth over loyalty in 93.48% of cases, long-term over short-term interests in 83.69%, community over the individual in 72.37%, and justice over mercy in 68.49%. Larger models hold these choices even when told the outcomes are negative, and explicit statements of a preferred value steer models far more reliably than example cases. Because the same models agree with themselves only 60–80% of the time across different phrasings of a dilemma, these preferences sit atop a measurable sensitivity to wording — and that combination matters for any system that delegates tough choices to an LLM.","feed_headline":"LLMs favor truth over loyalty in 93% of ethical dilemmas","feed_subtitle":"Across 1,730 right-vs-right dilemmas and 20 models, explicit value rules steer choices better than examples.","key_machinery":"The engine of the study is Kidder's four-paradigm taxonomy of right-vs-right dilemmas, which supplies the four value conflicts — truth vs. loyalty, individual vs. community, short-term vs. long-term, justice vs. mercy — that every scenario must instantiate. On top of it sits a two-step generation pipeline: gpt-3.5-turbo first lists domains where such dilemmas arise, then gpt-3.5-turbo and claude-3-sonnet each write scenarios with two actions, one aligned with each value, plus positive and negative consequences for each action; the only filtering is format compliance and mention of both value words. The measurement apparatus is twelve prompt templates: four phrasings (direct, direct-reverse, compare, compare-reverse) probe the baseline preference and the model's symbol binding; four consequence permutations probe whether outcomes move the choice; and explicit-preference plus few-shot templates (one, five, ten examples) probe steerability. The analytical keystone is unanimity filtering: only scenarios where a model answers identically under all four baseline phrasings contribute to its preference statistics, an attempt to separate a stable moral lean from prompt noise.","core_discovery":"The central discovery is that LLMs show strong, repeatable value preferences when every option in a dilemma is morally defensible, and that those preferences survive changes in wording, option order, and even stated consequences. Restricting the analysis to twelve models that passed a self-consistency threshold, and to scenarios where a model gave the same answer under all four baseline prompt formulations, the authors report that truth is chosen over loyalty in 93.48% of cases, long-term over short-term interests in 83.69%, community over the individual in 72.37%, and justice over mercy in 68.49%. A second finding is scale-dependent: the largest, newest models flip their choice in only 5–10% of scenarios when consequences are reversed — a deontological stance, adhering to the initial rule regardless of outcomes — while smaller models flip in up to roughly 32% of scenarios. A third concerns steerability: an explicit statement that one value outranks another reverses the model's baseline choice in up to about 85% of scenarios, whereas demonstrating the preference through one to ten examples reverses it in at most about 44%. A fourth is a floor of fragility: even the best models agree with themselves only 60–80% of the time across four phrasings of the same dilemma, so the measured preferences sit on top of substantial prompt-formulation sensitivity.","pith_inferences":["Because the 1,730 scenarios were generated by just two models with no human validation of the value assignments, the measured preference rates could partly reflect the generation pipeline rather than LLM moral reasoning; a human-rater study of a random sample would separate the two.","The unanimity filter keeps only the clearest dilemmas for each model, so the reported percentages likely overstate how often these preferences appear in messy real-world cases; computing the same rates on the full, non-unanimous responses would give a lower-bound estimate.","Since the two generator models were built largely on English-heavy, Western-oriented corpora, the truth-over-loyalty dominance may encode a particular training distribution; regenerating the corpus with models trained on other cultural data could show whether the profile is universal or inherited.","The explicit-beats-examples result suggests a practical alignment recipe — state the value hierarchy plainly — but also invites a follow-up question the paper does not ask: whether models can detect when a stated preference and a scenario's content pull in opposite directions."],"forward_implications":["Deploying an LLM where value trade-offs are routine — health care, journalism, human resources, governance — means inheriting a specific moral lean (truth over loyalty, community over individual, long-term over short-term) that the model will not necessarily surface or justify on its own.","Operators who want a particular ethical stance should state it explicitly in the prompt: explicit value guidelines reversed models' choices in up to about 85% of scenarios, while in-context examples managed at most about 44%.","Larger models' deontological stability means outcome information will not talk them out of an initial moral choice, so applications that want consequence-sensitive decisions should not rely on large models' default behavior.","Self-agreement of only 60–80% across prompt phrasings means any single-prompt answer to an ethical question carries a formulation-dependent noise floor; agreement checks across phrasings are a cheap way to flag unreliable answers.","The justice-vs-mercy pair shows the thinnest consensus (68.49%) and the most variation across both consequence and preference manipulations, marking it as the most context-dependent and steerable of the four conflicts."],"supporting_citations":[{"why":"Supplies the four-paradigm taxonomy of right-vs-right dilemmas that defines the value conflicts the whole corpus is built on.","marker":"Kidder [1996]"},{"why":"Documents LLM sensitivity to prompt formatting, motivating the sensitivity question and the four prompt variations.","marker":"[Sclar et al., 2024]"},{"why":"Shows option-order bias in LLM choices, motivating the reversed-order prompts that detect symbol-binding failures.","marker":"[Wang et al., 2023]"},{"why":"The ETHICS benchmark representing the right-vs-wrong moral-temptation paradigm this study contrasts with right-vs-right dilemmas.","marker":"[Hendrycks et al., 2021]"},{"why":"Introduces statistical measures for the moral beliefs encoded in LLMs, the measurement approach this paper adapts to value conflicts.","marker":"[Scherrer et al., 2023]"},{"why":"Pioneers in-context ethical policies for LLMs, the approach the paper's explicit-preference and few-shot experiments build on.","marker":"[Rao et al., 2023]"},{"why":"Connects consistency to moral integrity in dilemma studies, the standard the paper's consistency analysis targets.","marker":"[Arvanitis and Kalliris, 2020]"}],"fun_headline_variants":["LLMs pick truth over loyalty 93% of the time","Bigger LLMs stick to rules even when consequences flip","Explicit moral rules steer LLMs better than examples","LLMs have consistent moral preferences but 20-40% phrasing sensitivity","Larger LLMs ignore consequences more often in dilemmas"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the AI-generated scenarios and their paired actions really instantiate the intended value conflict — that 'Action A' genuinely expresses truth and 'Action B' genuinely expresses loyalty in each of the 1,730 dilemmas — since the corpus was filtered only for format compliance and for mentioning the two value words, with no human validation reported.","fun_headline_variants_meta":{"raw":{"variants":["LLMs pick truth over loyalty 93% of the time","Bigger LLMs stick to rules even when consequences flip","Explicit moral rules steer LLMs better than examples","LLMs have consistent moral preferences but 20-40% phrasing sensitivity","Larger LLMs ignore consequences more often in dilemmas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000465,"raw_usage":{"total_tokens":2358,"prompt_tokens":1021,"completion_tokens":1337,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":637,"completion_tokens_details":{"reasoning_tokens":1254}},"tokens_in":637,"tokens_out":1337,"duration_ms":10395,"temperature":1.0,"reasoning_tokens":1254,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:45:15.412533+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent human raters classify a random sample of the 1,730 dilemmas: which value each action instantiates, and whether the scenario truly poses a right-vs-right conflict. If raters frequently disagree with the generator's intended assignments, or if the measured preference rates (truth over loyalty at 93.48%, for example) fail to replicate on the subset of scenarios that pass human validation, then the reported percentages are artifacts of the generation pipeline rather than stable properties of LLM moral reasoning.","supporting_citations":[{"cited_title":"Aligning AI with shared human values","cited_arxiv_id":null,"evidence_quote":"The ETHICS benchmark representing the right-vs-wrong moral-temptation paradigm this study contrasts with right-vs-right dilemmas."},{"cited_title":"Evaluating the moral beliefs encoded in LLM s","cited_arxiv_id":null,"evidence_quote":"Introduces statistical measures for the moral beliefs encoded in LLMs, the measurement approach this paper adapts to value conflicts."},{"cited_title":"Consistency and moral integrity: A self-determination theory perspective","cited_arxiv_id":null,"evidence_quote":"Connects consistency to moral integrity in dilemma studies, the standard the paper's consistency analysis targets."}],"review_version":1}