Pith. sign in

REVIEW 4 major objections 6 minor 45 references

The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read LLMs display shifting, non-transitive moral preferences.

desk verdict New multi-step dilemma benchmark with a useful descriptive finding, but the headline non-transitivity claim doesn't survive the aggregation assumption. read the letter →

arxiv 2505.18154 v1 pith:BOY2NPJV submitted 2025-05-23 cs.CL cs.CY

classification cs.CLcs.CY
keywords multi-stepmoraldilemmasLLMreasoningvaluealignmentfoundationstheorySchwartzbasicvaluespreferencetransitivitycontext-dependentethicsdynamicevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Multi-step Moral Dilemmas (MMDs), a dataset of 3,302 five-stage scenarios that escalate the same ethical conflict step by step. Using this benchmark, it evaluates nine LLMs and claims that the models' value preferences shift as dilemmas intensify, with care remaining a stable anchor while fairness rises and sanctity becomes more volatile. The central claim is that LLMs do not rely on stable moral principles; instead, pairwise comparisons of value choices produce non-transitive cycles such as care over sanctity, sanctity over fairness, and care roughly tied with fairness. The paper argues this pattern indicates context-driven statistical imitation rather than a coherent axiological hierarchy. If right, static single-question moral evaluations miss a core feature of LLM ethical behavior.

What carries the argument

The central object is the Multi-step Moral Dilemmas (MMDs) dataset: 3,302 five-step scenarios generated from Moral Stories norms, with each step adding a new moral tension to the same narrative. The critical mechanism is the 'causal context' prompting protocol, where each new dilemma includes the full history of prior steps and the model's own earlier choices; this is what lets the analysis attribute changes to value dynamics rather than isolated judgments. Value labels come from a consensus mapping of each action to Moral Foundations Theory and Schwartz value dimensions by three LLMs with human adjudication of disagreements. The preference structure is then probed by pairwise win rates and transitivity checks across value triads such as care, sanctity, and fairness.

What would settle it

Compute pairwise win rates within each individual five-stage dilemma separately and test transitivity within each dilemma; if nearly all individual dilemmas are transitive while the aggregate across dilemmas is not, the paper's intransitivity conclusion would be an averaging artifact. Alternatively, run the same model on a single triad of purpose-built dilemmas that isolate the care/sanctity/fairness cycle and see whether the cycle still appears when scenario content is held constant.

Watch

Extended reading notes

Core claim

The paper claims that LLM moral judgment is not organized by consistent, globally ordered values. Across the MMDs benchmark, models maintain their directional orientation on dimensions like care and fairness but systematically adjust preference strength as steps escalate. In pairwise win-rate comparisons, local cycles appear: for example, care beats sanctity, sanctity beats fairness, and care nearly ties fairness, a triad that violates transitivity. The paper interprets this as evidence that models generate value preferences through context-driven statistical imitation rather than appealing to stable principles, while also noting unambiguous trade-offs like care strongly outranking loyalty and liberty.

Load-bearing premise

The claim of non-transitive preferences assumes that win rates averaged across many different dilemmas can be read as one revealed preference relation for each model.

Editorial extensions

If this is right

  • Static, single-step moral evaluations will overstate the consistency of LLM values, since preference strength changes with escalation.
  • Value-alignment safety audits should treat context and history as first-class inputs rather than averaging over isolated judgments.
  • Rank shifts across steps supply a diagnostic: dimensions like authority lose inter-model agreement as dilemmas intensify, while liberty converges.
  • Benchmarks for moral reasoning should report non-transitivity alongside accuracy, because cycles reveal when a model is pattern-matching rather than reasoning from principles.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One implication the paper leaves implicit: aggregate win rates computed over many heterogeneous dilemmas can form Condorcet-style cycles even when every individual dilemma is transitive, so the intransitivity evidence would be stronger if computed per-dilemma.
  • A testable extension would be to feed each model a single triad of dilemmas built from the same scenario family and check whether the cycle care > sanctity > fairness > care persists under controlled contexts.
  • The result suggests a practical safety heuristic: when a model's pairwise value preferences cycle, interventions that target a single principle may backfire, since the model's ranking is locally context-dependent.
  • Connecting to social choice theory, treating LLM value aggregation as a voting rule predicts that the order in which dilemmas are presented can change which value wins, a hypothesis the dataset could test directly by permuting step order.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Multi-step Moral Dilemmas (MMDs), a dataset of 3,302 five-stage moral dilemma chains generated by GPT-4o from Moral Stories norms, with each step presenting two actions that are labeled with moral values by three LLMs under a consensus protocol. The authors evaluate nine LLMs under a "causal context" protocol in which each model sees prior steps and its own prior choices, and they report two sets of findings: (i) descriptive findings that value preference strengths shift as dilemmas escalate while rank orderings remain mostly stable, and (ii) a stronger interpretive claim that pairwise win-rate comparisons reveal non-transitive moral preferences (e.g., care > sanctity, sanctity > fairness, yet care ≈ fairness), which they take to indicate that LLMs do not rely on stable moral principles but engage in context-driven statistical imitation. The central non-transitivity claim rests on aggregate win rates pooled across all dilemmas and steps.

Significance. The MMDs dataset is a substantial new resource for studying dynamic moral judgment in LLMs, and the descriptive finding that aggregate value preference strengths shift across escalating steps is interesting and likely reproducible. The paper is commendable for releasing a large benchmark, for making the evaluation protocol explicit, and for reporting human verification of a sample of value labels. If the non-transitivity claim could be established at the level of individual dilemmas or with appropriate statistical controls, the work would be a meaningful contribution to debates about the coherence of LLM moral reasoning. However, as presented, the evidence for the headline claim about non-transitive and unstable moral principles is not sufficient; the analysis conflates aggregate mixture phenomena with individual-level preference relations. The paper's value lies more in the dataset and the dynamic evaluation protocol than in the strong theoretical conclusion.

major comments (4)
  1. [Section 4.2, Table 4 and Fig. 4] The non-transitivity claim is built on win rates pooled across all 3,302 dilemmas and, in the overall panels, across all five steps. Aggregate pairwise proportions over heterogeneous contexts cannot establish intransitivity of a single revealed preference relation: even if every individual dilemma were perfectly transitive, the mixture of context-specific value weights can produce aggregate cycles such as care>sanctity, sanctity>fairness, care≈fairness. No within-dilemma or within-step transitivity test, mixed-effects model, or per-dilemma variance estimate is reported. Since Finding 3 and the conclusion rest on this analysis, the paper must either provide a proper transitivity test that accounts for the nested structure of the data or substantially weaken the claim to a statement about aggregate win rates in this benchmark.
  2. [Section 4.2, Table 4] The inference of a 'care≈fairness' tie from observed win rates of 0.50–0.54 is not supported without confidence intervals and sample sizes. With an unknown denominator, 0.52 may be statistically indistinguishable from 0.5; if so, the purported cycle collapses. The same issue affects other near-threshold comparisons, including the Schwartz triads in Appendix C.2. The paper should report exact counts, confidence intervals, and formal tests of difference from 0.5 for the relevant pairwise win rates.
  3. [Section 3.2 and Appendix D] The value labels that drive all pairwise comparisons are produced by three LLMs (GPT-4o-mini, DeepSeek-V3, GLM-4-Plus), and models from the same family or the same systems are among the nine evaluated models. Human verification covers only 120 dilemmas, with 80.3% agreement for MFT and 83.5% for Schwartz; a non-negligible fraction of labels may be misclassified, and the paper does not report per-step or per-value-dimension agreement. This label noise can bias win-rate estimates in unknown directions. The authors should report the distribution of human agreement across value dimensions and steps, and either perform a sensitivity analysis restricted to high-agreement labels or otherwise justify that the label noise cannot account for the reported cycles.
  4. [Section 3.3] The claim that the causal-context protocol is superior to full-context and no-context protocols is asserted without quantitative evidence: the text states that full context produces 'rigid' behavior and causal context 'closely approximates human moral development', but no tables, scores, or statistical tests are given for this comparison. Since this protocol choice is central to the experimental design, the paper should either provide the supporting measurements or clearly mark these statements as qualitative observations.
minor comments (6)
  1. [Abstract and Section 3.1] The paper uses both '3,302 five-stage dilemmas' and '33,020 value dilemmas' (Table 1) without clarifying that the latter counts each individual step/action annotation; please align the terminology.
  2. [Table 1] The Schwartz value listed as 'Simulation' should be 'Stimulation'.
  3. [Table 3] The 'Consistency' column (High/Medium) lacks a definition or threshold; please specify how these categories are derived from the reported average rho values.
  4. [Appendix B.1] The sentence 'GLM-4-Plus initially favors security with a score of 0.086 over benevolence at 0.149 in Step 1' appears to state the reverse of the intended comparison, since 0.086 is not greater than 0.149.
  5. [Appendix C.2, Table 12] The row labeled 'Self < Stimulation' reports values 0.20, 0.40, 0.40 that are not consistently interpretable as win rates of self-direction over stimulation; clarify the direction and the numbers.
  6. [Ethical Statement] The ethical statement says 'we do not perform human annotations and tests', but Appendix D describes 12 human evaluators validating 120 dilemmas; this contradiction should be corrected.

Circularity Check

1 steps flagged · score 4.0 of 10

Partial circularity: the non-transitivity finding is measured with value labels assigned by the same models whose preferences are being evaluated.

  1. self definitional [Section 3.2 (Consensus-Based Model Value Mapping), Section 4 (Value Preference Analysis), Table 4]
    "we employ a consensus-based approach using three LLMs: GPT-4o-mini, GLM-4-Plus, and DeepSeek-V3"

    DeepSeek-V3 and GLM-4-Plus are both annotators in the consensus value-mapping step and are also among the nine evaluated models listed in Section 4. The pairwise win rates in Table 4 (e.g., DeepSeek care>sanctity 0.59) are computed from choices whose care/sanctity labels were assigned by a consensus that includes DeepSeek itself, and similarly for GLM-4-Plus. Thus the measured value preference is partly produced by the same system that defines which action counts as which value. The non-transitivity used to support Finding 3 is therefore not an independent measurement for these models, but is partially self-defined by the annotator-evaluator overlap.

full rationale

The paper's central claim that LLMs exhibit non-transitive and shifting moral preferences rests on value labels attached to actions and on choices made by the evaluated models. For two evaluated models (DeepSeek-V3 and GLM-4-Plus), the value labels are not externally imposed: these models are themselves part of the three-model consensus that assigns the MFT and Schwartz labels. Consequently, the pairwise win rates and the resulting 'local intransitivity' triads in Table 4 and Appendix C.2 are partly constructed from the same systems' own judgments. The human verification in Appendix D (80.3% agreement on 120 cases) provides partial external grounding, but it covers only a small sample and does not remove the self-referential loop for the full dataset. No formal equation-level circularity or self-citation chain was found; the aggregation of win rates across heterogeneous dilemmas is a statistical validity concern rather than a circularity. The dynamic preference shifts across steps are more directly observed from model choices and retain some independent content, but the strongest theoretical conclusion ('context-driven statistical imitation') relies on the non-transitivity analysis, which is the part most affected by the annotator-evaluator overlap. Overall, the derivation is not fully equivalent to its inputs, but the central finding is partially circular.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of LLM-generated value labels, the assumption that forced binary choices in hypothetical dilemmas reveal underlying value preferences, and the aggregation of win rates across contexts as if they formed a single preference ordering. No free parameters were fitted, but the analysis depends heavily on the LLM annotation pipeline and on the unexamined revealed-preference assumption.

assumptions (4)
  • domain assumption The consensus value labels produced by three LLMs (majority agreement with manual adjudication for ties) are a valid ground truth for the moral value expressed by each action.
    Used throughout Section 3.2 and the evaluation; only 120 of 3,302 dilemmas were human-validated (Appendix D), so the bulk of the labels are LLM-generated.
  • domain assumption A model's choice between two actions reveals a preference for the value label attached to that action.
    The preference analysis in Section 4 assumes that mapping from choices to value preferences is unproblematic, without considering that the same action could be chosen for reasons unrelated to the annotated value.
  • ad hoc to paper Aggregate pairwise win rates across different dilemmas can be treated as a single revealed preference relation for transitivity analysis.
    Table 4 and Section 4.2 infer intransitivity from average win rates computed over heterogeneous scenarios; this is only valid if preferences are context-invariant, which the paper itself argues against.
  • domain assumption Moral Foundations Theory and Schwartz's Theory of Basic Values are the appropriate complete taxonomies for modeling moral values.
    Adopted from prior literature (Graham et al., 2013; Schwartz, 2012); the authors acknowledge in the Limitations that this privileges Western-centric constructs and may underrepresent collectivist ethics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas." pith.science (2026). https://pith.science/paper/BOY2NPJV

@misc{pith2026250518154,
  author       = {Pith},
  title        = {Pith review of: The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BOY2NPJV}},
  note         = {Machine review of arXiv:2505.18154}
}
read the original abstract

Ethical decision-making is a critical aspect of human judgment, and the growing use of LLMs in decision-support systems necessitates a rigorous evaluation of their moral reasoning capabilities. However, existing assessments primarily rely on single-step evaluations, failing to capture how models adapt to evolving ethical challenges. Addressing this gap, we introduce the Multi-step Moral Dilemmas (MMDs), the first dataset specifically constructed to evaluate the evolving moral judgments of LLMs across 3,302 five-stage dilemmas. This framework enables a fine-grained, dynamic analysis of how LLMs adjust their moral reasoning across escalating dilemmas. Our evaluation of nine widely used LLMs reveals that their value preferences shift significantly as dilemmas progress, indicating that models recalibrate moral judgments based on scenario complexity. Furthermore, pairwise value comparisons demonstrate that while LLMs often prioritize the value of care, this value can sometimes be superseded by fairness in certain contexts, highlighting the dynamic and context-dependent nature of LLM ethical reasoning. Our findings call for a shift toward dynamic, context-aware evaluation paradigms, paving the way for more human-aligned and value-sensitive development of LLMs.

Figures

Figures reproduced from arXiv: 2505.18154 by the authors.

Figure 1
Figure 1. Comparison of existing value evaluation protocols and ours for LLMs. Instead of asking a single question or situating an isolated moral dilemma, our proposed MMDs framework sets a multi-step moral dilemma questionnaire to progressively induce models into stronger and more complex ethical conflicts to expose their underlying value priorities. Email(s): {wuya23s,shengqiang18z,wangdanding,sunyifan23z,wangzhengjia21b,bu… view at source ↗
Figure 2
Figure 2. ① Moral Dilemmas Generation: A five-level dilemma series (S1–S5) is generated, each with context (Ctx), decision (D), action (A), and action (B). ② Model Value Mapping: Decisions and actions are mapped to values such as Liberty, Care, Fairness, Loyalty, and Sanctity. ③ LLM Value Evaluation: A language model evaluates the values, producing scores V A 1 –V A 5 and V B 1 –V B 5 . ④ Value Preference Analysis: Reveals mo… view at source ↗
Figure 3
Figure 3. The preference and ranking change of nine LLMs across six value dimensions: care, fairness, authority, sanctity, loyalty, and liberty. The left panels depict the preference scores over five steps (Step 1 to Step 5). Preference scores are determined by the proportion of times a model selects a specific moral dimension relative to the total occurrences at each step, normalized within a range of -0.5 to 0.5. A positive… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Win rates of pairwise comparisons between the six value dimensions from MFT, with a total of 15 dimension pairs. The X-axis represents these dimension pairs (e.g., care vs fairness indicates the win rate of care over fairness). Results are shown for Step 1, Step 5, and…
Figure 5
Figure 5. Figure 5: Preference and ranking scores of various models across ten value dimensions: self-direction, stimulation, hedonism, achievement, power, security, conformity, tradition, benevolence, universalism. The left panels depict preference scores over five steps (Step 1 to Step …
Figure 6
Figure 6. Figure 6: Results are intermediate steps (Steps 2–4), Win rates of pairwise comparisons between the six value dimensions from MFT, with a total of 15 dimension pairs. The X-axis represents these dimension pairs (e.g., care vs fairness indicates the win rate of care over fairness…
Figure 7
Figure 7. Figure 7: Win rates of pairwise comparisons between the ten value dimensions from Schwartz’s Theory of Basic Values, with a total of 45 dimension pairs. The X-axis represents these dimension pairs (e.g., power vs hedonism indicates the win rate of power over hedonism). 24 [PITH…
Figure 8
Figure 8. Figure 8: Screenshots of the Value Dimension Validation Questionnaire 25 [PITH_FULL_IMAGE:figures/full_fig_p025_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

45 extracted references · 27 canonical work pages

  1. [1]

    Marwa Abdulhai, Gregory Serapio-Garcia, Cl \'e ment Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023. Moral foundations of large language models. arXiv preprint arXiv:2310.15337

  2. [2]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. GPT-4 Technical Report . arXiv preprint arXiv:2303.08774

  3. [3]

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Man \'e . 2016. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565

  4. [4]

    Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-Fran c ois Bonnefon, and Iyad Rahwan. 2018. The Moral Machine Experiment . Nature, 563(7729):59--64

  5. [5]

    Albert Bandura. 1999. https://doi.org/10.1207/s15327957pspr0303_3 Moral Disengagement in the Perpetration of Inhumanities . Personality and social psychology review, 3(3):193--209

  6. [6]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610--623

  7. [7]

    Reuben Binns, Max Van Kleek, Michael Veale, Ulrik Lyngs, Jun Zhao, and Nigel Shadbolt. 2018. 'It's Reducing a Human Being to a Percentage' Perceptions of Justice in Algorithmic Decisions. In Proceedings of the 2018 Chi conference on human factors in computing systems, pages 1--14

  8. [8]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1--45

Show all 45 references
  1. [9]

    Yu Ying Chiu, Liwei Jiang, and Yejin Choi. 2025. https://openreview.net/forum?id=PGhiPGBf47 DailyDilemmas: Revealing Value Preferences of LLM s with Quandaries of Daily Life . In The Thirteenth International Conference on Learning Representations

  2. [10]

    Norman Daniels. 2007. Just Health: Meeting Health Needs Fairly . Cambridge University Press

  3. [11]

    Jeffrey Dastin. 2022. Amazon Scraps Secret AI Recruiting Tool that Showed Bias against Women . In Ethics of data and analytics, pages 296--299. Auerbach Publications

  4. [12]

    DeepSeek-AI . 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437

  5. [13]

    Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, and Ning Gu. 2024. https://openreview.net/forum?id=m3RRWWFaVe Denevil: Towards deciphering and navigating the ethical values of large language models via instruction learning . In The Twelfth International Conference on ...

  6. [14]

    Hwang, Maxwell Forbes, and Yejin Choi

    Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.54 Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences . In Proceedings of the 2021 Conference on Empirical Methods...

  7. [15]

    Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.48 Social Chemistry 101: Learning to Reason about Social and Moral Norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language...

  8. [16]

    Batya Friedman, Peter H Kahn, Alan Borning, and Alina Huldtgren. 2013. Value sensitive design and information systems. Early engagement and new technologies: Opening up the laboratory, pages 55--95

  9. [17]

    Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013. https://doi.org/10.1016/B978-0-12-407236-7.00002-4 Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism . In Advances in experimental social psychology, vol...

  10. [18]

    Joshua D Greene, R Brian Sommerville, Leigh E Nystrom, John M Darley, and Jonathan D Cohen. 2001. An fMRI Investigation of Emotional Engagement in Moral Judgment . Science, 293(5537):2105--2108

  11. [19]

    Dorit Hadar-Shoval, Kfir Asraf, Yonathan Mizrachi, Yuval Haber, and Zohar Elyoseph. 2024. Assessing the alignment of large language models with human values for mental health integration: cross-sectional study using Schwartz’s theory of basic values. JMIR Mental Health, 11:e55988

  12. [20]

    Jonathan Haidt. 2013. The Righteous Mind: Why Good People are Divided by Politics and Religion . New York Pantheon, 50:86--88

  13. [21]

    Jonathan Haidt and Jesse Graham. 2007. When Morality Opposes Justice: Conservatives Have Moral Intuitions that Liberals may not Recognize . Social justice research, 20(1):98--116

  14. [22]

    Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. 2025. Values in the wild: Discovering and analyzing values in real-world language model interactions. arXiv preprint arXiv:2504.15236

  15. [23]

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. 2023. AI Alignment: A Comprehensive Survey. CoRR

  16. [24]

    Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al. 2021. Can Machines Learn Morality? The Delphi Experiment . arXiv preprint arXiv:2110.07574

  17. [25]

    Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Sch \"o lkopf. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/b654d6150630a5ba5df7a55621390daf-Paper-Conference.pdf Wh...

  18. [26]

    Shelly Kagan. 2018. Normative ethics. Routledge

  19. [27]

    Bostrom Nick. 2014. Superintelligence: Paths, dangers, strategies . Oxford University Press, Oxford

  20. [28]

    Ritesh Noothigattu, Snehalkumar Gaikwad, Edmond Awad, Sohan Dsouza, Iyad Rahwan, Pradeep Ravikumar, and Ariel Procaccia. 2018. https://doi.org/10.1609/aaai.v32i1.11512 A Voting-Based System for Ethical Decision Making . In Proceedings of the AAAI Conference on Artificial Intel...

  21. [29]

    Joan Plepi, Charles Welch, and Lucie Flek. 2024. Perspective Taking through Generating Responses to Conflict Situations. In Findings of the Association for Computational Linguistics ACL 2024, pages 6482--6497

  22. [30]

    Peter Railton. 2017. Moral learning: Conceptual foundations and normative relevance. Cognition, 167:172--190

  23. [31]

    Shalom H Schwartz. 2012. An Overview of the Schwartz Theory of Basic Values . Online readings in Psychology and Culture, 2(1):11

  24. [32]

    Emily Sheng, Kai - Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. Societal Biases in Language Generation: Progress and Challenges . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natu...

  25. [33]

    Gabriel Simmons. 2023. https://doi.org/10.18653/v1/2023.acl-srw.40 Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: Student Research ...

  26. [34]

    Rafael Souza, Jia-Hao Lim, and Alexander Davis. 2024. Enhancing AI-Driven Psychological Consultation: Layered Prompts with Large Language Models. arXiv preprint arXiv:2408.16276

  27. [35]

    Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.naacl-main.56 On the Machine Learning of Ethical Judgments from Natural Language . In Proceedings of the 2022 Conference of the North America...

  28. [36]

    Judith Jarvis Thomson. 1976. Killing, Letting Die, and the Trolley Problem . The Monist, pages 204--217

  29. [37]

    Eugene Volokh. 2002. https://doi.org/10.2307/1342743 The Mechanisms of the Slippery Slope . Harvard Law Review, 116(4):1026--1137

  30. [38]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models . In Advances in Neural Information Processing Systems, volume 35, pages 24824--24837

  31. [39]

    Jing Yao, Xiaoyuan Yi, Yifan Gong, Xiting Wang, and Xing Xie. 2024. https://doi.org/10.18653/v1/2024.naacl-long.486 Value FULCRA : Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Value . In Proceedings of the 2024 Conference of the North American ...

  32. [40]

    Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Liutao Liutao, and Deyi Xiong. 2024. https://doi.org/10.18653/v1/2024.findings-acl.703 CM oral E val: A Moral Evaluation Benchmark for C hinese Large...

  33. [41]

    Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. SafetyBench: Evaluating the Safety of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguist...

  34. [42]

    Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng. 2023. Rethinking Machine Ethics--Can LLMs Perform Moral Reasoning through the Lens of Moral Theories? arXiv preprint arXiv:2308.15399

  35. [43]

    Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. https://doi.org/10.18653/v1/2022.acl-long.261 The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics...

  36. [44]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  37. [45]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.