REVIEW 3 major objections 6 minor 178 references
LLMs keep the same moral features under noise: counts change, meaning stays put.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 05:43 UTC pith:MQWFWGYE
load-bearing objection Useful invariance method and a real benchmark; the optimistic moral-sensitivity claim is only as strong as embedding similarity can make it. the 3 major comments →
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the MORPH-1K benchmark of one thousand moral vignettes, eight contemporary LLMs, under five classes of morally irrelevant noise, frequently change the number of features they list, yet the semantic content of those features remains stable: mean cosine similarities of about 0.80–0.86 clear per-model empirical floors of about 0.57–0.69 computed on unrelated vignette pairs. Count-level variance coexists with semantic invariance, which the authors read as format and granularity shifting while moral content holds.
What carries the argument
MORPH-1K plus an invariance test: paired clean and perturbed vignettes whose moral structure is held fixed by design and validated; models list morally relevant features; responses are embedded with a fixed sentence embedder and compared by cosine similarity against a per-model floor from randomly paired unrelated cases.
Load-bearing premise
That high cosine similarity between pooled feature-list embeddings is enough to say the model is tracking the same morally relevant features, rather than shared generic moral language, rewording, or a coarse embedder.
What would settle it
A controlled condition or weaker model in which feature lists clearly switch moral content under a validated irrelevant perturbation yet still score above the unrelated-pair floor, or human raters consistently judge high-similarity pairs as different in moral content.
If this is right
- Moral-sensitivity evaluation can scale without fresh human baselines or LLM judges for every new vignette.
- Count of listed features is a poor standalone robustness metric; semantic stability must be checked separately.
- The same clean-versus-irrelevant-noise invariance test can be reused in legal, clinical, and other contested judgment domains.
- Claims of moral competence under noise should distinguish format shifts from genuine content shifts.
Where Pith is reading between the lines
- If the method is adopted, labs can regression-test moral feature stability on every model release without re-running expensive human studies.
- The count-versus-semantics split suggests post-training may be shaping verbosity and list structure more than moral content selection.
- A natural next stress test is multi-turn or multi-severity noise packs that compound distractors until similarity falls to the floor.
- Invariance above the floor still needs occasional quality anchors on clean cases so template-like but stable answers are not mistaken for sensitivity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an invariance-based evaluation of LLM moral sensitivity that avoids human baselines and LLM-as-judge. It introduces MORPH-1K, a stratified 1,000-vignette benchmark spanning 50 Moral Foundations Theory foundation-pole combinations across four social domains, plus three classes of designed-irrelevant noise (textual distractors, irrelevant detail additions, chat histories). Models list morally relevant features on clean and perturbed cases; stability is measured by cosine similarity of Qwen3 Embedding 8B encodings of pooled feature lists, with per-model floors from 500 unrelated vignette pairs. Across eight contemporary LLMs, feature counts often change significantly under noise (Wilcoxon, Bonferroni α=0.0012), but mean similarities remain ≈0.80–0.86, well above floors ≈0.57–0.69. A small-model check (Qwen2.5 0.5B) produces lower scores near the floor. The authors read this as evidence that noise affects format/granularity more than moral content, and argue the framework generalizes to other domains where relevant vs. irrelevant features can be separated by design.
Significance. The methodological target is important: scalable behavioural evaluation of moral sensitivity without contaminated crowd baselines or a judge that presupposes the capability under test. The two-dimensional generation scaffold, eMFD filtering, bidirectional similarity, per-model empirical floors, and discriminative small-model check are concrete contributions that other alignment evaluations can reuse. If the invariance reading holds, the optimistic contrast with recent pessimistic moral-competence results would matter for both evaluation practice and deployment risk assessment. Even if the moral-sensitivity interpretation is only partly secured, the paper still offers a useful middle path between full normative adjudication and purely outcome-based benchmarks, with clear falsifiability conditions (scores at or below the unrelated-pair floor).
major comments (3)
- [§3.2 Semantic Similarity Analysis; §4 Results; §5 Limitations] §3.2 and Results (Figs. 2–3, Table 2): The central claim—that models identify substantially the same morally relevant features under noise—rests on high cosine similarity of pooled feature-list embeddings relative to unrelated-vignette floors. Those floors bound similarity of responses to different moral cases; they do not bound similarity of two responses that share generic moral vocabulary (care, fairness, loyalty, etc.) while missing or swapping vignette-specific content. The paper correctly notes (Discussion, Limitations) that invariance is necessary not sufficient and that a template-like model would score as invariant, but Abstract/Introduction still treat clearance of the floor as evidence of genuine feature identity. A load-bearing addition is needed: e.g., (i) a content-level audit (human or structured rubric) on a stratified subsample comparing clean vs. perturbed feature sets
- [§3.2 Transformations; Morally Irrelevant Details] §3.2 Transformations and validation: Preservation of moral structure under perturbation is the design premise of the invariance test. Textual distractors and chat histories are external insertions; irrelevant-detail and “contextually irrelevant moral features” edits are LLM-generated and accepted mainly via eMFD moral-to-non-moral ratio filters. Expert sequential sampling is reported only for 30 irrelevant-detail vignettes with “almost total agreement.” eMFD is a word-level dictionary and can miss structural shifts (who is harmed, which obligation is at stake). For the claim that response stability tracks moral content rather than surface form, the paper needs either broader dual-expert validation across all noise types (with reported agreement and rejection rates) or an automated structural check beyond eMFD ratios. Currently the strongest stress condition is the least thoroughly valida
- [§3.2 Vignette Generation; Collecting Morally Relevant Features] §3.2 Vignette Generation and model suite: Themes are produced with Claude Opus 4.6; vignettes and noise edits with GPT-5.4; both models (and close relatives) appear in the evaluated set. Generation and evaluation on overlapping model families risks style-matched feature lists that inflate clean–perturbed similarity independent of moral sensitivity. At minimum, report a leave-generator-out analysis (exclude GPT-5.4 / Claude from main tables, or regenerate a held-out slice with a non-evaluated model) and state whether floors and main means change. Without that, the optimistic multi-model result is partly confounded by generator–evaluatee overlap.
minor comments (6)
- [Abstract; §1] Abstract and §1 claim to “address and resolve” the scaling problem for behavioural moral evaluation. The contribution is better framed as a complementary invariance test; “resolve” overstates what necessary-but-not-sufficient stability can deliver.
- [Figure 3; Table 3] Figure 3 caption states “All values in the range above 0.61 empirical floor,” but floors are per-model (Table 3: 0.57–0.69). Align the figure annotation with per-model thresholds used in the text.
- [§4; Appendix H] §4 reports significance of feature-count differences but not direction or effect sizes. Even brief signed rank statistics or median deltas per condition would make the “format vs content” interpretation more testable.
- [§3.2 Semantic Similarity Analysis] Clarify the feature-level similarity procedure in §3.2: “For each feature, we take the maximum score and take the average of all features” is easy to misread relative to “combine all features… then encode the model’s base-case response.” State whether embeddings are of the full concatenated list or of individual features with max-matching.
- [§2; §5 Generalisability] Related Work could briefly situate the invariance idea against robustness/invariance testing outside moral domains (e.g., fairness under demographic noise, clinical decision stability) to strengthen the generalizability claim in §5.
- [§1; References] Typos/consistency: “bemorally competent” spacing (§1); arXiv id and model release dates in the manuscript should be checked against final camera-ready metadata.
Circularity Check
No derivation-by-construction circularity; only mild self-citation for the clean-case baseline premise, not for the invariance measurement itself.
specific steps
-
self citation load bearing
[§1 Introduction (logic of clean baseline + invariance)]
"First, we draw on existing human-baseline work to establish that the model performs well on clean base cases, that is, moral vignettes presented without noise or distraction (Aharoni et al., 2024; Dillion et al., 2023, 2025; Kilov et al., 2025; Scherrer et al., 2023; Chiu et al., 2025). This gives us a well-founded starting point: the model is sensitive to the right moral features when those features are presented clearly. The question then becomes whether that sensitivity is preserved under perturbation"
The optimistic competence reading (not the raw similarity numbers) treats clean-case adequacy as established partly by the authors’ own prior work. That is mild self-citation on the interpretive bridge from invariance to moral sensitivity; it is not load-bearing for the invariance statistics themselves, which are independently measured and do not reduce to that citation by construction.
full rationale
MORPH-1K’s central claim is empirical, not a first-principles derivation: under fixed embeddings, cosine similarity of pooled feature lists between clean and perturbed vignettes exceeds per-model floors estimated from 500 unrelated domain/foundation pairs. Floors are a control, not a fit that forces the main scores (≈0.80–0.86) to clear them. The invariance criterion is a methodological operationalization (necessary but not sufficient, as the paper states), not a self-definitional loop in which the measured quantity is algebraically identical to an input parameter. Vignette generation (Claude themes, GPT-5.4 vignettes) and eMFD filtering shape the stimulus set and are validity/confound concerns, not reductions of the reported similarities to their inputs by construction. The only mild circularity-adjacent element is interpretive: the stronger claim that clean-case competence plus invariance implies noisy-case moral sensitivity leans partly on prior human-baseline work including Kilov et al. (2025), but that premise is multi-cited and is not required for the raw invariance statistics to stand. Score 1 reflects that minor self-citation load on interpretation only; the measurement chain is self-contained against external benchmarks and does not match fitted-input-as-prediction, uniqueness-from-authors, or ansatz-via-self-citation patterns.
Axiom & Free-Parameter Ledger
free parameters (5)
- per-model empirical floor (mean cosine on 500 unrelated pairs)
- eMFD moral-to-non-moral ratio and foundation-probability filters
- stratified cell quota (≈5 vignettes per foundation-combo × domain cell)
- Bonferroni α = 0.0012 for Wilcoxon feature-count tests
- embedding model and pooling rule (Qwen3 Embedding 8B; max-then-average over features; 3 samples)
axioms (6)
- domain assumption Moral Foundations Theory (five foundations × poles) plus four proximal-to-distal social domains form an adequate sampling scaffold for general moral competence.
- ad hoc to paper If two prompts preserve the same morally relevant structure, a morally sensitive model should identify substantially the same features; invariance under designed-irrelevant noise is therefore evidence of moral sensitivity.
- domain assumption Prior human-baseline studies establish that the evaluated models already identify the right features on clean vignettes.
- domain assumption eMFD scores and limited expert review can validate that distractors and detail additions do not change morally salient content.
- domain assumption Sentence-embedding cosine similarity is a valid, sufficiently discriminative measure of semantic equivalence of moral feature lists.
- standard math Standard statistical comparisons (Wilcoxon signed-rank, Bonferroni) and distributional semantics (Sentence Transformers) are appropriate tools for these response pairs.
invented entities (2)
-
MORPH-1K benchmark (1,000 procedurally generated foundation×domain vignettes with noise variants)
no independent evidence
-
Invariance-based moral-sensitivity evaluation (clean vs perturbed semantic stability without human/LLM judge)
no independent evidence
read the original abstract
Moral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. In this paper, we offer a new evaluation of LLM moral sensitivity and in doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensive and sometimes metaethically dubious comparisons with a human baseline, without adopting an LLM judge that must be assumed to have the very capability that you are attempting to evaluate. Our central question is this: can LLMs successfully identify the morally relevant features of noisy cases, in which various kinds of morally irrelevant information have been introduced to distract the respondent? To explore this, we introduce \textbf{MORPH-1K (MOral Robustness under Perturbed Hypotheticals)}, a procedurally-generated 1,000-case benchmark spanning 50 moral foundation-pole combinations across four social domains. MORPH-1K is paired with a suite of textual noise elements, along with a method for validating that the distractors do not change the morally salient content of the case. We apply MORPH-1K to eight contemporary LLMs, and show that while morally irrelevant perturbations often changed the number of features listed, the semantic content of those features remained stable across all noise conditions, with similarity scores above our calibrated floor threshold. More broadly, our invariance framework extends to evaluative domains where ground truth is difficult to specify but relevant and irrelevant features can be separated by design.
Figures
Reference graph
Works this paper leans on
-
[1]
Shaw, Andrew and Hahn, Christina and Rasgaitis, Catherine and Mishra, Yash and Liu, Alisa and Jaques, Natasha and Tsvetkov, Yulia and Zhang, Amy X. , date =. Are Language Models Sensitive to Morally Irrelevant Distractors? , url =. 2026 , keywords =. doi:10.48550/arXiv.2602.09416 , abstract =
-
[2]
Science , volume =
Performance of a large language model on the reasoning tasks of a physician , author =. Science , volume =. 2026 , month = apr, doi =
2026
-
[3]
2026 , month = apr, url =
How people ask Claude for personal guidance , author =. 2026 , month = apr, url =
2026
-
[4]
2025 , url =
Miles McCain and Ryn Linthicum and Chloe Lubinski and Alex Tamkin and Saffron Huang and Michael Stern and Kunal Handa and Esin Durmus and Tyler Neylon and Stuart Ritchie and Kamya Jagadish and Paruul Maheshwary and Sarah Heck and Alexandra Sanderford and Deep Ganguli , title =. 2025 , url =
2025
-
[5]
International Journal of Law in Context , year =
Terzidou, Kalliopi , title =. International Journal of Law in Context , year =. doi:10.1017/S1744552325000047 , url =
-
[6]
2026 , month = apr, isbn =
2026
-
[7]
Franco, Mirko and Gaggi, Ombretta and Palazzi, Claudio E. , title =. ACM Trans. Web , month = may, articleno =. 2025 , issue_date =. doi:10.1145/3700789 , abstract =
doi:10.1145/3700789 2025
-
[8]
2025 , doi =
General practitioners' adoption of generative artificial intelligence in clinical practice in the UK: An updated online survey , author =. 2025 , doi =
2025
-
[9]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Beyond Verdicts: Evaluating Language Model Moral Competence , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , month = mar, doi =
2026
-
[10]
2026 , month = feb, url =
Janjeva, Ardi and Ashurst, Carolyn and Hennessy, Rick , title =. 2026 , month = feb, url =
2026
-
[11]
Nature , volume =
A roadmap for evaluating moral competence in large language models , author =. Nature , volume =. 2026 , month = feb, doi =
2026
-
[12]
Neuro-Symbolic Models of Human Moral Judgment:
Kwon, Joe and Tenenbaum, Josh and Levine, Sydney , booktitle =. Neuro-Symbolic Models of Human Moral Judgment:. 2023 , url =
2023
-
[13]
Agley, Jon , date =. Planning for New Threats to Online Research Data Validity: The Issue of Computer-Using Agents , issn =. Evaluation & the Health Professions , publisher =. doi:10.1177/01632787251367407 , abstract =
-
[14]
Crimston, Charlie R. and Bain, Paul G. and Hornsey, Matthew J. and Bastian, Brock , date =. Moral expansiveness: Examining variability in the extension of the moral world , volume =. Journal of Personality and Social Psychology , publisher =. 2016 , keywords =. doi:10.1037/pspp0000086 , abstract =
-
[15]
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , url =
Burns, Collin and Izmailov, Pavel and Kirchner, Jan Hendrik and Baker, Bowen and Gao, Leo and Aschenbrenner, Leopold and Chen, Yining and Ecoffet, Adrien and Joglekar, Manas and Leike, Jan and Sutskever, Ilya and Wu, Jeff , date =. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , url =. 2023 , keywords =. doi:10.48550/a...
-
[16]
, date =
Firth, J.R. , date =. A Synopsis of Linguistic Theory, 1930-1955 , url =
1930
-
[17]
Applied Economic Perspectives and Policy , author =
Battling bots: Experiences and strategies to mitigate fraudulent responses in online surveys , volume =. Applied Economic Perspectives and Policy , author =. 2023 , langid =. doi:10.1002/aepp.13353 , abstract =
-
[18]
Harris, Zellig S. , date =. Distributional Structure , volume =. 1954 , note =. doi:10.1080/00437956.1954.11659520 , pages =
-
[19]
Behavior Research Methods , author =
The extended Moral Foundations Dictionary (. Behavior Research Methods , author =. 2021 , langid =. doi:10.3758/s13428-020-01433-0 , shorttitle =
-
[20]
2018 , langid =
Irving, Geoffrey and Christiano, Paul and Amodei, Dario , date =. 2018 , langid =
2018
-
[21]
and Lazar, Seth , date =
Kilov, Daniel and Hendy, Caroline and Yanik Guyot, Secil and Snoswell, Aaron J. and Lazar, Seth , date =. Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in. 2025 , langid =
2025
-
[22]
Essays on moral development: Vol
Kohlberg, Lawrence , date =. Essays on moral development: Vol. 1. The philosophy of moral development , isbn =
-
[23]
Distributed Representations of Words and Phrases and their Compositionality , volume =
Mikolov, Tomas and Sutskever, Ilya and Chen, Kai and Corrado, Greg S and Dean, Jeff , date =. Distributed Representations of Words and Phrases and their Compositionality , volume =. Advances in Neural Information Processing Systems , publisher =
-
[24]
Panizza, Folco and Kyrychenko, Yara and Roozenbeek, Jon , date =. Survey-taking. Nature , publisher =. 2026 , langid =. doi:10.1038/d41586-026-00386-2 , abstract =
-
[25]
Recognising, Anticipating, and Mitigating
Rilla, Raluca and Werner, Tobias and Yakura, Hiromu and Rahwan, Iyad and Nussberger, Anne-Marie , date =. Recognising, Anticipating, and Mitigating. 2025 , keywords =. doi:10.48550/arXiv.2508.01390 , abstract =
-
[26]
The Expanding Circle: Ethics and Sociobiology , isbn =
Singer, Peter , date =. The Expanding Circle: Ethics and Sociobiology , isbn =
-
[27]
Venugopalan, Hari and Munir, Shaoor and Ahmed, Shuaib and Wang, Tangbaihe and King, Samuel T. and Shafiq, Zubair , date =. Proceedings of the 2025. doi:10.1145/3730567.3732919 , series =
-
[28]
and Gordon, Andrew and Rothschild, David and West, Robert , date =
Veselovsky, Veniamin and Ribeiro, Manoel Horta and Cozzolino, Philip J. and Gordon, Andrew and Rothschild, David and West, Robert , date =. Prevalence and Prevention of Large Language Model Use in Crowd Work – Communications of the. 2025 , langid =
2025
-
[29]
Westwood, Sean J. , date =. The potential existential threat of large language models to online survey research , volume =. Proceedings of the National Academy of Sciences , publisher =. doi:10.1073/pnas.2518075122 , abstract =
-
[30]
and Zhang, Xiangliang , date =
Ye, Jiayi and Wang, Yanbo and Huang, Yue and Chen, Dongping and Zhang, Qihui and Moniz, Nuno and Gao, Tian and Geyer, Werner and Huang, Chao and Chen, Pin-Yu and Chawla, Nitesh V. and Zhang, Xiangliang , date =. Justice or Prejudice? Quantifying Biases in. 2024 , keywords =. doi:10.48550/arXiv.2410.02736 , abstract =
-
[31]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , date =. Judging. 2023 , keywords =. doi:10.48550/arXiv.2306.05685 , series =
-
[33]
Dillion, Danica and Tandon, Niket and Gu, Yuling and Gray, Kurt , date =. Can. Trends in Cognitive Sciences , publisher =. 2023 , keywords =. doi:10.1016/j.tics.2023.04.008 , pages =
-
[34]
Scientific Reports , publisher =
Dillion, Danica and Mondal, Debanjan and Tandon, Niket and Gray, Kurt , date =. Scientific Reports , publisher =. 2025 , langid =. doi:10.1038/s41598-025-86510-0 , abstract =
-
[35]
2023 , title =
Scherrer, Nino and Shi, Claudia and Feder, Amir and Blei, David M , langid =. 2023 , title =
2023
-
[36]
Moral Foundations of Large Language Models , url =
Abdulhai, Marwa and Serapio-García, Gregory and Crepy, Clement and Valter, Daria and Canny, John and Jaques, Natasha , editor =. Moral Foundations of Large Language Models , url =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =. doi:10.18653/v1/2024.emnlp-main.982 , abstract =
-
[38]
Differences in the Moral Foundations of Large Language Models , url =
Kirgis, Peter , date =. Differences in the Moral Foundations of Large Language Models , url =. 2025 , keywords =. doi:10.48550/arXiv.2511.11790 , abstract =
-
[39]
2026 , langid =
Gemini 3.1 Pro Preview , url =. 2026 , langid =
2026
-
[40]
2026 , langid =
Claude Opus 4.6 , url =. 2026 , langid =
2026
-
[41]
2026 , langid =
Grok 4.20 , url =. 2026 , langid =
2026
-
[42]
2026 , langid =
Qwen3.6 Plus , url =. 2026 , langid =
2026
-
[43]
2026 , langid =
Nemotron 3 Super , url =. 2026 , langid =
2026
-
[44]
2026 , langid =
Gemma 4 31B , url =. 2026 , langid =
2026
-
[45]
2023 , langid =
Zhao, Wenting and Ren, Xiang and Hessel, Jack and Cardie, Claire and Choi, Yejin and Deng, Yuntian , date =. 2023 , langid =
2023
-
[46]
Bavaresco, Anna and Bernardi, Raffaella and Bertolazzi, Leonardo and Elliott, Desmond and Fernández, Raquel and Gatt, Albert and Ghaleb, Esam and Giulianelli, Mario and Hanna, Michael and Koller, Alexander and Martins, Andre and Mondorf, Philipp and Neplenbroek, Vera and Pezzelle, Sandro and Plank, Barbara and Schlangen, David and Suglia, Alessandro and S...
-
[47]
Chiang, Cheng-Han and Chen, Wei-Chih and Kuan, Chun-Yi and Yang, Chienchou and Lee, Hung-yi , editor =. Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course , url =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , publisher =. doi:10.18653/v1/2024.emnlp-main....
-
[48]
Huang, Hui and Bu, Xingyuan and Zhou, Hongli and Qu, Yingqi and Liu, Jing and Yang, Muyun and Xu, Bing and Zhao, Tiejun , editor =. An Empirical Study of. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-acl.306 , shorttitle =
-
[49]
Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =
Kocmi, Tom and Federmann, Christian , editor =. Large Language Models Are State-of-the-Art Evaluators of Translation Quality , url =. Proceedings of the 24th Annual Conference of the European Association for Machine Translation , publisher =
-
[50]
Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =
Koo, Ryan and Lee, Minhwa and Raheja, Vipul and Park, Jong Inn and Kim, Zae Myung and Kang, Dongyeop , editor =. Benchmarking Cognitive Biases in Large Language Models as Evaluators , url =. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2024.findings-acl.29 , abstract =
-
[51]
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , editor =. G-Eval:. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , publisher =. doi:10.18653/v1/2023.emnlp-main.153 , abstract =
-
[52]
Assistant-Guided Mitigation of Teacher Preference Bias in
Liu, Zhuo and Li, Moxin and Deng, Xun and Wang, Qifan and Feng, Fuli , editor =. Assistant-Guided Mitigation of Teacher Preference Bias in. Findings of the Association for Computational Linguistics:. doi:10.18653/v1/2025.findings-emnlp.510 , abstract =
-
[53]
Large Language Models are not Fair Evaluators , url =
Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Kong, Lingpeng and Liu, Qi and Liu, Tianyu and Sui, Zhifang , editor =. Large Language Models are not Fair Evaluators , url =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher...
-
[54]
Xu, Wenda and Zhu, Guanglei and Zhao, Xuandong and Pan, Liangming and Li, Lei and Wang, William , editor =. Pride and Prejudice:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , publisher =. doi:10.18653/v1/2024.acl-long.826 , abstract =
-
[55]
Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , date =. Aligning. 2023 , keywords =. doi:10.48550/arXiv.2008.02275 , abstract =
-
[56]
Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =
Curry, Oliver Scott and Mullins, Daniel Austin and Whitehouse, Harvey , date =. Is It Good to Cooperate?: Testing the Theory of Morality-as-Cooperation in 60 Societies , volume =. Current Anthropology , publisher =. doi:10.1086/701478 , abstract =
-
[57]
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models , url =
Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren , date =. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models , url =. 2025 , keywords =. doi:10.48550/arXiv.2506.05176 , abstract =
-
[58]
Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =
Haidt, Jonathan and Joseph, Craig , year =. Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues , booktitle =
-
[59]
and Ditto, Peter H
Graham, Jesse and Haidt, Jonathan and Koleva, Sena and Motyl, Matt and Iyer, Ravi and Wojcik, Sean P. and Ditto, Peter H. , year =. Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism , journal =
-
[60]
Sentence-. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , author =. 2019 , pages =
2019
-
[61]
Kaakinen, Johanna K. and Werlen, Egon and Kammerer, Yvonne and Acartürk, Cengiz and Aparicio, Xavier and Baccino, Thierry and Ballenghein, Ugo and Bergamin, Per and Castells, Núria and Costa, Armanda and Falé, Isabel and Mégalakaki, Olga and Fernández, Susana Ruiz , urldate =. 2022 , date =. doi:10.1371/journal.pone.0274480 , shorttitle =
-
[62]
Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer. Beyond Accuracy: Behavioral Testing of NLP Models with C heck L ist. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.442
-
[63]
Measuring and Improving Consistency in Pretrained Language Models
Elazar, Yanai and Kassner, Nora and Ravfogel, Shauli and Ravichander, Abhilasha and Hovy, Eduard and Sch. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics. 2021. doi:10.1162/tacl_a_00410
-
[64]
2025 , eprint=
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes , author=. 2025 , eprint=
2025
-
[65]
Qwen2.5: A Party of Foundation Models , url =
Qwen , month =. Qwen2.5: A Party of Foundation Models , url =
-
[66]
Aharoni, Eyal and Fernandes, Sharlene and Brady, Daniel J. and Alexander, Caelan and Criner, Michael and Queen, Kara and Rando, Javier and Nahmias, Eddy and Crespo, Victor , year=. Attributions toward artificial agents in a modified Moral Turing Test , volume=. Scientific Reports , publisher=. doi:10.1038/s41598-024-58087-7 , number=
-
[67]
Large-scale moral machine experiment on large language models , volume=
Zaim bin Ahmad, Muhammad Shahrul and Takemoto, Kazuhiro , editor=. Large-scale moral machine experiment on large language models , volume=. PLOS One , publisher=. 2025 , month=May, pages=. doi:10.1371/journal.pone.0322776 , number=
-
[68]
Concrete Problems in AI Safety , author =. 2016 , eprint =. doi:10.48550/arXiv.1606.06565 , url =
-
[69]
and Goldstein, Simon and Salib, Peter , year=
Arbel, Yonathan A. and Goldstein, Simon and Salib, Peter , year=. How to Count AIs: Individuation and Liability for AI Agents , url=. doi:10.2139/ssrn.6273198 , publisher=
-
[70]
2026 , eprint =
Trust as Monitoring: Evolutionary Dynamics of User Trust and AI Developer Behaviour , author =. 2026 , eprint =
2026
-
[71]
Baum, Seth D. , year=. Social choice ethics in artificial intelligence , volume=. AI & SOCIETY , publisher=. doi:10.1007/s00146-017-0760-1 , number=
-
[72]
Algorithmic Accountability and Public Reason , volume=
Binns, Reuben , year=. Algorithmic Accountability and Public Reason , volume=. Philosophy & Technology , publisher=. doi:10.1007/s13347-017-0263-5 , number=
-
[73]
AI Consciousness: A Centrist Manifesto , url=
Birch, Jonathan , year=. AI Consciousness: A Centrist Manifesto , url=. doi:10.31234/osf.io/af7c9_v1 , publisher=
-
[74]
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
Boerstler, Kyle and Keswani, Vijay and Chan, Lok and Borg, Jana Schaich and Conitzer, Vincent and Heidari, Hoda and Sinnott-Armstrong, Walter , keywords =. On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2408.02862 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2408.02862 2024
-
[75]
SaGE: Evaluating Moral Consistency in Large Language Models
Bonagiri, Vamshi Krishna and Vennam, Sreeram and Govil, Priyanshul and Kumaraguru, Ponnurangam and Gaur, Manas , keywords =. SaGE: Evaluating Moral Consistency in Large Language Models , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2402.13709 , url =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2402.13709 2024
-
[76]
2026 , eprint =
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment , author =. 2026 , eprint =
2026
-
[77]
Manipulating the Perceived Personality Traits of Language Models , url=
Caron, Graham and Srivastava, Shashank , year=. Manipulating the Perceived Personality Traits of Language Models , url=. doi:10.18653/v1/2023.findings-emnlp.156 , booktitle=
-
[78]
Harms from Increasingly Agentic Algorithmic Systems , url=
Chan, Alan and Salganik, Rebecca and Markelius, Alva and Pang, Chris and Rajkumar, Nitarshan and Krasheninnikov, Dmitrii and Langosco, Lauro and He, Zhonghao and Duan, Yawen and Carroll, Micah and Lin, Michelle and Mayhew, Alex and Collins, Katherine and Molamohammadi, Maryam and Burden, John and Zhao, Wanru and Rismani, Shalaleh and Voudouris, Konstantin...
-
[79]
From Persona to Personalization: A Survey on Role-Playing Language Agents , publisher =
Chen, Jiangjie and Wang, Xintao and Xu, Rui and Yuan, Siyu and Zhang, Yikai and Shi, Wei and Xie, Jian and Li, Shuang and Yang, Ruihan and Zhu, Tinghui and Chen, Aili and Li, Nianqi and Chen, Lida and Hu, Caiyu and Wu, Siye and Ren, Scott and Fu, Ziquan and Xiao, Yanghua , keywords =. From Persona to Personalization: A Survey on Role-Playing Language Agen...
-
[80]
Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =
Chen, Runjin and Arditi, Andy and Sleight, Henry and Evans, Owain and Lindsey, Jack , keywords =. Persona Vectors: Monitoring and Controlling Character Traits in Language Models , publisher =. 2025 , copyright =. doi:10.48550/ARXIV.2507.21509 , url =
-
[81]
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =
Chiu, Yu Ying and Jiang, Liwei and Choi, Yejin , keywords =. DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life , publisher =. 2024 , copyright =. doi:10.48550/ARXIV.2410.02683 , url =
-
[82]
Chiu, Yu Ying and Lee, Michael S. and Calcott, Rachel and Handoko, Brandon and de Font-Reaulx, Paul and Rodriguez, Paula and Zhang, Chen Bo Calvin and Han, Ziwen and Sehwag, Udari Madhushani and Maurya, Yash and Knight, Christina Q and Lloyd, Harry R. and Bacus, Florence and Mazeika, Mantas and Liu, Bing and Choi, Yejin and Gordon, Mitchell L and Levine, ...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.