REVIEW 3 major objections 5 minor 117 references
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read The yes-no bias of large language models on moral dilemmas is an artifact of answer order and the word “no,” not a change in moral judgment.
desk verdict Clean factorial decomposition of the yes-no bias into order + lexical surface pulls, with logical attachment ~0 under label swap and a coherent graded stance as independent axis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Crossed symmetrization: every logically irrelevant factor (verb, printed order, answer label, scale, anchor, wording, pole) is flipped in balanced pairs and the factors are crossed so their contributions separate. The symmetric part of each flip-pair estimates stance θ; the antisymmetric part is the artifact. The minimal model P = σ((θ ± m)/s) then summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, distinct from sampling temperature.
What would settle it
If the same crossed battery showed frontier models’ graded stance itself swinging as much as the reported binary artifact when only the rating question’s surface form changed, or if fully arbitrary answer labels still produced a large verdict-attached bias, the claim that the scale is format-invariant and the bias is surface-only would fail.
Extended reading notes
Core claim
Frontier models carry a coherent internal moral scale: graded ratings of the same dilemmas stay nearly format-invariant under crossed, logically equivalent framings. Forcing the judgment through yes/no overlays a decomposable format artifact—an order bias toward the last-printed option plus a lexical pull toward the word “no”—large mainly in Claude models and smaller under extended reasoning. With arbitrary answer labels the verdict-attached logical bias is approximately zero for every frontier model; the pull follows the printed surface, not the verdict it carries.
Load-bearing premise
That the averaged graded rating under many equivalent framings really is the model’s stable moral stance, so it can serve as the fixed axis against which yes/no artifacts are measured.
Editorial extensions
If this is right
- Graded, multi-frame elicitation recovers a model’s moral stance more cleanly than any single forced yes/no.
- Single-format binary readouts confound stance with surface format and should not be read as direct measures of value.
- Extended reasoning typically shrinks both cross-form incoherence and framing susceptibility where the artifact is large.
- The same battery applies unchanged to any dilemma set and binary format, so the decomposition can be rerun on other value domains.
- A yes-no bias reported without crossing verb, order, and label cannot be attributed to moral judgment versus surface pull.
Reading between the lines
- Safety gates and LLM-as-judge systems that force yes/no may be measuring label and order attachments rather than the intended policy.
- The recency-type order bias (opposite classic human primacy) is likely to appear in many multiple-choice readouts, not only moral items.
- Interventions that claim to reduce format sensitivity can be audited by tracking m and s rather than raw accuracy alone.
- Small models can look “coherent” by being indifferent; any coherence number without a discrimination check can flatter an empty scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the amplified yes–no bias of LLMs on moral dilemmas is a format artifact of the binary readout, not a shift in moral judgment. Using a crossed-symmetrization psychometric battery (graded ratings I-1, free choice I-2, forced binary I-3) on twenty dilemmas largely from Cheung et al., it recovers a nearly format-invariant graded stance θ for frontier models (cross-form incoherence σ_repro = 0.12–0.21). The forced yes/no readout decomposes, by the verb × order identity (Eq. 1), into an order bias toward the last-printed option plus a lexical pull toward the word “no,” concentrated in Claude models and shrinking under extended reasoning. Swapping yes/no for arbitrary labels (A/B) drives the verdict-attached logical component to ≈0 for frontier models, while surface label and order attachments remain. A logistic summary P = σ((θ ± m)/s) yields portable framing susceptibility m and moral decisiveness s.
Significance. If the result holds, it reframes a growing literature on LLM moral and survey biases: single-framing yes/no verdicts confound stance with surface form, and evaluations that ask once systematically misread format artifacts as value shifts. The methodological contribution—crossed symmetrization that separates order, lexical, and logical channels, with independent graded θ as the stance axis—is portable beyond these dilemmas and is a genuine advance over uncrossed multi-prompt consistency rates. Strengths include a definitional decomposition (Eq. 1), a clean label-swap identification, pre-registered exclusion of bipolar cells, matched 12-form baselines, salt replications establishing deterministic open-weight incoherence, convergent free-choice checks, and an explicit, reproducible analysis pipeline. The (m, s) parameterization and the demonstration that deliberation shrinks |m| are useful, falsifiable summaries for future work.
major comments (3)
- [Abstract; Results “Lexical, not logical”; Fig. 5] Abstract and Results (“Lexical, not logical” / Fig. 5): the abstract states that the verdict-attached logical bias “proves ≈0 for every frontier model.” Methods and Fig. 5 correctly qualify that six of seven A/B logic CIs span zero (exception −0.02) and that Haiku’s logic channels are wide (±0.2–0.3) and underpowered, so the claim is a bound there. Align the abstract and significance statement with that qualification; “≈0 where precisely measured; a bound for Haiku” is what the data support.
- [Materials and Methods (I-3, refusal handling); Results Fig. 2] Methods I-3 / refusal handling: Haiku refuses on 28%/35% of core verb-flip trials and Flash-Lite withholds on 24%. Bias estimates condition on non-refusal after cell exclusion. The paper shows family concentration is not a pure refusal artifact (Flash-Lite ≈0 bias despite high withholding), but does not report whether refusal rates differ systematically by verb, printed order, or label. Differential refusal by frame would make exclusion non-ignorable for the order/lexical split. Please report refusal rates by the crossed factors (or a sensitivity analysis that bounds the bias under plausible missingness) for the high-refusal configurations.
- [Results “An internal moral scale exists”; Methods I-2] Results I-2 convergent validity (r = 0.82–0.90 with graded θ): free-choice verdicts are extracted by Claude Opus 4.8, which shares a vendor family with several subjects. The authors flag circularity risk and ship transcripts, but the main-text correlations are presented as primary convergent evidence for the latent scale. Either re-extract a subset with an independent judge (or human coding) and report agreement, or move the I-2 correlations to a clearly caveated secondary check so the θ claim does not rest on same-family judging.
minor comments (5)
- [Fig. 1b caption] Fig. 1b: variance-share bars (anchor-direction solid, scale hatched) combine in quadrature; a one-line reminder in the caption that shares are not linearly additive would prevent misreading stacked heights.
- [Methods; throughout Results] Notation: b, z, θ, m, s, and σ_repro are introduced across Results and Methods; a short symbol table in Methods or SI would help readers track the [−1, +1] convention and the distinction between descriptive incoherence spreads and bootstrap CIs.
- [Results “The standard readout overlays a format artifact”; Methods “Cheung-verbatim control”] The Cheung-verbatim K01 control (order-balanced residual +0.01, CI spanning zero) is important for engaging the source finding; consider elevating one sentence of it into the main Results paragraph on decomposition rather than leaving it mostly in Methods.
- [Results opening; Discussion scopes] SI Appendix is heavily referenced for open-weight degeneracy (Nemotron), salt floors, and G coefficients; ensure the main text’s “small open-weight models fail in model-specific ways” is self-contained enough that a reader who skips SI still sees the discrimination-vs-coherence distinction (G vs raw σ_repro).
- [Introduction; Methods model panel] Typos / polish: “theyes–no bias” spacing in the Introduction; occasional missing spaces after em-dashes in the compiled text; confirm that “GPT-5.5” and model snapshot IDs are the intended public names at submission.
Circularity Check
No significant circularity: independent instruments, definitional crossing identity used transparently as measurement, and fitted (m,s) as descriptive summary rather than forced prediction.
full rationale
The paper's load-bearing chain does not reduce inputs to outputs by construction. Stance θ is recovered from the graded instrument I-1 (24 unipolar conditions; no yes/no labels or printed order), while binary artifacts are measured on a separate forced-binary instrument I-3; free-choice I-2 supplies an independent convergent check (r = 0.82–0.90). The identity b_apparent = b_order + b_lexical is stated as exact and definitional from the verb×order crossing (Eq. 1), which is standard factorial measurement, not a disguised prediction: the scientific content is that the summands are separately measurable and that the verdict-attached (logical) component collapses to ≈0 under fully arbitrary A/B labels—an empirical result of the label-swap design, not forced by the yes/no token identity. The logistic P = σ((θ ± m)/s) is a fitted descriptive summary of observed flip-pairs (bowtie/ridge geometry), with θ as an external regressor from I-1; m and s are not inserted into the central claim by definition, and the paper flags errors-in-variables attenuation and clip-limited s. No self-citation chain, uniqueness theorem, or ansatz from the same author underwrites the result. Materials are from Cheung et al.; the decomposition and logical-null are new measurements on those materials. Honest non-finding: score 0.
Assumptions & free parameters
free parameters (3)
- framing susceptibility m (per bias channel and model configuration)
- moral decisiveness s (per bias channel and model configuration)
- stance heta per dilemma
assumptions (5)
- domain assumption Logically equivalent but operationally independent graded elicitations that agree recover a latent stance (convergent validity).
- standard math Canonical logistic link p = σ(( heta ± m)/s) from continuous latent to binary choice.
- standard math Balanced 2-vs-2 splits of the four verb imes order cells separate order from lexical contributions exactly (Eq. 1).
- domain assumption Arbitrary answer labels (A/B) carry the verdict without carrying English yes/no lexical content, so the logic projection isolates verdict attachment.
- domain assumption Sampling-corrected between-form variance σ_repro is the dominant uncertainty and the proper error bar on stance.
invented entities (2)
-
framing susceptibility m and moral decisiveness s
-
crossed-symmetrization psychometric battery (I-1/I-2/I-3)
independent evidence
Cite this review
Pith. "Pith review of The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment." pith.science (2026). https://pith.science/paper/VWRNES4Y
@misc{pith2026260705552,
author = {Pith},
title = {Pith review of: The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment},
year = {2026},
howpublished = {\url{https://pith.science/paper/VWRNES4Y}},
note = {Machine review of arXiv:2607.05552}
}
abstract
Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = \sigma((\theta \pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
The framing of decisions and the psychology of choice.Science, 211(4481):453–458, 1981
Amos Tversky and Daniel Kahneman. The framing of decisions and the psychology of choice.Science, 211(4481):453–458, 1981. doi: 10.1126/science.7455683
-
[2]
Choices, values, and frames.American Psychologist, 39(4):341–350, 1984
Daniel Kahneman and Amos Tversky. Choices, values, and frames.American Psychologist, 39(4):341–350, 1984. doi: 10. 1037/0003-066X.39.4.341
1984
-
[3]
Irwin P. Levin, Sandra L. Schneider, and Gary J. Gaeth. All frames are not created equal: A typology and critical analysis of framing effects.Organizational Behavior and Human Decision Processes, 76(2):149–188, 1998. doi: 10.1006/obhd.1998.2804
-
[4]
A systematic review of risky-choice framing effects.EXCLI Journal, 22:1012–1031, 2023
Anton K¨ uhberger. A systematic review of risky-choice framing effects.EXCLI Journal, 22:1012–1031, 2023. doi: 10.17179/ excli2023-6169
2023
-
[5]
Lewis Petrinovich and Patricia O’Neill. Influence of wording and framing effects on moral intuitions.Ethology and Sociobiology, 17(3):145–171, 1996. doi: 10.1016/0162-3095(96)00041-6
-
[6]
Order effects in moral judgment.Philosophical Psychology, 25(6):813–836,
Alex Wiegmann, Yasmina Okan, and Jonas Nagel. Order effects in moral judgment.Philosophical Psychology, 25(6):813–836,
-
[7]
doi: 10.1080/09515089.2011.631995
-
[8]
Eric Schwitzgebel and Fiery Cushman. Expertise in moral reasoning? order effects on moral judgment in professional philosophers and non-philosophers.Mind & Language, 27(2): 135–153, 2012. doi: 10.1111/j.1468-0017.2012.01438.x
Show all 117 references
-
[9]
Philosophers’ biased judgments persist despite training, expertise and reflection.Cog- nition, 141:127–137, 2015
Eric Schwitzgebel and Fiery Cushman. Philosophers’ biased judgments persist despite training, expertise and reflection.Cog- nition, 141:127–137, 2015. doi: 10.1016/j.cognition.2015.04.015
2015 doi
-
[10]
Large lan- guage models show amplified cognitive biases in moral decision- making.Proc
Vanessa Cheung, Maximilian Maier, and Falk Lieder. Large lan- guage models show amplified cognitive biases in moral decision- making.Proc. Natl. Acad. Sci. U.S.A., 122(25):e2412015122,
-
[11]
doi: 10.1073/pnas.2412015122
-
[12]
Robustness of large language models in moral judgements.Royal Society Open Science, 12 (4):241229, 2025
Soyoung Oh and Vera Demberg. Robustness of large language models in moral judgements.Royal Society Open Science, 12 (4):241229, 2025. doi: 10.1098/rsos.241229
2025 doi
-
[13]
The greatest good benchmark: Measuring LLMs’ alignment with utilitarian moral dilemmas
Giovanni Franco Gabriel Marraffini, Andr´ es Cotton, No´ e Fabi´ an Hsueh, Axel Fridman, Juan Wisznia, and Luciano del Corro. The greatest good benchmark: Measuring LLMs’ alignment with utilitarian moral dilemmas. InProceedings of the 2024 Conference on Empirical Methods in Na...
2024 doi
-
[14]
Murukannaiah, and Munindar P
Jiaqing Yuan, Pradeep K. Murukannaiah, and Munindar P. Singh. Right vs. right: Can LLMs make tough choices?, 2024
2024
-
[15]
Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei. Evaluating the moral beliefs encoded in LLMs. In37th Con- ference on Neural Information Processing Systems (NeurIPS 2023), 2023
2023
-
[16]
Varshney
Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, and Lav R. Varshney. Many LLMs are more utilitarian than one. In39th Conference on Neural Information Processing Systems (NeurIPS 2025), 2025
2025
-
[17]
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. InInternational Conference on Learning Representations (ICLR), 2024
2024
-
[18]
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. Large language models sensitivity to the order of options in multiple-choice questions. In Kevin Duh, Helena Gomez, and Steven Bethard, editors,Findings of the Association for Computational Linguis- tics: NAACL 2024, pages 2006–2017, Mex...
2024
-
[19]
doi: 10.18653/ v1/2024.findings-naacl.130
Association for Computational Linguistics. doi: 10.18653/ v1/2024.findings-naacl.130. URL https://aclanthology.org/ 2024.findings-naacl.130/
2024
-
[20]
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large language models are not robust multiple choice selectors. InInternational Conference on Learning Representa- tions (ICLR), 2024
2024
-
[21]
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-D¨ unner. Questioning the survey responses of large language models. InAdvances in Neural Information Processing Systems (NeurIPS), volume 37, 2024
2024
-
[22]
Do LLMs exhibit human-like response biases? a case study in survey design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. Do LLMs exhibit human-like response biases? a case study in survey design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024. doi: 10.1162/tacl a 00685
2024 doi
-
[23]
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models
Paul R¨ ottger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schuetze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual Me...
2024 doi
-
[24]
State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024. doi: 10.1162/ tacl a 00681
2024
-
[25]
Promptrobust: Towards evaluating the 12 robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zeek Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, and Xing Xie. Promptrobust: Towards evaluating the 12 robustness of large language models on adversarial prompts. In Proceedings of the 1st ACM Worksho...
2024 doi
-
[26]
In-context impersonation reveals large lan- guage models’ strengths and biases
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. In-context impersonation reveals large lan- guage models’ strengths and biases. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023
2023
-
[27]
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning AI with shared human values. InInternational Conference on Learning Representations (ICLR 2021), 2021
2021
-
[28]
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Et- zioni, Maarten Sap, Regina Rini, and Yejin Choi. Can machines learn morality? the delphi experim...
2021
-
[29]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. InAd- vances in Neural Information Proces...
2023
-
[30]
Whose opinions do language models reflect? InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 29971–30004, 2023
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 29971–30004, 2023
2023
-
[31]
The effect of sampling temper- ature on problem solving in large language models
Matthew Renze and Erhan Guven. The effect of sampling temper- ature on problem solving in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 7346–7356, 2024. doi: 10.18653/v1/2024.findings-emnlp. 432
2024 doi
-
[32]
The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human...
2025
-
[33]
Hashimoto, and Tobias Gerstenberg
Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B. Hashimoto, and Tobias Gerstenberg. MoCa: Mea- suring human-language model alignment on causal and moral judgment tasks. InAdvances in Neural Information Processing Systems, volume 36, pages 78360–78393...
2023
-
[34]
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garc´ ıa, Clement Crepy, Daria Valter, John Canny, and Natasha Jaques. Moral foundations of large language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17737– 17752, Miami, Florida, USA,...
2024 doi
-
[35]
URLhttps://aclanthology.org/2024.emnlp-main.982/
2024
-
[36]
Jos´ e Luiz Nunes, Guilherme F. C. F. Almeida, Marcelo de Araujo, and Simone D. J. Barbosa. Are large language models moral hypocrites? a study based on moral foundations. InProceed- ings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 1074–1087, 2024. d...
2024 doi
-
[37]
The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393,
Kazuhiro Takemoto. The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393,
-
[38]
doi: 10.1098/rsos.231393
-
[39]
When to make exceptions: Exploring language models as accounts of human moral judg- ment
Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Joshua Tenen- baum, and Bernhard Sch¨ olkopf. When to make exceptions: Exploring language models as accounts of human moral judg- ment. In36th Conference on Neural Informat...
2022
-
[40]
Dorit Hadar-Shoval, Kfir Asraf, Yonathan Mizrachi, Yuval Haber, and Zohar Elyoseph. Assessing the alignment of large language models with human values for mental health integration: Cross- sectional study using Schwartz’s theory of basic values.JMIR Mental Health, 11:e55988, 2...
2024 doi
-
[41]
MoralBench: Moral evaluation of LLMs, 2024
Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. MoralBench: Moral evaluation of LLMs, 2024
2024
-
[42]
Campbell and Donald W
Donald T. Campbell and Donald W. Fiske. Convergent and dis- criminant validation by the multitrait-multimethod matrix.Psy- chological Bulletin, 56(2):81–105, 1959. doi: 10.1037/h0046016
1959 doi
-
[43]
Cronbach and Paul E
Lee J. Cronbach and Paul E. Meehl. Construct validity in psychological tests.Psychological Bulletin, 52(4):281–302, 1955. doi: 10.1037/h0040957
1955 doi
-
[44]
Academic Press, New York, 1981
Howard Schuman and Stanley Presser.Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Academic Press, New York, 1981. ISBN 0126313504
1981
-
[45]
Billiet and McKee J
Jaak B. Billiet and McKee J. McClendon. Modeling acquiescence in measurement models for two balanced sets of items.Structural Equation Modeling: A Multidisciplinary Journal, 7(4):608–628,
-
[46]
doi: 10.1207/S15328007SEM0704 5
-
[47]
Yves Van Vaerenbergh and Troy D. Thomas. Response styles in survey research: A literature review of antecedents, consequences, and remedies.International Journal of Public Opinion Research, 25(2):195–217, 2013. doi: 10.1093/ijpor/eds021
2013 doi
-
[48]
Krosnick and Duane F
Jon A. Krosnick and Duane F. Alwin. An evaluation of a cognitive theory of response-order effects in survey measurement. Public Opinion Quarterly, 51(2):201–219, 1987. doi: 10.1086/ 269029
1987
-
[49]
Krosnick
Jon A. Krosnick. Response strategies for coping with the cogni- tive demands of attitude measures in surveys.Applied Cognitive Psychology, 5(3):213–236, 1991. doi: 10.1002/acp.2350050305
1991 doi
-
[50]
Nelder.Generalized Linear Models
Peter McCullagh and John A. Nelder.Generalized Linear Models. Chapman and Hall, London, 2nd edition, 1989
1989
-
[51]
Danish Institute for Educational Research, Copenhagen, 1960
Georg Rasch.Probabilistic Models for Some Intelligence and Attainment Tests. Danish Institute for Educational Research, Copenhagen, 1960. Reissued 1980, Chicago: University of Chicago Press, foreword by Benjamin D. Wright
1960
-
[52]
Some latent trait models and their use in inferring an examinee’s ability
Allan Birnbaum. Some latent trait models and their use in inferring an examinee’s ability. In Frederic M. Lord and Melvin R. Novick, editors,Statistical Theories of Mental Test Scores, pages 397–479. Addison-Wesley, Reading, MA, 1968
1968
-
[53]
Lord.Applications of Item Response Theory to Prac- tical Testing Problems
Frederic M. Lord.Applications of Item Response Theory to Prac- tical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980
1980
-
[54]
Embretson and Steven P
Susan E. Embretson and Steven P. Reise.Item Response The- ory for Psychologists. Multivariate Applications Book Series. Lawrence Erlbaum Associates, Mahwah, NJ, 2000
2000
-
[55]
large language models show amplified cogni- tive biases in moral decision-making
Maximilian Maier, Vanessa Cheung, and Falk Lieder. Code and data for analyses in “large language models show amplified cogni- tive biases in moral decision-making”. Open Science Framework, https://osf.io/3kvjd/, 2025. Deposited 4 November 2025
2025
-
[56]
Rothkopf, and Kristian Kersting
Patrick Schramowski, Cigdem Turan, Nico Andersen, Con- stantin A. Rothkopf, and Kristian Kersting. Large pre-trained language models contain human-like biases of what is right and wrong to do.Nature Machine Intelligence, 4(3):258–268, 2022. doi: 10.1038/s42256-022-00458-8
2022 doi
-
[57]
Moral mimicry: Large language models pro- duce moral rationalizations tailored to political identity
Gabriel Simmons. Moral mimicry: Large language models pro- duce moral rationalizations tailored to political identity. InPro- ceedings of the 61st Annual Meeting of the Association for Com- putational Linguistics (Volume 4: Student Research Workshop), pages 282–297, Toronto, C...
-
[58]
Guilherme F. C. F. Almeida, Jos´ e Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de Ara´ ujo. Exploring the psychol- ogy of LLMs’ moral and legal reasoning.Artificial Intelligence, 333:104145, August 2024. doi: 10.1016/j.artint.2024.104145
2024 doi
-
[59]
Language model alignment in multilingual trolley problems
Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez, Francesco Ortu, Andr´ as Strausz, 13 Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, and Bernhard Sch¨ olkopf. Language model alignment in multilingual trolley problems. InInternational ...
2025
-
[60]
Who is GPT-3? An exploration of personality, values and demographics
Maril` u Miotto, Nicola Rossberg, and Bennett Kleinberg. Who is GPT-3? An exploration of personality, values and demographics. InProceedings of the Fifth Workshop on Natural Language Pro- cessing and Computational Social Science (NLP+CSS), pages 218–227, Abu Dhabi, UAE, Novemb...
2022 doi
-
[61]
The self-perception and po- litical biases of ChatGPT.Human Behavior and Emerging Technologies, 2024:1–9, 2024
J´ erˆ ome Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, Moritz Roidl, and Markus Pauly. The self-perception and po- litical biases of ChatGPT.Human Behavior and Emerging Technologies, 2024:1–9, 2024. doi: 10.1155/2024/7115633
2024 doi
-
[62]
CMoralEval: A moral evaluation benchmark for Chinese large language mod- els
Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Tao Liu, and Deyi Xiong. CMoralEval: A moral evaluation benchmark for Chinese large language mod- els. InFindings of the Association for Computationa...
2024 doi
-
[63]
Llm ethics benchmark: A three-dimensional assessment system for evaluating moral reasoning in large language models.Scientific Reports, 15:34642,
Junfeng Jiao, Saleh Afroogh, Abhejay Murali, Kevin Chen, David Atkinson, and Amit Dhurandhar. Llm ethics benchmark: A three-dimensional assessment system for evaluating moral reasoning in large language models.Scientific Reports, 15:34642,
-
[64]
doi: 10.1038/s41598-025-18489-7
-
[65]
Evaluating moral beliefs across LLMs through a pluralistic framework
Xuelin Liu, Yanfei Zhu, Shucheng Zhu, Pengyuan Liu, Ying Liu, and Dong Yu. Evaluating moral beliefs across LLMs through a pluralistic framework. InFindings of the Association for Com- putational Linguistics: EMNLP 2024, pages 4740–4760, Miami, Florida, USA, November 2024. Asso...
2024 doi
-
[66]
Cultural value alignment in large language mod- els: A prompt-based analysis of Schwartz values in Gemini, ChatGPT, and DeepSeek, 2025
Robin Segerer. Cultural value alignment in large language mod- els: A prompt-based analysis of Schwartz values in Gemini, ChatGPT, and DeepSeek, 2025
2025
-
[67]
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1...
2024 doi
-
[68]
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamil˙ e Lukoˇ si¯ ut˙ e, Karina Nguyen, et al. Discovering language model behaviors with model-written evaluations. InFindings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434, 2023
2023
-
[69]
Bowman, Newton Cheng, Esin Dur- mus, Zac Hatfield-Dodds, Scott R
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Dur- mus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang,...
2024
-
[70]
Primacy effect of ChatGPT
Yiwei Wang, Yujun Cai, Muhao Chen, Yuxuan Liang, and Bryan Hooi. Primacy effect of ChatGPT. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, pages 108–115, Singapore, December 2023. ...
2023 doi
-
[71]
Prompt perturbations reveal human-like biases in large language model survey responses.arXiv preprint arXiv:2507.07188, 2025
Jens Rupprecht, Georg Ahnert, and Markus Strohmaier. Prompt perturbations reveal human-like biases in large language model survey responses.arXiv preprint arXiv:2507.07188, 2025
2025
-
[72]
Acquiescence bias in large language models
Daniel Braun. Acquiescence bias in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2025, 2025
2025
-
[73]
Thilo Hagendorff, Ishita Dasgupta, Marcel Binz, Stephanie C. Y. Chan, Andrew Lampinen, Jane X. Wang, Zeynep Akata, and Eric Schulz. Machine psychology.arXiv preprint arXiv:2303.13988, 2023
2023 arXiv
-
[74]
Using cognitive psychology to understand GPT-3.Proceedings of the National Academy of Sci- ences, 120(6):e2218523120, 2023
Marcel Binz and Eric Schulz. Using cognitive psychology to understand GPT-3.Proceedings of the National Academy of Sci- ences, 120(6):e2218523120, 2023. doi: 10.1073/pnas.2218523120
2023 doi
-
[75]
Wang, and Eric Schulz
Julian Coda-Forno, Marcel Binz, Jane X. Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofPMLR, pages 9076–9108, 2024
2024
-
[76]
Argyle, Ethan C
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023. doi: 10.1017/pan.2023.2
2023 doi
-
[77]
Arriaga, and Adam Tauman Kalai
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 337–371, 2023
2023
-
[78]
Yeager, Christopher J
Dorottya Demszky, Diyi Yang, David S. Yeager, Christopher J. Bryan, Margarett Clapper, Susannah Chandhok, Johannes C. Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann John- son, Michaela Jones, Danielle Krettek-Cobb, Leslie Lai, Nirel JonesMitchell, Desmond C. Ong, Carol S....
2023 doi
-
[79]
Talking about large language models.Com- munications of the ACM, 67(2):68–79, 2024
Murray Shanahan. Talking about large language models.Com- munications of the ACM, 67(2):68–79, 2024. doi: 10.1145/ 3624724
2024
-
[80]
Lisa Messeri and M. J. Crockett. Artificial intelligence and illusions of understanding in scientific research.Nature, 627: 49–58, 2024. doi: 10.1038/s41586-024-07146-0
2024 doi
-
[81]
Can AI language models replace human participants?Trends in Cognitive Sciences, 27(7):597–600, 2023
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can AI language models replace human participants?Trends in Cognitive Sciences, 27(7):597–600, 2023. doi: 10.1016/j.tics. 2023.04.008
2023 doi
-
[82]
AI language model rivals expert ethicist in perceived moral expertise.Scientific Reports, 15:4084, 2025
Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. AI language model rivals expert ethicist in perceived moral expertise.Scientific Reports, 15:4084, 2025. doi: 10.1038/ s41598-025-86510-0
2025
-
[83]
Brady, Caelan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo
Eyal Aharoni, Sharlene Fernandes, Daniel J. Brady, Caelan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo. Attributions toward artificial agents in a modified Moral Turing Test.Scientific Reports, 14: 8458, 2024. doi: 10.1038/s41598-024-58087-7
2024 doi
-
[84]
The moral machine experiment.Nature, 563(7729): 59–64, 2018
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-Fran¸ cois Bonnefon, and Iyad Rahwan. The moral machine experiment.Nature, 563(7729): 59–64, 2018. doi: 10.1038/s41586-018-0637-6
2018 doi
-
[85]
Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier
Max Pellert, Clemens M. Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. AI psychometrics: Assess- ing the psychological profiles of large language models through psychometric inventories.Perspectives on Psychological Science, 19(5):808–826, 2024. doi: 10.11...
2024 doi
-
[86]
Ullman, Fernando Martinez-Plumed, Joshua B
Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Ru- tar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernand...
-
[87]
doi: 10.1126/science.adf6369
-
[88]
Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. AI and the ev- erything in the whole wide world benchmark. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS 2021), 2021
2021
-
[89]
Do large language model benchmarks test reliability?, 2025
Joshua Vendrow, Edward Vendrow, Sara Beery, and Aleksander Madry. Do large language model benchmarks test reliability?, 2025. 14
2025
-
[90]
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, Mar´ ıa Grandury, Simeng Han, Valentin Hof...
2025
-
[91]
Lost in bench- marks? rethinking large language model benchmarking with item response theory
Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Wei Bao, Jian Dong, Bing Xu, Conghui Zhu, Hailong Cao, and Tiejun Zhao. Lost in bench- marks? rethinking large language model benchmarking with item response theory. InProceedings of the AA...
2026
-
[92]
Establishing construct validity in LLM capa- bility benchmarks requires nomological networks, 2026
Timo Freiesleben. Establishing construct validity in LLM capa- bility benchmarks requires nomological networks, 2026
2026
-
[93]
Six fallacies in substituting large language models for human participants.Advances in Methods and Practices in Psychological Science, 8(3), 2025
Zhicheng Lin. Six fallacies in substituting large language models for human participants.Advances in Methods and Practices in Psychological Science, 8(3), 2025. doi: 10.1177/ 25152459251357566
2025
-
[94]
Being blind (or not) to scenarios used in sacrificial dilemmas: the influence of factual and contextual information on moral responses.Frontiers in Psychology, 15:1477825, 2024
Robin Carron, Emmanuelle Brigaud, Royce Anders, and Nathalie Blanc. Being blind (or not) to scenarios used in sacrificial dilemmas: the influence of factual and contextual information on moral responses.Frontiers in Psychology, 15:1477825, 2024. doi: 10.3389/fpsyg.2024.1477825
2024 doi
-
[95]
Cohen and Philip T
Dale J. Cohen and Philip T. Quinlan. Why moral judgements change across variations of trolley-like problems.British Journal of Psychology, 2025. doi: 10.1111/bjop.12782. Published online 18 February 2025; volume/issue/pages not yet assigned (re- checked via Crossref 2026-07-03)
2025 doi
-
[96]
Christensen and Antoni Gomila
Julia F. Christensen and Antoni Gomila. Moral dilemmas in cognitive neuroscience of moral decision-making: A principled review.Neuroscience & Biobehavioral Reviews, 36(4):1249–1264,
-
[97]
doi: 10.1016/j.neubiorev.2012.02.008
2012 doi
-
[98]
How stable are moral judgments?Review of Philosophy and Psychology, 14(4): 1377–1403, 2023
Paul Rehren and Walter Sinnott-Armstrong. How stable are moral judgments?Review of Philosophy and Psychology, 14(4): 1377–1403, 2023. doi: 10.1007/s13164-022-00649-7
2023 doi
-
[99]
Discrepancies between judgment and choice of action in moral dilemmas.Frontiers in Psychology, 4:250, 2013
S´ ebastien Tassy, Olivier Oullier, Julien Mancini, and Bruno Wicker. Discrepancies between judgment and choice of action in moral dilemmas.Frontiers in Psychology, 4:250, 2013. doi: 10.3389/fpsyg.2013.00250
2013 doi
-
[100]
What we say and what we do: The relationship between real and hypothetical moral choices
Oriel FeldmanHall, Dean Mobbs, Davy Evans, Lucy Hiscox, Lauren Navrady, and Tim Dalgleish. What we say and what we do: The relationship between real and hypothetical moral choices. Cognition, 123(3):434–441, 2012. doi: 10.1016/j.cognition.2012. 02.001
2012 doi
-
[101]
Francis, Charles Howard, Ian S
Kathryn B. Francis, Charles Howard, Ian S. Howard, Michaela Gummerum, Giorgio Ganis, Grace Anderson, and Sylvia Terbeck. Virtual morality: Transitioning from moral judgment to moral action?PLOS ONE, 11(10):e0164374, 2016. doi: 10.1371/ journal.pone.0164374
2016
-
[102]
Self-reports: How the questions shape the answers.American Psychologist, 54(2):93–105, 1999
Norbert Schwarz. Self-reports: How the questions shape the answers.American Psychologist, 54(2):93–105, 1999. doi: 10. 1037/0003-066x.54.2.93
1999
-
[103]
Rips, and Kenneth A
Roger Tourangeau, Lance J. Rips, and Kenneth A. Rasinski.The Psychology of Survey Response. Cambridge University Press, Cambridge, UK, 2000. ISBN 0521572460
2000
-
[104]
Bradburn, and Norbert Schwarz
Seymour Sudman, Norman M. Bradburn, and Norbert Schwarz. Thinking about Answers: The Application of Cognitive Processes to Survey Methodology. Jossey-Bass, San Francisco, 1996. ISBN 0787901202
1996
-
[105]
Greene, R
Joshua D. Greene, R. Brian Sommerville, Leigh E. Nystrom, John M. Darley, and Jonathan D. Cohen. An fmri investigation of emotional engagement in moral judgment.Science, 293(5537): 2105–2108, 2001. doi: 10.1126/science.1062872
2001 doi
-
[106]
How (and where) does moral judgment work?Trends in Cognitive Sciences, 6(12): 517–523, 2002
Joshua Greene and Jonathan Haidt. How (and where) does moral judgment work?Trends in Cognitive Sciences, 6(12): 517–523, 2002. doi: 10.1016/S1364-6613(02)02011-9
2002 doi
-
[107]
Action, outcome, and value: A dual-system framework for morality.Personality and Social Psychology Review, 17(3):273–292, 2013
Fiery Cushman. Action, outcome, and value: A dual-system framework for morality.Personality and Social Psychology Review, 17(3):273–292, 2013. doi: 10.1177/1088868313495594
2013 doi
-
[108]
The intuitive greater good: Testing the corrective dual process model of moral cognition
Bence Bago and Wim De Neys. The intuitive greater good: Testing the corrective dual process model of moral cognition. Journal of Experimental Psychology: General, 148(10):1782– 1801, 2019. doi: 10.1037/xge0000533
2019 doi
-
[109]
Guy Kahane, Jim A. C. Everett, Brian D. Earp, Lucius Caviola, Nadira S. Faber, Molly J. Crockett, and Julian Savulescu. Be- yond sacrificial harm: A two-dimensional model of utilitarian psychology.Psychological Review, 125(2):131–164, 2018. doi: 10.1037/rev0000093
2018 doi
-
[110]
Hippler, Elisabeth Noelle-Neumann, and Leslie Clark
Norbert Schwarz, B¨ arbel Kn¨ auper, Hans-J. Hippler, Elisabeth Noelle-Neumann, and Leslie Clark. Rating scales: Numeric values may change the meaning of scale labels.Public Opinion Quarterly, 55(4):570–582, 1991. doi: 10.1086/269282
1991 doi
-
[111]
K¨ uhnel
Jan Karem H¨ ohne, Dagmar Krebs, and Steffen-M. K¨ uhnel. Measurement properties of completely and end labeled unipo- lar and bipolar scales in Likert-type questions on income (in)equality.Social Science Research, 97:102544, 2021. doi: 10.1016/j.ssresearch.2021.102544
2021 doi
-
[112]
K¨ uhnel
Jan Karem H¨ ohne, Dagmar Krebs, and Steffen-M. K¨ uhnel. Mea- suring income (in)equality: Comparing survey questions with unipolar and bipolar scales in a probability-based online panel. Social Science Computer Review, 40(1):108–123, 2022. doi: 10.1177/0894439320902461
2022 doi
-
[113]
Burdick, Connie M
Richard K. Burdick, Connie M. Borror, and Douglas C. Mont- gomery. A review of methods for measurement systems capability analysis.Journal of Quality Technology, 35(4):342–354, 2003. doi: 10.1080/00224065.2003.11980232
2003 doi
-
[114]
Cronbach, Nageswari Rajaratnam, and Goldine C
Lee J. Cronbach, Nageswari Rajaratnam, and Goldine C. Gleser. Theory of generalizability: A liberalization of reliability theory. British Journal of Statistical Psychology, 16(2):137–163, 1963. doi: 10.1111/j.2044-8317.1963.tb00206.x
1963 doi
-
[115]
Cronbach, Goldine C
Lee J. Cronbach, Goldine C. Gleser, Harinder Nanda, and Nageswari Rajaratnam.The Dependability of Behavioral Mea- surements: Theory of Generalizability for Scores and Profiles. Wiley, New York, 1972
1972
-
[116]
Shavelson and Noreen M
Richard J. Shavelson and Noreen M. Webb.Generalizability Theory: A Primer. Sage Publications, Newbury Park, CA, 1991
1991
-
[117]
Brennan.Generalizability Theory
Robert L. Brennan.Generalizability Theory. Statistics for Social and Behavioral Sciences. Springer, New York, NY, 2001. doi: 10.1007/978-1-4757-3456-0. 15
2001 doi
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.