Pith. sign in

REVIEW 3 major objections 5 minor 117 references

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read The yes-no bias of large language models on moral dilemmas is an artifact of answer order and the word “no,” not a change in moral judgment.

desk verdict Clean factorial decomposition of the yes-no bias into order + lexical surface pulls, with logical attachment ~0 under label swap and a coherent graded stance as independent axis. read the letter →

arxiv 2607.05552 v1 pith:VWRNES4Y submitted 2026-07-06 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords largelanguagemodelsmoraljudgmentpsychometricsframingeffectsAIevaluationyes-nobiasanswerordercrossedsymmetrization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Frontier language models appear to flip moral verdicts under tiny wording changes, including an amplified yes-no bias that people do not show. This paper argues that the flip is not a change in what the model values. When the same dilemmas are rated on graded scales under many logically equivalent framings, frontier models keep a stable internal stance. The large bias appears only when the answer is forced through yes/no: it splits into a pull toward the last-printed option and a pull toward the word “no.” Replace those words with arbitrary labels and the bias attached to the actual verdict vanishes—the models are not drawn toward rejecting, only toward the printed surface. Measuring what a model values therefore requires crossing the frames of the question, not asking once.

What carries the argument

Crossed symmetrization: every logically irrelevant factor (verb, printed order, answer label, scale, anchor, wording, pole) is flipped in balanced pairs and the factors are crossed so their contributions separate. The symmetric part of each flip-pair estimates stance θ; the antisymmetric part is the artifact. The minimal model P = σ((θ ± m)/s) then summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, distinct from sampling temperature.

What would settle it

If the same crossed battery showed frontier models’ graded stance itself swinging as much as the reported binary artifact when only the rating question’s surface form changed, or if fully arbitrary answer labels still produced a large verdict-attached bias, the claim that the scale is format-invariant and the bias is surface-only would fail.

Watch

Extended reading notes

Core claim

Frontier models carry a coherent internal moral scale: graded ratings of the same dilemmas stay nearly format-invariant under crossed, logically equivalent framings. Forcing the judgment through yes/no overlays a decomposable format artifact—an order bias toward the last-printed option plus a lexical pull toward the word “no”—large mainly in Claude models and smaller under extended reasoning. With arbitrary answer labels the verdict-attached logical bias is approximately zero for every frontier model; the pull follows the printed surface, not the verdict it carries.

Load-bearing premise

That the averaged graded rating under many equivalent framings really is the model’s stable moral stance, so it can serve as the fixed axis against which yes/no artifacts are measured.

Editorial extensions

If this is right

  • Graded, multi-frame elicitation recovers a model’s moral stance more cleanly than any single forced yes/no.
  • Single-format binary readouts confound stance with surface format and should not be read as direct measures of value.
  • Extended reasoning typically shrinks both cross-form incoherence and framing susceptibility where the artifact is large.
  • The same battery applies unchanged to any dilemma set and binary format, so the decomposition can be rerun on other value domains.
  • A yes-no bias reported without crossing verb, order, and label cannot be attributed to moral judgment versus surface pull.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Safety gates and LLM-as-judge systems that force yes/no may be measuring label and order attachments rather than the intended policy.
  • The recency-type order bias (opposite classic human primacy) is likely to appear in many multiple-choice readouts, not only moral items.
  • Interventions that claim to reduce format sensitivity can be audited by tracking m and s rather than raw accuracy alone.
  • Small models can look “coherent” by being indifferent; any coherence number without a discrimination check can flatter an empty scale.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that the amplified yes–no bias of LLMs on moral dilemmas is a format artifact of the binary readout, not a shift in moral judgment. Using a crossed-symmetrization psychometric battery (graded ratings I-1, free choice I-2, forced binary I-3) on twenty dilemmas largely from Cheung et al., it recovers a nearly format-invariant graded stance θ for frontier models (cross-form incoherence σ_repro = 0.12–0.21). The forced yes/no readout decomposes, by the verb × order identity (Eq. 1), into an order bias toward the last-printed option plus a lexical pull toward the word “no,” concentrated in Claude models and shrinking under extended reasoning. Swapping yes/no for arbitrary labels (A/B) drives the verdict-attached logical component to ≈0 for frontier models, while surface label and order attachments remain. A logistic summary P = σ((θ ± m)/s) yields portable framing susceptibility m and moral decisiveness s.

Significance. If the result holds, it reframes a growing literature on LLM moral and survey biases: single-framing yes/no verdicts confound stance with surface form, and evaluations that ask once systematically misread format artifacts as value shifts. The methodological contribution—crossed symmetrization that separates order, lexical, and logical channels, with independent graded θ as the stance axis—is portable beyond these dilemmas and is a genuine advance over uncrossed multi-prompt consistency rates. Strengths include a definitional decomposition (Eq. 1), a clean label-swap identification, pre-registered exclusion of bipolar cells, matched 12-form baselines, salt replications establishing deterministic open-weight incoherence, convergent free-choice checks, and an explicit, reproducible analysis pipeline. The (m, s) parameterization and the demonstration that deliberation shrinks |m| are useful, falsifiable summaries for future work.

major comments (3)
  1. [Abstract; Results “Lexical, not logical”; Fig. 5] Abstract and Results (“Lexical, not logical” / Fig. 5): the abstract states that the verdict-attached logical bias “proves ≈0 for every frontier model.” Methods and Fig. 5 correctly qualify that six of seven A/B logic CIs span zero (exception −0.02) and that Haiku’s logic channels are wide (±0.2–0.3) and underpowered, so the claim is a bound there. Align the abstract and significance statement with that qualification; “≈0 where precisely measured; a bound for Haiku” is what the data support.
  2. [Materials and Methods (I-3, refusal handling); Results Fig. 2] Methods I-3 / refusal handling: Haiku refuses on 28%/35% of core verb-flip trials and Flash-Lite withholds on 24%. Bias estimates condition on non-refusal after cell exclusion. The paper shows family concentration is not a pure refusal artifact (Flash-Lite ≈0 bias despite high withholding), but does not report whether refusal rates differ systematically by verb, printed order, or label. Differential refusal by frame would make exclusion non-ignorable for the order/lexical split. Please report refusal rates by the crossed factors (or a sensitivity analysis that bounds the bias under plausible missingness) for the high-refusal configurations.
  3. [Results “An internal moral scale exists”; Methods I-2] Results I-2 convergent validity (r = 0.82–0.90 with graded θ): free-choice verdicts are extracted by Claude Opus 4.8, which shares a vendor family with several subjects. The authors flag circularity risk and ship transcripts, but the main-text correlations are presented as primary convergent evidence for the latent scale. Either re-extract a subset with an independent judge (or human coding) and report agreement, or move the I-2 correlations to a clearly caveated secondary check so the θ claim does not rest on same-family judging.
minor comments (5)
  1. [Fig. 1b caption] Fig. 1b: variance-share bars (anchor-direction solid, scale hatched) combine in quadrature; a one-line reminder in the caption that shares are not linearly additive would prevent misreading stacked heights.
  2. [Methods; throughout Results] Notation: b, z, θ, m, s, and σ_repro are introduced across Results and Methods; a short symbol table in Methods or SI would help readers track the [−1, +1] convention and the distinction between descriptive incoherence spreads and bootstrap CIs.
  3. [Results “The standard readout overlays a format artifact”; Methods “Cheung-verbatim control”] The Cheung-verbatim K01 control (order-balanced residual +0.01, CI spanning zero) is important for engaging the source finding; consider elevating one sentence of it into the main Results paragraph on decomposition rather than leaving it mostly in Methods.
  4. [Results opening; Discussion scopes] SI Appendix is heavily referenced for open-weight degeneracy (Nemotron), salt floors, and G coefficients; ensure the main text’s “small open-weight models fail in model-specific ways” is self-contained enough that a reader who skips SI still sees the discrimination-vs-coherence distinction (G vs raw σ_repro).
  5. [Introduction; Methods model panel] Typos / polish: “theyes–no bias” spacing in the Introduction; occasional missing spaces after em-dashes in the compiled text; confirm that “GPT-5.5” and model snapshot IDs are the intended public names at submission.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: independent instruments, definitional crossing identity used transparently as measurement, and fitted (m,s) as descriptive summary rather than forced prediction.

full rationale

The paper's load-bearing chain does not reduce inputs to outputs by construction. Stance θ is recovered from the graded instrument I-1 (24 unipolar conditions; no yes/no labels or printed order), while binary artifacts are measured on a separate forced-binary instrument I-3; free-choice I-2 supplies an independent convergent check (r = 0.82–0.90). The identity b_apparent = b_order + b_lexical is stated as exact and definitional from the verb×order crossing (Eq. 1), which is standard factorial measurement, not a disguised prediction: the scientific content is that the summands are separately measurable and that the verdict-attached (logical) component collapses to ≈0 under fully arbitrary A/B labels—an empirical result of the label-swap design, not forced by the yes/no token identity. The logistic P = σ((θ ± m)/s) is a fitted descriptive summary of observed flip-pairs (bowtie/ridge geometry), with θ as an external regressor from I-1; m and s are not inserted into the central claim by definition, and the paper flags errors-in-variables attenuation and clip-limited s. No self-citation chain, uniqueness theorem, or ansatz from the same author underwrites the result. Materials are from Cheung et al.; the decomposition and logical-null are new measurements on those materials. Honest non-finding: score 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 2 invented entities

The central claim rests on a psychometric identification strategy (crossed symmetrization + independent graded instrument) plus a standard logistic link. Free parameters are the fitted (m, s) and the estimated heta; axioms are ordinary measurement and GLM assumptions plus the domain claim that format-invariant graded ratings recover a latent stance. No new physical entities are postulated; m and s are summary statistics of observed flip-pairs.

free parameters (3)
  • framing susceptibility m (per bias channel and model configuration)
    Horizontal offset fitted by nonlinear least squares to each flip-pair; story-independent summary of the artifact size.
  • moral decisiveness s (per bias channel and model configuration)
    Scale parameter of the logistic link, fitted jointly with m; interpreted as how sharply binary verdicts track graded stance.
  • stance heta per dilemma
    Estimated as a_action - a_complement after normalization and anchor flip across the 24 unipolar I-1 conditions; chosen coordinate on [-1, +1].
assumptions (5)
  • domain assumption Logically equivalent but operationally independent graded elicitations that agree recover a latent stance (convergent validity).
    Invoked in Results "An internal moral scale exists" to treat heta as the independent axis for bias tests.
  • standard math Canonical logistic link p = σ(( heta ± m)/s) from continuous latent to binary choice.
    Used in "The artifact has structure" and Methods "The two-parameter fit"; standard IRT/GLM link, not derived from first principles here.
  • standard math Balanced 2-vs-2 splits of the four verb imes order cells separate order from lexical contributions exactly (Eq. 1).
    Definitional identity once cells are measured; content is that the summands are separately meaningful.
  • domain assumption Arbitrary answer labels (A/B) carry the verdict without carrying English yes/no lexical content, so the logic projection isolates verdict attachment.
    Identification claim for Fig. 5; rests on the labels being valence-free for the models tested.
  • domain assumption Sampling-corrected between-form variance σ_repro is the dominant uncertainty and the proper error bar on stance.
    Methods "Cross-form incoherence"; Gauge R&R style correction applied throughout.
invented entities (2)
  • framing susceptibility m and moral decisiveness s
    purpose: Portable two-parameter summary of any binary format artifact relative to graded stance.
    Introduced as the minimal model P=σ(( heta±m)/s); they are fitted summaries, not postulated mechanisms with independent existence claims.
  • crossed-symmetrization psychometric battery (I-1/I-2/I-3) independent evidence
    purpose: Separate logical, lexical, and order contributions that coincide on a single yes/no token.
    The instrument itself is the methodological contribution; it is a design, not a physical entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment." pith.science (2026). https://pith.science/paper/VWRNES4Y

@misc{pith2026260705552,
  author       = {Pith},
  title        = {Pith review of: The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VWRNES4Y}},
  note         = {Machine review of arXiv:2607.05552}
}
abstract

Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = \sigma((\theta \pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.

Figures

Figures reproduced from arXiv: 2607.05552 by the authors.

Figure 1
Figure 1. Frontier models carry a coherent internal moral scale; a small open-weight model, without deliberation, does not. (a) Stance θ per dilemma (abscissa: the 20 dilemmas, ordered by the frontier-mean consensus, gray band), small multiples per model family; filled markers = extended reasoning (“think”), open = direct; error bars = per-item cross-form incoherence σrepro (a descriptive spread, not a CI). Open diamonds: hum… view at source ↗
Figure 2
Figure 2. The apparent yes–no bias is a confound: it splits exactly into order bias + lexical. Per model: the apparent single-framing yes/no bias (purple) and its decomposition, by the crossing identity, into an order bias (toward the last-printed option; orange) and a lexical pull (toward the word “no”; red). Negative = toward “no.” The artifact is substantial only for the Claude models and shrinks under extended reasoning (… view at source ↗
Figure 3
Figure 3. Story-resolved structure: the artifact is a horizontal offset on a logistic, and one fit explains both bias channels. Rows: the four Claude configurations. Columns: the stance sigmoid z(θ) with the two per-bias fits overlaid (their overlap is the internal consistency check); then, per bias (order, lexical), the bowtie (bias versus stance z) and the ridge (bias versus θ). Points: dilemmas (n = 18–20 per configuration… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Two portable parameters: framing susceptibility m and moral decisiveness s. (a–c) The model. A flip-pair responds as p± = σ((θ ± m)/s): the mean of the pair is the stance, the signed difference is the bias, and the frame’s pull is the horizontal offset m (m < 0 = towar…
Figure 5
Figure 5. Figure 5: The yes/no pull is lexical, not logical: swapping the answer label removes the word-attached pull. (a) The verb×label×order projections per answer-label family, per model: order bias / label (surface-label pull) / logic (verdict-attached pull carried by a non-yes/no-wo…
Figure 6
Figure 6. Figure 6: The pull is graded by the label’s yes/no-ness—on average—and it is model-specific. Thirteen answer-label pairs, ordered non-lexical (A/B, 1/2, $/%, +/−) → quasi-lexical (Y/N; chk = check/cross marks; thumb = thumbs-up/down; T/F = true/false) → lexical (ja/nein, oui/non…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

117 extracted references · 23 canonical work pages

  1. [1]

    The framing of decisions and the psychology of choice.Science, 211(4481):453–458, 1981

    Amos Tversky and Daniel Kahneman. The framing of decisions and the psychology of choice.Science, 211(4481):453–458, 1981. doi: 10.1126/science.7455683

  2. [2]

    Choices, values, and frames.American Psychologist, 39(4):341–350, 1984

    Daniel Kahneman and Amos Tversky. Choices, values, and frames.American Psychologist, 39(4):341–350, 1984. doi: 10. 1037/0003-066X.39.4.341

  3. [3]

    Levin, Sandra L

    Irwin P. Levin, Sandra L. Schneider, and Gary J. Gaeth. All frames are not created equal: A typology and critical analysis of framing effects.Organizational Behavior and Human Decision Processes, 76(2):149–188, 1998. doi: 10.1006/obhd.1998.2804

  4. [4]

    A systematic review of risky-choice framing effects.EXCLI Journal, 22:1012–1031, 2023

    Anton K¨ uhberger. A systematic review of risky-choice framing effects.EXCLI Journal, 22:1012–1031, 2023. doi: 10.17179/ excli2023-6169

  5. [5]

    Influence of wording and framing effects on moral intuitions.Ethology and Sociobiology, 17(3):145–171, 1996

    Lewis Petrinovich and Patricia O’Neill. Influence of wording and framing effects on moral intuitions.Ethology and Sociobiology, 17(3):145–171, 1996. doi: 10.1016/0162-3095(96)00041-6

  6. [6]

    Order effects in moral judgment.Philosophical Psychology, 25(6):813–836,

    Alex Wiegmann, Yasmina Okan, and Jonas Nagel. Order effects in moral judgment.Philosophical Psychology, 25(6):813–836,

  7. [7]

    doi: 10.1080/09515089.2011.631995

  8. [8]

    Expertise in moral reasoning? order effects on moral judgment in professional philosophers and non-philosophers.Mind & Language, 27(2): 135–153, 2012

    Eric Schwitzgebel and Fiery Cushman. Expertise in moral reasoning? order effects on moral judgment in professional philosophers and non-philosophers.Mind & Language, 27(2): 135–153, 2012. doi: 10.1111/j.1468-0017.2012.01438.x

Show all 117 references
  1. [9]

    Philosophers’ biased judgments persist despite training, expertise and reflection.Cog- nition, 141:127–137, 2015

    Eric Schwitzgebel and Fiery Cushman. Philosophers’ biased judgments persist despite training, expertise and reflection.Cog- nition, 141:127–137, 2015. doi: 10.1016/j.cognition.2015.04.015

  2. [10]

    Large lan- guage models show amplified cognitive biases in moral decision- making.Proc

    Vanessa Cheung, Maximilian Maier, and Falk Lieder. Large lan- guage models show amplified cognitive biases in moral decision- making.Proc. Natl. Acad. Sci. U.S.A., 122(25):e2412015122,

  3. [11]

    doi: 10.1073/pnas.2412015122

  4. [12]

    Robustness of large language models in moral judgements.Royal Society Open Science, 12 (4):241229, 2025

    Soyoung Oh and Vera Demberg. Robustness of large language models in moral judgements.Royal Society Open Science, 12 (4):241229, 2025. doi: 10.1098/rsos.241229

  5. [13]

    The greatest good benchmark: Measuring LLMs’ alignment with utilitarian moral dilemmas

    Giovanni Franco Gabriel Marraffini, Andr´ es Cotton, No´ e Fabi´ an Hsueh, Axel Fridman, Juan Wisznia, and Luciano del Corro. The greatest good benchmark: Measuring LLMs’ alignment with utilitarian moral dilemmas. InProceedings of the 2024 Conference on Empirical Methods in Na...

  6. [14]

    Murukannaiah, and Munindar P

    Jiaqing Yuan, Pradeep K. Murukannaiah, and Munindar P. Singh. Right vs. right: Can LLMs make tough choices?, 2024

  7. [15]

    Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei. Evaluating the moral beliefs encoded in LLMs. In37th Con- ference on Neural Information Processing Systems (NeurIPS 2023), 2023

  8. [16]

    Varshney

    Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, and Lav R. Varshney. Many LLMs are more utilitarian than one. In39th Conference on Neural Information Processing Systems (NeurIPS 2025), 2025

  9. [17]

    Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. InInternational Conference on Learning Representations (ICLR), 2024

  10. [18]

    Large language models sensitivity to the order of options in multiple-choice questions

    Pouya Pezeshkpour and Estevam Hruschka. Large language models sensitivity to the order of options in multiple-choice questions. In Kevin Duh, Helena Gomez, and Steven Bethard, editors,Findings of the Association for Computational Linguis- tics: NAACL 2024, pages 2006–2017, Mex...

  11. [19]

    doi: 10.18653/ v1/2024.findings-naacl.130

    Association for Computational Linguistics. doi: 10.18653/ v1/2024.findings-naacl.130. URL https://aclanthology.org/ 2024.findings-naacl.130/

  12. [20]

    Large language models are not robust multiple choice selectors

    Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large language models are not robust multiple choice selectors. InInternational Conference on Learning Representa- tions (ICLR), 2024

  13. [21]

    Questioning the survey responses of large language models

    Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-D¨ unner. Questioning the survey responses of large language models. InAdvances in Neural Information Processing Systems (NeurIPS), volume 37, 2024

  14. [22]

    Do LLMs exhibit human-like response biases? a case study in survey design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024

    Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. Do LLMs exhibit human-like response biases? a case study in survey design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024. doi: 10.1162/tacl a 00685

  15. [23]

    Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models

    Paul R¨ ottger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schuetze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. InProceedings of the 62nd Annual Me...

  16. [24]

    State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024

    Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024. doi: 10.1162/ tacl a 00681

  17. [25]

    Promptrobust: Towards evaluating the 12 robustness of large language models on adversarial prompts

    Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zeek Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, and Xing Xie. Promptrobust: Towards evaluating the 12 robustness of large language models on adversarial prompts. In Proceedings of the 1st ACM Worksho...

  18. [26]

    In-context impersonation reveals large lan- guage models’ strengths and biases

    Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. In-context impersonation reveals large lan- guage models’ strengths and biases. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023

  19. [27]

    Aligning AI with shared human values

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning AI with shared human values. InInternational Conference on Learning Representations (ICLR 2021), 2021

  20. [28]

    Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Et- zioni, Maarten Sap, Regina Rini, and Yejin Choi. Can machines learn morality? the delphi experim...

  21. [29]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. InAd- vances in Neural Information Proces...

  22. [30]

    Whose opinions do language models reflect? InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 29971–30004, 2023

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 29971–30004, 2023

  23. [31]

    The effect of sampling temper- ature on problem solving in large language models

    Matthew Renze and Erhan Guven. The effect of sampling temper- ature on problem solving in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 7346–7356, 2024. doi: 10.18653/v1/2024.findings-emnlp. 432

  24. [32]

    The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism

    Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human...

  25. [33]

    Hashimoto, and Tobias Gerstenberg

    Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B. Hashimoto, and Tobias Gerstenberg. MoCa: Mea- suring human-language model alignment on causal and moral judgment tasks. InAdvances in Neural Information Processing Systems, volume 36, pages 78360–78393...

  26. [34]

    Moral foundations of large language models

    Marwa Abdulhai, Gregory Serapio-Garc´ ıa, Clement Crepy, Daria Valter, John Canny, and Natasha Jaques. Moral foundations of large language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17737– 17752, Miami, Florida, USA,...

  27. [35]

    URLhttps://aclanthology.org/2024.emnlp-main.982/

  28. [36]

    Jos´ e Luiz Nunes, Guilherme F. C. F. Almeida, Marcelo de Araujo, and Simone D. J. Barbosa. Are large language models moral hypocrites? a study based on moral foundations. InProceed- ings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 1074–1087, 2024. d...

  29. [37]

    The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393,

    Kazuhiro Takemoto. The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393,

  30. [38]

    doi: 10.1098/rsos.231393

  31. [39]

    When to make exceptions: Exploring language models as accounts of human moral judg- ment

    Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Joshua Tenen- baum, and Bernhard Sch¨ olkopf. When to make exceptions: Exploring language models as accounts of human moral judg- ment. In36th Conference on Neural Informat...

  32. [40]

    Dorit Hadar-Shoval, Kfir Asraf, Yonathan Mizrachi, Yuval Haber, and Zohar Elyoseph. Assessing the alignment of large language models with human values for mental health integration: Cross- sectional study using Schwartz’s theory of basic values.JMIR Mental Health, 11:e55988, 2...

  33. [41]

    MoralBench: Moral evaluation of LLMs, 2024

    Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. MoralBench: Moral evaluation of LLMs, 2024

  34. [42]

    Campbell and Donald W

    Donald T. Campbell and Donald W. Fiske. Convergent and dis- criminant validation by the multitrait-multimethod matrix.Psy- chological Bulletin, 56(2):81–105, 1959. doi: 10.1037/h0046016

  35. [43]

    Cronbach and Paul E

    Lee J. Cronbach and Paul E. Meehl. Construct validity in psychological tests.Psychological Bulletin, 52(4):281–302, 1955. doi: 10.1037/h0040957

  36. [44]

    Academic Press, New York, 1981

    Howard Schuman and Stanley Presser.Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Academic Press, New York, 1981. ISBN 0126313504

  37. [45]

    Billiet and McKee J

    Jaak B. Billiet and McKee J. McClendon. Modeling acquiescence in measurement models for two balanced sets of items.Structural Equation Modeling: A Multidisciplinary Journal, 7(4):608–628,

  38. [46]

    doi: 10.1207/S15328007SEM0704 5

  39. [47]

    Yves Van Vaerenbergh and Troy D. Thomas. Response styles in survey research: A literature review of antecedents, consequences, and remedies.International Journal of Public Opinion Research, 25(2):195–217, 2013. doi: 10.1093/ijpor/eds021

  40. [48]

    Krosnick and Duane F

    Jon A. Krosnick and Duane F. Alwin. An evaluation of a cognitive theory of response-order effects in survey measurement. Public Opinion Quarterly, 51(2):201–219, 1987. doi: 10.1086/ 269029

  41. [49]

    Krosnick

    Jon A. Krosnick. Response strategies for coping with the cogni- tive demands of attitude measures in surveys.Applied Cognitive Psychology, 5(3):213–236, 1991. doi: 10.1002/acp.2350050305

  42. [50]

    Nelder.Generalized Linear Models

    Peter McCullagh and John A. Nelder.Generalized Linear Models. Chapman and Hall, London, 2nd edition, 1989

  43. [51]

    Danish Institute for Educational Research, Copenhagen, 1960

    Georg Rasch.Probabilistic Models for Some Intelligence and Attainment Tests. Danish Institute for Educational Research, Copenhagen, 1960. Reissued 1980, Chicago: University of Chicago Press, foreword by Benjamin D. Wright

  44. [52]

    Some latent trait models and their use in inferring an examinee’s ability

    Allan Birnbaum. Some latent trait models and their use in inferring an examinee’s ability. In Frederic M. Lord and Melvin R. Novick, editors,Statistical Theories of Mental Test Scores, pages 397–479. Addison-Wesley, Reading, MA, 1968

  45. [53]

    Lord.Applications of Item Response Theory to Prac- tical Testing Problems

    Frederic M. Lord.Applications of Item Response Theory to Prac- tical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980

  46. [54]

    Embretson and Steven P

    Susan E. Embretson and Steven P. Reise.Item Response The- ory for Psychologists. Multivariate Applications Book Series. Lawrence Erlbaum Associates, Mahwah, NJ, 2000

  47. [55]

    large language models show amplified cogni- tive biases in moral decision-making

    Maximilian Maier, Vanessa Cheung, and Falk Lieder. Code and data for analyses in “large language models show amplified cogni- tive biases in moral decision-making”. Open Science Framework, https://osf.io/3kvjd/, 2025. Deposited 4 November 2025

  48. [56]

    Rothkopf, and Kristian Kersting

    Patrick Schramowski, Cigdem Turan, Nico Andersen, Con- stantin A. Rothkopf, and Kristian Kersting. Large pre-trained language models contain human-like biases of what is right and wrong to do.Nature Machine Intelligence, 4(3):258–268, 2022. doi: 10.1038/s42256-022-00458-8

  49. [57]

    Moral mimicry: Large language models pro- duce moral rationalizations tailored to political identity

    Gabriel Simmons. Moral mimicry: Large language models pro- duce moral rationalizations tailored to political identity. InPro- ceedings of the 61st Annual Meeting of the Association for Com- putational Linguistics (Volume 4: Student Research Workshop), pages 282–297, Toronto, C...

  50. [58]

    Guilherme F. C. F. Almeida, Jos´ e Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de Ara´ ujo. Exploring the psychol- ogy of LLMs’ moral and legal reasoning.Artificial Intelligence, 333:104145, August 2024. doi: 10.1016/j.artint.2024.104145

  51. [59]

    Language model alignment in multilingual trolley problems

    Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez, Francesco Ortu, Andr´ as Strausz, 13 Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, and Bernhard Sch¨ olkopf. Language model alignment in multilingual trolley problems. InInternational ...

  52. [60]

    Who is GPT-3? An exploration of personality, values and demographics

    Maril` u Miotto, Nicola Rossberg, and Bennett Kleinberg. Who is GPT-3? An exploration of personality, values and demographics. InProceedings of the Fifth Workshop on Natural Language Pro- cessing and Computational Social Science (NLP+CSS), pages 218–227, Abu Dhabi, UAE, Novemb...

  53. [61]

    The self-perception and po- litical biases of ChatGPT.Human Behavior and Emerging Technologies, 2024:1–9, 2024

    J´ erˆ ome Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, Moritz Roidl, and Markus Pauly. The self-perception and po- litical biases of ChatGPT.Human Behavior and Emerging Technologies, 2024:1–9, 2024. doi: 10.1155/2024/7115633

  54. [62]

    CMoralEval: A moral evaluation benchmark for Chinese large language mod- els

    Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Tao Liu, and Deyi Xiong. CMoralEval: A moral evaluation benchmark for Chinese large language mod- els. InFindings of the Association for Computationa...

  55. [63]

    Llm ethics benchmark: A three-dimensional assessment system for evaluating moral reasoning in large language models.Scientific Reports, 15:34642,

    Junfeng Jiao, Saleh Afroogh, Abhejay Murali, Kevin Chen, David Atkinson, and Amit Dhurandhar. Llm ethics benchmark: A three-dimensional assessment system for evaluating moral reasoning in large language models.Scientific Reports, 15:34642,

  56. [64]

    doi: 10.1038/s41598-025-18489-7

  57. [65]

    Evaluating moral beliefs across LLMs through a pluralistic framework

    Xuelin Liu, Yanfei Zhu, Shucheng Zhu, Pengyuan Liu, Ying Liu, and Dong Yu. Evaluating moral beliefs across LLMs through a pluralistic framework. InFindings of the Association for Com- putational Linguistics: EMNLP 2024, pages 4740–4760, Miami, Florida, USA, November 2024. Asso...

  58. [66]

    Cultural value alignment in large language mod- els: A prompt-based analysis of Schwartz values in Gemini, ChatGPT, and DeepSeek, 2025

    Robin Segerer. Cultural value alignment in large language mod- els: A prompt-based analysis of Schwartz values in Gemini, ChatGPT, and DeepSeek, 2025

  59. [67]

    Large language models are not fair evaluators

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1...

  60. [68]

    Discovering language model behaviors with model-written evaluations

    Ethan Perez, Sam Ringer, Kamil˙ e Lukoˇ si¯ ut˙ e, Karina Nguyen, et al. Discovering language model behaviors with model-written evaluations. InFindings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434, 2023

  61. [69]

    Bowman, Newton Cheng, Esin Dur- mus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Dur- mus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang,...

  62. [70]

    Primacy effect of ChatGPT

    Yiwei Wang, Yujun Cai, Muhao Chen, Yuxuan Liang, and Bryan Hooi. Primacy effect of ChatGPT. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, pages 108–115, Singapore, December 2023. ...

  63. [71]

    Prompt perturbations reveal human-like biases in large language model survey responses.arXiv preprint arXiv:2507.07188, 2025

    Jens Rupprecht, Georg Ahnert, and Markus Strohmaier. Prompt perturbations reveal human-like biases in large language model survey responses.arXiv preprint arXiv:2507.07188, 2025

  64. [72]

    Acquiescence bias in large language models

    Daniel Braun. Acquiescence bias in large language models. InFindings of the Association for Computational Linguistics: EMNLP 2025, 2025

  65. [73]

    Thilo Hagendorff, Ishita Dasgupta, Marcel Binz, Stephanie C. Y. Chan, Andrew Lampinen, Jane X. Wang, Zeynep Akata, and Eric Schulz. Machine psychology.arXiv preprint arXiv:2303.13988, 2023

  66. [74]

    Using cognitive psychology to understand GPT-3.Proceedings of the National Academy of Sci- ences, 120(6):e2218523120, 2023

    Marcel Binz and Eric Schulz. Using cognitive psychology to understand GPT-3.Proceedings of the National Academy of Sci- ences, 120(6):e2218523120, 2023. doi: 10.1073/pnas.2218523120

  67. [75]

    Wang, and Eric Schulz

    Julian Coda-Forno, Marcel Binz, Jane X. Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofPMLR, pages 9076–9108, 2024

  68. [76]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023. doi: 10.1017/pan.2023.2

  69. [77]

    Arriaga, and Adam Tauman Kalai

    Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofPMLR, pages 337–371, 2023

  70. [78]

    Yeager, Christopher J

    Dorottya Demszky, Diyi Yang, David S. Yeager, Christopher J. Bryan, Margarett Clapper, Susannah Chandhok, Johannes C. Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann John- son, Michaela Jones, Danielle Krettek-Cobb, Leslie Lai, Nirel JonesMitchell, Desmond C. Ong, Carol S....

  71. [79]

    Talking about large language models.Com- munications of the ACM, 67(2):68–79, 2024

    Murray Shanahan. Talking about large language models.Com- munications of the ACM, 67(2):68–79, 2024. doi: 10.1145/ 3624724

  72. [80]

    Lisa Messeri and M. J. Crockett. Artificial intelligence and illusions of understanding in scientific research.Nature, 627: 49–58, 2024. doi: 10.1038/s41586-024-07146-0

  73. [81]

    Can AI language models replace human participants?Trends in Cognitive Sciences, 27(7):597–600, 2023

    Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can AI language models replace human participants?Trends in Cognitive Sciences, 27(7):597–600, 2023. doi: 10.1016/j.tics. 2023.04.008

  74. [82]

    AI language model rivals expert ethicist in perceived moral expertise.Scientific Reports, 15:4084, 2025

    Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. AI language model rivals expert ethicist in perceived moral expertise.Scientific Reports, 15:4084, 2025. doi: 10.1038/ s41598-025-86510-0

  75. [83]

    Brady, Caelan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo

    Eyal Aharoni, Sharlene Fernandes, Daniel J. Brady, Caelan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo. Attributions toward artificial agents in a modified Moral Turing Test.Scientific Reports, 14: 8458, 2024. doi: 10.1038/s41598-024-58087-7

  76. [84]

    The moral machine experiment.Nature, 563(7729): 59–64, 2018

    Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-Fran¸ cois Bonnefon, and Iyad Rahwan. The moral machine experiment.Nature, 563(7729): 59–64, 2018. doi: 10.1038/s41586-018-0637-6

  77. [85]

    Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier

    Max Pellert, Clemens M. Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. AI psychometrics: Assess- ing the psychological profiles of large language models through psychometric inventories.Perspectives on Psychological Science, 19(5):808–826, 2024. doi: 10.11...

  78. [86]

    Ullman, Fernando Martinez-Plumed, Joshua B

    Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Ru- tar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernand...

  79. [87]

    doi: 10.1126/science.adf6369

  80. [88]

    Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna

    Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. AI and the ev- erything in the whole wide world benchmark. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS 2021), 2021

  81. [89]

    Do large language model benchmarks test reliability?, 2025

    Joshua Vendrow, Edward Vendrow, Sara Beery, and Aleksander Madry. Do large language model benchmarks test reliability?, 2025. 14

  82. [90]

    Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, Mar´ ıa Grandury, Simeng Han, Valentin Hof...

  83. [91]

    Lost in bench- marks? rethinking large language model benchmarking with item response theory

    Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Wei Bao, Jian Dong, Bing Xu, Conghui Zhu, Hailong Cao, and Tiejun Zhao. Lost in bench- marks? rethinking large language model benchmarking with item response theory. InProceedings of the AA...

  84. [92]

    Establishing construct validity in LLM capa- bility benchmarks requires nomological networks, 2026

    Timo Freiesleben. Establishing construct validity in LLM capa- bility benchmarks requires nomological networks, 2026

  85. [93]

    Six fallacies in substituting large language models for human participants.Advances in Methods and Practices in Psychological Science, 8(3), 2025

    Zhicheng Lin. Six fallacies in substituting large language models for human participants.Advances in Methods and Practices in Psychological Science, 8(3), 2025. doi: 10.1177/ 25152459251357566

  86. [94]

    Being blind (or not) to scenarios used in sacrificial dilemmas: the influence of factual and contextual information on moral responses.Frontiers in Psychology, 15:1477825, 2024

    Robin Carron, Emmanuelle Brigaud, Royce Anders, and Nathalie Blanc. Being blind (or not) to scenarios used in sacrificial dilemmas: the influence of factual and contextual information on moral responses.Frontiers in Psychology, 15:1477825, 2024. doi: 10.3389/fpsyg.2024.1477825

  87. [95]

    Cohen and Philip T

    Dale J. Cohen and Philip T. Quinlan. Why moral judgements change across variations of trolley-like problems.British Journal of Psychology, 2025. doi: 10.1111/bjop.12782. Published online 18 February 2025; volume/issue/pages not yet assigned (re- checked via Crossref 2026-07-03)

  88. [96]

    Christensen and Antoni Gomila

    Julia F. Christensen and Antoni Gomila. Moral dilemmas in cognitive neuroscience of moral decision-making: A principled review.Neuroscience & Biobehavioral Reviews, 36(4):1249–1264,

  89. [97]

    doi: 10.1016/j.neubiorev.2012.02.008

  90. [98]

    How stable are moral judgments?Review of Philosophy and Psychology, 14(4): 1377–1403, 2023

    Paul Rehren and Walter Sinnott-Armstrong. How stable are moral judgments?Review of Philosophy and Psychology, 14(4): 1377–1403, 2023. doi: 10.1007/s13164-022-00649-7

  91. [99]

    Discrepancies between judgment and choice of action in moral dilemmas.Frontiers in Psychology, 4:250, 2013

    S´ ebastien Tassy, Olivier Oullier, Julien Mancini, and Bruno Wicker. Discrepancies between judgment and choice of action in moral dilemmas.Frontiers in Psychology, 4:250, 2013. doi: 10.3389/fpsyg.2013.00250

  92. [100]

    What we say and what we do: The relationship between real and hypothetical moral choices

    Oriel FeldmanHall, Dean Mobbs, Davy Evans, Lucy Hiscox, Lauren Navrady, and Tim Dalgleish. What we say and what we do: The relationship between real and hypothetical moral choices. Cognition, 123(3):434–441, 2012. doi: 10.1016/j.cognition.2012. 02.001

  93. [101]

    Francis, Charles Howard, Ian S

    Kathryn B. Francis, Charles Howard, Ian S. Howard, Michaela Gummerum, Giorgio Ganis, Grace Anderson, and Sylvia Terbeck. Virtual morality: Transitioning from moral judgment to moral action?PLOS ONE, 11(10):e0164374, 2016. doi: 10.1371/ journal.pone.0164374

  94. [102]

    Self-reports: How the questions shape the answers.American Psychologist, 54(2):93–105, 1999

    Norbert Schwarz. Self-reports: How the questions shape the answers.American Psychologist, 54(2):93–105, 1999. doi: 10. 1037/0003-066x.54.2.93

  95. [103]

    Rips, and Kenneth A

    Roger Tourangeau, Lance J. Rips, and Kenneth A. Rasinski.The Psychology of Survey Response. Cambridge University Press, Cambridge, UK, 2000. ISBN 0521572460

  96. [104]

    Bradburn, and Norbert Schwarz

    Seymour Sudman, Norman M. Bradburn, and Norbert Schwarz. Thinking about Answers: The Application of Cognitive Processes to Survey Methodology. Jossey-Bass, San Francisco, 1996. ISBN 0787901202

  97. [105]

    Greene, R

    Joshua D. Greene, R. Brian Sommerville, Leigh E. Nystrom, John M. Darley, and Jonathan D. Cohen. An fmri investigation of emotional engagement in moral judgment.Science, 293(5537): 2105–2108, 2001. doi: 10.1126/science.1062872

  98. [106]

    How (and where) does moral judgment work?Trends in Cognitive Sciences, 6(12): 517–523, 2002

    Joshua Greene and Jonathan Haidt. How (and where) does moral judgment work?Trends in Cognitive Sciences, 6(12): 517–523, 2002. doi: 10.1016/S1364-6613(02)02011-9

  99. [107]

    Action, outcome, and value: A dual-system framework for morality.Personality and Social Psychology Review, 17(3):273–292, 2013

    Fiery Cushman. Action, outcome, and value: A dual-system framework for morality.Personality and Social Psychology Review, 17(3):273–292, 2013. doi: 10.1177/1088868313495594

  100. [108]

    The intuitive greater good: Testing the corrective dual process model of moral cognition

    Bence Bago and Wim De Neys. The intuitive greater good: Testing the corrective dual process model of moral cognition. Journal of Experimental Psychology: General, 148(10):1782– 1801, 2019. doi: 10.1037/xge0000533

  101. [109]

    Guy Kahane, Jim A. C. Everett, Brian D. Earp, Lucius Caviola, Nadira S. Faber, Molly J. Crockett, and Julian Savulescu. Be- yond sacrificial harm: A two-dimensional model of utilitarian psychology.Psychological Review, 125(2):131–164, 2018. doi: 10.1037/rev0000093

  102. [110]

    Hippler, Elisabeth Noelle-Neumann, and Leslie Clark

    Norbert Schwarz, B¨ arbel Kn¨ auper, Hans-J. Hippler, Elisabeth Noelle-Neumann, and Leslie Clark. Rating scales: Numeric values may change the meaning of scale labels.Public Opinion Quarterly, 55(4):570–582, 1991. doi: 10.1086/269282

  103. [111]

    K¨ uhnel

    Jan Karem H¨ ohne, Dagmar Krebs, and Steffen-M. K¨ uhnel. Measurement properties of completely and end labeled unipo- lar and bipolar scales in Likert-type questions on income (in)equality.Social Science Research, 97:102544, 2021. doi: 10.1016/j.ssresearch.2021.102544

  104. [112]

    K¨ uhnel

    Jan Karem H¨ ohne, Dagmar Krebs, and Steffen-M. K¨ uhnel. Mea- suring income (in)equality: Comparing survey questions with unipolar and bipolar scales in a probability-based online panel. Social Science Computer Review, 40(1):108–123, 2022. doi: 10.1177/0894439320902461

  105. [113]

    Burdick, Connie M

    Richard K. Burdick, Connie M. Borror, and Douglas C. Mont- gomery. A review of methods for measurement systems capability analysis.Journal of Quality Technology, 35(4):342–354, 2003. doi: 10.1080/00224065.2003.11980232

  106. [114]

    Cronbach, Nageswari Rajaratnam, and Goldine C

    Lee J. Cronbach, Nageswari Rajaratnam, and Goldine C. Gleser. Theory of generalizability: A liberalization of reliability theory. British Journal of Statistical Psychology, 16(2):137–163, 1963. doi: 10.1111/j.2044-8317.1963.tb00206.x

  107. [115]

    Cronbach, Goldine C

    Lee J. Cronbach, Goldine C. Gleser, Harinder Nanda, and Nageswari Rajaratnam.The Dependability of Behavioral Mea- surements: Theory of Generalizability for Scores and Profiles. Wiley, New York, 1972

  108. [116]

    Shavelson and Noreen M

    Richard J. Shavelson and Noreen M. Webb.Generalizability Theory: A Primer. Sage Publications, Newbury Park, CA, 1991

  109. [117]

    Brennan.Generalizability Theory

    Robert L. Brennan.Generalizability Theory. Statistics for Social and Behavioral Sciences. Springer, New York, NY, 2001. doi: 10.1007/978-1-4757-3456-0. 15

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.