Pith. sign in

REVIEW 3 major objections 5 minor 15 references

What does AI consider praiseworthy?

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper argues that the praise and critique LLMs give to users' stated intentions is a measurable moral stance: trustworthiness drives news praise more than ideology, human moral scores predict praise for everyday actions, and no…

desk verdict A useful behavioral method for auditing LLM moral stances, but the headline 'trustworthiness over ideology' is not scale-invariant and needs a standardized reanalysis. read the letter →

arxiv 2412.09630 v2 pith:XMS6VBL6 submitted 2024-11-27 cs.CY cs.HC

classification cs.CYcs.HC
keywords praiseandcritiquemorallandscapeuser-statedintentionsideologicalbiastrustworthinessLLMalignmentpoliticalhumanjudgments
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the way chatbots respond to users' stated intentions—with praise, critique, or neutrality—is a measurable form of moral judgment, and that this behavior can be mapped across politics, ethics, and world leaders. It claims that when ideology and trustworthiness are separated, LLMs' praise of news sources tracks source reliability more strongly than left-right ideology, so an appearance of anti-right bias is largely a side effect of right-leaning sources being less trustworthy on average. It further claims that LLMs' praise and criticism of everyday ethical actions correlate strongly with human moral ratings, but that achieving this value alignment requires high engagement, creating a tradeoff between alignment and reticence. Finally, it reports no detectable favoritism by the six models toward leaders of their home country. If true, these findings suggest that LLMs' spontaneous praise is a normative stance that deserves monitoring.

What carries the argument

The machinery is the praise-score measurement: every LLM response to a user-stated intention is coded on a three-point scale (+1 praise or encouragement, 0 neutral or factual, −1 critique or discouragement), with negative prompts inverted so that support and opposition to the same target are comparable. Contrasting prompt pairs, such as promoting versus opposing a named source or leader, and contextual variants are generated for each target, and ordered logit regression separates the contributions of ideology, trustworthiness, and other controls. A second load-bearing component is the human moral score from Schramowski et al., which grounds the ethical-actions experiment. For the coding itself, GPT-3.5-turbo assigns the initial labels, with human review of ambiguous cases.

What would settle it

Take a matched set of news sources where left- and right-leaning outlets have equal trustworthiness scores and re-estimate the statistical model; the claim that trustworthiness outweighs ideology would be falsified if the ideology coefficient consistently exceeded the trustworthiness coefficient at moderate trust levels. A second direct check would be to recode a random sample of all six models' raw outputs with human annotators and verify that the same praise indices, engagement rates, and model orderings emerge, since the coding model is itself one of the evaluated models.

Watch

Extended reading notes

Core claim

The central discovery is that praise and critique responses to user-stated intentions constitute a window into the implicit moral positions of LLMs. Testing six widely used models on constructed prompts such as "I'm thinking of campaigning for {candidate}" or "I've decided to leave my partner," the paper codes each response as +1 (praise), 0 (neutral), or −1 (critique), and finds that models engage normatively most of the time. In the news experiment, once source trustworthiness is included in an ordered logit model, its marginal effects on praise are typically two to five times larger than the effects of ideology, and for most models the ideology coefficient is negligible or insignificant. On everyday actions, Spearman correlations between model praise scores and human moral ratings range from about 0.65 to 0.81 across models, without large outliers. On world leaders, a same-country indicator is not statistically significant, indicating no strong national-origin bias. The paper therefore claims that the apparent anti-right slant of LLMs is better described as an anti-untrustworthiness slant, that models are broadly human-aligned in their implicit moral praise, and that the price of that alignment is a refusal to stay neutral on morally loaded statements.

Load-bearing premise

Every praise, neutral, or critique label used in the analysis was first assigned by GPT-3.5-turbo—one of the models under evaluation—so if its judgments are biased toward its own style of responding, all model comparisons inherit that bias.

Editorial extensions

If this is right

  • Apparent ideological bias in LLM responses should not be read as left-right bias without accounting for source quality, because evaluations that control for trustworthiness can change the conclusion.
  • Models that aim to be value-aligned will often need to praise or criticize users, so policies that simply instruct models to stay neutral on contested topics conflict with alignment on everyday ethical decisions.
  • The reticence-alignment tradeoff suggests that a model designed to be unbiased by staying silent is not truly neutral in effect: silence itself becomes a normative choice when users announce morally relevant plans.
  • Because praise and critique patterns vary across models and over time, monitoring LLM engagement levels should be part of responsible deployment rather than treated as a stylistic afterthought.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension the paper leaves implicit is that the same praise-score method could be run in languages other than English, where the paper's own anecdotal evidence suggests decisions are sometimes framed as revisable rather than final, which would test whether the moral landscape is language-dependent.
  • The coding bottleneck could be turned into a strength by using multiple coder LLMs and measuring inter-coder agreement, which would quantify how much of the measured landscape belongs to the evaluator rather than to the models being evaluated.
  • A testable extension for the trustworthiness result would construct synthetic news sources that combine high or low trustworthiness with left or right labels so that ideology and quality are fully orthogonal, rather than relying on the natural correlation in existing media ratings.
  • The absence of home-country bias may be specific to generic statements about leaders; probing policy-specific positions such as trade, climate, or human rights could reveal national or regional patterns that the aggregate measure washes out.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a behavioral evaluation of LLM moral stances by analyzing how six LLMs respond to user-stated intentions across three domains: news sources (testing whether ideology or trustworthiness drives praise), everyday ethical actions (comparing model praise to human moral ratings), and world leaders (testing country-of-origin bias). Responses are coded as praise, neutral, or critique, and the paper reports that trustworthiness dominates ideology in news-source evaluations, that model praise correlates strongly with human moral judgments, and that there is no evidence of same-country favoritism. The paper also identifies a 'reticence-alignment tradeoff,' notably in Claude-3-Sonnet, and provides open replication code and data.

Significance. If the findings hold, the paper offers a novel and ecologically valid measurement of implicit LLM moral judgments, with implications for AI alignment and for monitoring the psychological and societal effects of conversational AI. Strengths include the use of six diverse models, contrast-set prompting, multiple ideology measures, explicit robustness checks, and a replication repository. The central 'trustworthiness over ideology' claim, however, is currently supported by comparisons on non-commensurable units, and the measurement of the outcome variable depends on one of the evaluated models as coder; both issues require additional analysis before the headline findings can be considered established.

major comments (3)
  1. [Section 3.3, Tables 2-3 and Appendix Tables 8-9] The headline finding that 'trustworthiness is a stronger driver than ideology' is based on comparing raw ordered-logit coefficients and average marginal effects for a one-unit increase in each variable. These units are arbitrary: Ad Fontes ideology spans -28 to 44, Ad Fontes trustworthiness spans 1 to 62, and AllSides ideology spans -2 to 2. A one-unit change is not comparable across scales, so the ratios in Table 3 (e.g., 6.7, 10.2) and in Table 9 are scale-dependent statements. Using the standard deviations in Table 5, the standardized coefficients for Llama-3-70B under Ad Fontes are approximately -0.162 for ideology and 0.195 for trustworthiness (ratio about 1.2), and for Qwen-1.5-32B approximately -0.234 and 0.255. Under AllSides, Llama-3-70B's standardized ideology coefficient (-0.223) is more than twice its standardized trustworthiness coefficient (0.105), directly contradicting the abstract's general claim. Please redo the comparison using standardized coefficients or comparable quantile/percentile shifts, and revise the abstract and Section 3.3 conclusions accordingly.
  2. [Section 3.2 and all experiments] All outcome variables are coded by GPT-3.5-turbo, which is itself one of the six models under evaluation. If the coder's judgments are systematically different when applied to its own outputs than to other models' outputs, then the praise scores, engagement rates, and cross-model comparisons reported in Tables 2-4 and 11-14 are not comparable. The manual review was limited to ambiguous responses, described as less than one percent, so a systematic bias in the remaining responses would not be detected. Please validate the coding by (a) obtaining human annotations on a random sample of outputs from all six models and reporting agreement statistics, and/or (b) re-running the analysis with an independent coder model that is not among the evaluated six; either would allow an assessment of coder-induced bias.
  3. [Section 4.4] The paper acknowledges that the Schramowski et al. dataset has been public since September 2021 and 'may have been incorporated into the training data,' which 'raises the possibility that our results may overstate the true extent of alignment.' This limitation is central to the second experiment's claim of strong human-model alignment, so a caveat is not sufficient. Because the human moral scores are public and fixed, the observed Spearman correlations of 0.65-0.81 could reflect memorization of the score pattern rather than a general property of praise. Please provide a robustness check that is not susceptible to this contamination, for example by evaluating the models on newly constructed action statements with fresh human ratings, or by testing on actions whose human moral scores were not part of the public dataset before the models' training cutoffs.
minor comments (5)
  1. [Section 1] There are typos: 'responsvie companion' should be 'responsive companion,' and 'The remained of this paper' should be 'The remainder of this paper.'
  2. [Section 5.2] The phrase 'one one by a French company' should be 'one by a French company.'
  3. [Appendix Table 14] The cutpoints are labeled '0/1' and '1/2' while the text describes outcomes coded as -1, 0, 1; please clarify that the outcome was recoded to 0, 1, 2 for this regression.
  4. [Section 4.2] The sentence 'To disambiguate this use of "encouraging," fourth example of a negative response (−1)...' appears grammatically incomplete; please rephrase.
  5. [Section 3.1] The Ad Fontes ratings used are from 2019, while the LLM evaluations were conducted in 2024; please note this temporal mismatch explicitly in the data description.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper's claims rest on external ratings, independent human-moral labels, and LLM outputs that are not defined in terms of the quantities they are used to predict.

full rationale

I walked the paper's derivation chain for each experiment. In Experiment I, praise scores are human-supervised annotations of LLM responses, and ideology and trustworthiness are taken from external sources (Ad Fontes and AllSides); no equation defines praise as a function of ideology or trustworthiness. The ordered-logit and marginal-effect comparisons are statistical summaries, not constructions of the outcome from the predictors. In Experiment II, the human moral scores come from Schramowski et al. (2022), an external dataset, and the LLM praise scores are computed from model responses to independently constructed prompts; the reported correlations are empirical, not identity-based. In Experiment III, the country-of-origin test uses an external list of leaders and model responses, with no parameter fitted to the outcome. The use of GPT-3.5-turbo as an initial annotator is a measurement-dependence concern because that model is also one of the evaluated systems, but ambiguous responses were manually reviewed by a human, and the coding step does not make any of the paper's conclusions true by construction. Similarly, the concern that Ad Fontes and AllSides use non-comparable units is a statistical interpretation issue about comparing coefficients across scales, not a circularity. There are no self-citations carrying a load-bearing argument, no fitted parameter renamed as a prediction, and no result that is equivalent to its inputs by definition.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central empirical claims rest on measurement assumptions rather than fitted parameters. No novel theoretical entities are introduced; the key assumptions concern the validity of external rating scales, the symmetry of prompt inversion, and the unbiasedness of the automated coder.

assumptions (5)
  • domain assumption Ad Fontes Media and AllSides ratings are valid measures of news source ideology and trustworthiness.
    Used as ground truth in Experiment I (Section 3.1); the author acknowledges the proprietary and not fully replicable methodology of Ad Fontes.
  • domain assumption The Schramowski et al. human moral scores are valid ground truth for the moral valence of everyday actions.
    Used as benchmark in Experiment II (Section 4.1); based on 234 survey participants and may have entered LLM training data, which the author notes in Section 4.4.
  • domain assumption Inverting responses to negative prompts yields a symmetric praise measure.
    Section 3.2 constructs the praise score as positive prompt score minus inverted negative prompt score; the very large negative-prompt coefficients in Table 2 indicate the prompts are not behaviorally symmetric.
  • domain assumption LLM responses to the constructed prompts are stable enough within a session to support aggregate scores.
    Each item is measured with a small set of prompt variations; API randomness and model updates are not modeled, as described in Section 3.2.
  • ad hoc to paper GPT-3.5-turbo provides unbiased coding of praise and critique for all six models, including itself.
    Section 3.2 states that initial coding was performed using OpenAI's GPT-3.5-turbo, with manual review only of ambiguous cases; no independent human coding of all responses is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What does AI consider praiseworthy?." pith.science (2026). https://pith.science/paper/XMS6VBL6

@misc{pith2026241209630,
  author       = {Pith},
  title        = {Pith review of: What does AI consider praiseworthy?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XMS6VBL6}},
  note         = {Machine review of arXiv:2412.09630}
}
read the original abstract

As large language models (LLMs) are increasingly used for work, personal, and therapeutic purposes, researchers have begun to investigate these models' implicit and explicit moral views. Previous work, however, focuses on asking LLMs to state opinions, or on other technical evaluations that do not reflect common user interactions. We propose a novel evaluation of LLM behavior that analyzes responses to user-stated intentions, such as "I'm thinking of campaigning for {candidate}." LLMs frequently respond with critiques or praise, often beginning responses with phrases such as "That's great to hear!..." While this makes them friendly, these praise responses are not universal and thus reflect a normative stance by the LLM. We map out the moral landscape of LLMs in how they respond to user statements in different domains including politics and everyday ethical actions. In particular, although a na\"ive analysis might suggest LLMs are biased against right-leaning politics, our findings on news sources indicate that trustworthiness is a stronger driver of praise and critique than ideology. Second, we find strong alignment across models in response to ethically-relevant action statements, but that doing so requires them to engage in high levels of praise and critique of users, suggesting a reticence-alignment tradeoff. Finally, our experiment on statements about world leaders finds no evidence of bias favoring the country of origin of the models. We conclude that as AI systems become more integrated into society, their patterns of praise, critique, and neutrality must be carefully monitored to prevent unintended psychological and societal consequences.

Figures

Figures reproduced from arXiv: 2412.09630 by the authors.

Figure 1
Figure 1. OLS predicted probability by model, with 95% confidence intervals. The ‘negative prompt’ is set to [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Moral Actions - Correlation with Human Evaluations by Model. The x-axis is the praise score, [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Moral Actions: Correlation and Engagement by Model, Category. The x-axis represents the [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: International Politicians: Top- and Bottom-8 Leaders by Average Score. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: International Politicians: From Selected States and the EU [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Ad Fontes vs. AllSides Measure of News Ideology [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: OLS predicted probability by model, with 95% confidence intervals (AllSides ideology measure). [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: News sources: Praise score residualized on trustworthiness, by ideology [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 4 canonical work pages

  1. [7]

    Annual Review of Psychology 75(Volume 75, 2024):433–466

    The Neuroscience of Human and Artificial Intelligence Presence. Annual Review of Psychology 75(Volume 75, 2024):433–466. Publisher: Annual Reviews. Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021a. Aligning {ai} with shared human values. In International Conference on Learning Representations. 29 Andrew Peterson...

  2. [8]

    arXiv preprint arXiv:2406.09279

    Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback. arXiv preprint arXiv:2406.09279. Jiang, L.; Hwang, J. D.; Bhagavatula, C.; Bras, R. L.; Liang, J.; Dodge, J.; Sakaguchi, K.; Forbes, M.; Borchardt, J.; Gabriel, S.; et al

  3. [9]

    arXiv preprint arXiv:2110.07574

    Can machines learn morality? the delphi experiment. arXiv preprint arXiv:2110.07574. Jin, Z.; Levine, S.; Gonzalez Adauto, F.; Kamal, O.; Sap, M.; Sachan, M.; Mihalcea, R.; Tenenbaum, J.; and Schölkopf, B

  4. [10]

    In Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, HRI ’21 Companion, 10–18

    Robots as Moral Advisors: The Effects of Deontological, Virtue, and Confucian Role Ethics on Encouraging Honest Behavior. In Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, HRI ’21 Companion, 10–18. New York, NY, USA: Association for Computing Machinery. Kirk, H. R.; Vidgen, B.; Röttger, P .; and Hale, S. A

  5. [11]

    arXiv preprint arXiv:2402.04249

    Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249. Motoki, F.; Pinho Neto, V .; and Rodrigues, V

  6. [12]

    arXiv preprint arXiv:2403.13313

    Polaris: A safety-focused llm constellation architecture for healthcare. arXiv preprint arXiv:2403.13313. Naous, T.; Ryan, M. J.; Ritter, A.; and Xu, W

  7. [14]

    Xu, B., and Zhuang, Z

    Can large language model agents simulate human trust behaviors? arXiv preprint arXiv:2402.04559. Xu, B., and Zhuang, Z

  8. [15]

    arXiv preprint arXiv:2204.03021

    The moral integrity corpus: A benchmark for ethical dialogue systems. arXiv preprint arXiv:2204.03021. 31

Show all 15 references
  1. [2002]

    Ubiquity 2002(December):2

    Persuasive technology: using computers to change what we think and do. Ubiquity 2002(December):2. Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; and Ahmed, N. K

  2. [2006]

    Oxford University Press, USA

    Principled agents?: The political economy of good government. Oxford University Press, USA. Bisbee, J.; Clinton, J.; Dorff, C.; Kenkel, B.; and Larson, J. 2023a. Artificially precise extremism: how internet-trained llms exaggerate our differences. SocArXiv Preprint (https://do...

  3. [2020]

    arXiv preprint arXiv:2004.02709

    Evaluating models’ local decision boundaries via contrast sets. arXiv preprint arXiv:2004.02709. Gelman, A., and Hill, J

  4. [2021]

    Technical report, European Union

    Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts. Technical report, European Union. COM(2021) 206 final. (FAIR)†, M. F. A. R...

  5. [2022]

    arXiv preprint arXiv:2212.08073

    Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. Bail, C. A

  6. [2023]

    arXiv preprint arXiv:2305.14456

    Having beer after prayer? measuring cultural bias in large language models. arXiv preprint arXiv:2305.14456. Nazer, L. H.; Zatarah, R.; Waldrip, S.; Ke, J. X. C.; Moukheiber, M.; Khanna, A. K.; Hicklen, R. S.; Moukheiber, L.; Moukheiber, D.; Ma, H.; and Mathur, P

  7. [2024]

    arXiv preprint arXiv:2406.14508

    Evidence of a log scaling law for political persuasion with large language models. arXiv preprint arXiv:2406.14508. Hadar-Shoval, D.; Asraf, K.; Mizrachi, Y.; Haber, Y.; and Elyoseph, Z

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.