{"id":"7bcc8f09-5315-4a76-8c87-cc5855d95a81","arxiv_id":"2412.09630","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM praise and critique of user-stated intentions are driven more by source trustworthiness than ideology, align broadly with human moral scores, and show no country-of-origin bias.","lead":"This paper measures how six AI chatbots praise, criticize, or stay neutral when users say they plan to do something, from supporting a politician to harming animals. It finds that news source trustworthiness matters more than left-right ideology, and that models align with human moral judgments only by giving lots of praise and criticism.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 'trustworthiness over ideology' rests on comparing coefficients on non-comparable units; standardized reanalysis is needed.","rationale":"The reader's conditional verdict is appropriate, but the most load-bearing weakness is not the GPT-3.5-turbo coding issue highlighted by the reader. Even granting perfect labels, the paper's headline result about trustworthiness versus ideology is not established by the reported statistics because it compares coefficients and marginal effects on arbitrarily scaled variables. This is directly the central claim in the abstract: 'trustworthiness is a stronger driver of praise and critique than ideology.' The paper ships code and data, which is genuinely helpful, and the non-finding on home-country bias and the human-alignment results are separate contributions. But the news-source experiment is the basis of the strongest claim, and the scale-dependence of the comparison is a concrete, fixable, and central flaw. A standardized reanalysis could either confirm the claim or substantially weaken it; the verdict should therefore remain conditional pending that check. I partially agree with the reader because their coding concern is legitimate and worth fixing, but it is not the decisive issue for the paper's main conclusion.","tokens_in":23928,"tokens_out":5841,"duration_ms":55308,"concrete_test":"Re-run the Experiment I ordered-logit and OLS regressions with ideology and trustworthiness both standardized (z-scored, or scaled to equal ranges) before fitting, using the public replication code and data at the GitHub/OSF repositories. Compute the standardized coefficients and the average marginal effects per one-standard-deviation change for both Ad Fontes and AllSides measures, for all six models covering Tables 2, 3, 7, 8, and 9. If, after standardization, trustworthiness no longer has a larger effect than ideology for Llama-3-70B and Qwen-1.5-32B, and if the AllSides measure makes ideology dominant for Llama-3-70B, the abstract's claim that trustworthiness is the stronger driver must be weakened or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that trustworthiness is a stronger driver of praise and critique than ideology. The evidence for this in Experiment I is presented in Section 3.3, Tables 2 and 3, and the AllSides robustness tables in the appendix. In every case, the comparison is made between raw ordered-logit coefficients and average marginal effects for a one-unit change in ideology versus a one-unit change in trustworthiness. These units are arbitrary and not commensurable: Ad Fontes ideology runs from -28 to 44, Ad Fontes trustworthiness from 1 to 62, and AllSides ideology from -2 to 2. A one-unit shift on one scale is not equivalent to a one-unit shift on the other, so ratios such as 'trustworthiness is five times stronger' in Table 3 are scale-dependent statements, not substantive findings. Using the standard deviations reported in Table 5 (ideology SD=18.03, trustworthiness SD=14.98 for Ad Fontes), the standardized effect of trustworthiness shrinks relative to the raw coefficient comparison by about a factor of 1.2. For Llama-3-70B (raw ordered-logit coefficients -0.009 versus 0.013) and Qwen-1.5-32B (-0.013 versus 0.017 in Table 2), standardized coefficients become nearly equal, contradicting the 'often by a factor of 2 or more' claim. Under the AllSides measure, the mismatch is even more severe: Table 8 gives Llama-3-70B an ideology coefficient of -0.156 and a trustworthiness coefficient of 0.007, which after standardization implies trustworthiness is weaker than ideology for that model. The central claim therefore needs a scale-invariant reanalysis before it can be accepted. The GPT-3.5-turbo coding concern identified by the reader is real but secondary; this is an internal statistical validity issue that affects the headline result even if the labels are perfect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a behavioral evaluation of LLM moral stances by analyzing how six LLMs respond to user-stated intentions across three domains: news sources (testing whether ideology or trustworthiness drives praise), everyday ethical actions (comparing model praise to human moral ratings), and world leaders (testing country-of-origin bias). Responses are coded as praise, neutral, or critique, and the paper reports that trustworthiness dominates ideology in news-source evaluations, that model praise correlates strongly with human moral judgments, and that there is no evidence of same-country favoritism. The paper also identifies a 'reticence-alignment tradeoff,' notably in Claude-3-Sonnet, and provides open replication code and data.","tokens_in":24275,"tokens_out":8654,"duration_ms":73597,"significance":"If the findings hold, the paper offers a novel and ecologically valid measurement of implicit LLM moral judgments, with implications for AI alignment and for monitoring the psychological and societal effects of conversational AI. Strengths include the use of six diverse models, contrast-set prompting, multiple ideology measures, explicit robustness checks, and a replication repository. The central 'trustworthiness over ideology' claim, however, is currently supported by comparisons on non-commensurable units, and the measurement of the outcome variable depends on one of the evaluated models as coder; both issues require additional analysis before the headline findings can be considered established.","major_comments":[{"comment":"The headline finding that 'trustworthiness is a stronger driver than ideology' is based on comparing raw ordered-logit coefficients and average marginal effects for a one-unit increase in each variable. These units are arbitrary: Ad Fontes ideology spans -28 to 44, Ad Fontes trustworthiness spans 1 to 62, and AllSides ideology spans -2 to 2. A one-unit change is not comparable across scales, so the ratios in Table 3 (e.g., 6.7, 10.2) and in Table 9 are scale-dependent statements. Using the standard deviations in Table 5, the standardized coefficients for Llama-3-70B under Ad Fontes are approximately -0.162 for ideology and 0.195 for trustworthiness (ratio about 1.2), and for Qwen-1.5-32B approximately -0.234 and 0.255. Under AllSides, Llama-3-70B's standardized ideology coefficient (-0.223) is more than twice its standardized trustworthiness coefficient (0.105), directly contradicting the abstract's general claim. Please redo the comparison using standardized coefficients or comparable quantile/percentile shifts, and revise the abstract and Section 3.3 conclusions accordingly.","section":"Section 3.3, Tables 2-3 and Appendix Tables 8-9"},{"comment":"All outcome variables are coded by GPT-3.5-turbo, which is itself one of the six models under evaluation. If the coder's judgments are systematically different when applied to its own outputs than to other models' outputs, then the praise scores, engagement rates, and cross-model comparisons reported in Tables 2-4 and 11-14 are not comparable. The manual review was limited to ambiguous responses, described as less than one percent, so a systematic bias in the remaining responses would not be detected. Please validate the coding by (a) obtaining human annotations on a random sample of outputs from all six models and reporting agreement statistics, and/or (b) re-running the analysis with an independent coder model that is not among the evaluated six; either would allow an assessment of coder-induced bias.","section":"Section 3.2 and all experiments"},{"comment":"The paper acknowledges that the Schramowski et al. dataset has been public since September 2021 and 'may have been incorporated into the training data,' which 'raises the possibility that our results may overstate the true extent of alignment.' This limitation is central to the second experiment's claim of strong human-model alignment, so a caveat is not sufficient. Because the human moral scores are public and fixed, the observed Spearman correlations of 0.65-0.81 could reflect memorization of the score pattern rather than a general property of praise. Please provide a robustness check that is not susceptible to this contamination, for example by evaluating the models on newly constructed action statements with fresh human ratings, or by testing on actions whose human moral scores were not part of the public dataset before the models' training cutoffs.","section":"Section 4.4"}],"minor_comments":[{"comment":"There are typos: 'responsvie companion' should be 'responsive companion,' and 'The remained of this paper' should be 'The remainder of this paper.'","section":"Section 1"},{"comment":"The phrase 'one one by a French company' should be 'one by a French company.'","section":"Section 5.2"},{"comment":"The cutpoints are labeled '0/1' and '1/2' while the text describes outcomes coded as -1, 0, 1; please clarify that the outcome was recoded to 0, 1, 2 for this regression.","section":"Appendix Table 14"},{"comment":"The sentence 'To disambiguate this use of \"encouraging,\" fourth example of a negative response (−1)...' appears grammatically incomplete; please rephrase.","section":"Section 4.2"},{"comment":"The Ad Fontes ratings used are from 2019, while the LLM evaluations were conducted in 2024; please note this temporal mismatch explicitly in the data description.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper is that the behavioral method is the real contribution, and the headline claim about trustworthiness versus ideology is overstated. The author proposes measuring implicit moral stances by coding how LLMs praise, stay neutral, or critique user-stated intentions ('I'm thinking of campaigning for X'), across news sources, everyday actions, and world leaders, with six models. That is a genuinely useful and distinct lens, and the paper ships code and data, which sets a good example.\n\nThe news-source experiment is the strongest part. The contrast-set design, the inversion of negative prompts, the engagement rates, and the Ad Fontes/AllSides robustness checks are thoughtful. The observation that the naive correlation between praise and right-wing ideology is mediated by trustworthiness is interesting and worth taking seriously.\n\nThe soft spot is that the central comparison rests on non-comparable units. Table 3 reports raw ordered-logit coefficients and AMEs for a one-unit change in ideology versus a one-unit change in trustworthiness. Those scales are arbitrary: Ad Fontes ideology runs -28 to 44 and trustworthiness 1 to 62. Standardizing with the SDs in Table 5, the 'factor of 2 or more' for Llama-3-70B and Qwen shrinks to near parity; under AllSides, Llama-3's ideology coefficient is actually larger than trustworthiness. The author does acknowledge the Llama-3/AllSides exception in the text, but the abstract and conclusion generalize too strongly. A standardized reanalysis, or a comparison in terms of meaningful differences (e.g., 1 SD), is needed before the headline can be trusted.\n\nThe second issue is the coding. All responses were coded by GPT-3.5-turbo, which is itself one of the evaluated models. The author says ambiguous cases were manually reviewed and final codes assigned by a human, and that ambiguous cases were rare (<1%), but there is no validation against an independent human-coded sample or inter-coder agreement. That is a fixable measurement dependency, not a fatal flaw.\n\nThe moral-action experiment is plausible, and the author explicitly acknowledges the training-data contamination possibility, which is good practice. The world-leader experiment is a reasonable null result.\n\nWho should read this: anyone studying AI alignment evaluation, chatbot behavior auditing, or political bias in LLMs. It deserves a serious referee because the method is novel and the replication package is solid, but I would ask for the standardized reanalysis and coding validation before accepting the central claim.","headline":"A useful behavioral method for auditing LLM moral stances, but the headline 'trustworthiness over ideology' is not scale-invariant and needs a standardized reanalysis.","tokens_in":24777,"tokens_out":3542,"would_cite":false,"duration_ms":31836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the praise and critique LLMs give to users' stated intentions is a measurable moral stance: trustworthiness drives news praise more than ideology, human moral scores predict praise for everyday actions, and no…","keywords":["praise and critique","moral landscape","user-stated intentions","ideological bias","trustworthiness","LLM alignment","political bias","human moral judgments"],"falsifier":"Take a matched set of news sources where left- and right-leaning outlets have equal trustworthiness scores and re-estimate the statistical model; the claim that trustworthiness outweighs ideology would be falsified if the ideology coefficient consistently exceeded the trustworthiness coefficient at moderate trust levels. A second direct check would be to recode a random sample of all six models' raw outputs with human annotators and verify that the same praise indices, engagement rates, and model orderings emerge, since the coding model is itself one of the evaluated models.","tokens_in":23746,"feed_emoji":"🧭","tokens_out":4799,"duration_ms":41950,"temperature":0.7,"pith_summary":"This paper argues that the way chatbots respond to users' stated intentions—with praise, critique, or neutrality—is a measurable form of moral judgment, and that this behavior can be mapped across politics, ethics, and world leaders. It claims that when ideology and trustworthiness are separated, LLMs' praise of news sources tracks source reliability more strongly than left-right ideology, so an appearance of anti-right bias is largely a side effect of right-leaning sources being less trustworthy on average. It further claims that LLMs' praise and criticism of everyday ethical actions correlate strongly with human moral ratings, but that achieving this value alignment requires high engagement, creating a tradeoff between alignment and reticence. Finally, it reports no detectable favoritism by the six models toward leaders of their home country. If true, these findings suggest that LLMs' spontaneous praise is a normative stance that deserves monitoring.","feed_headline":"Trustworthiness, not ideology, drives chatbot praise","feed_subtitle":"Across six models, praise of news sources tracks source reliability, aligns with human moral ratings, and shows no home-country bias.","key_machinery":"The machinery is the praise-score measurement: every LLM response to a user-stated intention is coded on a three-point scale (+1 praise or encouragement, 0 neutral or factual, −1 critique or discouragement), with negative prompts inverted so that support and opposition to the same target are comparable. Contrasting prompt pairs, such as promoting versus opposing a named source or leader, and contextual variants are generated for each target, and ordered logit regression separates the contributions of ideology, trustworthiness, and other controls. A second load-bearing component is the human moral score from Schramowski et al., which grounds the ethical-actions experiment. For the coding itself, GPT-3.5-turbo assigns the initial labels, with human review of ambiguous cases.","core_discovery":"The central discovery is that praise and critique responses to user-stated intentions constitute a window into the implicit moral positions of LLMs. Testing six widely used models on constructed prompts such as \"I'm thinking of campaigning for {candidate}\" or \"I've decided to leave my partner,\" the paper codes each response as +1 (praise), 0 (neutral), or −1 (critique), and finds that models engage normatively most of the time. In the news experiment, once source trustworthiness is included in an ordered logit model, its marginal effects on praise are typically two to five times larger than the effects of ideology, and for most models the ideology coefficient is negligible or insignificant. On everyday actions, Spearman correlations between model praise scores and human moral ratings range from about 0.65 to 0.81 across models, without large outliers. On world leaders, a same-country indicator is not statistically significant, indicating no strong national-origin bias. The paper therefore claims that the apparent anti-right slant of LLMs is better described as an anti-untrustworthiness slant, that models are broadly human-aligned in their implicit moral praise, and that the price of that alignment is a refusal to stay neutral on morally loaded statements.","pith_inferences":["An extension the paper leaves implicit is that the same praise-score method could be run in languages other than English, where the paper's own anecdotal evidence suggests decisions are sometimes framed as revisable rather than final, which would test whether the moral landscape is language-dependent.","The coding bottleneck could be turned into a strength by using multiple coder LLMs and measuring inter-coder agreement, which would quantify how much of the measured landscape belongs to the evaluator rather than to the models being evaluated.","A testable extension for the trustworthiness result would construct synthetic news sources that combine high or low trustworthiness with left or right labels so that ideology and quality are fully orthogonal, rather than relying on the natural correlation in existing media ratings.","The absence of home-country bias may be specific to generic statements about leaders; probing policy-specific positions such as trade, climate, or human rights could reveal national or regional patterns that the aggregate measure washes out."],"forward_implications":["Apparent ideological bias in LLM responses should not be read as left-right bias without accounting for source quality, because evaluations that control for trustworthiness can change the conclusion.","Models that aim to be value-aligned will often need to praise or criticize users, so policies that simply instruct models to stay neutral on contested topics conflict with alignment on everyday ethical decisions.","The reticence-alignment tradeoff suggests that a model designed to be unbiased by staying silent is not truly neutral in effect: silence itself becomes a normative choice when users announce morally relevant plans.","Because praise and critique patterns vary across models and over time, monitoring LLM engagement levels should be part of responsible deployment rather than treated as a stylistic afterthought."],"supporting_citations":[{"why":"Supplies the human moral scores that the ethical-actions experiment aligns against.","marker":"Schramowski et al. (2022)"},{"why":"Describes the Constitutional AI approach used to explain Claude-3-sonnet's reticence and the reticence-alignment tradeoff.","marker":"Bai et al. (2022)"},{"why":"Introduces contrast sets, the method behind the opposing prompt variants used to make praise scores comparable.","marker":"Gardner et al. (2020)"},{"why":"Provides the ordered logit regression framework for treating the three-level praise codes as ordinal outcomes.","marker":"Gelman and Hill (2007)"},{"why":"Earlier evidence of ChatGPT political bias that the paper challenges by adding trustworthiness controls.","marker":"Motoki, Pinho Neto, and Rodrigues (2024)"},{"why":"Documents the scale of companion-AI use, motivating why praise behavior toward users matters.","marker":"Hadero (2024)"},{"why":"Anecdotal case of a chatbot endorsing a user's stated harmful intention, motivating the behavioral lens on user-stated intentions.","marker":"Singleton, Gerken, and McMahon (2023)"},{"why":"Evidence on LLM political persuasion, supporting the concern that praise and critique can influence users.","marker":"Hackenburg et al. (2024)"}],"fun_headline_variants":["Chatbot praise tracks trustworthiness, not ideology","AI praise is driven by trust, not political slant","LLM praise aligns with human moral ratings","Praise from AI: trust matters more than politics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every praise, neutral, or critique label used in the analysis was first assigned by GPT-3.5-turbo—one of the models under evaluation—so if its judgments are biased toward its own style of responding, all model comparisons inherit that bias.","fun_headline_variants_meta":{"raw":{"variants":["Chatbot praise tracks trustworthiness, not ideology","AI praise is driven by trust, not political slant","LLM praise aligns with human moral ratings","Praise from AI: trust matters more than politics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000469,"raw_usage":{"total_tokens":2396,"prompt_tokens":1063,"completion_tokens":1333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":1273}},"tokens_in":679,"tokens_out":1333,"duration_ms":8146,"temperature":1.0,"reasoning_tokens":1273,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:10:48.051834+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a matched set of news sources where left- and right-leaning outlets have equal trustworthiness scores and re-estimate the statistical model; the claim that trustworthiness outweighs ideology would be falsified if the ideology coefficient consistently exceeded the trustworthiness coefficient at moderate trust levels. A second direct check would be to recode a random sample of all six models' raw outputs with human annotators and verify that the same praise indices, engagement rates, and model orderings emerge, since the coding model is itself one of the evaluated models.","supporting_citations":[],"review_version":1}