REVIEW 3 major objections 5 minor 1 cited by
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read LLM-generated rationales for or against a commonsense answer systematically shift both human and LLM plausibility ratings, with CON arguments lowering scores more than PRO arguments raise them.
desk verdict The human 'PRO raises, CON lowers' claim rests on a between-pool baseline that the paper never defends; the LLM results are clean, and the question is important enough that it deserves a referee, but the human absolute shifts need a same-pool NO condition before they're credible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a controlled rationale-condition comparison built around a 1–5 Likert plausibility scale. Each sampled question-answer pair is rated under four conditions: no rationale, a PRO rationale (an argument for the answer's plausibility), a CON rationale (an argument against it), and both. The PRO and CON rationales are generated by a single LLM selected through a small human preference study; they introduce no new evidence, only highlight circumstances that would make the answer more or less likely. The argument rests on comparing mean ratings and rating distributions across conditions, then regressing rating change on initial plausibility and rationale type to isolate the anchorin
What would settle it
A matched or within-participant replication study: the same annotators rate the same question-answer pairs both with no rationale and with PRO or CON rationales, with order counterbalanced. If the mean shifts shrink to zero or reverse under this design, the original between-pool comparison, not the rationale content, was driving the reported changes.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that plausibility is not stable: for the same question-and-answer pair, the mean rating given by human judges moves in the direction of a short LLM-generated rationale. A PRO rationale reliably increases the mean rating of a distractor answer, a CON rationale reliably decreases ratings of both gold and distractor answers, and showing both produces intermediate shifts that lean negative. Humans and LLM raters follow the same general pattern, with LLMs—especially models from the same family that wrote the rationales—showing larger swings. The one sharp divergence is that human ratings of gold answers fall when a PRO rationale is added, which t
Load-bearing premise
The load-bearing premise is that the earlier NO-rationale ratings, collected from a different annotator pool in prior work, are directly comparable to the new rationale-condition ratings; if the two pools use the 1–5 scale differently, the reported shifts mix annotator differences with rationale effects.
Editorial extensions
If this is right
- If correct, LLM-generated explanations can influence human judgment even in commonsense domains where laypeople are already competent, not just in technical or unfamiliar tasks.
- The directionality is predictable: PRO rationales raise ratings, CON rationales lower them, and CON arguments have a larger per-argument effect—so opposition is more persuasive than support.
- Initial plausibility anchors the effect, meaning persuasive arguments will move already-uncertain answers the most and are unlikely to inflate already-high confidence further.
- LLM raters move more strongly than humans, and models related to the rationale generator move most, suggesting self-preference in model-based evaluation is a real confound.
- The protocol offers a method for using LLM-generated text as controlled stimuli in studying human judgment, while also implying practical safeguards are needed where LLM explanations reach people.
Reading between the lines
- Because the NO-rationale baseline came from an earlier, separate annotator group, a within-participant replication might reveal that part of the measured shift is between-group scale use rather than rationale persuasion.
- The puzzling drop in human gold-answer ratings under PRO could be tested by rewording the prompt to say the rationale supports but does not cap likelihood; if the drop disappears, the effect is a framing artifact rather than genuine persuasion.
- The negative lean of PRO+CON suggests that presenting both sides does not neutralize bias; a balanced-argument interface may still push judgments downward, with implications for debate-style AI systems.
- A testable extension would vary the source, style, or stated confidence of the rationale and measure whether persuasiveness persists or decays over delay, separating immediate compliance from genuine belief change.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Building on the plausibility-rating framework of Palta et al. (2024), the authors ask whether LLM-generated rationales—PRO (favoring an answer), CON (against an answer), or PRO+CON—change human and LLM plausibility ratings for answer choices in two commonsense multiple-choice benchmarks (SIQA and CQA). They sample 200 (q,a) pairs, generate rationales with GPT-4o (selected by a small preference study), and collect 3,000 new human judgments plus 13,600 LLM judgments from 17 models across four conditions (NO, PRO, CON, PRO+CON). They report that PRO rationales generally raise mean ratings and CON rationales lower them, with mixed effects for PRO+CON, and use chi-squared tests and OLS regression to support significance and an anchoring effect of initial ratings. The paper frames the results as evidence that LLM rationales can systematically sway human plausibility judgments, with implications for human-AI interaction and for using LLMs to probe human cognition.
Significance. If the human results were clean, this would be a valuable contribution: it demonstrates a concrete, measurable channel by which LLM-generated arguments can shift human plausibility judgments in everyday reasoning, with implications for trust, belief change, and evaluation methodology. The paper has real strengths: it commits to releasing data and code, uses a transparent prompt protocol, evaluates a broad set of 17 LLMs, and explicitly addresses self-preference by splitting OpenAI and non-OpenAI models. The LLM-side comparisons are internally controlled and are the more secure part of the paper. However, the central human comparison is threatened by the use of an external NO-rationale baseline from a different annotator pool, and the anchoring analysis is confounded by boundary/regression-to-the-mean mechanics. The significance of the human-side findings therefore depends on re-analysis with a same-pool control or a clear reframing of the claims.
major comments (3)
- [§3, Table 1, Tables 3–4] The NO-rationale human baseline is taken from Palta et al. (2024), a different annotator pool with different recruitment criteria (A.4 lists US-based, English primary, no literacy difficulties, undergraduate degree, 99–100% approval, 50/50 gender; Palta et al.'s criteria are not reported). Every human Δ in Table 3 and the first two bullets of §1 compare new PRO/CON/PRO+CON ratings against this external baseline. If the two pools use the 1–5 scale differently, the 'PRO raises, CON lowers' pattern is not identified as a rationale effect. The chi-squared tests in §3.1 compare pooled distributions across pools and do not control for rater identity. The internal PRO-vs-CON contrast is more defensible, but the absolute claim vs. a no-rationale state is not. The Limitations section never mentions this baseline-comparability threat. I recommend collecting a same-pool NO condition (or a calibrati
- [§5, Table 4] The OLS regression uses the change in rating (Δ = post − pre) as the dependent variable and the NO rating as a predictor. Because the NO rating is bounded on a 1–5 Likert scale, low initial ratings have more room to increase and high initial ratings more room to decrease. The consistently negative coefficient on NO rating is therefore at least partly a floor/ceiling or regression-to-the-mean artifact. This undermines the 'strong anchoring effect' bullet in §1 and the interpretation in §5 that a higher initial assessment causes smaller changes. A null model (e.g., random rematching of NO ratings to within-condition changes, or a within-item residualized-change regression) is needed to separate anchoring from scale-boundary mechanics. The claim that anchoring is 'more pronounced for distractors' also needs standard errors or a formal test of the gold-vs-distractor coefficient difference.
- [§3.1 and §4.1] The chi-squared tests are reported without clarifying the unit of analysis or accounting for clustering. Human ratings are nested within (q,a,r) items (five raters per item), and LLM ratings are nested within models. Treating individual responses as independent across conditions and models will inflate significance; the reported p-values (e.g., 0.069, 2.17E−9) should be accompanied by effect sizes and robust standard errors. For the LLM claims, pooling all 17 models into 'OpenAI' vs 'Non-OpenAI' and running chi-squared tests across conditions conflates model-level variation with condition effects. A model-level analysis (e.g., per-model deltas with a mixed-effects model or a sign test) would be more informative. The blanket statement in §4.1 that 'all p-values<0.0001' is not supported by the reporting given the absence of full test statistics.
minor comments (5)
- [Abstract and §1, bullet 2] The claim that 'PRO rationales raise mean plausibility ratings from humans and LLM judges alike' is overstated for humans: in both SIQA and CQA, human mean ratings for gold-label answers drop under PRO (SIQA: −0.26; CQA: −0.44, Table 3). The later bullet acknowledges the exception, but the abstract and the first bullet should carry the same caveat.
- [Table 3] The column headers 'Overall Gold Label Distractor' are ambiguous; please use explicit sub-columns (e.g., Overall / Gold / Distractor) or a clearer multi-level header to make it obvious which mean and Δ correspond to each answer type.
- [§2.1] The preference study used only 4 annotators to select GPT-4o from among four models. Please report vote counts, any tie-breaking rule, and inter-annotator agreement, and note that a 4-annotator majority vote is low-powered. This does not invalidate the full-scale study's goal of showing that at least one LLM can produce persuasive rationales, but the selection step should be described with its uncertainty.
- [§3 and §4] The paper should explicitly distinguish the human NO baseline (from Palta et al. 2024, external) from the LLM NO condition (generated by the same models in the same run). The LLM-side comparison is internally controlled; making that contrast explicit would help readers calibrate the strength of the human-side claims.
- [Throughout] Minor presentation issues: 'LLaMa' and 'LLaMA' are used inconsistently; '2.17E−9' should be typeset as '2.17×10⁻⁹'; the phrase 'Your answer should just be a complete option' repeats 'answer'/'option' awkwardly in Prompts A.3–A.6; and Table 5 omits standard errors for the rationale-length coefficients.
Circularity Check
No derivation-chain circularity; the only self-citation is reuse of prior NO-rationale ratings, which is data, not a derived prediction.
full rationale
The paper's central claim is empirical: LLM-generated PRO and CON rationales shift human and LLM plausibility ratings. The reported deltas are computed against NO-rationale ratings from Palta et al. (2024), a prior study by overlapping authors. This is a self-citation, and it is load-bearing for the human deltas, but it is not circular in a derivation sense: those baseline ratings are an external empirical dataset, not a quantity defined in terms of the present paper's outputs. No equation or definition makes the predicted rating change equal to an input. The rationale-generation prompts explicitly instruct the model to argue for or against plausibility, but whether judges are actually swayed is measured independently, so the effect is not entailed by construction. The model-preference study selects GPT-4o, but the paper acknowledges this and splits OpenAI vs. Non-OpenAI models to address self-preference; this is a selection effect, not circular reasoning. The internal PRO-vs-CON contrasts rest on newly collected data and do not depend on the prior baseline. The differing-annotator NO baseline is a validity threat for the absolute deltas, but that is an experimental-design concern, not a circularity of the derivation chain. Other references to the authors' earlier work concern conventions (Likert scale) or dataset construction, not load-bearing uniqueness claims. The paper's observations are not predictions derived from fitted parameters, and no ansatz is smuggled in via self-citation. Overall, the central result has independent empirical content and does not reduce to its inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption Likert scale ratings (1-5) are a valid operationalization of plausibility
- domain assumption GPT-4o rationales represent LLM rationales for the purpose of the claim
- domain assumption The PRO and CON rationales do not introduce new evidence; they only highlight possible circumstances
- domain assumption US-based English-speaking Prolific workers with college degrees represent human common sense
Cite this review
Pith. "Pith review of Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility." pith.science (2026). https://pith.science/paper/452EOEK7
@misc{pith2026251008091,
author = {Pith},
title = {Pith review of: Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility},
year = {2026},
howpublished = {\url{https://pith.science/paper/452EOEK7}},
note = {Machine review of arXiv:2510.08091}
}
read the original abstract
We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments for or against an answer, in particular, using rationales generated by LLMs. We collect 3,000 plausibility judgments from humans and another 13,600 judgments from LLMs. Overall, we observe increases and decreases in mean human plausibility ratings in the presence of LLM-generated PRO and CON rationales, respectively, suggesting that, on the whole, human judges find these rationales convincing. Experiments with LLMs reveal similar patterns of influence. Our findings demonstrate a novel use of LLMs for studying aspects of human cognition, while also raising practical concerns that, even in domains where humans are ``experts'' (i.e., common sense), LLMs have the potential to exert considerable influence on people's beliefs.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
From Plausible to Actionable: A Position on LLM Self-Explanations
Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.
Reference graph
Works this paper leans on
-
[1]
AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, and 14 others. 2025. Yi: Open foundation models by 01.ai.Preprint, arXiv:2403.04652. Vicky Arnold, Philip A Collier, Stewart A Leech, and Steve ...
arXiv 2025
-
[2]
Primary language must be English
-
[3]
Must not have any literacy difficulties
-
[4]
Must have attained a minimum of an under- graduate level degree
-
[5]
In Findings of the Association for Computational Lin- guistics: EMNLP 2024, pages 3451–3473, Miami, Florida, USA
Plausibly problematic questions in multiple- choice benchmarks for commonsense reasoning. In Findings of the Association for Computational Lin- guistics: EMNLP 2024, pages 3451–3473, Miami, Florida, USA. Association for Computational Lin- guistics. Shramay Palta and Rachel Rudinger. 2023. FORK: A bite-sized test set for probing culinary cultural biases in...
2024
-
[6]
InThe Thirty-eighth Annual Conference on Neural Information Processing Systems
LLM evaluators recognize and favor their own generations. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems. Bhargavi Paranjape, Julian Michael, Marjan Ghazvininejad, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Prompting contrastive explana- tions for commonsense reasoning tasks. InFindings of the Association for Computat...
2021
-
[7]
Large language models help humans verify truthfulness – except when they are convincingly wrong. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers), pages 1459–1474, Mexico City, Mexico. Association for Computational Linguistics. Dirk D ...
2024
-
[9]
Must be located in the United States
Show all 14 references
-
[13]
Must have an approval rate between 99− 100%on Prolific
-
[14]
curtains drew back
We use a 50−50 split of male and female 5 5Gender as indicated on Prolific. 12 Dataset Humans OpenAI Models Non-OpenAI Models Gold Label Distractor Gold Label Distractor Gold Label Distractor SIQA −0.0025−0.0021 −0.0068−0.0058 −0.0074−0.0023 CQA −0.0042−0.0037 −0.0113−0.0046 −...
-
[2011]
InAAAI Spring Symposium: Logical Formalizations of Com- monsense Reasoning
The winograd schema challenge. InAAAI Spring Symposium: Logical Formalizations of Com- monsense Reasoning. Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah Smith, Yejin Choi, and Hannaneh Hajishirzi. 2023. Vera: A general-purpose plausibility estimation model for commonsense stat...
2023 arXiv
-
[2016]
InProceed- ings of the 14th European Conference on Computer Vision (ECCV), Amsterdam, Netherlands
Generating visual explanations. InProceed- ings of the 14th European Conference on Computer Vision (ECCV), Amsterdam, Netherlands. Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Pi- queras...
2022 arXiv
-
[2021]
InProceedings of the 2021 Con- ference on Empirical Methods in Natural Language Processing, pages 10266–10284, Online and Punta Cana, Dominican Republic
Measuring association between labels and free-text rationales. InProceedings of the 2021 Con- ference on Empirical Methods in Natural Language Processing, pages 10266–10284, Online and Punta Cana, Dominican Republic. Association for Compu- tational Linguistics. Sheng Zhang, Ra...
2021
-
[2024]
InProceedings of the 2024 Conference on Empiri- cal Methods in Natural Language Processing, pages 9483–9502, Miami, Florida, USA
Susu box or piggy bank: Assessing cultural commonsense knowledge between Ghana and the US. InProceedings of the 2024 Conference on Empiri- cal Methods in Natural Language Processing, pages 9483–9502, Miami, Florida, USA. Association for Computational Linguistics
2024
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.