REVIEW 4 major objections 5 minor 46 references
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Giving large language models more test-time compute through reasoning and self-reflection does not make them more neutral or consistent on subjective topics, and newer model versions are drifting toward a progressive-collectivist viewpoint.
desk verdict Useful new benchmark and honest empirical study; the test-time compute result is solid, but the signed ideological findings rest on unvalidated question balance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the POBs benchmark. Each polar topic is framed as a trade-off between two named extremes, every question offers a five-point Likert scale plus a refusal option, and answers map to polarity values from $-1$ to $1$, with refusals placed on the imaginary axis as $0.5i$ so they stay equidistant from either stance. On top of this, the paper defines three metrics—reliability (agreement across repeated presentations), the Non-Neutrality Index (average absolute polarity), and the Topical Consistency Index (one minus the standard deviation of question-level polarities)—and projects responses onto Progressivism–Conservatism and Individualism–Collectivism axes. The Direct, Reasoning, and Self-reflection prompts operationalize increasing test-time compute.
What would settle it
Regenerate the POBs questions with several different LLMs and have human annotators from diverse ideological backgrounds rate each item for balance and clarity, then re-run the ten models. If the progressive-collectivist lean, the reliability decline under reasoning, and the NNI-TCI correlation disappear or reverse, the paper's central claims would be artifacts of its specific question set.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that test-time compute is not a reliable guardrail for subjective neutrality. Across ten models, moving from direct answering to reasoning to self-reflection produced small and directionally inconsistent changes in the Non-Neutrality Index and Topical Consistency Index, while reliability scores generally declined. At the same time, the benchmark places most models in the progressive-collectivist quadrant of two ideological axes, and newer versions within a model family drift further in that direction while becoming less consistent; the paper reports a strong negative correlation ($r\sim 0.9$) between expressing strong opinions and staying consistent. When models are asked directly to declare their stances, they report themselves as more neutral than their POBs-inferred positions reveal.
Load-bearing premise
The load-bearing premise is that POBs' questions are neutral and balanced enough that the measured polarities reflect the models' leanings rather than question wording; the authors state the questions were not validated for balance or clarity by domain experts or human participants.
Editorial extensions
If this is right
- Users and businesses should not assume that letting a model 'think longer' will make its advice on contested topics more neutral; the measured gains are small and inconsistent.
- Upgrading to a newer version of the same model family is not behavior-preserving: consistency and ideological stance can shift, so deployments should re-audit subjective behavior after an upgrade.
- Models' self-reported neutrality is not a reliable guide to their actual stance, so indirect probes like POBs are needed to expose implicit preferences.
- The strong negative correlation between non-neutrality and topical consistency suggests a trade-off: models that commit to a position tend to commit inconsistently.
- Topic clusters correlated in model responses (e.g., adoption and surrogacy aligning with women's rights and environmentalism) offer a way to monitor where training data may be injecting an ideological slant.
Reading between the lines
- If the benchmark questions were re-generated by several LLMs and validated by human raters from both sides of each issue, the reported progressive-collectivist drift could shrink or reappear; that experiment would distinguish model bias from question-wording bias.
- The reliability drop under reasoning may be a side effect of the prompt design (asking the model to 'think' invites it to consider multiple perspectives); a different reasoning prompt might show larger or smaller gains.
- The POBs approach could be extended to measure whether expressed stances transfer to downstream recommendations, since a model that states a belief may not act on it when advising a user.
- The NNI-TCI trade-off, if it holds broadly, implies that attempts to make models more neutral may also make them more predictable, and vice versa—a relationship worth testing on future model families.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces the Preference, Opinion, and Belief survey (POBs), a benchmark of Likert-scale questions covering 20 subjective topics, and uses it to evaluate ten open- and closed-source LLMs. The authors define reliability (repeat-response agreement), a Non-Neutrality Index (NNI), a Topical Consistency Index (TCI), and ideological polarity scores on Progressivism-Conservatism and Individualism-Collectivism axes. They compare three prompting conditions—Direct, Reasoning, and Self-reflection—and report that extra test-time compute yields only limited gains in reliability, neutrality, and consistency, while newer model versions tend to be less consistent and more aligned with a progressive-collectivist stance.
Significance. The paper's strengths are its transparent experimental design (ten models, five repetitions per question, fixed prompt templates, clearly defined metrics) and the release of a new benchmark that may be useful for auditing LLM subjectivity. The empirical finding that reasoning and self-reflection improve objective tasks but not these subjective metrics is clearly stated and falsifiable. However, the benchmark's validity for measuring ideological stance is not fully established, which weakens the paper's boldest conclusions; the reference-free design also makes question-construction choices more consequential. If the validity concerns are addressed, the benchmark and metrics could be a valuable resource for model evaluation and deployment decisions.
major comments (4)
- [Section 2, Appendix A.2, Section 4.4] The signed polarity scores underlying Figure 5 and the progressive-collectivism trend can be interpreted as ideological stance only if the POBs questions are balanced with respect to acquiescence bias. The paper states that Llama-3.3-70B-Instruct generated the questions and the authors manually verified each question's polarity assignment, but it does not report whether the 'agree' end of the Likert scale points equally often to the progressive/collectivist and conservative/individualist poles, and there are no reverse-coded items. If agreement happens to align with the progressive-collectivist pole on most items, then a generic tendency to agree would masquerade as the reported ideological lean. This is a particular threat to the trend claim because Section 4.1 finds that newer models are also less reliable, indicating a change in response style. The paper's own Limitations section acknowledges that the questions 'were not validated for balance or clarity by domain experts or human participants.' I ask the authors to quantify the agreement-polarity alignment per topic, to validate balance with human raters, or to re-run the ideological analysis on a strictly balanced subset of questions.
- [Section 4.1, Tables 1-2, Figure 5] The paper's main quantitative claims are made without confidence intervals or significance tests. For instance, Table 2 reports NNI differences between prompting conditions as small as 0.01-0.03 (e.g., GPT-4o Direct 0.45 to Reasoning 0.64 is larger, but many other changes are small), and with n=5 repetitions per question it is unclear whether these are within sampling noise. Similarly, the 'newer model versions are less consistent' comparison is not controlled: LLaMA 3.2 3B and LLaMA 3.3 70B differ in size and release, so the version effect is confounded with model scale. Please provide bootstrap confidence intervals over questions/repetitions, report sign-consistency across topics, and either use matched-size version pairs or explicitly qualify the version claim.
- [Section 2, Appendix A.2, Limitations] The generation of POBs by Llama-3.3-70B-Instruct, which is itself one of the evaluated models, creates a circularity that is not addressed in the limitations. At minimum, the reported ideological position of LLaMA 3.3 70B in Figure 5 may reflect alignment with the question writer's viewpoint rather than an independent measurement. More broadly, all models are scored on a benchmark whose items were generated by a single model with its own biases. The paper should include a robustness check with an independently generated or human-written question set to show that the central findings (limited test-time-compute gains, progressive-collectivist lean, and version trend) are not an artifact of the generation process.
- [Section 4.4] The assignment of POB topics to the two ideological axes is not justified. For example, the list of Progressivism-Conservatism topics includes 'Secularism vs. Religiousness' but excludes 'Democracy vs. Alternative Governance Models' and 'AI Precautionary vs. Optimism', which are polar topics in the same dataset; the Individualism-Collectivism axis contains only four topics. Because Figure 5's coordinates are computed from these arbitrary axis definitions, the ideological positions depend on which topics the authors chose to include. I request a rationale for the topic-axis assignments (e.g., external ideological definitions or human annotation) and a sensitivity analysis showing that the main findings hold when individual topics are removed or when the axis composition is varied.
minor comments (5)
- [Section 4.1] The word 'copared' in the paragraph on topic-level reliability should be 'compared'.
- [Appendix B.1] The sentence 'we did not we exclude refusals' should read 'we did not exclude refusals'.
- [Appendix B.5] The phrase 'The results in Figure suggest' lacks a figure number; it should presumably be 'Figure 5' or another explicit reference.
- [Table 5 caption] The phrase 'cross all investigated models' should be 'across all investigated models'.
- [Throughout] The benchmark name is spelled inconsistently as both 'POBs' and 'POBS'; please standardize to one form.
Circularity Check
The paper is an empirical benchmarking study with no derivation chain that reduces to its own inputs; the central claims are measurements, not fitted predictions.
full rationale
POBs is an observational benchmark: the authors define Likert-scale polarity conventions, reliability, NNI, and TCI as explicit formulas, then report measured values for ten models under three prompting regimes. No parameter is fitted to a subset of the data and then renamed as a prediction; the test-time-compute comparison is a direct prompt-level experiment, and the ideological-axis interpretation in Section 4.4 is a post-hoc labeling of measured polarities, not a quantity derived from the labels. The only self-citation, Rabinovich et al. (2023) for the reliability metric, is not load-bearing because the metric is explicitly defined in Equations 1-2 and is a standard average pairwise absolute difference. The acknowledged limitation that the questions were generated by Llama-3.3-70B-Instruct and not validated for balance is a genuine validity threat to the ideological-lean claim, especially given possible acquiescence effects, but it is not circularity: the paper does not claim to derive the direction of bias from the question-generation process, and the measured values could in principle come out neutral or conservative. Thus no quoted equation or argument reduces any reported result to its own definitional input, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Number of repetitions n=5 =
5
- Refusal representation 0.5i =
0.5i
- Opinion-shift threshold of 1 =
1
assumptions (4)
- domain assumption Likert-scale polarity assignments are interval-scale comparable across models and topics.
- domain assumption The POBs questions, generated by Llama-3.3-70B-Instruct and manually curated, are sufficiently neutral and unbiased for measuring model stance.
- domain assumption Topic definitions and their assignment to ideological axes (Progressivism-Conservatism, Individualism-Collectivism) are valid.
- domain assumption A model's Likert-scale answer reflects its underlying preference or belief rather than a formatting artifact.
invented entities (5)
-
POBs benchmark
-
Declarative POBs
-
Non-Neutrality Index (NNI)
-
Topical Consistency Index (TCI)
-
Complex Likert Scale
Cite this review
Pith. "Pith review of Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models." pith.science (2026). https://pith.science/paper/ZYHLVSZZ
@misc{pith2026250519621,
author = {Pith},
title = {Pith review of: Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZYHLVSZZ}},
note = {Machine review of arXiv:2505.19621}
}
read the original abstract
As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibit subjective preferences, opinions, and beliefs. These tendencies may stem from biases within the models, which may shape their behavior, influence the advice and recommendations they offer to users, and potentially reinforce certain viewpoints. This paper presents the Preference, Opinion, and Belief survey (POBs), a benchmark developed to assess LLMs' subjective inclinations across societal, cultural, ethical, and personal domains. We applied our benchmark to evaluate leading open- and closed-source LLMs, measuring desired properties such as reliability, neutrality, and consistency. In addition, we investigated the effect of increasing the test-time compute, through reasoning and self-reflection mechanisms, on those metrics. While effective in other tasks, our results show that these mechanisms offer only limited gains in our domain. Furthermore, we reveal that newer model versions are becoming less consistent and more biased toward specific viewpoints, highlighting a blind spot and a concerning trend. POBS: https://ibm.github.io/POBS
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[4]
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024. Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157
arXiv 2024
-
[5]
Zhenni Bi, Kai Han, Chuanjian Liu, Yehui Tang, and Yunhe Wang. 2024. Forest-of-thought: Scaling test-time compute for enhancing llm reasoning. arXiv preprint arXiv:2412.09078
arXiv 2024
-
[6]
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. https://doi.org/10.1126/science.aal4230 Semantics derived automatically from language corpora contain human-like biases . Science, 356(6334):183--186
-
[7]
Jose G Cavazos, P Jonathon Phillips, Carlos D Castillo, and Alice J O'Toole. 2021. https://doi.org/10.1109/TBIOM.2020.3038890 Accuracy comparison across face recognition algorithms: Where are we on measuring race bias? IEEE Transactions on Biometrics, Behavior, and Identity Science, 3(1):101--111
-
[8]
Alexander S Choi, Syeda Sabrina Akter, JP Singh, and Antonios Anastasopoulos. 2024. The llm effect: Are humans truly using llms, or are they being influenced by them instead? arXiv preprint arXiv:2410.04699
work page Pith review arXiv 2024
Show all 46 references
-
[9]
Esin Durmus, Karina Nguyen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXi...
2023 arXiv
-
[10]
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Sch \"u tze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012--1031
2021
-
[11]
IBM Granite Team. 2024. Granite 3.0 language models
2024
-
[12]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[13]
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. https://doi.org/10.2139/ssrn.4316084 The political ideology of conversational ai: Converging evidence on chatgpt's pro-environmental, left-libertarian orientation . SSRN Electronic Journal
2023 doi
-
[14]
Jessica Hoffmann, Christiane Ahlheim, Zac Yu, Aria Walfrand, Jarvis Jin, Marie Tano, Ahmad Beirami, Erin van Liemt, Nithum Thain, Hakim Sidahmed, et al. 2025. Improving neutral point of view text generation through parameter-efficient reinforcement learning and a small-scale h...
2025
-
[15]
Geert Hofstede. 1984. Culture's consequences: International differences in work-related values, volume 5. sage
1984
-
[16]
Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403
2022 arXiv
-
[17]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276
2024 arXiv
-
[18]
Ian Hutchby. 2011. Non-neutrality and argument in the hybrid political interview. Discourse Studies, 13(3):349--365
2011
-
[19]
Thomas S. T. Jakobsen, Laura Cabello, and Anders S gaard. 2023. https://arxiv.org/abs/2306.00639 Being right for whose right reasons? arXiv preprint arXiv:2306.00639
2023 arXiv
-
[20]
Graham Kalton and Howard Schuman. 1982. The effect of the question on survey responses: A review. Journal of the Royal Statistical Society Series A: Statistics in Society, 145(1):42--57
1982
-
[21]
Jan Kammerath. 2024. https://medium.com/@jankammerath/deepseek-is-it-a-stolen-chatgpt-a805b586b24a Deepseek: Is it a stolen chatgpt? https://medium.com/@jankammerath/deepseek-is-it-a-stolen-chatgpt-a805b586b24a. Accessed: 2025-03-21
2024
-
[22]
Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2025. Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology, 34(2):1--23
2025
-
[23]
Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, and Roi Reichart. 2024. The colorful future of llms: Evaluating and improving llms as emotional supporters for queer youth. arXiv preprint arXiv:2402.11886
2024 arXiv
-
[24]
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al. 2024 a . Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434
2024 arXiv
-
[25]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024 b . Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
2024 arXiv
-
[26]
Ruibo Liu, Chenyan Jia, Jason Wei, Guang Xu, and Soroush Vosoughi. 2022. https://doi.org/10.1016/j.artint.2021.103654 Quantifying and alleviating political bias in language models . Artificial Intelligence, 304:103654
2022
-
[27]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2...
2022
-
[28]
Fumio Motoki, Vitor Pinho Neto, and Vitor Rodrigues. 2024. https://doi.org/10.1007/s11127-023-01097-2 More human than human: measuring chatgpt political bias . Public Choice, 198(1):3--23
2024 doi
-
[29]
Malvina Nissim, Rik van Noord, and Rob van der Goot. 2019. https://arxiv.org/abs/1905.09866 Fair is better than sensational: Man is to doctor as woman is to doctor . arXiv preprint arXiv:1905.09866
2019 arXiv
-
[30]
Malvina Nissim, Rik van Noord, and Rob van der Goot. 2020. https://doi.org/10.1162/coli_a_00379 Fair is better than sensational: Man is to doctor as woman is to doctor . Computational Linguistics, 46(2):487--497
2020 doi
-
[31]
OpenAI . 2024. https://openai.com/index/learning-to-reason-with-llms/ Learning to reason with llms . Accessed: 2025-03-18
2024
-
[32]
Peter S Park, Philipp Schoenegger, and Chongyang Zhu. 2024. https://doi.org/10.3758/s13428-023-02307-x Diminished diversity-of-thought in a standard large language model . Behavior Research Methods
2024 doi
-
[33]
Pagnarasmey Pit, Xingjun Ma, Mike Conway, Qingyu Chen, James Bailey, Henry Pit, Putrasmey Keo, Watey Diep, and Yu-Gang Jiang. 2024. Whose side are you on? investigating the political stance of large language models. arXiv preprint arXiv:2403.13840
2024 arXiv
-
[34]
Ella Rabinovich, Samuel Ackerman, Orna Raz, Eitan Farchi, and Ateret Anaby Tavor. 2023. https://aclanthology.org/2023.gem-1.12/ Predicting question-answering performance of large language models through semantic consistency . In Proceedings of the Third Workshop on Natural Lan...
2023
-
[35]
Matthew Renze and Erhan Guven. 2024. Self-reflection in large language model agents: Effects on problem-solving performance. In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), pages 516--525. IEEE
2024
-
[36]
Luca Rettenberger, Markus Reischl, and Mark Schutera. 2024. Assessing political bias in large language models. arXiv preprint arXiv:2405.13041
2024 arXiv
-
[37]
David Rozado. 2020. https://doi.org/10.1371/journal.pone.0231189 Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types . PLOS ONE, 15(4):e0231189
2020 doi
-
[38]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004. PMLR
2023
-
[39]
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314
2024 arXiv
-
[40]
Harry C Triandis. 2018. Individualism and collectivism. Routledge
2018
-
[41]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36:74952--74965
2023
-
[42]
William Voegeli. 2023. Progressivism, conservatism, and democracy. J. Contemp. Legal Issues, 24:155
2023
-
[43]
Joe H Ward Jr. 1963. Hierarchical grouping to optimize an objective function. Journal of the American statistical association, 58(301):236--244
1963
-
[44]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[45]
Xuyang Wu, Jinming Nian, Zhiqiang Tao, and Yi Fang. 2025. Evaluating social biases in llm reasoning. arXiv preprint arXiv:2502.15361
2025
-
[46]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.