Pith. sign in

REVIEW 4 major objections 5 minor 46 references

Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Giving large language models more test-time compute through reasoning and self-reflection does not make them more neutral or consistent on subjective topics, and newer model versions are drifting toward a progressive-collectivist viewpoint.

desk verdict Useful new benchmark and honest empirical study; the test-time compute result is solid, but the signed ideological findings rest on unvalidated question balance. read the letter →

arxiv 2505.19621 v1 pith:ZYHLVSZZ submitted 2025-05-26 cs.AI cs.CL

classification cs.AIcs.CL
keywords POBsbenchmarktest-timecomputeLLMsubjectivityideologicalbiasreliabilityneutralityself-reflectionopinionbenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Giving large language models more time to think—through explicit reasoning and self-reflection before answering—does not make them more neutral, reliable, or consistent on subjective and controversial topics. To test this, the authors build the Preference, Opinion, and Belief Survey (POBs), a benchmark of Likert-scale questions spanning 20 societal, cultural, ethical, and personal topics, and run it on ten open- and closed-source models under three prompting regimes: direct answering, reasoning, and self-reflection. The extra compute changes answers but delivers only limited and inconsistent gains: it does not reliably reduce bias or improve consistency, and it often lowers reliability. The paper also finds that newer versions of the same model family tend to be less consistent and more strongly aligned with a progressive-collectivist viewpoint, and that models underreport their own biases when asked directly.

What carries the argument

The engine is the POBs benchmark. Each polar topic is framed as a trade-off between two named extremes, every question offers a five-point Likert scale plus a refusal option, and answers map to polarity values from $-1$ to $1$, with refusals placed on the imaginary axis as $0.5i$ so they stay equidistant from either stance. On top of this, the paper defines three metrics—reliability (agreement across repeated presentations), the Non-Neutrality Index (average absolute polarity), and the Topical Consistency Index (one minus the standard deviation of question-level polarities)—and projects responses onto Progressivism–Conservatism and Individualism–Collectivism axes. The Direct, Reasoning, and Self-reflection prompts operationalize increasing test-time compute.

What would settle it

Regenerate the POBs questions with several different LLMs and have human annotators from diverse ideological backgrounds rate each item for balance and clarity, then re-run the ten models. If the progressive-collectivist lean, the reliability decline under reasoning, and the NNI-TCI correlation disappear or reverse, the paper's central claims would be artifacts of its specific question set.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that test-time compute is not a reliable guardrail for subjective neutrality. Across ten models, moving from direct answering to reasoning to self-reflection produced small and directionally inconsistent changes in the Non-Neutrality Index and Topical Consistency Index, while reliability scores generally declined. At the same time, the benchmark places most models in the progressive-collectivist quadrant of two ideological axes, and newer versions within a model family drift further in that direction while becoming less consistent; the paper reports a strong negative correlation ($r\sim 0.9$) between expressing strong opinions and staying consistent. When models are asked directly to declare their stances, they report themselves as more neutral than their POBs-inferred positions reveal.

Load-bearing premise

The load-bearing premise is that POBs' questions are neutral and balanced enough that the measured polarities reflect the models' leanings rather than question wording; the authors state the questions were not validated for balance or clarity by domain experts or human participants.

Editorial extensions

If this is right

  • Users and businesses should not assume that letting a model 'think longer' will make its advice on contested topics more neutral; the measured gains are small and inconsistent.
  • Upgrading to a newer version of the same model family is not behavior-preserving: consistency and ideological stance can shift, so deployments should re-audit subjective behavior after an upgrade.
  • Models' self-reported neutrality is not a reliable guide to their actual stance, so indirect probes like POBs are needed to expose implicit preferences.
  • The strong negative correlation between non-neutrality and topical consistency suggests a trade-off: models that commit to a position tend to commit inconsistently.
  • Topic clusters correlated in model responses (e.g., adoption and surrogacy aligning with women's rights and environmentalism) offer a way to monitor where training data may be injecting an ideological slant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the benchmark questions were re-generated by several LLMs and validated by human raters from both sides of each issue, the reported progressive-collectivist drift could shrink or reappear; that experiment would distinguish model bias from question-wording bias.
  • The reliability drop under reasoning may be a side effect of the prompt design (asking the model to 'think' invites it to consider multiple perspectives); a different reasoning prompt might show larger or smaller gains.
  • The POBs approach could be extended to measure whether expressed stances transfer to downstream recommendations, since a model that states a belief may not act on it when advising a user.
  • The NNI-TCI trade-off, if it holds broadly, implies that attempts to make models more neutral may also make them more predictable, and vice versa—a relationship worth testing on future model families.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces the Preference, Opinion, and Belief survey (POBs), a benchmark of Likert-scale questions covering 20 subjective topics, and uses it to evaluate ten open- and closed-source LLMs. The authors define reliability (repeat-response agreement), a Non-Neutrality Index (NNI), a Topical Consistency Index (TCI), and ideological polarity scores on Progressivism-Conservatism and Individualism-Collectivism axes. They compare three prompting conditions—Direct, Reasoning, and Self-reflection—and report that extra test-time compute yields only limited gains in reliability, neutrality, and consistency, while newer model versions tend to be less consistent and more aligned with a progressive-collectivist stance.

Significance. The paper's strengths are its transparent experimental design (ten models, five repetitions per question, fixed prompt templates, clearly defined metrics) and the release of a new benchmark that may be useful for auditing LLM subjectivity. The empirical finding that reasoning and self-reflection improve objective tasks but not these subjective metrics is clearly stated and falsifiable. However, the benchmark's validity for measuring ideological stance is not fully established, which weakens the paper's boldest conclusions; the reference-free design also makes question-construction choices more consequential. If the validity concerns are addressed, the benchmark and metrics could be a valuable resource for model evaluation and deployment decisions.

major comments (4)
  1. [Section 2, Appendix A.2, Section 4.4] The signed polarity scores underlying Figure 5 and the progressive-collectivism trend can be interpreted as ideological stance only if the POBs questions are balanced with respect to acquiescence bias. The paper states that Llama-3.3-70B-Instruct generated the questions and the authors manually verified each question's polarity assignment, but it does not report whether the 'agree' end of the Likert scale points equally often to the progressive/collectivist and conservative/individualist poles, and there are no reverse-coded items. If agreement happens to align with the progressive-collectivist pole on most items, then a generic tendency to agree would masquerade as the reported ideological lean. This is a particular threat to the trend claim because Section 4.1 finds that newer models are also less reliable, indicating a change in response style. The paper's own Limitations section acknowledges that the questions 'were not validated for balance or clarity by domain experts or human participants.' I ask the authors to quantify the agreement-polarity alignment per topic, to validate balance with human raters, or to re-run the ideological analysis on a strictly balanced subset of questions.
  2. [Section 4.1, Tables 1-2, Figure 5] The paper's main quantitative claims are made without confidence intervals or significance tests. For instance, Table 2 reports NNI differences between prompting conditions as small as 0.01-0.03 (e.g., GPT-4o Direct 0.45 to Reasoning 0.64 is larger, but many other changes are small), and with n=5 repetitions per question it is unclear whether these are within sampling noise. Similarly, the 'newer model versions are less consistent' comparison is not controlled: LLaMA 3.2 3B and LLaMA 3.3 70B differ in size and release, so the version effect is confounded with model scale. Please provide bootstrap confidence intervals over questions/repetitions, report sign-consistency across topics, and either use matched-size version pairs or explicitly qualify the version claim.
  3. [Section 2, Appendix A.2, Limitations] The generation of POBs by Llama-3.3-70B-Instruct, which is itself one of the evaluated models, creates a circularity that is not addressed in the limitations. At minimum, the reported ideological position of LLaMA 3.3 70B in Figure 5 may reflect alignment with the question writer's viewpoint rather than an independent measurement. More broadly, all models are scored on a benchmark whose items were generated by a single model with its own biases. The paper should include a robustness check with an independently generated or human-written question set to show that the central findings (limited test-time-compute gains, progressive-collectivist lean, and version trend) are not an artifact of the generation process.
  4. [Section 4.4] The assignment of POB topics to the two ideological axes is not justified. For example, the list of Progressivism-Conservatism topics includes 'Secularism vs. Religiousness' but excludes 'Democracy vs. Alternative Governance Models' and 'AI Precautionary vs. Optimism', which are polar topics in the same dataset; the Individualism-Collectivism axis contains only four topics. Because Figure 5's coordinates are computed from these arbitrary axis definitions, the ideological positions depend on which topics the authors chose to include. I request a rationale for the topic-axis assignments (e.g., external ideological definitions or human annotation) and a sensitivity analysis showing that the main findings hold when individual topics are removed or when the axis composition is varied.
minor comments (5)
  1. [Section 4.1] The word 'copared' in the paragraph on topic-level reliability should be 'compared'.
  2. [Appendix B.1] The sentence 'we did not we exclude refusals' should read 'we did not exclude refusals'.
  3. [Appendix B.5] The phrase 'The results in Figure suggest' lacks a figure number; it should presumably be 'Figure 5' or another explicit reference.
  4. [Table 5 caption] The phrase 'cross all investigated models' should be 'across all investigated models'.
  5. [Throughout] The benchmark name is spelled inconsistently as both 'POBs' and 'POBS'; please standardize to one form.

Circularity Check

0 steps flagged · score 0.0 of 10

The paper is an empirical benchmarking study with no derivation chain that reduces to its own inputs; the central claims are measurements, not fitted predictions.

full rationale

POBs is an observational benchmark: the authors define Likert-scale polarity conventions, reliability, NNI, and TCI as explicit formulas, then report measured values for ten models under three prompting regimes. No parameter is fitted to a subset of the data and then renamed as a prediction; the test-time-compute comparison is a direct prompt-level experiment, and the ideological-axis interpretation in Section 4.4 is a post-hoc labeling of measured polarities, not a quantity derived from the labels. The only self-citation, Rabinovich et al. (2023) for the reliability metric, is not load-bearing because the metric is explicitly defined in Equations 1-2 and is a standard average pairwise absolute difference. The acknowledged limitation that the questions were generated by Llama-3.3-70B-Instruct and not validated for balance is a genuine validity threat to the ideological-lean claim, especially given possible acquiescence effects, but it is not circularity: the paper does not claim to derive the direction of bias from the question-generation process, and the measured values could in principle come out neutral or conservative. Thus no quoted equation or argument reduces any reported result to its own definitional input, and the appropriate circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 5 invented entities

The central claims rest on the benchmark design, the chosen metric parameters, and the author-defined topic/axis mappings. None are fitted to data, but each introduces a modeling choice that could influence the results. The weakest point is the assumption that LLM-generated, author-curated questions are neutral enough to measure bias without a human or expert baseline.

free parameters (3)
  • Number of repetitions n=5 = 5
    The reliability and polarity metrics average over n=5 repeated responses per question. This value is chosen, not fitted, and affects the stability of all reported scores.
  • Refusal representation 0.5i = 0.5i
    The authors place refusals on the imaginary axis at 0.5i to keep them equidistant from agreement and disagreement. This is a representational choice that influences reliability distances.
  • Opinion-shift threshold of 1 = 1
    In Section 4.4, a 'substantial opinion change' is defined as a polarity shift greater than 1. The threshold value is arbitrary and affects the reported shift percentages in Figure 9.
assumptions (4)
  • domain assumption Likert-scale polarity assignments are interval-scale comparable across models and topics.
    Polarity values (-1, -0.5, 0, 0.5, 1) are treated as equidistant in reliability and NNI/TCI calculations (Section 4.1-4.2). There is no evidence that models perceive these options as equally spaced.
  • domain assumption The POBs questions, generated by Llama-3.3-70B-Instruct and manually curated, are sufficiently neutral and unbiased for measuring model stance.
    The authors acknowledge in Appendix A.2 that questions were generated by an LLM and not validated by domain experts. The observed progressive-collectivism could be an artifact of question wording.
  • domain assumption Topic definitions and their assignment to ideological axes (Progressivism-Conservatism, Individualism-Collectivism) are valid.
    Section 4.4 maps POBs topics to two axes based on the authors' judgment and cited literature (Voegeli, Triandis). Alternative mappings would change the reported ideological positions.
  • domain assumption A model's Likert-scale answer reflects its underlying preference or belief rather than a formatting artifact.
    The entire analysis treats the selected letter as an expression of stance. Prompt phrasing, role identity, or refusal options could confound this.
invented entities (5)
  • POBs benchmark
    purpose: A survey dataset of 20 topics with Likert-scale questions to measure LLM subjective preferences, opinions, and beliefs.
    POBs is introduced in this paper and not independently validated by other studies. Its neutrality and completeness are assumed.
  • Declarative POBs
    purpose: A small survey that directly asks a model to self-report its alignment on each polar topic, used to compare self-declared vs revealed stances.
    Introduced in Appendix B.5; its single-question-per-topic format has no external validation.
  • Non-Neutrality Index (NNI)
    purpose: A metric quantifying the average absolute polarity of model responses within a topic.
    Simple averaging of absolute polarities; no external benchmark establishes its validity as a bias measure.
  • Topical Consistency Index (TCI)
    purpose: A metric measuring how consistent a model's stance is across questions within the same polar topic.
    Computed as 1 minus the standard deviation of question-level average polarities; not externally validated.
  • Complex Likert Scale
    purpose: A representation placing refusals on the imaginary axis (0.5i) to keep them distinct from neutral responses in distance calculations.
    This is a representational choice introduced here; its effect on scores has not been studied elsewhere.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models." pith.science (2026). https://pith.science/paper/ZYHLVSZZ

@misc{pith2026250519621,
  author       = {Pith},
  title        = {Pith review of: Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZYHLVSZZ}},
  note         = {Machine review of arXiv:2505.19621}
}
read the original abstract

As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibit subjective preferences, opinions, and beliefs. These tendencies may stem from biases within the models, which may shape their behavior, influence the advice and recommendations they offer to users, and potentially reinforce certain viewpoints. This paper presents the Preference, Opinion, and Belief survey (POBs), a benchmark developed to assess LLMs' subjective inclinations across societal, cultural, ethical, and personal domains. We applied our benchmark to evaluate leading open- and closed-source LLMs, measuring desired properties such as reliability, neutrality, and consistency. In addition, we investigated the effect of increasing the test-time compute, through reasoning and self-reflection mechanisms, on those metrics. While effective in other tasks, our results show that these mechanisms offer only limited gains in our domain. Furthermore, we reveal that newer model versions are becoming less consistent and more biased toward specific viewpoints, highlighting a blind spot and a concerning trend. POBS: https://ibm.github.io/POBS

Figures

Figures reproduced from arXiv: 2505.19621 by the authors.

Figure 1
Figure 1. Examples of model responses to Likert-scale [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. NNI vs. TCI across different prompting approaches. A strong negative correlation indicates that models become more inconsistent as they express stronger opinions. Newer versions within a model fam￾ily exhibit lower neutrality and reduced consistency. to topic t, i.e., over all questions q ∈ Qt . We use the average polarity to disregard the variance in answers polarity between different repetitions. T CIt(m) = 1 − ST… view at source ↗
Figure 3
Figure 3. Visualizing NNI vs. TCI for polar topics in [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: Ideological stances of models on the Progres [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: The Complex Likert Scale. Demonstrating the relative distances between answers in the complex plane; Strong (-1, 1) and weak responses (-0.5, 0.5), Neutral (0) and Refused (0.5i). B.2 Topical Correlation and Clustering The dendrogram heatmap in [PITH_FULL_IMAGE:figure…
Figure 7
Figure 7. Figure 7: Heatmap of model distance Based on polarity [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Models’ impartiality. The percentage of neu [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: The percentage of substantial opinion change [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Reliability of model responses across different topics. Following the definition of a question-level [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]
Figure 11
Figure 11. Figure 11: Topics where LLMs exhibit the highest NNI in their response to direct prompt, showing the relative [PITH_FULL_IMAGE:figures/full_fig_p016_11.png]
Figure 12
Figure 12. Figure 12: Ranking of topical consistency of models in direct prompting, while showing the relative model [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 13
Figure 13. Figure 13: Heatmap of models’ response average polarity by topic. The polarity of responses is displayed along [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

46 extracted references · 16 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024. Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157

  5. [5]

    Zhenni Bi, Kai Han, Chuanjian Liu, Yehui Tang, and Yunhe Wang. 2024. Forest-of-thought: Scaling test-time compute for enhancing llm reasoning. arXiv preprint arXiv:2412.09078

  6. [6]

    Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. https://doi.org/10.1126/science.aal4230 Semantics derived automatically from language corpora contain human-like biases . Science, 356(6334):183--186

  7. [7]

    Jose G Cavazos, P Jonathon Phillips, Carlos D Castillo, and Alice J O'Toole. 2021. https://doi.org/10.1109/TBIOM.2020.3038890 Accuracy comparison across face recognition algorithms: Where are we on measuring race bias? IEEE Transactions on Biometrics, Behavior, and Identity Science, 3(1):101--111

  8. [8]

    Alexander S Choi, Syeda Sabrina Akter, JP Singh, and Antonios Anastasopoulos. 2024. The llm effect: Are humans truly using llms, or are they being influenced by them instead? arXiv preprint arXiv:2410.04699

Show all 46 references
  1. [9]

    Esin Durmus, Karina Nguyen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXi...

  2. [10]

    Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Sch \"u tze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012--1031

  3. [11]

    IBM Granite Team. 2024. Granite 3.0 language models

  4. [12]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  5. [13]

    Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. https://doi.org/10.2139/ssrn.4316084 The political ideology of conversational ai: Converging evidence on chatgpt's pro-environmental, left-libertarian orientation . SSRN Electronic Journal

  6. [14]

    Jessica Hoffmann, Christiane Ahlheim, Zac Yu, Aria Walfrand, Jarvis Jin, Marie Tano, Ahmad Beirami, Erin van Liemt, Nithum Thain, Hakim Sidahmed, et al. 2025. Improving neutral point of view text generation through parameter-efficient reinforcement learning and a small-scale h...

  7. [15]

    Geert Hofstede. 1984. Culture's consequences: International differences in work-related values, volume 5. sage

  8. [16]

    Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403

  9. [17]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  10. [18]

    Ian Hutchby. 2011. Non-neutrality and argument in the hybrid political interview. Discourse Studies, 13(3):349--365

  11. [19]

    Thomas S. T. Jakobsen, Laura Cabello, and Anders S gaard. 2023. https://arxiv.org/abs/2306.00639 Being right for whose right reasons? arXiv preprint arXiv:2306.00639

  12. [20]

    Graham Kalton and Howard Schuman. 1982. The effect of the question on survey responses: A review. Journal of the Royal Statistical Society Series A: Statistics in Society, 145(1):42--57

  13. [21]

    Jan Kammerath. 2024. https://medium.com/@jankammerath/deepseek-is-it-a-stolen-chatgpt-a805b586b24a Deepseek: Is it a stolen chatgpt? https://medium.com/@jankammerath/deepseek-is-it-a-stolen-chatgpt-a805b586b24a. Accessed: 2025-03-21

  14. [22]

    Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2025. Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology, 34(2):1--23

  15. [23]

    Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, and Roi Reichart. 2024. The colorful future of llms: Evaluating and improving llms as emotional supporters for queer youth. arXiv preprint arXiv:2402.11886

  16. [24]

    Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al. 2024 a . Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434

  17. [25]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024 b . Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437

  18. [26]

    Ruibo Liu, Chenyan Jia, Jason Wei, Guang Xu, and Soroush Vosoughi. 2022. https://doi.org/10.1016/j.artint.2021.103654 Quantifying and alleviating political bias in language models . Artificial Intelligence, 304:103654

  19. [27]

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2...

  20. [28]

    Fumio Motoki, Vitor Pinho Neto, and Vitor Rodrigues. 2024. https://doi.org/10.1007/s11127-023-01097-2 More human than human: measuring chatgpt political bias . Public Choice, 198(1):3--23

  21. [29]

    Malvina Nissim, Rik van Noord, and Rob van der Goot. 2019. https://arxiv.org/abs/1905.09866 Fair is better than sensational: Man is to doctor as woman is to doctor . arXiv preprint arXiv:1905.09866

  22. [30]

    Malvina Nissim, Rik van Noord, and Rob van der Goot. 2020. https://doi.org/10.1162/coli_a_00379 Fair is better than sensational: Man is to doctor as woman is to doctor . Computational Linguistics, 46(2):487--497

  23. [31]

    OpenAI . 2024. https://openai.com/index/learning-to-reason-with-llms/ Learning to reason with llms . Accessed: 2025-03-18

  24. [32]

    Peter S Park, Philipp Schoenegger, and Chongyang Zhu. 2024. https://doi.org/10.3758/s13428-023-02307-x Diminished diversity-of-thought in a standard large language model . Behavior Research Methods

  25. [33]

    Pagnarasmey Pit, Xingjun Ma, Mike Conway, Qingyu Chen, James Bailey, Henry Pit, Putrasmey Keo, Watey Diep, and Yu-Gang Jiang. 2024. Whose side are you on? investigating the political stance of large language models. arXiv preprint arXiv:2403.13840

  26. [34]

    Ella Rabinovich, Samuel Ackerman, Orna Raz, Eitan Farchi, and Ateret Anaby Tavor. 2023. https://aclanthology.org/2023.gem-1.12/ Predicting question-answering performance of large language models through semantic consistency . In Proceedings of the Third Workshop on Natural Lan...

  27. [35]

    Matthew Renze and Erhan Guven. 2024. Self-reflection in large language model agents: Effects on problem-solving performance. In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), pages 516--525. IEEE

  28. [36]

    Luca Rettenberger, Markus Reischl, and Mark Schutera. 2024. Assessing political bias in large language models. arXiv preprint arXiv:2405.13041

  29. [37]

    David Rozado. 2020. https://doi.org/10.1371/journal.pone.0231189 Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types . PLOS ONE, 15(4):e0231189

  30. [38]

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004. PMLR

  31. [39]

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314

  32. [40]

    Harry C Triandis. 2018. Individualism and collectivism. Routledge

  33. [41]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36:74952--74965

  34. [42]

    William Voegeli. 2023. Progressivism, conservatism, and democracy. J. Contemp. Legal Issues, 24:155

  35. [43]

    Joe H Ward Jr. 1963. Hierarchical grouping to optimize an objective function. Journal of the American statistical association, 58(301):236--244

  36. [44]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  37. [45]

    Xuyang Wu, Jinming Nian, Zhiqiang Tao, and Yi Fang. 2025. Evaluating social biases in llm reasoning. arXiv preprint arXiv:2502.15361

  38. [46]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.