Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Against 360 human participants on three classic decision tasks, five leading LLMs usually matched or exceeded humans and approached near-optimal play, while fitted cognitive models show the learning processes behind those choices are…

desk verdict The performance findings are credible and worth engaging; the process-level claim that LLMs learn in a fundamentally non-human way is confounded by the LLMs receiving a complete transcript of past rounds while humans rely on memory. read the letter →

arxiv 2506.16163 v1 pith:OI2R4JZT submitted 2025-06-19 cs.AI

classification cs.AI
keywords largelanguagemodelsdecision-makingIowaGamblingTaskCambridgeWisconsinCardSortingcomputationalcognitivemodelingriskanduncertaintyhuman-AIcomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models make decisions the way humans do. It answers: mostly better on scores, but not the same way underneath. Five leading LLMs were run through three established psychology tasks—the Iowa Gambling Task, the Cambridge Gambling Task, and the Wisconsin Card Sorting Task—alongside 360 new human participants, with both groups receiving the same instructions. Most models matched or beat the humans and approached optimal-strategy benchmarks, yet computational models fitted to their choices show faster learning rates, more deterministic choices, stronger outcome sensitivity, and little of the flexible risk adjustment humans display. The authors read this as evidence that LLMs are capable, rational, outcome-driven decision agents rather than cognitive replicas of humans, which matters wherever such systems are delegated real decisions.

What carries the argument

The central machinery is a matched human–LLM comparison on three classic experimental psychology tasks, with two optimal-play heuristics (upper confidence bound and $\epsilon$-greedy for the Iowa Gambling Task; expected-utility maximization for the other two) providing near-optimal benchmarks. The process-level claims are carried by three computational cognitive models: the prospect valence learning model with a decay reinforcement-learning rule (learning rate $A$, choice consistency $c$, outcome sensitivity $\alpha$, loss aversion $\lambda$) for the Iowa Gambling Task; the cumulative model (type bias $c$, probability distortion $\alpha$, risk aversion $\rho$, choice consistency $\gamma$) for the Cambridge Gambling Task; and the sequential learning model (reward sensitivity $r$, punishment sensitivity $p$, choice consistency $d$) for the Wisconsin Card Sorting Task. Posterior estimates of these parameters are compared between humans and each LLM, and the divergence in those estimates is what supports the claim that similar or better scores are reached through a different decision process.

What would settle it

Run human participants on the same three tasks while showing them a continuously updated, written transcript of all their previous choices and outcomes—the exact information each LLM receives—and refit the same cognitive models; if human posterior estimates then approach the LLM estimates, the central claim that the divergence is intrinsic to LLM decision processes would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that five leading LLMs (GPT-4o, o4-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, and DeepSeek-R1) are near-optimal decision-makers on standard psychological tasks while the learning behavior behind their choices is non-human. Across three tasks and 360 human participants, most LLMs surpassed or matched humans and approached the performance of upper-confidence-bound, $\epsilon$-greedy, or expected-utility-maximizing baselines, with isolated exceptions such as Gemini underperforming humans under risk and DeepSeek matching rather than beating humans under set-shifting. Fitted cognitive models nonetheless show a systematic divergence: LLMs learn faster, choose more deterministically, are more sensitive to outcomes and penalties, distort probabilities more strongly, and barely adjust their bets as risk changes, whereas humans learn more slowly and noisily and scale their risk-taking flexibly. The authors conclude that LLMs behave as consistent, rational, outcome-driven agents, not as substitutes for human cognition.

Load-bearing premise

The argument that LLMs learn fundamentally differently from humans assumes that the LLMs' access to a complete written transcript of every past choice and outcome, while humans rely on their own memory, is not the real cause of the fitted differences in learning rate, choice consistency, and outcome sensitivity.

Editorial extensions

If this is right

  • Four of the five LLMs (GPT-4o, Claude, o4-mini, and DeepSeek) matched or exceeded human performance across all three tasks, with several approaching the near-optimal benchmarks on each task.
  • The fitted parameters imply that LLMs learn faster, choose more deterministically, and react more strongly to outcomes—especially penalties—than humans do, so their high scores rest on a different learning profile.
  • LLMs showed little or no risk adjustment in betting as odds changed, a flat pattern that in human subjects is a trans-diagnostic marker of cognitive dysfunction; using LLMs as stand-ins in probabilistic-reasoning research can therefore mislead.
  • Robustness checks with GPT-4o indicate the decision patterns persisted across temperature changes, payoff rescaling, context reframing, and demographic role-play, suggesting the non-human style is not an artifact of one prompt formulation.
  • Post-experiment surveys showed that participants were reluctant to accept AI assistance even after being told the LLMs had scored higher, pointing to algorithm aversion as a practical barrier to delegation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the authors leave implicit is to give human participants the same continuously updated written transcript of past choices and outcomes that the LLMs receive; if human parameter estimates then move toward the LLM estimates, the divergence could be attributed to memory and information access rather than to intrinsic non-human learning.
  • The near-total invariance of GPT-4o's strategy to demographic role-play implies that LLM outputs cannot stand in for the between-subject heterogeneity that human samples show; this strengthens the paper's caution against LLM substitution in behavioral research into a measurable claim about flattened variance.
  • Because flat risk adjustment is a clinical marker, the three-task battery could be repurposed the other way: LLM behavior as an anchor for pure expected-utility maximization, against which clinical deviations in human patients could be quantified.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript compares the decision-making of five large language models (GPT-4o, o4-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, and DeepSeek-R1) with 360 human participants across three psychological tasks: the Iowa Gambling Task (uncertainty), the Cambridge Gambling Task (risk), and the Wisconsin Card Sorting Task (set-shifting). The authors report that LLMs generally match or exceed human performance and approach near-optimal strategies, while computational cognitive models fitted to choices (PVL-Decay RL, Cumulative Model, and Sequential Learning Model) reveal posterior parameter estimates that differ significantly from humans, which they interpret as evidence of fundamentally non-human learning processes. Robustness checks vary temperature, prompt framing, role-play, and score scaling, and a post-experiment survey probes attitudes toward AI assistance.

Significance. If the conclusions hold, this is a valuable contribution: it is a pre-registered, multi-task benchmark of frontier LLMs against a newly recruited human sample, it uses established computational models to characterize behavior mechanistically, and it ships data and code. The performance findings are interesting even on their own. However, the paper's central second claim—that LLMs show fundamentally different learning processes—rests on a comparison that is confounded by a substantial information asymmetry between the LLM and human conditions. Because the robustness checks vary many prompt features but never the availability of the full outcome transcript, the process-level conclusion is not yet established. With additional matched-information conditions or a re-framed claim, the paper could be made defensible.

major comments (4)
  1. [Methods (Large Language Models); Supplementary Note 1] The decision-making prompt gives LLMs a complete transcript of every previous choice, reward, penalty, and (for WCST) the model's own reasoning text, while human participants only see the outcome of the current round and must maintain the rest in memory (Methods; Fig. S12–S15). This asymmetry directly confounds the process-level comparison. For example, in the PVL-Decay model, access to a verbatim history can push the learning-rate parameter A toward its upper bound (LLM A ≈ 0.86–0.99 vs. humans ≈ 0.71; Table S1d) and inflate outcome sensitivity α, without implying any difference in the learning process. The authors' statement that LLMs were 'just like human participants' because they were not instructed on how to use the provided information conflates information access with instruction. The robustness checks in Supplementary Note 3 vary temperature, framing, role-play, and score transformations but never restrict or remove the historical transcript, so they do not address this confound. This undermines the central claim that LLM decisions 'diverged fundamentally' from human decision processes.
  2. [Results (Decision-making under uncertainty); Fig. 1A] The near-optimality claim is also affected by the same asymmetry: a model that receives the full outcome history is performing offline inference over a complete record, whereas humans learn online from one outcome per trial. Thus the statement that 'LLMs managed to identify advantageous decks earlier and adapted their choices more effectively' is an observation about the prompt design rather than about the LLMs' intrinsic decision-making ability. The authors should either run a human condition with an equivalent written history (e.g., a paper-based running summary of past picks and outcomes) or an LLM condition with a recency-limited window, and then re-examine whether performance and parameter differences persist. At minimum, all performance and process claims should be explicitly re-framed as conditional on access to the complete history.
  3. [Supplementary Note 1 (Wisconsin Card Sorting Task prompt)] In the WCST decision prompt, the LLM is re-presented with its own previous reasoning text (<reason_{i}>) together with the feedback. This is a scaffolding intervention not available to human participants: the model can consult a written record of its own hypotheses while updating its attention weights, which can inflate the fitted reward sensitivity r, punishment sensitivity p, and choice consistency d in the Sequential Learning Model (Supplementary Note 2). The robustness checks do not include a condition that removes the reasoning replay, so this potential confound remains untested. The authors should add such a control or explicitly justify why replaying the model's own reasoning is equivalent to a human's memory of it.
  4. [Tables S1d, S2c, S3d] The reported LLM posterior means often have near-zero standard deviations (e.g., GPT α = 0.9522, SD = 0.0089 in Table S1d; LLM r SD = 0.0000 in Table S3d), indicating that repeated sessions of the same model with the same prompt yield nearly identical parameter estimates. As a result, Mann-Whitney U tests against human posteriors are significant almost by construction, and the enormous Cohen's d values (e.g., r: d = –18.875 in Table S3c) are not informative about the magnitude of cognitive differences. The authors should report the number of LLM sessions and, if the low variance is genuine, use a statistical comparison that accounts for it (e.g., testing whether the human posterior distribution lies within a tolerance of the LLM distribution) before claiming fundamental differences.
minor comments (6)
  1. [Methods (Large Language Models)] The main text does not report the number of decision-making sessions or repetitions per LLM; Tables S1–S3 give W statistics but not n for the LLM groups. Please state the sample size per model explicitly in the Methods.
  2. [Figure 2 caption; Results (CGT)] The model name is spelled inconsistently: 'GPTo4m' is used in most places, but 'GPT4o4m' appears in the Figure 2 caption and in the Results paragraph describing parameter estimates. Please use a consistent name for the o4-mini model throughout.
  3. [References] Reference [35] is malformed: 'Strategy, A.: Deciding advantageously before knowing the' is incomplete and should cite Bechara, A., Damasio, H., Tranel, D., and Damasio, A.R. (1997), Science 275, 1293–1295.
  4. [Methods (Robustness Checks)] The robustness checks are conducted only on GPT-4o, not on the other four LLMs; this limits the generality of the claim that decision-making patterns are robust across prompt variations. Please state this limitation in the main text or add at least one additional model to the robustness analyses.
  5. [Data Availability] The pre-registration IDs (AsPredicted #182473, #186129, and #203115) are mentioned but no URLs or analysis plans are provided. Please include links to the pre-registrations in the Data Availability section so that readers can verify the planned analyses.
  6. [Supplementary Note 3] The text states that CGT decision quality held 'except for economics and medicine,' but Fig. S6 is not annotated to indicate which panels are the exceptions. Please clarify which variants deviate and why.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: performance benchmarks are external (UCB, epsilon-greedy, EU-max) and the cognitive-model parameters are fitted comparative measurements, not predictions derived from fitted constants.

full rationale

The paper's two central claims are that LLMs approach optimal performance and that their fitted decision-process parameters differ from humans'. Neither claim reduces to its own inputs. The performance comparisons use external benchmarks specified independently of LLM data: Upper Confidence Bound and epsilon-greedy for the Iowa Gambling Task (Methods), and expected-utility-maximization strategies for the Cambridge Gambling Task and Wisconsin Card Sorting Task (Methods; Figs. 1A, 2A, 3A). The process-level comparison fits standard published cognitive models (PVL-decay reinforcement learning [53], cumulative model [54], sequential learning model [56]) to choice sequences from both LLMs and humans, then compares posterior parameter estimates; this is a comparative measurement and interpretation, not a prediction validated against the same fitted constants. The only apparent self-citation (reference [22], an LLM experiment on strategic games) is cited as one example among several and is not load-bearing. The main methodological vulnerability is that LLMs receive a full transcript of past choices and outcomes in the decision-making prompt while humans rely on memory; this is a potential confound in interpreting parameter differences, but it does not constitute circularity because no equation or fitted value is defined in terms of the paper's conclusions. No circular step can be exhibited with the quoted reduction required by the analysis rules.

Assumptions & free parameters 11 free parameters · 4 assumptions · 0 invented entities

The central comparison rests on the assumptions that the three psychological tasks measure the intended decision dimensions when administered to LLMs, that the standard cognitive models can be meaningfully fitted to LLM choices, and that the reworded materials do not change the constructs. These are domain assumptions inherited from the machine-psychology literature rather than established facts for LLMs. The cognitive-model parameters are free parameters fitted to the data; they are not derived from first principles, but they are also not used to generate out-of-sample predictions in this paper.

free parameters (11)
  • IGT learning rate A (PVL-decay RL model) = LLM medians 0.86-0.99; humans 0.70
    Estimated from choice data in Table S1(d); higher A in LLMs supports the claim of faster expectation updating.
  • IGT choice consistency c = LLM medians 0.41-0.90; humans 0.35
    Table S1(d); used to argue LLMs choose the learned best option more deterministically.
  • IGT outcome sensitivity alpha = LLM medians 0.82-0.95; humans 0.31
    Table S1(d); used to argue LLMs are more sensitive to outcome magnitudes.
  • IGT loss aversion lambda = LLM medians 0.72-1.87; humans 0.59
    Table S1(d); used to argue most LLMs react more strongly to penalties than humans.
  • CGT type bias c for red boxes = LLM medians 0.46-0.58; humans 0.50
    Table S2(c); used to show response biases differ, with Gemini and DeepSeek strongly biased in opposite directions.
  • CGT probability distortion alpha = LLM medians 4.33-4.99; humans 3.11
    Table S2(c); higher distortion is interpreted as LLMs treating high probabilities as even more certain.
  • CGT risk aversion rho = LLM medians 0.002-692.5; humans 6.43
    Table S2(c); extreme spread across models indicates either genuine preference differences or difficulties in identifying this parameter for LLM data.
  • CGT choice consistency gamma = LLM medians 5.49-65.02; humans 12.53
    Table S2(c); high values for GPTo4m and DeepSeek underlie the claim of deterministic, expected-value-aligned betting.
  • WCST reward sensitivity r = LLM medians 0.9995-0.9997; humans 0.9975
    Table S3(d); differences, though tiny in absolute terms, are statistically significant and used to argue faster adaptation to positive feedback.
  • WCST punishment sensitivity p = LLM medians 0.977-0.9997; humans 0.955
    Table S3(d); used to argue LLMs adapt faster to negative feedback, with DeepSeek the exception.
  • WCST choice consistency d = LLM medians 0.29-0.55; humans 0.36
    Table S3(d); used to support the claim of generally more deterministic rule following by LLMs.
assumptions (4)
  • domain assumption Validity of IGT, CGT, and WCST as measures of uncertainty, risk, and set-shifting.
    The paper adapts these tasks and assumes they isolate the stated constructs; this is standard in psychology but not independently validated for LLMs.
  • domain assumption Applicability of PVL, Cumulative, and Sequential Learning models to LLM choices.
    The paper fits human cognitive models to LLM response patterns and interprets the fitted parameters as reflecting underlying processes; there is no evidence these models describe LLM computation beyond providing a curve fit.
  • domain assumption Rewording and payoff redesign preserve the essence of the original tasks.
    To reduce memorization effects, the authors reworded descriptions and altered payoff structures, assuming the psychological challenge for the LLMs remains equivalent to the standard tasks.
  • domain assumption The 360-participant sample and the five chosen LLMs are representative enough for the claims about humans and LLMs.
    Human participants are university students in Xi'an, mean age 18.95; LLMs are a convenience set of five commercial APIs. The paper generalizes beyond these samples without explicit justification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior." pith.science (2026). https://pith.science/paper/OI2R4JZT

@misc{pith2026250616163,
  author       = {Pith},
  title        = {Pith review of: Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OI2R4JZT}},
  note         = {Machine review of arXiv:2506.16163}
}
read the original abstract

Human decision-making belongs to the foundation of our society and civilization, but we are on the verge of a future where much of it will be delegated to artificial intelligence. The arrival of Large Language Models (LLMs) has transformed the nature and scope of AI-supported decision-making; however, the process by which they learn to make decisions, compared to humans, remains poorly understood. In this study, we examined the decision-making behavior of five leading LLMs across three core dimensions of real-world decision-making: uncertainty, risk, and set-shifting. Using three well-established experimental psychology tasks designed to probe these dimensions, we benchmarked LLMs against 360 newly recruited human participants. Across all tasks, LLMs often outperformed humans, approaching near-optimal performance. Moreover, the processes underlying their decisions diverged fundamentally from those of humans. On the one hand, our finding demonstrates the ability of LLMs to manage uncertainty, calibrate risk, and adapt to changes. On the other hand, this disparity highlights the risks of relying on them as substitutes for human judgment, calling for further inquiry.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Artificially intelligent agents in the social and behavioral sciences: A history and outlook

    cs.AI 2025-10 conditional novelty 2.0 of 10

    AI and social science have co-evolved for 75 years through rapid technological adoption and slower scientific consolidation, with direct human-focused AI studies still scarce.

Reference graph

Works this paper leans on

69 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    The Free Press, Glencoe (1947)

    Simon, H.A.: Administrative Behavior: A Study of Decision-Making Processes in Administrative Organization. The Free Press, Glencoe (1947)

  2. [2]

    Turban, E., Watkins, P.R.: Integrating expert systems and decision support systems. MIS Q. 10(2), 121–136 (1986)

  3. [3]

    Preprint arXiv:2503.23674 (2025)

    Jones, C.R., Bergen, B.K.: Large Language Models Pass the Turing Test. Preprint arXiv:2503.23674 (2025)

  4. [4]

    Proceedings of the National Academy of Sciences 121(9), 2313925121 (2024)

    Mei, Q., Xie, Y., Yuan, W., Jackson, M.O.: A turing test of whether ai chatbots are behaviorally similar to humans. Proceedings of the National Academy of Sciences 121(9), 2313925121 (2024)

  5. [5]

    arXiv preprint arXiv:2108.07258 (2021)

    Bommasani, R., Hudson, D.A., Adeli, E., Altman, R., Arora, S., Arx, S., Bernstein, M.S., Bohg, J., Bosselut, A., Brunskill, E., et al.: On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)

  6. [6]

    Handler, A., Larsen, K.R., Hackathorn, R.: Large language models present new questions for decision support. Int. J. Info. Manag. 79, 102811 (2024)

  7. [7]

    Preprint arXiv:2405.01769 (2024)

    Chen, Z., Ma, J., Zhang, X., Hao, N., Yan, A., Nourbakhsh, A., Yang, X., McAuley, J., Petzold, L.R., Wang, W.Y.: A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law. Preprint arXiv:2405.01769 (2024)

  8. [8]

    Technology section (2023)

    Taylor, L.: Colombian judge says he used ChatGPT in ruling. Technology section (2023). https://www. theguardian.com/technology/2023/feb/03/colombia-judge-chatgpt-ruling Accessed 2025-05-08

Show all 69 references
  1. [9]

    Bloomberg Green Daily newsletter article (2024)

    Alba, D., Yin, L.: Uncovering What OpenAI’s GPT Sees When Used For Ranking Resumes. Bloomberg Green Daily newsletter article (2024). https://www.bloomberg.com/news/newsletters/2024-03-08/ companies-should-think-twice-before-using-generative-ai-in-hiring Accessed 2025-05-08

  2. [10]

    Accessed 2025-05-11 (2025)

    Singla, A., Sukharevsky, A., Yee, L., Chui, M., Hall, B.: The state of AI: How organizations are rewiring to capture value. Accessed 2025-05-11 (2025). https://www.mckinsey.com/capabilities/quantumblack/ our-insights/the-state-of-ai

  3. [11]

    Science 381(6654), 187–192 (2023)

    Noy, S., Zhang, W.: Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654), 187–192 (2023)

  4. [12]

    Nature 623(7987), 474–477 (2023)

    Extance, A.: Chatgpt has entered the classroom: how llms could transform education. Nature 623(7987), 474–477 (2023)

  5. [13]

    Nature Communications 15(1), 8236 (2024)

    Williams, C.Y., Miao, B.Y., Kornblith, A.E., Butte, A.J.: Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nature Communications 15(1), 8236 (2024)

  6. [14]

    Goh, E., Gallo, R.J., Strong, E., Weng, Y., Kerman, H., Freed, J.A., Cool, J.A., Kanjee, Z., Lane, K.P., Parsons, A.S., et al.: Gpt-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med. 31, 1233–1238 (2025)

  7. [15]

    Nature 620(7972), 172–180 (2023)

    Singhal, K., Azizi, S., Tu, T., Mahdavi, S.S., Wei, J., Chung, H.W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al.: Large language models encode clinical knowledge. Nature 620(7972), 172–180 (2023)

  8. [16]

    Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., Vielhauer, J., Makowski, M., Braren, R., Kaissis, G., et al.: Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30(9), 2613–2622 (2024)

  9. [17]

    In: Moens, M.-F., Huang, X., Specia, L., Yih, S.W.-t

    Chen, Z., Chen, W., Smiley, C., Shah, S., Borova, I., Langdon, D., Moussa, R., Beane, M., Huang, T.-H., Routledge, B., Wang, W.Y.: FinQA: A dataset of numerical reasoning over financial data. In: Moens, M.-F., Huang, X., Specia, L., Yih, S.W.-t. (eds.) Proceedings of the 2021 ...

  10. [18]

    Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., Xiao, M., Li, D., Dai, Y., Feng, D., et al.: Finben: A holistic financial benchmark for large language models. Adv. Neural Inf. Process. Syst. 37, 95716–95743 (2024)

  11. [19]

    Transact

    Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., Anandkumar, A.: Voyager: An open-ended embodied agent with large language models. Transact. Mach. Learn. Res. 3 (2024)

  12. [20]

    arXiv preprint arXiv:2403.03186 (2024)

    Tan, W., Zhang, W., Xu, X., Xia, H., Ding, Z., Li, B., Zhou, B., Yue, J., Jiang, J., Li, Y., et al.: Cradle: Empowering foundation agents towards general computer control. arXiv preprint arXiv:2403.03186 (2024)

  13. [21]

    online ahead of print in Nat

    Akata, E., Schulz, L., Coda-Forno, J., Oh, S.J., Bethge, M., Schulz, E.: Playing repeated games with large language models. online ahead of print in Nat. Hum. Behav. (2025)

  14. [22]

    arXiv preprint arXiv:2410.03724 (2024)

    Wang, Z., Song, R., Shen, C., Yin, S., Song, Z., Battu, B., Shi, L., Jia, D., Rahwan, T., Hu, S.: Large language models overcome the machine penalty when acting fairly but not when acting selfishly or altruistically. arXiv preprint arXiv:2410.03724 (2024)

  15. [23]

    Wiley Interdiscip

    Johnson, J.G., Busemeyer, J.R.: Decision making under risk and uncertainty. Wiley Interdiscip. Rev. Cogn. Sci. 1(5), 736–749 (2010)

  16. [24]

    Sage, ??? (2010)

    Hastie, R., Dawes, R.M.: Rational Choice in an Uncertain World: The Psychology of Judgment and Decision Making. Sage, ??? (2010)

  17. [25]

    Einhorn, H.J., Hogarth, R.M.: Behavioral decision theory: Processes of judgement and choice. Annu. Rev. Psychol. 32(1981), 53–88 (1981)

  18. [26]

    Knight, F.H.: Risk, Uncertainty and Profit vol. 31. Houghton Mifflin, ??? (1921)

  19. [27]

    Camerer, C., Weber, M.: Recent developments in modeling preferences: Uncertainty and ambiguity. J. Risk Uncertain. 5, 325–370 (1992)

  20. [28]

    Science 236(4801), 537–543 (1987)

    Machina, M.J.: Decision-making in the presence of risk. Science 236(4801), 537–543 (1987)

  21. [29]

    In: Handbook of the Fundamentals of Financial Decision Making: Part I, pp

    Kahneman, D., Tversky, A.: Prospect theory: An analysis of decision under risk. In: Handbook of the Fundamentals of Financial Decision Making: Part I, pp. 99–127. World Scientific, ??? (2013)

  22. [30]

    Acta Psychol

    Brehmer, B.: Dynamic decision making: Human control of complex systems. Acta Psychol. 81(3), 211–241 (1992)

  23. [31]

    Cambridge University Press, ??? (1993)

    Payne, J.W., Bettman, J.R., Johnson, E.J.: The Adaptive Decision Maker. Cambridge University Press, ??? (1993)

  24. [32]

    (ed.): Decision Making Under Uncertainty: Theory and Application

    Kochenderfer, M.J. (ed.): Decision Making Under Uncertainty: Theory and Application. MIT Press, Cambridge MA (2015)

  25. [33]

    Ruggeri, K., Al ´ ı, S., Berge, M.L., Bertoldo, G., Bjørndal, L.D., Cortijos-Bernabeu, A., Davison, C., Demi´ c, E., Esteban-Serna, C., Friedemann, M., et al.: Replicating patterns of prospect theory for decision under risk. Nat. Hum. Behav. 4(6), 622–633 (2020)

  26. [34]

    Uddin, L.Q.: Cognitive and behavioural flexibility: neural mechanisms and clinical considerations. Nat. Rev. Neurosci. 22(3), 167–179 (2021)

  27. [35]

    Science 275, 1293–1293 (1997)

    Strategy, A.: Deciding advantageously before knowing the. Science 275, 1293–1293 (1997)

  28. [36]

    Sacr´ e, P., Kerr, M.S., Subramanian, S., Fitzgerald, Z., Kahn, K., Johnson, M.A., Niebur, E., Eden, U.T., Gonz´ alez-Mart ´ ınez, J.A., Gale, J.T.,et al.: Risk-taking bias in human decision-making is encoded via a right–left brain push–pull system. Proc. Natl. Acad. Sci. USA ...

  29. [37]

    Konishi, S., Nakajima, K., Uchida, I., Kameyama, M., Nakahara, K., Sekihara, K., Miyashita, Y.: Transient activation of inferior prefrontal cortex during cognitive set shifting. Nat. Neurosci.1(1), 80–84 (1998)

  30. [38]

    Proceedings of the National 12 Academy of Sciences 120(6), 2218523120 (2023)

    Binz, M., Schulz, E.: Using cognitive psychology to understand gpt-3. Proceedings of the National 12 Academy of Sciences 120(6), 2218523120 (2023)

  31. [39]

    Preprint arXiv:2303.13988 (2023)

    Hagendorff, T.: Machine psychology: Investigating emergent capabilities and behavior in large language models using psychological methods. Preprint arXiv:2303.13988 (2023)

  32. [40]

    Proceedings of the National Academy of Sciences 121(45), 2405460121 (2024)

    Kosinski, M.: Evaluating large language models in theory of mind tasks. Proceedings of the National Academy of Sciences 121(45), 2405460121 (2024)

  33. [41]

    Proceedings of the National Academy of Sciences 121(24), 2317967121 (2024)

    Hagendorff, T.: Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences 121(24), 2317967121 (2024)

  34. [42]

    Chen, Y., Liu, T.X., Shan, Y., Zhong, S.: The emergence of economic rationality of gpt. Proc. Natl. Acad. Sci. USA 120(51), 2316205120 (2023)

  35. [43]

    Proceedings of the National Academy of Sciences 122(20), 2501823122 (2025)

    Lehr, S.A., Saichandran, K.S., Harmon-Jones, E., Vitali, N., Banaji, M.R.: Kernels of selfhood: Gpt-4o shows humanlike patterns of cognitive dissonance moderated by free choice. Proceedings of the National Academy of Sciences 122(20), 2501823122 (2025)

  36. [44]

    Cognition 50(1-3), 7–15 (1994)

    Bechara, A., Damasio, A.R., Damasio, H., Anderson, S.W.: Insensitivity to future consequences following damage to human prefrontal cortex. Cognition 50(1-3), 7–15 (1994)

  37. [45]

    Neuropsychopharmacology20(4), 322–339 (1999)

    Rogers, R.D., Everitt, B., Baldacchino, A., Blackshaw, A.J., Swainson, R., Wynne, K., Baker, N., Hunter, J., Carthy, T., Booker, E.,et al.: Dissociable deficits in the decision-making cognition of chronic amphetamine abusers, opiate abusers, patients with focal damage to prefr...

  38. [46]

    Berg, E.A.: A simple objective technique for measuring flexibility in thinking. J. Gen. Psychol. 39(1), 15–22 (1948)

  39. [47]

    PNAS Nexus 3(7), 245 (2024)

    Abdurahman, S., Atari, M., Karimi-Malekabadi, F., Xue, M.J., Trager, J., Park, P.S., Golazizian, P., Omrani, A., Dehghani, M.: Perils and opportunities in using large language models in psychological research. PNAS Nexus 3(7), 245 (2024)

  40. [48]

    Wang, A., Morgenstern, J., Dickerson, J.P.: Large language models that replace human participants can harmfully misportray and flatten identity groups. Nat. Mach. Intell. 7, 400–411 (2025)

  41. [49]

    Science 380(6650), 1108–1109 (2023)

    Grossmann, I., Feinberg, M., Parker, D.C., Christakis, N.A., Tetlock, P.E., Cunningham, W.A.: Ai and the transformation of social science research. Science 380(6650), 1108–1109 (2023)

  42. [50]

    Bull, P.N., Tippett, L.J., Addis, D.R.: Decision making in healthy participants on the iowa gambling task: new insights from an operant approach. Front. Psychol. 6, 391 (2015)

  43. [51]

    Auer, P., Cesa-Bianchi, N., Fischer, P.: Finite-time analysis of the multiarmed bandit problem. Mach. Learn. 47, 235–256 (2002)

  44. [52]

    PhD thesis (1989)

    Watkins, C.J.C.H., et al.: Learning from delayed rewards. PhD thesis (1989)

  45. [53]

    Ahn, W.-Y., Busemeyer, J.R., Wagenmakers, E.-J., Stout, J.C.: Comparison of decision learning models using the generalization criterion method. Cogn. Sci. 32(8), 1376–1402 (2008)

  46. [54]

    Drug Alcohol Depend

    Romeu, R.J., Haines, N., Ahn, W.-Y., Busemeyer, J.R., Vassileva, J.: A computational model of the cambridge gambling task with applications to substance use disorders. Drug Alcohol Depend. 206, 107711 (2020)

  47. [55]

    Gl¨ ascher, J., Adolphs, R., Tranel, D.: Model-based lesion mapping of cognitive control using the wisconsin card sorting test. Nat. Commun. 10(1), 20 (2019)

  48. [56]

    Bishara, A.J., Kruschke, J.K., Stout, J.C., Bechara, A., McCabe, D.P., Busemeyer, J.R.: Sequential learning models for the wisconsin card sort task: Assessing processes in substance dependent individuals. J. Math. Psychol. 54(1), 5–13 (2010)

  49. [57]

    Brain 131(5), 1311–1322 (2008)

    Clark, L., Bechara, A., Damasio, H., Aitken, M., Sahakian, B., Robbins, T.: Differential effects of insular 13 and ventromedial prefrontal cortex lesions on risky decision-making. Brain 131(5), 1311–1322 (2008)

  50. [58]

    Effah, R., Ioannidis, K., Grant, J.E., Chamberlain, S.: Exploring decision-making performance in young adults with mental health disorders: a comparative study using the cambridge gambling task. Psychol. Med. 54(9), 1890–1896 (2024)

  51. [59]

    Byrnes, J.P., Miller, D.C., Schafer, W.D.: Gender differences in risk taking: A meta-analysis. Psychol. Bull. 125(3), 367 (1999)

  52. [60]

    Weber, E.U., Hsee, C.: Cross-cultural differences in risk perception, but cross-cultural similarities in attitudes towards perceived risk. Manag. Sci. 44(9), 1205–1217 (1998)

  53. [61]

    Science 211(4481), 453–458 (1981)

    Tversky, A., Kahneman, D.: The framing of decisions and the psychology of choice. Science 211(4481), 453–458 (1981)

  54. [62]

    Tymula, A., Rosenberg Belmaker, L.A., Ruderman, L., Glimcher, P.W., Levy, I.: Like cognitive function, decision making across the life span shows profound age-related changes. Proc. Natl. Acad. Sci. USA 110(42), 17143–17148 (2013)

  55. [63]

    Dietvorst, B.J., Simmons, J.P., Massey, C.: Algorithm aversion: people erroneously avoid algorithms after seeing them err. J. Exp. Psychol. Gen. 144(1), 114 (2015)

  56. [64]

    Castelo, N., Bos, M.W., Lehmann, D.R.: Task-dependent algorithm aversion. J. Market. Res. 56(5), 809–825 (2019)

  57. [65]

    Karata¸ s, M., Cutright, K.M.: Thinking about god increases acceptance of artificial intelligence in decision-making. Proc. Natl. Acad. Sci. USA 120(33), 2218961120 (2023)

  58. [66]

    In: International Conference on Machine Learning, pp

    Zhao, Z., Wallace, E., Feng, S., Klein, D., Singh, S.: Calibrate before use: Improving few-shot perfor- mance of language models. In: International Conference on Machine Learning, pp. 12697–12706 (2021). PMLR

  59. [67]

    Journal of Behavioral and Experimental Finance 9, 88–97 (2016)

    Chen, D.L., Schonger, M., Wickens, C.: otree—an open-source platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance 9, 88–97 (2016)

  60. [68]

    Carpenter, B., Gelman, A., Hoffman, M.D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M.A., Guo, J., Li, P., Riddell, A.: Stan: A probabilistic programming language. J. Stat. Softw.76, 1–32 (2017)

  61. [69]

    Card A has 2 green flowers

    Ahn, W.-Y., Haines, N., Zhang, L.: Revealing neurocomputational mechanisms of reinforcement learning and decision-making with the hBayesDM package. Comp. Psychiatr. 1, 24–57 (2017) 14 Supplementary note 1: Prompt for LLMs The prompts provided to the LLMs consist of two parts: ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.