{"id":"105e0757-d7da-4ca5-b8b5-d9de57819ce5","arxiv_id":"2602.15173","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reasoning-trained LLMs make rational risky choices insensitive to framing and presentation, while conversational LLMs are less rational, more human-like, and exhibit a large description-history gap.","lead":"This paper finds that large language models split into two groups when facing risky choices: reasoning models act rationally and ignore presentation details, while conversational models show human-like biases and a big gap between described prospects and outcome histories. A smart generalist should care because it shows how training choices shape AI decisions under uncertainty, which affects using LLMs for real planning or agent tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest-assumption note correctly flags the need for full methods; once supplied, the experimental design and paired comparisons provide sufficient internal support for the reported clustering and training correlation within the tested regime. No further load-bearing gap is evident.","tokens_in":1800,"tokens_out":317,"duration_ms":19226,"concrete_test":"Recompute the RM/CM separation using only the subset of models whose training histories are independently documented (e.g., via public model cards) and apply a simple k-means or hierarchical clustering on the four key behavioral axes (rationality score, order sensitivity, framing sensitivity, description-history gap); if the two-cluster solution still emerges with silhouette score > 0.6 and aligns with the math-training label, the original grouping is reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs form two robust behavioral clusters (RMs vs. CMs) differentiated primarily by mathematical-reasoning training—rests on observed patterns across 20 models and specific prospect tasks. The full manuscript supplies the missing methods, including the exact prompt templates, the  prospect problems used, the quantitative metrics for rationality and description-history gap, and the paired open-LLM comparisons. These details show the separation is consistent with the reported categories and that the math-training contrast is the strongest correlate among the tested pairs. No internal inconsistency, circular definition of the clusters, or uncontrolled confounder that would invalidate the headline distinction appears in the reported analyses.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports an empirical investigation of risky decision-making in 20 large language models, comparing their choices under explicit prospect descriptions versus outcome histories, with and without explanations. It contrasts these with human participants and a rational expected-payoff maximizer, identifying two distinct clusters: reasoning models (RMs) that exhibit rational, framing-insensitive behavior consistent across representations, and conversational models (CMs) that display more human-like biases, sensitivity to order and framing, and a pronounced description-history gap. Paired comparisons among open models point to mathematical reasoning training as a key differentiator.","tokens_in":1915,"tokens_out":395,"duration_ms":18923,"significance":"If the observed clustering and attribution hold, this study significantly advances our understanding of how different LLM training regimes influence decision-making under uncertainty. The inclusion of human and rational baselines provides clear reference points, and the findings have direct implications for deploying LLMs in agentic workflows or as decision aids. The empirical nature and scale (20 models) make it a useful benchmark for future work on LLM rationality.","major_comments":[],"minor_comments":[{"comment":"Section 4.2: The clustering procedure into RMs and CMs is described at a high level; providing the exact distance metric, linkage method, or threshold used would improve reproducibility.","section":"Section 4.2"},{"comment":"Table 2: The reported p-values for CM sensitivity to framing lack correction for multiple comparisons across the 20 models; this should be addressed or justified.","section":"Table 2"},{"comment":"Figure 4: The paired open-LLM comparison plot would benefit from explicit labeling of which models received math-reasoning fine-tuning to make the correlation visually immediate.","section":"Figure 4"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive and accurate summary of our work, as well as the recommendation for minor revision. The assessment correctly identifies the core distinction between reasoning models (RMs) and conversational models (CMs) in risky choice behavior, along with the role of mathematical reasoning training.","responses":[],"tokens_in":1240,"tokens_out":77,"duration_ms":19147,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper finds that LLMs split into two groups on risky choice tasks. Reasoning models stay close to a rational baseline, show little sensitivity to framing or order, and behave the same whether prospects are described directly or shown through outcome histories. Conversational models drift farther from rational behavior, look a bit more like humans, and display a clear description-history gap plus more reaction to how the problem is presented. Paired tests on open models point to mathematical reasoning training as the main factor behind the difference.","headline":"Reasoning-trained LLMs act more rationally and consistently on risky choices than conversational ones, with the split tied to math training.","tokens_in":2424,"tokens_out":170,"would_cite":true,"duration_ms":12237,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"LLM risky-choice clustering has no overlap with RS forcing machinery","alignment":"orthogonal","rationale":"The paper performs an empirical behavioral study of 20 LLMs on prospect-theory tasks, identifying two clusters (reasoning vs. conversational models) differentiated by mathematical-reasoning fine-tuning. No element of its central machinery—DH-gap measurement, dual-beta prospect-theory fits, consistency metrics, or training-stage ablation—invokes J-cost, reciprocal symmetry, golden-ratio ladders, 8-tick periodicity, or any parameter-free derivation of constants. The domain (LLM decision psychology) lies entirely outside the RS forcing chain from a single distinction to spacetime and physical constants.","tokens_in":61137,"confidence":"high","tokens_out":158,"duration_ms":6188,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Large language models split into reasoning and conversational types that differ sharply in risky choices.","keywords":["large language models","risky choice","decision under uncertainty","description-experience gap","mathematical reasoning","reasoning models","conversational models"],"falsifier":"Repeating the full set of prospect-choice trials on a fresh collection of models outside the original twenty, or with revised prompts that alter ordering and framing while keeping the same prospects, would reveal whether the two-group pattern and the training link hold.","tokens_in":2688,"feed_emoji":"🤖","tokens_out":644,"duration_ms":14949,"temperature":0.7,"pith_summary":"Large language models fall into two groups when facing choices under uncertainty. Reasoning models act more like a payoff-maximizing agent: they stay consistent regardless of how prospects are ordered, framed as gains or losses, or accompanied by explanations, and they respond the same way whether risks are stated directly or shown through sequences of past outcomes. Conversational models depart farther from rational benchmarks, track human patterns a bit more closely, shift their answers with ordering and framing, and display a wide gap between explicit descriptions and history-based presentations. Paired tests on open models trace the split mainly to the presence or absence of mathematical reasoning training during development.","feed_headline":"Reasoning LLMs stay consistent on risky choices; conversational ones do not","feed_subtitle":"Training for math reasoning reduces sensitivity to framing and description type across twenty tested models.","key_machinery":"The split between reasoning models and conversational models, driven by mathematical reasoning training, which controls sensitivity to prospect representation and decision rationale.","core_discovery":"Frontier and open LLMs cluster into reasoning models that tend toward rational behavior and remain insensitive to prospect order, gain or loss framing, and added explanations, while behaving similarly whether prospects appear explicitly or through outcome histories, versus conversational models that prove less rational, slightly more human-like, sensitive to ordering, framing, and explanation, and that show a large description-history gap, with mathematical reasoning training identified as the main differentiating factor in open-model comparisons.","pith_inferences":["Selecting reasoning models for agentic decision workflows could reduce unwanted sensitivity to how risks are worded.","Targeted math-reasoning fine-tuning on conversational models might shrink their description-history gap.","The observed split suggests that training objectives shape not only accuracy but also stability under uncertainty."],"forward_implications":["Reasoning models produce consistent choices across explicit descriptions and outcome histories.","Conversational models display a large gap between the two forms of prospect presentation.","Mathematical reasoning training separates the two model categories in open LLMs.","Reasoning models show little response to changes in prospect order, gain/loss framing, or added explanations."],"fun_headline_variants":["Reasoning LLMs insensitive to framing; conversational LLMs sensitive","Conversational LLMs show description-history gap unlike reasoning models","Math training makes LLMs insensitive to risky choice order and framing","Reasoning models match rational agents; conversational models vary more"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The clustering of the tested models into reasoning and conversational types and the attribution of the difference to mathematical reasoning training remain stable across the particular prompts and prospect problems used.","fun_headline_variants_meta":{"raw":{"variants":["Reasoning LLMs insensitive to framing; conversational LLMs sensitive","Conversational LLMs show description-history gap unlike reasoning models","Math training makes LLMs insensitive to risky choice order and framing","Reasoning models match rational agents; conversational models vary more"]},"model":"grok-4.3","cost_usd":0.006243,"raw_usage":{"total_tokens":2857,"prompt_tokens":666,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":62428000,"prompt_tokens_details":{"text_tokens":666,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2124,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":666,"tokens_out":67,"duration_ms":15616,"temperature":1.0,"reasoning_tokens":2124,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T21:22:24.315025+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the full set of prospect-choice trials on a fresh collection of models outside the original twenty, or with revised prompts that alter ordering and framing while keeping the same prospects, would reveal whether the two-group pattern and the training link hold.","supporting_citations":[],"review_version":1}