Pith. sign in

REVIEW 4 cited by

Incoherent Probability Judgments in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16646 v2 pith:TK4GLUKA submitted 2024-01-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords judgmentsprobabilityllmsmodelsautoregressivebayesiancoherentdeviations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Autoregressive Large Language Models (LLMs) trained for next-word prediction have demonstrated remarkable proficiency at producing coherent text. But are they equally adept at forming coherent probability judgments? We use probabilistic identities and repeated judgments to assess the coherence of probability judgments made by LLMs. Our results show that the judgments produced by these models are often incoherent, displaying human-like systematic deviations from the rules of probability theory. Moreover, when prompted to judge the same event, the mean-variance relationship of probability judgments produced by LLMs shows an inverted-U-shaped like that seen in humans. We propose that these deviations from rationality can be explained by linking autoregressive LLMs to implicit Bayesian inference and drawing parallels with the Bayesian Sampler model of human probability judgments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    LLM probability estimates violate the law of total probability across partitions, and subgroup-aggregated estimates often beat direct population-level estimates (the macro fallacy).

  2. Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs struggle to use passive observations for reverse engineering, but active intervention improves performance, largely through the process of generating queries rather than the data obtained.

  3. Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Behavioral and neural representations of risk are aligned via regression to produce steering vectors that reliably shift LLM risk preferences across tasks.

  4. Let's Think Var-by-Var: Large Language Models Enable Ad Hoc Probabilistic Reasoning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A pipeline that turns LLM-proposed variables and moment constraints into an ad hoc log-linear probability model answers guesstimation questions about as well as direct prompting, and no better.

Pith tools