Pith. sign in

REVIEW 11 cited by

Centaur: a foundation model of human cognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20268 v3 pith:F4F5OQQ7 submitted 2024-10-26 cs.LG

classification cs.LG
keywords humanmodelbehaviorcentaurmodelscomputationalbeencaptures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational models, we currently do not have one model that captures the human mind in its entirety. A first step in this direction is to create a model that can predict human behavior in a wide range of settings. Here we introduce Centaur, a computational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the-art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 participants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model's internal representations become more aligned with human neural activity after finetuning. Taken together, our results demonstrate that it is possible to discover computational models that capture human behavior across a wide range of domains. We believe that such models provide tremendous potential for guiding the development of cognitive theories and present a case study to demonstrate this.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Informing AI Policy Assessment using Large-Scale Simulation of Interventions

    cs.CY 2026-04 conditional novelty 6.5 of 10

    A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.

  2. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  3. Mixture of Cognitive Experts in Large Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.

  4. Can Vision Language Models Learn Intuitive Physics from Interaction?

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Training VLMs through interaction (GRPO) does not yield generalizable physical intuitions beyond within-task performance, matching—not exceeding—supervised fine-tuning.

  5. Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A hybrid language-model and probabilistic-program architecture predicts human judgments on novel open-world reasoning vignettes better than language-model-only baselines.

  6. Aligning LLM with human travel choices: a persona-based embedding learning approach

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.

  7. Automated scientific minimization of regret

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ASMR couples the Centaur foundation model with Qwen3 to flag and repair a simple cognitive model's failures, reaching Centaur-level AIC after five automated iterations.

  8. Be.FM: Open Foundation Models for Human Behavior

    cs.AI 2025-05 reject novelty 5.0 of 10

    Be.FM fine-tunes Llama models on behavioral data and claims improved behavior prediction, but its headline evaluation is compromised by testing on the same data it trained on.

  9. Using LLMs to Advance the Cognitive Science of Collectives

    q-bio.NC 2025-05 conditional novelty 5.0 of 10

    A position paper arguing that LLMs can help cognitive scientists study collective behavior along structural, interactional, and individual complexity axes, with cautions about bias and reproducibility.

  10. Towards Measurement Theory for Artificial Intelligence

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A formal measurement theory for AI, built from representational measurement theory, measure theory, metrology, and psychometrics, would make evaluations of AI systems commensurable and scientifically grounded.

  11. Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

    cs.AI 2025-10 conditional novelty 1.0 of 10

    This paper is a survey: it organizes existing foundation-model work in neuroscience into five application domains and lists public datasets, without presenting new experiments.

Pith tools