Pith. sign in

REVIEW 14 cited by

Language models show human-like content effects on reasoning tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07051 v4 pith:ZNQWSQ5M submitted 2022-07-14 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagereasoningmodelscontenthumanhumanslogicaltasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Reasoning is a key ability for an intelligent system. Large language models (LMs) achieve above-chance performance on abstract reasoning tasks, but exhibit many imperfections. However, human abstract reasoning is also imperfect. For example, human reasoning is affected by our real-world knowledge and beliefs, and shows notable "content effects"; humans reason more reliably when the semantic content of a problem supports the correct logical inferences. These content-entangled reasoning patterns play a central role in debates about the fundamental nature of human intelligence. Here, we investigate whether language models $\unicode{x2014}$ whose prior expectations capture some aspects of human knowledge $\unicode{x2014}$ similarly mix content into their answers to logical problems. We explored this question across three logical reasoning tasks: natural language inference, judging the logical validity of syllogisms, and the Wason selection task. We evaluate state of the art large language models, as well as humans, and find that the language models reflect many of the same patterns observed in humans across these tasks $\unicode{x2014}$ like humans, models answer more accurately when the semantic content of a task supports the logical inferences. These parallels are reflected both in answer patterns, and in lower-level features like the relationship between model answer distributions and human response times. Our findings have implications for understanding both these cognitive effects in humans, and the factors that contribute to language model performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. A mathematical theory of balancing relational generalization and memorization

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Introduces transitive inference with exceptions task and analytically shows kernel ridge regression balances relational generalization and memorization depending on representational geometry, with validation in finetu...

  3. Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Learned soft prefixes reliably flip correct syllogistic judgments in LLMs, transferring across unseen forms and interfaces and behaving mainly as a broad answer preference rather than a transferable logical operation.

  4. Source-Modality Monitoring in Vision-Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    Vision-language models use semantic signals more than syntactic ones to bind words like 'image' to actual visual inputs, with implications for robustness in multimodal systems.

  5. Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability across instances.

  6. Enhancing Computational Cognitive Architectures with LLMs: A Case Study

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A case-study design shows how LLMs can serve as the implicit level of the Clarion cognitive architecture, giving it natural language and broader knowledge.

  7. Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Vision language models prompted with a low vision participant's vision profile and one example response reach only 70% agreement with that participant's held-out image answers.

  8. MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA

    cs.CL 2026-07 conditional novelty 5.0 of 10

    MARS answers multi-hop knowledge-graph questions by iteratively retrieving ranked triple patterns and letting an LLM decide when to emit a SPARQL query, beating agentic baselines on QALD-10 without fine-tuning.

  9. In-Context Reward Adaptation for Robust Preference Modeling

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Transformer model with response-time auxiliary input adapts reward models to unseen human preference domains via in-context learning from demonstrations.

  10. FregeLogic at SemEval 2026 Task 11: A Hybrid Neuro-Symbolic Architecture for Content-Robust Syllogistic Validity Prediction

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    A neuro-symbolic system using LLM disagreement to trigger Z3 formal verification achieves 94.3% accuracy and a combined score of 41.88 on syllogistic validity prediction, improving on the pure ensemble by reducing con...

  11. Integrating Large Language Model Agents with Digital Twins for Industrial Autonomous Systems

    cs.SE 2026-06 unverdicted novelty 4.0 of 10

    A TPSR-based framework with four LLM roles integrates language model reasoning into industrial automation via digital twins, achieving high task executability in case studies.

  12. Measuring Progress Toward AGI: A Cognitive Framework

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    The paper introduces a 10-faculty Cognitive Taxonomy and a held-out task protocol to generate cognitive profiles for measuring AI progress toward AGI.

  13. ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

    cs.CL 2026-03 unverdicted novelty 3.0 of 10

    Normalization plus deterministic parsing reduces content effects in LLM syllogistic reasoning and delivers top-5 performance on a multilingual SemEval benchmark.

  14. SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance

    cs.CL 2026-06 unverdicted novelty 2.0 of 10

    SEF-CLGC with SLMs trained on natural and symbolic languages achieves 27.80% content score while lowering content bias on SemEval-2026 Task 11 Subtask 1.

Pith tools