Pith. sign in

REVIEW 4 cited by

Emergent Analogical Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09196 v3 pith:WA6BXK6S submitted 2022-12-19 cs.AI cs.CL

classification cs.AIcs.CL
keywords modelshumanlanguagelargeabilitygpt-3analogicalanalogy
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent advent of large language models has reinvigorated debate over whether human cognitive capacities might emerge in such generic models given sufficient training data. Of particular interest is the ability of these models to reason about novel problems zero-shot, without any direct training. In human cognition, this capacity is closely tied to an ability to reason by analogy. Here, we performed a direct comparison between human reasoners and a large language model (the text-davinci-003 variant of GPT-3) on a range of analogical tasks, including a non-visual matrix reasoning task based on the rule structure of Raven's Standard Progressive Matrices. We found that GPT-3 displayed a surprisingly strong capacity for abstract pattern induction, matching or even surpassing human capabilities in most settings; preliminary tests of GPT-4 indicated even better performance. Our results indicate that large language models such as GPT-3 have acquired an emergent ability to find zero-shot solutions to a broad range of analogy problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiG-bench: Discovery in Games

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A 70-game interactive benchmark where agents must discover hidden rules and objectives, with human beatability on every game and frontier models failing on the hardest tiers.

  2. Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Two-phase pretraining, wide web data first and high-quality data second, improves average downstream accuracy by 3.4% over random ordering and 17% over natural token distribution, and the best 1T-scale blend transfers...

  3. Investigating social alignment via mirroring in a system of interacting language models

    cs.MA 2024-12 conditional novelty 5.0 of 10

    A simulation of 30 interacting language models shows that the range of communication shapes opinion clusters, and higher rates of mirroring (agreement) delays or prevents consensus.

  4. Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

    cs.AI 2024-12 conditional novelty 3.0 of 10

    A roadmap paper argues that reproducing o1 hinges on four RL components, policy initialization, reward design, search, and learning, and frames existing open-source o1 projects as instances of this roadmap.

Pith tools