Pith. sign in

REVIEW 7 cited by

Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.05230 v1 pith:V4YSCVKF submitted 2023-12-08 cs.AI cs.CLcs.CVcs.LGcs.RO

classification cs.AIcs.CLcs.CVcs.LGcs.RO
keywords modelsreasoninglanguageworldagentplanningcapabilitieselements
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite their tremendous success in many applications, large language models often fall short of consistent reasoning and planning in various (language, embodied, and social) scenarios, due to inherent limitations in their inference, learning, and modeling capabilities. In this position paper, we present a new perspective of machine reasoning, LAW, that connects the concepts of Language models, Agent models, and World models, for more robust and versatile reasoning capabilities. In particular, we propose that world and agent models are a better abstraction of reasoning, that introduces the crucial elements of deliberate human-like reasoning, including beliefs about the world and other agents, anticipation of consequences, goals/rewards, and strategic planning. Crucially, language models in LAW serve as a backend to implement the system or its elements and hence provide the computational power and adaptability. We review the recent studies that have made relevant progress and discuss future research directions towards operationalizing the LAW framework.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    LLMs estimate pulley mechanical advantage above chance via a pulley-counting heuristic, but fail to distinguish functional from connected-but-nonfunctional systems, indicating brittle world-model use.

  2. Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning

    cs.AI 2024-12 conditional novelty 6.0 of 10

    LLM Theory of Mind benchmarks focus on logical inference at a fixed mentalizing depth, but overlook the prior step of deciding whether and how deep to mentalize, which the paper argues should be measured with interact...

  3. GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GEM generates controllable future RGB and depth ego-vision frames, conditioned on ego-trajectories, sparse object tokens, and human poses, across driving, egocentric, and drone domains.

  4. PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making

    cs.AI 2024-11 reject novelty 6.0 of 10

    An LLM can generate executable world-model components that, combined with MCTS, outperform LLM-as-policy in GOPS and match it in Taboo, though the evaluation under-supports the partial-observability claim.

  5. Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A position paper advocating large-scale training of event forecasting LLMs, with proposals for label selection, counterfactual training data, auxiliary rewards, and multi-source datasets.

  6. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A systematic survey and cross-benchmark evaluation showing that multimodal LLMs can recognize humor artifacts but still struggle to interpret the intended meaning and mechanisms of visual humor.

  7. Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

    cs.AI 2025-07 conditional novelty 2.0 of 10

    A broad survey arguing that AGI requires modular, memory-augmented, embodied architectures rather than scaled-up token prediction, with a brief proposal to decompose intelligence into five components.

Pith tools