Pith. sign in

REVIEW 7 cited by

Language Models as Agent Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.01681 v1 pith:3UQ4YQUC submitted 2022-12-03 cs.CL cs.MA

classification cs.CLcs.MA
keywords languagemodelsagentsagentarguecommunicativecontextdocuments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models (LMs) are trained on collections of documents, written by individual human agents to achieve specific goals in an outside world. During training, LMs have access only to text of these documents, with no direct evidence of the internal states of the agents that produced them -- a fact often used to argue that LMs are incapable of modeling goal-directed aspects of human language production and comprehension. Can LMs trained on text learn anything at all about the relationship between language and use? I argue that LMs are models of intentional communication in a specific, narrow sense. When performing next word prediction given a textual context, an LM can infer and represent properties of an agent likely to have produced that context. These representations can in turn influence subsequent LM generation in the same way that agents' communicative intentions influence their language. I survey findings from the recent literature showing that -- even in today's non-robust and error-prone models -- LMs infer and use representations of fine-grained communicative intentions and more abstract beliefs and goals. Despite the limited nature of their training data, they can thus serve as building blocks for systems that communicate and act intentionally.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 11 citations worldwide. Full citation record

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures

    cs.AI 2026-05 conditional novelty 6.0 of 10

    On multi-step procedural tasks, LoRA fine-tuning underperforms full fine-tuning at every rank tested because procedural knowledge requires high-rank weight updates.

  3. Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Pairwise matrices for SAEs demonstrate that single-feature inspection mislabels causal axes, with joint suppression and matched-geometry controls revealing distinct output regimes not captured by single-feature or ran...

  4. What Does it Mean for a Neural Network to Learn a "World Model"?

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Defines a world model as a simple commutative-diagram factorization through an intermediate representation, with conditions that the model be learned and emergent rather than inherited from input or output.

  5. Thinking beyond the anthropomorphic paradigm benefits LLM research

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.

  6. Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

    cs.AI 2025-10 conditional novelty 5.0 of 10

    Human vividness-rating networks are correlated across populations and cluster by questionnaire context, whereas LLM-derived networks are mostly degenerate single-clusters, showing a human-LLM divergence in imagined-sc...

  7. Modular Speaker Architecture: A Framework for Sustaining Responsibility and Contextual Integrity in Multi-Agent AI Communication

    cs.AI 2025-06 reject novelty 3.0 of 10

    MSA is a modular role, responsibility, and context-validation framework for LLM dialogue; its pilot study reports higher annotation scores for MSA-active segments, but without random assignment, baselines, data, or code.

Pith tools