Pith. sign in

REVIEW 4 cited by

The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05859 v1 pith:RDYRHDKH submitted 2024-08-11 cs.AI

classification cs.AI
keywords cognitiveinterpretationrepresentationsbehaviorlearnedmodelsalgorithmsbeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Artificial neural networks have long been understood as "black boxes": though we know their computation graphs and learned parameters, the knowledge encoded by these weights and functions they perform are not inherently interpretable. As such, from the early days of deep learning, there have been efforts to explain these models' behavior and understand them internally; and recently, mechanistic interpretability (MI) has emerged as a distinct research area studying the features and implicit algorithms learned by foundation models such as large language models. In this work, we aim to ground MI in the context of cognitive science, which has long struggled with analogous questions in studying and explaining the behavior of "black box" intelligent systems like the human brain. We leverage several important ideas and developments in the history of cognitive science to disentangle divergent objectives in MI and indicate a clear path forward. First, we argue that current methods are ripe to facilitate a transition in deep learning interpretation echoing the "cognitive revolution" in 20th-century psychology that shifted the study of human psychology from pure behaviorism toward mental representations and processing. Second, we propose a taxonomy mirroring key parallels in computational neuroscience to describe two broad categories of MI research, semantic interpretation (what latent representations are learned and used) and algorithmic interpretation (what operations are performed over representations) to elucidate their divergent goals and objects of study. Finally, we elaborate the parallels and distinctions between various approaches in both categories, analyze the respective strengths and weaknesses of representative works, clarify underlying assumptions, outline key challenges, and discuss the possibility of unifying these modes of interpretation under a common framework.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

    cs.AI 2026-08 conditional novelty 7.0 of 10

    An agentic system called Mechanist autonomously discovers mechanisms of AI behavior, including hidden safety risks and separable internal 'belief heads', and uses them to steer model outputs.

  2. Thinking beyond the anthropomorphic paradigm benefits LLM research

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.

  3. AI Behavioral Science

    cs.HC 2025-08 unverdicted novelty 3.0 of 10

    A position paper outlines a research agenda for AI Behavioral Science built on three pillars: assessing AI behavior, using AI as a behavioral science tool, and understanding human-AI ecosystems.

  4. Seeing Through Risk: A Symbolic Approximation of Prospect Theory

    cs.AI 2025-04 reject novelty 2.0 of 10

    A logistic regression with Prospect-Theory-inspired features is validated on synthetic data generated from the same logistic equation, making the validation circular.

Pith tools