REVIEW 2 cited by
Propositional Interpretability in Artificial Intelligence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mechanistic interpretability is the program of explaining what AI systems are doing in terms of their internal mechanisms. I analyze some aspects of the program, along with setting out some concrete challenges and assessing progress to date. I argue for the importance of propositional interpretability, which involves interpreting a system's mechanisms and behavior in terms of propositional attitudes: attitudes (such as belief, desire, or subjective probability) to propositions (e.g. the proposition that it is hot outside). Propositional attitudes are the central way that we interpret and explain human beings and they are likely to be central in AI too. A central challenge is what I call thought logging: creating systems that log all of the relevant propositional attitudes in an AI system over time. I examine currently popular methods of interpretability (such as probing, sparse auto-encoders, and chain of thought methods) as well as philosophical methods of interpretation (including those grounded in psychosemantics) to assess their strengths and weaknesses as methods of propositional interpretability.
Forward citations
Cited by 2 Pith papers
-
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
An LLM navigation agent encodes a coarse spatial map and multi-step plans in its activations, and reasoning shifts these representations from broad environment information to immediate action selection.
-
Explaining Neural Networks with Reasons
A new interpretability method computes 'reasons vectors' from neuron activations and measures how strongly each neuron supports propositions about the input, with experiments on MNIST, Adult, and SST2.
Discussion (0). Continue with ORCID to comment.