Introduces BeliefTrack benchmark diagnosing three CBM failures in LLMs and shows RL with belief-state rewards cuts failure rates by 70.9% while representation steering cuts them by 46.1%.
human-interpretable
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.AI 2years
2026 2representative citing papers
Re-examination of two LLM introspection paradigms with new controls shows models lack privileged access to internal states, performing equivalently with input-only classifiers or near chance on relabeled tasks.
citing papers explorer
-
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Introduces BeliefTrack benchmark diagnosing three CBM failures in LLMs and shows RL with belief-state rewards cuts failure rates by 70.9% while representation steering cuts them by 46.1%.
-
Can LLMs Introspect? A Reality Check
Re-examination of two LLM introspection paradigms with new controls shows models lack privileged access to internal states, performing equivalently with input-only classifiers or near chance on relabeled tasks.