REVIEW 1 cited by
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider the questions of whether or not large language models (LLMs) have beliefs, and, if they do, how we might measure them. First, we evaluate two existing approaches, one due to Azaria and Mitchell (2023) and the other to Burns et al. (2022). We provide empirical results that show that these methods fail to generalize in very basic ways. We then argue that, even if LLMs have beliefs, these methods are unlikely to be successful for conceptual reasons. Thus, there is still no lie-detector for LLMs. After describing our empirical results we take a step back and consider whether or not we should expect LLMs to have something like beliefs in the first place. We consider some recent arguments aiming to show that LLMs cannot have beliefs. We show that these arguments are misguided. We provide a more productive framing of questions surrounding the status of beliefs in LLMs, and highlight the empirical nature of the problem. We conclude by suggesting some concrete paths for future work.
Forward citations
Cited by 1 Pith paper
-
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
Modeling the full token-by-token trajectory of LLM hidden states with neural ODEs, CDEs, and SDEs improves hallucination detection by over 14% AUC on a constructed true/false benchmark, though gains shrink on QA datasets.
Discussion (0). Continue with ORCID to comment.