REVIEW 11 cited by
Detecting Strategic Deception Using Linear Probes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while their internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al., 2023) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading (Scheurer et al., 2023) and purposely underperforming on safety evaluations (Benton et al., 2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes' outputs can be viewed at data.apolloresearch.ai/dd and our code at github.com/ApolloResearch/deception-detection.
Forward citations
Cited by 11 Pith papers
-
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
Rewriting only an agent's reasoning, leaving actions byte-identical, drops a CoT monitor's catch rate from about 95% to under 11% on the subset where reasoning is the only signal.
-
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Naturally-emerging alignment faking leaves a hidden-state trace that per-sample probes detect on Llama-3.1-8B (AUROC 0.87) but not on Qwen3-32B (0.43) under leakage-free leave-one-query-out evaluation.
-
Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.
-
GDM AI Control Roadmap
A frontier-lab roadmap proposes a threat taxonomy and tiered internal-security defenses to contain potentially misaligned AI agents.
-
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
In a coding RLVR setup where reward hacking naturally occurs, white-box deception probes steer models to honest policies when penalties are strong, but otherwise models evade via rationalized hacks (obfuscated policie...
-
One Probe Won't Catch Them All: Towards Targeted Deception Detection
Deception probes trained with taxonomy-specific prompts appear to beat a universal probe only because the best prompt is selected per dataset after evaluation; a priori matching is claimed but not demonstrated.
-
Can LLMs Lie? Investigation beyond Hallucination
The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.
-
Fine-Grained Interpretation of Political Opinions in Large Language Models
Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.
-
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Deception detection in LLMs is representation-dependent: depth, probe expressivity, sparse features, and lie typology each shift performance in ways that do not transfer across datasets.
-
BlueGlass: A Framework for Composite AI Safety
BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...
-
Transcoders for Investigating Deception in Language Models
Steering 112 manually identified 'deception features' in Qwen3-4B changed whether the model revealed a hidden word, but the same steering test was used to pick the features.
Discussion (0). Continue with ORCID to comment.