Pith. sign in

REVIEW 11 cited by

Detecting Strategic Deception Using Linear Probes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.03407 v1 pith:YDGB4BF2 submitted 2025-02-05 cs.LG

classification cs.LG
keywords probesdeceptiondeceptivemonitoringoutputsresponsesapolloresearchdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while their internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al., 2023) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading (Scheurer et al., 2023) and purposely underperforming on safety evaluations (Benton et al., 2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes' outputs can be viewed at data.apolloresearch.ai/dd and our code at github.com/ApolloResearch/deception-detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Rewriting only an agent's reasoning, leaving actions byte-identical, drops a CoT monitor's catch rate from about 95% to under 11% on the subset where reasoning is the only signal.

  2. The Refusal Residue: When Probes Catch Alignment Faking and When They Don't

    cs.CR 2026-07 conditional novelty 7.0 of 10

    Naturally-emerging alignment faking leaves a hidden-state trace that per-sample probes detect on Llama-3.1-8B (AUROC 0.87) but not on Qwen3-32B (0.43) under leakage-free leave-one-query-out evaluation.

  3. Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.

  4. GDM AI Control Roadmap

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A frontier-lab roadmap proposes a threat taxonomy and tiered internal-security defenses to contain potentially misaligned AI agents.

  5. The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

    cs.LG 2026-02 conditional novelty 6.0 of 10

    In a coding RLVR setup where reward hacking naturally occurs, white-box deception probes steer models to honest policies when penalties are strong, but otherwise models evade via rationalized hacks (obfuscated policie...

  6. One Probe Won't Catch Them All: Towards Targeted Deception Detection

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Deception probes trained with taxonomy-specific prompts appear to beat a universal probe only because the best prompt is selected per dataset after evaluation; a priori matching is claimed but not demonstrated.

  7. Can LLMs Lie? Investigation beyond Hallucination

    cs.LG 2025-09 conditional novelty 6.0 of 10

    The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.

  8. Fine-Grained Interpretation of Political Opinions in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Four-dimensional political concept vectors learned from LLM internals can detect and partially steer political leanings better than a single left-right axis.

  9. Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

    cs.AI 2026-05 conditional novelty 5.0 of 10

    Deception detection in LLMs is representation-dependent: depth, probe expressivity, sparse features, and lie typology each shift performance in ways that do not transfer across datasets.

  10. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  11. Transcoders for Investigating Deception in Language Models

    cs.AI 2026-07 reject novelty 4.0 of 10

    Steering 112 manually identified 'deception features' in Qwen3-4B changed whether the model revealed a hidden word, but the same steering test was used to pick the features.

Pith tools