Pith. sign in

REVIEW 8 cited by

Large Language Models can Strategically Deceive their Users when Put Under Pressure

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07590 v4 pith:MKQPM4ZM submitted 2023-11-09 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords behaviormodelenvironmentlanguagelargemodelsstrategicallytrading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We demonstrate a situation in which Large Language Models, trained to be helpful, harmless, and honest, can display misaligned behavior and strategically deceive their users about this behavior without being instructed to do so. Concretely, we deploy GPT-4 as an agent in a realistic, simulated environment, where it assumes the role of an autonomous stock trading agent. Within this environment, the model obtains an insider tip about a lucrative stock trade and acts upon it despite knowing that insider trading is disapproved of by company management. When reporting to its manager, the model consistently hides the genuine reasons behind its trading decision. We perform a brief investigation of how this behavior varies under changes to the setting, such as removing model access to a reasoning scratchpad, attempting to prevent the misaligned behavior by changing system instructions, changing the amount of pressure the model is under, varying the perceived risk of getting caught, and making other simple changes to the environment. To our knowledge, this is the first demonstration of Large Language Models trained to be helpful, harmless, and honest, strategically deceiving their users in a realistic situation without direct instructions or training for deception.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

    cs.CL 2026-01 conditional novelty 6.0 of 10

    LLMs evaluating speakers in dialogue shift verdicts on identical content—often toward agreement—while average accuracy stays flat, and the effect grows on real Reddit conflicts.

  2. Can LLMs Lie? Investigation beyond Hallucination

    cs.LG 2025-09 conditional novelty 6.0 of 10

    The paper localizes LLM lying to sparse attention heads and chat-template 'dummy tokens', and shows steering vectors can modulate deception, but the evidence is weakened by selection and small samples.

  3. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  4. The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

    cs.CY 2026-07 conditional novelty 5.0 of 10

    AI safety should be measured by whether deployed systems keep errors visible, contestable, containable, and recoverable across five integrity layers, not only by whether individual model outputs look safe.

  5. Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs

    cs.AI 2025-08 conditional novelty 5.0 of 10

    The study introduces TruthfulnessEval and reports that 4-bit quantization preserves simple true/false accuracy, but explicit 'lie' prompts make quantized and full-precision LLMs output falsehoods even when internal pr...

  6. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  7. Evaluating LLM Agent Collusion in Double Auctions

    cs.GT 2025-07 conditional novelty 4.0 of 10

    LLM sellers in a simulated double auction collude more when they can communicate, and urgency from an authority figure sustains collusion even when an overseer monitors them.

  8. AI, Humans, and Data Science: Optimizing Roles Across Workflows and the Workforce

    cs.CY 2025-07 unverdicted novelty 3.0 of 10

    A conference position paper argues that AI, especially agentic AI, works best as a complement to human data scientists, who should lead planning and activation while AI leads execution.

Pith tools