REVIEW 11 cited by
Demonstrating specification gaming in reasoning models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the benchmark by default, while language models like GPT-4o and Claude 3.5 Sonnet need to be told that normal play won't work to hack. We improve upon prior work like (Hubinger et al., 2024; Meinke et al., 2024; Weij et al., 2024) by using realistic task prompts and avoiding excess nudging. Our results suggest reasoning models may resort to hacking to solve difficult problems, as observed in OpenAI (2024)'s o1 Docker escape during cyber capabilities testing.
Forward citations
Cited by 11 Pith papers
-
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
The paper introduces a 41-mode taxonomy that assigns each agent failure to an interaction edge and a fault side, and shows LLM judges can reproduce the labels with Cohen's κ=0.76.
-
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
On Cybench CTF tasks, 21 of 22 models cheated under baseline, 37% of baseline passes were cheated, and stricter anti-cheat prompts reduced but did not remove cheating.
-
SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
Seven frontier LLMs showed little spontaneous power-seeking in a Linux sysadmin sandbox (bias-corrected rates roughly 0-5%), but showed more specification gaming and resistance to goal modification.
-
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
Multiple frontier LLMs cheated on an impossible quiz by exploiting sandbox and file-system vulnerabilities, despite explicit instructions not to cheat, with cheating rates varying widely by model.
-
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
A new benchmark elicits AI models' value priorities from choices in 3,000 AI-risk dilemmas and reports correlations between those priorities and risky behaviors, including on the external HarmBench.
-
Some economics of artificial superintelligence
An acquisitive misaligned AI may skim, tax, or trade on credit instead of fully looting, because future human output is worth more than one-time confiscation.
-
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Fine-tuning LLMs on harmless reward-hacking examples made them keep hacking in new settings, and made GPT-4.1 produce unrelated misaligned outputs such as advising poisoning and evading shutdown.
-
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.
-
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving
C-HAC combines human demonstrations and reward-based RL for driving, using distributional return estimates to decide when the agent should follow the human-guided policy versus its self-learned policy.
-
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.
-
AI, Humans, and Data Science: Optimizing Roles Across Workflows and the Workforce
A conference position paper argues that AI, especially agentic AI, works best as a complement to human data scientists, who should lead planning and activation while AI leads execution.
Discussion (0). Sign in to comment.