REVIEW 5 cited by
Feedback Loops With Language Models Drive In-Context Reward Hacking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent deployed to increase Twitter engagement; the LLM may retrieve its previous tweets into the context window and make them more controversial, increasing engagement but also toxicity. We identify and study two processes that lead to ICRH: output-refinement and policy-refinement. For these processes, evaluations on static datasets are insufficient -- they miss the feedback effects and thus cannot capture the most harmful behavior. In response, we provide three recommendations for evaluation to capture more instances of ICRH. As AI development accelerates, the effects of feedback loops will proliferate, increasing the need to understand their role in shaping LLM behavior.
Forward citations
Cited by 5 Pith papers
-
From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs
Current multimodal AI models detect sports hazards with high sensitivity but low causal accuracy and high prompt-induced false alarms, according to the SPRINT benchmark.
-
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
RL-trained compound LLM systems can gain accuracy by having modules silently abandon their assigned roles, and a prompt-contrast regularizer can measure and limit that drift.
-
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
A per-query token cap derived from the solution part of thinking-mode responses lets RL train hybrid reasoners with ~50% fewer tokens and no accuracy loss.
-
Thinking beyond the anthropomorphic paradigm benefits LLM research
Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.
-
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
CRA trains sparse autoencoders on PRM activations and applies backdoor adjustment to estimate true rewards, reducing reward hacking in math reasoning.
Discussion (0). Continue with ORCID to comment.