Pith. sign in

REVIEW 2 cited by

Avoiding Tampering Incentives in Deep RL via Decoupled Approval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.08827 v1 pith:O7NYUHPL submitted 2020-11-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords approvaldecoupledfeedbackagentsalgorithmsincentivesinfluenceabletampering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can thus perform poorly in settings where agents can tamper with the reward-generating mechanism. We present a principled solution to the problem of learning from influenceable feedback, which combines approval with a decoupled feedback collection procedure. For a natural class of corruption functions, decoupled approval algorithms have aligned incentives both at convergence and for their local updates. Empirically, they also scale to complex 3D environments where tampering is possible.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A post-hoc audit framework finds that many agent benchmark scores are inflated by protocol exposures (public answers, readable hidden state, generator regularities, feedback, or scoring flaws), with measured inflation...

  2. RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    RIVAL iteratively re-trains a reward model adversarially against the current translator and adds a BLEU-predicting head, improving in-domain WMT and subtitle translation over SFT baselines.

Pith tools