Pith. sign in

REVIEW 1 cited by

Efficient iterative policy optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.08967 v1 pith:ARSV4NVL submitted 2016-12-28 cs.AI cs.LGcs.RO

classification cs.AIcs.LGcs.RO
keywords policygoodnumberupdatesachieveapproximatingboundsconcave
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We tackle the issue of finding a good policy when the number of policy updates is limited. This is done by approximating the expected policy reward as a sequence of concave lower bounds which can be efficiently maximized, drastically reducing the number of policy updates required to achieve good performance. We also extend existing methods to negative rewards, enabling the use of control variates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SFT on curated data is a lower bound on a sparse-reward RL objective, and an importance-weighted variant, iw-SFT, tightens the bound and beats plain SFT on AIME 2024 and GPQA.

Pith tools