Pith. sign in

REVIEW 2 cited by

Regularized Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.03691 v2 pith:6ZPHRC4I submitted 2020-10-07 cs.LG

classification cs.LG
keywords expertregularizedsolutionsbehaviorinverselearnerlearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Inverse Reinforcement Learning (IRL) aims to facilitate a learner's ability to imitate expert behavior by acquiring reward functions that explain the expert's decisions. Regularized IRL applies strongly convex regularizers to the learner's policy in order to avoid the expert's behavior being rationalized by arbitrary constant rewards, also known as degenerate solutions. We propose tractable solutions, and practical methods to obtain them, for regularized IRL. Current methods are restricted to the maximum-entropy IRL framework, limiting them to Shannon-entropy regularizers, as well as proposing the solutions that are intractable in practice. We present theoretical backing for our proposed IRL method's applicability for both discrete and continuous controls, empirically validating our performance on a variety of tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inverse Reinforcement Learning using Revealed Preferences and Passive Stochastic Optimization

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A three-chapter monograph that uses Afriat's theorem and Bayesian revealed preference tests for inverse reinforcement learning, plus a passive Langevin dynamics algorithm for real-time reward reconstruction.

  2. Reward Models in Deep Reinforcement Learning: A Survey

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A structured survey of reward modeling in deep RL, proposing a three-axis taxonomy and reviewing applications and evaluation methods.

Pith tools