Pith. sign in

REVIEW 3 cited by

Scalable Bayesian Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.06483 v2 pith:LCTU3VPE submitted 2021-02-12 cs.LG

classification cs.LG
keywords learningrewardbayesianmethodswellalongsideapproximatebeyond
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bayesian inference over the reward presents an ideal solution to the ill-posed nature of the inverse reinforcement learning problem. Unfortunately current methods generally do not scale well beyond the small tabular setting due to the need for an inner-loop MDP solver, and even non-Bayesian methods that do themselves scale often require extensive interaction with the environment to perform well, being inappropriate for high stakes or costly applications such as healthcare. In this paper we introduce our method, Approximate Variational Reward Imitation Learning (AVRIL), that addresses both of these issues by jointly learning an approximate posterior distribution over the reward that scales to arbitrarily complicated state spaces alongside an appropriate policy in a completely offline manner through a variational approach to said latent reward. Applying our method to real medical data alongside classic control simulations, we demonstrate Bayesian reward inference in environments beyond the scope of current methods, as well as task performance competitive with focused offline imitation learning algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributional Inverse Reinforcement Learning

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    DistIRL recovers reward distributions and risk-aware policies from offline demonstrations by minimizing first-order stochastic dominance violations between agent and expert returns.

  2. Detection of coordinated fleet vehicles in route choice urban games. Part I. Inverse fleet assignment theory

    math.OC 2025-06 conditional novelty 6.0 of 10

    Fleet vehicle flows can be recovered from total flows exactly when the fleet strategy is more selfish than altruistic, and the paper shows failure cases for social and altruistic fleets.

  3. On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A variational autoencoder over skill sequences produces label-free trajectory embeddings that separate and imitate policies of different ability levels in MuJoCo control tasks.

Pith tools