REVIEW 4 cited by
Imitation Learning via Off-Policy Distribution Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When performing imitation learning from expert demonstrations, distribution matching is a popular approach, in which one alternates between estimating distribution ratios and then using these ratios as rewards in a standard reinforcement learning (RL) algorithm. Traditionally, estimation of the distribution ratio requires on-policy data, which has caused previous work to either be exorbitantly data-inefficient or alter the original objective in a manner that can drastically change its optimum. In this work, we show how the original distribution ratio estimation objective may be transformed in a principled manner to yield a completely off-policy objective. In addition to the data-efficiency that this provides, we are able to show that this objective also renders the use of a separate RL optimization unnecessary.Rather, an imitation policy may be learned directly from this objective without the use of explicit rewards. We call the resulting algorithm ValueDICE and evaluate it on a suite of popular imitation learning benchmarks, finding that it can achieve state-of-the-art sample efficiency and performance.
Forward citations
Cited by 4 Pith papers
-
Distributional Inverse Reinforcement Learning
DistIRL recovers reward distributions and risk-aware policies from offline demonstrations by minimizing first-order stochastic dominance violations between agent and expert returns.
-
Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
VfO trains a state-value function on action-free expert demonstrations mixed with lower-quality background data, then uses advantage-weighted regression on the background data to improve the agent, approaching oracle ...
-
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
VAC is a new actor-critic method with a single optimistic objective and a provably near-optimal regret bound in linear Markov decision processes.
-
Diffusion-Modeled Reinforcement Learning for Carbon and Risk-Aware Microgrid Optimization
DiffCarl, a diffusion-actor variant of SAC with carbon pricing and CVaR risk terms, is reported to lower microgrid operating cost by 2.3-30.1% versus baselines, though the paper's own numbers contradict its 28.7% carb...
Discussion (0). Continue with ORCID to comment.