Pith. sign in

REVIEW 2 cited by

ProMP: Proximal Meta-Policy Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.06784 v4 pith:QXWIRMY4 submitted 2018-10-16 cs.LG stat.ML

classification cs.LGstat.ML
keywords assignmentcreditmeta-policymeta-rlpre-adaptationalgorithmbehaviorduring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Credit assignment in Meta-reinforcement learning (Meta-RL) is still poorly understood. Existing methods either neglect credit assignment to pre-adaptation behavior or implement it naively. This leads to poor sample-efficiency during meta-training as well as ineffective task identification strategies. This paper provides a theoretical analysis of credit assignment in gradient-based Meta-RL. Building on the gained insights we develop a novel meta-learning algorithm that overcomes both the issue of poor credit assignment and previous difficulties in estimating meta-policy gradients. By controlling the statistical distance of both pre-adaptation and adapted policies during meta-policy search, the proposed algorithm endows efficient and stable meta-learning. Our approach leads to superior pre-adaptation policy behavior and consistently outperforms previous Meta-RL algorithms in sample-efficiency, wall-clock time, and asymptotic performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer

    cs.LG 2025-05 conditional novelty 6.0 of 10

    M3DT combines a Decision Transformer with grouped, separately trained expert modules and a learned router, achieving better normalized scores than baselines across 10 to 160 multi-task RL tasks.

  2. Coreset-Based Task Selection for Sample-Efficient Meta-Reinforcement Learning

    math.OC 2025-02 conditional novelty 6.0 of 10

    A derivative-free coreset task-selection algorithm for MAML-RL trains on a small weighted task subset and provably reduces sample complexity by O(1/epsilon), provided the task-selection bias is small.

Pith tools