Pith. sign in

REVIEW 1 cited by

Provably Improved Context-Based Offline Meta-RL with Attention and Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.10774 v2 pith:IMYLM7G2 submitted 2021-02-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningalgorithmstaskattentioncontext-basedcontrastivemeta-rloffline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Meta-learning for offline reinforcement learning (OMRL) is an understudied problem with tremendous potential impact by enabling RL algorithms in many real-world applications. A popular solution to the problem is to infer task identity as augmented state using a context-based encoder, for which efficient learning of robust task representations remains an open challenge. In this work, we provably improve upon one of the SOTA OMRL algorithms, FOCAL, by incorporating intra-task attention mechanism and inter-task contrastive learning objectives, to robustify task representation learning against sparse reward and distribution shift. Theoretical analysis and experiments are presented to demonstrate the superior performance and robustness of our end-to-end and model-free framework compared to prior algorithms across multiple meta-RL benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CausalCOMRL: Context-Based Offline Meta-Reinforcement Learning with Causal Representation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    CausalCOMRL learns task representations through a causal VAE plus mutual information and contrastive losses, claiming improved out-of-distribution generalization in context-based offline meta-RL.

Pith tools