Pith. sign in

REVIEW 9 cited by

Learning Invariant Representations for Reinforcement Learning without Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10742 v2 pith:GNISDVH6 submitted 2020-06-18 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningbisimulationmethodrepresentationsdistancesinformationinvariancelatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations that both provide for effective downstream control and invariance to task-irrelevant details. Bisimulation metrics quantify behavioral similarity between states in continuous MDPs, which we propose using to learn robust latent representations which encode only the task-relevant information from observations. Our method trains encoders such that distances in latent space equal bisimulation distances in state space. We demonstrate the effectiveness of our method at disregarding task-irrelevant information using modified visual MuJoCo tasks, where the background is replaced with moving distractors and natural videos, while achieving SOTA performance. We also test a first-person highway driving task where our method learns invariance to clouds, weather, and time of day. Finally, we provide generalization results drawn from properties of bisimulation metrics, and links to causal inference.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 77 citations worldwide. Full citation record

  1. When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  2. Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces

    cs.LG 2025-07 conditional novelty 7.0 of 10

    For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.

  3. Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A self-supervised auxiliary loss combining weak and strong augmentations, an adversarial discriminator, and inverse-then-forward latent dynamics improves both data efficiency and zero-shot generalization in vision-based RL.

  4. Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Across noisy DeepMind Control tasks, explicit bisimulation-metric losses add little denoising benefit beyond plain self-prediction and feature normalization, which dominate performance.

  5. Next-Latent Prediction Transformers Learn Compact World Models

    cs.LG 2025-11 unverdicted novelty 6.0 of 10

    NextLat augments next-token prediction with latent next-state prediction, theoretically converging latents to belief states and showing empirical gains in world modeling, reasoning, planning, and faster inference via ...

  6. Towards Empowerment Gain through Causal Structure Learning in Model-Based RL

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A model-based RL framework that alternates causal structure learning with empowerment-driven exploration, plus a curiosity reward, improves sample efficiency and asymptotic performance in six environments.

  7. Hierarchical Successor Representation for Robust Transfer

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Successor representations built from temporally extended options are less sensitive to policy changes and, after non-negative matrix factorization, yield sparse, topologically interpretable features that speed transfe...

  8. Mask-based Predictive Representations for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Mask-based predictive representations (MPR) as an auxiliary self-supervised task improve sample efficiency of vision-based RL over prior SOTA on continuous and discrete control benchmarks.

  9. Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A decoupled hierarchical RL framework with a rule-based low-level policy and DeepMDP state abstraction outperforms PPO on two custom discrete grid environments, but with a single baseline and sparse experimental detail.

Pith tools