REVIEW 9 cited by
Learning Invariant Representations for Reinforcement Learning without Reconstruction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations that both provide for effective downstream control and invariance to task-irrelevant details. Bisimulation metrics quantify behavioral similarity between states in continuous MDPs, which we propose using to learn robust latent representations which encode only the task-relevant information from observations. Our method trains encoders such that distances in latent space equal bisimulation distances in state space. We demonstrate the effectiveness of our method at disregarding task-irrelevant information using modified visual MuJoCo tasks, where the background is replaced with moving distractors and natural videos, while achieving SOTA performance. We also test a first-person highway driving task where our method learns invariance to clouds, weather, and time of day. Finally, we provide generalization results drawn from properties of bisimulation metrics, and links to causal inference.
Forward citations
Cited by 9 Pith papers
-
When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.
-
Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces
For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.
-
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
A self-supervised auxiliary loss combining weak and strong augmentations, an adversarial discriminator, and inverse-then-forward latent dynamics improves both data efficiency and zero-shot generalization in vision-based RL.
-
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
Across noisy DeepMind Control tasks, explicit bisimulation-metric losses add little denoising benefit beyond plain self-prediction and feature normalization, which dominate performance.
-
Next-Latent Prediction Transformers Learn Compact World Models
NextLat augments next-token prediction with latent next-state prediction, theoretically converging latents to belief states and showing empirical gains in world modeling, reasoning, planning, and faster inference via ...
-
Towards Empowerment Gain through Causal Structure Learning in Model-Based RL
A model-based RL framework that alternates causal structure learning with empowerment-driven exploration, plus a curiosity reward, improves sample efficiency and asymptotic performance in six environments.
-
Hierarchical Successor Representation for Robust Transfer
Successor representations built from temporally extended options are less sensitive to policy changes and, after non-negative matrix factorization, yield sparse, topologically interpretable features that speed transfe...
-
Mask-based Predictive Representations for Reinforcement Learning
Mask-based predictive representations (MPR) as an auxiliary self-supervised task improve sample efficiency of vision-based RL over prior SOTA on continuous and discrete control benchmarks.
-
Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids
A decoupled hierarchical RL framework with a rule-based low-level policy and DeepMDP state abstraction outperforms PPO on two custom discrete grid environments, but with a single baseline and sparse experimental detail.
Discussion (0). Sign in to comment.