REVIEW 7 cited by
Bridging State and History Representations: Understanding Self-Predictive RL
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Representations are at the core of all deep reinforcement learning (RL) methods for both Markov decision processes (MDPs) and partially observable Markov decision processes (POMDPs). Many representation learning methods and theoretical frameworks have been developed to understand what constitutes an effective representation. However, the relationships between these methods and the shared properties among them remain unclear. In this paper, we show that many of these seemingly distinct methods and frameworks for state and history abstractions are, in fact, based on a common idea of self-predictive abstraction. Furthermore, we provide theoretical insights into the widely adopted objectives and optimization, such as the stop-gradient technique, in learning self-predictive representations. These findings together yield a minimalist algorithm to learn self-predictive representations for states and histories. We validate our theories by applying our algorithm to standard MDPs, MDPs with distractors, and POMDPs with sparse rewards. These findings culminate in a set of preliminary guidelines for RL practitioners.
Forward citations
Cited by 7 Pith papers
-
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
Across noisy DeepMind Control tasks, explicit bisimulation-metric losses add little denoising benefit beyond plain self-prediction and feature normalization, which dominate performance.
-
Hierarchical Latent Prediction for Language Models
HiLP adds a hierarchical latent prediction objective to LM pretraining, improving coding and multi-step reasoning benchmarks and speculative decoding acceptance, with zero inference-time overhead.
-
Can We Really Learn One Representation to Optimize All Rewards?
Finite-dimensional FB representations cannot exactly encode all rewards in continuous control; a new one-step FB variant that fits the behavioral policy converges better and beats FB on average.
-
Hadamax Encoding: Elevating Performance in Model-Free Atari
Hadamax, a Hadamard-product and max-pooling encoder, improves PQN's median human-normalized Atari-57 score by about 80% with no algorithmic changes.
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
The University AI Didn't Replace -- Rethinking Universities in the AI Era
Universities stuck in informal AI use must move to strategic integration that redesigns learning around AI-supported reasoning and aligns policy, workload, and recognition.
-
Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
The informed asymmetric actor-critic lets a critic condition on arbitrary state-dependent privileged signals with an unbiased policy gradient, plus HSCIC and return-prediction tests for signal selection.
Discussion (0). Continue with ORCID to comment.