REVIEW 9 cited by
Goal-Conditioned Reinforcement Learning: Problems and Solutions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Goal-conditioned reinforcement learning (GCRL), related to a set of complex RL problems, trains an agent to achieve different goals under particular scenarios. Compared to the standard RL solutions that learn a policy solely depending on the states or observations, GCRL additionally requires the agent to make decisions according to different goals. In this survey, we provide a comprehensive overview of the challenges and algorithms for GCRL. Firstly, we answer what the basic problems are studied in this field. Then, we explain how goals are represented and present how existing solutions are designed from different points of view. Finally, we make the conclusion and discuss potential future prospects that recent researches focus on.
Forward citations
Cited by 9 Pith papers
-
DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention
State-conditioning an offline goal-conditioned agent's goal embedding via a near-identity gated residual improves navigation (Dual 28 to 82% on AntMaze-large), and the gain is carried by the gate, not the difference-a...
-
Equivariant Goal Conditioned Contrastive Reinforcement Learning
Equivariant Contrastive RL imposes C8 rotation symmetry on the critic and actor, improving sample efficiency and goal generalization in simulated manipulation.
-
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
OIR relabels failed trajectories via an LLM into open-ended instructions and trains a unified instruction-following policy, outperforming PQN and ELLM on Craftax.
-
Reachability Weighted Offline Goal-conditioned Resampling
RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.
-
A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer
One diffusion policy trained via energy-guided RL solves multi-shape block pushing without demos and transfers zero-shot to real robots under varied conditions.
-
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.
-
Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects
A model-based RL policy trained entirely in simulation, paired with an online visual servo, controls DLO shapes in 2D with ~2 mm real-world error and 30/30 success, including opposite-curvature deformations.
-
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
Bourbaki (7B), an MCTS-based system with self-generated subgoal rewards, solves 26/658 PutnamBench problems, beating the prior 7B best of 23 at a larger sample budget.
-
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning
VSC-RL combines VLM-generated subgoals with a subgoal-conditioned AWR-style RL objective and claims improved sample efficiency over DigiRL and WebRL on AitW and WebArena-Lite.
Discussion (0). Continue with ORCID to comment.