Pith. sign in

REVIEW 9 cited by

Goal-Conditioned Reinforcement Learning: Problems and Solutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.08299 v3 pith:Y2H7JSOL submitted 2022-01-20 cs.AI cs.LG

classification cs.AIcs.LG
keywords differentgcrlgoalsproblemssolutionsagentgoal-conditionedlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Goal-conditioned reinforcement learning (GCRL), related to a set of complex RL problems, trains an agent to achieve different goals under particular scenarios. Compared to the standard RL solutions that learn a policy solely depending on the states or observations, GCRL additionally requires the agent to make decisions according to different goals. In this survey, we provide a comprehensive overview of the challenges and algorithms for GCRL. Firstly, we answer what the basic problems are studied in this field. Then, we explain how goals are represented and present how existing solutions are designed from different points of view. Finally, we make the conclusion and discuss potential future prospects that recent researches focus on.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention

    cs.LG 2026-07 conditional novelty 6.0 of 10

    State-conditioning an offline goal-conditioned agent's goal embedding via a near-identity gated residual improves navigation (Dual 28 to 82% on AntMaze-large), and the gain is carried by the gate, not the difference-a...

  2. Equivariant Goal Conditioned Contrastive Reinforcement Learning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Equivariant Contrastive RL imposes C8 rotation symmetry on the critic and actor, improving sample efficiency and goal generalization in simulated manipulation.

  3. Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    OIR relabels failed trajectories via an LLM into open-ended instructions and trains a unified instruction-following policy, outperforming PQN and ELLM on Craftax.

  4. Reachability Weighted Offline Goal-conditioned Resampling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.

  5. A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer

    cs.RO 2026-07 conditional novelty 5.0 of 10

    One diffusion policy trained via energy-guided RL solves multi-shape block pushing without demos and transfers zero-shot to real robots under varied conditions.

  6. FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

    cs.AI 2026-07 accept novelty 5.0 of 10

    FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.

  7. Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects

    cs.RO 2026-02 conditional novelty 5.0 of 10

    A model-based RL policy trained entirely in simulation, paired with an online visual servo, controls DLO shapes in 2D with ~2 mm real-world error and 30/30 success, including opposite-curvature deformations.

  8. Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Bourbaki (7B), an MCTS-based system with self-generated subgoal rewards, solves 26/658 PutnamBench problems, beating the prior 7B best of 23 at a larger sample budget.

  9. Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning

    cs.LG 2025-02 reject novelty 4.0 of 10

    VSC-RL combines VLM-generated subgoals with a subgoal-conditioned AWR-style RL objective and claims improved sample efficiency over DigiRL and WebRL on AitW and WebArena-Lite.

Pith tools