Pith. sign in

REVIEW 13 cited by

OGBench: Benchmarking Offline Goal-Conditioned RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20092 v2 pith:57I7V4MY submitted 2024-10-26 cs.LG cs.AI

classification cs.LGcs.AI
keywords algorithmsofflineogbenchcapabilitiesgcrlgoal-conditionedbenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Offline goal-conditioned reinforcement learning (GCRL) is a major problem in reinforcement learning (RL) because it provides a simple, unsupervised, and domain-agnostic way to acquire diverse behaviors and representations from unlabeled data without rewards. Despite the importance of this setting, we lack a standard benchmark that can systematically evaluate the capabilities of offline GCRL algorithms. In this work, we propose OGBench, a new, high-quality benchmark for algorithms research in offline goal-conditioned RL. OGBench consists of 8 types of environments, 85 datasets, and reference implementations of 6 representative offline GCRL algorithms. We have designed these challenging and realistic environments and datasets to directly probe different capabilities of algorithms, such as stitching, long-horizon reasoning, and the ability to handle high-dimensional inputs and stochasticity. While representative algorithms may rank similarly on prior benchmarks, our experiments reveal stark strengths and weaknesses in these different capabilities, providing a strong foundation for building new algorithms. Project page: https://seohong.me/projects/ogbench

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A single ~21M JEPA checkpoint trained with Brownian-bridge state flow and edge-aligned action-state noise sampling serves planning, behaviour cloning, and inverse dynamics without retraining.

  2. Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Bilinear contrastive critics remain good compatibility rankers but are unsafe to maximize for action selection; cosine bounding does not fix value decalibration, while Bellman TD-Q does.

  3. Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

    cs.LG 2026-07 conditional novelty 6.0 of 10

    On six robot-manipulation tasks, offline Q-pretraining does not accelerate online RL fine-tuning from a pretrained policy, while seeding the replay buffer with rollouts from an ensemble of policies (IPE) improves fina...

  4. Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Minimum predicted cost selection fails under proposal overgeneration; reconstructing actions from adjacent low-cost prefixes (ASAR) raises Cube carry-and-release success by ~19–28 points.

  5. Patch Policy: Efficient Embodied Control via Dense Visual Representations

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Patch Policy shows that frozen dense ViT patch tokens, consumed through a block-causal attention mask, let lightweight robot policies beat pooled-feature policies and even a fine-tuned 7B vision-language-action model.

  6. SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows

    cs.RO 2026-02 unverdicted novelty 6.0 of 10

    SERNF fine-tunes dexterous manipulation policies on real hardware by pairing normalizing-flow policies with action-chunked critics and conservative off-policy RL.

  7. Compositional Diffusion with Guided Search for Long-Horizon Planning

    cs.RO 2025-12 conditional novelty 6.0 of 10

    CDGS adds population-based search and likelihood-based pruning to compositional diffusion, enabling long-horizon planning from short-horizon models across robot manipulation, panoramas, and video.

  8. Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks

    cs.RO 2025-07 conditional novelty 6.0 of 10

    The paper introduces MTBench, a GPU-accelerated benchmark for massively parallel multi-task RL, and reports experiments suggesting on-policy methods outperform off-policy baselines while value learning limits MTRL per...

  9. VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

    cs.CL 2026-07 conditional novelty 5.0 of 10

    A two-level induction procedure—active-probe sketch selection plus multi-step rollout fitting—recovers executable code world models that improve CEM planning over prior code baselines on four LeWM tasks.

  10. Mollified Value Learning

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Mollified Value Learning regularizes offline goal-conditioned value estimates with a Feynman-Kac expectation version of the viscous HJB equation instead of a pointwise Eikonal constraint.

  11. Value Flows

    cs.LG 2025-10 reject novelty 5.0 of 10

    Value Flows fits the full return distribution in RL with a flow-matching critic and reweights its learning objective by estimated return variance; the central theoretical guarantee does not follow from the stated equations.

  12. Learning The Minimum Action Distance

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A state-embedding method learns asymmetric minimum-action distances from state-only trajectories and beats existing representation methods on tested environments.

  13. Information-Based Exploration via Random Features for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Random-feature Gaussian-process information gain is turned into a closed-form exploration bonus for PPO that matches RND/VIME/#Explo on 12 control, navigation, and sparse-locomotion tasks, with error bounds on the app...

Pith tools