Pith. sign in

REVIEW 11 cited by

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06718 v3 pith:LJYR2RAN submitted 2022-10-13 cs.LG

classification cs.LG
keywords hybridofflineonlinealgorithmdatasetefficienthy-qiteration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider a hybrid reinforcement learning setting (Hybrid RL), in which an agent has access to an offline dataset and the ability to collect experience via real-world online interaction. The framework mitigates the challenges that arise in both pure offline and online RL settings, allowing for the design of simple and highly effective algorithms, in both theory and practice. We demonstrate these advantages by adapting the classical Q learning/iteration algorithm to the hybrid setting, which we call Hybrid Q-Learning or Hy-Q. In our theoretical results, we prove that the algorithm is both computationally and statistically efficient whenever the offline dataset supports a high-quality policy and the environment has bounded bilinear rank. Notably, we require no assumptions on the coverage provided by the initial distribution, in contrast with guarantees for policy gradient/iteration methods. In our experimental results, we show that Hy-Q with neural network function approximation outperforms state-of-the-art online, offline, and hybrid RL baselines on challenging benchmarks, including Montezuma's Revenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage

    cs.LG 2026-02 conditional novelty 8.0 of 10

    Q*-realizability plus Bellman completeness is insufficient for sample-efficient offline RL under partial coverage, and a new decision-estimation framework recovers and improves existing bounds.

  2. A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

    cs.LG 2026-07 reject novelty 6.0 of 10

    A tabular hybrid RL framework with shifted transition dynamics, claiming near-optimal regret and suboptimality bounds under a known per-state-action bias bound.

  3. Multi-Turn On-Policy Distillation with Prefix Replay

    cs.LG 2026-07 conditional novelty 6.0 of 10

    ReOPD offline-distills multi-turn agentic LLMs via teacher-prefix replay plus step-decay sampling, matching online OPD accuracy at ≥4× speed with zero tool calls.

  4. Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach

    stat.ML 2025-10 conditional novelty 6.0 of 10

    A two-stage RL framework learns value-function envelopes from offline data and uses them to shape online exploration, yielding regret bounds that improve as offline data grows.

  5. The Three Regimes of Offline-to-Online Reinforcement Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Offline-to-online fine-tuning works best when the chosen method preserves the stronger of the pretrained policy or the offline dataset; a three-regime taxonomy organizes these choices.

  6. Decentralized Relaxed Smooth Optimization with Gradient Descent Methods

    math.OC 2025-08 unverdicted novelty 6.0 of 10

    A decentralized gradient descent method with adaptive clipping is claimed to reach best-known convergence rates for convex and nonconvex problems under (L0,L1)-smoothness without knowing the constants.

  7. Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.

  8. Online Pre-Training for Offline-to-Online Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A new 'online pre-training' phase trains a second value function that is then blended with the offline one during fine-tuning, improving offline-to-online RL across D4RL benchmarks.

  9. Reinforcement Learning via Implicit Imitation Guidance

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A reinforcement learning method that learns a state-dependent covariance from expert-policy action differences and uses it as exploration noise, improving sample efficiency on sparse-reward continuous control tasks.

  10. Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SquareχPO, a square-loss variant of χPO, achieves optimal 1/sqrt(n) suboptimality under label privacy and Huber corruption for offline direct alignment with general function classes.

  11. Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only

    cs.LG 2025-05 conditional novelty 6.0 of 10

    PORL fine-tunes a pre-trained policy online by first collecting epsilon-greedy interaction data and learning a fresh Q-function from scratch, removing the need for pre-trained critics or offline data.

Pith tools