Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2210.06718.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:45.725844Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:50:12.306741Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 29cab017-bd44-464d-b574-9ca054a70db0 · inbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · inbound
Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e703aba-876f-44e6-abea-a1f2da435afd · inbound
EXPO: Stable Reinforcement Learning with Expressive Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 767f32c3-af73-45db-9a21-063a76a9d3bb · inbound
Online Pre-Training for Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526b2c15-4c8c-4ea7-8dee-d7cd56523cfc · inbound
Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 162
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d71233-b41f-407d-9024-bd9fa7794d14 · inbound
Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61191bb3-ed59-4c20-bad3-3192f7dd2288 · inbound
The Three Regimes of Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2dc4c0c-1335-4d2d-9833-b78d661542cd · inbound
Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfaca071-6b06-4792-aa97-c858755680ae · inbound
On the Sample Complexity of Differentially Private Policy Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5fe5774a-48df-4fd5-a21a-3bdccbd329b0 · inbound
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b234c6c-1aac-4e6c-8ffb-f70253b147a7 · inbound
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64a95906-4a9a-49a6-9fd2-4af53865b724 · inbound
WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2de8b1eb-8296-4ad6-b7c6-3fdafcb22eda · inbound
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 975e4bed-bd57-4f19-abf5-fb3ea33ed49f · inbound
Fisher Decorator: Refining Flow Policy via a Local Transport Map Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49d113e4-ba79-4d98-8959-46d1516a6769 · inbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · inbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b2b7717-b975-48fe-9716-aed5571caef9 · inbound
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdc32471-d64f-439d-bd74-0d8826c39e84 · inbound
SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 796d031f-96a3-44c4-a061-5a62bb85114e · inbound
Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe4e1566-4459-4778-935c-a9e6bb1344db · inbound
Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb0c0acb-c715-42a3-90ff-ba797ec27a00 · inbound
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation acb0340c-2a99-43bc-9711-85dba150d4d4 · inbound
Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 634974ff-59c9-40ec-a2f8-321d59b8e1f4 · inbound
COOPO: Cyclic Offline-Online Policy Optimization Algorithm Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c40f1fc8-4888-48af-ab4b-b102ebdd2e46 · inbound
When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87b49c29-d6ad-4e76-8c37-c3c35c24bd8a · inbound
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c84440eb-076d-4640-b8b7-a7d1d2dbf63f · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 282
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa1f197-80c9-41b7-b6f9-9f1cb3891903 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 283
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae795ad-3718-4759-94ea-ce72fd92563d · inbound
A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.