Pith. sign in

Paper Citation Record · LEDGER

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2210.06718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.06718 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:41.065666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:12.306741Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ee7bc27-b8b3-44e2-be21-0238b4b01c53 · inbound

From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment cites this paper.

From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:41.065666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:27:41.065666Z digest=sha256:ceb0a9b77f54458a6c49e60345df72009e29a5c032baee1305c1df422fe9d7b1

Observation 29cab017-bd44-464d-b574-9ca054a70db0 · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:45.725844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:45.725844Z digest=sha256:f25392a3d2b300bb061c9e474e1228b8fde8dd4383cc0c71aae6f34551b92ab7

Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.439564Z digest=sha256:74e3b40a3343a57ec82fe350414ed79c5b346581491b3920f4b179e0bf74d9f1

Observation 0b408ba0-19b1-410c-bd43-e0e121d4bbbf · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.383535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.383535Z digest=sha256:6a46e792c7a72972de9118c7033e3b26ea3b6a3cb49ba0d8128ed42eeb17c1b6

Observation 3e703aba-876f-44e6-abea-a1f2da435afd · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.259804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:3433a30220489b46697bf080d839d6ab2804615bd369de5eb3f1c7918e341521

Observation 767f32c3-af73-45db-9a21-063a76a9d3bb · inbound

Online Pre-Training for Offline-to-Online Reinforcement Learning cites this paper.

Online Pre-Training for Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:02.465774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:02.465774Z digest=sha256:e80495552f8b21afbe9abe2a94eab7661013651085ca5939c7c47ac004b11ba2

Observation 526b2c15-4c8c-4ea7-8dee-d7cd56523cfc · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:59.023990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:59.023990Z digest=sha256:56c0e0bb039d9b16450b7a6acf0a946b0cd35c2d5bbc81d47cbacf178d27eae2

Observation 54d71233-b41f-407d-9024-bd9fa7794d14 · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:30.243849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:30.243849Z digest=sha256:e780d45618cdab3d0c34bb22c36f3695e4dd45dc371107de77875c16e3ae56db

Observation 61191bb3-ed59-4c20-bad3-3192f7dd2288 · inbound

The Three Regimes of Offline-to-Online Reinforcement Learning cites this paper.

The Three Regimes of Offline-to-Online Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:08.709075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:08.709075Z digest=sha256:bdc3922cad7a94e4658602e280bb0b07935143ee59a75b8b3459f241b2928076

Observation f2dc4c0c-1335-4d2d-9833-b78d661542cd · inbound

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach cites this paper.

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:45:50.812731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:45:50.812731Z digest=sha256:217da45c24f901477ad7b2cbdcb021b6cd41cd4aaeb30e458fbe365bb0d0cc35

Observation bfaca071-6b06-4792-aa97-c858755680ae · inbound

On the Sample Complexity of Differentially Private Policy Optimization cites this paper.

On the Sample Complexity of Differentially Private Policy Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:45:54.791452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:43:58.646419Z digest=sha256:3cf5e35f2d823667ed25fd79ed4acf777f62ba95805fdb8c580f2486150c492c

Observation 5fe5774a-48df-4fd5-a21a-3bdccbd329b0 · inbound

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage cites this paper.

On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:05:19.691967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:05:19.691967Z digest=sha256:ef32eff4234fa786a5ff4f8cd0a77b74e0fb16f75854a4cd8297e7508a13a795

Observation 0b234c6c-1aac-4e6c-8ffb-f70253b147a7 · inbound

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models cites this paper.

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:02.892690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T16:46:30.674244Z digest=sha256:837fd0ea2c47a673f763fa4ed41541cf6627d2c31cf0fdbb09c2a5c2c644214b

Observation 64a95906-4a9a-49a6-9fd2-4af53865b724 · inbound

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning cites this paper.

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.307546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:18:05.636131Z digest=sha256:d62de77c1cf5bca45118eb7a7718a7e471f9d15a0f1db3099c36840f067bcf78

Observation 2de8b1eb-8296-4ad6-b7c6-3fdafcb22eda · inbound

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation cites this paper.

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:24.163245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:55:12.531822Z digest=sha256:63a1c539ebffdbab9a16a775df2cdf9ea7787db948152577788a60a653a4ca70

Observation 975e4bed-bd57-4f19-abf5-fb3ea33ed49f · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.640676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:a1b3987788b719726f0d955ce342db697bcab80aac8553f50063cd524c7ece43

Observation 49d113e4-ba79-4d98-8959-46d1516a6769 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:12.446351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:2327a8a8f56a2109ad1fe1da3370153da35d752fae33f3091a27b1bcf1fc72e6

Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.342064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:0975d6e195d3c248d4225d23237c73f5957d3d70e46a808cfc7dff90c87e9438

Observation 0b2b7717-b975-48fe-9716-aed5571caef9 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:08.508849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:54:30.895137Z digest=sha256:43f01d3d922c7b6e798fac6840201528c245599b30cada58d61974316303e972

Observation fdc32471-d64f-439d-bd74-0d8826c39e84 · inbound

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data cites this paper.

SOPE: Stabilizing Off-Policy Evaluation for Online RL with Prior Data Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:19:56.740644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T09:15:11.280343Z digest=sha256:7a37ca6285673f509898ae3c45a770ce81a462c0c258e1f69c06814215ed35a4

Observation 796d031f-96a3-44c4-a061-5a62bb85114e · inbound

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift cites this paper.

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.988012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:41:20.147645Z digest=sha256:f815639deb4a281c464e929c64786d96c6734a661d5268f04409f3b91dbd7f6e

Observation fe4e1566-4459-4778-935c-a9e6bb1344db · inbound

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift cites this paper.

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:09:45.449771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:07:34.344964Z digest=sha256:c9166d9ce9bf915d3183c8e11d2cfe5a2a78ca5e0dbb44028a9ecde52ba80b56

Observation fb0c0acb-c715-42a3-90ff-ba797ec27a00 · inbound

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization cites this paper.

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:33:27.192556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:32:42.972836Z digest=sha256:58a987110264546133f388d47677b958516e5cf8c6020c845ba0c0ce203f9662

Observation acb0340c-2a99-43bc-9711-85dba150d4d4 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.850150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:1fc39ef70f7169a91ac4c8f6ec70084f5f07b751c8acd384742b8a8c0cb127e7

Observation 634974ff-59c9-40ec-a2f8-321d59b8e1f4 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.031118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:e3fc191ddfcf9c13e074765ad608e0c594c796009c1781b0a7bb586938062c15

Observation c40f1fc8-4888-48af-ab4b-b102ebdd2e46 · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.983617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:84ec38272d8b048e7aadf1e5ea0b4aa7d72cbbd1ebc9ad4598ab366fda0b542c

Observation 87b49c29-d6ad-4e76-8c37-c3c35c24bd8a · inbound

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation cites this paper.

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.308569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T19:29:00.117285Z digest=sha256:a8468938e293fd7d8756ddde49bd8e78565a03b2b8b1e022359290ae4c928b06

Observation c84440eb-076d-4640-b8b7-a7d1d2dbf63f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 282

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:b90bf85c473cc7736fc2e003578a667f4e937929283e539feac6d59f40c46b87

Observation 8aa1f197-80c9-41b7-b6f9-9f1cb3891903 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 283

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:05.419064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:05.419064Z digest=sha256:5bf100a1ce1743ae9aab15a1e6579314bfcc6a7ffbcaf0af69a38eeb81d142a3

Observation 3ae795ad-3718-4759-94ea-ce72fd92563d · inbound

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics cites this paper.

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:13:35.351631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:13:35.351631Z digest=sha256:bfcff6ca4b622a3cbfe00c6ee2bc133ef6f9e5303144691c10977861dbb0a728