Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning with Offline Data

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2302.02948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02948 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.188595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.181580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07b8cd2f-ab37-4319-9b16-b246748fc2b7 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.436566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:640b1f9f23949d2b34730d69a5215f05eed76c8f435f2bec88a8c4a788accf25

Observation da1fdad5-7c8d-4786-9ebf-e5c3fdaf1061 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own Efficient Online Reinforcement Learning with Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:44:02.559846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:a252d412549ae0175ee8c7e7823f0fad41d2d7c38f86605016da080246f60818

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:34633ec5717134600a9949099f742271b8ef68a505d83097c5fb90d6492df43d

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:fd25f0699d5a989b40c914a5b6f540f10426334638dad7d35c991a8bf919037a

Observation 19b8f177-1080-4532-812f-1a5351fb1d6d · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:28.644931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:28.644931Z digest=sha256:a068054e03bb9926fe4918594007265ffd041c486ed4f0d322907ae6334fca68

Observation 6af8edd1-46b1-4419-882f-c5d3db83d8a5 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning with Offline Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:26.920260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:26.920260Z digest=sha256:8345caaf837ffc4562545e391b9fd10632554ce44e2e14caa2930b8304b64f3c

Observation 8ea21e37-3b0e-49ab-9fd5-8a67e61b2199 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.613254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:04:41.766964Z digest=sha256:b7bf298ea09a61834b23d8add1535e1e6ad567e0e168ea6b618cbda1e167fcf6

Observation 7a7d4669-fe08-4290-ab25-458ba3ca73e6 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T21:40:44.030470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:40:44.030470Z digest=sha256:f166b61000560de2b45234cddf727ede413ec1944019b5f278c2a8597ce5a781

Observation 93d52c32-20b8-454a-bc34-a47bfcea627a · inbound

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation cites this paper.

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:06:52.898450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:06:52.898450Z digest=sha256:06e7a686793800fbffbf220e57e55559e02ea07351e625cefa1d15ddcf7f0f53

Observation 44b8beff-9de2-4cae-95ce-d09ce9fe70b4 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.108166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:84147086719eb9acfd536e731e963f0f9e774bc0f4765e130246f7bc70a65323

Observation f2bb1677-e9ad-4059-94aa-336f814d5aae · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.704833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:421d5c673fb18cdc721624067d4dfaa119c8d20b899f2e97471ee760b3bfad57

Observation e4ef87bb-5bf0-42bd-9790-d38f19d2060e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.164807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:b709c3b31e308d09cb99d1a5dd7fd82d2f8fe7aa1c2eda5b7a45f313075a4f8a

Observation 4302685f-03af-401c-bfd5-d9e21e29b59e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.785098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:e8f6359b4e404db28ab0217be3d96df6df1893442741648fe6df9ca1df67449c

Observation d19e3344-8788-4200-ae8b-6b7fb000eb31 · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Efficient Online Reinforcement Learning with Offline Data

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.803777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:058a51d0ca0b314bfce816e61cb864934d1d21a1eeedc1a84c38b9b1e27562c2

Observation 3122a81d-378e-485c-97e4-c1ebfff421a5 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Efficient Online Reinforcement Learning with Offline Data

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.514884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:31a131ac459c449209195a5b524040b4200cec3a29660d18d450234b0d7aaf91

Observation 3bf457a2-ff38-4369-8ac5-5166c4c1a1ca · inbound

An Introduction to Causal Reinforcement Learning cites this paper.

An Introduction to Causal Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.183066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T00:17:57.091481Z digest=sha256:5b18a7decfc3825909a37f09d114c6e84347c6f1d830a78390974b34043a1c5c

Observation 63993894-2958-492b-a60f-a338fef6ccec · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning with Offline Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:2e45edcb91ea850dfb004a5046d22691c4d8f204dc5c6d6dd70a850c670594d4

Observation b0587974-934a-47ec-a15b-e711c698d88d · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:21.826493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:21.826493Z digest=sha256:3560a884690d812e1d81a33906a90df75eb158fa5628afc8ebda2176329524db

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:8aab9e8bdfdca8b34a7e2df62c7071ce32821cdf33ea1a12f0cc7ea1bd109856