Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning with Offline Data

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2302.02948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.02948 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:45.188595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.181580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07b8cd2f-ab37-4319-9b16-b246748fc2b7 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:48:36.436566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:c3af086843808b39e938d1c1a6f4db911413badc3dd5472cc154c4937f8fde21

Observation da1fdad5-7c8d-4786-9ebf-e5c3fdaf1061 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own Efficient Online Reinforcement Learning with Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:44:02.559846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:44bc91dfaad6f2a46c38f114307b6ef3149e5661a1b85bf81b6fed5e74f347f5

Observation 9ca01919-e7c7-4a62-9b55-695b0f199b9b · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.188595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.188595Z digest=sha256:34633ec5717134600a9949099f742271b8ef68a505d83097c5fb90d6492df43d

Observation 54098eb7-f739-4b53-b9d7-90eaa0d274cc · inbound

Reinforcement Learning via Implicit Imitation Guidance cites this paper.

Reinforcement Learning via Implicit Imitation Guidance Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:46.315134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:46.315134Z digest=sha256:fd25f0699d5a989b40c914a5b6f540f10426334638dad7d35c991a8bf919037a

Observation 19b8f177-1080-4532-812f-1a5351fb1d6d · inbound

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods cites this paper.

Decentralized Relaxed Smooth Optimization with Gradient Descent Methods Efficient Online Reinforcement Learning with Offline Data

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:34:28.644931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:34:28.644931Z digest=sha256:444581a299b58aa89db9c4f6114461c23bfb0518650cef73750eba260f0a74d8

Observation 6af8edd1-46b1-4419-882f-c5d3db83d8a5 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning with Offline Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:26.920260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:26.920260Z digest=sha256:8345caaf837ffc4562545e391b9fd10632554ce44e2e14caa2930b8304b64f3c

Observation 8ea21e37-3b0e-49ab-9fd5-8a67e61b2199 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.613254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:04:41.766964Z digest=sha256:33f3d15d1706244d2e7fe70919ae290cd5676023c45f96f7adb2c08e510dad25

Observation 7a7d4669-fe08-4290-ab25-458ba3ca73e6 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T21:40:44.030470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:40:44.030470Z digest=sha256:f166b61000560de2b45234cddf727ede413ec1944019b5f278c2a8597ce5a781

Observation 93d52c32-20b8-454a-bc34-a47bfcea627a · inbound

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation cites this paper.

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation Efficient Online Reinforcement Learning with Offline Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:06:52.898450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:06:52.898450Z digest=sha256:06e7a686793800fbffbf220e57e55559e02ea07351e625cefa1d15ddcf7f0f53

Observation 44b8beff-9de2-4cae-95ce-d09ce9fe70b4 · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.108166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:30deffff3ca4e30f2f2a3341c0e4fecd8ec7c21b5b739f72bd4289a69c4251eb

Observation f2bb1677-e9ad-4059-94aa-336f814d5aae · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Efficient Online Reinforcement Learning with Offline Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.704833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:4022d2c28fe7d53cd7e2b9bb5a79db3d5fe3e4acfdf373959538552525453abf

Observation e4ef87bb-5bf0-42bd-9790-d38f19d2060e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:08.164807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:32:16.746824Z digest=sha256:932a08b5dcc68e455ca0ecb1b0d75b55a3c71065a387267dec4582ff27c9c3ee

Observation 4302685f-03af-401c-bfd5-d9e21e29b59e · inbound

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking cites this paper.

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking Efficient Online Reinforcement Learning with Offline Data

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.785098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:53:29.468764Z digest=sha256:adc85819fc676b5bfb5bd52877d3b32ba50a3f9c61dbd4948dc18685ecfdaba2

Observation d19e3344-8788-4200-ae8b-6b7fb000eb31 · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Efficient Online Reinforcement Learning with Offline Data

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.803777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:e1eb0468c596efcb5723d297ee696b040fc8beb464901fccbb652899acab50c7

Observation 3122a81d-378e-485c-97e4-c1ebfff421a5 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Efficient Online Reinforcement Learning with Offline Data

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.514884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:7ef328a5a4eb5bda3f70d16d4c428a1c5dac7c6ff73a6b828d213aad358242b3

Observation 3bf457a2-ff38-4369-8ac5-5166c4c1a1ca · inbound

An Introduction to Causal Reinforcement Learning cites this paper.

An Introduction to Causal Reinforcement Learning Efficient Online Reinforcement Learning with Offline Data

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.183066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:17:57.091481Z digest=sha256:1e39d954612fff89f76e8bd5ee4c40db52b39e673ee8d0256c1a3309f1ea6c17

Observation 63993894-2958-492b-a60f-a338fef6ccec · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning with Offline Data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:2e45edcb91ea850dfb004a5046d22691c4d8f204dc5c6d6dd70a850c670594d4

Observation b0587974-934a-47ec-a15b-e711c698d88d · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:21.826493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:21.826493Z digest=sha256:40bad374ba9209a0b3fa339ce797fe70d6dd56e1870f80c3b524b97e7136cf71

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:be2f5e57e69b834c0f756c106a74fdc3b7972ea07b129533ce300d85cd25fc68