Pith. sign in

Paper Citation Record · LEDGER

Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2412.07762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07762 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:20:33.609824Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:45.826205Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2193e37-82f2-4ada-8ea5-42471c41e9db · inbound

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy cites this paper.

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:20:33.609824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:20:33.609824Z digest=sha256:bb985774aeea58aba7c197452a550ab3bb477c3ffdd80eb4f72e54341d9f41fb

Observation 526a60c2-719a-4d3d-a26a-1709df4352be · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.275862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.275862Z digest=sha256:0f9a982b5c1e544ad932ca64a6fe25ea61150bdbda938d6617652a2bc906e2bf

Observation 0b282f09-1121-4413-bd59-10267b12cb72 · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:56.493723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:56.493723Z digest=sha256:090a22be770f4afb43a3fe6dda54b480c6acb7d06beb349cdf4286e6a9d032f9

Observation ba3f8c42-beed-4bc0-8510-eae4208a51e4 · inbound

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training cites this paper.

SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:06.334933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:53:06.334933Z digest=sha256:8a764a73393d953e59214061363d6296cd968230167c0a255b14feba5a475cf3

Observation 010eb31b-5ee2-44fc-adbd-f9db4dea1fe8 · inbound

Reinforcement Learning with Action Chunking cites this paper.

Reinforcement Learning with Action Chunking Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:22:06.470744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:18:20.960945Z digest=sha256:23952673a84c12ca809b4758a75e91cf08af9398952017a189d280084177c04d

Observation 960fc39a-6c78-45a9-8cd2-88d8b797036e · inbound

The Three Regimes of Offline-to-Online Reinforcement Learning cites this paper.

The Three Regimes of Offline-to-Online Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:10.173421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:10.173421Z digest=sha256:ab2daee4720a65dc4d944f533d47aa4689e28ca8dbfe84453e5d4dde2f109776

Observation 38c6f8fb-2a80-4507-9e75-29f66546dac2 · inbound

Value Flows cites this paper.

Value Flows Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T11:01:35.297180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:01:35.297180Z digest=sha256:02e86164c6e78b35538ece9b41ee6eacd76e2ba0ed1487bdf71fc0c5009a99ed

Observation 9e952ef1-391a-48c6-972b-c985f4848809 · inbound

SpikeATac: A Multimodal Tactile Finger with Taxelized Dynamic Sensing for Dexterous Manipulation cites this paper.

SpikeATac: A Multimodal Tactile Finger with Taxelized Dynamic Sensing for Dexterous Manipulation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:51.446369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:51.446369Z digest=sha256:cbf35fcac6b285588603ec1b82a50afa25abe04878307af043374271846a509d

Observation 80f8f9bf-10fd-4a81-ba9f-f5588dec47e3 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:49:58.825253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T11:48:13.368844Z digest=sha256:421ac65e79cc89eb7c50f324a33f0067b88c11ad5880d22b7e7eb26d54994c6e

Observation 3f1340c9-6380-40cb-8b8e-a34b9a9ae262 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:55:04.124155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:54:57.866685Z digest=sha256:b65280c2b1d988af2a1c7c1b97d4ef453c63aaa27d934cda19f7d82462f4c750

Observation 9903d06e-5fa8-4ab2-8ab6-764ca48aae88 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:49:54.444164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:46fdb37aa87f5e1e14dfc6cb805f18143b7dc208ade66621bf1b8d104b50b7fa

Observation f3a14684-4401-4cdf-a74f-3785b27866ff · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:55.143469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:1fa44745882da701097916323cdf70a366aa92420303ad559dde05dda5ee6209

Observation b39afcf6-a0d7-482b-9729-1e050cb9cd33 · inbound

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning cites this paper.

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.325903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:18:05.636131Z digest=sha256:8a9ad39e74f8b6764a673e72a4e0581a2bf2599a84929b6f56341ff2c7b2c2b4

Observation 2b6c922f-9247-4bcc-8025-1453d1d53ec8 · inbound

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation cites this paper.

Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:24.122109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:55:12.531822Z digest=sha256:56001739421dec848c4d8c237746656cbf98bb03537f1b92944737fda1ce1fc8

Observation 3e7e6bcc-1572-4028-95e5-a15a9bf995d8 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.426634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:25aa2f4414d26cf591b84f40f781067146729001101806b3ec7e00cdcd81865a

Observation 5b4de9fb-0cae-4a8d-a72f-9485efdf5270 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.633737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:763b22dd3fe05fbd9322723da9c730b76bc7153af4fd543eb2471c428d2c7f91

Observation fdd7bff5-439e-4469-9192-caf94b27fbe2 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.827910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:4d29315e8d1d7bccdd3d618420e30c29a0a7215727db65495bd3e9d50d8bd0d4

Observation 9c94a5ab-9180-4f85-b548-bbb8d6e8912d · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.054667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:72f0ccd46aa84b8d717cc61f9a74e9757e51379dbe6ed49fa2b3014805cf7ffc

Observation cf6d3ac2-d0b0-43a4-b0d5-d4ae7b410429 · inbound

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies cites this paper.

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T00:23:04.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:23:04.763773Z digest=sha256:7e3577e59b076ed50b51442ff740ee1fb67a12075af28ec78658a983bee51996

Observation ebdd9374-8d1a-4aef-bda1-94e5980ee6db · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:24.369491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:24.369491Z digest=sha256:8cfa6aaadbc562e9a40436941378930b43e62d116a3954531be43fd76ba23e61

Observation dce039d1-cbdd-4a12-8533-f5f4580e6e34 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.958427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.958427Z digest=sha256:1b815f86821c39da0d15826f65fd3175f68a1ed9a4c73f5d00170d9beee16a09