Pith. sign in

Paper Citation Record · LEDGER

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

As of 21 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 31 inbound Pith citation observations for arXiv:2412.06685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06685 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:28:38.548957Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:23:38.206985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.649113Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3682650-6a4a-4d7c-a290-040be540b661 · outbound

This paper cites Abdolmaleki, J.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Abdolmaleki, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.762199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.777772Z digest=sha256:a501cba9e069a94086ee791f9f6d218ac9c1b7b6f46214dcf956af6824077f3e

Observation 57210eb1-0b26-4073-b9cd-de722e8ef3cd · outbound

This paper cites DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.783068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.783068Z digest=sha256:144bf4dcbc58c491e12cf158ade6782574322e5517c3b3ab05db40be1bc882b5

Observation 3756bff2-2cb0-4f3d-bbd5-ab83f380eccc · outbound

This paper cites Efficient online reinforcement learning with offline data.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Efficient online reinforcement learning with offline data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.788366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.788366Z digest=sha256:ff4f158c414e7d177c2cbdcbad42a8afca603ceb128d79c7d3ea303aa8436fd0

Observation c640ceba-ea63-478d-9769-dc38c584847c · outbound

This paper cites Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.793143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.793143Z digest=sha256:80e7fca15d0f5c7d89da3446d913f96d474f016fde863b475fae60671cb3ab10

Observation 85351a09-9f2b-4bdf-9146-2437a62c1183 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone FireAct: Toward Language Agent Fine-tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.798320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.798320Z digest=sha256:43be55520560f8c117ecddaaacec32c6c50ad36dc1ad5d135d976ac22777fd1b

Observation a9c44608-52ce-4d84-8189-988d0b3f0e39 · outbound

This paper cites Diffusionpolicy: Visuomotorpolicylearningviaactiondiffusion.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusionpolicy: Visuomotorpolicylearningviaactiondiffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.638278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.803294Z digest=sha256:d012a6045eb70135ab43ad8c75d59c0843e764e8ada0395f2762392a54f00608

Observation 7f289261-15e4-4165-a12b-3283755a8f5b · outbound

This paper cites Bridge data: Boosting generalization of robotic skills with cross-domain datasets.Robotics: Science and Systems, 2022.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridge data: Boosting generalization of robotic skills with cross-domain datasets.Robotics: Science and Systems, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.503358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.808438Z digest=sha256:1bef373b9ec9492cadb866f89ad8868b913043802791c88e7b0a92ed6aa6df94

Observation cea8ecc9-0983-428c-ba04-b08daa2027e2 · outbound

This paper cites Stop regressing: Training value 16 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone functions via classification for scalable deep rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Stop regressing: Training value 16 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone functions via classification for scalable deep rl

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.332367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.813070Z digest=sha256:5773c20524bb7ed1966712a15d12ef3be8dbdd028086585f688309fdc9232bc2

Observation fe6e76d3-abec-4f35-a97c-05ea505deb62 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.818160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.818160Z digest=sha256:886645211d5349c761f7532d8d2d2f458e28f5a29c6f3df73ff3fc00ad5f58d5

Observation 4b8af79e-74b6-4c28-8c47-cd9e726ab9b7 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone A minimalist approach to offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.822367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.822367Z digest=sha256:5242d01a81c593f8c575c20bbb1da40d1f0d93ab5f49f55a16cb5f1dd58a226f

Observation 86c6269a-6684-48ff-aac8-9c76364803f2 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Addressing function approximation error in actor-critic methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.827415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.827415Z digest=sha256:01dd9ea872c2a79a7716cc809192119ac4954305961d078c2aa94a43a2b93ed8

Observation f7603abf-cff8-4374-bcc1-f8de43aa4872 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.832515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.832515Z digest=sha256:6d87b1a175652f760a404169c6f2c0c13a2ada786e75174f01082050eab8aea3

Observation 33f97c0a-d888-424c-84bd-df3892676277 · outbound

This paper cites Emaq: Expected- max q-learning operator for simple yet effective offline and online rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Emaq: Expected- max q-learning operator for simple yet effective offline and online rl

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.285717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.836857Z digest=sha256:0d2e00298119cdb00e76d75d5d0c7a2839846277f0c276d60e6e6342b47fca6f

Observation dfeb17e4-84a0-4cd7-b97a-b32763da0934 · outbound

This paper cites Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.269847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.841410Z digest=sha256:340bb639c80ec945b4d62e876112dcaa7b125d59b6c2bcc28a85cc1e1ef3e041

Observation 9e049b36-5eb2-4ef9-8322-49db2d15f830 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.846301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.846301Z digest=sha256:27d150525bf3dcfe881b7c524af637a507824c00559c784b3cb9a76c84ce9c45

Observation fcd3e3b7-9c83-4ba5-b143-31e20c2cc21c · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.851080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.851080Z digest=sha256:68d56abf61753281e020cba042c6cf1dc7b40f4db1e2c213dfdc4c6367ce4823

Observation 1419ef9a-d790-446c-9adb-562c0f9be061 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.855741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.855741Z digest=sha256:53b7f0acf0123bfe63264978311314b57681504598516ebeffeedf3b8d7e5dbc

Observation 900932f6-9fa5-4bda-80de-f8db9fe56cad · outbound

This paper cites Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.896663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.896663Z digest=sha256:1897ff078ee6df531fbd7b35fa12d41e01e99770655e67953e63bb9de680acfa

Observation 8c8718a4-beaf-486d-9b04-86ee4a729b54 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:37.935059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:37.935059Z digest=sha256:12ac1106e3a3a90c62713eb6094d6ee2978ab46556c32f05073a6b5a47f32b71

Observation 35a0a684-f0c9-47c0-8fd4-84020e7182dc · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning as one big sequence modeling problem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.213089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:37.975759Z digest=sha256:38b6e81564193d2a05a8c2a967696232cea81b8362d22dd248762b6510d6670e

Observation 7ac1b492-06d7-4e66-9575-0adf6385fb92 · outbound

This paper cites Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.126016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.015147Z digest=sha256:384999f6ff18e20e546476cdb98fce73f5e554dd282b6879de6c15483c85d3ea

Observation 1f6e5670-a8d5-48f2-ac50-1c6e6ab01acf · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.085466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.085466Z digest=sha256:9bd8d36d2ddaaa50bdca8e92e412ed4e62c6ce99b671faba3c5898f7af0c62fc

Observation 393c17c9-5904-4e36-bf7a-a34d0973a03f · outbound

This paper cites Adam: A method for stochastic optimization.International Conference on Learning Representations (ICLR), 2015.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Adam: A method for stochastic optimization.International Conference on Learning Representations (ICLR), 2015

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.110467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.151870Z digest=sha256:3a652718b41c32f77edd8fb4df9cb140ff05192c595c02b803e631b0e751ac10

Observation 6e224e8b-02cc-4e5e-ac33-cc3e93572b68 · outbound

This paper cites Offline reinforcement learning with implicit q- learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning with implicit q- learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.095697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.156808Z digest=sha256:73adf5867b15c7d9b074eeda24a8993015579c0fecb171fb00b66fde50191fdf

Observation 79359524-3e4b-46fc-b2d5-2a8a0875a164 · outbound

This paper cites Kumar, X.B.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Kumar, X.B

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.077671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.162001Z digest=sha256:063042564c8b22dd1d76b8efb60d24efc7684e6ee468709e6170ddca2ef4a825

Observation 775538a6-0dc3-460d-9935-cbb9cf9f8c34 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33:1179–1191, 2020.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33:1179–1191, 2020

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.166784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.166784Z digest=sha256:23a597019464567a765afdf5a01fe853fdc5a2991b7dfc05eb46525efa9de5f4

Observation c8ea5083-aa4e-4af1-98d5-343aae6bfc45 · outbound

This paper cites Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.172066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.172066Z digest=sha256:effdfdb6d06734820061563fc538034c8bfc4173bb866959148ae71f0afe6d5f

Observation e3a74265-7d7d-4d62-b4b4-a1440a509201 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.176724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.176724Z digest=sha256:ca515c1df38c463b2bc0711797f1865c7ec75ecae2e796ddab233b58c84b2920

Observation b542df0b-9d74-405b-a4fb-5849c2da2e5d · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:40.025368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.181708Z digest=sha256:25afdcd9e74ba96fe2d18064a661b9d5166800bac9b84096711f30feafc14408

Observation 6be748c2-c87b-41ba-a477-ed2861044d62 · outbound

This paper cites Continuous control with deep reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Continuous control with deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.185750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.185750Z digest=sha256:cc9b0d56814a90553f9c5ed776385c20e00860d9c133f942065cd7c55150240d

Observation ccef55f1-8c98-40b7-8b32-6d2f4de27bbd · outbound

This paper cites Leveraging exploration in off-policy algorithms via normalizing flows.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Leveraging exploration in off-policy algorithms via normalizing flows

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.915108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.190192Z digest=sha256:ca25b55247d1eeb6f62fc6a97ccae9bb39fefe655131e17fd3a684eeaaede74a

Observation f33dd6af-30ba-42af-b754-8a1c475586e0 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.194784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.194784Z digest=sha256:d493dbf1a10727121bb42f681ab2ebefa632dd497537635a712ebdaf7974b5f1

Observation 81a73538-74a0-4250-8201-8018fa5efa2e · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.198810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.198810Z digest=sha256:8f08c626605f331c8d357b4587ef33be705c9a4aa8c56357878f86f46f8a9c75

Observation 8658d600-0fce-4e4f-9475-4523f84572db · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.Conference on Robot Learning (CoRL), 2024.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Steering your generalists: Improving robotic foundation models via value guidance.Conference on Robot Learning (CoRL), 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.827158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.203147Z digest=sha256:97cd63fd2dcacfcccc10c8e012602f25fa0031d1fe84ccd16ac6d9a6f3b298b3

Observation e0c1692d-7dc3-4467-b535-c87bef76b780 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.811700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.207403Z digest=sha256:394ccb17caf4579022057b9791fa86792b859d5e7c9f1ff9feebafcbc0bdc264

Observation d113f657-f8b1-4641-86b1-34531597c672 · outbound

This paper cites Greedy actor-critic: A new conditional cross-entropy method for policy improvement.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Greedy actor-critic: A new conditional cross-entropy method for policy improvement

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.795425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.212797Z digest=sha256:2f6933f4ff5e15fb50444e4e28f7cbc1e1e5482723044ab5627e15954cebcc3b

Observation b2e2af66-2abf-4477-8231-f1dd1d53af56 · outbound

This paper cites Self-imitation learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Self-imitation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.778973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.217614Z digest=sha256:02c19dc4f42cbe8457cdc718d2f496df604f5a2101e2140444ec0b1945d343b2

Observation 081fa1d1-61be-4a29-b247-bedcb98d486d · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.222122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.222122Z digest=sha256:51e4dc2eb49b88370b1c4b24f64e801ba73dbfaa317082ed76b8338168338e52

Observation 5bc20772-9a24-4b5c-b8bd-184f68796180 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.226671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.226671Z digest=sha256:88f114088f95047ec2c4e5a810ee042dcec3d8ad5d546e0f80714fa0dd4662d8

Observation 1d9f211b-d82d-4b83-b548-731c999f232d · outbound

This paper cites Peters and S.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Peters and S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.647040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.230916Z digest=sha256:1328953a9dd052ab6cdd15fff13e4cc0e95d795d08b9706073f74066fdbfe043

Observation a24d8794-b64c-4072-ac60-d338d34c7b5f · outbound

This paper cites Relative entropy policy search.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relative entropy policy search

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.566816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.235007Z digest=sha256:5dda74a53f30bebb903994bf19114da602d7f4fbde67429c3b6935b369d6fcac

Observation baeddc20-2273-4533-9e6f-f4778a91a6f2 · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning a diffusion model policy from rewards via q-score matching

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.550340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.239427Z digest=sha256:a8ce0a1cc6604d4eee264a809d5baa7537b2ee77f4742ec80ebbcb9b5f03a0d8

Observation 56fe718b-1dd6-47f2-9a18-cc015507edaf · outbound

This paper cites Diffusion Policy Policy Optimization.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion Policy Policy Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.244395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.244395Z digest=sha256:7b5a24265fbd234667f5fab1dd99f388ae249836fa5433dace2542ac3fb5c985

Observation dc0a1f4b-0a67-4d01-a0b4-770a9365f531 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.249232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.249232Z digest=sha256:45499b04336da315c9ccc80b50d43dc33513dbcb8c75c04b675ea418c30cc5ab

Observation 79a038d1-d0e9-46fc-adb9-92ee0820caa6 · outbound

This paper cites Grac: Self- guided and self-regularized actor-critic.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Grac: Self- guided and self-regularized actor-critic

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.326396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.257838Z digest=sha256:822c239bbc4af84e6ddaf580a66de07d1b5a4148a95bbfba4ed5875beeb9e7f6

Observation 221685e3-8468-4755-a020-05eddb470901 · outbound

This paper cites Skill-based model-based reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Skill-based model-based reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.302833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.332453Z digest=sha256:a192b98a03f0bb8048f733b34e10f4b16d1cf0fd8d473db0b31672b50957692c

Observation d14fd665-77f4-409e-8587-75aaf4caee95 · outbound

This paper cites Q-Learning for Continuous Actions with Cross-Entropy Guided Policies.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-Learning for Continuous Actions with Cross-Entropy Guided Policies

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.396922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.396922Z digest=sha256:45242d9cfa47f90485436b9194fcb6f2940e3f4f4981a67afc94eb8331ca0b76

Observation a22cd976-2b68-4d26-be2f-644f254f58bc · outbound

This paper cites Hybrid RL:UsingbothofflineandonlinedatacanmakeRLefficient.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Hybrid RL:UsingbothofflineandonlinedatacanmakeRLefficient

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.284548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.459011Z digest=sha256:bea753ada056fd71c375980232f690709bb63c6d5bc311ff86b1ca6dcbd7f634

Observation 24994787-2415-4d81-a67c-61527fce5091 · outbound

This paper cites Second edition, 2018.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Second edition, 2018

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.501183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.501183Z digest=sha256:3bbb9c73c34f41757ff6e853926f44275db8ad3da6d993675c314007d72b4f7c

Observation b74b0c52-2c29-4f23-813d-c70d0a960b5d · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.256265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.505556Z digest=sha256:f73c9cb209e6f7aba7e62160e7710e4a1a05cac0100ab385f45e98d6f068b7e5

Observation ddd37579-adeb-407f-a972-a93d3dffe64c · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridgedata v2: A dataset for robot learning at scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.510594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.510594Z digest=sha256:785369dd6f7c3fda013950ca5f3dd92a88c54ca5e1319d2e02dfc1dcbc3302f4

Observation 0b474557-84da-40f9-a582-bb79fa41abf9 · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.515332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.515332Z digest=sha256:e4da1c0930603975ad43e0dd25356402804c82ac7b6b8513e98114ca7cebbe8e

Observation 61a61b21-4098-445b-a622-ebf21a549ae2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.519972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.519972Z digest=sha256:a123f86278fd6fc60c37e41a3a1e62af9e3397d968518842e6fa9823ffbccc77

Observation 5ed0b98b-a5f9-42a4-a95c-3ce9e28b4a5f · outbound

This paper cites V-former: Offline RL with temporally-extended actions, 2024.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone V-former: Offline RL with temporally-extended actions, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:39.139198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.524739Z digest=sha256:c071137992c95845c55b1d86fd7e9d5537bb7b6b715b8e40d12300785e17fb03

Observation 3a02aeaf-8436-4d36-8fe7-741fe74598f3 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.974717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.529793Z digest=sha256:37031c0e9ad0adcc160bb0d2f63391d8276e65a292c58e1ce70253e917429528

Observation 88d1e079-8e37-4248-a4c5-d907d1663a40 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T19:28:38.534595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:28:38.534595Z digest=sha256:43f1eee205a40f3c37cba1cda3094323bdd122dad1ce92d4efc5f5b74aae78e5

Observation 98f2ede7-b571-46b7-9152-e41df897f2bd · outbound

This paper cites Mastering visual continuous control: Improved data-augmented reinforcement learning.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Mastering visual continuous control: Improved data-augmented reinforcement learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.904003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.539705Z digest=sha256:80b39f08bdc0417ae35bb955d8fcfe6693151b69f33973758e7e48e5ffe38071

Observation 739b6e11-9382-4951-b6b5-b9243dde2068 · outbound

This paper cites Autonomous improvement of instruction following skills via foundation models.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Autonomous improvement of instruction following skills via foundation models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.887226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.544436Z digest=sha256:86628c3c14f97ffecf45a0b8ea25f8191f8cf26d7f1a3a0264cdf12eb83fc601

Observation d7aaaa09-ce8a-477e-8bb3-68901f58d8f2 · outbound

This paper cites -v0” antmaze datasets from D4RL, but Fu et al.[9] deprecated the “-v0.

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone -v0” antmaze datasets from D4RL, but Fu et al.[9] deprecated the “-v0

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:28:38.870057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:28:38.548957Z digest=sha256:b25a721389fa497c4e15635e174a19ab0547ee4d4dead5be045545041a87ffe5

Pith citing papers

Observation 8c7411f3-fee2-422d-9eaf-f55ce5b57aa2 · inbound

Flow Q-Learning cites this paper.

Flow Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:10.467494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:54:10.467494Z digest=sha256:9425f77d896b7d58800011d007c386ba6a24b379375a541879995cb9d1845782

Observation 130da234-bcdd-4f93-8ff5-e34a16e1b8b0 · inbound

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy cites this paper.

ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:20:33.596522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:20:33.596522Z digest=sha256:1e9e53d62407aeb340d7e284078e0618c74e1674ed26f8347d0fac50c93fdcc6

Observation 4c5de78b-2d75-439a-83d5-9af270e71cc1 · inbound

Exploratory Diffusion Model for Unsupervised Reinforcement Learning cites this paper.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.339718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.339718Z digest=sha256:dc1d23f0e79f7d641e19b6d092b03952a4e707cde38fa8ae38dc04d94d851aaf

Observation 0ad9fb00-7442-47ff-b970-fb7420bd510e · inbound

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning cites this paper.

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:38.206985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:23:38.206985Z digest=sha256:6c13369ff61820df6814dcfc00c569c9d296eb95f4f1d5f07906715599525f7f

Observation a5446eea-762d-424c-bc91-e942572d7a78 · inbound

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners cites this paper.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.612205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.612205Z digest=sha256:d7f0235eb338d7b6b898f8cfc6345189257ce84e7d139f2e536befceff7f3650

Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · inbound

Diffusion Guidance Is a Controllable Policy Improvement Operator cites this paper.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.792421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.792421Z digest=sha256:e85d3c699820035a16285ddb34b84d7381c2763cf94df418d728306d78a5bee7

Observation 4335a470-040a-42bf-9ca8-fbf1a0758d2f · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.258768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:575232909ff853851c5e8d54f1427f0641d06252f6c2b69b4d886a0f1a3aa7c8

Observation cdd8c86c-0b3b-4f1a-af6f-f9b4914be986 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.611023Z digest=sha256:9738d5889730504e193cdf65b63be011ceb26238c49393ea5cbca8cc34a7749c

Observation 609b4a33-52c6-4b10-aecf-d739fcbb9438 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.359504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:2d76302ef19f5e8f9d1d45b2291cc95f8556265c3c89d4461f1b8f4405554e3d

Observation 88261a3b-d278-416d-8adc-03952f33d58e · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:26.050159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:b7e1de65bfd8c3d0f616cb59c33ed021bceae0b4407aaedbc67b4f0e8a8b92bd

Observation 6e7b3928-7f87-4e28-8f30-79961d2b7c67 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:49:58.836520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T11:48:13.368844Z digest=sha256:ffb02ae94ecabc0ecd5e6735ba1fe960ebb295507e75875210c4ed846692b2dd

Observation 12e3cd77-d007-4137-8d44-38141bb4cb74 · inbound

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies cites this paper.

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:55:04.159093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T11:54:57.866685Z digest=sha256:55235ab69f5414aa1251e5cef2b6119095bb1e6873b79ae508c2801a4c81baa0

Observation 00521669-6dec-42f6-945a-8f2473994a8e · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:998a1ed5e3de9544a957900ea9e382faa11dff80bee2e38f4b31c39c61ee9736

Observation e59a8f7b-b208-4f10-bd55-49962d8816da · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.630018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:2ee2e061b43248321e34148dbc35f1b9d5b0e7832c23c27025890c44e2ed8276

Observation dd0e3994-6bfe-48ab-98aa-0a0d3fe5c70d · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.480995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:603b6061f98f9929d40900f414a37ed435918d91b9f619d1478ea0036e62a999

Observation c6b3ca7a-5ffb-44a1-9f66-0bf9626af0c8 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.674048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:2e9af557d5394811cb99942671ecaa00a850e19433b4920d2deea1e4d950b114

Observation eaafebcc-fb68-4ffd-9c74-5bdd74bfe522 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:53.258341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:b4e0881e510715f0a82f5af3e05441ba2020618ee3e3b6bc2ff7e9b4ee1b33f5

Observation f5d912b0-9f41-46d9-acd5-ea7348a7f017 · inbound

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding cites this paper.

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:56.223208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:21:23.325806Z digest=sha256:fffffeffea717dcbc7e40af300fd18975de442f02e11e9b84c7497bce434e7f2

Observation 06a164a0-6f2d-4115-bdb6-49fcb209186a · inbound

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding cites this paper.

Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.522744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T07:41:12.038619Z digest=sha256:190d8be81ab8411346d378690358759cae0f09cb4bf46e2defa8e38e4c86071f

Observation a8263df0-2b6d-435c-8b68-aa9a8721d04d · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.469693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:28b086476d56ad0fc3066499dba53e217e2d05efb40f2fd0a7b98b2d1bc1f01b

Observation 72bf615c-edb6-42d7-bd65-6b99f9a55500 · inbound

Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning cites this paper.

Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:04:43.127249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T08:02:27.951445Z digest=sha256:ef42328e3c9d071d06d4cdc10aa7df5a3bf9c287c585673f0c5bf9b2477cdfe8

Observation 7eb4df5d-9408-47fd-b34e-d984144a8237 · inbound

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models cites this paper.

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.959872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:10:08.682307Z digest=sha256:5300b2e9d04cb2cb5c110fa51e192bffc4859db73fe0ffd5e79f171d2bef238d

Observation 15bae689-3df4-4612-b061-054739eb52de · inbound

MODIP: Efficient Model-Based Optimization for Diffusion Policies cites this paper.

MODIP: Efficient Model-Based Optimization for Diffusion Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.560909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T13:44:19.708550Z digest=sha256:8370f1502b2b7a582942fb93a3d01329a7656e57a9b4a3a8fda883f34b909c7e

Observation 2126534e-497d-41c7-bdb2-2c6513451b7d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.921976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:8a15852d58647e7434b67a944d5523ede12cead957672b916765d0eb3d69b826

Observation 5e6891cb-20f9-4ec3-972d-18881248cdac · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.726202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:294b4827e2f512dd0c28f485b75c46351694c76b6a5acd51bbcb970d6b8d4f21

Observation bf5b6a4e-37b3-4eb4-9983-2349a56664a5 · inbound

DiPOD: Diffusion Policy Optimization without Drifting Apart cites this paper.

DiPOD: Diffusion Policy Optimization without Drifting Apart Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.948665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:b7f645f7c48c7a80575d763358f5e1f233ca3ad14407f3a13f6120fcbfaafd68

Observation d861aa65-49e0-4a40-a343-920ed4c51fe0 · inbound

Reversal Q-Learning cites this paper.

Reversal Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T18:38:49.871380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T02:30:24.951689Z digest=sha256:2d6d35a42da1c8501538b5e41ebaea1563910700fa11cf45e0fa4b0547a7fab8

Observation b93a7f98-09d0-4631-8508-e0dd233a94dc · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.650888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:2b5024291c5e1705db10104acaa40a4f5e38b42a4b25fa01e6a7ca26875f561f

Observation 9733d33a-8d5b-4d57-9a76-cd74781f6036 · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.657378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:513b0589022439c9c6923f6990e354572d4babbd9b778b67dbaf740f002ba6ea

Observation 9f8c20cc-bc8e-4874-bb85-8fbe4589b096 · inbound

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models cites this paper.

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T00:12:42.173815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:12:42.173815Z digest=sha256:ba5cd3e203e0225f769ec6c529c7dd6db0e576464a15bd709e5fc247d31e4dbc

Observation ade8d8dd-59e8-4d12-b661-e9019c6d7f85 · inbound

Adaptation of Generalist Robot Policies with Minimal Data cites this paper.

Adaptation of Generalist Robot Policies with Minimal Data Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:18:48.614940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:18:48.614940Z digest=sha256:0f72835159bd77f3e2cd5429aaa9463a141e34e3fe88c993553456fb8f6b0af8