Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:28:38.548957Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 31 inbound Pith citation observations for arXiv:2412.06685.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:28:38.548957Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:23:38.206985Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:59:44.649113Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3682650-6a4a-4d7c-a290-040be540b661 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Abdolmaleki, J
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 57210eb1-0b26-4073-b9cd-de722e8ef3cd · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3756bff2-2cb0-4f3d-bbd5-ab83f380eccc · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Efficient online reinforcement learning with offline data
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c640ceba-ea63-478d-9769-dc38c584847c · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85351a09-9f2b-4bdf-9146-2437a62c1183 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone FireAct: Toward Language Agent Fine-tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c44608-52ce-4d84-8189-988d0b3f0e39 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusionpolicy: Visuomotorpolicylearningviaactiondiffusion
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f289261-15e4-4165-a12b-3283755a8f5b · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridge data: Boosting generalization of robotic skills with cross-domain datasets.Robotics: Science and Systems, 2022
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cea8ecc9-0983-428c-ba04-b08daa2027e2 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Stop regressing: Training value 16 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone functions via classification for scalable deep rl
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fe6e76d3-abec-4f35-a97c-05ea505deb62 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8af79e-74b6-4c28-8c47-cd9e726ab9b7 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone A minimalist approach to offline reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c6269a-6684-48ff-aac8-9c76364803f2 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Addressing function approximation error in actor-critic methods
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7603abf-cff8-4374-bcc1-f8de43aa4872 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Off-policy deep reinforcement learning without exploration
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f97c0a-d888-424c-84bd-df3892676277 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Emaq: Expected- max q-learning operator for simple yet effective offline and online rl
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dfeb17e4-84a0-4cd7-b97a-b32763da0934 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9e049b36-5eb2-4ef9-8322-49db2d15f830 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd3e3b7-9c83-4ba5-b143-31e20c2cc21c · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1419ef9a-d790-446c-9adb-562c0f9be061 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900932f6-9fa5-4bda-80de-f8db9fe56cad · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8718a4-beaf-486d-9b04-86ee4a729b54 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone LoRA: Low-Rank Adaptation of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a0a684-f0c9-47c0-8fd4-84020e7182dc · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning as one big sequence modeling problem
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ac1b492-06d7-4e66-9575-0adf6385fb92 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1f6e5670-a8d5-48f2-ac50-1c6e6ab01acf · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone OpenVLA: An Open-Source Vision-Language-Action Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393c17c9-5904-4e36-bf7a-a34d0973a03f · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Adam: A method for stochastic optimization.International Conference on Learning Representations (ICLR), 2015
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6e224e8b-02cc-4e5e-ac33-cc3e93572b68 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline reinforcement learning with implicit q- learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 79359524-3e4b-46fc-b2d5-2a8a0875a164 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Kumar, X.B
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 775538a6-0dc3-460d-9935-cbb9cf9f8c34 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33:1179–1191, 2020
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ea5083-aa4e-4af1-98d5-343aae6bfc45 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a74265-7d7d-4d62-b4b4-a1440a509201 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b542df0b-9d74-405b-a4fb-5849c2da2e5d · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning multimodal behaviors from scratch with diffusion policy gradient
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6be748c2-c87b-41ba-a477-ed2861044d62 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Continuous control with deep reinforcement learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccef55f1-8c98-40b7-8b32-6d2f4de27bbd · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Leveraging exploration in off-policy algorithms via normalizing flows
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f33dd6af-30ba-42af-b754-8a1c475586e0 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a73538-74a0-4250-8201-8018fa5efa2e · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8658d600-0fce-4e4f-9475-4523f84572db · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Steering your generalists: Improving robotic foundation models via value guidance.Conference on Robot Learning (CoRL), 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e0c1692d-7dc3-4467-b535-c87bef76b780 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d113f657-f8b1-4641-86b1-34531597c672 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Greedy actor-critic: A new conditional cross-entropy method for policy improvement
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b2e2af66-2abf-4477-8231-f1dd1d53af56 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Self-imitation learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 081fa1d1-61be-4a29-b247-bedcb98d486d · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Is Value Learning Really the Main Bottleneck in Offline RL?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc20772-9a24-4b5c-b8bd-184f68796180 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9f211b-d82d-4b83-b548-731c999f232d · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Peters and S
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a24d8794-b64c-4072-ac60-d338d34c7b5f · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Relative entropy policy search
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation baeddc20-2273-4533-9e6f-f4778a91a6f2 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Learning a diffusion model policy from rewards via q-score matching
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 56fe718b-1dd6-47f2-9a18-cc015507edaf · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion Policy Policy Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0a1f4b-0a67-4d01-a0b4-770a9365f531 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a038d1-d0e9-46fc-adb9-92ee0820caa6 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Grac: Self- guided and self-regularized actor-critic
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 221685e3-8468-4755-a020-05eddb470901 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Skill-based model-based reinforcement learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d14fd665-77f4-409e-8587-75aaf4caee95 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a22cd976-2b68-4d26-be2f-644f254f58bc · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Hybrid RL:UsingbothofflineandonlinedatacanmakeRLefficient
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 24994787-2415-4d81-a67c-61527fce5091 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Second edition, 2018
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74b0c52-2c29-4f23-813d-c70d0a960b5d · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Preference fine-tuning of llms should leverage suboptimal, on-policy data
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ddd37579-adeb-407f-a972-a93d3dffe64c · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Bridgedata v2: A dataset for robot learning at scale
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b474557-84da-40f9-a582-bb79fa41abf9 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Diffusion policies as an expressive policy class for offline reinforcement learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a61b21-4098-445b-a622-ebf21a549ae2 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed0b98b-a5f9-42a4-a95c-3ce9e28b4a5f · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone V-former: Offline RL with temporally-extended actions, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a02aeaf-8436-4d36-8fe7-741fe74598f3 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88d1e079-8e37-4248-a4c5-d907d1663a40 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Policy Representation via Diffusion Probability Model for Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f2ede7-b571-46b7-9152-e41df897f2bd · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Mastering visual continuous control: Improved data-augmented reinforcement learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 739b6e11-9382-4951-b6b5-b9243dde2068 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Autonomous improvement of instruction following skills via foundation models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7aaaa09-ce8a-477e-8bb3-68901f58d8f2 · outbound
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone -v0” antmaze datasets from D4RL, but Fu et al.[9] deprecated the “-v0
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8c7411f3-fee2-422d-9eaf-f55ce5b57aa2 · inbound
Flow Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130da234-bcdd-4f93-8ff5-e34a16e1b8b0 · inbound
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5de78b-2d75-439a-83d5-9af270e71cc1 · inbound
Exploratory Diffusion Model for Unsupervised Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad9fb00-7442-47ff-b970-fb7420bd510e · inbound
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5446eea-762d-424c-bc91-e942572d7a78 · inbound
Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · inbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4335a470-040a-42bf-9ca8-fbf1a0758d2f · inbound
Steering Your Diffusion Policy with Latent Space Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cdd8c86c-0b3b-4f1a-af6f-f9b4914be986 · inbound
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609b4a33-52c6-4b10-aecf-d739fcbb9438 · inbound
$\pi^{*}_{0.6}$: a VLA That Learns From Experience Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88261a3b-d278-416d-8adc-03952f33d58e · inbound
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6e7b3928-7f87-4e28-8f30-79961d2b7c67 · inbound
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 12e3cd77-d007-4137-8d44-38141bb4cb74 · inbound
HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 00521669-6dec-42f6-945a-8f2473994a8e · inbound
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59a8f7b-b208-4f10-bd55-49962d8816da · inbound
Reinforcement Learning via Value Gradient Flow Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dd0e3994-6bfe-48ab-98aa-0a0d3fe5c70d · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c6b3ca7a-5ffb-44a1-9f66-0bf9626af0c8 · inbound
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eaafebcc-fb68-4ffd-9c74-5bdd74bfe522 · inbound
Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f5d912b0-9f41-46d9-acd5-ea7348a7f017 · inbound
Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 06a164a0-6f2d-4115-bdb6-49fcb209186a · inbound
Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a8263df0-2b6d-435c-8b68-aa9a8721d04d · inbound
Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 72bf615c-edb6-42d7-bd65-6b99f9a55500 · inbound
Target-Aligned Bellman Backup for Cross-domain Offline Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7eb4df5d-9408-47fd-b34e-d984144a8237 · inbound
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 15bae689-3df4-4612-b061-054739eb52de · inbound
MODIP: Efficient Model-Based Optimization for Diffusion Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2126534e-497d-41c7-bdb2-2c6513451b7d · inbound
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5e6891cb-20f9-4ec3-972d-18881248cdac · inbound
Improving Robotic Generalist Policies via Flow Reversal Steering Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bf5b6a4e-37b3-4eb4-9983-2349a56664a5 · inbound
DiPOD: Diffusion Policy Optimization without Drifting Apart Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d861aa65-49e0-4a40-a343-920ed4c51fe0 · inbound
Reversal Q-Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b93a7f98-09d0-4631-8508-e0dd233a94dc · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9733d33a-8d5b-4d57-9a76-cd74781f6036 · inbound
Adapting Generalist Robot Policies with Semantic Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f8c20cc-bc8e-4874-bb85-8fbe4589b096 · inbound
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade8d8dd-59e8-4d12-b661-e9019c6d7f85 · inbound
Adaptation of Generalist Robot Policies with Minimal Data Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.