Pith. sign in

Paper Citation Record · LEDGER

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

As of 12 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2607.27203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27203 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:27:44.056272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc8d6680-afd6-4aa9-abff-3fbe039a3cd8 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning with Offline Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.851239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.851239Z digest=sha256:8aab9e8bdfdca8b34a7e2df62c7071ce32821cdf33ea1a12f0cc7ea1bd109856

Observation 435b5e08-e59e-4b49-8514-ea7445d15ea1 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.855135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.855135Z digest=sha256:8f4125948f3524e6d4859ef6aac04853b6ad07dba6164a3e71201cb03bd1ff09

Observation 49b28c9e-8205-4ea0-b167-2c60e716a353 · outbound

This paper cites Reinforcement Learning via Implicit Imitation Guidance.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Reinforcement Learning via Implicit Imitation Guidance

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.858412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.858412Z digest=sha256:d9a51da8739631e96fad8b7eba2f06c3f47fffd289ab7d151f0694d9d17c9474

Observation 3de1e9ba-ad4e-4dc2-8962-641f2c84653a · outbound

This paper cites EXPO: Stable Reinforcement Learning with Expressive Policies.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? EXPO: Stable Reinforcement Learning with Expressive Policies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.862543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.862543Z digest=sha256:48dd581c68d9804e1fcbf5e2987309ccdcfde5f59fc23eb0624da7ac72a7415a

Observation b753a510-00a3-4072-b11f-db5da5f4506f · outbound

This paper cites EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:27:44.406212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.865881Z digest=sha256:8ce784569e82dec413d9dd1897b436595f50fced6e838a2f6a3779203f30b9b3

Observation 4f1dd5b7-39d8-4d00-bd48-d23ff9b982ad · outbound

This paper cites Tql: Scaling q-functions with transformers by preventing attention collapse, 2026 b.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Tql: Scaling q-functions with transformers by preventing attention collapse, 2026 b

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.869331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.869331Z digest=sha256:4108e2d03c54c0065cf7bd5ed1b4a106d96b0df659c5b76763e3e22c064b20d2

Observation dfb8630c-ab78-40c1-9a6e-8839ca96ed43 · outbound

This paper cites Value Flows.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Value Flows

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.872344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.872344Z digest=sha256:7ba90c8092640c2b6f5caef4923da0c80854f5a67de2efbfc5c42313a3d6af83

Observation 1ba61b2f-ad5d-4c3f-90f1-50f313f9c14c · outbound

This paper cites A Minimalist Approach to Offline Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? A Minimalist Approach to Offline Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.875428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.875428Z digest=sha256:d79a6cec49e95f3a9e85e111fbd15e59844c88e6c3ea8200565e362f54d8b3da

Observation db5f9166-ea8d-42c8-b127-c020905a6e98 · outbound

This paper cites Off-Policy Deep Reinforcement Learning without Exploration.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Off-Policy Deep Reinforcement Learning without Exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.878960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.878960Z digest=sha256:8fb7ce9dd9753d61148e63d12523d7460ca0a42c0bee9ce59b36a91e2c060550

Observation e76a3389-5aca-4375-83dd-25a19774da10 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.882699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.882699Z digest=sha256:0c401beb8fdf67b60ce673735c9719c1d1f4555af9acf61d97a0f73a83c15a2d

Observation 30ce3aad-ceb7-4613-88d2-e5bc3e2ef94e · outbound

This paper cites an unresolved cited work.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:27:44.708940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.889168Z digest=sha256:5548b505c19e4d9c7dbc96f874b2a4391a082781162eae6d3e43fbf37dc3e6f4

Observation 4ef21e19-4a49-49a0-bd4e-28aec6a260ac · outbound

This paper cites Residual Reinforcement Learning for Robot Control.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Residual Reinforcement Learning for Robot Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.895668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.895668Z digest=sha256:0d87da0a913ebf700c789945680b68e0c24d1e244496d3c845ed406987d4cd92

Observation 8732657e-48fb-40c0-8468-94f37b81e5f2 · outbound

This paper cites RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.899296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.899296Z digest=sha256:66fb430a90c4b9a7032cbcce5be1c46257a63d3cfe85828adc08642fa9d87e87

Observation 3b1a3d1c-7b99-4f72-a8d7-3ff78b3bd4d5 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.903146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.903146Z digest=sha256:4b18f2b4a8fc4e21ea928760bb0cda714f420cedf858e5fe86ac280cd9225413

Observation 58830bf1-7272-4d8e-91da-9c5b9b5759ea · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.909990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.909990Z digest=sha256:8900977bcce22e269116f4d21f2ef71e03259fc7e7d5cdd83e478eca4bb3eaef

Observation bc29c0e7-e20b-4a77-80af-d7e93fb93bb6 · outbound

This paper cites Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.913078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.913078Z digest=sha256:a42eda09235830ec6ad302da4145df47385212e63e1f662eb90739ba639de62c

Observation 9ed5d8e6-24f2-4902-a4e4-8c77af3b4e91 · outbound

This paper cites Training language models to follow instructions with human feedback.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.916046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.916046Z digest=sha256:8305361a60e646145bdb92b47f5488f0fabe1ec5e74ab1cc73ba3d366954a7e2

Observation 59243467-6b8b-404f-be6d-410361a58dae · outbound

This paper cites OGBench: Benchmarking Offline Goal-Conditioned RL.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? OGBench: Benchmarking Offline Goal-Conditioned RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.918926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.918926Z digest=sha256:8ac375bb27dfaf494d6d1373223bde9e878dea786fcd6bc44a4236b78681b030

Observation 56944067-1641-4a97-b8df-a49a69282f82 · outbound

This paper cites Flow Q-Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Flow Q-Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.921979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.921979Z digest=sha256:427daea73f085cc21186ac7d5ad38dc71d7d3a1dbc93d01d4ba3eb9befd89a7f

Observation 5f3bcbff-1be5-4146-9a4e-5b34e7f19672 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.925054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.925054Z digest=sha256:20ed7e0cd82f4b1cff2a1be1e9bbb168e223499b781fe604855091ac77275bd7

Observation 7d59be1e-05dd-4196-8f89-cffe8033cd6b · outbound

This paper cites Diffusion Policy Policy Optimization.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Diffusion Policy Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.928017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.928017Z digest=sha256:6ca18da6d3d262108c3f75bbf7e01dd3ce468cd436fc70078eb87b8281a6b6a5

Observation c59d1e37-14ea-4bb7-b8ce-d3f0f8f81e93 · outbound

This paper cites Learning from demonstration.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning from demonstration

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.700434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.931195Z digest=sha256:91aae0acf2fe229574e3fa38d7aa4c908911cf388b2b304f4b96f50433b93a66

Observation 32783b5f-8530-4ddd-ac87-dceba0a00845 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.933855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.933855Z digest=sha256:640f716add4267f1044101c9d6b41303367592af58ffaeebf5ff0efd754d9224

Observation 63bc5cb3-043d-4cd5-89ed-347195cd729c · outbound

This paper cites Residual Policy Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Residual Policy Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.936921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.936921Z digest=sha256:489d256201e757938bda9cf3c7bbe9f21263df90df444f8d9b9d9584d7d8746b

Observation cb0d603a-dfea-4cfe-a598-f91434d33dae · outbound

This paper cites Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.940482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.940482Z digest=sha256:8a138caf7f9f2004c60bb28b652debd6846bf3c38735dcd416431900203fba6c

Observation 3bb83959-8101-4961-9f60-cad24838613f · outbound

This paper cites Jump-start reinforcement learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Jump-start reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.691778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.943635Z digest=sha256:3b6a28ba935c0af738c63d09d101fab231416afc34b4f650135ea8db191979b5

Observation 5cb8093f-1c55-45d5-b11b-e15ab67eeaa8 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.946235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.946235Z digest=sha256:b78af7c2ed65a45ddb4afce900d51782500bdd238cab594ac95b4aac7e3036c6

Observation b9a10551-8f68-4ff2-845b-e4b2c0c6343a · outbound

This paper cites Posterior behavioral cloning: Pretraining bc policies for efficient rl finetuning, 2025.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Posterior behavioral cloning: Pretraining bc policies for efficient rl finetuning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.949288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.949288Z digest=sha256:99b74685663b1c0abef16db0f8d859188c931bdc5bcbcdf853a099541f6df3ae

Observation 927ceae2-6df2-4eb7-b7a2-05f2ce14d0a3 · outbound

This paper cites Hybrid policy optimization from imperfect demonstrations.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Hybrid policy optimization from imperfect demonstrations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.682421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.952021Z digest=sha256:8dc4bf5da63449aa0d74cf81c023f25e60d4c1d931d4b0ec87c438b70cb03ae6

Observation 94deceb8-0c11-4433-b87a-394c4ee5677b · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.955043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.955043Z digest=sha256:e211bff6312cc67169927bd8bd268d801a92ac13cc1b6b095794c8f88ce7c224

Observation dce039d1-cbdd-4a12-8533-f5f4580e6e34 · outbound

This paper cites Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.958427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.958427Z digest=sha256:e1574cee34b446b5341a7b7c6725db13778e69942502f6781ab2548bc5d4432a

Observation 1066cbd0-4f29-44c3-9c82-b26fb9b31e4c · outbound

This paper cites 2024 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2024 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.673796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.961450Z digest=sha256:fdfdc26a80df1e547d666433e2f975c0ae53c61a5523eda75c36d564172687c3

Observation 7835ab6d-fdf0-4e5a-af30-89b276e89431 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.665295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.964400Z digest=sha256:db94fd3cf000686f8abc4e2ba5d76fe6dbd7e0d583231d312f9b034fd3dba99e

Observation 78f74e7c-1372-412e-805e-151b93895833 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.656865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.967979Z digest=sha256:fff3b7d4ff03d860f72843e9aeb22bca12793d0ebac4ce5e1775af0d3530ec9c

Observation ff6ca7d4-31ee-4b3f-86c9-dae97b4d45b1 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.970826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.970826Z digest=sha256:6b379035ba80533d0bc3380dde1fb383ff7d198db53ea873c2b6f3edffe8b43a

Observation 39bfc212-b23c-4369-ad9b-b4b6baa6585e · outbound

This paper cites Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.973679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.973679Z digest=sha256:90bc4efbb8fe339fcac1ac75516c32f7a01cb05695e20bb24d003ac2fde8aefe

Observation 8694ecfc-f619-4c79-bc33-0cc52e54f1d3 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.641172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.976727Z digest=sha256:33f9792d4117d40e728c215d3c899f32f107408f19695052586415e9490405eb

Observation 0d33480c-c323-4ef5-881f-ccfe952540f4 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.632942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.979698Z digest=sha256:1790b3c9a5e1603ff5b70606f6fab3485bdce48ca54fa5cf0fca0260c3d637e5

Observation e789b70c-4153-4194-9bf6-a7c476d0876e · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.624224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.983043Z digest=sha256:817a83234c55e82c73a7d947cf48ef3ae7620fa0236ca4b651585d7e5ef08922

Observation d1980857-ee60-4eba-9aa3-c30134053f99 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.615817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.985785Z digest=sha256:1b3e9bdf3e06c7c03237d30c954b0f0c2b5fdf2c3bf7ae8990927a0796e35b53

Observation dcf5a4ef-cd04-41e1-9252-af049d922457 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.988561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.988561Z digest=sha256:49c4ff2602e0a7638f7fcc1d10c487e6b712eac759c5deb9a0a341f1dd943af4

Observation d2787269-2f10-4c59-a4c2-3dec0ccc7b6f · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.602667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:43.991903Z digest=sha256:a9063aed581d1a5f9e6a39876ff07bd569347d9204e606cdb84af992a1af39a2

Observation 6e3502c1-8b5f-4acf-8f1a-3b3868634f40 · outbound

This paper cites 2022 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2022 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.994697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.994697Z digest=sha256:15760cd349cd5bfc8ea8425c7c76272102d00549506230c6faff0c3d4e381e83

Observation 261c0147-f982-4990-b942-df7184291f2e · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:43.997602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:43.997602Z digest=sha256:710e833d8c76d9b05c54321eec5ce931440c4b86e86f22fb6f8e629a23093518

Observation 96736635-cc01-4463-a940-97b4f349a94e · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.589458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.000430Z digest=sha256:a6771259bb6ee3a9ab8e75088346724dc18851d7e7b30afc9133177a9214d024

Observation e588344d-82c4-4db6-9258-71db439f99c9 · outbound

This paper cites 2019 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2019 , eprint=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.004064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.004064Z digest=sha256:8a221e3188b44025c268ee3cfb2007f0dfb251802f94b0a02fcac3c8a7922967

Observation 3340c488-d9d8-4969-8f98-e65a4360125b · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.575786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.006852Z digest=sha256:a3e85c9cde3f5245ea35c2f9f8ab52b2ceebfabaedcfdd2f3892775a244fccf9

Observation 6e818da6-7c3f-4ff9-85f4-d8474843bbfa · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.566579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.009721Z digest=sha256:eaed004ea5491280cded11d4567284d70df677d2d7ccdeed5607b1b81e7ba94b

Observation ff2fc180-aa2b-46f3-ae96-a689cd150192 · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.558137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.012520Z digest=sha256:dd557373d1a9394b5fd3e91cb7d41debd6c3200a68e515e8c0dfcef1e2898b15

Observation abc7736b-dbb5-47b9-a59f-a4c8830b68ad · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.549631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.015214Z digest=sha256:ff275ee39c5cc34367ebd8f3e433f6dffe0eba0fa0c85a313f4ef464803c184d

Observation 8e547edc-336b-4161-8db1-bd7993eca08b · outbound

This paper cites 2019 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2019 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.541037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.018125Z digest=sha256:5d1e419fad660f08439949870a751a6fcef35e81f6b4fffdcbc3a42bab1b50b5

Observation d034223c-a59b-4923-ac80-a71d0870a309 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.532567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.021235Z digest=sha256:375f7aadff2790e58aa98d9fe07f807c04abad8b643727abcc3b54e87c7d0e1c

Observation c8d85fc6-f76f-49cd-a26c-0f1333988e7a · outbound

This paper cites 2021 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2021 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.023977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.023977Z digest=sha256:08eb4d6e9ec07930bb82dac9eacc36e7360f598d0c84aca26b60bfc0058e0167

Observation 47dd6549-d3ae-4431-af91-e43b699b93cc · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Imitation Bootstrapped Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.027011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.027011Z digest=sha256:2a2cd7a4351dcd406927bafc9302b542301603f7220368449ec34ffbdd5fc4ad

Observation 3f7ed219-f687-47a8-997d-2073b0fb3a45 · outbound

This paper cites 2017 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2017 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.029744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.029744Z digest=sha256:0ce8f63ada30dbd06f2c97c67b6d4c8dc36143f7375b93397225e90f6ea4a2dd

Observation b2d80ac2-af79-401f-8317-cec902a0e667 · outbound

This paper cites Learning from Demonstration , url =.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Learning from Demonstration , url =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.513790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.032496Z digest=sha256:12c7a9d35f5ec6e6c62e1ee9f6b68308ea8555321d51285343e9e8a92da34142

Observation a4a62e94-6014-4bc9-8471-5a6f91f827f0 · outbound

This paper cites 2018 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2018 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.505056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.035984Z digest=sha256:1e40147c9a596bace60915c5085922f122d8ee7972221cc01ae352bd1aa304a5

Observation c5580547-0659-4f5a-bcce-2ef08c83c0a9 · outbound

This paper cites 2024 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:44.038696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:27:44.038696Z digest=sha256:80d61e0f2d1e9165bad7d864b658357736834b858fcfb60d2bc35f1a2a6e08d3

Observation 388ee01e-ae08-48d0-a270-2a8eb908f0e5 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.490436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.041526Z digest=sha256:2c3fa61dbea3f694c5e9c1431e00fb3cb3317bc3546cfdc76e52d0a96cf1a67f

Observation f4d91b64-1675-4b94-a1a9-8fd81f082c41 · outbound

This paper cites 2026 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2026 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.480664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.044253Z digest=sha256:da13fcad74cf130871bb7d0afc0cf316bf9df82200ad1d0e18cf136080353ea3

Observation b68ff781-7502-41e6-912b-c022c4320ba4 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.471999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.047941Z digest=sha256:a0a92720ffc87a8129450dd6971c054a33a93f72e089fabb4dc7bf1f6e0e172a

Observation c548dbad-505e-48e7-9df4-4407752df131 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.462445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.050890Z digest=sha256:53b2512c178e9c3a35583680866f5647c2b8447ca0db810f07103e4983c8f64b

Observation 494a7ca2-e0ed-46dc-8428-a1cd07b0dfa0 · outbound

This paper cites 2023 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2023 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.453888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.053550Z digest=sha256:316748d3abe59377d092b3c2fe4d68477715089f414628e7fcab4060c709ed23

Observation 48b63200-b705-4a67-beb3-e8378cfcfaf2 · outbound

This paper cites 2025 , eprint=.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:27:44.445380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T04:27:44.056272Z digest=sha256:0ab81535b23358db70311c93ba093885344dcef9bd955d6125f92753c1ceb5a9

Pith citing papers

No inbound Pith citation observations are available.