Pith. sign in

Paper Citation Record · LEDGER

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL

As of 13 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2412.18855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18855 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:30:49.187659Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3be00b87-80c9-4e0d-83fb-77693bf1b357 · outbound

This paper cites Constrained policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.935676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.935676Z digest=sha256:9a16b979fe6c5204987d1c4ab68d932d064ad2739f0f5fa9b62c6afddb7bfde7

Observation b8e8779d-b6d3-40c8-8977-23a777fb2748 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning: Theory and algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.087909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:48.941106Z digest=sha256:a195b165728033cee8ab130188a30479810466a200d546b69c262a36cd3c0989

Observation c1eb6107-27ea-47da-b206-13c9cf8efe86 · outbound

This paper cites Reincarnating reinforcement learning: Reusing prior computation to accelerate progress.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reincarnating reinforcement learning: Reusing prior computation to accelerate progress

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.071758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:48.946054Z digest=sha256:d6a26122715214daa51db454a926368c3843731a1fb156b9c0ef8564e9f51570

Observation aeba181c-5585-4ce2-88a1-953ce5a8e638 · outbound

This paper cites Efficient Online Reinforcement Learning with Offline Data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Efficient Online Reinforcement Learning with Offline Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.950682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.950682Z digest=sha256:b4c89599ac55d01d42e512b6cf598d149b1709d723a5880256a15bcee7592341

Observation f57d225d-8fdb-48d5-9a7b-b1e140eff677 · outbound

This paper cites Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.955499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.955499Z digest=sha256:265a6742dc917bab2d4cb6e9d63c458436ff98ccc6050a7a5f4263ec377a228a

Observation 931c0e32-c4c3-48b9-a8da-f1bc7699846c · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.960378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.960378Z digest=sha256:6283a7d6ead359e9ef8fa46202eecb48398d5d8967226ab97d59499c8d1cb55c

Observation 8d97be98-45b7-48a3-8e26-6d41cd6f76c7 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.965286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.965286Z digest=sha256:038961e291f1b9d4e987448022ca287c841ec552a5acb831234aae46742f58a0

Observation 16a0bb77-79d9-41f5-bf6e-e2a850b4f462 · outbound

This paper cites Uncertainty-aware model-based offline reinforcement learning for automated driving.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-aware model-based offline reinforcement learning for automated driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.047352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:48.969948Z digest=sha256:53c29a19108e6880edbbd813ecfa3cfdce440234bf56a5850b58526211bf39e9

Observation 1be9be94-cf64-49a2-a916-7783e452b959 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.974180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.974180Z digest=sha256:cbf3475dc072688ed7107ec3547d9994fe88bc4bc7d77e5702556835cb47a193

Observation 3b822ef4-6f86-4a48-95c5-9a5f7ddfd70f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A minimalist approach to offline reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.978973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.978973Z digest=sha256:df8989336a13f3a6ff9815e07b12f6993d40426c07bcfc9ee41c621afa4f6e80

Observation 26167c32-88bf-4702-b62e-2c88e1012874 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Addressing function approximation error in actor-critic methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.983533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.983533Z digest=sha256:3c8fcdcd64bb35041f71131eada082466117449ede5267202c57239b011ed982

Observation 3e884ace-2ae7-4635-b915-da21ba897fbc · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Off-policy deep reinforcement learning without exploration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.987827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.987827Z digest=sha256:b43002d58c518fa207f040ae42e47913c3dcde94e268c8720bbd730fa436ca36

Observation 94ccdcad-0bf9-47a6-ab8c-1e72494dfeba · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Extreme Q-Learning: MaxEnt RL without Entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:48.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:48.991946Z digest=sha256:f8d83e753d85971cdac6d9ee4e79e34cb864f72b342d947a9d6a19c24cb8b956

Observation 1078a651-dda5-442a-99a1-292e83de04f2 · outbound

This paper cites Open and real-world human-ai coordination by heterogeneous training with communication.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Open and real-world human-ai coordination by heterogeneous training with communication

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:50.001318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:48.997192Z digest=sha256:7570d466075dfc121132f4c3f4b61036d0d87712fe999f4ca598c6fb18194ed8

Observation 791cd18e-9715-4197-8547-41090e436c30 · outbound

This paper cites A simple unified uncertainty-guided framework for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A simple unified uncertainty-guided framework for offline-to-online reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.001554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.001554Z digest=sha256:d20c255a0de7e654aa2a89813492f4ecdcedc3fb916c11bf845c095edf638bcc

Observation a2a59874-9d4a-42a2-8e17-a87128268627 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reinforcement learning with deep energy-based policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.005983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.005983Z digest=sha256:50ad45d0f50f5c9649394b08348bf77bca136e2133c43c236c10fac83a0b4f4b

Observation 680d33f4-22a6-4f31-8ff5-940f5489d238 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.010070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.010070Z digest=sha256:9c5fc83c90340b6bbf2d5d8b9d34739e0288c8918840b38ac177f42696e6757d

Observation f919f741-e582-43b1-9b2d-b1242342896c · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.014463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.014463Z digest=sha256:f71b7553ff7356ec750844a6e47e83fb9a931536d8c64144231c69677678e634

Observation 516b2635-e169-418f-987d-e79741e2376e · outbound

This paper cites Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uncertainty-driven pessimistic q-ensemble for offline-to- online reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.969216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.019462Z digest=sha256:0a354f0b0811cd55a0e5453c8223caf4fe62f600c548b4a8c16cc12d8a421eb4

Observation 64af4ca2-64c0-40f6-9c16-df69ed7322cc · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline reinforcement learning with implicit q-learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.023652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.023652Z digest=sha256:95b60d185a52cf10261b4e1728b35cb532160416a6cb2dbf0649a786b0a6d9b0

Observation 6746a8bb-fd32-4958-a3e1-2deb673aa3ab · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Conservative q-learning for offline reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.027727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.027727Z digest=sha256:33f5f58ca440c56fcb6e6273fa373310f2ff48b51369b2eb176caa89ac47f3ae

Observation 17fb3d7f-9744-436b-8fe3-685d92aaf3ba · outbound

This paper cites Batch policy learning under constraints.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Batch policy learning under constraints

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.935505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.031806Z digest=sha256:6832fa49fc546fdba8a9023f589b4bef26dce90ce790638c56bc8966d0666834

Observation c00cb2d8-d4c6-46e5-b035-6c9dc9e02144 · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.919106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.035910Z digest=sha256:e72103d976c13aafea4c88c22c3d8317a770c0ccb803b5edad00681bcf30c9fa

Observation 391ffc41-f4b9-468d-a5cc-564f956456eb · outbound

This paper cites Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.039985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.039985Z digest=sha256:0deea8a67c176669e9f2444a01df1f6ca079227516c2c932dea6157159da7245

Observation 7b857664-28ff-4841-a166-1eed0a315c88 · outbound

This paper cites PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.044417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.044417Z digest=sha256:a2204db25b12cf5e483f564ac5259b5b1013efdf76fdcd5cf18eeef122e6f345

Observation 82fd9920-06d9-4eac-9227-6fc5976b813a · outbound

This paper cites Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.048977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.048977Z digest=sha256:a43aa55631c8ad9c0b5f426a475eb592c0e3477f9809efafd86052b8405617af

Observation 53a18ba4-5f14-4d3c-8f35-9dd8a1d5391c · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Mildly conservative q-learning for offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.903853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.053729Z digest=sha256:3133b2dd721011f4b592df57c7874157db1da0f24adf0a8ca820a3a5f763a615

Observation 9dfaeb79-4d21-4164-9f6e-3c6368a1aa84 · outbound

This paper cites MOORe: Model-based Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL MOORe: Model-based Offline-to-Online Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.411742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.057905Z digest=sha256:38c4d0819b034e1b0081fd8b3e06fc60e12c1149f2ab07e5d686ea0b4cda51b8

Observation fbd3e492-d978-437d-938e-9749de752dea · outbound

This paper cites Supported trust region optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported trust region optimization for offline reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.887854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.062641Z digest=sha256:442382bb595e10cd04c130cb5377ea688ce05dd6760b3a06b12553a806d545a4

Observation 1281f0c8-a005-4ff5-b632-a1173f4816a4 · outbound

This paper cites Fine-tuning offline policies with optimistic action selection.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Fine-tuning offline policies with optimistic action selection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.872845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.066875Z digest=sha256:2d89df423dfb5d7de011a79acae64115ccc3e1d3622a3e851b0476e5cd4b8dda

Observation 5d037d44-40f3-4ef0-96c3-9441caeb1c8b · outbound

This paper cites Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Planning to Go Out-of-Distribution in Offline-to-Online Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.391005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.071160Z digest=sha256:be363898848e169160013da5a17eb35826136d1f5fa82c2ef90bb1b43f8f216f

Observation 87cd898a-545a-4798-8d34-cfe6b3d63c7c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.075601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.075601Z digest=sha256:366cd5dc5c7ba825d25787da88dfa3f28a91a31d284a608cc9cfaceba7523ee3

Observation 2e2f4430-3cb6-4f79-ba2e-f0e4e1694785 · outbound

This paper cites Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.080017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.080017Z digest=sha256:607fb333138b8d8b52eca4d1cb54ce5fecf40b6a14b0085c251c2d3907b6da1e

Observation 77df57cc-3353-487f-a5cf-429697734043 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.084656Z digest=sha256:f0ed14fce08f9b88c2274c7ae144aa214012f58764460598ae83253407ad1d14

Observation 16c7240c-2801-434c-9ca7-525d0eecf8fa · outbound

This paper cites Trust region policy optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Trust region policy optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.089965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.089965Z digest=sha256:49cb338ed977403564caa39672fe8aa1e63b78c791a171163cf49cbe2ebe91e0

Observation 3f1115b8-aee4-4409-a4d4-0693a3fdc85d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.094279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.094279Z digest=sha256:b9212267f23119f6dede4cf8321a70b0c69f7cf29be3dc3efbe48926e53433b3

Observation 3b2e1149-fdc4-4d38-8b3e-976e37b5566f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.098890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.098890Z digest=sha256:898f396c66c1ba29310791fa25f16ad54402e8fe73fada41a91a6ace792246c6

Observation 88530470-1a24-44bc-9604-bad92fa4e204 · outbound

This paper cites Sutton and AndrewG.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Sutton and AndrewG

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.831044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.103172Z digest=sha256:6e682e5b6c0ea1f039cab10fa67c90a2a4d30d43ef1b96f8e271e34dded60001

Observation 601c7bc8-c6a0-41ef-9765-2c427eebd7dc · outbound

This paper cites Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Lever- aging factored action spaces for efficient offline reinforcement learning in healthcare

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.815625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.107419Z digest=sha256:877b0b5ad0af6664aeb8a64e814ed5056f042a68d110f7645711e50e78ea55a5

Observation a778597d-eb97-48f1-a963-2c861bc24650 · outbound

This paper cites CORL: Research-oriented deep offline reinforcement learning library.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL CORL: Research-oriented deep offline reinforcement learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.801226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.111884Z digest=sha256:f3ed63e7184bed73a79d6e153e216e88fd4b31861120265b9851490adda46eba

Observation 01c0eaf4-3bea-4485-811b-50926caddbbc · outbound

This paper cites Reward Constrained Policy Optimization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Reward Constrained Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.116005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.116005Z digest=sha256:b27d1453c480231cb3d781dd4b4d6e9b5f2e277c7346be70f40df3469eb01922

Observation d2f83805-5139-4b0d-a893-8b0679a025e8 · outbound

This paper cites Jump-start reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Jump-start reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.786819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.120558Z digest=sha256:5eb6c1d5baef3be56fce4b96a872a019d7e2f4c359226ca2b3f7878fe8fd6cfb

Observation d77c9109-c886-4784-b58a-52e7e42c26cf · outbound

This paper cites Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Train once, get a family: State-adaptive balances for offline-to- online reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.772852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.125050Z digest=sha256:6e4d7cfcf7546f8609419eb68deeecfac37b46cd02c87778c612850bf49bc1d2

Observation 7cafb6e8-c05c-48d4-ba66-1963ea3253e7 · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Supported policy optimization for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.758387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.129229Z digest=sha256:675ffbf598240dc1b5fad1cbbb0f392cde229b3ff54c0fd874e85997110b2600

Observation d7ec5233-1756-45db-b3f8-afee46f24507 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.133453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.133453Z digest=sha256:28e959e8df34c4c4b3c7fefb794edf34f13d2ba9090e97b62a41438841f31401

Observation 7759f3da-00c4-4d39-b472-4b77da252b4d · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL A policy-guided imitation approach for offline reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.733631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.137747Z digest=sha256:294c65c7b15776bfc831d093eb6a1747b1465c6dec598fde095b934d3c15b0d0

Observation 20d6d736-a044-495b-8099-7d0881f00018 · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.142024Z digest=sha256:dea029edadaea6c3252d095c20de3f0acc44052c0d16ce4104310a19d494403c

Observation 7bafb258-e4a9-481d-8b7f-b1ff5d9a276c · outbound

This paper cites Actor-critic alignment for offline-to-online reinforcement learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Actor-critic alignment for offline-to-online reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.717995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.146683Z digest=sha256:a442d4ee9af59c01ae118ab75741d50fd2eadd38591d49a365369234132c552b

Observation 4058ab69-5d44-424e-977b-1bd44d22b625 · outbound

This paper cites Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:30:49.280625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.151220Z digest=sha256:1f0b541df1ab1d5a4a1b13f38ad5718a6ac14d96d89aaffab557afd5c7233295

Observation f05126e9-b6dd-412d-9435-c2f8ec460bd6 · outbound

This paper cites Policy Expansion for Bridging Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Policy Expansion for Bridging Offline-to-Online Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.155903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.155903Z digest=sha256:5dd7781fbbfba7d69cce4751a31f0156db9265b4d54d086b4ced034fb23dd97b

Observation 36075a83-cc30-4203-aa73-eade42be6aea · outbound

This paper cites ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.160669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.160669Z digest=sha256:4b63cea41fd205b59701ace40e4cab97d57121f6c92c37be324d5421da211459

Observation f1c5647e-6306-4562-ab2d-7b441ae543d8 · outbound

This paper cites Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.165225Z digest=sha256:5c5d8cc55750bc06a9cac1f202c8e677f5a8e1bc0be3c0f2f0e4599c5cd2e38e

Observation 652d11f2-dc55-47c1-92c3-2f48f7e03ee6 · outbound

This paper cites Online decision transformer.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Online decision transformer

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:30:49.702843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.170538Z digest=sha256:d3fc0a88afcdf090dc097827b52416899110b81a4788c7a26bde356eb8e27118

Observation 6a610778-6fdc-45ac-94b5-8d919abe8ac1 · outbound

This paper cites However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL However, our O2SAC still outperforms it with less computational cost during online fine-tuning and less requirements for offline policy

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.688112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.175429Z digest=sha256:8b365b8955134406d5ef2fec054493a169c4cbd3fd6f95a28a694e932d138542

Observation 2de3f5ca-53a8-465d-90c2-41cd7ee63c22 · outbound

This paper cites For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL For model-based O2O RL, [31] explores regions with high uncertainty and returns in learned model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.673069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.182233Z digest=sha256:ea89b4928779b651e90a28a868493acf5e8f2e376510322ebd13f2583cfe998e

Observation 292accf8-fd76-46a6-88f4-c3b9448ede9b · outbound

This paper cites Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Moreover, [4] find that LayerNorm is favourable for efficient online RL with offline data

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:30:49.658434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:30:49.187659Z digest=sha256:cbe17d952d56649717bae238620d29a2c44bc526199e8765cb070c8140a91813

Pith citing papers

No inbound Pith citation observations are available.