Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.700024Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2504.13368.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.700024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:15:33.500556Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T13:20:25.624340Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b342661-463f-48c8-a8ef-cf03dd5036a1 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ebd34b7-bae4-47c8-8391-03904004dbbe · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f32b47-c730-41e2-986d-79a27a826be9 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c5f9e2b2-8542-4e3f-8f5a-58b6cdc60273 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d90d0a-18cd-432d-aa92-d9225025a03a · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec01e9ea-5c25-453c-a674-83646d202d6c · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3934a4-3298-49f8-9965-e07a2f740d3d · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 617dfe91-2cec-43d2-8e17-4f1fa250f7b7 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Score-Based Generative Modeling through Stochastic Differential Equations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ee2d5f-1c58-4ae7-bb69-e5c3ef572370 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0256fe-b7f8-435e-a6c3-a3eeb64e433b · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b7d0b8-9b42-47d7-80de-5d9d625d917b · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning State deviation correction for offline reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c2cd90d9-4ab6-4407-bb02-78c47372e643 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db81075a-fd52-40f1-a1e6-31e9be0ab235 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Deep Residual Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a39bd8f-c25c-42fe-a576-4af9b312790f · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning The regularization term aims at imposing visitation distribution constraints (Nachum & Dai, 2020; Lee et al., 2021; Mao et al., 2024a)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2af0adbf-4237-492f-ab9e-0cae667c72e1 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d939c2d7-5ede-4e54-8796-55978a24dd37 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning The number of total transitions of the noisy dataset is 1, 000,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cc305ca4-ba3a-4175-99f7-483d3a5319f4 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 954ca0b2-2b1d-4d5c-9fc2-1b06da0667b8 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
Reference 1960
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4ba7dc5a-5a45-4db6-ab8b-bca1f4cf824c · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
Reference 1970
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03bfdfd2-fe44-4ea7-9718-884314e52c63 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Reinforcement Learning via Fenchel-Rockafellar Duality
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bb28c8-75ff-49e2-8b63-09e93078fd93 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Off-policy deep reinforcement learning without exploration
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dc791905-98fc-49e8-b565-b45c55949573 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Harnessing Mixed Offline Reinforcement Learning Datasets via Trajectory Weighting
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7dd64b-96ab-44cc-84cf-2c8c0f035a0c · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Residual algorithms: Reinforcement learning with function approximation
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 83c1896e-56f2-4494-997c-83046a098000 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Is Value Learning Really the Main Bottleneck in Offline RL?
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc30dcd-a736-4d31-8c6a-aa2bc058f138 · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Dealing with the unknown: Pessimistic offline reinforcement learning
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 62368e0a-39bd-40c5-b908-d52c4694333f · outbound
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e4dfae-42b7-4f45-8e16-afb9d4a6634c · inbound
Semi-gradient DICE for Offline Constrained Reinforcement Learning An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a1765c-5ccd-42f4-9528-f69f63c67c47 · inbound
Dichotomous Diffusion Policy Optimization An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d266f3-efc1-4b59-853c-1d004699c332 · inbound
Reinforcement Learning via Value Gradient Flow An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.