Pith. sign in

Paper Citation Record · LEDGER

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2504.13368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13368 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.700024Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:15:33.500556Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:20:25.624340Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b342661-463f-48c8-a8ef-cf03dd5036a1 · outbound

This paper cites an unresolved cited work.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:17:38.958027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.686214Z digest=sha256:ffbf99c2f13d199589c0672278356ca4bcb7b884fa186549dcb933780f3939d0

Observation 5ebd34b7-bae4-47c8-8391-03904004dbbe · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.616065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.616065Z digest=sha256:482709fdbecbbb8fc309fedf1826b246b2cee8400a8144b1c8acac0448ca53dc

Observation e1f32b47-c730-41e2-986d-79a27a826be9 · outbound

This paper cites an unresolved cited work.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:17:38.935708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.692945Z digest=sha256:e28edf30fdd0215d36eadd7a9c71dec459d45fef135b210668a0ea0e182f0357

Observation c5f9e2b2-8542-4e3f-8f5a-58b6cdc60273 · outbound

This paper cites When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.623944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.623944Z digest=sha256:a202e79f7ed34c8d8f8fc907796795485ffc41535f42b11acfaba4004e385247

Observation 13d90d0a-18cd-432d-aa92-d9225025a03a · outbound

This paper cites COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.627839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.627839Z digest=sha256:f075a95875168b78d926ef5f8bbaacc0358cd7b53e1771e9e1bf90714e41d42f

Observation ec01e9ea-5c25-453c-a674-83646d202d6c · outbound

This paper cites PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.631616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.631616Z digest=sha256:1b022b22f1401988578a69af66239932e136888fc0dcf3f8490fb28487d5db86

Observation bb3934a4-3298-49f8-9965-e07a2f740d3d · outbound

This paper cites Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.642669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.642669Z digest=sha256:4dd51383ec3704d58aa30c1ef5dd424c8b25b76015759f3fb4b140cd25a0bf2f

Observation 617dfe91-2cec-43d2-8e17-4f1fa250f7b7 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Score-Based Generative Modeling through Stochastic Differential Equations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.661193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.661193Z digest=sha256:991134b3a6089b38dc0813a827be9a5967ddf240dc9582fe9292a18fd88c8916

Observation 01ee2d5f-1c58-4ae7-bb69-e5c3ef572370 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.664480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.664480Z digest=sha256:4b06ab7cc3bcde1155a42b633245ec60ca1b00418da8f449bf300aa7b75e23f6

Observation aa0256fe-b7f8-435e-a6c3-a3eeb64e433b · outbound

This paper cites Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.668063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.668063Z digest=sha256:d834df5e916ebf08cfa055247f22791fb3e3e36d6aba9a1f1958f5e1e04d0589

Observation 21b7d0b8-9b42-47d7-80de-5d9d625d917b · outbound

This paper cites State deviation correction for offline reinforcement learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning State deviation correction for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:38.982010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.671311Z digest=sha256:4f68b08dab594dcf1b10e5dbf69db09ae50500cafdbe38b8fe60f559ec696c9e

Observation c2cd90d9-4ab6-4407-bb02-78c47372e643 · outbound

This paper cites SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.675196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.675196Z digest=sha256:e2d762973a241c95e171028828c40efcc92442b80c64fd6ba7097c62f8e399c1

Observation db81075a-fd52-40f1-a1e6-31e9be0ab235 · outbound

This paper cites Deep Residual Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Deep Residual Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.678696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.678696Z digest=sha256:bff79f0528fc2c1b2f79bd7ba8de348d874fea2d3b1cd8e0c6e99ffe13866c61

Observation 4a39bd8f-c25c-42fe-a576-4af9b312790f · outbound

This paper cites The regularization term aims at imposing visitation distribution constraints (Nachum & Dai, 2020; Lee et al., 2021; Mao et al., 2024a).

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning The regularization term aims at imposing visitation distribution constraints (Nachum & Dai, 2020; Lee et al., 2021; Mao et al., 2024a)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:38.970648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.682358Z digest=sha256:19d131119c0cd33a60a50b03cbcb175446cfdc5e1d172ccfb870e24ded7fb660

Observation 2af0adbf-4237-492f-ab9e-0cae667c72e1 · outbound

This paper cites an unresolved cited work.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:17:38.946805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.689474Z digest=sha256:e4f0d307a668fc25cecbecbc3dc57ff61d1de4a6e0dee36b1a9a2225b08b1add

Observation d939c2d7-5ede-4e54-8796-55978a24dd37 · outbound

This paper cites The number of total transitions of the noisy dataset is 1, 000,.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning The number of total transitions of the noisy dataset is 1, 000,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:38.924813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.696319Z digest=sha256:0f9139a1a742d89e29a84750c3d4ac81e9ef5721dfdbdd4e2c403cc89ca12309

Observation cc305ca4-ba3a-4175-99f7-483d3a5319f4 · outbound

This paper cites an unresolved cited work.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:17:38.913488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.700024Z digest=sha256:c4afd4aaad68e927f654c8f4dcff7ff72acdb0f79134c4082007e0a27b1dbaeb

Observation 954ca0b2-2b1d-4d5c-9fc2-1b06da0667b8 · outbound

This paper cites ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update

Reference 1960

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:17:38.825428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.647300Z digest=sha256:d06a58692cccc05d22f20003cbf948087a11107e0326e07b7dbe6dd4b931f123

Observation 4ba7dc5a-5a45-4db6-ab8b-bca1f4cf824c · outbound

This paper cites SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning

Reference 1970

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.657712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.657712Z digest=sha256:31569ef56d0824c8e6c21ddeae556a8460dab6465e9ebcdf4bcaccef0eddac76

Observation 03bfdfd2-fe44-4ea7-9718-884314e52c63 · outbound

This paper cites Reinforcement Learning via Fenchel-Rockafellar Duality.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Reinforcement Learning via Fenchel-Rockafellar Duality

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.650706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.650706Z digest=sha256:5032afa77beee26887d335778adacae11ef8169b3ff252e61ac8f3b4b4c46eef

Observation b2bb28c8-75ff-49e2-8b63-09e93078fd93 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Off-policy deep reinforcement learning without exploration

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:39.003582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.612678Z digest=sha256:a4983af281f0cbbec1c98017aaf8bdb6d64c74d3efc5ed17ddb5160dfc167e7c

Observation dc791905-98fc-49e8-b565-b45c55949573 · outbound

This paper cites Harnessing Mixed Offline Reinforcement Learning Datasets via Trajectory Weighting.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Harnessing Mixed Offline Reinforcement Learning Datasets via Trajectory Weighting

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.620256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.620256Z digest=sha256:6d4e0f151902624a9740da41be11c8d96149824969a046aca97d3a7dc1874a5e

Observation bf7dd64b-96ab-44cc-84cf-2c8c0f035a0c · outbound

This paper cites Residual algorithms: Reinforcement learning with function approximation.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Residual algorithms: Reinforcement learning with function approximation

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:39.014454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.608735Z digest=sha256:77144aef042c45e08805476809c13135ad8da0f5970d80954664aea6ecb8397c

Observation 83c1896e-56f2-4494-997c-83046a098000 · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.654314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.654314Z digest=sha256:9b067b653c369a3cd7c22c9f892a83dcb661d3f4de9c420c783c69cf7f5853be

Observation adc30dcd-a736-4d31-8c6a-aa2bc058f138 · outbound

This paper cites Dealing with the unknown: Pessimistic offline reinforcement learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Dealing with the unknown: Pessimistic offline reinforcement learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:17:38.992654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T12:17:38.635338Z digest=sha256:50d5f8809e33535c45618eae0358a640237e2374444dcf188dad658b2094ff02

Observation 62368e0a-39bd-40c5-b908-d52c4694333f · outbound

This paper cites SelfBC: Self Behavior Cloning for Offline Reinforcement Learning.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning SelfBC: Self Behavior Cloning for Offline Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.638637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.638637Z digest=sha256:fdbc5e262eb3ad08435ad176744b4c67f0498a16c98aa92143b14ca580ea438b

Pith citing papers

Observation 76e4dfae-42b7-4f45-8e16-afb9d4a6634c · inbound

Semi-gradient DICE for Offline Constrained Reinforcement Learning cites this paper.

Semi-gradient DICE for Offline Constrained Reinforcement Learning An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:15:33.500556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:15:33.500556Z digest=sha256:a9cd7609dec664c0bd3ed04a55d790f9011a6a7b367026871e487ec4331fb42d

Observation c6a1765c-5ccd-42f4-9528-f69f63c67c47 · inbound

Dichotomous Diffusion Policy Optimization cites this paper.

Dichotomous Diffusion Policy Optimization An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.831475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.831475Z digest=sha256:9f9c6b06f86544d2a1e05d4609fda0bc4548590c4f8c511907274fd1b2851301

Observation 48d266f3-efc1-4b59-853c-1d004699c332 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.626504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:2e10fbc92aff881687972f8a361a66b74923aef6bc2229e69aaf32b0894b0eb7