Pith. sign in

Paper Citation Record · LEDGER

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2303.15810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.15810 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:46.427182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.940883Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 369249ae-67b4-40de-85d7-24c8df8067e2 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T13:48:36.611406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:c2f4b205f2838cd9af11283a36bafc7d0d6dc564988a3f8854e0eaa2f5329c51

Observation 5795690e-4847-4996-ba18-6a6adf07d008 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.427182Z digest=sha256:e571b838f7a3909a6016a608350dfb5cf887548fde3e9788b28aeb0a412fcb13

Observation 0c67fefc-1819-42ba-9415-5af47f1e6ea1 · inbound

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning cites this paper.

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:52:15.285443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:48:28.980868Z digest=sha256:402515a30b4b64a9b15602b5bd82b90e5124860c922871ecb98d75c246a4813b

Observation a73e2a5f-1b28-4238-8d50-5f4e3eb823c2 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:27.299725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:27.299725Z digest=sha256:915aa4062da3c0c0b5ef3f25225e5e06dee5a2e11452a42f1e342eb3f6f1d563

Observation c79745c4-080c-4a8b-9206-447d4de34d79 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:069dd58273e1584977ab2faad596f2816c28e91f6fde81092fa1bbf653bc7deb

Observation 93f16d80-24d1-46ed-b157-e537401a1850 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.622646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:6f817e255741a5920c9f01ece14fff25d25071e6e634256b48acdcb9fd47803b

Observation 9b03b8b9-acfa-4d1a-92c4-483729c5c34a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.618145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:6a0dd00722a21b46c244a66c283b3d9a65cf73cbf390b7fb4d6e14b2e874f21f

Observation 48544bef-20c6-44af-af27-d789c3d80a4b · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.631554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:4b22cff3fdef4d2d3290429f62378c0b9add634dca42ac7d400ce475b2dc8544

Observation 38d69b95-3dee-4ae3-ae56-96d2a6576225 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:45.840469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:fecc12d9dade0ebfa00c07205775f7ea8d56060887d1ea66fdbeed98b992f1dc

Observation 8b6381b3-ad22-4a49-82f8-610a1812c1e6 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.990694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:6d56a6230d696c47dea7da01ba392c08c75afa38c066a2dd3ba589e8fbf67c9e

Observation 1cf9ae9f-5987-4cd4-9237-5260c287e35d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.942518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:5ca2a6fc9da21e3c9a66f080db77ac5a7a1a751195b9d24a47504dcec3e8ccb2