Pith. sign in

Paper Citation Record · LEDGER

Score Regularized Policy Optimization through Diffusion Behavior

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.07297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.07297 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.067287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.357585Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d21a81c5-2540-4161-b310-5a0870610646 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.479870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:6e87e158e6e62994dd24ab418d03312b429dfc7674afd185471740f4f9b85c88

Observation f7d77ab3-7008-4262-a7e8-eb0341d5621b · inbound

Offline Reinforcement Learning with Penalized Action Noise Injection cites this paper.

Offline Reinforcement Learning with Penalized Action Noise Injection Score Regularized Policy Optimization through Diffusion Behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.067287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.067287Z digest=sha256:06d9a24e565b77e763187c55b7aa5de09bb80dafc40a11be2af39df8c554cc0f

Observation d6fb52d4-c22f-446b-b8bb-f4bf0807ccaa · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Score Regularized Policy Optimization through Diffusion Behavior

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:2769ebb2148cfcc9b156ccc8912acaa78df3e03ecb207882b27c7d9e6ad5bfff

Observation 1d192dac-cf14-4ec6-9a65-16360f99cb38 · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Score Regularized Policy Optimization through Diffusion Behavior

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.558176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:52b75e0de1dd9c039532746c76ce80f9c1662391d8ca09dbdcced16fdaedea7d

Observation f5beaba2-02ee-4f9c-a0dd-0ae027d9e460 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.773020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:43f14fc8af539ae31dbd1eee1a6e831208c4d9adc50d072f6b9a3a79a8002f60

Observation cb921262-cf60-4f3b-8af1-a231886ac316 · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:48.333368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:cc535395d20866cc3579da72ccd695f407b71096a26689813d92f3778ffb9190

Observation db9e2e38-2a12-4fde-a37c-0685532fee09 · inbound

Path-Coupled Bellman Flows for Distributional Reinforcement Learning cites this paper.

Path-Coupled Bellman Flows for Distributional Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:16.247181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:12:58.528130Z digest=sha256:bc5c045aa3bf46561c10065701e21b087a44b50140bfec852de90adf9bd083eb

Observation 3dfc10f2-7f7f-4631-bedb-5c612630c352 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Score Regularized Policy Optimization through Diffusion Behavior

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:43:30.999353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:ccf164bad8d3c750d319087db16671f62d306b07e8d1f8c6034d3e98250d5e9c

Observation f1b200e0-2084-477a-9a5a-43de682006b0 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Score Regularized Policy Optimization through Diffusion Behavior

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.397932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:7dd313ae7da4facddfcc7c70dd1cdb87f9dd55537fa5afb9535fb516218a2801

Observation dc086fae-516b-4571-ac7b-7ddde9870d3e · inbound

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning cites this paper.

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.359291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T13:57:31.180928Z digest=sha256:3a89fd12279f0b473f804047b521ae45ea6e379b1ab5258a7e4ceb284cd7150a