Pith. sign in

Paper Citation Record · LEDGER

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2312.11752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11752 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:44.592129Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.937596Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 16c179a2-3336-4052-94c8-4f91798adf1a · inbound

Diffusion Policy Policy Optimization cites this paper.

Diffusion Policy Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.975258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b32d746867104855cb5380f99e538944725afdf23ecd76fe75b292285c50cdfb

Observation 07e6a49b-df04-46df-a177-0ec39454bb38 · inbound

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation cites this paper.

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:48:30.730504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:48:30.730504Z digest=sha256:511050e1893f32977460ebbf4dd245dacae7aaef68836cab64adf991c08ad773

Observation b399cbd0-d316-495e-9e4b-ca909ce2e3a7 · inbound

Generative Diffusion Modeling: A Practical Handbook cites this paper.

Generative Diffusion Modeling: A Practical Handbook Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:49:35.453168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:49:35.453168Z digest=sha256:da0152dd0927d92b84570b6ada34e2bc642a6e160805b55338c42220eaab836f

Observation f5b3d552-8eee-4bcf-b728-92f8aaadc188 · inbound

Efficient Online Reinforcement Learning for Diffusion Policy cites this paper.

Efficient Online Reinforcement Learning for Diffusion Policy Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:28:41.908138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:28:41.908138Z digest=sha256:a03cfaf3dd1b5b118f5170a76ef1ff033fb9c413379087ec3c7e08c6edbae2fc

Observation 26c18d38-9d6d-48d8-8cfc-eb8af3de670c · inbound

Habitizing Diffusion Planning for Efficient and Effective Decision Making cites this paper.

Habitizing Diffusion Planning for Efficient and Effective Decision Making Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:40:16.106279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:40:16.106279Z digest=sha256:c70647b3de7064b1fd65bf6534789ff7905c97be05bb56a6d85662396dfa396e

Observation 42042865-f0ce-403b-a6f5-260f47240610 · inbound

Exploratory Diffusion Model for Unsupervised Reinforcement Learning cites this paper.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.374992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.374992Z digest=sha256:5ea59794076804f1a5e781a31588b048ced683267888ee55aaa118bab2b1d920

Observation cb614448-9894-4ff5-914d-2d667dd83192 · inbound

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion cites this paper.

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:09:51.254192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:09:51.254192Z digest=sha256:15df16f760381ff30ab5e998c3ac5ffed2dc5df73383f73798bcad19ccfcbd3e

Observation 2821fadf-5e2e-48bb-86da-83d8fda249a7 · inbound

Adaptive Diffusion Policy Optimization for Robotic Manipulation cites this paper.

Adaptive Diffusion Policy Optimization for Robotic Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:44.592129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:00:44.592129Z digest=sha256:9794104917eea1adf4acf4655e2329b5751eec7f20e59185ecc0af7e4dc252e3

Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · inbound

Flow-Based Policy for Online Reinforcement Learning cites this paper.

Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.346736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.346736Z digest=sha256:91b9363e92a67b18df4a9f6f90b70972ff7c8eb24b731573b81d0d5ebc7cf6bb

Observation 856d2f9c-7937-4712-9664-e6fde3acddc0 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.490841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:7887754b6bdc221765415670d43245cc52f6eb96a299fb98ad2b3c8234818073

Observation f35fc052-b97b-48f5-a52d-0b7a802fb943 · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.226296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:c81041e790a72b3a116d6c45af7cd9e374d7c3129f85c142c2498ccda1735c1f

Observation 307a410a-5ad7-4481-83ff-e91af1ade30c · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.645809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.645809Z digest=sha256:eaeceeb3f9d77914d6d650364ac2b83e3c94a2f801010eaa5423212fd2d76ccd

Observation a7318975-2a5e-4318-ba51-97ae2e6af7ab · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:09.827237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:09.827237Z digest=sha256:adaf652dcd4c10f1b587faadf6231cb5d356e9df58c9d3c2afa3f55ab60e352e

Observation 47c4faa7-5cc4-4405-9729-7c73890d20fd · inbound

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? cites this paper.

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:57:33.094604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:55:31.706717Z digest=sha256:e141136a65bfbe776fc17382ee0acddee7323ac3346974ebed770c63769b41cb

Observation 65ae67e8-1760-4953-a541-4addab9581ab · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:6072a145d37f8048d4add63eba85b5da2989944343e431c70e54f72196d6db7d

Observation 7ba5bfc9-19c5-4378-97af-882b25215429 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.680725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:70a61407b5d14776cb36aa1d549cc37e7bd37bf7a32ba7d10921567c691e6937

Observation cdef4ab6-d0d9-4680-9c2b-775c94be44c3 · inbound

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning cites this paper.

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:55.078354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T01:08:17.624506Z digest=sha256:250600627f05dccf8ffc2078ac69ffb856709ac96a1d210a8844f233fb8bff8f

Observation 35a20015-daaf-4adc-843d-4d1405965043 · inbound

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing cites this paper.

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:18:59.880117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T20:15:44.030714Z digest=sha256:3f7776f55f64b0e70035ac89dbe34395b111be6ae6715b0a5056c64745945e54

Observation 7782a234-9094-4d3b-9cda-6b1745b293cf · inbound

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization cites this paper.

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:01.348420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:46:27.179341Z digest=sha256:a3dedbaedcc5171580d9555374c960a765a2372d76ae01d34f0b9b15dab26180

Observation da7e45d9-72c8-4795-bf24-8f1fcf9b86a8 · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.524120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:591f817cdb4efee791e8adb9e7d08de1ea850b08162a108bc0068f9bc2891e5a

Observation 1ba195ea-06a1-4644-a323-f277eed15319 · inbound

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance cites this paper.

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.577987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T07:21:56.382343Z digest=sha256:03effb6852006a328f8b1ac3a84cbb5f541275d3fdf92fdcd5fe69ee7a1b99a7

Observation 1114145b-2ae2-47de-a7e7-4574ea8f6ad3 · inbound

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cites this paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:08.545297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:47:10.062975Z digest=sha256:dfdbde2b671244bfea60938535d21097663e2c35d57fa7bd76c2f5e8f39ea525

Observation 6886e5a8-f91d-4a62-829d-c7acab63fbe4 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.939467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:a366fc416bd1b66456954a533b8fb6b19adceb9619a2d8281b5ea70b1981eb4a

Observation 560ffcdb-62ca-4960-95ce-b0b3da32b78b · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:929c3d0570ebe84a1d3aaf362fcba7efe4a5e4be298471e6e46853e8543d986f

Observation 131fc597-58d5-4fb6-a993-89847e6b741a · inbound

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners cites this paper.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.881499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.881499Z digest=sha256:3981e594c04326e17fbaafd43a86eef2f7e318cf9ba2a35a4480a61f28f97f2f