Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2312.11752.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.346736Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:17:36.937596Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 16c179a2-3336-4052-94c8-4f91798adf1a · inbound
Diffusion Policy Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 07e6a49b-df04-46df-a177-0ec39454bb38 · inbound
Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b399cbd0-d316-495e-9e4b-ca909ce2e3a7 · inbound
Generative Diffusion Modeling: A Practical Handbook Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b3d552-8eee-4bcf-b728-92f8aaadc188 · inbound
Efficient Online Reinforcement Learning for Diffusion Policy Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c18d38-9d6d-48d8-8cfc-eb8af3de670c · inbound
Habitizing Diffusion Planning for Efficient and Effective Decision Making Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42042865-f0ce-403b-a6f5-260f47240610 · inbound
Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb614448-9894-4ff5-914d-2d667dd83192 · inbound
Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · inbound
Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856d2f9c-7937-4712-9664-e6fde3acddc0 · inbound
Steering Your Diffusion Policy with Latent Space Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f35fc052-b97b-48f5-a52d-0b7a802fb943 · inbound
EXPO: Stable Reinforcement Learning with Expressive Policies Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 307a410a-5ad7-4481-83ff-e91af1ade30c · inbound
Flow Matching Policy Gradients Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7318975-2a5e-4318-ba51-97ae2e6af7ab · inbound
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c4faa7-5cc4-4405-9729-7c73890d20fd · inbound
How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65ae67e8-1760-4953-a541-4addab9581ab · inbound
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba5bfc9-19c5-4378-97af-882b25215429 · inbound
Reinforcement Learning via Value Gradient Flow Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cdef4ab6-d0d9-4680-9c2b-775c94be44c3 · inbound
Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35a20015-daaf-4adc-843d-4d1405965043 · inbound
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7782a234-9094-4d3b-9cda-6b1745b293cf · inbound
Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation da7e45d9-72c8-4795-bf24-8f1fcf9b86a8 · inbound
Adversarial Dual On-Policy Distillation from Expressive Teacher Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ba195ea-06a1-4644-a323-f277eed15319 · inbound
Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1114145b-2ae2-47de-a7e7-4574ea8f6ad3 · inbound
GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6886e5a8-f91d-4a62-829d-c7acab63fbe4 · inbound
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 560ffcdb-62ca-4960-95ce-b0b3da32b78b · inbound
A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131fc597-58d5-4fb6-a993-89847e6b741a · inbound
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.