Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:30.265947Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 34 inbound Pith citation observations for arXiv:2505.23458.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:30.265947Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:42:22.223628Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.122193Z
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 274acc93-49ec-4872-82e3-7fc79ccf1802 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies for out-of-distribution generalization in offline reinforcement learning.IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e25982c-fdbb-4d15-a09c-5f11796d28ca · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Building normalizing flows with stochastic interpolants
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c8a5edd-a8ff-4ab6-bfc9-6f2bd24dac69 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Uncertainty-based offline reinforcement learning with diversified q-ensemble
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d631dc44-c4c3-4791-a139-0e2f163d7972 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Hindsight experience replay
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d261bb7-52e2-42b6-9e15-7cd904e73583 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bcf97e-f0ce-4eb1-80b2-b42e8b204c9b · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Zero-shot robotic manipulation with pretrained image-editing diffusion models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 714c8e21-4f4a-42e1-88a9-710d62b6844c · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Whitney, Rajesh Ranganath, and Joan Bruna
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8d9265a-1c4c-48e5-9356-e219f334da01 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning via high-fidelity generative behavior modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74bde7e8-c041-4700-92e2-603a020bc4e2 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Score regularized policy optimization through diffusion behavior
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 802a466e-2092-4a69-b502-dbd8eef22ece · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Aligning diffusion behaviors with q-functions for efficient continuous control
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31dbe662-30bc-4e13-a288-4e83934ed286 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Abbeel, A
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e5c6a4c-cd75-47ac-a79d-c4ac3e03c3fc · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies creating a trust region for offline reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f99072af-9eb1-449b-aa3c-70647f1bfc58 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy: Visuomotor policy learning via action diffusion
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642ed821-922b-4b5d-a86d-7d811cc1397f · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator da Silva
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6660750a-a6d4-4d59-a499-5734479665b2 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Using expectation-maximization for reinforcement learning.Neural Computation, 9:271–278, 1997
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6baffa69-bd7b-42ef-9032-355ef9e9716a · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-based reinforcement learning via q-weighted variational policy optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2ae4034-f701-45e1-b67f-ce913f910fe1 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Consistency models as a rich and efficient policy class for reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dfa2dba-d65f-4a2f-b469-1a53c1577de1 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations (ICLR), 2022
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c051a217-5a4e-4391-aacb-46e6b944f61f · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Imitating past successes can be very suboptimal
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5256e09e-c9a8-4162-a405-424367c63127 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive learning as goal-conditioned reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab7d341c-fbe0-4bb5-bcbe-8476cdff4e70 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69bb8e44-52f8-4eaf-8a65-28adba2ab38f · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator A minimalist approach to offline reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d237293-255e-4cc9-bb01-6abf487dfdd2 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Addressing function approximation error in actor-critic methods
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a683e37b-e065-4cb8-a17c-da924fc5e915 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Murphy, and Tim Salimans
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00d8e094-27f7-4188-ae32-abe1b12d04a9 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Extreme q-learning: Maxent rl without entropy
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a58fd3f-0340-408a-aa74-34bfd3a72932 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Learning to reach goals via iterated supervised learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f936e271-bb74-451b-9a60-626cda99cec8 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Closing the gap between td learning and supervised learning–a generalisation point of view
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b890cd14-cc24-450f-8c7a-cb04512583e1 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Explaining and harnessing adversarial examples
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af6ac9e-9dba-4833-81fe-212a580d3a1a · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a609deab-73a5-4a8b-8378-70e300f433dc · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80827a7-1c30-4b8b-b0e5-1a18224f6796 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b45d15d-891e-4072-840a-4d8ea3a004fd · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Aligniql: Policy alignment in implicit q-learning through constrained optimization.ArXiv, abs/2405.18187, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cdc18df-65ac-4108-9745-750ac70b7148 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Gaussian Error Linear Units (GELUs)
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53feb01d-6d0a-48da-8d91-cd9763944c6d · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Classifier-Free Diffusion Guidance
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b83f3ae-92e2-49cb-86b3-f8b5cf6b2d18 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Denoising diffusion probabilistic models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79375b08-a493-42d7-b055-4c6e4e25eb82 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Tenenbaum, and Sergey Levine
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04dfc25a-5712-416c-ad4c-bb68502e31b6 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Efficient diffusion policies for offline reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0a23421-8d9a-4410-a78f-015dc9ea282a · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Kingma and Jimmy Ba
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1d601c-2551-4ac9-8bc4-695c4a5526b3 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning with implicit q-learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b2c0ff-aedb-4e00-8e0e-01f2ac6f42ab · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-conditioned diffusion: Offline rl via generalization.OpenReview, 2023
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b68cdd13-872c-4b13-9aac-37ecd00af636 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Reward-Conditioned Policies
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab444d0-758f-4913-84b3-886758f587d9 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Tucker, and Sergey Levine
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54fb33ee-57a4-4ae5-ba20-ccd08657ac2d · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Learning multimodal behaviors from scratch with diffusion policy gradient
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f14ddcfc-602d-4fa8-96ae-14e7d2bddbfb · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Lillicrap, Jonathan J
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7778eff-81cd-448d-adc2-3761d12556a8 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Flow matching for generative modeling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40be38c3-64eb-46d2-be15-b5110bc091d2 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Flow Matching Guide and Code
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15fd0fc-e658-40a5-9f74-445ad8470372 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Flow straight and fast: Learning to generate and transfer data with rectified flow
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4ac65c-1a4a-45fd-89c7-228952e56585 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 186a8a5e-bffe-49d5-8301-714d5d616d02 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Learning latent plans from play
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b39392f-0dca-4393-9244-8e326a2cd847 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad88bfd8-f5c4-412d-b5d6-4f3426b19802 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Mish: A self regularized non-monotonic activation function
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc9e6f52-cd4f-4fc8-b454-080ea9473496 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63304ea1-0b86-4062-ae8a-489c57bf65e0 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Anti-exploration by random network distillation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aff38423-441a-48b5-a61c-d9a8530723e8 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Is value learning really the main bottleneck in offline rl? InNeural Information Processing Systems (NeurIPS), 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8f0d878-ac34-4dda-bbd3-57460468c85f · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Ogbench: Benchmarking offline goal-conditioned rl
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c62e4a-7ec3-413d-b363-d7a933bb5ccd · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Flow q-learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c802dd2-e7f0-4e3a-a7c6-b8d2d3ce315d · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6fe649-1620-49cb-9503-9196a9bba391 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Reinforcement learning by reward-weighted regression for operational space control
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea3880fc-0749-47e5-bbe3-6c9b3a1ad795 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Learning a diffusion model policy from rewards via q-score matching
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31f9eb45-0ce3-4031-be9e-a5a4fe774f09 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy policy optimization
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddb9c88b-0d86-4d36-9e6d-b96d693575d6 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Trust region policy optimization
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5794becc-5d9e-4908-a0c2-dc1fb1b1ee4d · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Proximal Policy Optimization Algorithms
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63e00e8-fe76-4f37-977e-baef64f090e5 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8c70b07-4674-4b80-ab56-1f72b7ca04a8 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Deep unsupervised learning using nonequilibrium thermodynamics
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae689195-a547-4cfa-a0ef-fcc6fba41660 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Generative modeling by estimating gradients of the data distribution
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5871684f-043c-42a1-a6d1-2b1e394593d7 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Sutton and Andrew G
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b9578f0-d2a8-466d-8172-62933aaaca77 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Policy gradient methods for reinforcement learning with function approximation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e82727e7-5032-43ee-9501-5bc00cdd03b7 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Revisiting the minimalist approach to offline reinforcement learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8204b323-f5b4-4fe3-a956-2fbf79d60459 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Learning one representation to optimize all rewards
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44ea1e02-a513-4732-a729-ed5ce1f0510c · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies as an expressive policy class for offline reinforcement learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 594f885e-0130-465b-a3d4-70f16a80e98d · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7adf455-ba68-41b8-ab50-8952b9fb234c · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Behavior Regularized Offline Reinforcement Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afcc1f24-a588-4efc-ba1a-105e0309a826 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Offline rl with no ood actions: In-sample learning via implicit value regularization
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 481580db-6b44-4321-8133-98cbe73ed38e · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 749c8f5a-2bf8-4433-b336-ba5affe8f977 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Representation via Diffusion Probability Model for Reinforcement Learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 062be4fa-d7d1-4753-afa9-7000e4ee3e37 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7214780-13b5-4915-a75a-6de01e8e131c · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd15f4a3-d5a5-42db-ba9d-5a09e4d6c2c5 · outbound
Diffusion Guidance Is a Controllable Policy Improvement Operator Energy-weighted flow matching for offline reinforce- ment learning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56e107ec-6e8c-4c85-a1b2-0b52646da785 · inbound
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e09793a3-be45-46dd-ab6f-0509b118ac7c · inbound
DiffusionNFT: Online Diffusion Reinforcement with Forward Process Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c674f213-37de-4d0d-bacb-bec5bad07779 · inbound
$\pi^{*}_{0.6}$: a VLA That Learns From Experience Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21705f32-f9ce-49ff-a85b-0490e878683c · inbound
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba2875a-4da8-40c8-be8f-2d6167778d44 · inbound
Dichotomous Diffusion Policy Optimization Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b108cbb-24e0-4cb4-a360-d36d9a30a11f · inbound
RISE: Self-Improving Robot Policy with Compositional World Model Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c50286fe-3ad8-49ce-8090-0f3f49962bf5 · inbound
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6feff226-b30e-4dc5-ad74-e425c919bf0f · inbound
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dae9f6f-7c72-481c-868b-5f8291ec48e8 · inbound
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c53f7d-15fd-40a5-b96b-befcb10abac6 · inbound
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99916ee1-1754-44d0-ac18-3fa2ad6c56d5 · inbound
Update-Free On-Policy Steering via Verifiers Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3c10a8-86bf-4b35-95a9-20bf1f2df124 · inbound
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a92b01-d2b5-4d81-a199-b009fea7e531 · inbound
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116dcf5c-f634-4dfc-a916-2a6d8186b23f · inbound
Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 751150a7-56e0-4b66-bdbd-90071755782a · inbound
Reinforcement Learning via Value Gradient Flow Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b4e3dca-8494-4ee4-a417-c80a3138ea27 · inbound
Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d03a8337-83c4-4fc6-99fc-a10f2e1ca7cb · inbound
Refining Compositional Diffusion for Reliable Long-Horizon Planning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 160a7291-e759-4891-abc5-4bf75cc0c274 · inbound
JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a00fd679-7e6d-4d2a-ac68-b370845d1fa2 · inbound
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dd19e58-12e4-4496-a756-b80bf09c29d3 · inbound
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f04b9ec-b5c1-47d0-bfe7-c1b9dbe2fb5f · inbound
Scaling by Diversified Experience for Vision-Language-Action Models Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52822f58-c9ec-4ed4-8971-7b239830740f · inbound
DexPIE: Stable Dexterous Policy Improvement from Real-World Experience Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf1732f0-4678-40d3-8402-69278e232546 · inbound
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b5bc7f2-3940-4368-a58f-6b6c11111dbd · inbound
Improving Robotic Generalist Policies via Flow Reversal Steering Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee17dbea-1e62-424d-b08f-aa196c8abe32 · inbound
Reversal Q-Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 423458cd-c8b6-4838-9d59-e3ce23867c09 · inbound
Robot Self-Improvement via Human-Video Dynamics Models Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4df7f6ef-7196-4271-bc01-16b0a71e1671 · inbound
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac60c372-34e8-4bca-9091-38a5ec11690e · inbound
ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da2aa04d-e3cb-46c6-a4ee-83dbef6125dd · inbound
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9f5b974-9bec-4850-9588-93545f270202 · inbound
Controllable Sim Agents with Behavior Latents Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac80c2fb-d7eb-4c50-935b-3f66ea3d10e2 · inbound
TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2819e81-cf32-47a0-9c56-7434a032f4f6 · inbound
VINE: Taming Generative Control Policies for Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08ee7e7-79ca-4e0b-ac81-af9759d37e18 · inbound
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3b3865-8852-4d00-93fa-510cccdb0c58 · inbound
CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.