Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.393031Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2506.12811.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.393031Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:16:11.130271Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:19:37.801846Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80413add-a181-4e55-9b7a-6e83205e7cef · outbound
Flow-Based Policy for Online Reinforcement Learning A markovian decision process.Journal of mathematics and mechanics, pages 679–684, 1957
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0c6a2b-544c-4413-98e7-e7294918d8b6 · outbound
Flow-Based Policy for Online Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc5feda-359b-423e-9737-d00e0c3f5221 · outbound
Flow-Based Policy for Online Reinforcement Learning SE(3)-Stochastic Flow Matching for Protein Backbone Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c287922-e198-44e5-8940-a6a508c11558 · outbound
Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef08daae-3b0b-4451-8425-2ea98617bb5b · outbound
Flow-Based Policy for Online Reinforcement Learning Neural ordinary differential equations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31978adb-d075-48f7-854b-74808710fc45 · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a40291-4d40-46d6-9601-55cf0b6294b0 · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f1818f-1dcf-4766-a046-14487dd8c0bd · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d9509c2-c3b1-49a4-a133-df57efb2fe73 · outbound
Flow-Based Policy for Online Reinforcement Learning Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02ead63-07d8-4762-a248-384541f0024a · outbound
Flow-Based Policy for Online Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15083ea9-9872-4702-8804-b9f3787aad9a · outbound
Flow-Based Policy for Online Reinforcement Learning A minimalist approach to offline reinforcement learning.Advances in neural information processing systems, 34:20132–20145, 2021
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c959ac2-172f-4f83-9b2e-2015d712f9af · outbound
Flow-Based Policy for Online Reinforcement Learning Addressing function approximation error in actor-critic methods
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c705be-1d97-4367-a78f-dbb4d2de7100 · outbound
Flow-Based Policy for Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3827596-99a4-4a14-9fce-efbbcb812d70 · outbound
Flow-Based Policy for Online Reinforcement Learning TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f9fb3e0-761f-4815-824b-d8207e8b40d6 · outbound
Flow-Based Policy for Online Reinforcement Learning DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae535516-f9c2-4d21-8c61-7b6ce41b57c3 · outbound
Flow-Based Policy for Online Reinforcement Learning Denoising diffusion probabilistic models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b4f240-ded0-4118-bfef-25dae157c7d9 · outbound
Flow-Based Policy for Online Reinforcement Learning AlphaFold Meets Flow Matching for Generating Protein Ensembles
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee6799d-6930-46f9-bf39-28a7b8d6d488 · outbound
Flow-Based Policy for Online Reinforcement Learning Efficient diffusion policies for offline reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b2ecf5d1-dff0-4a92-87bb-79fa9c817978 · outbound
Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff883579-478f-462f-85f6-075c11acc940 · outbound
Flow-Based Policy for Online Reinforcement Learning Stabilizing off-policy q-learning via bootstrapping error reduction.Advances in neural information processing systems, 32, 2019
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5767e56e-2700-4e23-ab8c-2e2619586369 · outbound
Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5ddba2-7544-404d-8157-2145da7564d2 · outbound
Flow-Based Policy for Online Reinforcement Learning Flow Matching for Generative Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d2fc83-3023-49ca-9ec2-322a9f364cbf · outbound
Flow-Based Policy for Online Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5559add1-5eab-4a63-a9f3-207b008c0301 · outbound
Flow-Based Policy for Online Reinforcement Learning Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05275907-e1bd-401e-ba72-8a83ef446b96 · outbound
Flow-Based Policy for Online Reinforcement Learning Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451216f4-161f-49d8-b82f-54a353574301 · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f19a43-4b22-43f6-9662-5bb718a1ff0d · outbound
Flow-Based Policy for Online Reinforcement Learning Self-imitation learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6664bdf-b841-4192-9764-287892c1f78d · outbound
Flow-Based Policy for Online Reinforcement Learning The difficulty of passive learning in deep reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c089d76-81fb-4215-aa34-7ef7fee7e4dd · outbound
Flow-Based Policy for Online Reinforcement Learning Is Value Learning Really the Main Bottleneck in Offline RL?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5376564-e237-4d21-84e6-5269287a3572 · outbound
Flow-Based Policy for Online Reinforcement Learning Flow Q-Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b82270a-0f0d-4c66-a73e-b4ae207c82f3 · outbound
Flow-Based Policy for Online Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa8bec0-d654-4340-9d63-c84b28a59fbd · outbound
Flow-Based Policy for Online Reinforcement Learning Reinforcement learning by reward-weighted regression for operational space control
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · outbound
Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9e0f0b-acce-4343-af7e-f9f6d3418e63 · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion Policy Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c40c83-3c3b-49ea-acb3-7b3e3307e811 · outbound
Flow-Based Policy for Online Reinforcement Learning HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df987991-127d-473b-a2b7-c300bbf8f0a8 · outbound
Flow-Based Policy for Online Reinforcement Learning Deterministic policy gradient algorithms
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6191128b-e248-4af0-99fa-4a7ef1b68612 · outbound
Flow-Based Policy for Online Reinforcement Learning Score-Based Generative Modeling through Stochastic Differential Equations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bb8b3c-75fa-437c-b6aa-14c2b38539ab · outbound
Flow-Based Policy for Online Reinforcement Learning Reinforcement learning.Journal of Cognitive Neuroscience, 11(1): 126–134, 1999
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca32db8c-638b-4247-8b9d-8dbecc0ca744 · outbound
Flow-Based Policy for Online Reinforcement Learning DeepMind Control Suite
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1637be-3421-46af-80c1-20a836495baa · outbound
Flow-Based Policy for Online Reinforcement Learning Improving and generalizing flow-based generative models with minibatch optimal transport
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc033cc0-5e72-42ef-a15e-3b9b41de0543 · outbound
Flow-Based Policy for Online Reinforcement Learning Springer, 2008
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b9530c-34af-4332-a469-378be282eb02 · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43cc9353-aed6-4360-8a61-8545b6fe35fa · outbound
Flow-Based Policy for Online Reinforcement Learning Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bff938-c31d-4619-9847-279cfb70f603 · outbound
Flow-Based Policy for Online Reinforcement Learning Policy Representation via Diffusion Probability Model for Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9093019-2300-4f7a-b23c-d42f708e8bf6 · outbound
Flow-Based Policy for Online Reinforcement Learning Energy-Weighted Flow Matching for Offline Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f2938f-aa5c-4a56-b8e3-6a47e1da3c77 · inbound
What Does Flow Matching Bring To TD Learning? Flow-Based Policy for Online Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7a5816e2-7622-49d0-9370-8686d0084056 · inbound
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data Flow-Based Policy for Online Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1274de69-c92b-4002-9014-8ce57310826e · inbound
Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Flow-Based Policy for Online Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 53d83936-0ddb-46c2-85f7-48147fa6f40b · inbound
ReFPO: Reflow Regularization for Flow Matching Policy Gradients Flow-Based Policy for Online Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 450a06e7-e743-494f-bce1-df4265c25f7d · inbound
Dual-Flow Reinforcement Learning with State-Aware Exploration Flow-Based Policy for Online Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.