Pith. sign in

Paper Citation Record · LEDGER

Flow-Based Policy for Online Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2506.12811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12811 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.393031Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:16:11.130271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:37.801846Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80413add-a181-4e55-9b7a-6e83205e7cef · outbound

This paper cites A markovian decision process.Journal of mathematics and mechanics, pages 679–684, 1957.

Flow-Based Policy for Online Reinforcement Learning A markovian decision process.Journal of mathematics and mechanics, pages 679–684, 1957

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.214500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.214500Z digest=sha256:d0b8b3d9e9faa60378691aa3c31572a0c038c8a1e48ecda4fef03b281e86e4fa

Observation 9b0c6a2b-544c-4413-98e7-e7294918d8b6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Flow-Based Policy for Online Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.218998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.218998Z digest=sha256:e824b8d51bf04a8f039c06131e7fca72877ddde43240824d7f15130c1d09fae9

Observation 6cc5feda-359b-423e-9737-d00e0c3f5221 · outbound

This paper cites SE(3)-Stochastic Flow Matching for Protein Backbone Generation.

Flow-Based Policy for Online Reinforcement Learning SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.225162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.225162Z digest=sha256:359066953aae8673f32746dacf0a0ed0263763fdf51aa5d15cfdbbc81ccf5c0d

Observation 9c287922-e198-44e5-8940-a6a508c11558 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.230392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.230392Z digest=sha256:9d3d23d0760456093d81f5a7444a195da71ba117983475c30870236147a0f667

Observation ef08daae-3b0b-4451-8425-2ea98617bb5b · outbound

This paper cites Neural ordinary differential equations.

Flow-Based Policy for Online Reinforcement Learning Neural ordinary differential equations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.234827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.234827Z digest=sha256:1fee4e3350e254b07c8096655f3316f5d2c0ea3fc900a875b178f7e7786965b2

Observation 31978adb-d075-48f7-854b-74808710fc45 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023.

Flow-Based Policy for Online Reinforcement Learning Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.239064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.239064Z digest=sha256:5f760109a97427a6cd5948b6c56ef1f6c4f7a125a3b285f5c13fe59c396d1e3c

Observation e6a40291-4d40-46d6-9601-55cf0b6294b0 · outbound

This paper cites Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021.

Flow-Based Policy for Online Reinforcement Learning Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.243560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.243560Z digest=sha256:06a7ac5aa8a86a7402945fd6312540a7afb60885c929eab9a2a18af3e462900b

Observation 37f1818f-1dcf-4766-a046-14487dd8c0bd · outbound

This paper cites Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization.

Flow-Based Policy for Online Reinforcement Learning Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.248611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.248611Z digest=sha256:ef97394e2699db78d5995ab4c62d68600ec84f867c939480b6fa170aaa871c98

Observation 2d9509c2-c3b1-49a4-a133-df57efb2fe73 · outbound

This paper cites Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.253038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.253038Z digest=sha256:b8a3861c4010c0d299f1dec7a42c5f8efa410faf8bcf89e9485b7fdd602799f4

Observation d02ead63-07d8-4762-a248-384541f0024a · outbound

This paper cites Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization.

Flow-Based Policy for Online Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.257162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.257162Z digest=sha256:92e6fa8524ad2f84c7373564848e613748bb4ad990f683a081ed4cc53a5d71f9

Observation 15083ea9-9872-4702-8804-b9f3787aad9a · outbound

This paper cites A minimalist approach to offline reinforcement learning.Advances in neural information processing systems, 34:20132–20145, 2021.

Flow-Based Policy for Online Reinforcement Learning A minimalist approach to offline reinforcement learning.Advances in neural information processing systems, 34:20132–20145, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.261156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.261156Z digest=sha256:ee9fb008c5ba7fc5b41f11cc8fd62622f04bc8e52dfd4d9c719b558202ff76ca

Observation 1c959ac2-172f-4f83-9b2e-2015d712f9af · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Flow-Based Policy for Online Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.264659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.264659Z digest=sha256:12ff9513cde389149e6d1cc68cea73bcbff3050731c0a12faf427616dab37cbc

Observation 66c705be-1d97-4367-a78f-dbb4d2de7100 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Flow-Based Policy for Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.268177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.268177Z digest=sha256:5e962ea69f959096428e8ef58cc981ae2bbdb64a4d6d5c4895e5a66104ca788b

Observation e3827596-99a4-4a14-9fce-efbbcb812d70 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Flow-Based Policy for Online Reinforcement Learning TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.271871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.271871Z digest=sha256:57e4f3fc3e5e4fb733512a63da028dde9c2f5ace1afe8a3b535bd0d2b7d32e01

Observation 5f9fb3e0-761f-4815-824b-d8207e8b40d6 · outbound

This paper cites DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.275799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.275799Z digest=sha256:beefe63dfb7be0e60cc0f60fec4c26cb5021f8c2de0971699f17bb72f9a2ffe9

Observation ae535516-f9c2-4d21-8c61-7b6ce41b57c3 · outbound

This paper cites Denoising diffusion probabilistic models.

Flow-Based Policy for Online Reinforcement Learning Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.279335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.279335Z digest=sha256:5301121f5815c0bbdd48f8f34db5dd11a6a6595da65246df1b01e98f09bb8abe

Observation a4b4f240-ded0-4118-bfef-25dae157c7d9 · outbound

This paper cites AlphaFold Meets Flow Matching for Generating Protein Ensembles.

Flow-Based Policy for Online Reinforcement Learning AlphaFold Meets Flow Matching for Generating Protein Ensembles

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.282865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.282865Z digest=sha256:4a1ffcd2039240931e99c9622a2667d1819ee30ebb5359a90ab48a85174d04a6

Observation eee6799d-6930-46f9-bf39-28a7b8d6d488 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning Efficient diffusion policies for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.838771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.286738Z digest=sha256:a042260b0f5d9723be6d9ac1a88784be9d730f99bf7d384525dd90cb648b2715

Observation b2ecf5d1-dff0-4a92-87bb-79fa9c817978 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.291276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.291276Z digest=sha256:2c3265199fec45e151b8101abdb8b80b16ddcc4a2e431aaf8b235d5e48a2fbc5

Observation ff883579-478f-462f-85f6-075c11acc940 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.Advances in neural information processing systems, 32, 2019.

Flow-Based Policy for Online Reinforcement Learning Stabilizing off-policy q-learning via bootstrapping error reduction.Advances in neural information processing systems, 32, 2019

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.826534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.295541Z digest=sha256:56a787a89d096ae8c8bf1b75548d95aebad14a72ed440710446000463fcf81dc

Observation 5767e56e-2700-4e23-ab8c-2e2619586369 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.299421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.299421Z digest=sha256:540bdcaf56340d4913ff8d78908880a8fe8bcc0ff9c7da7f957078bf8c5f6529

Observation 5f5ddba2-7544-404d-8157-2145da7564d2 · outbound

This paper cites Flow Matching for Generative Modeling.

Flow-Based Policy for Online Reinforcement Learning Flow Matching for Generative Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.304223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.304223Z digest=sha256:2f0b0a4a29c47fb2da8c6dd0120cc27a2007f99cfd4ffa6e9b0b6a16a8def6b3

Observation 16d2fc83-3023-49ca-9ec2-322a9f364cbf · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Flow-Based Policy for Online Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.308765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.308765Z digest=sha256:e4d6465139e66a0d69b8d7de4d6e39b60fc07007cb05bb1a9403803be1a1930e

Observation 5559add1-5eab-4a63-a9f3-207b008c0301 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.312759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.312759Z digest=sha256:1a02f77cd8578e900af50663d653bc11f7dd6301f5269032901a4e314e360363

Observation 05275907-e1bd-401e-ba72-8a83ef446b96 · outbound

This paper cites Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL.

Flow-Based Policy for Online Reinforcement Learning Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.316738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.316738Z digest=sha256:9f9f01c24e688b93f2a8e5637500132970129faf9bedbcbc0f42a1e0cc5363c4

Observation 451216f4-161f-49d8-b82f-54a353574301 · outbound

This paper cites Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.320535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.320535Z digest=sha256:339487d64e7548ebf9648b62d3d6ac0895bd895cfa21b60617fc3004821c8de9

Observation b5f19a43-4b22-43f6-9662-5bb718a1ff0d · outbound

This paper cites Self-imitation learning.

Flow-Based Policy for Online Reinforcement Learning Self-imitation learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.324469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.324469Z digest=sha256:1ce60ddd575f3818bcc866573c2120f7d22b70189f77f045364a12cf9fcff412

Observation e6664bdf-b841-4192-9764-287892c1f78d · outbound

This paper cites The difficulty of passive learning in deep reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning The difficulty of passive learning in deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.797418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.327981Z digest=sha256:b0f610fc6eaa937a4f03ccdbe9f249554aff1089c449d7eafb3376da5625eac5

Observation 4c089d76-81fb-4215-aa34-7ef7fee7e4dd · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Flow-Based Policy for Online Reinforcement Learning Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.331493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.331493Z digest=sha256:ebf3031ffa085542f4ca7cf97d383832b82d52b3684562cf29dc1f20cd9d3e16

Observation b5376564-e237-4d21-84e6-5269287a3572 · outbound

This paper cites Flow Q-Learning.

Flow-Based Policy for Online Reinforcement Learning Flow Q-Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.335423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.335423Z digest=sha256:a4150c45c8bb60e49ed06576cff3361ec407e71c5f94f5b8c624e7faf340649e

Observation 7b82270a-0f0d-4c66-a73e-b4ae207c82f3 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.339230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.339230Z digest=sha256:243c75be79189c804501cda6b4a092bd74ae07c41a4cf7bdff34551ea58c3428

Observation 4fa8bec0-d654-4340-9d63-c84b28a59fbd · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Flow-Based Policy for Online Reinforcement Learning Reinforcement learning by reward-weighted regression for operational space control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.343045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.343045Z digest=sha256:e54a76ef16b1c6db6d83b551099ed4ca49b19a9dc3611f9471e8f5587cd4d73a

Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.346736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.346736Z digest=sha256:cae778a012aadcfa27af44d91540bcf1e1ea099b9d87d71d33c0c4e70ec0e5d6

Observation 9e9e0f0b-acce-4343-af7e-f9f6d3418e63 · outbound

This paper cites Diffusion Policy Policy Optimization.

Flow-Based Policy for Online Reinforcement Learning Diffusion Policy Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.350660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.350660Z digest=sha256:a1393a900f6b4921f0063dfb0304bc7d958cc75954934c114195898a1da94fc2

Observation 84c40c83-3c3b-49ea-acb3-7b3e3307e811 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Flow-Based Policy for Online Reinforcement Learning HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.354763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.354763Z digest=sha256:ce16bba2b2def34bf2057f0a3f529b6b2b2184e6ef0176960589249525081d30

Observation df987991-127d-473b-a2b7-c300bbf8f0a8 · outbound

This paper cites Deterministic policy gradient algorithms.

Flow-Based Policy for Online Reinforcement Learning Deterministic policy gradient algorithms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.777805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.358447Z digest=sha256:d37a38c83bb69f956dabe6608608bf70825f9e21088ba4e3bd9db107e5426943

Observation 6191128b-e248-4af0-99fa-4a7ef1b68612 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Flow-Based Policy for Online Reinforcement Learning Score-Based Generative Modeling through Stochastic Differential Equations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.362097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.362097Z digest=sha256:4f1c135ba442f73a025b9f160ecb2a223b2481c5718c00a650c44548c0c15bf5

Observation 22bb8b3c-75fa-437c-b6aa-14c2b38539ab · outbound

This paper cites Reinforcement learning.Journal of Cognitive Neuroscience, 11(1): 126–134, 1999.

Flow-Based Policy for Online Reinforcement Learning Reinforcement learning.Journal of Cognitive Neuroscience, 11(1): 126–134, 1999

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.765713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.365717Z digest=sha256:74954d1cb93873e746c62375ed65e1a55af555d25289e515f853a77409f48724

Observation ca32db8c-638b-4247-8b9d-8dbecc0ca744 · outbound

This paper cites DeepMind Control Suite.

Flow-Based Policy for Online Reinforcement Learning DeepMind Control Suite

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.369356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.369356Z digest=sha256:68b016a0065e34cb2ba9dbbedf31cc21525f2603c991da35e420c7c1ef4c916f

Observation bd1637be-3421-46af-80c1-20a836495baa · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Flow-Based Policy for Online Reinforcement Learning Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.373204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.373204Z digest=sha256:b4f32f6dcdea76854dc9006247df43abf87cb2311841012f3ba50d5cde5b81ed

Observation bc033cc0-5e72-42ef-a15e-3b9b41de0543 · outbound

This paper cites Springer, 2008.

Flow-Based Policy for Online Reinforcement Learning Springer, 2008

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.377185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.377185Z digest=sha256:3e80510107f12f6385f5796feebc3bd2e1837654373ec36608a0dd41fc916c10

Observation 39b9530c-34af-4332-a469-378be282eb02 · outbound

This paper cites Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024.

Flow-Based Policy for Online Reinforcement Learning Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.744527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:13:00.380879Z digest=sha256:b62664be0a2dbe620bcdfc0adc11b99805bd2ed34c3c1e3b762e2bd4dcc87b4e

Observation 43cc9353-aed6-4360-8a61-8545b6fe35fa · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.385359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.385359Z digest=sha256:d5659d146e7986cc868848c543c8f235a8b768d4da9984d07d1276978e50b9ab

Observation 12bff938-c31d-4619-9847-279cfb70f603 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.389144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.389144Z digest=sha256:d009d54681acbd0bfd4928bb47ba7c1b1ab2ddf4c85e61d827c531b1f3ee321c

Observation f9093019-2300-4f7a-b23c-d42f708e8bf6 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.393031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.393031Z digest=sha256:142867fc24d04e97e3ccb25fe91d63c1a7f7651482a2425498e14b967fc4bc0d

Pith citing papers

Observation a3f2938f-aa5c-4a56-b8e3-6a47e1da3c77 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Flow-Based Policy for Online Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.961170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:4d70716d43d03620300638a979c9ab63d931e9d0bb7515a630597ce8878b51e2

Observation 7a5816e2-7622-49d0-9370-8686d0084056 · inbound

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data cites this paper.

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data Flow-Based Policy for Online Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.772811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T14:24:58.919333Z digest=sha256:472f3450adc77e69685ec1304dc08edf93a7f0a0a21153728168809127fc428e

Observation 1274de69-c92b-4002-9014-8ce57310826e · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Flow-Based Policy for Online Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.879087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:a18950b25b852aa81eca704b7af27f4576fec4de3ee80595dd6abb58cef902b6

Observation 53d83936-0ddb-46c2-85f7-48147fa6f40b · inbound

ReFPO: Reflow Regularization for Flow Matching Policy Gradients cites this paper.

ReFPO: Reflow Regularization for Flow Matching Policy Gradients Flow-Based Policy for Online Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.803584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T14:38:54.754049Z digest=sha256:456a92b1b9367f27cb4317ce6a6a2c0dbbc63b8cc8cd215307995d039f93af77

Observation 450a06e7-e743-494f-bce1-df4265c25f7d · inbound

Dual-Flow Reinforcement Learning with State-Aware Exploration cites this paper.

Dual-Flow Reinforcement Learning with State-Aware Exploration Flow-Based Policy for Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.792831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T07:16:11.130271Z digest=sha256:02064be9c45b73b310cec3821687155e342d09d0aba6b2b17c26c7eaa40c5a52