Pith. sign in

Paper Citation Record · LEDGER

Flow-Based Policy for Online Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2506.12811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12811 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.393031Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:16:11.130271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:37.801846Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80413add-a181-4e55-9b7a-6e83205e7cef · outbound

This paper cites A markovian decision process.Journal of mathematics and mechanics, pages 679–684, 1957.

Flow-Based Policy for Online Reinforcement Learning A markovian decision process.Journal of mathematics and mechanics, pages 679–684, 1957

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.214500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.214500Z digest=sha256:b46c6661591186410214adb57a937b349bcbc692f9b15d52e06d2fdb65a97bd2

Observation 9b0c6a2b-544c-4413-98e7-e7294918d8b6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Flow-Based Policy for Online Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.218998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.218998Z digest=sha256:2fd8812c00407201581b48fd0870817809c496c7e2f1b33f94d31c258a96c43d

Observation 6cc5feda-359b-423e-9737-d00e0c3f5221 · outbound

This paper cites SE(3)-Stochastic Flow Matching for Protein Backbone Generation.

Flow-Based Policy for Online Reinforcement Learning SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.225162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.225162Z digest=sha256:c225d8e92e7cb02d005faa6d15994e2ff534c816f0be72807ac8ae458384fd39

Observation 9c287922-e198-44e5-8940-a6a508c11558 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.230392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.230392Z digest=sha256:7ad2189f0ad554a802edef68e76c93d587262f04be0430325b3e3655d0f04c8e

Observation ef08daae-3b0b-4451-8425-2ea98617bb5b · outbound

This paper cites Neural ordinary differential equations.

Flow-Based Policy for Online Reinforcement Learning Neural ordinary differential equations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.234827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.234827Z digest=sha256:3ebfdda8ff9fe4980c71fe0560a60e107d22394bcd892cdaae6ba90907226aa7

Observation 31978adb-d075-48f7-854b-74808710fc45 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023.

Flow-Based Policy for Online Reinforcement Learning Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, page 02783649241273668, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.239064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.239064Z digest=sha256:02c8750bd09f1f79b65fbe40f916cbd522aabab1861aca155c5188f7228cc452

Observation e6a40291-4d40-46d6-9601-55cf0b6294b0 · outbound

This paper cites Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021.

Flow-Based Policy for Online Reinforcement Learning Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.243560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.243560Z digest=sha256:d3d214b84ed66da31a5be1ec439154fd7cd3a5d8e3099171c3b23003518a1bf4

Observation 37f1818f-1dcf-4766-a046-14487dd8c0bd · outbound

This paper cites Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization.

Flow-Based Policy for Online Reinforcement Learning Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.248611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.248611Z digest=sha256:9c6c7fd42ddb5302734263e720c4654d19ceb48be25a883fccc13e5f4cbf6bf1

Observation 2d9509c2-c3b1-49a4-a133-df57efb2fe73 · outbound

This paper cites Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.253038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.253038Z digest=sha256:46f39b487f7015e82ff2f40f3684a7550f9cf87e6fa45644afb91d0af13a9e48

Observation d02ead63-07d8-4762-a248-384541f0024a · outbound

This paper cites Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization.

Flow-Based Policy for Online Reinforcement Learning Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.257162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.257162Z digest=sha256:f319890ff0eac13d77af042db239e321fa787630e6787eb27624ade507776226

Observation 15083ea9-9872-4702-8804-b9f3787aad9a · outbound

This paper cites A minimalist approach to offline reinforcement learning.Advances in neural information processing systems, 34:20132–20145, 2021.

Flow-Based Policy for Online Reinforcement Learning A minimalist approach to offline reinforcement learning.Advances in neural information processing systems, 34:20132–20145, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.261156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.261156Z digest=sha256:e70f9014f8af3cc73c1ff55e8e97b84a36aa0c7ed7155a7f30dc26d2e99b0cbf

Observation 1c959ac2-172f-4f83-9b2e-2015d712f9af · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Flow-Based Policy for Online Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.264659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.264659Z digest=sha256:f528d7faa3dd00d173375019c12c4ec4cf004e99133f605d222fefedfceae6dc

Observation 66c705be-1d97-4367-a78f-dbb4d2de7100 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Flow-Based Policy for Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.268177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.268177Z digest=sha256:cf8c7593ae8a3de185e9b814d08c42fa7218ea9bb5c1d3ef0a2580f814aba6a2

Observation e3827596-99a4-4a14-9fce-efbbcb812d70 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Flow-Based Policy for Online Reinforcement Learning TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.271871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.271871Z digest=sha256:c938aa27293e4313a8640146931bcd4b72dde09525c8f8cb5ba99472dc37b30c

Observation 5f9fb3e0-761f-4815-824b-d8207e8b40d6 · outbound

This paper cites DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.275799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.275799Z digest=sha256:8af827338c70a6b2837452f05839795d51b32ff387df4260ed81a626834ed2b5

Observation ae535516-f9c2-4d21-8c61-7b6ce41b57c3 · outbound

This paper cites Denoising diffusion probabilistic models.

Flow-Based Policy for Online Reinforcement Learning Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.279335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.279335Z digest=sha256:2312a2de033d994f338f8e676ebe5dd37ba2fa23c90e205742f094d6dbe66fd7

Observation a4b4f240-ded0-4118-bfef-25dae157c7d9 · outbound

This paper cites AlphaFold Meets Flow Matching for Generating Protein Ensembles.

Flow-Based Policy for Online Reinforcement Learning AlphaFold Meets Flow Matching for Generating Protein Ensembles

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.282865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.282865Z digest=sha256:db824eb09e18a266aa009d0f88b3ae9c6260c24ca04e3e717fa4ed6fcd4f7250

Observation eee6799d-6930-46f9-bf39-28a7b8d6d488 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning Efficient diffusion policies for offline reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.838771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.286738Z digest=sha256:a114bf30da890803e60861f31f7fa512a44664e0c86408039db05c7538ced658

Observation b2ecf5d1-dff0-4a92-87bb-79fa9c817978 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.291276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.291276Z digest=sha256:339a85e06ec049c74cf9954ca3c6d88fc84138f2d62e2247d2daa67d079860fc

Observation ff883579-478f-462f-85f6-075c11acc940 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.Advances in neural information processing systems, 32, 2019.

Flow-Based Policy for Online Reinforcement Learning Stabilizing off-policy q-learning via bootstrapping error reduction.Advances in neural information processing systems, 32, 2019

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.826534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.295541Z digest=sha256:20855dd29fec46874d5a29836464b24316e0ccf4f58718610cc088bac3fdf8a6

Observation 5767e56e-2700-4e23-ab8c-2e2619586369 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Flow-Based Policy for Online Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.299421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.299421Z digest=sha256:b11cbeb34b70695fb42dd46e76a01425b157ecdcadccbe84f3f703009ece7318

Observation 5f5ddba2-7544-404d-8157-2145da7564d2 · outbound

This paper cites Flow Matching for Generative Modeling.

Flow-Based Policy for Online Reinforcement Learning Flow Matching for Generative Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.304223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.304223Z digest=sha256:af8331265e68a9fe04151b415c39a0c3e2691b3e95f8a516a6d024a6cc381aec

Observation 16d2fc83-3023-49ca-9ec2-322a9f364cbf · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Flow-Based Policy for Online Reinforcement Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.308765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.308765Z digest=sha256:a654faabda18bd10cbfaa189934bc496081194e2fc1dc666d8331f56e88b4da8

Observation 5559add1-5eab-4a63-a9f3-207b008c0301 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.312759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.312759Z digest=sha256:fa96abf439deae78e5a04fba05f7c4a12725f7606cc2220f75a33cde02d06c61

Observation 05275907-e1bd-401e-ba72-8a83ef446b96 · outbound

This paper cites Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL.

Flow-Based Policy for Online Reinforcement Learning Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.316738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.316738Z digest=sha256:cb7f8754763dafca94af7757c17d4c614696d37e89bd82950d7c13e65a64fff7

Observation 451216f4-161f-49d8-b82f-54a353574301 · outbound

This paper cites Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.320535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.320535Z digest=sha256:7dc6026eb718993fcc67d887769febfdb03274a77da93f6d5cd101bd6f6d9c5e

Observation b5f19a43-4b22-43f6-9662-5bb718a1ff0d · outbound

This paper cites Self-imitation learning.

Flow-Based Policy for Online Reinforcement Learning Self-imitation learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.324469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.324469Z digest=sha256:db3ffd289aeff3b5b4c9e4965bec96a68ef3034bac8437381d91b1eb8f371c2c

Observation e6664bdf-b841-4192-9764-287892c1f78d · outbound

This paper cites The difficulty of passive learning in deep reinforcement learning.

Flow-Based Policy for Online Reinforcement Learning The difficulty of passive learning in deep reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.797418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.327981Z digest=sha256:22e57b19b28ef74c202c9ab0788b559e090fd2f8d2a8bf385c0332aa0d9e86b1

Observation 4c089d76-81fb-4215-aa34-7ef7fee7e4dd · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Flow-Based Policy for Online Reinforcement Learning Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.331493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.331493Z digest=sha256:36e944504060dc6fb497072a98886b152a70686cf47933ad4c3f2cebf37ec547

Observation b5376564-e237-4d21-84e6-5269287a3572 · outbound

This paper cites Flow Q-Learning.

Flow-Based Policy for Online Reinforcement Learning Flow Q-Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.335423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.335423Z digest=sha256:71b2259cba87914e987f514def50d163862a8330cb4cb33f19bf003845e0ab1e

Observation 7b82270a-0f0d-4c66-a73e-b4ae207c82f3 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.339230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.339230Z digest=sha256:05c679a27ecb773c9dd160d6526e27f2a80ce605bb44a09e3f4780ba9edef83b

Observation 4fa8bec0-d654-4340-9d63-c84b28a59fbd · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Flow-Based Policy for Online Reinforcement Learning Reinforcement learning by reward-weighted regression for operational space control

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.343045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.343045Z digest=sha256:0ea048761277fb6fee8406ebe161f06bba06d46a3514dc5230aaea206925b29d

Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.346736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.346736Z digest=sha256:8cd528e1af3f9be30a0dfb6c2a488a6a4ad1f2c29097977fb5924247727b1fba

Observation 9e9e0f0b-acce-4343-af7e-f9f6d3418e63 · outbound

This paper cites Diffusion Policy Policy Optimization.

Flow-Based Policy for Online Reinforcement Learning Diffusion Policy Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.350660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.350660Z digest=sha256:532fbe2cadbcdeb1376fdf360d988399f3fb7d339ad81df256c930d845a6a43b

Observation 84c40c83-3c3b-49ea-acb3-7b3e3307e811 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Flow-Based Policy for Online Reinforcement Learning HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.354763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.354763Z digest=sha256:f9c02d716b4348fc8efd8476513d4f8d64745f0ded838ef9915a4b06d6939917

Observation df987991-127d-473b-a2b7-c300bbf8f0a8 · outbound

This paper cites Deterministic policy gradient algorithms.

Flow-Based Policy for Online Reinforcement Learning Deterministic policy gradient algorithms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.777805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.358447Z digest=sha256:b51b8043c8c9364c7fe5f47d7abb38f07eecdc31beeeab4d569ca60e9c72359d

Observation 6191128b-e248-4af0-99fa-4a7ef1b68612 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Flow-Based Policy for Online Reinforcement Learning Score-Based Generative Modeling through Stochastic Differential Equations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.362097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.362097Z digest=sha256:0858d7210dd4793c67521a5cf24dea985d2d67d287108a21562cd216770ab2e2

Observation 22bb8b3c-75fa-437c-b6aa-14c2b38539ab · outbound

This paper cites Reinforcement learning.Journal of Cognitive Neuroscience, 11(1): 126–134, 1999.

Flow-Based Policy for Online Reinforcement Learning Reinforcement learning.Journal of Cognitive Neuroscience, 11(1): 126–134, 1999

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.765713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.365717Z digest=sha256:6d9d677ef414f6cb1faa28961c1a6de334eb217439edc4489981715e26505f67

Observation ca32db8c-638b-4247-8b9d-8dbecc0ca744 · outbound

This paper cites DeepMind Control Suite.

Flow-Based Policy for Online Reinforcement Learning DeepMind Control Suite

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.369356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.369356Z digest=sha256:0b2f58a98bc2d05201416258e2a0ae38086d88ceaa3ab9ed8f9e71a3ef61fa74

Observation bd1637be-3421-46af-80c1-20a836495baa · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Flow-Based Policy for Online Reinforcement Learning Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.373204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.373204Z digest=sha256:46f18bf4b2dd51f7563a040e93d6e28d8c3274c4d554fe0922801cf9d1e6c8a0

Observation bc033cc0-5e72-42ef-a15e-3b9b41de0543 · outbound

This paper cites Springer, 2008.

Flow-Based Policy for Online Reinforcement Learning Springer, 2008

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.377185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.377185Z digest=sha256:5a8d673f0f60cc8a3ea1ff3c80b026f6c0aa4b1c40989453fbce091c64eab6cf

Observation 39b9530c-34af-4332-a469-378be282eb02 · outbound

This paper cites Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024.

Flow-Based Policy for Online Reinforcement Learning Diffusion actor-critic with entropy regulator.Advances in Neural Information Processing Systems, 37:54183–54204, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:00.744527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:13:00.380879Z digest=sha256:e0c179f52ba441cd8cfd4513d43451c93e7ac51f3a5fedbe4f4fb68b74755fd3

Observation 43cc9353-aed6-4360-8a61-8545b6fe35fa · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.385359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.385359Z digest=sha256:f501aee564217962692e02aa18b19a7e3e4004d9ff3067695acd0c2886215c97

Observation 12bff938-c31d-4619-9847-279cfb70f603 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.389144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.389144Z digest=sha256:ef25e372ff7f8a143b7698afb79d52aa9762879ed1632fb638a43dd18689fb9d

Observation f9093019-2300-4f7a-b23c-d42f708e8bf6 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Flow-Based Policy for Online Reinforcement Learning Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.393031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.393031Z digest=sha256:67cf6bf3692dd495a0d91def74b81883cbbc2ba4c49bbd7bca48ba0266b331d9

Pith citing papers

Observation a3f2938f-aa5c-4a56-b8e3-6a47e1da3c77 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Flow-Based Policy for Online Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.961170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:8995cf83918623c08022991dc61715cf41f53dde99bcb1d61cc031f739dd50ab

Observation 7a5816e2-7622-49d0-9370-8686d0084056 · inbound

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data cites this paper.

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data Flow-Based Policy for Online Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.772811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T14:24:58.919333Z digest=sha256:13bdc5be547691a80ce1ad4fb18ed34dd8b8c2cedf63f653dd020a10c5ad6758

Observation 1274de69-c92b-4002-9014-8ce57310826e · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Flow-Based Policy for Online Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.879087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:c561cd4907de990ee32e49d24823c226ea1cee994188b5fe88a3cf38472061d5

Observation 53d83936-0ddb-46c2-85f7-48147fa6f40b · inbound

ReFPO: Reflow Regularization for Flow Matching Policy Gradients cites this paper.

ReFPO: Reflow Regularization for Flow Matching Policy Gradients Flow-Based Policy for Online Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.803584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T14:38:54.754049Z digest=sha256:a1151133f5413655f5ded120eb508f39d0f24102b26b32b6ec2bbe2481ca51ae

Observation 450a06e7-e743-494f-bce1-df4265c25f7d · inbound

Dual-Flow Reinforcement Learning with State-Aware Exploration cites this paper.

Dual-Flow Reinforcement Learning with State-Aware Exploration Flow-Based Policy for Online Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.792831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T07:16:11.130271Z digest=sha256:1c919a939d04af0fdabf3bab74dbdb6376f56a1a8accb03e7225fd81c6784d7d