Pith. sign in

Paper Citation Record · LEDGER

Decision Flow Policy Optimization

As of 20 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2505.20350.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20350 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:42.023257Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact7
  • verified fuzzy17
  • unresolved71
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bb54068-3566-43d6-9ef6-5368890198c5 · outbound

This paper cites Diffusion policies for out-of-distribution generalization in offline reinforcement learning.

Decision Flow Policy Optimization Diffusion policies for out-of-distribution generalization in offline reinforcement learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.794468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.794468Z digest=sha256:8c74908246f0a91ce20074b5b107469796630157b0053366c9ec381a30dfee80

Observation f119d52d-6872-42d2-8739-dad7a8a1e273 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

Decision Flow Policy Optimization Is Conditional Generative Modeling all you need for Decision-Making?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.841381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.841381Z digest=sha256:bc7357cbd96188a21664550dbaf2a9919ef7a69c1982056e6213a7e64530d06c

Observation de253d64-9fad-452f-b44c-a5243b4d6ea9 · outbound

This paper cites Let Offline RL Flow: Training Conservative Agents in the Latent Space of Normalizing Flows.

Decision Flow Policy Optimization Let Offline RL Flow: Training Conservative Agents in the Latent Space of Normalizing Flows

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.925222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.925222Z digest=sha256:0d1080b7f422e3d56ffdb91ee23a5360915a622fa42b06d20d635c5d7839ce03

Observation 0f129e7f-1e58-4fca-a9d3-2fb211336e0d · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Decision Flow Policy Optimization Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:35.996839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:35.996839Z digest=sha256:90d9592c35add3bf2b8a41b43b9a9f9dde69a104c18c306d552aac9dd0b4102e

Observation b4df2d69-3fe9-4ec9-81b5-adf5c3b19d50 · outbound

This paper cites Model-Based Offline Planning.

Decision Flow Policy Optimization Model-Based Offline Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.027635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.027635Z digest=sha256:6582e4b0d52d869d3135ff708b3fac9abc8c567534a9e4d794824dee9c7f26c4

Observation 0efb84d6-7b3e-45e6-ad7b-bd796887085b · outbound

This paper cites The peano-baker series.

Decision Flow Policy Optimization The peano-baker series

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.087639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.087639Z digest=sha256:db095d2a966865ba89713b27b9f26ef0bdbfb16b4932f36ca30e5d3fe326b39d

Observation f669e893-c37a-4c18-aedb-6dc3c9e72cc4 · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Decision Flow Policy Optimization Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.145203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.145203Z digest=sha256:86151ce9da3629463de5a73ca898b398c610acbfe4a0fdcd42f7224686f59d19

Observation 84d6bdb3-f2f8-4b43-b4c0-e1357b662ae1 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Decision Flow Policy Optimization Training Diffusion Models with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.208602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.208602Z digest=sha256:c6c205990b684dc9cb64940e0c67103054f007f5d0e2c582b494b963e28519ca

Observation 348dc640-46d7-4a64-9952-8c7840faa9e3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Decision Flow Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.260626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.260626Z digest=sha256:215c77cec9ea1ab4d493a326d2d10ebeccbd9565081cb3921c4ab3d535657a03

Observation 7bcaa2fe-154a-4da1-8370-bcc01f7a57c5 · outbound

This paper cites Reinforcement Learning for Generative AI: A Survey.

Decision Flow Policy Optimization Reinforcement Learning for Generative AI: A Survey

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:43.903572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:36.325840Z digest=sha256:91d46c4cf51b0083fefeee6a3c9f151a238097d00c9619677d258d0f738d2cc7

Observation 820a9767-3a44-4d6b-9503-ab3a35904761 · outbound

This paper cites Simple Hierarchical Planning with Diffusion.

Decision Flow Policy Optimization Simple Hierarchical Planning with Diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.384955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.384955Z digest=sha256:2e377a442dfcb3c9e3ea9098f48cd31e30739685c0fb38019b0804d165c493fb

Observation ad21c487-7e76-4c69-aab3-0e6db50e2a36 · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Decision Flow Policy Optimization Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.514447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.514447Z digest=sha256:dd5a5e42dfad6fcf14d6995585c1e670723d70abd96323d11e552454f2808c6a

Observation 7bbc88ad-efbd-4cf0-9dda-be2de6e039af · outbound

This paper cites Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions.

Decision Flow Policy Optimization Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.669857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.669857Z digest=sha256:4089a743e455033f5da687c66d60731af0fc9549ab07e362bb2290880f4a805a

Observation 3c723759-4267-4c73-ac31-05085568bd52 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Decision Flow Policy Optimization Decision transformer: Reinforcement learning via sequence modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.788091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.788091Z digest=sha256:f428499fb19f72f1d4c53be411f953b3ad388dcd12a1f3be883e145b01dee79d

Observation 2278a0db-9356-4ad5-bd8a-7920cd691edf · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Decision Flow Policy Optimization Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.878337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.878337Z digest=sha256:eaf6ff0182936853b5b6ea7de1903fe507fd9a56f57fca781a0fcaefd826f222

Observation db435bce-f951-41d5-9933-609c4375bfae · outbound

This paper cites Flow Matching in Latent Space.

Decision Flow Policy Optimization Flow Matching in Latent Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.912260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.912260Z digest=sha256:a51892fe3d82a40dc82616deb00fb72021aa18827cb9c34cdb93c6d63a74323c

Observation db110ff5-eb32-42fb-b950-82196006e03a · outbound

This paper cites Fisher flow matching for generative modeling over discrete data.

Decision Flow Policy Optimization Fisher flow matching for generative modeling over discrete data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:36.964361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:36.964361Z digest=sha256:8c52d72494764cc907805db9bd7d67da1fe6d6037299bfab2d23669387c34dcc

Observation de1150a3-cfae-485b-a4eb-b33f5c0b678b · outbound

This paper cites Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization.

Decision Flow Policy Optimization Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.020291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.020291Z digest=sha256:e16a73fcf0bec0061062e9de550bf4e86a14d588f6694a6e0b77f81cb3e79346

Observation 738948f2-c8df-4a37-bca2-b4a02c58691d · outbound

This paper cites DiffuserLite: Towards Real-time Diffusion Planning.

Decision Flow Policy Optimization DiffuserLite: Towards Real-time Diffusion Planning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.065822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.065822Z digest=sha256:cd13e6df74fb5c247d6dd4add9276efb5cb33412df08a9aadac5efbeb849fc3a

Observation 96595d25-2083-4c04-a978-81d105de798d · outbound

This paper cites Probabilistic number theory I: Mean-value theorems, volume 239.

Decision Flow Policy Optimization Probabilistic number theory I: Mean-value theorems, volume 239

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.130184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.130184Z digest=sha256:3f1d24888ae9660a8be520a8e7650d266f804946c788f29a3204f4ada22e9e6f

Observation 25e213c1-9081-4b08-b16a-346397d1a105 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Decision Flow Policy Optimization Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.190416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.190416Z digest=sha256:639c8d592aa37923eb03e13a1935610c06aea32abd09e79f3a1008334184fc11

Observation e0ada8b7-b884-46c2-a7b3-cd3f3d362923 · outbound

This paper cites A reinforcement learning diffusion decision model for value-based decisions.

Decision Flow Policy Optimization A reinforcement learning diffusion decision model for value-based decisions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.230498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.230498Z digest=sha256:6ffe59f925e18799dc74cc92581cf22336c6204b2435161f98c2004bb09d1b66

Observation 4e7cf567-5aba-4e08-8562-5dae2e2047e3 · outbound

This paper cites Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.

Decision Flow Policy Optimization Reinforcement learning for generative ai: State of the art, opportunities and open research challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.281105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.281105Z digest=sha256:60c253709ce96358cfda7f93521d6c026b6947b465a7a2ad97ceca2f7d947ae8

Observation 9af8a4e9-a85c-4332-b706-7a7bec547408 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Decision Flow Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.372971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.372971Z digest=sha256:ee4285d3a51b92d591216377ea30d454461cc4f38d46bb4326b66f7ef244d8e4

Observation 9a03472b-99bd-44ef-9143-fee30503e1db · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Decision Flow Policy Optimization A minimalist approach to offline reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.438497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.438497Z digest=sha256:1266e4f75f88183eaf3635fd72a7c7cb188bd2a9e2bee725ec7d522636b1885d

Observation e125cbdf-408d-4e94-88ba-db8bc3189dc0 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Decision Flow Policy Optimization Off-policy deep reinforcement learning without exploration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.510693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.510693Z digest=sha256:c827312c37a631c70c3b67781a3bba77aef9b59dcdcc60aa5bf352462fdb0fb0

Observation d953dd6a-2031-45a3-9913-27c10a576f89 · outbound

This paper cites Generalized Decision Transformer for Offline Hindsight Information Matching.

Decision Flow Policy Optimization Generalized Decision Transformer for Offline Hindsight Information Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.550897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.550897Z digest=sha256:5001101d8fff554be0234542fec6c986948521a99dc72fdd1c646a7b8bdc87b1

Observation 2dff5418-ebe0-492d-be11-b3819b4cd23c · outbound

This paper cites Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning.

Decision Flow Policy Optimization Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.619351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.619351Z digest=sha256:1ebd5572cba9b778f4ca0406da2affed1ef4aca47ce016f940da93591695fdfd

Observation 2206e9e9-7177-4d91-8879-450d1184d261 · outbound

This paper cites Discrete flow matching.

Decision Flow Policy Optimization Discrete flow matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.684967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.684967Z digest=sha256:9bc927e48dda845a50f170690f57bea8fa34740aa3c24c84148dd701043631f1

Observation d631f224-5059-4175-9157-3e54022fb3d9 · outbound

This paper cites Offline rl policies should be trained to be adaptive.

Decision Flow Policy Optimization Offline rl policies should be trained to be adaptive

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.741192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.741192Z digest=sha256:7167c07b608b1f07f8072a7456cd8ae135ea2d200e4766f2de8729ddfa4e0d1a

Observation a5047400-14cb-44ea-968e-41a56053beaf · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Decision Flow Policy Optimization Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.793878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.793878Z digest=sha256:cd47310405a9e6c20edc9a952afae553be72397d2978ff92740f79711e984dfb

Observation 51eff6bf-9a58-4084-b442-0edb90f51bf6 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Decision Flow Policy Optimization IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.857764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.857764Z digest=sha256:d8ef6f4a61e02a82d1f49985ce18c192ea93f691e811a3ec8fef5cf583c69bfb

Observation 56f6ab40-9ae6-4d94-a4f1-86474d6ada37 · outbound

This paper cites Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning.

Decision Flow Policy Optimization Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.931785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.931785Z digest=sha256:1a14cfffe458783163017d1e85ad63e2d8a6cc256794cdc98d80ba148aadf356

Observation 45995f37-337d-4382-bcb8-5e9c76826337 · outbound

This paper cites Lectures on Lipschitz analysis.

Decision Flow Policy Optimization Lectures on Lipschitz analysis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:37.998904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:37.998904Z digest=sha256:6e964adff9d1d2d3f0d802f4bb7d953fba3f1b5447f1c3f4ded4bdab9a28ee8f

Observation a942017a-8f99-480c-bb0a-77b6a401fe28 · outbound

This paper cites Flow++: Improving flow-based generative models with variational dequantization and architecture design.

Decision Flow Policy Optimization Flow++: Improving flow-based generative models with variational dequantization and architecture design

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.278627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.069389Z digest=sha256:345c693825fab97a06e285245fdf09acbfb9d093c695a0f66a7bf36918c9f0a4

Observation 20f550fd-9574-465a-818a-2a7c68710f0b · outbound

This paper cites Instructed Diffuser with Temporal Condition Guidance for Offline Reinforcement Learning.

Decision Flow Policy Optimization Instructed Diffuser with Temporal Condition Guidance for Offline Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.121882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.121882Z digest=sha256:2a38b82cdc42ecd60a01e0cf6714667af2d29426c77a3716773cb4f64966ed9c

Observation a05039dd-c937-4793-acdb-53693627ce24 · outbound

This paper cites On Transforming Reinforcement Learning by Transformer: The Development Trajectory.

Decision Flow Policy Optimization On Transforming Reinforcement Learning by Transformer: The Development Trajectory

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:43.445599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.165967Z digest=sha256:c9d987baa195259c67c679c1e36bf7ca39b340a3a01300eb91f6acfd2817cebe

Observation b280403c-910d-4944-adf1-5da6f0f07d93 · outbound

This paper cites Graph Decision Transformer.

Decision Flow Policy Optimization Graph Decision Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.225712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.225712Z digest=sha256:e50f7a376b8b43a850fd5740c84789017adc6faa6ebfbef1c31f7a60612feaea

Observation f3e6866e-bb55-41c4-947e-cd9b8066ff5e · outbound

This paper cites Adaflow: Imitation learning with variance- adaptive flow-based policies.

Decision Flow Policy Optimization Adaflow: Imitation learning with variance- adaptive flow-based policies

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.112420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.289200Z digest=sha256:bed7f0c1867279b3501d81160582d35deeb8e5d80c4f819a513e71fd56b8df83

Observation ede41e49-49ce-4c59-8bfd-0440c2f19fb2 · outbound

This paper cites Diffusion models as optimizers for efficient planning in offline rl.

Decision Flow Policy Optimization Diffusion models as optimizers for efficient planning in offline rl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:46.003051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.338249Z digest=sha256:1ca9e65f1df90ff8629fb1c6c806bcf6ff8a894e6b886531240d05a6b223ed40

Observation f402d366-30cd-4120-8488-c0978b239947 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.

Decision Flow Policy Optimization Offline reinforcement learning as one big sequence modeling problem

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.883231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.382943Z digest=sha256:50b593dfa77b276df7bae80552da4a062f8fb6e7a6fe25a125cd29f9b67844c9

Observation cecf9b24-9199-4a7a-a826-8c0511ff1539 · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Decision Flow Policy Optimization Planning with Diffusion for Flexible Behavior Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.445055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.445055Z digest=sha256:475d3cd7f95db108756fcbbd4c68a8d5080f68a776df1f852046bb8d70a29684

Observation 4861505a-1e2d-487c-a132-c445086446ce · outbound

This paper cites Efficient Planning in a Compact Latent Action Space.

Decision Flow Policy Optimization Efficient Planning in a Compact Latent Action Space

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.497294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.497294Z digest=sha256:a4d2783f2a269e220a304382ed65fdc63f0ea199c6009b3f1ad904b4ecfc2bb5

Observation 2d8098ab-c4e2-459f-9614-4cfac6062659 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Decision Flow Policy Optimization Pyramidal flow matching for efficient video generative modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.547205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.547205Z digest=sha256:8f5500322ff5dd0a943e9ea0982ac05e2ddd66a520ce16d9d50686328e4f88c9

Observation eb7e562b-9b99-433e-b89a-5d65b1f14354 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Decision Flow Policy Optimization Efficient diffusion policies for offline reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.709405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.624432Z digest=sha256:1fa2f3fa8c3fdef37c5a403c4e1490c83c313542260bf47af5dc18f6d5ed9732

Observation 147e7268-7bad-41a7-84cd-ca58fb1e5aa2 · outbound

This paper cites Morel: Model-based offline reinforcement learning.

Decision Flow Policy Optimization Morel: Model-based offline reinforcement learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.689398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.689398Z digest=sha256:80d8e7ce6494108ce4ddb5f2a85f9c05ba3e2b5ce76ec96c88df4b7c820c772c

Observation aade2317-0a92-436b-a2d0-33ca8966987f · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Decision Flow Policy Optimization Offline Reinforcement Learning with Implicit Q-Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.742099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.742099Z digest=sha256:3625a79b17d8453431314812adb530c5f459ec45d261d49f18e13757374ece99

Observation bb3f2554-6bce-47e0-a4e2-e4fbd730a15e · outbound

This paper cites Stabilizing off- policy q-learning via bootstrapping error reduction.

Decision Flow Policy Optimization Stabilizing off- policy q-learning via bootstrapping error reduction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.799102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.799102Z digest=sha256:39a94507c452e247bd96cd66ca1219898a43a500c0c5c69c610d95c8257338f1

Observation 913dcc55-4788-431b-bf37-de8794841bb8 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Decision Flow Policy Optimization Conservative q-learning for offline reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.842872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.842872Z digest=sha256:650974c8bd06179a7e7985233a5787807d031361ddabf85ab3cf782982f1b9f7

Observation d507ba1f-1aec-4ef2-b44b-8ca0ebbcad97 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Decision Flow Policy Optimization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.891334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.891334Z digest=sha256:86b8a4e4e60b2aaeb7e4badcee7e8701da384eb5225251d424bdf142b9d0f83f

Observation 4c22e361-d8f6-477d-9dbc-b58ce2e550ac · outbound

This paper cites DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching.

Decision Flow Policy Optimization DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:38.939358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:38.939358Z digest=sha256:162f7d2bfe7e76e5fd2745523c5c151eefaccc9334d9585fb72111cfbcd6f10b

Observation 1a19bf54-4dc7-4172-b70c-84e0a8517511 · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Decision Flow Policy Optimization Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.570026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:38.995078Z digest=sha256:37965055f8614963c68baad2656a2a0e3db3b999af3df507f51d86bea948df0c

Observation 65e40278-984b-4041-8211-9b2dc7e9e0a3 · outbound

This paper cites Efficient Planning with Latent Diffusion.

Decision Flow Policy Optimization Efficient Planning with Latent Diffusion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.045014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.045014Z digest=sha256:1a2fdb80dfa1e8843135e8858f3610174804afa21f01c668fb79d3c337ac472b

Observation 38d867f4-4693-4e1b-81ca-6b0a6d249a75 · outbound

This paper cites Hierarchical diffusion for offline decision making.

Decision Flow Policy Optimization Hierarchical diffusion for offline decision making

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.411100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:39.111910Z digest=sha256:9a0514b2eebc33a9735ed3cdc681d76a95288c1f478c355776589e1fde3106c6

Observation cfe9aa23-86c1-407b-90fc-015a29e29dcd · outbound

This paper cites Generative models in decision making: A survey.

Decision Flow Policy Optimization Generative models in decision making: A survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.162569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.162569Z digest=sha256:4a130a9d947d0b6090faa304e5e48e8ad5c1f595293d8134814cdfc0277ea23d

Observation 32e09caf-2a45-4896-9ecb-97a749b894da · outbound

This paper cites Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis.

Decision Flow Policy Optimization Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.229974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:39.252580Z digest=sha256:be8ee92d3bb99d6e2cb8f2fc1250b684a565ce7529621126642b49b8b27dd341

Observation f3f5e4cf-2988-4a3b-8c56-6046e29e676e · outbound

This paper cites AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners.

Decision Flow Policy Optimization AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.290915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.290915Z digest=sha256:412080055c4177e9343eae4c296f03e74ead450b887fde1dedb78b0639d34c08

Observation 59e99e51-59f7-4256-8a97-fd6895b8c622 · outbound

This paper cites Dataset distillation for offline reinforcement learning.

Decision Flow Policy Optimization Dataset distillation for offline reinforcement learning

Reference 58

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:20:43.045019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:39.370378Z digest=sha256:57eeb3f5f6ba1f518d6c91ac28551ef08860585eeb296e9efbbc5d0cfe8f6081

Observation ab098d77-bed5-4008-b1fa-e314f1024513 · outbound

This paper cites Flow Matching for Generative Modeling.

Decision Flow Policy Optimization Flow Matching for Generative Modeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.454108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.454108Z digest=sha256:9ec1fdfe81b9099146dd19f64a681cbef22a0e4edfc54b34a64bd0ab05591cd0

Observation b8d5be13-e190-4420-89b9-8fb618ada7a5 · outbound

This paper cites Flow Matching Guide and Code.

Decision Flow Policy Optimization Flow Matching Guide and Code

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.518713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.518713Z digest=sha256:f7b22848a4ad41622ffa8483187aa550bc6b2947876318251a5db34331811c9d

Observation 9233db61-bb90-43e3-804e-88b4e89d4d0b · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

Decision Flow Policy Optimization Generative Pre-training for Speech with Flow Matching

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.553836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.553836Z digest=sha256:bddd45c956a2d014865cf58bf671f95cc82c83e2fc02aa758c92362877548315

Observation 4f762be6-8ef3-4a64-8971-1ee7b7732092 · outbound

This paper cites SelfBC: Self Behavior Cloning for Offline Reinforcement Learning.

Decision Flow Policy Optimization SelfBC: Self Behavior Cloning for Offline Reinforcement Learning

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.862741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:39.589213Z digest=sha256:dd06f5a1c417aaa161351906e24e85e04060b764f9b3f051cd18d10ee7a17a59

Observation bfb96da5-a6e1-454f-9a76-5ed920a99fc4 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Decision Flow Policy Optimization Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.668066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.668066Z digest=sha256:41cf0d6ef0dc68603a5a2d26b27f40c3355753dd05f6f37771297a0a1d00f256

Observation 5d8b44d6-4635-4d16-b8a7-05a935b8b799 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Decision Flow Policy Optimization Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.718510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.718510Z digest=sha256:ab9e753f4238ec138698e86faa0999f304ed850ab60359d6356adffdbcecfb16

Observation 8ad00c81-8a7f-4c6f-aa43-aa704d9760cd · outbound

This paper cites Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning.

Decision Flow Policy Optimization Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.813467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.813467Z digest=sha256:4443bd9d97224f2c3d68414b6b338f2417925be7096c40b0498f622276e5c6e2

Observation 38f3ad79-c04a-4fb3-9f11-329072bf7a95 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Decision Flow Policy Optimization AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.855798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.855798Z digest=sha256:8dbd724be4c0198cd3b40c0874206c61792360bb29ef5f8d62a3ec01cdddc96c

Observation 85234d8d-2dcc-48bf-9bb7-ecccdb2ee88d · outbound

This paper cites Normalizing flows for probabilistic modeling and inference.

Decision Flow Policy Optimization Normalizing flows for probabilistic modeling and inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:39.960749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:39.960749Z digest=sha256:fe3b1f3c3c16fbd9549a8fa9c84b199cdf0e133e0c5bf661b3099e33bba8fb02

Observation c9adb103-e45c-4af9-8025-061ce3b280dd · outbound

This paper cites Flow Q-Learning.

Decision Flow Policy Optimization Flow Q-Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.029877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.029877Z digest=sha256:4eb569e18658d08adddee029eb236b65dc8663b67100eddd3085f03a4928fdd5

Observation 7573d0f4-41ed-4f64-90d9-6fbf94a9db52 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Decision Flow Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.119405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.119405Z digest=sha256:bb680b8fd99eac13a54187a15a43e083f7f0a2b64c291c7c83b11dd9253a80e3

Observation b6a752ef-5d03-40be-9253-bb2ee4a3ab12 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Decision Flow Policy Optimization Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.163607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.163607Z digest=sha256:f5afcc76b89d0c6bedaa05cd06acb45dbb839e377dce5861f708607299adcdc2

Observation 294b33b5-01f2-46d9-8400-56e849367032 · outbound

This paper cites FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching.

Decision Flow Policy Optimization FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.249756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.249756Z digest=sha256:46fdd547d77815a1287031eb24393beda71bbc264c1779d8e9df0b5455d59109

Observation e4601625-e387-4bc5-850e-c7d99470c718 · outbound

This paper cites Offline reinforcement learning as anti-exploration.

Decision Flow Policy Optimization Offline reinforcement learning as anti-exploration

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.103384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.346131Z digest=sha256:dbb104420e1ebb78fec795d24ac74d5d75122786e89f41921b6b03ce42e96bfb

Observation 1359d8ca-070a-490e-b627-436869ba5b16 · outbound

This paper cites Flow matching imitation learning for multi-support manipulation.

Decision Flow Policy Optimization Flow matching imitation learning for multi-support manipulation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:45.033689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.424534Z digest=sha256:ab00d8c3f3454af7a1fe1aa22aabf4e81e542b60c380a5b09939674727b5d847

Observation 3dd87819-0a97-400f-b2ca-18147394877c · outbound

This paper cites Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning.

Decision Flow Policy Optimization Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.480698Z digest=sha256:923b96884bb1e9b03ddb1e1851fa6face7491cba9aab58837807f626a69e0381

Observation 95c2ff49-a3ee-44e9-b473-67f0eae07cd3 · outbound

This paper cites Video prediction by modeling videos as continu- ous multi-dimensional processes.

Decision Flow Policy Optimization Video prediction by modeling videos as continu- ous multi-dimensional processes

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.881671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.557491Z digest=sha256:4178e4dfc350a2f7e52ef8289040e32173f28159e880becd4be30475a1e9cf58

Observation a9556fd5-b4ff-43c3-ac7f-117dbc136841 · outbound

This paper cites Ensemble reinforcement learning: A survey.

Decision Flow Policy Optimization Ensemble reinforcement learning: A survey

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.775098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.620346Z digest=sha256:e9bf8da36ec1f37df993dad70aa687cfd1ed3ffca10510b7d48f57861f1825a4

Observation 21c57fae-dcc9-4d7d-b384-60b51ad6ebd2 · outbound

This paper cites Flowllm: Flow matching for material generation with large language models as base distributions.

Decision Flow Policy Optimization Flowllm: Flow matching for material generation with large language models as base distributions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.668710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.692841Z digest=sha256:b9488bf317b0085dabb787321ede570aeefad91eb01ab6adbe374c5ffe7f97df

Observation c137d2e6-7826-482f-a691-963b72e1096a · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Decision Flow Policy Optimization Reinforcement learning: An introduction, volume 1

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.739672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.739672Z digest=sha256:dc3fd9cd43d88739ef6cecd666e63cbd16179da746c0fcdcfbe042b13725bb00

Observation f4f26b48-b58e-421b-a12c-173f4c8ed829 · outbound

This paper cites Imitationflow: Learning deep stable stochastic dynamic systems by normalizing flows.

Decision Flow Policy Optimization Imitationflow: Learning deep stable stochastic dynamic systems by normalizing flows

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.526585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:40.800133Z digest=sha256:7b8f48b23594982fce3902cd65bb40b496fb8673b85cc0eded5b154b2f873adb

Observation 791942d1-952a-4327-8c45-833414bec49b · outbound

This paper cites Deep reinforcement learning with double q-learning.

Decision Flow Policy Optimization Deep reinforcement learning with double q-learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:40.893058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:40.893058Z digest=sha256:080473425d368146bf8f51d434eb7605bd668500bc4d706b06c29709d85a40d3

Observation 4141acbe-6d66-4c89-9103-23b54f1a66c8 · outbound

This paper cites Matrix calculus operations and taylor expansions.

Decision Flow Policy Optimization Matrix calculus operations and taylor expansions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.377627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:41.012416Z digest=sha256:4cfb695f41b2f51c9b04f7b6380fe0b9a2a00f510982c6d8669a04403b22b91e

Observation 4abbb693-68da-4591-a9cf-10b94f6997cf · outbound

This paper cites Bootstrapped Transformer for Offline Reinforcement Learning.

Decision Flow Policy Optimization Bootstrapped Transformer for Offline Reinforcement Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.534167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:41.057275Z digest=sha256:86a6386a59cedc05fbb84afdd634ce3211c5862523275486a7b23ad0e4c84014

Observation e161bd7c-6c5e-48ef-a647-89c905ddf6f5 · outbound

This paper cites Prioritized Generative Replay.

Decision Flow Policy Optimization Prioritized Generative Replay

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.144029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.144029Z digest=sha256:b1dafc1c81431ea4d5468f40c17afd30203911d51d06c6e5db7c151820d93b78

Observation d52a9e49-a01c-4f4b-9059-477a16ecc566 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Decision Flow Policy Optimization Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.247019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.247019Z digest=sha256:bb29806ea14164503dd95e70754c05d21c6a8f30d2411d43b53a43220b1ab215

Observation b99e1af5-4281-4452-af04-637c5620b629 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Decision Flow Policy Optimization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.291215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:41.312573Z digest=sha256:be4ca0d6f25ddced72e710c57a7002bbf89360d98ff70048201e085cda72b4e3

Observation 432d8a0e-306f-423c-b0bf-dd14ab8a6b5e · outbound

This paper cites A Behavior Regularized Implicit Policy for Offline Reinforcement Learning.

Decision Flow Policy Optimization A Behavior Regularized Implicit Policy for Offline Reinforcement Learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.361128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.361128Z digest=sha256:3aabb234577c7fae6b57825f6f761025b3aec11f066f8cb9eb689f6f7826dea6

Observation 9bc740fc-e567-476a-b66e-cd6ed73ada50 · outbound

This paper cites Policy-to-language: Train llms to explain decisions with flow-matching generated rewards.

Decision Flow Policy Optimization Policy-to-language: Train llms to explain decisions with flow-matching generated rewards

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.429401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.429401Z digest=sha256:1375e60276bf45693097ab1175a0f8ae55d107ddc884096e74b83755e96a2cf7

Observation fb0554b1-a476-4222-a6be-78aec2d74647 · outbound

This paper cites Flow to control: Offline reinforcement learning with lossless primitive discovery.

Decision Flow Policy Optimization Flow to control: Offline reinforcement learning with lossless primitive discovery

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:44.222293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:41.499645Z digest=sha256:ad0d0b9637d0f831c7b2afd65f4b7fbb37f47fc058af0ffcd5d1b7dd808fe83e

Observation f99fefd8-8556-41f7-a0d0-c7de39ef5974 · outbound

This paper cites Mopo: Model-based offline policy optimization.

Decision Flow Policy Optimization Mopo: Model-based offline policy optimization

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.575848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.575848Z digest=sha256:e510cef8ee724e804aab8fa94666b3f2e8ccfc19afe93250c9e4b99d07b461a9

Observation 1598d88e-dd30-43c9-beef-1edfb68bf0ce · outbound

This paper cites Combo: Conservative offline model-based policy optimization.

Decision Flow Policy Optimization Combo: Conservative offline model-based policy optimization

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.631335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.631335Z digest=sha256:45a09eb4814e803d68ec49320c1ec2ce80c21290a04952e07de3d7c936f2600a

Observation 57bd9d08-6527-423c-9fec-719af8a1f20a · outbound

This paper cites Affordance-based robot manipulation with flow matching.

Decision Flow Policy Optimization Affordance-based robot manipulation with flow matching

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.678599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.678599Z digest=sha256:69cfed1b2b390671d699b747952eb974378a0c05c2286445913fbf539cb55f37

Observation 0840203d-2792-4814-b7b9-6059b49ff997 · outbound

This paper cites SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning.

Decision Flow Policy Optimization SaFormer: A Conditional Sequence Modeling Approach to Offline Safe Reinforcement Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.739960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.739960Z digest=sha256:7f3e4aa4ac998b90fd4400e9ef73383d163d62259c193d101c29c7feee56de0a

Observation 863b8cb8-2f9d-4b82-9b87-409e89188567 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Decision Flow Policy Optimization Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.812673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.812673Z digest=sha256:4e91e0150573a014e9d3e2e096034008179e92739ed4291836d739ca9daf5d8e

Observation 1319d61f-646e-43b1-acc6-e301f0bad3a3 · outbound

This paper cites Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning.

Decision Flow Policy Optimization Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:20:42.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:20:41.895384Z digest=sha256:992b612d26e646d2780ec814f65ae741dcd1aa17cf2de1a60f110f7b162d1e5c

Observation edd50545-d00f-4770-bc6e-9e6cce26abe9 · outbound

This paper cites Guided Flows for Generative Modeling and Decision Making.

Decision Flow Policy Optimization Guided Flows for Generative Modeling and Decision Making

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:41.964655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:41.964655Z digest=sha256:d4e2dcd0b8d732218ac3df4d52b426858f1eb4ed0b36e7b9ca8c8822d1e78b00

Observation f4db1893-30e5-461d-acef-c364e040e909 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Decision Flow Policy Optimization Open-Sora: Democratizing Efficient Video Production for All

Reference 96

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:20:42.023257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:42.023257Z digest=sha256:df0790737871fb2a8ee757bc07391056a78059f4d307a89f7486b5f7fa0098d0

Pith citing papers

No inbound Pith citation observations are available.