Pith. sign in

Paper Citation Record · LEDGER

Score-Based One-step MeanFlow Policy Optimization

As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2605.23365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23365 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:21:52.748214Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:36:10.359049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T14:37:15.953758Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact9
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e16b1662-262e-4316-8611-47a11a2f0a76 · outbound

This paper cites Is Conditional Generative Modeling all you need for Decision-Making?.

Score-Based One-step MeanFlow Policy Optimization Is Conditional Generative Modeling all you need for Decision-Making?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.060227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:73839f83bd00041685b7c49b364018d772955c7e0f8df4a22428a83bfa3b2e28

Observation fe2a9ef8-634b-4dca-a98e-0ff79b082d4e · outbound

This paper cites Iterated denoising energy matching for sampling from boltzmann densities.

Score-Based One-step MeanFlow Policy Optimization Iterated denoising energy matching for sampling from boltzmann densities

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.169906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:ed28a06ec74c6570bf393afe76037f5a6a504dfbaf44fb9f89b078d0020a3cf9

Observation e835e503-70fd-45ff-b5ee-bae68fe6b10f · outbound

This paper cites Score regularized policy optimization through diffusion behavior.

Score-Based One-step MeanFlow Policy Optimization Score regularized policy optimization through diffusion behavior

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.150860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:9fc63dd1e0488eea948a6c66028bef3cbd03ebad9a706bdc6b4c90bbc37b39c2

Observation dac36487-d07e-49c8-9cf7-6fcfb7e5e17d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Score-Based One-step MeanFlow Policy Optimization Diffusion policy: Visuomotor policy learning via action diffusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.154017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:b5d9fd254e568aa70273ab5a10f0b37a0498758e6b615d9b66d8f7510221b3e5

Observation c6566601-a61d-4164-a854-fefbedd48cc9 · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy optimization.

Score-Based One-step MeanFlow Policy Optimization Diffusion-based reinforcement learning via q-weighted variational policy optimization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.137898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:2624ca5dd62d398fad76e6f5615fff2a93da671bda8184740975b75b71d344da

Observation c4fac339-7b83-440c-9ee2-ed955f3510f0 · outbound

This paper cites One step diffusion via shortcut models.

Score-Based One-step MeanFlow Policy Optimization One step diffusion via shortcut models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.156901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:69c445d4999cbc2473c03da6a95d99c64b8636c6f5d42d35fc9d773b003e3a45

Observation 4aebde4d-8f20-4261-b539-b27e5d06e10a · outbound

This paper cites Mean flows for one-step generative modeling.

Score-Based One-step MeanFlow Policy Optimization Mean flows for one-step generative modeling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.160118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:b727f6666136ca4068e8ec1cbce4f8caa6c2c363742932f34a6a42a893ea1912

Observation 80ca5bd5-9953-4cc2-b317-dd9e1246883c · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Score-Based One-step MeanFlow Policy Optimization Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.166758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:3d5e3d3b8b8a0886ffbf30fdac6ca5669cd0ad846dc5673ba8d24477c0874f33

Observation e145d41d-209f-4da1-aa6a-88860d1d8024 · outbound

This paper cites Denoising diffusion probabilistic models.

Score-Based One-step MeanFlow Policy Optimization Denoising diffusion probabilistic models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.173743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:9eb0e012bb2de5ee1b87174c85f6d5898a020db26c60ca786741b663b5acd060

Observation fb889358-7f67-45a3-a65c-8679d722f27d · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Score-Based One-step MeanFlow Policy Optimization Planning with Diffusion for Flexible Behavior Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.049945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:808a9229d1377e5fcd42623260f6634e23e71ddd5531831a0ace304bca015b02

Observation b6ac2a15-732d-442c-b150-1d29aa4051a7 · outbound

This paper cites Prior-guided diffusion planning for offline reinforcement learning.

Score-Based One-step MeanFlow Policy Optimization Prior-guided diffusion planning for offline reinforcement learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:26:39.035451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:6ba5ffeddb75855cc7cbc1ca0c1cc2d7be5d0c8ea3f7085ea7200a2e252196b8

Observation 492b20f6-9de9-4291-bda5-490666b1ca5e · outbound

This paper cites an unresolved cited work.

Score-Based One-step MeanFlow Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-25T11:56:57.122735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:482235909b6c4f3f049484ebad2faddfeaa4d52efd1856b01d9bb690e2f6c3f9

Observation 07029e77-d08d-4d2a-bac3-0bbb973597dc · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Score-Based One-step MeanFlow Policy Optimization Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.108656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:47d611bd0e5701561a825a0cf549bc96f879d04c531aad362d5aa73aac14263a

Observation d8c9407d-4691-446e-8b3d-fdb98f9f369a · outbound

This paper cites Simplifying, stabilizing and scaling continuous-time consistency models.

Score-Based One-step MeanFlow Policy Optimization Simplifying, stabilizing and scaling continuous-time consistency models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.180228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:c8ca5b261efef949781f66b83d8e91cf495ac9617e099a63b1070f4c6a1561a8

Observation e791643d-1a99-4c84-a9fc-2adffdbb1539 · outbound

This paper cites Efficient online reinforcement learning for diffusion policy.

Score-Based One-step MeanFlow Policy Optimization Efficient online reinforcement learning for diffusion policy

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.186599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:4bf936e9e901e9d08e35c1bd220fdc762ded08e12bfee3931b8afb23e2d70a8b

Observation 0b95ac95-a5d2-49c9-ab48-91457004b53e · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Score-Based One-step MeanFlow Policy Optimization Learning a diffusion model policy from rewards via q-score matching

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.141016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:bb3525abc19609d0cd384f3fe8779c1a2aa8eef75f2d551eda1ff23d5d67a1fd

Observation dfd762db-b4d8-4092-912e-22a5e0bae85f · outbound

This paper cites Diffusion Policy Policy Optimization.

Score-Based One-step MeanFlow Policy Optimization Diffusion Policy Policy Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.040336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:d4bc76adcbe64152dd05013c42e0db70dcf96243659e6ddd22ba7da68ab76dbf

Observation d6c1805b-c865-41ad-a931-fde69f2eb2e1 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Score-Based One-step MeanFlow Policy Optimization Progressive distillation for fast sampling of diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.119449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:85e271e0c049d412283d99df629cda8c5d708f4c509b33a4190074163fcf89de

Observation b8694883-1729-4fa4-a207-7653430550c3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Score-Based One-step MeanFlow Policy Optimization Proximal Policy Optimization Algorithms

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.045159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:5157840a41c95b60263aadf8648f11a493c73e39d3b11d7b262674fa88a77832

Observation f15f3bab-bf9b-4d07-bee2-f661700fecd3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Score-Based One-step MeanFlow Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.054865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8136e616f98888fca3250f728007d5ede0d671f6bdd41575c0a298ea8a026d26

Observation 7333592f-794c-4f02-bfc1-5aca43672304 · outbound

This paper cites Denoising Diffusion Implicit Models.

Score-Based One-step MeanFlow Policy Optimization Denoising Diffusion Implicit Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:25:23.791074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:f0abcb3b832327a861b7334de88f542cf9daad6035140365fba16e3401648224

Observation 763f04c7-b259-44b1-b986-4d16d979675a · outbound

This paper cites Consistency models.

Score-Based One-step MeanFlow Policy Optimization Consistency models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.126168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8eaf6ffa86763d3acfac15bb764a92a2b1bd9aa557759b2229661f5d8b726f12

Observation 8b9826e1-7c37-4249-807e-6aacddd5eddb · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.Advances in Neural Information Processing Systems, 32.

Score-Based One-step MeanFlow Policy Optimization Generative modeling by estimating gradients of the data distribution.Advances in Neural Information Processing Systems, 32

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.177176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:54519c229ee735ae496f0aad106f8be7445a3ef06b2a716731933c4f2eea29b2

Observation 724ac68d-ab5e-4319-a08a-376234bc98d7 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Score-Based One-step MeanFlow Policy Optimization Score-based generative modeling through stochastic differential equations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.147342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:40a5ec14b5fd863ae54e94000d68f8be934706b2dc727d02cb69bb18f078e41c

Observation 589c7b6b-05c1-41fd-a090-a312414a711e · outbound

This paper cites MIT press Cambridge.

Score-Based One-step MeanFlow Policy Optimization MIT press Cambridge

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.115682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:0bf6f61c3bd1b068d6a16c685620120bc531a6638d1dd40503f25aa882b3464a

Observation 0d1e467d-bd2a-4a05-968b-fa0232021a69 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Score-Based One-step MeanFlow Policy Optimization Mujoco: A physics engine for model-based control

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.183032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:82be16e096acc045b9abf51ac94613d5dfa1ec8167132fccc0436788633151e9

Observation 2fb51365-d222-4e43-98f5-7ba26eb10f64 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Score-Based One-step MeanFlow Policy Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.029853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:c0d8f39eda814822c4f68950a8f248cc3bfffb6b8be1d67c3b12bae91e99c109

Observation 8ccdbf9e-525a-4d85-b8e7-04ee27ca024f · outbound

This paper cites Diffusion actor-critic with entropy regulator.

Score-Based One-step MeanFlow Policy Optimization Diffusion actor-critic with entropy regulator

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.163332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:5c4364a42b4ad4200ddb27c14f055aabbf0076df9afeac478ecb9eb2d5ddb62e

Observation 6fcb2bcd-c38f-441e-b0a7-ca7b2faff9cb · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Score-Based One-step MeanFlow Policy Optimization Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.130845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:6bfa9b8fb686877fec87a6e5a558f423cb95e4f7d0255ac31fb1da0e6e93dd09

Observation c49ab6de-e540-40d5-97ee-f5641be5a5e2 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Score-Based One-step MeanFlow Policy Optimization Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:26:39.024863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:480d59370830b273db5f924acbed74c298a219db73d064bdda32810d07dec5aa

Observation e1970c2d-ee3b-4ff8-a082-f1acb63f6d56 · outbound

This paper cites One-step diffusion with distribution matching distillation.

Score-Based One-step MeanFlow Policy Optimization One-step diffusion with distribution matching distillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.144048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:43c60f2dcd5b6aa61394064a63d2fafc58b434021464f87ddb5447e0d221f7de

Observation 1ffc8f5f-c9b4-4a12-86ba-e4dc119b479b · outbound

This paper cites Mean flow policy with instantaneous velocity constraint for one-step action generation.

Score-Based One-step MeanFlow Policy Optimization Mean flow policy with instantaneous velocity constraint for one-step action generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.134704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:cf8bcab122a402cc1cbf52d80dfe147c780d8e9ebeb3fe0b2a0ec35fabf83e99

Observation 23a9e2cc-d14d-47f3-95c9-589d4565155f · outbound

This paper cites The final reward is given by the normalized mixture density, producing a smooth multimodal reward landscape with values in[0,1].

Score-Based One-step MeanFlow Policy Optimization The final reward is given by the normalized mixture density, producing a smooth multimodal reward landscape with values in[0,1]

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T11:56:57.112390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:f9b5d89fdd05bcadd9fe0810003106e16741b536c86473c172cc7dfbfe7c8bc2

Pith citing papers

Observation c8575dd3-97fd-4132-989d-5617a15fd9d5 · inbound

Expressivity and Statistical Trade-offs in Diffusion Policy Learning cites this paper.

Expressivity and Statistical Trade-offs in Diffusion Policy Learning Score-Based One-step MeanFlow Policy Optimization

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:37:15.956480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T14:36:10.359049Z digest=sha256:af058e9e35fa7ac5eefb0b590662c8e927f84cae8e7fcc5c53fdef030165de3a