Pith. sign in

Paper Citation Record · LEDGER

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.05084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05084 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:50:32.825035Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d6fe352-65b3-4fc9-a831-d405f70baa40 · outbound

This paper cites Robot Skill Adaptation via Soft Actor-Critic Gaussian Mixture Models.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Robot Skill Adaptation via Soft Actor-Critic Gaussian Mixture Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T05:50:32.959913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.772675Z digest=sha256:59c3c7023b6a99f415d47bcf50c4cb06623521b7b17b64a1e9684af85585fe8b

Observation aec1490e-e744-4695-8222-7d85c8b81fc2 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Progressive Distillation for Fast Sampling of Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.777893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.777893Z digest=sha256:ba5993c1c73b4b5421c74452bcbf4f54bff30c8b84aea00968941b43a295f186

Observation c58f4d21-ee59-48bb-8465-c6e5955f8c29 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.782847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.782847Z digest=sha256:00e451e4490746d33a1d8ecf69798ac4026dd1e9b95db945c79013f0d52aa997

Observation 8bc2b8c3-8f39-4b8d-9e48-ce42806e35a5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.787581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.787581Z digest=sha256:45fb17e95dc8325563d73e9bc95411dfe9b19a8aec06b507655401aba5378c4e

Observation 8c8e7c4d-e0a2-4592-8bee-573369f9246b · outbound

This paper cites Diffusion Actor-Critic with Entropy Regulator.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Diffusion Actor-Critic with Entropy Regulator

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.792214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.792214Z digest=sha256:047e2bac960ee6e0bf9eadff06e54c6180e9b886df3379c3016b1dd25ea7d1c2

Observation f86f061b-1113-4748-8b5d-c12bbe3b94c6 · outbound

This paper cites D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.796809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.796809Z digest=sha256:0da867ebff67af572b15289889b236e07f38f21125b33f10c6bbfb250762ccd1

Observation 039b18b9-b8b4-43cf-8529-8f76ffed2b58 · outbound

This paper cites an unresolved cited work.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:50:33.115174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.806731Z digest=sha256:b2406578dad545cdbaf103432350d0200fcca061277be5812053a892e994b611

Observation 10bd7a1f-1fa2-4a8a-90a1-661a50480be0 · outbound

This paper cites distill toK ′=5.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control distill toK ′=5

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:50:33.099910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.811378Z digest=sha256:473ce8be1f457166fbcbcb3d479f704d88535d8077627b1d7e115d184773cc42

Observation 4cfb098c-6bbc-4262-b812-fb744fdd4d1e · outbound

This paper cites (5) with ¯h(t) = Eh∼H[h(t)].

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control (5) with ¯h(t) = Eh∼H[h(t)]

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:50:33.083329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.815906Z digest=sha256:503929addf96df797322f93bc564cf599e906ba77150bb7090021d2f3b36c558

Observation c400fc68-4903-4464-b5a9-5ad237fc9371 · outbound

This paper cites •Diffusion-QL: Diffusion policy withQ-weighted behavioral cloning loss (Wang et al., 2022).

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control •Diffusion-QL: Diffusion policy withQ-weighted behavioral cloning loss (Wang et al., 2022)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:50:33.068241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.820519Z digest=sha256:cd7d5bd39cba00d135b558e9a33361d1ce0a910a52d05b5c9ef3ad1597e33420

Observation 80532993-068c-4edc-ab2d-e0ad91274796 · outbound

This paper cites URLhttps://ojs.aaai.org/aimagazine/ index.php/aimagazine/article/view/1232.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control URLhttps://ojs.aaai.org/aimagazine/ index.php/aimagazine/article/view/1232

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.801772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.801772Z digest=sha256:55bf3f875908391ca8ec8b3b30343e2a425cfe9da10f66a6367604d6db5a5a19

Observation 4d0613af-ee6b-4e8a-af83-0e58defe5d01 · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Planning with Diffusion for Flexible Behavior Synthesis

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.761485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.761485Z digest=sha256:4eade68f58950812d275826348a5027bf1c3046b96c0de9714c4cfe6f252e0c8

Observation f80ad7ee-1301-4306-b5b5-7ba0be90dc70 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Adaptive Computation Time for Recurrent Neural Networks

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.756288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.756288Z digest=sha256:7e1d3bb366b18d20af56067b33dc6774f968d27e7badd831a01296d7537a6069

Observation e3a32a21-b656-4168-8f59-49706f80350a · outbound

This paper cites Deep Reinforcement Learning at the Edge of the Statistical Precipice.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Deep Reinforcement Learning at the Edge of the Statistical Precipice

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.745229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.745229Z digest=sha256:8be0f99245f9a80a44b5d42ba09a2126d7d842701a221f5655f305438b1bbe6f

Observation 9b4be681-ab94-4d11-9827-346af1a173db · outbound

This paper cites •FQL: FlowQ-learning (Wildberger et al., 2023); the flow-based policy is trained with a reflow objective to enable one-step action generation.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control •FQL: FlowQ-learning (Wildberger et al., 2023); the flow-based policy is trained with a reflow objective to enable one-step action generation

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:50:33.052308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:50:32.825035Z digest=sha256:cd9680dadcb489536855ecc847a67b4d9a4ed991edbe6eec00846610f44a8a9b

Observation d165c37c-29ca-4fdb-aedb-e34e36565f09 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.751225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.751225Z digest=sha256:e1baefada881cbaccb68febe3733a449495ce7be3669300fa523bef4edbc0931

Observation 15117198-4f4d-479b-8af2-c7be9bf41625 · outbound

This paper cites Distributional Soft Actor-Critic with Diffusion Policy.

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Distributional Soft Actor-Critic with Diffusion Policy

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T05:50:32.767520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:50:32.767520Z digest=sha256:d990ea8424c4b2e308a151fa73fee2ee495204edc88684707e895cb4e3b5172e

Pith citing papers

No inbound Pith citation observations are available.