Pith. sign in

Paper Citation Record · LEDGER

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.04372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04372 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:22:38.305872Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd3082b0-8c74-4f96-9715-d129ba5474e2 · outbound

This paper cites (2017).First-order methods in optimization.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology (2017).First-order methods in optimization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.695460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.237900Z digest=sha256:729899da4655f66f6c8401bc34e541c50e2e225110e3c15ca3a2ef1104f59244

Observation c0a7c2e3-2b2c-4633-a31a-a452fca5ac6c · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Training Diffusion Models with Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.242023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.242023Z digest=sha256:4374f4678be0fcdad572a0167e7331f46ad8678ae9fc5758c2ae201307a44584

Observation 2d7a47e9-e20b-47e4-a3af-42b94d32b2f8 · outbound

This paper cites Classifier-Free Guidance is a Predictor-Corrector.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Classifier-Free Guidance is a Predictor-Corrector

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.246154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.246154Z digest=sha256:8dc9adc78684bbae440f60cf2a804a67745b1c00ebef51709791c655cdb83164

Observation f5316555-f017-4eb3-997b-825b26f673a9 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.686324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.249522Z digest=sha256:6a978b713a9b440f0cf5290837c328fd1ca14c6616ae486c69d8d696195759c6

Observation 7cb71b9e-b355-49db-8a2a-5e754e480e46 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.252791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.252791Z digest=sha256:fd8ef2d14cfe7dfdbb10289a7f0f9afb56465ef3e8b3ce39e2b6506269311abf

Observation 2bc5fd5c-f9d7-43fd-801f-395a9864940d · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.676443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.256587Z digest=sha256:c390769f21770fcc925a3f0fa070157f8f6759525c187c1c26b7bd62fabf733f

Observation d25f6725-d7f1-4125-aef6-f7c103f4a153 · outbound

This paper cites and Nichol, A.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Nichol, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.665920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.259858Z digest=sha256:5d9cbbb651206d2799b6bd564ef09406d2669602e3f0f28fe39a5fbd4e4a6f23

Observation 52fd0fb8-f6fb-40c0-8e73-058ed5d255e2 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.655548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.262830Z digest=sha256:e9720467d772d7328e0eb3529511bf8afc5c694967ecce0570ddda1f66b57f0b

Observation 87b7f893-9cbd-424d-822b-08be91ae1a17 · outbound

This paper cites Reward-Directed Score-Based Diffusion Models via q-Learning.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Reward-Directed Score-Based Diffusion Models via q-Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.265706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.265706Z digest=sha256:3ae94c95edf767b85717638c6a20297cd9fb9eb52ffb53e31769f3ad7cd73207

Observation 85aa4d1d-64c5-48f1-a16c-1f46a4dd569d · outbound

This paper cites and Salimans, T.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Salimans, T

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.645855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.269327Z digest=sha256:867ce220740c11c2721c8a83cfc18075706cd84c9ce26b086b331ea00ff2a6b2

Observation a6e5b84c-2ae1-426f-a044-fad4ddd19f8a · outbound

This paper cites and Martin, J.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Martin, J

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.635292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.272663Z digest=sha256:29cd72816140ff2612620acef1d159a5f950bc08a34a469163ba70e4620c22bb

Observation cb943ee2-cf26-40a3-b4b8-b7879a812195 · outbound

This paper cites Unfamiliar Finetuning Examples Control How Language Models Hallucinate.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.275970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.275970Z digest=sha256:77d0191b9a4b7cef8b3feae49937294568c34c21696cf6a238c0afb1df5bbdb0

Observation f72b04ac-3d8e-4a12-99b1-9aa8f46544ba · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.279540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.279540Z digest=sha256:9a90a911e92c97277779c303ed28d2c08292858b73aedd0c5b17a04afe34a60c

Observation 4add29e6-a4f2-4a2b-bfaf-c9620c485b75 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.282693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.282693Z digest=sha256:4fa00cca61dee307097652e4f623f0683d092830ebe8450d6b88fc6cad796380

Observation 22f5eaa6-73ef-45f9-8a11-9d7b3ef3a84d · outbound

This paper cites A Score-Based Density Formula, with Applications in Diffusion Generative Models.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology A Score-Based Density Formula, with Applications in Diffusion Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:22:38.401166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.285588Z digest=sha256:40b4008b7871bb35fba18a9d15a4b698d100b77f4cbee8b16a2b83b702f62402

Observation 30c8b382-262b-445b-bdfe-55857c7d060b · outbound

This paper cites s1: Simple test-time scaling.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology s1: Simple test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.289020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.289020Z digest=sha256:690cf02bc8efdd183d2cd127ebdcc2da8bd28dd0e24099986ec6ceb9f718d9b9

Observation 205eccfc-630f-4c06-9fda-1d0ef19f56d4 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.624367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T10:22:38.292172Z digest=sha256:f3d2feee0496671a1228943cddf0fa9cef614fe1aaf10feabedc199dfd758d9c

Observation 816f2ba0-a9dc-4797-a14e-7bedb121976a · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.294929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.294929Z digest=sha256:c5ee421c12bcfc7b02348e9a7abcb758cef8b3637c93710063dabf1f4b33db17

Observation b4018f26-4601-459e-a03d-bbf657a72bf4 · outbound

This paper cites Soft Best-of-n Sampling for Model Alignment.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Soft Best-of-n Sampling for Model Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.298599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.298599Z digest=sha256:7524e1e101f7e999d3641c65368423dbc2c55de84f15c0993ffc733d20352538

Observation 54f756f7-d16c-4324-b4af-d85ed1f561e5 · outbound

This paper cites Learning to Reason without External Rewards.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Learning to Reason without External Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.301973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.301973Z digest=sha256:e244bc7caf861e84050238501dc2ed456ff8ffcadf3abca49d50d3534ab7bb45

Observation 9b663c7f-160e-49b9-bad5-fbae983d3217 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Fine-Tuning Language Models from Human Preferences

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.305872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.305872Z digest=sha256:e3cc42405f27d7a7aec33f18e44a70e4a1f5dd00b6a859e1e491bc5b6f2fabfc

Pith citing papers

No inbound Pith citation observations are available.