Pith. sign in

Paper Citation Record · LEDGER

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.04372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04372 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:22:38.305872Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd3082b0-8c74-4f96-9715-d129ba5474e2 · outbound

This paper cites (2017).First-order methods in optimization.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology (2017).First-order methods in optimization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.695460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.237900Z digest=sha256:c7aee5a23ea1e902e0c2071d62be4b81c2e966fe9868d630cfdb9e084e68e65a

Observation c0a7c2e3-2b2c-4633-a31a-a452fca5ac6c · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Training Diffusion Models with Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.242023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.242023Z digest=sha256:50feb8bb7196d7e103dac2ac47e6772f3cd9b8e994185ac88f2e0c9a28a08f66

Observation 2d7a47e9-e20b-47e4-a3af-42b94d32b2f8 · outbound

This paper cites Classifier-Free Guidance is a Predictor-Corrector.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Classifier-Free Guidance is a Predictor-Corrector

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.246154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.246154Z digest=sha256:98f70e888a33db1f20e8453d26fe9f821f8ca8f26eb70c7411207ea782376698

Observation f5316555-f017-4eb3-997b-825b26f673a9 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.686324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.249522Z digest=sha256:bd70190b67ee93ca2c91653afcf4976ced0b1160d097beb4b47ad2c6c81cb04b

Observation 7cb71b9e-b355-49db-8a2a-5e754e480e46 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.252791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.252791Z digest=sha256:06dd7b96af158897e63c1ba6394a1118bb0ae3e624d599b872f1d9c7488ac687

Observation 2bc5fd5c-f9d7-43fd-801f-395a9864940d · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.676443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.256587Z digest=sha256:b7fec7d11567cdd57956ee35a3b53870994338a1917d3fa8916b217aff36f382

Observation d25f6725-d7f1-4125-aef6-f7c103f4a153 · outbound

This paper cites and Nichol, A.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Nichol, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.665920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.259858Z digest=sha256:81d711dbc32b75a114ecde0271f6737072da1decefc85265d5a62ca521729d64

Observation 52fd0fb8-f6fb-40c0-8e73-058ed5d255e2 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.655548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.262830Z digest=sha256:c9c3a7c31814f7a4f9f4172b2a1fa84947549d9a4b4939a4666e71751711881f

Observation 87b7f893-9cbd-424d-822b-08be91ae1a17 · outbound

This paper cites Reward-Directed Score-Based Diffusion Models via q-Learning.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Reward-Directed Score-Based Diffusion Models via q-Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.265706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.265706Z digest=sha256:ca0baf7b8e1e6340bdcd4d76b4d1989f74aabd03dfaf538d13d04f82c677a77f

Observation 85aa4d1d-64c5-48f1-a16c-1f46a4dd569d · outbound

This paper cites and Salimans, T.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Salimans, T

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.645855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.269327Z digest=sha256:3b1f0cf019c78061c6b8f80bd7e07c6d921bd6acb90ce10e6c84016031718f35

Observation a6e5b84c-2ae1-426f-a044-fad4ddd19f8a · outbound

This paper cites and Martin, J.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology and Martin, J

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:22:38.635292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.272663Z digest=sha256:30569a3af9edf1db9ffbb61d2a1e6bea7b701d53cc5fb1d1d7e9ad12202147e6

Observation cb943ee2-cf26-40a3-b4b8-b7879a812195 · outbound

This paper cites Unfamiliar Finetuning Examples Control How Language Models Hallucinate.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.275970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.275970Z digest=sha256:d9e057c37d0648633b0a6a7a32107da8cd426519dab1ceec21a80e680f513530

Observation f72b04ac-3d8e-4a12-99b1-9aa8f46544ba · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.279540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.279540Z digest=sha256:cd2767553e1e64899bd3e2eae0cc27b01e3376b2478b9c5064f090aae25ed0ec

Observation 4add29e6-a4f2-4a2b-bfaf-c9620c485b75 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.282693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.282693Z digest=sha256:f37cec8f64f916eaf9c4a7cf866c709656c84e5c22d952c61609374f3b489199

Observation 22f5eaa6-73ef-45f9-8a11-9d7b3ef3a84d · outbound

This paper cites A Score-Based Density Formula, with Applications in Diffusion Generative Models.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology A Score-Based Density Formula, with Applications in Diffusion Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:22:38.401166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.285588Z digest=sha256:d98db4bf27d807daf379ffd79bd19f664a92365fc1f945f84f512cccfea28af5

Observation 30c8b382-262b-445b-bdfe-55857c7d060b · outbound

This paper cites s1: Simple test-time scaling.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology s1: Simple test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.289020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.289020Z digest=sha256:68539c45e499889eacb3440d2b4ee2cb266b088efce5f0b2a8fb1496998c1e3b

Observation 205eccfc-630f-4c06-9fda-1d0ef19f56d4 · outbound

This paper cites an unresolved cited work.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:22:38.624367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T10:22:38.292172Z digest=sha256:b7f51cde1c295743b088b4da1b116dc0397d8dba853dcacc6951355aa810675d

Observation 816f2ba0-a9dc-4797-a14e-7bedb121976a · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.294929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.294929Z digest=sha256:9e91e2db16e01a01a7ccf8b1af6a82403598baf64e3bf8afe0d4fc47d4537fa7

Observation b4018f26-4601-459e-a03d-bbf657a72bf4 · outbound

This paper cites Soft Best-of-n Sampling for Model Alignment.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Soft Best-of-n Sampling for Model Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.298599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.298599Z digest=sha256:93128bf66d3bf3bc4a0a5dc75dd79c336f1dd71920418a6a9a548d87a48b85a1

Observation 54f756f7-d16c-4324-b4af-d85ed1f561e5 · outbound

This paper cites Learning to Reason without External Rewards.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Learning to Reason without External Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.301973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.301973Z digest=sha256:224733a1af70b319860b7d4662ac664dce853b9ef3e45bf90fd231e24c7e33cf

Observation 9b663c7f-160e-49b9-bad5-fbae983d3217 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology Fine-Tuning Language Models from Human Preferences

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:22:38.305872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:22:38.305872Z digest=sha256:2cd97fb017b68ad64ba286654288b18dd013ee687678ea053ec3e826b9294c3e

Pith citing papers

No inbound Pith citation observations are available.