Pith. sign in

Paper Citation Record · LEDGER

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2505.19196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19196 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:26.232105Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:33.490433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.983158Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a59cfb8-d6c1-4a4a-98d1-7b9a0c718506 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Deep unsupervised learning using nonequilibrium thermodynamics,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.824017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.059212Z digest=sha256:795a75adb75e376c74ab5283324c090e8d6019049703bb4d78af679b14090b62

Observation b9b42b1d-c934-4fe3-a11f-9245bc2f86dc · outbound

This paper cites Denoising diffusion probabilistic models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Denoising diffusion probabilistic models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.062747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.062747Z digest=sha256:9c792b6718d43cfb4cafae130f45221a987cfa3e693ebf219007e01131f9a26f

Observation f9f7afaf-0e9b-44a9-8d97-aa9f714e2f5e · outbound

This paper cites Denoising Diffusion Implicit Models.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Denoising Diffusion Implicit Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.066590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.066590Z digest=sha256:c740ff78b1873d6d0133af83c5397b8288cc57f1a1e94146f15cf188340b66ea

Observation 616f4244-65ae-41b0-af70-5fc41b127016 · outbound

This paper cites Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.069859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.069859Z digest=sha256:b8bdfbbfc0d62ac98402de54f07594272a300b24be55f43b00e8b76c3f973147

Observation c9ba1b6e-ed7a-4f4a-8446-0f66494ba7c5 · outbound

This paper cites Generative Adversarial Networks.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Generative Adversarial Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.072893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.072893Z digest=sha256:6a44fe68104af7b09c8c7b79bb178e146b8660d2d72be7a157b3d848ccec4129

Observation 36fc2aa5-ff92-4347-bb87-46614c818290 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Learning transferable visual models from natural language supervision,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.808598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.076062Z digest=sha256:42f5aad3efee228b58024e2d6c193831cc845cc8d5174d89f14bba3a29afe774

Observation a8cf128a-03e1-4295-ba87-01beadc4c0b3 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.797986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.079004Z digest=sha256:507efa30708d6a4aa5dc9f76ba5f90a05985d9f53290405de27758821d04317a

Observation c2c730e0-bca1-4d26-9780-15e0f95475fe · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.788543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.081959Z digest=sha256:26d05df3ce07df40b5507b9dd0c276633371abc9b65537f0f4f7b5f45fce4ba8

Observation 31e3e112-8969-48c6-af7e-3303037df51c · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Microsoft COCO: Common Objects in Context

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.084718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.084718Z digest=sha256:e074aa74167bd4168c27b8b53947f081ce0105a05320eb04c1810b50549662a0

Observation 5035b145-6b7c-4329-8388-b38f41b8f4b9 · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Imagenet: A large-scale hierarchical image database,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.087407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.087407Z digest=sha256:09db1bfb3a54ecb6627aaf107309f1327c2ae402bad535498803082e56e3dcdd

Observation f5969777-9350-40a4-a3bd-dc39466c7730 · outbound

This paper cites High-resolution image synthe- sis with latent diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning High-resolution image synthe- sis with latent diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.773507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.091066Z digest=sha256:9ef81e1c2088dae49fc8ef3aff054990bbd2e32d4c747aae0419a4f7a110464a

Observation 54a42102-2d77-47b5-891a-c44c17bdd6c9 · outbound

This paper cites Improving image generation with better captions,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Improving image generation with better captions,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.764568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.093961Z digest=sha256:67d32bcff6d5c1174288460d07dbea10aad8c807156ff3c575a646abdde629d2

Observation f7a42b79-f648-40cd-9bd7-d5abb12dc7c7 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.756226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.096690Z digest=sha256:e4d0e28e447550df61d01df877dfcc5f3d49d26717eb2e935e608cbd3408c8ee

Observation 73f0cfd5-16ce-492f-b66b-8a3922494028 · outbound

This paper cites Training-free structured diffusion guidance for compositional text-to-image synthesis,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training-free structured diffusion guidance for compositional text-to-image synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.746068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.099271Z digest=sha256:bb9bf98e8e9b1972866835143450bb170ef3e7a8600b573088c65d88084f2439

Observation ac972a64-6cde-41a2-b6f0-db927337c602 · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.101775Z digest=sha256:e969de95f10d9bebc0de95bbf60deb83c0577e617e3dce290e8ccef013ab79e2

Observation 83dfc91e-b25e-4b5a-997d-134f4b12d7c2 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.104634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.104634Z digest=sha256:af6a8458894a55240340d6561b89cbd2aa6e76eec697637f4e75590f5c7338ae

Observation d9ec2f17-2af9-4049-87ef-f4529c2dd380 · outbound

This paper cites Imagereward: learning and evaluating human preferences for text-to-image generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Imagereward: learning and evaluating human preferences for text-to-image generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.737462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.107547Z digest=sha256:b61f26e300e7ca1ed3719e324f3f8b9caeae99998fc2c6b53781346d32b0dfc9

Observation b4ae061a-1de1-4a54-b92f-fc9145c7f647 · outbound

This paper cites Pick-a-pic: an open dataset of user preferences for text-to-image generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Pick-a-pic: an open dataset of user preferences for text-to-image generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.728358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.110441Z digest=sha256:ae598efd6dd268d5097c6020f5ac6bbe1e05d274d4770df32a7c8fb53bb7bebb

Observation 9c0f8328-a04a-43ae-9372-84076f1c2057 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.112892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.112892Z digest=sha256:a47b9567e91eb4603699a1de0836a0a43cf5400db6a31401004d583ad06fa463

Observation ffc72921-f5e6-4523-9723-d23778d31cd0 · outbound

This paper cites Dpok: reinforcement learning for fine-tuning text-to-image diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Dpok: reinforcement learning for fine-tuning text-to-image diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.719666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.115498Z digest=sha256:af41ca442421578dc40d070b02bf7d6c5ee313379c433194bc041ef083a6e150

Observation a49943ae-083e-4f04-924f-c5f98e89a191 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training Diffusion Models with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.118205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.118205Z digest=sha256:80a146dd0010b6dbbcdeb89dfdf07bc3cd8cc46430999934a98d5b645c19ab08

Observation 3d82d039-20cc-42ab-9944-ced4ffa73f96 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.120984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.120984Z digest=sha256:95bc3f77e9a94ef3d74c302d83f358e80c9f1713566fdc6b0404829ba09debf1

Observation f5d1d4d8-172b-4707-b6b7-55870db39f66 · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Direct prefer- ence optimization: Your language model is secretly a reward model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.710368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.123805Z digest=sha256:0223596e31b1882bb87fe8da4999d1cc59a4113512df84e3cb942ff448956dbc

Observation 0af8666d-0e7e-4f19-a634-d622111e675f · outbound

This paper cites Stimulating diffusion model for image denoising via adaptive embedding and ensembling,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Stimulating diffusion model for image denoising via adaptive embedding and ensembling,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.701020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.126311Z digest=sha256:16f7ba502df57e4063740651201899f80b7929a88e6f9fe60dae2e2146291665

Observation baa652e8-00ce-48a7-bf95-b584a9bdb855 · outbound

This paper cites Blue noise for diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Blue noise for diffusion models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.691274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.129059Z digest=sha256:5360d5af2a4a266563dda3d9d9b9069599b615d974e501092113e9ba21900d4a

Observation 7a764e69-28ab-4476-bac1-3702711a3a7c · outbound

This paper cites Boosting Diffusion Models with Moving Average Sampling in Frequency Domain.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Boosting Diffusion Models with Moving Average Sampling in Frequency Domain

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.132556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.132556Z digest=sha256:828d2a178b2975c521cd40ead0e1e271eb4569d39c6edc322a45ba5ba906a00e

Observation 80e01d5d-5291-4403-a1f0-853487a63e6a · outbound

This paper cites FreSca: Scaling in Frequency Space Enhances Diffusion Models.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning FreSca: Scaling in Frequency Space Enhances Diffusion Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.136256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.136256Z digest=sha256:e8b5b91500162a932807561889dedb97bf347f5da3995844de5316f57befd72b

Observation 293a2962-4691-4cb0-bce0-88f8b86fb19b · outbound

This paper cites A dense reward view on aligning text-to-image diffusion with preference,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning A dense reward view on aligning text-to-image diffusion with preference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.682210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.139255Z digest=sha256:664764e199bc5efe32e0e309317fb1d85ee715b63f613a0ad4e91eaae85bba00

Observation 1f77c922-b895-41d0-a9a4-06d761f17ebe · outbound

This paper cites Confronting reward overoptimization for diffusion models: A perspective of inductive and primacy biases,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Confronting reward overoptimization for diffusion models: A perspective of inductive and primacy biases,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.672920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.141926Z digest=sha256:780bd5ee88a24707e89b429a1e6e37fa42111e4cd43e8bf18a49ddcacffb2fc2

Observation 7ddc2152-916d-4b1e-b35e-103cb7023a17 · outbound

This paper cites Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.144841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.144841Z digest=sha256:f6641ed5249dfb307dba6673c6bd69aba3925ded5c55d544b967c3fae6a48179

Observation 11dccaa8-3582-4b83-b462-f53b6771e70d · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Diffusion models beat gans on image synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.663616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.147939Z digest=sha256:d624bf46010bbc8f979adf7472c41798743f799708ba3965b4c912abe148d9f5

Observation 186e990f-a787-4d59-ab3d-7b884ffe7fd5 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Classifier-Free Diffusion Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.150649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.150649Z digest=sha256:43eb0c09289196c41be8dd7c361667b2c05b57554c3710f0a34459a9eeaba434

Observation edd2dd8b-19b6-4ce7-bea8-43db8f0eb1bd · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Aligning Text-to-Image Models using Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.154161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.154161Z digest=sha256:24ff1ca080fda7ad2cc06fe969707e5bb26591e792a51d8d7a8b202b2d474175

Observation 7795060c-6414-4eb9-821b-c97d09fe8840 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.158191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.158191Z digest=sha256:9de583578318362ba09712c01d7e0ff5875e7bdc7e5be7e827521131c66fcf25

Observation 03783dd0-c41f-4932-9fd6-780499cd33e4 · outbound

This paper cites Optimizing ddpm sampling with shortcut fine-tuning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Optimizing ddpm sampling with shortcut fine-tuning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.653284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.161300Z digest=sha256:9908ab74d877b2f9ae2c088143b38336fc0e02271f7bb369da403017b5de8b7d

Observation 44e5fb01-4654-4a95-ae6e-54baca475030 · outbound

This paper cites Deep reward supervisions for tuning text-to-image diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Deep reward supervisions for tuning text-to-image diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.644330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.164256Z digest=sha256:2757c224883f78e7937cc7f3cd5537f1950ea56d5e0396c493a128d642e7d4ff

Observation 44045a80-0c4d-4cde-b168-b89721392dd0 · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Using human feedback to fine-tune diffusion models without any reward model,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.634403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.167113Z digest=sha256:a583696016bc1284a76769ca41e57c669c724e52e1a69a174ad69ce66f47d9b2

Observation 4e8fbaae-74bc-41a7-8ce7-76a0602f52ed · outbound

This paper cites Diffusion model alignment using direct preference optimization,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Diffusion model alignment using direct preference optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.549875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.170322Z digest=sha256:d398a524f5152c9fd52d8d9c149b0f9b9e0049feff9d15072edc7f2c052398ee

Observation 357d6f4e-48c6-430d-9274-161ff9e8f422 · outbound

This paper cites Steps toward artificial intelligence,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Steps toward artificial intelligence,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.540929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.173109Z digest=sha256:18d2da4e4e0be9afe8872005d1e4d58b8c3e58e30b0979e7f34e569cc70d33a7

Observation d0ad225d-b9e1-4768-80e0-10a61a269b37 · outbound

This paper cites Temporal credit assignment in reinforcement learning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Temporal credit assignment in reinforcement learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.530218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.176690Z digest=sha256:3bcbdd6b15fe548a7fb779bf2d6780c9f1590317a3d90a798f9c81c573e3f875

Observation 65d6c08a-74bd-40f4-9270-b8aac917d6a7 · outbound

This paper cites Learning guidance rewards with trajectory-space smooth- ing,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Learning guidance rewards with trajectory-space smooth- ing,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.520203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.179657Z digest=sha256:b266b57e7dd069dfbc5691448fc60b913b9ae8490820a4039185dce705901630

Observation 33f72736-f8f7-4112-a81b-668036b499ca · outbound

This paper cites Harutyunyan, W.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Harutyunyan, W

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.510627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.182674Z digest=sha256:99694d3908436dfb97a24c2b7299583723c8201ce99639d25fda0812b382d52f

Observation 3724d282-7afc-42bc-aa97-5da828ff567a · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.502040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.185401Z digest=sha256:9a46089b42f9e882efc4e3bed25f2a049a0a90c2b46a9d24e3412208aa94065b

Observation d0a81003-26d7-43ce-ba62-ae59e4f1d1e5 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.188530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.188530Z digest=sha256:88ad27fe3c386fbc1aee0d63412b78f67c0ec16fe79162fea4023632e4423ee1

Observation c99752e1-3a18-426a-8a2b-ebd7f2dcc812 · outbound

This paper cites DPO meets PPO: Reinforced token optimization for RLHF,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning DPO meets PPO: Reinforced token optimization for RLHF,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.493068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.192604Z digest=sha256:cc25230a7b2f66dd2913b1f814474245772c8dd630a8e8010547d7e885b8842a

Observation 7a6199c3-bf2b-4c68-ba3a-3e3b22308795 · outbound

This paper cites R3HF: Reward redistribution for enhancing reinforcement learning from human feedback,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning R3HF: Reward redistribution for enhancing reinforcement learning from human feedback,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.484041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.195874Z digest=sha256:a4d18bdd0c05d807e764b5652673e44aa79d8cfde9020c39dd3d31f96fec1a11

Observation 9469099a-82cb-40f7-a438-e3935c0b55f9 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Dense reward for free in reinforcement learning from human feedback,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.474408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.198621Z digest=sha256:011d0b146cedfa812ef0cf90cec88328717d7f8e685dac10e0036db8ec3ef50f

Observation 1be93a43-71a6-4aa1-8b09-4246ea680c60 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Simple statistical gradient-following algorithms for connectionist reinforcement learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.453992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.204535Z digest=sha256:44be0b416c089f71ed5131459133d7fcaba77540518f39a735aeae2b680e58a1

Observation 69347cfb-4e3b-468c-9ed9-27beddb7a661 · outbound

This paper cites Auto-Encoding Variational Bayes.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Auto-Encoding Variational Bayes

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.207681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.207681Z digest=sha256:683483c5e45a1e26954fb44d80f7c90ecfefc53e2688cb2447f0857504164ce3

Observation f6951719-dd51-470b-898e-5a1b0c4726a2 · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Emerging properties in self-supervised vision transformers,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.444511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.210873Z digest=sha256:2141d1c8ff21dceb5c0bf05623ddf63f2a380be71efddcccf352c56d1700d1f6

Observation dabf1354-2074-4903-b190-7c30f48f6d6d · outbound

This paper cites DiffSim: Taming Diffusion Models for Evaluating Visual Similarity.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning DiffSim: Taming Diffusion Models for Evaluating Visual Similarity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.214334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.214334Z digest=sha256:5d4a92a4fa08c3b88557be070503f97441aa05d0c9552566d65ed04df95e06a8

Observation 92df4940-22bb-4894-8c43-65ab8291f10c · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning LoRA: Low-rank adaptation of large language models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.217670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.217670Z digest=sha256:708e5b62f9570da2c7039cc168448faef724fd92feeb0b56a09547cde19fc076

Observation 1762ec38-ef64-4b33-9f98-a3eb686a4f25 · outbound

This paper cites Policy gradient methods for rein- forcement learning with function approximation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Policy gradient methods for rein- forcement learning with function approximation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.429414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.220474Z digest=sha256:2176886ea1e7c942a1fab3b99415a57bdfe518786d9c6d121c41013bf8bd5a9a

Observation f8186217-4635-4ead-9ac7-55d49981bf93 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.224156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.224156Z digest=sha256:d157e60c623959a920b065ff32f165a4eeb1b9976be12a516a925ae6365f71a5

Observation e00ca729-7998-4e5e-9493-2baae7aac08c · outbound

This paper cites Training deep nets with sublinear memory cost,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training deep nets with sublinear memory cost,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.419888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.229122Z digest=sha256:ff092ce52b91c3420e709ce0a13a3d2529f005099e4352add21b31624bf505c2

Observation f5acd2f4-921b-43af-b595-29fdb03e2953 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training Deep Nets with Sublinear Memory Cost

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.232105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.232105Z digest=sha256:1e40ce5ad58fd11c74828d0ad1d7d32a314bbfe09ddece56cb5749ee5de7fdc4

Observation eb825529-10da-4a0b-bf1e-6db2928fdaf5 · outbound

This paper cites Available: https://openreview.net/forum?id=eyxVRMrZ4m.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Available: https://openreview.net/forum?id=eyxVRMrZ4m

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.463485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:25:26.201653Z digest=sha256:45b26422aec3b2f3f0629c14ebfae4c143a7963f2fa5c665c5c21dbb6da45878

Pith citing papers

Observation 45e9e2b8-63b4-4484-a920-a61a0bdab5c2 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:33.490433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:33.490433Z digest=sha256:b22a14f2dd69643c478ac95770836407d71d8e9af03c97e9de1c77542170496f

Observation fb275552-9d13-432f-8f4c-003c320e670b · inbound

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion cites this paper.

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T21:09:36.040879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:09:36.040879Z digest=sha256:0f9ef3386f198521964ae7c4d8fbc340a12589f67139f4ff0d77660a0b4cd0a2

Observation 71e0a791-1444-4142-a587-66be70140e9f · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.786800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:f4a189d676d04815a6723ca2069cb63788fc6db2ed572a140c8c9b21bc58081c

Observation e2f44700-b3bf-4fee-b294-a22df8a5e414 · inbound

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation cites this paper.

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.984555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:00:16.005187Z digest=sha256:b40d149c452bb14d3ce92b05435c1fc3c8a73fcb3e2bb4e537cb54964e38230f