Pith. sign in

Paper Citation Record · LEDGER

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

As of 19 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2605.15855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15855 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:48:17.049547Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:34:44.686328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:23:55.444823Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact17
  • verified fuzzy56
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bcb7de9-18ad-47d6-a362-0833bbd1c3e4 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Training Diffusion Models with Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.308257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f76a21cef6576280cc66c7f7eec7afb01aee7816006333afd714cc7c489d55e0

Observation ccfbb829-fc36-466e-acac-6930ea490396 · outbound

This paper cites A sur- vey on generative diffusion models.IEEE transactions on knowledge and data engineering, 36(7):2814–2830.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A sur- vey on generative diffusion models.IEEE transactions on knowledge and data engineering, 36(7):2814–2830

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.842876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7009dc5010aeeb70fee18cd02c2c17b1239a4e2123442fde6c2a88705d61e009

Observation 1cd3e7a2-3dc1-4e11-b13c-2c4d5a69fbd2 · outbound

This paper cites An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.302386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:42791de36dcb2e9783dd469b341d9d36e338eccc6163fd618bd25aaa87b6468e

Observation f4680a21-ac6f-434d-bdbf-2c55ae1100dd · outbound

This paper cites Dif- fusiondet: Diffusion model for object detection.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dif- fusiondet: Diffusion model for object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.841045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:76e5c36de50932b95c76d5929efa959e4a515fca4d82071be2700833e2582e7b

Observation bd2739b0-60cc-4a61-b68e-89f31175bda5 · outbound

This paper cites Diffusion models in vision: A survey.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion models in vision: A survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.839134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7eb007ed45cd5d16eb3363b8cde5dff1ba3bac2dc0e4853c6a95b98b4a355d03

Observation e39bdd7b-6892-43cb-a10e-85be05c618c2 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:48:57.332773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:10417951aa804d04df7b2a94157728a79a947ab99fa014bee45d7815d2c5efc8

Observation cc18749b-a09b-4589-be41-c605af67c372 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.837363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:08e52b52ed3b57a2523f6e31affe77a64c9ef36c2e6c1db93deea80be118b597

Observation 135ba8d6-484a-43f4-873e-e5abd472cc61 · outbound

This paper cites Re- inforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Sys- tems, 36.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Re- inforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Sys- tems, 36

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.835544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7e5ab7d35129a613e8f8ab0b6092f0a686f4e3bc48a177db74ad1601799a764f

Observation 2e0a1158-f96f-4f29-8c63-2cc891fbc0e3 · outbound

This paper cites Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.Journal of Artificial Intelligence Research, 79:417–446.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.Journal of Artificial Intelligence Research, 79:417–446

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.833747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:475c7074273d6eabd9dde963c1fe45ee71fe382c27f5775599c83709a40d8a02

Observation 2545ebc2-28a1-4326-b73e-c60f6f9a2b8f · outbound

This paper cites Re- flective policy optimization.International Conference on Machine Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Re- flective policy optimization.International Conference on Machine Learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.831612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4d4f78d49cb51f300bd06772a64f435c896c3a69e9d18c4d91cda1e4a056a107

Observation 3b7a36d4-23a0-49ab-8fc6-b763b125d5b4 · outbound

This paper cites Scaling laws for reward model overoptimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Scaling laws for reward model overoptimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.829903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c2bbb5bc668fdb88c9c7d15c01837457d9bc1b78acf9201e5025222e0141b6e8

Observation 2fdfeb35-f613-4952-8c49-d2c8a44c354d · outbound

This paper cites Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.317233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:04d3a6b7c482582917434095c9df25d7143372834d2d2be25b59dca5f5604670

Observation ce1b4e95-c112-4510-8d56-4328a4dbab63 · outbound

This paper cites Dealing with Sparse Rewards in Reinforcement Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dealing with Sparse Rewards in Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:dff86f913c7b5dba642699ef2f1beeb91a1dfb095beb67495a68402fcb94d175

Observation 2fcf6550-5140-41b0-a1d7-923102585f33 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in Neural Information Processing Systems, 30.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in Neural Information Processing Systems, 30

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.817030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:44fb535d7138186e27e814288c36623a64e9d90e54b1cce6b2485444b1bd6330

Observation 1a24358b-e2d5-47ca-a75c-33039563bae6 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.821871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:90fa289ba0fd2530e2588115f3b51a40fb23030f61ac1d16be3d9ea2061c39ee

Observation 03ce1bc0-f3c7-4728-9eac-b4c1c98b4d93 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.341223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c7849b3401c5d806318420301e365df125ca7df36b16146ee7632408e3f624e8

Observation 820bd6b8-1e48-4294-b46e-a63293967081 · outbound

This paper cites Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.828131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:5f2dd0b2acbac2f06f74577be8989b44297e45efdf2ad5d8aaeab0dd5cf099ef

Observation feb5cc73-2e30-4ccf-8c7a-595a64f29687 · outbound

This paper cites Reward hacking in reinforcement learning and rlhf: A multidisciplinary exami- nation of vulnerabilities, mitigation strategies, and alignment challenges.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Reward hacking in reinforcement learning and rlhf: A multidisciplinary exami- nation of vulnerabilities, mitigation strategies, and alignment challenges

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.798412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e9196bdbf59a29953e7866dd10db106fa9a716b3f42055681b311d8785ba509c

Observation a3325f06-d978-4d3b-94cc-1951244ba7e8 · outbound

This paper cites Dif- fusion reward: Learning rewards via conditional video dif- fusion.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dif- fusion reward: Learning rewards via conditional video dif- fusion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.800132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c6a55ea447286ec4a9450b97601860db5c2e3f9b6bc3d08771b59b0b3f1b9d78

Observation 24a48e48-3fb1-4844-9d57-78672cc942a7 · outbound

This paper cites Diffusion model-based image editing: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion model-based image editing: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.819867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:931f9fbc84cf30860f3b94610b050536246832c9ad152d0b9281cd0a174e71f1

Observation e0bfdc6d-64e5-40ec-bf8f-0591e3837307 · outbound

This paper cites Measuring Diversity in Co-creative Image Generation.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Measuring Diversity in Co-creative Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.338492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:04b8a2acd769651f502949db6a4561cc1adaf66b0d46ebda78c3606f7489e93a

Observation e9e903b1-bb29-417e-bbd6-b938137d6e03 · outbound

This paper cites Holodiffusion: Training a 3d diffusion model using 2d images.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Holodiffusion: Training a 3d diffusion model using 2d images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.796486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:803918daa610b60acf5967b56b3e415c72f94bc3966ea60910b5a2b07443de53

Observation 659cff2e-f249-4ff5-9109-cf4335076c5f · outbound

This paper cites Test-time Alignment of Diffusion Models without Reward Over-optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Test-time Alignment of Diffusion Models without Reward Over-optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.344161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:721063d26a060aba088231327cb0c80314c12ec7a36f3db30f3e480efa0f5c01

Observation 8d51bc0f-df24-4786-b634-890e7d98c517 · outbound

This paper cites Variational diffusion models.Advances in neural infor- mation processing systems, 34:21696–21707.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Variational diffusion models.Advances in neural infor- mation processing systems, 34:21696–21707

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.794774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e2c9dd7a46b788d5ee3b7b4bdeb193d6272020e5693eba4bd93307471925c542

Observation a8720526-3bcf-4794-87e9-993099d1ca9c · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in Neural Information Processing Systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in Neural Information Processing Systems

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.790557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:fcef2d81a8dc4336eacd6d7306e18623a490fb8de21fb7ed6fcba02830a99ece

Observation 4fd0689c-1a7a-4ba4-8c8a-ff92eae7f2a1 · outbound

This paper cites Improved precision and recall met- ric for assessing generative models.Advances in Neural In- formation Processing Systems, 32.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Improved precision and recall met- ric for assessing generative models.Advances in Neural In- formation Processing Systems, 32

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.824089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:a53d644e192d831db3e8b0eb67224e6b15e7873549e46d4bdb121bd71c5be37c

Observation 9335e1c7-c2a3-47d1-b080-5b63029c8805 · outbound

This paper cites Aligning diffusion mod- els by optimizing human utility.Advances in Neural Infor- mation Processing Systems, 37:24897–24925.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Aligning diffusion mod- els by optimizing human utility.Advances in Neural Infor- mation Processing Systems, 37:24897–24925

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.826112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7a21d96e3da7ffb282a8f1ffa689c2f1f774834861bcd2a9ef16cd64f71cd6c1

Observation c77785a1-63fa-4b32-874d-9484b904a438 · outbound

This paper cites Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.326419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:503fd1ae45882c793621c3ade386dcb232a5c324ed0d2c2be3a6d9176aae2dd6

Observation 447ea7be-bcc3-40ac-95bc-2aa795f89557 · outbound

This paper cites No-reference image quality assessment based on spatial and spectral entropies.Signal Processing: Image communica- tion, 29(8):856–863.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? No-reference image quality assessment based on spatial and spectral entropies.Signal Processing: Image communica- tion, 29(8):856–863

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.788614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:6866ca841eebef3544242764662d742931caf828b40bb3fd48498e108da8fd09

Observation df099ab6-8009-494e-89b6-045612f2282c · outbound

This paper cites Deepcache: Accelerating diffusion models for free.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deepcache: Accelerating diffusion models for free

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.777784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:965746da15b43a930cb3458167854791612593f74e76507b987398a174f8db87

Observation 0a993b5e-2a98-4313-9795-8538109036db · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.779881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:aec901d29da11e9bdc73072ca96e274c77883444b7348f4264aa307bdba9b6d9

Observation 737cf6e0-d5db-45d4-afdd-feaf4fb2bcec · outbound

This paper cites No-reference image quality assessment in the spa- tial domain.IEEE Transactions on Image Processing, 21 (12):4695–4708.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? No-reference image quality assessment in the spa- tial domain.IEEE Transactions on Image Processing, 21 (12):4695–4708

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.775807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:67a8eb966c51c2a1e468ae07f155a92aa9267c61c904335c0fd7a459149ac6a4

Observation 9805fa92-95c3-489a-b8a1-15e769819751 · outbound

This paper cites completely blind.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? completely blind

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.773502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e4ab6a2eca313564abd60a284f77e28e8451e1e1f014d07d82324cecd582ec22

Observation 27547f44-9da5-4373-9bda-b4c76be63de9 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.335429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:19d7d0b653e1f926a64aeb9b8127c8960d909e2763012fa9f4f9712cb15406af

Observation dbf1b86c-b20a-40c2-b057-910954ff97af · outbound

This paper cites Efficient Controllable Diffusion via Optimal Classifier Guidance.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Efficient Controllable Diffusion via Optimal Classifier Guidance

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.286646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:daa4fdb1967b64233af10bebe12bb6f2ebc4f7a9ebc8417fbcadedcf8a40a800

Observation e7a12db2-04fd-4761-8fe6-fed9875f37ce · outbound

This paper cites Markov decision processes.Handbooks in Operations Research and Management Science, 2:331– 434.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Markov decision processes.Handbooks in Operations Research and Management Science, 2:331– 434

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.782021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ef425f7333d686c85b4f8a664a3cf5e8b5d441015abc24cd14581693d60b85bb

Observation 407fdf8c-d242-46cf-b305-9554e7a866f5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Learn- ing transferable visual models from natural language super- vision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.784099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4b712c878878cf3fff5cbb2cdfc1db8b358e5e0115059fa968400d03fcb47027

Observation 2e3819f6-0f8c-4acb-bf67-aab2c2bd08f7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.320043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b35fb975e17164803d49361421c5acb29c1f829852531ba7cc370548c5f34363

Observation eb0541e8-4aa8-42d8-8158-e5e48ba79560 · outbound

This paper cites Learning by playing solving sparse reward tasks from scratch.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Learning by playing solving sparse reward tasks from scratch

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.764453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:0cd5c491d072b0daf1f5f1d9de29cd14c7a7a9b39dca4f9f33d35fe583899726

Observation 3f73eceb-cbfc-4da3-ace7-5ce04fe79f1c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.766741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:fa5a0941bf822ceb4179de6241894dd4c67f4536b046ece590c94cdb2d60d27b

Observation 6b4474c8-117b-4f4d-8e24-e7d418a1d37b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.771392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:adb469aafed3475287d27a95d4232806cda31de66f42a19d68545946378c562b

Observation a67cc0e1-303e-4690-822c-dd5590b4512b · outbound

This paper cites Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis.Journal of Computational and Applied Mathematics, 20:53–65.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis.Journal of Computational and Applied Mathematics, 20:53–65

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.744326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:fd7a2de43831f4ec61e65610fb6911f808e7650a78642c4235c8f85a75677411

Observation 1128b4b9-5073-4f4d-81ae-1a46f13ba6ed · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.757268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:5e3d395cd31b35d032ea2206760a39c287a678dff957851a5656c3ed514ea90c

Observation b15f4320-755c-4cb4-a9d8-9f5089e23cf9 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Improved techniques for training gans.Advances in neural information processing systems, 29

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.759596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b04252b9c3788f992a7cab0c6e277aa2b6d20bd8f7031e482aed2112ef83e361

Observation b5158451-e2c6-4c5d-8726-e0b2e8c470b9 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.754912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f2caf31fbd777e2b32b3916d0c6cd80510afbbb262c6293ed747903052450fc8

Observation ec9f9c63-5958-4602-85dc-abe14a9b4ca9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Proximal Policy Optimization Algorithms

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.329746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:236d0f7e8e510ee8917d3135b714563c8bcc52aac92d9b51e73c15562d1a34bb

Observation 77fb2339-f8e7-420c-8b4f-efd13fe239aa · outbound

This paper cites Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.762057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:cd26973cdc5ea061ce400f128733eae9b6e725d11a770b1dec73afbff284a279

Observation cd81aee5-8a30-43df-bf2e-624c424c7ed1 · outbound

This paper cites Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.768978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4870ea543f7916079051f93e04482fe215ff9b2eba7fa09d1c53fa1536c013b2

Observation 0b920364-e15d-46ee-8616-8ee9ec7651ae · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deep unsupervised learning using nonequilibrium thermodynamics

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.786570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:180c0739ef0bd7e9e88a09b5f7269ffd18c16090f81b7587f00c9ee12836d3de

Observation 0444539c-d75f-4c77-85a3-acfdf72f57b5 · outbound

This paper cites Inference-Time Alignment of Diffusion Models with Direct Noise Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Inference-Time Alignment of Diffusion Models with Direct Noise Optimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.291608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:72eb2d1c7df5adb051a2a2a29268f1f925d59b92ecef08e133fb69156b037192

Observation 63087ef6-f5ab-4105-ad07-bbb4736bc16d · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.305447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f8382917d586b70a954a42a0d5560921b776968ee8e85418de6b8b5cd3eb4902

Observation 2cac8b13-bfa6-4075-8bd0-de36f0c05fc2 · outbound

This paper cites Visualizing data using t-sne.Journal of machine learning research, 9 (11).

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Visualizing data using t-sne.Journal of machine learning research, 9 (11)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.752447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:0ed54efe17bfe33d2dde0e5449e39a4bf9c34bc70c99f8609963e85b97006e86

Observation abda76e1-ab2e-4fc0-bd6e-9d5f5107b27c · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion model alignment using direct preference optimization

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.749288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:383fd69321857d38fd2ba26290bc43809c1912389f7904d8e2b6b400cf1e7c3f

Observation 4c3336c2-a1ee-433d-9678-22bbdf32d63d · outbound

This paper cites Deep-reinforcement-learning-based autonomous uav navi- gation with sparse rewards.IEEE Internet of Things Journal, 7(7):6180–6190.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deep-reinforcement-learning-based autonomous uav navi- gation with sparse rewards.IEEE Internet of Things Journal, 7(7):6180–6190

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.746932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:654b1ccce196b397d7626d79aebc1aa1f3e454130e2e3fc1e79c0e9f8b7b8ade

Observation 257b259f-93f7-4fd9-a2f5-e179c0a14efc · outbound

This paper cites TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:59.936669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e53ab680ffcf435fa9066792b5c4b61dd162a8add434b952e20a727297421191

Observation 5eba3af8-bb76-46be-8b49-81d7b20884a7 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.298828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ff78cde2eb3e2f254a3c6db38c076f4b3ef5626429e4f35c456a52b36855b9b1

Observation 307ba49e-b2cf-4914-a7f6-4bb9562866c6 · outbound

This paper cites Diffir: Efficient diffusion model for image restoration.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffir: Efficient diffusion model for image restoration

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.880287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:255480d30c9b9d9746ad7e045b84625bc12c079e2f751fff6720392f06ad5afb

Observation 466da5dc-a5ed-4ebc-8ac9-0b7cbe1262b9 · outbound

This paper cites Dymo: Training-free diffusion model alignment with dynamic multi-objective scheduling.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dymo: Training-free diffusion model alignment with dynamic multi-objective scheduling

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.882159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ec288ecb479516d9a07cf6abfd542285b2f17b65e1496444ce792772238d3f71

Observation b49e4063-6029-4b37-943f-e0f5cfc219dd · outbound

This paper cites A survey on video dif- fusion models.ACM Computing Surveys, 57(2):1–42.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A survey on video dif- fusion models.ACM Computing Surveys, 57(2):1–42

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.874259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:680b2c09f0be198d455966a7203800f10e2b9c5afe7e2f627391e459455f1801

Observation 9c3585fe-e11f-4e7b-95a5-af1d66069cc0 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.876157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:251db98ee3287aac2a8a9e498fdb2eeaa00f6a1961bc3582f3c942c5805c0ea9

Observation 0064e5d0-b387-42e0-a4f7-cf8f0492bd0d · outbound

This paper cites Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.878023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e9966586c2043e62a75f4ce65322028ad89cc272573c8e5abcf9c6429d89217e

Observation d451510c-e91f-4077-90e8-442cce488153 · outbound

This paper cites Versatile diffusion: Text, images and variations all in one diffusion model.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Versatile diffusion: Text, images and variations all in one diffusion model

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.883994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:49d45d7c313b7644b861a77c4016cfad66baf7b854f8b3e8fc80eef9580f630d

Observation bb487ac0-3dd8-4dc8-9d17-1ab800ffaca7 · outbound

This paper cites The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.295846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c9d445021b6c00464f1d6833c4d70cd058bdd78be715b4a15a8592789c7ecf40

Observation cd1c5c24-7720-4129-95ea-f650676ffa48 · outbound

This paper cites Entropy-adaptive diffusion policy optimiza- tion with dynamic step alignment.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Entropy-adaptive diffusion policy optimiza- tion with dynamic step alignment

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.865650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:af8abb700d67be2f435f85114c54e7919258380945278f12ffff6c4e94134540

Observation 9c3382a6-86ba-495e-b24b-07294f3eaa40 · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Using human feedback to fine-tune diffusion models without any reward model

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.867425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:2741623db7e5c703fa7f46d2757a4d8bccc0a2408e34b8b7c91ed573950a35fc

Observation 1477834b-e9f6-49f8-8ff0-89fbe958803b · outbound

This paper cites A novel multi-step reinforcement learning method for solving reward hacking.Applied Intelligence, 49 (8):2874–2888.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A novel multi-step reinforcement learning method for solving reward hacking.Applied Intelligence, 49 (8):2874–2888

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.869645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:99a4d0496de6bbe9e07d3f7dcc22ab1ac155eb4640b943ed75887f1592081e10

Observation 3259f9a6-1279-4ed6-985d-a7fc22732510 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? The unreasonable effectiveness of deep features as a perceptual metric

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.862026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:877458eff65cc37b8f1c7fd1be7e17a12a82d062357cc3d42d06f6d818250d83

Observation 58892d7a-8ed7-445e-9942-f925e4f15021 · outbound

This paper cites Confronting reward overoptimiza- tion for diffusion models: A perspective of inductive and pri- macy biases.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Confronting reward overoptimiza- tion for diffusion models: A perspective of inductive and pri- macy biases

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.858446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:96b1ce60c216bcd023ff7f8b339682729326bbbe652075f33ac7034ad8984f9c

Observation c186831c-21fe-4e44-a8a6-00d7b8825565 · outbound

This paper cites Alphaholdem: High-performance artificial intelli- gence for heads-up no-limit poker via end-to-end reinforce- ment learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Alphaholdem: High-performance artificial intelli- gence for heads-up no-limit poker via end-to-end reinforce- ment learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.860290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:871c3a4d5ed9eb7f8347dc7afe1997d4a445e6c052e1de7ff00175e26d5cf2d2

Observation 89ac76cb-2e2d-47d0-aca5-6887a594870f · outbound

This paper cites 3d shape generation and completion through point-voxel diffusion.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? 3d shape generation and completion through point-voxel diffusion

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.863765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:9a051dbad16a32f70d2673921f44ca9ec937723769d3d017f6f04f0f36ae5033

Observation 1da7ee14-bb16-44f4-9687-cc58c01966fe · outbound

This paper cites Mixture of global and local experts with diffusion transformer for con- trollable face generation.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Mixture of global and local experts with diffusion transformer for con- trollable face generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.871500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:19b871910d0fa30f0124467ceb02bbb35957d692750a27cfea70235d81f7ee08

Observation b61469d5-4d07-4b3c-a32c-238dde804236 · outbound

This paper cites RL Fine-Tuning in Diffusion Models Existing diffusion models [4, 5, 24, 59, 62] primarily approximate the data distribution through denoising reconstruction loss.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? RL Fine-Tuning in Diffusion Models Existing diffusion models [4, 5, 24, 59, 62] primarily approximate the data distribution through denoising reconstruction loss

Reference 72

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T19:48:57.314460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:69362c0cbcaf4fd9ba1341ea6364247bea0ae4769e0b61fa25d38b141a92c89f

Observation e5304a71-38a6-474e-b7c1-7228977c2d35 · outbound

This paper cites •(1) Visualization Experiments.See Fig.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? •(1) Visualization Experiments.See Fig

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.854117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e9803146b4210eb9aa7ec7bf5c40807807645d1ff2238ee8e4f7d7e8cd95e993

Observation 565b7ff3-7f0e-43a3-a3a0-ff88e0c7121f · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.848394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:90700aef7c66e136b0d1bf0e6db927e2456644e6d8e6c3624563e07fd2b94866

Observation 499adcd8-b8ad-451c-8512-8c13e58f7071 · outbound

This paper cites Proof of Theorem 1.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Proof of Theorem 1

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.850498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:d5b8c3f96673fc455aa8849241e26193b6a6537db500afbe7b7240654e434581

Observation 068f1f5f-4784-4562-83e5-e23d0104f116 · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.846447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:fad7db9c22b2a54a94952188368c0446d23e2002ebb08c0d04ad7b66d8b02a42

Observation 356c7317-dbc0-47eb-ae78-f52bcb9075aa · outbound

This paper cites Therefore, componentwise, Cov x(i) t , x(j) t+τ = √¯αt ¯αt+τ Σij + r ¯αt+τ ¯αt (1−¯αt)δ ij.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Therefore, componentwise, Cov x(i) t , x(j) t+τ = √¯αt ¯αt+τ Σij + r ¯αt+τ ¯αt (1−¯αt)δ ij

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.852435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:591e3561f333306d63e2019dfffcb086decfec9aa26601d8bc8c5303a298ce7f

Observation a2d4b130-1e26-4dc0-9b74-8d9874997bac · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.844618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:6655d6120eb2964f84bbb86899f015dd66c88353c99766f36311936de2d7f0d6

Observation 499868ac-56b4-44b2-842b-e36e7024ba94 · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.855751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:632ed55b17dc5d4ce777d77c85e726eeb5dfeaef5441db45a58bbd38cd9b95a0

Pith citing papers

Observation 800e6608-2c88-48a3-80f0-49efb6a095a1 · inbound

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation cites this paper.

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:23:55.517186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T00:23:54.855176Z digest=sha256:68714cac683a94ef95de4c798cd0d86e5503252f98234cd963f3cdc908cf073e

Observation b303e609-b644-420d-8275-3eb5c8c5c215 · inbound

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models cites this paper.

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.884241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.884241Z digest=sha256:049099c902e0ed9e14420b149d677971b076802138270810e1ab8d769201c4f5

Observation deabb4af-042c-4d5f-be6f-48b1b3f49ed3 · inbound

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model cites this paper.

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:44.686328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:34:44.686328Z digest=sha256:7ea0eab05422e94e014b194edb1ac8fde4c4f07c7b1d4eee7891d67ef7faeb42

Observation ef9ea5d0-b9ea-4b21-b32f-8d6691654e0b · inbound

Diffusion Image Editing via Asynchronous Token Decoding cites this paper.

Diffusion Image Editing via Asynchronous Token Decoding Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:27:32.280009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:27:32.280009Z digest=sha256:9555ff130e8a6945a42634234b130530d09c0ef39c9d867efb42aa6d0a1c32b9