REVIEW 3 major objections 5 minor 1 cited by
Guidance in the Frequency Domain Enables High-Fidelity Sampling at Low CFG Scales
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Classifier-free guidance works better when low frequencies are guided gently and high frequencies strongly, a scheme the paper calls frequency-decoupled guidance (FDG), which improves FID and recall at low guidance scales.
desk verdict A useful plug-and-play tweak to CFG with a nice frequency-domain story, though the causal claim about which frequency hurts diversity is partly confounded by total guidance strength. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the frequency decomposition of the denoiser's clean-image prediction under a linear, invertible transform $\psi$, chosen as a single-level Laplacian pyramid: a coarse-to-fine decomposition that separates a signal into a low-frequency base and high-frequency detail. FDG computes $\psi$ on both the conditional and unconditional predictions, applies the CFG interpolation separately with scales $w_{\mathrm{low}}$ and $w_{\mathrm{high}}$ to the two bands, and inverts the pyramid to obtain the guided prediction. This carries the argument because it converts the single global guidance scale of CFG into two per-band scales, making the paper's claim that low and high frequencies should be treated differently directly testable. The implementation also includes a projection step that keeps the guidance difference aligned with the conditional prediction, but the core mechanism is the per-band scaling.
What would settle it
Swap the scales: run FDG with $w_{\mathrm{high}} < w_{\mathrm{low}}$. The paper predicts oversaturation, reduced diversity, and worse FID; if that configuration matches or beats uniform CFG, the claimed division of labor between frequency bands is wrong.
Extended reading notes
Core claim
The discovery is an asymmetric frequency account of CFG. Viewing the CFG update as an interpolation in the frequency domain, the paper derives that a uniform scale $w$ acts identically on the low- and high-frequency components of the guided prediction, even though the two bands play different roles. Empirically, raising the low-frequency scale is what collapses diversity and causes oversaturation, while raising the high-frequency scale sharpens detail and improves quality without hurting diversity. FDG therefore sets $w_{\mathrm{low}}$ close to or below the CFG scale and $w_{\mathrm{high}}$ at or above it, and the paper reports consistent FID and recall improvements over CFG across EDM2, DiT-XL/2, Stable Diffusion 2.1, XL, and 3.
Load-bearing premise
The load-bearing premise is that low- and high-frequency bands in the model's latent prediction correspond to overall structure versus fine detail in the final image; if a Laplacian pyramid in latent space does not track perceptual frequency, the explanation of why FDG helps is not established.
Editorial extensions
If this is right
- At low guidance scales, FDG can replace uniform CFG as a drop-in sampling rule, improving FID and recall without retraining or extra compute.
- Users can keep the diversity and natural color of low CFG scales while obtaining detail comparable to high CFG scales, by setting $w_{\mathrm{low}}$ low and $w_{\mathrm{high}}$ high.
- Time-gated guidance methods, such as guidance interval, can be understood and tuned through the frequency norms of the guidance signal, since their practical benefit is an implicit emphasis on high-frequency guidance.
- Distilled few-step models, where CFG often hurts, can still benefit from guidance through FDG, and text rendering in models like Stable Diffusion 3 improves because fine text detail lives in the high-frequency band.
Reading between the lines
- Editorial inference: if the frequency account is right, guidance schedules that increase $w_{\mathrm{high}}$ over sampling time, rather than using a fixed split, should push the quality-diversity frontier further; the paper's norm analysis suggests the effective balance shifts as denoising progresses.
- Editorial inference: the same per-band split could be applied to other guidance mechanisms and to non-image domains such as video or audio, where low-frequency structure and high-frequency texture have clear analogues.
- Editorial inference: because the split is performed in latent space, a useful stress test is whether the gains persist when the same split is applied in decoded pixel space; that would separate a genuine perceptual-frequency effect from a latent-representation artifact.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes classifier-free guidance (CFG) in the frequency domain and argues that low-frequency components of the guidance signal mainly govern global structure and condition alignment, while high-frequency components mainly contribute to visual fidelity. On this basis, the authors propose frequency-decoupled guidance (FDG), which applies separate CFG scales w_low and w_high to low- and high-frequency bands of the denoiser prediction. The central algebraic identity (Eq. 4), which follows from linearity of the frequency transform, is correct. The paper reports consistent FID and recall improvements over standard CFG across several models (EDM2, DiT-XL/2, Stable Diffusion 2.1, SDXL, SD3), plus improvements on human-preference and prompt-alignment metrics, and it claims FDG is a plug-and-play, no-retraining alternative to CFG.
Significance. If the empirical claims hold, FDG is a practically valuable contribution: it is simple, adds negligible sampling cost, works with pretrained models, and improves the quality-diversity trade-off relative to CFG. The paper also provides a plausible mechanistic story for why CFG hurts diversity and causes oversaturation, which is of independent interest. The main results are supported by experiments across multiple model families and datasets, and the central derivation is elementary and sound. However, the causal interpretation is currently under-supported by the ablations: the comparison in Fig. 5 does not control for total guidance strength, and the frequency decomposition is applied in latent space for the Stable Diffusion models without validating that latent-space bands correspond to perceptual image frequencies. These issues affect the paper's main explanatory claim, though the empirical gains of FDG may still be valid. The method is not circular: the derivation is a linearity identity, not a fitted quantity, and the per-model parameter choices are presented as design choices informed by sweeps.
major comments (3)
- [Section 5.2, Figure 5] The comparison between the arms (w_low=w, w_high=1) and (w_low=1, w_high=w) does not control for total guidance strength. From Eqs. (5)-(6), the injected guidance is (w-1)*psi_low[Delta] in the first arm and (w-1)*psi_high[Delta] in the second, where Delta = D_c - D_u. Figure 7 itself shows that ||psi_low[Delta]|| is much larger than ||psi_high[Delta]|| over most of the sampling trajectory. Lower recall, higher saturation, and worse FID in the low-frequency-only arm could therefore be explained by a larger L2 norm of the applied correction rather than by a frequency-specific effect. The same confound affects the prompt-alignment analysis in Figure 6. Please rerun the ablation with matched per-arm guidance norms (e.g., scaling w_high so that (w_high-1)||psi_high[Delta]|| = (w_low-1)||psi_low[Delta]|| at each step) and report recall, saturation, FID, and CLIP score as functions of the total injected norm; this is necessary to support the causal narrative that low-frequency guidance specifically harms diversity.
- [Section 5.1, Table 8] The FDG parameters w_low and w_high are reported per model, but the selection procedure is not described: there is no validation split, no sensitivity analysis, and no error bars over seeds or prompt sets. Because FDG introduces two free parameters while CFG has one, a fair comparison requires a specified protocol for choosing (w_low, w_high) and an estimate of variance. The reported FID differences for SDXL (25.23 vs 24.60) and SD2.1 (24.99 vs 23.33) are small and could fall within run-to-run noise. Please add a validation protocol, confidence intervals or standard errors, and at least a one-dimensional sensitivity sweep around the chosen settings.
- [Section 4, Algorithm 2] For the Stable Diffusion models, the Laplacian pyramid is applied directly to the latent-space predictions pred_cond and pred_uncond. The paper's mechanistic interpretation implicitly assumes that low- and high-frequency bands in the VAE latent space correspond to low- and high-frequency perceptual structure in the decoded image. This is not immediate because the latent space has a different spectral bias and the VAE decoder applies learned upsampling. Please validate the mapping, for example by decoding latent images whose frequency bands have been manipulated and measuring whether image-domain frequency content changes correspondingly, or by computing the same norm analysis on decoded predictions. Without this validation, the claim that 'low-frequency guidance governs global structure' is not established for latent diffusion models.
minor comments (5)
- [Algorithm 2] The pseudocode includes parallel_weights and a project() function that appear related to APG, but the main text does not describe when these are used in FDG. Please clarify whether they are part of FDG or optional add-ons, or remove them from the main algorithm.
- [Table 8] The heading 'Guidane parameters' contains a typo; it should read 'Guidance parameters'.
- [Table 2] The PickScore column reports values such as 0.45 vs 0.55, which the appendix describes as win probabilities, but the table caption does not explain this. Please state explicitly in the caption that these are win rates against the CFG baseline, not raw PickScore values.
- [Figure 5] The legend defines the three curves but does not explain what 'w' denotes for the CFG curve nor how the saturation metric is computed; please add these details to the caption.
- [Related work] Reference [62] applies frequency-aware guidance to diffusion models for image restoration; a brief sentence positioning FDG relative to this prior frequency-guided diffusion work would help the reader understand the novelty.
Circularity Check
No significant circularity: FDG's core update is a linearity identity, and the empirical claims rest on independent parameter sweeps rather than on fitted inputs or self-citation chains.
full rationale
The paper's central derivation is the frequency-domain rewriting of the CFG update: applying a linear, invertible transform psi to both sides of the standard CFG rule gives Eqs. (5)-(6), and setting w_low = w_high = w recovers exactly the original CFG update. This is an algebraic identity, not a prediction that is equivalent to its input. FDG's design — conservative low-frequency guidance and stronger high-frequency guidance — is motivated by empirical comparisons in Figures 5 and 6 and Table 8, where the scales are selected per model and then evaluated against CFG on held-out metrics (FID, precision, recall, ImageReward, HPSv2, PickScore, CLIP Score). No equation feeds its own output back as an input, and no fitted parameter is renamed as a prediction. The self-citations to the authors' prior work (CADS [47] and APG [49]) appear only in appendix combinations showing complementarity; they are not load-bearing for the FDG derivation and no uniqueness claim is imported from them. The most serious concern in the paper is experimental, not circular: the comparison in Figure 5 between w_low = w, w_high = 1 and w_low = 1, w_high = w does not equalize the total injected guidance, since ||psi_low[Delta]|| and ||psi_high[Delta]|| differ substantially (as the paper's own Figure 7 shows). That is a potential confound in the mechanistic attribution of diversity loss to low frequencies, but it is a validity threat, not a circularity: the claim is not true by definition of the method, and the main FID/recall gains of FDG are separately benchmarked against external baselines. Accordingly, the appropriate circularity score is low; the confound should be weighed under correctness risk rather than under circularity.
Assumptions & free parameters
free parameters (3)
- w_low (low-frequency guidance scale) =
1 for EDM2-S, EDM2-XL, DiT-XL/2; 3 for SD2.1 and SDXL; 1.5 for SD3 (Table 1); 3 or 1.5 in Table 2
- w_high (high-frequency guidance scale) =
2 to 3 for EDM2 and DiT; 7 for SD2.1 and SD3; 10 for SDXL; 12 for Table 2
- Pyramid level count =
1 (single-level) in main experiments
assumptions (5)
- standard math The frequency transform psi is linear and invertible (Laplacian pyramid or wavelet).
- domain assumption The denoiser output is a clean-image (x0) prediction in the model's working space, and frequency decomposition is applied directly to that space.
- domain assumption The default samplers and official checkpoints for each model are appropriate and are used unchanged.
- domain assumption FID, precision, recall, ImageReward, HPSv2, PickScore, and CLIP Score are valid proxies for image quality, diversity, and prompt alignment.
- domain assumption The behaviors observed in the frequency sweeps (low-frequency guidance reduces diversity, high-frequency guidance improves detail) generalize across datasets, models, and tasks.
Cite this review
Pith. "Pith review of Guidance in the Frequency Domain Enables High-Fidelity Sampling at Low CFG Scales." pith.science (2026). https://pith.science/paper/GHK4S6CN
@misc{pith2026250619713,
author = {Pith},
title = {Pith review of: Guidance in the Frequency Domain Enables High-Fidelity Sampling at Low CFG Scales},
year = {2026},
howpublished = {\url{https://pith.science/paper/GHK4S6CN}},
note = {Machine review of arXiv:2506.19713}
}
read the original abstract
Classifier-free guidance (CFG) has become an essential component of modern conditional diffusion models. Although highly effective in practice, the underlying mechanisms by which CFG enhances quality, detail, and prompt alignment are not fully understood. We present a novel perspective on CFG by analyzing its effects in the frequency domain, showing that low and high frequencies have distinct impacts on generation quality. Specifically, low-frequency guidance governs global structure and condition alignment, while high-frequency guidance mainly enhances visual fidelity. However, applying a uniform scale across all frequencies -- as is done in standard CFG -- leads to oversaturation and reduced diversity at high scales and degraded visual quality at low scales. Based on these insights, we propose frequency-decoupled guidance (FDG), an effective approach that decomposes CFG into low- and high-frequency components and applies separate guidance strengths to each component. FDG improves image quality at low guidance scales and avoids the drawbacks of high CFG scales by design. Through extensive experiments across multiple datasets and models, we demonstrate that FDG consistently enhances sample fidelity while preserving diversity, leading to improved FID and recall compared to CFG, establishing our method as a plug-and-play alternative to standard classifier-free guidance.
Figures
Figures from the paper (11 more)
Forward citations
Cited by 1 Pith paper
-
RelaxFlow: Text-Driven Amodal 3D Generation
A training-free dual-branch flow method uses multi-prior consensus and attention-logit low-pass relaxation to text-steer occluded 3D geometry while preserving the observed image.
Reference graph
Works this paper leans on
-
[1]
Cosmos world foundation model platform for physical ai
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025
arXiv 2025
-
[2]
Edify image: High-quality image generation with pixel space laplacian diffusion models
Yuval Atzmon, Maciej Bala, Yogesh Balaji, Tiffany Cai, Yin Cui, Jiaojiao Fan, Yunhao Ge, Siddharth Gururani, Jacob Huffman, Ronald Isaac, et al. Edify image: High-quality image generation with pixel space laplacian diffusion models. arXiv preprint arXiv:2411.07126, 2024
arXiv 2024
-
[3]
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. CoRR, abs/2211.01324,
-
[4]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets. CoRR, abs/2311.15127, 2023. doi: 10.48550/ARXIV .2311.15127. URL https://doi.org/10.48550/a...
-
[5]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023
2023
-
[6]
M. E. Brewster. An introduction to wavelets (charles k. chui). SIAM Rev., 35(2):312–313, 1993. doi: 10.1137/1035061. URL https://doi.org/10.1137/1035061
-
[7]
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https: //openreview.net/forum?id=B1xsqj09Fm
work page 2019
-
[8]
Peter J. Burt and Edward H. Adelson. The laplacian pyramid as a compact image code. volume 31, pages 532–540. 1983. doi: 10.1109/TCOM.1983.1095851. URL https://doi.org/10. 1109/TCOM.1983.1095851
arXiv 1983
Show all 71 references
-
[9]
Weiss, Mohammad Norouzi, and William Chan
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating gradients for waveform generation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net,
2021
-
[10]
Deep generative image models using a laplacian pyramid of adversarial networks
Emily L Denton, Soumith Chintala, Rob Fergus, et al. Deep generative image models using a laplacian pyramid of adversarial networks. Advances in neural information processing systems , 28, 2015
2015
-
[11]
Diffusion models beat gans on image syn- thesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image syn- thesis. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors,Advances in Neural Information Processing Systems 34: Annual Conferenc...
2021
-
[12]
Scaling rectified flow transform- ers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transform- ers for high-resolution image synthesis. arXiv preprint arXiv:2403.03206, 2024
2024 arXiv
-
[13]
SW AGAN: a style- based wavelet-driven generative model
Rinon Gal, Dana Cohen Hochberg, Amit Bermano, and Daniel Cohen-Or. SW AGAN: a style- based wavelet-driven generative model. ACM Trans. Graph., 40(4):134:1–134:11, 2021. doi: 10.1145/3450626.3459836. URL https://doi.org/10.1145/3450626.3459836
2021
-
[14]
Photorealistic video generation with diffusion models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662, 2023
2023 arXiv
-
[15]
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. CoRR, abs/2104.08718, 2021. URL https://arxiv.org/abs/2104.08718. 10
2021 arXiv
-
[16]
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vish- wanat...
2017
-
[17]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. CoRR, abs/2207.12598,
-
[18]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: An- nual Conference on Neural I...
2020
-
[19]
simple diffusion: End-to-end diffusion for high resolution images
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. CoRR, abs/2301.11093, 2023. doi: 10.48550/arXiv.2301.11093. URL https://doi.org/10.48550/arXiv.2301.11093
- [20]
-
[21]
Park, Tao Wang, Timo I
Qingqing Huang, Daniel S. Park, Tao Wang, Timo I. Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Havnø Frank, Jesse H. Engel, Quoc V . Le, William Chan, and Wei Han. Noise2music: Text-conditioned music generation with diffusion models. CoRR, ...
-
[22]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. 2022. URL https://openreview.net/forum?id= k7FuTOWMOc7
2022
-
[23]
Simpler diffusion (sid2): 1.5 FID on imagenet512 with pixel-space diffusion
Emiel Hoogeboom, Thomas Mensink, Jonathan Heek, Kay Lamerigts, Ruiqi Gao, and Tim Salimans. Simpler diffusion (sid2): 1.5 FID on imagenet512 with pixel-space diffusion. CoRR, abs/2410.19324, 2024. doi: 10.48550/ARXIV .2410.19324. URL https://doi.org/10.48550/arXiv. 2410.19324
-
[24]
Guiding a diffusion model with a bad version of itself
Tero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen, Timo Aila, and Samuli Laine. Guiding a diffusion model with a bad version of itself. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024. URL https://openreview.net/ forum?id=b...
2024
-
[25]
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. In Thirty-seventh Conference on Neural Information Processing Systems , 2023. URL https://openreview.net/ foru...
2023
-
[26]
Analyzing and improving the training dynamics of diffusion models, 2023
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models, 2023
2023
-
[27]
Improved precision and recall metric for assessing generative models
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett, editors, Advanc...
2019
-
[28]
Applying guidance in a limited interval improves sample and distribution quality in diffusion models
Tuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine, Timo Aila, and Jaakko Lehtinen. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. CoRR, abs/2404.07724, 2024. doi: 10.48550/ARXIV .2404.07724. URL https://doi.org...
-
[29]
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net, 2021. URL https://op...
2021
- [30]
-
[31]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In David J. Fleet, Tomás Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision - ECCV 2014...
2014 doi
-
[32]
Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers
Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers. arXiv preprint arXiv:2504.10483, 2025
2025
-
[33]
Pseudo numerical methods for diffusion models on manifolds
Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. Pseudo numerical methods for diffusion models on manifolds. 2022. URL https://openreview.net/forum?id=PlKWVd2yBkY
2022
-
[34]
Pseudo numerical methods for diffusion models on manifolds
Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. Pseudo numerical methods for diffusion models on manifolds. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum? id...
2022
-
[35]
Theodorou, Weili Nie, and Anima Anandkumar
Guan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A. Theodorou, Weili Nie, and Anima Anandkumar. I2sb: Image-to-image schrödinger bridge. CoRR, abs/2302.05872, 2023. doi: 10.48550/arXiv.2302.05872. URL https://doi.org/10.48550/arXiv.2302.05872
-
[36]
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. CoRR, abs/2211.01095, 2022. doi: 10.48550/arXiv.2211.01095. URL https://doi.org/10.48550/arXiv.2211.01095
-
[37]
A theory for multiresolution signal decomposition: the wavelet repre- sentation
Stephane G Mallat. A theory for multiresolution signal decomposition: the wavelet repre- sentation. IEEE transactions on pattern analysis and machine intelligence , 11(7):674–693, 1989
1989
-
[38]
Dpm- solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm- solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In NeurIPS, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/ 260a14acce2a89dad36adc8eefe7c59e-Abstr...
2022
-
[39]
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , volume 139 of Proceedings...
2021
-
[40]
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Cs...
2022
-
[41]
Kevin P. Murphy. Probabilistic Machine Learning: Advanced Topics. MIT Press, 2023. URL http://probml.github.io/book2
2023
-
[42]
SDXL: improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: improving latent diffusion models for high-resolution image synthesis. CoRR, abs/2307.01952, 2023. doi: 10.48550/ARXIV .2307.01952. URL https://doi.org/1...
-
[43]
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. CoRR, abs/2204.06125, 2022. doi: 10.48550/arXiv.2204.06125. URL https://doi.org/10.48550/arXiv.2204.06125. 12
- [44]
-
[45]
Ethics and creativity in computer vision
Negar Rostamzadeh, Emily Denton, and Linda Petrini. Ethics and creativity in computer vision. CoRR, abs/2112.03111, 2021. URL https://arxiv.org/abs/2112.03111
2021 arXiv
-
[46]
Bernstein, Alexander C
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei- Fei. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis., 115(3):211–252,
-
[47]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18- 24, 2022 , pages 10674–1...
2022
-
[48]
Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, and Romann M. Weber. LiteV AE: Lightweight and efficient variational autoencoders for latent diffusion models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024. URL https://...
2024
-
[49]
Seyedmorteza Sadat, Otmar Hilliges, and Romann M. Weber. Eliminating oversaturation and artifacts of high guidance scales in diffusion models. InThe Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=e2ONKX6qzJ
2025
-
[50]
Lee, Jonathan Ho, Tim Salimans, David J
Chitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee, Jonathan Ho, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Palette: Image-to-image diffusion models. In Munkht- setseg Nandigjav, Niloy J. Mitra, and Aaron Hertzmann, editors, SIGGRAPH ’22: Special Interest Group...
2022
-
[51]
Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, and Romann M. Weber. CADS: Unleashing the diversity of diffusion models through condition-annealed sampling. In The Twelfth International Conference on Learning Representations , 2024. URL https: //openreview.ne...
2024
-
[52]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=TIdIXIpzhoI
2022
-
[53]
Weiss, Niru Maheswaranathan, and Surya Ganguli
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. 37:2256–2265, 2015. URL http://proceedings.mlr.press/v37/sohl-dickstein15.html
2015
-
[54]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=St1giarCHLP
2021
-
[55]
Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with de...
2022
-
[56]
Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-...
2021
-
[57]
Computer Vision - Algorithms and Applications, Second Edition
Richard Szeliski. Computer Vision - Algorithms and Applications, Second Edition. Texts in Com- puter Science. Springer, 2022. ISBN 978-3-030-34371-2. doi: 10.1007/978-3-030-34372-9. URL https://doi.org/10.1007/978-3-030-34372-9. 13
2022 doi
-
[58]
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit Haim Bermano. Human motion diffusion model. 2023. URL https://openreview.net/pdf?id= SJ1kSyO2jwu
2023
-
[59]
Generative modeling by estimating gradients of the data dis- tribution
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data dis- tribution. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: A...
2019
-
[60]
Analysis of classifier-free guidance weight schedulers
Xi W ANG, Nicolas Dufour, Nefeli Andreou, Marie-Paule CANI, Victoria Fernandez Abrevaya, David Picard, and Vicky Kalogeiton. Analysis of classifier-free guidance weight schedulers. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview. net/for...
2024
-
[61]
Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis
Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341, 2023
2023 arXiv
-
[62]
Frequency-aware guidance for blind image restoration via diffusion models
Jun Xiao, Zihang Lyu, Hao Xie, Cong Zhang, Yakun Ju, Changjian Shui, and Kin-Man Lam. Frequency-aware guidance for blind image restoration via diffusion models. CoRR, abs/2411.12450, 2024. doi: 10.48550/ARXIV .2411.12450. URL https://doi.org/10.48550/arXiv. 2411.12450
-
[63]
Karen Liu
Jonathan Tseng, Rodrigo Castellon, and C. Karen Liu. EDGE: editable dance generation from music. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, V ancouver , BC, Canada, June 17-24, 2023, pages 448–458. IEEE, 2023. doi: 10.1109/ CVPR52729.2023.000...
2023
-
[64]
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-ric...
2022
-
[65]
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. In The Thirteenth International Conference on Learning Representations ,
-
[66]
Unipc: A unified predictor- corrector framework for fast sampling of diffusion models
Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. Unipc: A unified predictor- corrector framework for fast sampling of diffusion models. CoRR, abs/2302.04867, 2023. doi: 10.48550/arXiv.2302.04867. URL https://doi.org/10.48550/arXiv.2302.04867. 14 A Broader impact...
-
[67]
Imagereward: learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagereward: learning and evaluating human preferences for text-to-image generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems , p...
2023
-
[2015]
URL https://doi.org/10.1007/s11263-015-0816-y
doi: 10.1007/s11263-015-0816-y. URL https://doi.org/10.1007/s11263-015-0816-y
-
[2021]
URL https://openreview.net/forum?id=NsMLjcFaO8O
- [2022]
-
[2025]
URL https://openreview.net/forum?id=DJSZGGZYVi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.