REVIEW 3 major objections 5 minor 48 references
DiC: Rethinking Conv3x3 Designs in Diffusion Models
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read DiC, a diffusion model made entirely of stride-1 3x3 convolutions, claims to surpass transformer-based diffusion models on ImageNet generation while running faster.
desk verdict DiC is a real empirical architecture study with a likely-true core claim, but its headline margins lean on baseline numbers from the authors' own prior work and the abstract overstates the long-training comparison. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Conv3x3 Basic Block: two sequential stride-1 3x3 convolutions with GroupNorm, GELU activation, a residual shortcut, and conditioning injected into the second convolution. Around it, the hourglass encoder-decoder uses downsampling and upsampling so each 3x3 kernel sees a region equivalent to 6x6 or 12x12 pixels in the original image, compensating for the kernel's narrow receptive field. Sparse (strided) skip connections, applied every few blocks instead of every block, cut the channel-concatenation cost that dense skips impose at scale. Stage-specific embedding tables give each encoder/decoder stage its own condition vector aligned to that stage's feature dimension, and mid-block injection plus conditional gating (borrowed from DiT's AdaLN) let the condition modulate features more precisely. The paper counts Winograd acceleration as halving the effective FLOPs of the 3x3 stride-1 convolutions, which underpins the reported throughput advantage.
What would settle it
Independently re-run DiT-XL/2 and the other baselines under the same 400K-iteration ImageNet 256x256 setting with the same VAE and optimizer, and compute FID/IS on the same reference statistics; if DiT-XL/2's FID drops from 20.05 toward DiC's 13.11, or if DiC's numbers move when re-measured, the claimed margin collapses. A second check is to measure DiC and DiT throughput on the same GPU and kernel library, since DiC's reported Winograd FLOPs assume hardware and software support for that algorithm.
Extended reading notes
Core claim
The central claim is that a purely convolutional diffusion model can surpass diffusion transformers in both generation quality and throughput when its architecture is matched to the 3x3 kernel's receptive-field limits. DiC replaces the transformer block with two stride-1 3x3 convolutions, builds the network as an encoder-decoder hourglass, keeps only sparse skip connections between stages, and conditions each stage with its own embedding injected mid-block under an AdaLN-style gating scheme. On ImageNet 256x256 at 400K iterations with no guidance, DiC-XL reduces FID from 20.05 (DiT-XL/2) to 13.11 and raises IS from 66.74 to 100.15; with classifier-free guidance DiC-XL reaches FID 3.89 versus 6.24. Scaled further, DiC-H reaches FID 2.25 with guidance after 2M iterations, slightly better than DiT-XL/2's 2.27 after 7M iterations, at roughly 2.4 times the throughput. The paper also reports the same pattern at 512x512 resolution.
Load-bearing premise
The comparison against diffusion transformers depends on the baseline FID/IS numbers in the main table being obtained under exactly the same aligned 400K-iteration DiT training setting, but those numbers are inherited from a prior study co-authored by DiC's own first author rather than re-run here.
Editorial extensions
If this is right
- On the standard DiT 400K schedule, DiC-S, DiC-B, and DiC-XL all beat their DiT counterparts by large margins (for example FID 13.11 versus 20.05 at XL size), so the advantage holds across model scales.
- Because DiC's cost grows roughly linearly with image resolution while DiT's self-attention grows quadratically, the quality and speed gap widens at higher resolutions such as 512x512.
- DiC-H converges fast: with no guidance it reaches FID 9.73 at 600K steps, close to DiT-XL/2's 9.62 after 7M steps, suggesting CNN backbones need far less compute to reach a given quality.
- With classifier-free guidance and longer training, DiC-H reaches FID 2.25 on ImageNet 256x256, matching or slightly beating DiT-XL/2's 2.27 while training for 2M instead of 7M iterations and sampling at 2.4 times the throughput.
- Combining DiC with representation alignment (REPA/U-REPA) yields FID 1.74 after 1M iterations, exceeding the plain-training results and showing the architecture benefits from the same convergence accelerations as transformers.
Reading between the lines
- The conditioning recipe (stage-specific embeddings, mid-block injection) addresses a general mismatch in U-shaped networks where a single condition embedding table is shared across stages with different channel widths; this design lesson likely transfers to other encoder-decoder generative models, including U-shaped vision transformers and hybrid diffusion backbones.
- If the reported margins hold under independent replication, the current preference for isotropic transformer backbones in diffusion models may be a convenience rather than a necessity; simplified convolutional architectures could define a better efficiency frontier for real-time and high-resolution synthesis, where attention's quadratic cost is most punishing.
- The Winograd-based FLOP accounting suggests that reported GFLOPs for convolution-based models can be misleadingly high; comparing actual latency on target hardware, including which kernel libraries are used, matters more than theoretical FLOPs for deciding between CNN and transformer backbones.
- A testable prediction follows from the receptive-field argument: on very high-resolution images, where a few downsampling stages no longer give the deepest 3x3 kernels a view of the whole image, DiC's advantage should shrink unless the hourglass depth is increased, which would indicate that the receptive-field expansion is indeed the mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes DiC, a class-conditional latent diffusion model built entirely from stride-1 3x3 convolutions. The architecture is an encoder-decoder U-Net with sparse skip connections, stage-specific condition embeddings, mid-block condition injection, and conditional gating. The authors report that DiC outperforms DiT and other transformer baselines on ImageNet 256x256 and 512x512 in FID/IS at 400K iterations, while maintaining higher throughput, and they present ablations supporting each design choice.
Significance. If the results are reproducible, DiC would be a significant counterpoint to the field's shift toward isotropic transformers, demonstrating that a well-designed purely convolutional architecture can match or beat transformer backbones at comparable FLOPs while being faster in practice. The paper's strengths include a clean set of ablations (Tables 1 and 2) showing monotonic improvements, measured throughput numbers, and a clear progression from standard architectures to the final design. However, the external validity of the headline numbers is currently limited by the reliance on inherited baseline statistics and the absence of error bars.
major comments (3)
- [Sec. 4.2, Tables 4 and 6] The central claim that DiC surpasses existing diffusion transformers by considerable margins rests entirely on baseline FID/IS values taken from ref. [38], a paper co-authored by the first author. In particular, the DiT-XL/2 value of 20.05 used in Tables 4 and 6 differs from the 19.5 reported in the original DiT paper [30], so the comparison is not against the official numbers. The paper should either reproduce all baselines in the same codebase with the same 400K setting, or provide the exact checkpoint and sampling details for the inherited numbers. Without this, the margins such as 13.11 vs 20.05 are not yet established.
- [Sec. 4.1, Tables 3, 6, and 7] The FLOPs comparisons are inconsistent: DiC FLOPs are reported both raw and Winograd-adjusted, while DiT FLOPs are given only in raw form. Since Winograd applies only to 3x3 convolutions, the entries '116.1 (57.2)' next to '118.6' overstate the efficiency advantage. The measured throughput (TP) is a fairer metric and should be emphasized; the FLOPs tables should either use unadjusted FLOPs for all models or adjust all applicable operations. Additionally, the text in Sec. 4.3 saying 'DiT-XL/2 requires 524.7G FLOPs (with Winograd optimization)' appears to be a typo, as Winograd does not apply to DiT.
- [Sec. 4.2, Tables 4-7] All FID/IS values are single-run estimates without error bars or multiple seeds. Diffusion training is known to have run-to-run variance, and the close margins in Table 9 (DiC-H 2.25 vs DiT-XL/2 2.27) and the ordering of the ablations in Tables 1-2 could change with repeated runs. The authors should report at least two or three seeds for the key comparisons, or state the variance if fewer runs are available.
minor comments (5)
- [Sec. 4.3, Table 7] The table header and row alignment are confusing: for DiT-XL/2, the '16.2' appears to be throughput, not Winograd FLOPs, but it is placed in the Wino. column. Please reformat the table and correct the accompanying text.
- [Figure 5 caption] The caption reads 'cf g= 4'; this should be 'cfg=4'.
- [Sec. 4.2] The claim that PixArt-α-XL/2 is compared under the same class-conditional setting should be clarified, since PixArt-α is originally a text-to-image model; state explicitly how it was adapted per [38].
- [Sec. 3.3] The phrase 'prevents any leakage of label between stages' is imprecise; synchronized label-drop during training ensures consistent conditioning dropout, not information leakage. Rephrase for accuracy.
- [Supplementary, 'Credit'] The acknowledgment that baseline statistics in Table 6 are from [38] is good, but it should also appear prominently in the main text near the table.
Circularity Check
Headline transformer comparison is inherited from the authors' U-DiTs paper, making the central external claim self-referential; DiC's own measured results are otherwise independent.
-
self citation load bearing
[Section 4.2, 'Comparison with Diffusion Transformer Baselines' (Table 6); Supplementary 'Credit' paragraph]
"As different models use different settings, including training hyperparameters, the choice of samplers, training iterations et cetera, we adopt a universally-aligned setting (400K iterations on the DiT codebase) according to [38]. ... Credit. Baseline performance statistics in Tab. 6 are from [38], a work that measures the capability of Diffusion Transformers under the aligned standard setting of DiT."
The paper's headline claim that DiC surpasses existing diffusion transformers by considerable margins is carried by Table 6, whose external baseline FID/IS values are not measured in this manuscript but taken from [38], a U-DiTs paper co-authored by the present first author. The aligned 400K comparison protocol is itself adopted 'according to [38]' and the baseline statistics are explicitly credited to [38]. The central external comparison therefore reduces to the authors' own prior measurements: if those baseline runs are stale, misaligned, or not directly comparable, the reported margins are not independently established.
full rationale
This paper is an empirical architecture study, not a mathematical derivation, and most of its evidence is self-contained: DiC's FID/IS values are measured outputs of training runs, and the incremental design improvements (hourglass U-Net, sparse skips, stage-specific embeddings, mid-block injection, conditional gating, GELU) are validated by the paper's own ablations in Tables 1 and 2. No fitted parameter is renamed as a prediction, and no equation is defined in terms of its own output. However, the headline comparative claim against diffusion transformers is not independently benchmarked in this manuscript: the 'universally-aligned setting' is adopted according to [38], and all Table 6 baseline statistics are credited to [38], a prior paper co-authored by the present first author. Thus the central external comparison inherits both protocol and numbers from the authors' own earlier work, making the evidence chain for 'surpasses existing diffusion transformers by considerable margins' self-referential. The score of 4 reflects this partial, load-bearing self-citation while acknowledging that DiC's own measurements and internal ablations remain independent content; if the U-DiTs baseline runs are independently verified and reproducible, the circularity concern would drop to near zero.
Assumptions & free parameters
free parameters (5)
- Encoder-decoder depth configuration (DiC-XL) =
[7,7,8,7,7]
- Channel width (DiC-XL) =
384
- Sparse skip stride =
not disclosed (every few blocks)
- Condition injection position =
mid-block (second conv layer)
- Stage-specific embedding dimension =
matched to stage feature dims (+14.06M params)
assumptions (4)
- domain assumption FID and IS on ImageNet 256x256/512x512 are the relevant measures of generative quality.
- domain assumption Baseline statistics from U-DiTs [38] were produced under the identical aligned 400K-iteration DiT setting and are accurate.
- domain assumption Winograd acceleration reduces the effective FLOPs of every stride-1 3x3 convolution in DiC by 5/9 and does not incur overhead that changes the efficiency comparison.
- domain assumption Single-seed training runs produce FID differences large enough to be meaningful for the stated claims.
Cite this review
Pith. "Pith review of DiC: Rethinking Conv3x3 Designs in Diffusion Models." pith.science (2026). https://pith.science/paper/ADJZFUP3
@misc{pith2026250100603,
author = {Pith},
title = {Pith review of: DiC: Rethinking Conv3x3 Designs in Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/ADJZFUP3}},
note = {Machine review of arXiv:2501.00603}
}
read the original abstract
Diffusion models have shown exceptional performance in visual generation tasks. Recently, these models have shifted from traditional U-Shaped CNN-Attention hybrid structures to fully transformer-based isotropic architectures. While these transformers exhibit strong scalability and performance, their reliance on complicated self-attention operation results in slow inference speeds. Contrary to these works, we rethink one of the simplest yet fastest module in deep learning, 3x3 Convolution, to construct a scaled-up purely convolutional diffusion model. We first discover that an Encoder-Decoder Hourglass design outperforms scalable isotropic architectures for Conv3x3, but still under-performing our expectation. Further improving the architecture, we introduce sparse skip connections to reduce redundancy and improve scalability. Based on the architecture, we introduce conditioning improvements including stage-specific embeddings, mid-block condition injection, and conditional gating. These improvements lead to our proposed Diffusion CNN (DiC), which serves as a swift yet competitive diffusion architecture baseline. Experiments on various scales and settings show that DiC surpasses existing diffusion transformers by considerable margins in terms of performance while keeping a good speed advantage. Project page: https://github.com/YuchuanTian/DiC
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[38]
U-dits: Downsample tokens in u-shaped diffusion transformers, 2024
Yuchuan Tian, Zhijun Tu, Hanting Chen, Jie Hu, Chao Xu, and Yunhe Wang. U-dits: Downsample tokens in u-shaped diffusion transformers, 2024. 2, 6, 1
work page 2024
-
[30]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InIEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 4172–4182. IEEE, 2023. 1, 2, 3, 5, 6, 7
work page 2023
-
[1]
All are worth words: A vit backbone for diffusion models
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. All are worth words: A vit backbone for diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 22669–22679. IEEE, 2023. 1, 2, 4, 7
work page 2023
-
[2]
Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthe- sis, 2023
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthe- sis, 2023. 1, 2, 3, 7
work page 2023
-
[3]
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- Σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.CoRR, abs/2403.04692, 2024. 1, 2
arXiv 2024
-
[4]
Pixart- δ: Fast and controllable image generation with latent consistency models,
Junsong Chen, Yue Wu, Simian Luo, Enze Xie, Sayak Paul, Ping Luo, Hang Zhao, and Zhenguo Li. Pixart- δ: Fast and controllable image generation with latent consistency models,
-
[5]
Visionllama: A unified llama interface for vision tasks.CoRR, abs/2403.00522, 2024
Xiangxiang Chu, Jianlin Su, Bo Zhang, and Chunhua Shen. Visionllama: A unified llama interface for vision tasks.CoRR, abs/2403.00522, 2024. 3, 5, 6, 7, 1
arXiv 2024
-
[6]
Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z. Kaplan, and En- rico Shippole. Scalable high-resolution pixel-space image synthesis with hourglass diffusion transformers.CoRR, abs/2401.11605, 2024. 2, 1
arXiv 2024
Show all 48 references
-
[7]
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. CoRR, abs/1703.06211, 2017. 3
2017 arXiv
-
[8]
FlashAttention-2: Faster attention with better paral- lelism and work partitioning
Tri Dao. FlashAttention-2: Faster attention with better paral- lelism and work partitioning. InInternational Conference on Learning Representations (ICLR), 2024. 6
2024
-
[9]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InAdvances in Neural Information Processing Systems (NeurIPS), 2022. 6
2022
-
[10]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pages 248–255. IEEE...
2009
-
[11]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. InAdvances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 8780–8794, 20...
2021
-
[12]
Repvgg: Making vgg-style convnets great again
Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. Repvgg: Making vgg-style convnets great again. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 13733–13742, 2021. 2, 3
2021
-
[13]
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns.CoRR, abs/2203.06717, 2022
Xiaohan Ding, Xiangyu Zhang, Yizhuang Zhou, Jungong Han, Guiguang Ding, and Jian Sun. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns.CoRR, abs/2203.06717, 2022. 3
2022 arXiv
-
[14]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim En- tezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synth...
2024
-
[15]
Finder, Roy Amoyal, Eran Treister, and Oren Freifeld
Shahaf E. Finder, Roy Amoyal, Eran Treister, and Oren Freifeld. Wavelet convolutions for large receptive fields. CoRR, abs/2407.05848, 2024. 3
2024 arXiv
-
[16]
Mamba: Linear-time sequence model- ing with selective state spaces.CoRR, abs/2312.00752, 2023
Albert Gu and Tri Dao. Mamba: Linear-time sequence model- ing with selective state spaces.CoRR, abs/2312.00752, 2023. 2
2023 arXiv
-
[17]
Diffit: Diffusion vision transformers for image generation.CoRR, abs/2312.02139, 2023
Ali Hatamizadeh, Jiaming Song, Guilin Liu, Jan Kautz, and Arash Vahdat. Diffit: Diffusion vision transformers for image generation.CoRR, abs/2312.02139, 2023. 3, 7, 1
2023 arXiv
-
[18]
Deep residual learning for image recognition.CoRR, abs/1512.03385, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition.CoRR, abs/1512.03385, 2015. 2, 3
2015 arXiv
-
[19]
Denoising diffu- sion probabilistic models.CoRR, abs/2006.11239, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models.CoRR, abs/2006.11239, 2020. 1, 2, 4
2006 arXiv
-
[20]
simple diffusion: End-to-end diffusion for high resolution images
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, pages 13213– 13232. PMLR, 2023. 2, 1
2023
-
[21]
Scalable adaptive computation for iterative generation, 2023
Allan Jabri, David Fleet, and Ting Chen. Scalable adaptive computation for iterative generation, 2023. 2, 1
2023
-
[22]
Analyzing and improving the training dynamics of diffusion models
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models. InProc. CVPR, 2024. 2, 1
2024
-
[23]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Im- agenet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting he...
2012
-
[24]
Open-sora-plan, 2024
PKU-Yuan Lab and Tuzhan AI etc. Open-sora-plan, 2024. 1
2024
-
[25]
Flux, 2024
Black Forest Labs. Flux, 2024. 1
2024
-
[26]
Fast algorithms for convolutional neural net- works.CoRR, abs/1509.09308, 2015
Andrew Lavin. Fast algorithms for convolutional neural net- works.CoRR, abs/1509.09308, 2015. 2, 3, 5
2015 arXiv
-
[27]
Torchprofile, 2024
Zhijian Liu. Torchprofile, 2024. 6
2024
-
[28]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 11966–11976. IEEE, 2022. 3, 5
2022
-
[29]
Albergo, Nicholas M
Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, and Saining Xie. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers.CoRR, abs/2401.08740, 2024. 6, 1
2024 arXiv
-
[31]
Prajit Ramachandran, Barret Zoph, and Quoc V . Le. Search- ing for activation functions.CoRR, abs/1710.05941, 2017. 5
2017 arXiv
-
[32]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. InIEEE/CVF Con- ference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674– 1...
2022
-
[33]
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. InMedical Image Computing and Computer-Assisted Inter- vention - MICCAI 2015 - 18th International Conference Mu- nich, Germany, October 5 - 9, 2015, Proceedings...
2015
-
[34]
Very deep convo- lutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition. In3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. 2
2015
-
[35]
Todo: To- ken downsampling for efficient generation of high-resolution images, 2024
Ethan Smith, Nayan Saxena, and Aninda Saha. Todo: To- ken downsampling for efficient generation of high-resolution images, 2024. 1, 2
2024
-
[36]
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. InInternational Conference on Learning Representations,
-
[37]
Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024
Yao Teng, Yue Wu, Han Shi, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024. 2
2024
-
[39]
U-repa: Aligning diffusion u-nets to vits, 2025
Yuchuan Tian, Hanting Chen, Mengyu Zheng, Yuchen Liang, Chao Xu, and Yunhe Wang. U-repa: Aligning diffusion u-nets to vits, 2025. 8, 1
2025
-
[40]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkor- eit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need.CoRR, abs/1706.03762,
-
[41]
Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. DCN-M: improved deep & cross network for feature cross learning in web-scale learning to rank systems.CoRR, abs/2008.13535,
2008 arXiv
-
[42]
Internimage: Exploring large-scale vision foundation models with deformable convo- lutions, 2023
Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, Xiaogang Wang, and Yu Qiao. Internimage: Exploring large-scale vision foundation models with deformable convo- lutions, 2023. 3
2023
-
[43]
Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, and Song Han. Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024. 2
2024
-
[44]
Enhancing vi- sion transformer: Amplifying non-linearity in feedforward network module
Yixing Xu, Chao Li, Dong Li, Xiao Sheng, Fan Jiang, Lu Tian, Ashish Sirasao, and Emad Barsoum. Enhancing vi- sion transformer: Amplifying non-linearity in feedforward network module. InForty-first International Conference on Machine Learning, 2024. 5
2024
-
[45]
Jing Nathan Yan, Jiatao Gu, and Alexander M. Rush. Dif- fusion models without attention. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8239–8249, 2024. 2
2024
-
[46]
Representa- tion alignment for generation: Training diffusion transformers is easier than you think.CoRR, abs/2410.06940, 2024
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representa- tion alignment for generation: Training diffusion transformers is easier than you think.CoRR, abs/2410.06940, 2024. 8, 1
-
[47]
Open-sora: Democratizing efficient video production for all, 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, 2024. 1 DiC: Rethinking Conv3x3 Designs in Diffusion Models Supplementary Material
2024
-
[48]
Additional Experiments Further Details about Baselines.In Tab. 6, most baselines including PixArt-α, DiffiT [ 17], and DiT-LLaMA [ 5] are direct improvements over DiTs [ 30]; U-ViTs [1] are pub- lished earlier to DiTs, and we also find some coincidence between the hyperparamet...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.