REVIEW 37 references
The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the growing adoption of AI image generation, in conjunction with the ever-increasing environmental resources demanded by AI, we are urged to answer a fundamental question: What is the environmental impact hidden behind each image we generate? In this research, we present a comprehensive empirical experiment designed to assess the energy consumption of AI image generation. Our experiment compares 17 state-of-the-art image generation models by considering multiple factors that could affect their energy consumption, such as model quantization, image resolution, and prompt length. Additionally, we consider established image quality metrics to study potential trade-offs between energy consumption and generated image quality. Results show that image generation models vary drastically in terms of the energy they consume, with up to a 46x difference. Image resolution affects energy consumption inconsistently, ranging from a 1.3x to 4.7x increase when doubling resolution. U-Net-based models tend to consume less than Transformer-based one. Model quantization instead results to deteriorate the energy efficiency of most models, while prompt length and content have no statistically significant impact. Improving image quality does not always come at the cost of a higher energy consumption, with some of the models producing the highest quality images also being among the most energy efficient ones.
Reference graph
Works this paper leans on
-
[1]
Verdecchia, J
R. Verdecchia, J. Sallou, L. Cruz, A systematic review of Green AI, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 13 (4) (2023) e1507
2023
-
[2]
S. N. Gowda, X. Hao, G. Li, S. N. Gowda, X. Jin, L. Sevilla-Lara, Watt for what: Rethinking deep learning’s energy-performance relationship, arXiv preprint arXiv:2310.06522 (2023)
arXiv 2023
-
[3]
Casta ˜no, S
J. Casta ˜no, S. Mart ´ınez-Fern´andez, X. Franch, J. Bogner, Exploring the carbon footprint of hugging face’s ml models: A repository mining study, in: 2023 ACM/IEEE International Symposium on Empirical Software Engi- neering and Measurement (ESEM), IEEE, 2023, pp. 1–12. 17
2023
-
[4]
S. A. Budennyy, V . D. Lazarev, N. N. Zakharenko, A. N. Korovin, O. Plosskaya, D. V . Dimitrov, V . Akhripkin, I. Pavlov, I. V . Oseledets, I. S. Barsola, et al., Eco2ai: carbon emissions tracking of machine learning mod- els as the first step towards sustainable ai, in: Doklady Mathematics, V ol. 106, Springer, 2022, pp. S118–S128
2022
-
[5]
Razzhigaev, A
A. Razzhigaev, A. Shakhmatov, A. Maltseva, V . Arkhipkin, I. Pavlov, I. Ryabov, A. Kuts, A. Panchenko, A. Kuznetsov, D. Dimitrov, Kandinsky: An improved text-to-image synthesis with image prior and latent diffusion, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2023, pp. 286–295
2023
-
[6]
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y . Wu, Z. Wang, J. Kwok, P. Luo, H. Lu, et al., Pixart-α: Fast training of diffusion transformer for photorealis- tic text-to-image synthesis, arXiv preprint arXiv:2310.00426 (2023)
arXiv 2023
-
[7]
Heusel, H
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, S. Hochreiter, Gans trained by a two time-scale update rule converge to a local nash equilibrium, in: Neural Information Processing Systems, 2017. URL https://api.semanticscholar.org/CorpusID: 326772
2017
-
[8]
M. S. M. Sajjadi, O. Bachem, M. Lu ˇci´c, O. Bousquet, S. Gelly, Assessing Generative Models via Precision and Recall, in: Advances in Neural Infor- mation Processing Systems (NeurIPS), 2018
2018
Show all 37 references
-
[9]
Salimans, I
T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, X. Chen, Improved techniques for training gans, Advances in neural information pro- cessing systems 29 (2016)
2016
-
[10]
Hessel, A
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, Y . Choi, Clipscore: A reference-free evaluation metric for image captioning, arXiv preprint arXiv:2104.08718 (2021)
2021 arXiv
-
[11]
Szegedy, V
C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna, Rethinking the in- ception architecture for computer vision, in: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[12]
J. Chen, C. Ge, E. Xie, Y . Wu, L. Yao, X. Ren, et al., Pixart- σ: Weak- to-strong training of diffusion transformer for 4k text-to-image generation, arXiv preprint (2024). 18
2024
-
[13]
P. Gao, L. Zhuo, Z. Lin, C. Liu, J. Chen, R. Du, E. Xie, X. Luo, L. Qiu, Y . Zhang, et al., Lumina-t2x: Transforming text into any modality, resolu- tion, and duration via flow-based large diffusion transformers, arXiv preprint arXiv:2405.05945 (2024)
2024 arXiv
-
[14]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, B. Ommer, High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10684–10695
2022
-
[15]
Sauer, D
A. Sauer, D. Lorenz, A. Blattmann, R. Rombach, Adversarial diffusion dis- tillation, in: European Conference on Computer Vision, Springer, 2025, pp. 87–103
2025
-
[16]
Esser, S
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. M ¨uller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, et al., Scaling rectified flow transformers for high-resolution image synthesis, in: Forty-first International Conference on Machine Learning, 2024
2024
-
[17]
Chadebec, O
C. Chadebec, O. Tasar, E. Benaroche, B. Aubin, Flash diffusion: Accelerat- ing any conditional diffusion model for few steps image generation (2024). arXiv:2406.02347
2024 arXiv
-
[18]
Sohl-Dickstein, E
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, S. Ganguli, Deep unsupervised learning using nonequilibrium thermodynamics, in: F. Bach, D. Blei (Eds.), Proceedings of the 32nd International Conference on Machine Learning, V ol. 37 of Proceedings of Machine Learning Research,...
2015
-
[19]
J. Ho, A. Jain, P. Abbeel, Denoising diffusion probabilistic models, Ad- vances in neural information processing systems 33 (2020) 6840–6851
2020
-
[20]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sas- try, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual mod- els from natural language supervision, in: International conference on ma- chine learning, PMLR, 2021, pp. 8748–8763. 19
2021
-
[21]
Peebles, S
W. Peebles, S. Xie, Scalable diffusion models with transformers, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4195–4205
2023
-
[22]
Vaswani, Attention is all you need, Advances in Neural Information Pro- cessing Systems (2017)
A. Vaswani, Attention is all you need, Advances in Neural Information Pro- cessing Systems (2017)
2017
-
[23]
Alexey, An image is worth 16x16 words: Transformers for image recog- nition at scale, arXiv preprint arXiv: 2010.11929 (2020)
D. Alexey, An image is worth 16x16 words: Transformers for image recog- nition at scale, arXiv preprint arXiv: 2010.11929 (2020)
2020 arXiv
-
[24]
Ramesh, P
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, M. Chen, Hierarchi- cal text-conditional image generation with clip latents, arXiv preprint arXiv:2204.06125 1 (2) (2022) 3
2022 arXiv
-
[25]
Midjourney, https://www.midjourney.com/ (2022)
2022
-
[26]
Podell, Z
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M ¨uller, J. Penna, R. Rombach, Sdxl: Improving latent diffusion models for high- resolution image synthesis, arXiv preprint arXiv:2307.01952 (2023)
2023 arXiv
-
[27]
S. Lin, A. Wang, X. Yang, Sdxl-lightning: Progressive adversarial diffusion distillation, arXiv preprint arXiv:2402.13929 (2024)
2024 arXiv
-
[28]
Y . Ren, X. Xia, Y . Lu, J. Zhang, J. Wu, P. Xie, X. Wang, X. Xiao, Hyper-sd: Trajectory segmented consistency model for efficient image synthesis, arXiv preprint arXiv:2404.13686 (2024)
2024 arXiv
-
[29]
Gupta, V
Y . Gupta, V . V . Jaddipal, H. Prabhala, S. Paul, P. V . Platen, Progressive knowledge distillation of stable diffusion xl using layer level loss (2024). arXiv:2401.02677
2024 arXiv
-
[30]
S. Luo, Y . Tan, L. Huang, J. Li, H. Zhao, Latent consistency models: Syn- thesizing high-resolution images with few-step inference, arXiv preprint arXiv:2310.04378 (2023)
2023 arXiv
-
[31]
B. F. Labs, Flux, https://github.com/black-forest-labs/ flux (2023)
2023
-
[32]
Krizhevsky, G
A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images (2009). 20
2009
-
[33]
Abdin, J
M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl, et al., Phi-3 technical report: A highly capable language model locally on your phone, arXiv preprint arXiv:2404.14219 (2024)
2024 arXiv
-
[34]
Wind, Top 40 useful prompts for stable diffusion xl, https://medium.com/phygital/top-40-useful-prompts-for-stable-diffusion- xl-008c03dd0557, accessed: 2025-03-07 (Dec
D. Wind, Top 40 useful prompts for stable diffusion xl, https://medium.com/phygital/top-40-useful-prompts-for-stable-diffusion- xl-008c03dd0557, accessed: 2025-03-07 (Dec. 2023)
2025
-
[35]
Courty, V
B. Courty, V . Schmidt, S. Luccioni, Goyal-Kamal, MarionCoutarel, B. Feld, J. Lecourt, LiamConnell, A. Saboni, Inimaz, supatomic, M. L ´eval, L. Blanche, A. Cruveiller, ouminasara, F. Zhao, A. Joshi, A. Bogroff, H. de Lavoreille, N. Laskaris, E. Abati, D. Blank, Z. Wang, A. Ca...
2024 doi
-
[36]
C. Chen, J. Mo, IQA-PyTorch: Pytorch toolbox for image quality as- sessment, [Online]. Available: https://github.com/chaofengc/ IQA-PyTorch (2022)
2022
-
[37]
streetcar
M. F. Naeem, S. J. Oh, Y . Uh, Y . Choi, J. Yoo, Reliable fidelity and diversity metrics for generative models (2020). 21 Appendix A. Preliminary Experiment The results of the correlation between the semantic content of prompts and the energy consumed during the image generati...
2020
Discussion (0). Sign in to comment.