REVIEW 4 major objections 6 minor 57 references
Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Repeated prompt language measurably shrinks visual diversity in AI image generation.
desk verdict Solid large-scale evidence of prompt homogenization on CivitAI, but the visual-diversity link is confounded by shared generation settings and needs a controlled analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is token overlap: the fraction of shared vocabulary between two prompts is treated as the operating cause of visual similarity. To measure it, the paper truncates prompts to 20 tokens, applies MinHash-based Jaccard similarity to form high- and moderate-overlap clusters, and compares the mean pairwise cosine similarity of the corresponding image embeddings. Visual diversity itself is measured with the Vendi score, the effective number of distinct images in a set, which combines how many clusters of images exist and how evenly they are populated. The user taxonomy of consistent repeaters, occasional repeaters, and non-repeaters does the explanatory work of identifying who produces the repeated language.
What would settle it
Regenerate the images in the high- and moderate-overlap prompt clusters while holding model version, style add-on modules, random seed, sampler, and guidance scale fixed across all clusters; if the correlation between token overlap and image-embedding similarity falls to near zero, the claim that linguistic repetition drives visual uniformity fails.
Extended reading notes
Core claim
On its own terms, the paper establishes three linked facts. First, over seven months the platform's prompt vocabulary contracts: type-token ratio and effective vocabulary fall, compression rises, self-repetition rises, and high-similarity near-duplicates grow from about 68-73% to over 83% of submissions. Second, this contraction is driven largely by consistent repeaters, users who resubmit identical prompts, who are about 47% of users but contribute about 72% of prompts and over three-quarters of near-duplicate clusters. Third, textual overlap predicts visual sameness: prompt clusters sharing at least 16 of 20 tokens yield more similar image embeddings, and Vendi scores of generated images decline as lexical diversity falls after February 2024. The paper reads these facts as evidence that linguistic repetition reinforces less diverse representations, adding a user-driven layer of bias on top of model training bias.
Load-bearing premise
The central claim assumes that overlapping words in prompts are what make generated images look similar, even though images can differ in model version, style add-on modules, random seed, sampler, and guidance scale, all of which may be shared within prompt clusters.
Editorial extensions
If this is right
- If token overlap drives image sameness, then prompting tools that suggest fresh words or rare descriptors should measurably increase the visual spread of generated images.
- Because exact duplicates account for 40-50% of submissions and consistent repeaters dominate near-duplicate clusters, interventions aimed at heavy repeaters could change corpus-wide visual diversity more than general interventions.
- Text-only lexical metrics, especially the self-repetition score and compression ratio, can serve as cheap early-warning signals for expected drops in visual diversity, since they correlate with Vendi scores without needing to render images.
- The post-February 2024 drop tracks adoption of a new model's structured rating tags, so the homogenization effect is tied to community adoption of fixed prompt formulas and is likely to recur when major new models appear.
Reading between the lines
- A decisive confound remains untested: high-overlap prompt clusters may share model versions, style add-on modules, seeds, samplers, or guidance scales, so a controlled regeneration experiment holding those fixed is the direct way to confirm the language effect.
- The same repetition-to-homogenization loop likely occurs in other generative systems, including LLM writing tools, where template reuse and 'best practice' phrases could narrow output diversity independently of model bias.
- A practical extension: randomly substitute rare or novel descriptors into the dominant prompt templates and monitor Vendi scores; if rare-token variants measurably increase visual diversity, community tag recommendations could be designed as an antidote to homogenization.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper analyzes over six million prompts from the Civiverse dataset on CivitAI across seven months to study how prompt language evolves and whether it relates to visual diversity. The authors categorize users into consistent repeaters, occasional repeaters, and non-repeaters; compute lexical diversity metrics (TTR, ENW, SRS, CR); track semantic topics via MiniLM embeddings with HDBSCAN/UMAP; and measure visual diversity with Vendi scores. Their main reported findings are that prompt language becomes increasingly homogenized over time, that repeated prompts make up 40–50% of submissions, that semantic topic diversity remains relatively stable, and that higher token overlap in prompts correlates with higher image embedding similarity (r = 0.33, R² = 0.18), which they interpret as linguistic repetition reinforcing visual homogenization.
Significance. If the central claim held, this would be a valuable contribution to the sociotechnical study of text-to-image systems, showing that user behavior—not only model training data—can drive output homogenization. The paper's strengths include the large-scale real-world dataset, the use of multiple complementary lexical diversity metrics, the fixed-size sampling robustness checks for lexical trends, and the transparent description of the pipeline. The public Civiverse dataset also enables further verification. However, the headline claim rests on an uncontrolled observational correlation: image similarity is measured across images generated with different checkpoints, LoRA modules, samplers, CFG scales, and seeds, and the paper does not condition on any of these factors. The temporal post-February 2024 decline also coincides with the adoption of the Pony Diffusion model, so the language–visual diversity link is not isolated from model change. These issues make the central causal-style conclusion currently under-supported, though the descriptive lexical trends appear robust.
major comments (4)
- [§7.2 and §9] The central claim that lexical repetition in prompts drives visual homogeneity is not identified from generation configuration. Civiverse images were generated with different Stable Diffusion checkpoints, LoRA modules, samplers, CFG scales, and seeds, yet the cluster-level analysis in §7.2 correlates prompt token overlap with image embedding similarity without conditioning on any of these settings. This confound is concrete, not hypothetical: the dominant prompt clusters in Table 6 contain LoRA references such as lora_gothic_outfit06, and Appendix C.2 shows the emergence of 4-grams like 'style sdxllorapony diffusion v6'. Users who share a prompt template typically also share the checkpoint or LoRA that makes that template work, so the observed r = 0.33 could reflect shared model weights rather than linguistic repetition per se. The Discussion (§9) acknowledges 'model checkpoints, seed values, and hyperparameters' but does not test them. To support the paper's wording, the analysis should either stratify by generation settings or show that the token–image similarity relationship persists within fixed checkpoint/LoRA/sampler configurations.
- [§7.2 and Appendix E] The statistical evidence for the core correlation is weaker than the text suggests. The reported r = 0.33 with R² = 0.18 means only about 18% of variance in image similarity is explained by token similarity, and no confidence intervals or bootstrap estimates are reported. Appendix E shows the cluster-size-weighted R² drops to 0.145 for token similarity, which is a meaningful sensitivity result that is not discussed in the main text. Additionally, the sentence in Appendix E stating that the weighted regression using identical tokens had R² = 0.962 appears to be a typo, presumably 0.0962 or similar; this should be corrected. The paper should report uncertainty intervals and discuss the modest explained variance when framing the result as a 'clear correlation'.
- [§5.2 and §7.1] The temporal alignment between the post-February 2024 decline in lexical diversity and the decline in Vendi scores is presented as suggestive evidence of a language–visual diversity link, but this period coincides with the widespread adoption of Pony Diffusion XL, as the paper itself notes in §5.2. The Vendi score time series therefore cannot separate the effect of changing prompt language from the effect of changing underlying models. A minimal control would be to compute Vendi scores separately for images generated with the same checkpoint before and after the shift, or to restrict the time-series comparison to a stable model cohort. Without such a control, the claimed correspondence between lexical and visual diversity over time remains confounded.
- [§4 and §5.3] The user categories are defined from duplicate-submission behavior, and the same categories are then used to characterize lexical diversity of unique prompts in §5.3, which risks a form of circularity. For example, consistent repeaters are users who resubmit identical prompts, so it is not surprising that their unique prompts share more formulaic structure. This does not invalidate the descriptive finding, but the paper should clarify that the category-based lexical comparisons are descriptive properties of behaviorally defined groups rather than independent evidence about linguistic experimentation.
minor comments (6)
- [§4] The sentence beginning 'Additionally, categorized users into three distinct groups' is missing a subject and should read 'Additionally, we categorized users into three distinct groups.'
- [§7.1] The description of Vendi score sampling says '10,000 prompts and their corresponding images from each user category, consistent, occasional, and non-repeaters, per month, resulting in a total of 30,000 monthly samples'; the phrase 'and their corresponding images' should be clarified because images and prompts are not one-to-one in the duplicate-inclusive dataset, and it is unclear whether the 10,000 are sampled from unique prompts or from all prompts.
- [§6.1] The topic modeling paragraph states that GPT-4o proposes labels based on c-TF-IDF keywords, but no details are given about the prompt template, number of labels, or validation of the labels; adding this information would improve reproducibility.
- [Appendix C.2, Table 25] In Table 25, the 4-gram 'style sdxllorapony diffusion v6' appears with percentages that seem inconsistent (e.g., 36.85% in April 2024 while adjacent rows show 38.32% and 32.12%); these values should be checked for internal consistency.
- [General] The paper does not state whether analysis code or computed metric values are available; providing a reproducibility link or specifying data/software availability would strengthen the contribution.
- [Figure 8] The captions for Figures 8a and 8b refer to 'token-based' and 'semantic-based' image similarity, but the text also discusses CLIP text embedding similarity; the captions should be expanded to indicate which embedding spaces are compared in each panel.
Circularity Check
No circularity: the prompt-image similarity correlations are measured associations, not fitted or self-defined; the only self-citation is to a dataset used as data, not as a load-bearing conclusion.
full rationale
I find no circular step in the claimed derivation chain. The lexical metrics (TTR, SRS, CR, ENW) are computed on prompt text, while Vendi scores and pairwise image similarities are computed on image embeddings; neither quantity is fitted to the other, and the correlations (e.g., SRS vs. Vendi r=-0.536; token-vs-image similarity r=0.33) are empirical associations that could have been null or negative. The paper does not rename a fitted parameter as a prediction, and it does not import a uniqueness theorem or ansatz from the authors' prior work. The self-citation to the Civiverse dataset [38] is a data-source citation, and the present study re-derives its aggregate statistics from that data rather than importing a conclusion from the cited paper. The Discussion does acknowledge that model checkpoints, seeds, and hyperparameters vary and are not controlled; that is a genuine internal-validity caveat, but it is a confounding-factor concern, not circularity by construction. The Appendix E cross-model comparison (SDXL, SD3, FLUX) provides an external check that high token overlap produces visually similar outputs even with the model held fixed, which further supports the non-circular reading. In short, the strongest substantive claims are measured correlational findings whose inputs and outputs are operationally distinct, so the paper is self-contained with respect to circularity.
Assumptions & free parameters
free parameters (4)
- Jaccard similarity thresholds for near-duplicate clusters =
0.8 for high similarity; 0.5-0.7 for moderate similarity; 20-token truncation
- User category repetition threshold =
20% reuse or less for occasional repeaters
- Fixed-size sampling choices =
213,779 monthly sample; 90,521 per category; 10,000 images per category per month
- Minimum cluster size for semantic embedding clusters =
20
assumptions (6)
- standard math MinHash and locality-sensitive hashing yield valid approximate Jaccard similarities between token sets.
- standard math The Vendi score is a valid diversity measure for image embeddings.
- domain assumption Anonymized usernames in Civiverse uniquely identify users across months.
- domain assumption Exact character-for-character matching after cleaning and lemmatization captures user repetition behavior.
- domain assumption Removing 3% non-English prompts does not bias the lexical diversity trends.
- domain assumption The post-February 2024 lexical decline is attributable to the Pony Diffusion release and its tag system.
Cite this review
Pith. "Pith review of Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity." pith.science (2026). https://pith.science/paper/VERSQVCP
@misc{pith2026250414125,
author = {Pith},
title = {Pith review of: Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity},
year = {2026},
howpublished = {\url{https://pith.science/paper/VERSQVCP}},
note = {Machine review of arXiv:2504.14125}
}
read the original abstract
Following the initial excitement, Text-to-Image (TTI) models are now being examined more critically. While much of the discourse has focused on biases and stereotypes embedded in large-scale training datasets, the sociotechnical dynamics of user interactions with these models remain underexplored. This study examines the linguistic and semantic choices users make when crafting prompts and how these choices influence the diversity of generated outputs. Analyzing over six million prompts from the Civiverse dataset on the CivitAI platform across seven months, we categorize users into three groups based on their levels of linguistic experimentation: consistent repeaters, occasional repeaters, and non-repeaters. Our findings reveal that as user participation grows over time, prompt language becomes increasingly homogenized through the adoption of popular community tags and descriptors, with repeated prompts comprising 40-50% of submissions. At the same time, semantic similarity and topic preferences remain relatively stable, emphasizing common subjects and surface aesthetics. Using Vendi scores to quantify visual diversity, we demonstrate a clear correlation between lexical similarity in prompts and the visual similarity of generated images, showing that linguistic repetition reinforces less diverse representations. These findings highlight the significant role of user-driven factors in shaping AI-generated imagery, beyond inherent model biases, and underscore the need for tools and practices that encourage greater linguistic and thematic experimentation within TTI systems to foster more inclusive and diverse AI-generated content.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. 2024. Ai suggestions homogenize writing toward western styles and diminish cultural nuances. arXiv preprint arXiv:2409.11360 (2024)
arXiv 2024
-
[2]
Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. 2024. Ho- mogenization effects of large language models on human creative ideation. In Proceedings of the 16th Conference on Creativity & Cognition . 413–425
work page 2024
-
[3]
Muhammad Sidik Asyaky and Rila Mandala. 2021. Improving the perfor- mance of HDBSCAN on short text clustering by using word embedding and UMAP. In 2021 8th international conference on advanced informatics: Concepts, theory and applications (ICAICTA) . IEEE, 1–6
work page 2021
-
[4]
Sofian Audry. 2021. Art in the age of machine learning . Mit Press
work page 2021
-
[5]
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . 1493–1504
work page 2023
-
[6]
O’Reilly Media, Inc
Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit . " O’Reilly Media, Inc. "
2009
-
[7]
Andrei Z Broder. 2000. Identifying and filtering near-duplicate documents. In Annual symposium on combinatorial pattern matching . Springer, 1–10
work page 2000
-
[8]
Andrei Z Broder, Moses Charikar, Alan M Frieze, and Michael Mitzen- macher. 1998. Min-wise independent permutations. In Proceedings of the thirtieth annual ACM symposium on Theory of computing . 327–336
work page 1998
Show all 57 references
-
[9]
Ricardo JGB Campello, Davoud Moulavi, and Jörg Sander. 2013. Density- based clustering based on hierarchical density estimates. In Pacific-Asia conference on knowledge discovery and data mining . Springer, 160–172
2013
-
[10]
Eva Cetinic and James She. 2022. Understanding and creating art with AI: Review and outlook. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 18, 2 (2022), 1–22
2022
-
[11]
Li-Yuan Chiou, Peng-Kai Hung, Rung-Huei Liang, and Chun-Teng Wang
-
[12]
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3043–3054
2023
-
[13]
CivitAI. 2024. Civitai: The home of open-source generative AI. https: //civitai.com/
2024
-
[14]
Michael A Covington and Joe D McFall. 2010. Cutting the Gordian knot: The moving-average type–token ratio (MATTR). Journal of quantitative linguistics 17, 2 (2010), 94–100
2010
-
[15]
Kevin T Cunningham and Katarina L Haley. 2020. Measuring lexical diversity for discourse analysis in aphasia: Moving-average type–token ratio and word information measure. Journal of Speech, Language, and Hearing Research 63, 3 (2020), 710–721
2020
-
[16]
Werner Ebeling and Gregoire Nicolis. 1991. Entropy of symbolic sequences: the role of correlations. Europhysics Letters 14, 3 (1991), 191
1991
-
[17]
Otmar Ertl. 2020. Probminhash–a class of locality-sensitive hash algo- rithms for the (probability) jaccard similarity. IEEE Transactions on Knowl- edge and Data Engineering 34, 7 (2020), 3491–3506
2020
-
[18]
Sedigheh Eslami, Gerard de Melo, and Christoph Meinel. 2021. Does clip benefit visual question answering in the medical domain as much as it does in the general domain? arXiv preprint arXiv:2112.13906 (2021)
2021 arXiv
-
[19]
Dejan Grba. 2023. Renegade X: Poetic Contingencies in Computational Art. In Proceedings of xCoAx 2023, 11th Conference on Computation, Com- munication, Aesthetics & X. Edited by Mario Verdicchio, Miguel Carvalhais, Luisa Ribas and Andre Rangel. Porto: i2ADS Research Institute ...
2023
-
[20]
Dejan Grba. 2024. Art Notions in the Age of (Mis) anthropic AI. In Arts, Vol. 13. MDPI, 137
2024
-
[21]
Carla W Hess, Kelley P Ritchie, and Richard G Landry. 1984. The type- token ratio and vocabulary performance. Psychological Reports 55, 1 (1984), 12 51–57
1984
-
[22]
Thierry Hoquet. 2023. From the Modern Synthesis to the Other (Extended, Super, Postmodern. . . ) Syntheses. InUnderstanding Evolution in Darwin’s" Origin" The Emerging Context of Evolutionary Thinking . Springer, 397–413
2023
-
[23]
Jianqiu Ji, Jianmin Li, Shuicheng Yan, Qi Tian, and Bo Zhang. 2013. Min- max hash for jaccard similarity. In 2013 IEEE 13th International Conference on Data Mining. IEEE, 301–309
2013
-
[24]
S Kannan, S Karuppusamy, A Nedunchezhian, P Venkateshan, P Wang, N Bojja, and A Kejariwal. 2016. Big data analytics for social media. Big data: principles and paradigms, Cambridge-MA, Morgan Kaufmann-Elsevier (2016)
2016
-
[25]
Nithish Kannen, Arif Ahmad, Marco Andreetto, Vinodkumar Prabhakaran, Utsav Prabhu, Adji Bousso Dieng, Pushpak Bhattacharyya, and Shachi Dave. 2024. Beyond aesthetics: Cultural competence in text-to-image models. arXiv preprint arXiv:2407.06863 (2024)
2024 arXiv
-
[26]
Hyung-Kwon Ko, Gwanmo Park, Hyeon Jeon, Jaemin Jo, Juho Kim, and Jinwook Seo. 2023. Large-scale text-to-image generation models for visual artists’ creative works. In Proceedings of the 28th international conference on intelligent user interfaces . 919–933
2023
-
[27]
Takio Kurita. 2021. Principal component analysis (PCA). In Computer vision: a reference guide . Springer, 1013–1016
2021
-
[28]
Vivian Liu, Han Qiao, and Lydia Chilton. 2022. Opal: Multimodal image generation for news illustration. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–17
2022
-
[29]
Philip M McCarthy and Scott Jarvis. 2010. MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment. Behavior research methods 42, 2 (2010), 381–392
2010
-
[30]
Jon McCormack, Maria Teresa Llano, Stephen James Krol, and Nina Rajcic
-
[31]
Leland McInnes, John Healy, and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)
2018 arXiv
-
[32]
Kibum Moon, Adam Green, and Kostadin Kushlev. 2024. Homogenizing Effect of Large Language Model (LLM) on Creative Diversity: An Empirical Comparison
2024
-
[33]
Francisco Javier Moreno Arboleda, Felipe Cortés Noreña, and Benjamín Cruz Álvarez. 2022. On the Use of Minhash and Locality Sensitive Hashing for Detecting Similar Lyrics. Engineering Letters 30, 1 (2022)
2022
-
[34]
Grégoire Nicolis and Pierre Gaspard. 1994. Toward a probabilistic approach to complex systems. Chaos, Solitons & Fractals 4, 1 (1994), 41–57
1994
-
[35]
Jonas Oppenlaender. 2023. A taxonomy of prompt modifiers for text-to- image generation. Behaviour & Information Technology (2023), 1–14
2023
-
[36]
Ville Paananen, Jonas Oppenlaender, and Aku Visuri. 2023. Using text-to- image generation for architectural design ideation. International Journal of Architectural Computing (2023), 14780771231222783
2023
-
[37]
Vishakh Padmakumar and He He. 2023. Does Writing with Language Models Reduce Content Diversity? arXiv preprint arXiv:2309.05196 (2023)
2023 arXiv
-
[38]
Maria-Teresa De Rosa Palmini, Laura Wagner, and Eva Cetinic. 2024. Ci- viverse: A Dataset for Analyzing User Engagement with Open-Source Text-to-Image Models. arXiv preprint arXiv:2408.15261 (2024)
2024 arXiv
-
[39]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[40]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 1, 2 (2022), 3
2022 arXiv
-
[41]
N Reimers. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv preprint arXiv:1908.10084 (2019)
2019 arXiv
-
[42]
Brian Richards. 1987. Type/token ratios: What do they really tell us? Journal of child language 14, 2 (1987), 201–209
1987
-
[43]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695
2022
-
[44]
Nikita Salkar, Thomas Trikalinos, Byron C Wallace, and Ani Nenkova
-
[45]
Téo Sanchez. 2023. Examining the Text-to-Image Community of Practice: Why and How do People Prompt Generative AIs?. In Proceedings of the 15th Conference on Creativity and Cognition . 43–61
2023
-
[46]
Chantal Shaib, Joe Barrow, Jiuding Sun, Alexa F Siu, Byron C Wallace, and Ani Nenkova. 2024. Standardizing the measurement of text diversity: A tool and a comparative analysis of scores. arXiv preprint arXiv:2403.00553 (2024)
2024
-
[47]
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao. 2023. Glaze: Protecting artists from style mimicry by {Text-to-Image} models. In 32nd USENIX Security Symposium (USENIX Security 23). 2187–2204
2023
-
[48]
Paulus Setiawan Suryadjaja and Rila Mandala. 2021. Improving the per- formance of the extractive text summarization by a novel topic modeling and sentence embedding technique using SBERT. In 2021 8th International Conference on Advanced Informatics: Concepts, Theory and Applic...
2021
-
[49]
Mahan Tafreshipour, Aaron Imani, Eric Huang, Eduardo Almeida, Thomas Zimmermann, and Iftekhar Ahmed. 2024. Prompting in the Wild: An Empirical Study of Prompt Evolution in Software Repositories. arXiv preprint arXiv:2412.17298 (2024)
2024 arXiv
-
[50]
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems 33 (2020), 5776–5788
2020
-
[51]
Zijie J Wang, Evan Montoya, David Munechika, Haoyang Yang, Ben- jamin Hoover, and Duen Horng Chau. 2022. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. arXiv preprint arXiv:2210.14896 (2022)
2022 arXiv
-
[52]
Lukas RA Wilde. 2023. Generative imagery as media form and research field: Introduction to a new paradigm. (2023)
2023
-
[53]
Yutong Xie, Zhaoying Pan, Jinge Ma, Luo Jie, and Qiaozhu Mei. 2023. A prompt log analysis of text-to-image generation systems. In Proceedings of the ACM Web Conference 2023. 3892–3902
2023
-
[54]
hair, "
Lili Zhang, Xi Liao, Zaijia Yang, Baihang Gao, Chunjie Wang, Qiuling Yang, and Deshun Li. 2024. Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–25. 1...
2024
-
[2022]
In Proceedings of the conference
Self-repetition in abstractive neural summarizers. In Proceedings of the conference. Association for Computational Linguistics. Meeting, Vol. 2022. NIH Public Access, 341
2022
-
[2023]
In Proceedings of the 2023 ACM designing interactive systems conference
Designing with AI: an exploration of co-ideation with image genera- tors. In Proceedings of the 2023 ACM designing interactive systems conference. 1941–1954
2023
-
[2024]
In International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar)
No Longer Trending on Artstation: Prompt Analysis of Generative AI Art. In International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) . Springer, 279–295
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.