REVIEW 4 major objections 6 minor 29 references
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Text-to-image models systematically reproduce societal stereotypes across race, gender, age, and body type, an audit of 16,000 generated images finds.
desk verdict A broad, useful prompt-expansion audit with a genuinely new model, but the headline percentages lack the statistical support the text implies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The audit machinery is a prompt battery of 160 manually curated topics with multiple wording variations, run through two model families (Stable Diffusion 1.5 and Flux-1) with fixed sampling settings, followed by human visual labeling into coarse demographic bins (White/Black/East Asian/Latino/Middle Eastern/Other; Male/Female/Other; Young/Adult/Senior; Underweight/Average/Overweight). The comparison set of 8,000 search-engine images provides a baseline against which model skews are measured. The load-bearing step is the labeling: each generated image is assigned a demographic category by visual inspection, and the label ratios across prompt groups become the reported bias statistics.
What would settle it
Take a random sample of, say, 500 of the generated images from the paper's prompt set, have a diverse panel of labelers independently assign the same demographic categories, and compute agreement (e.g., Fleiss' kappa). If agreement is low (for instance below 0.6) for race, age, or somatotype, then the specific percentages are not stable. A second test: regenerate the same 160 topics with different random seeds and a different model version and check whether the sign and size of the gaps (e.g., high-income White 70% vs low-income 43%) reproduces; if the gaps vanish under reseeding, the findings may be sampling artifacts.
Extended reading notes
Core claim
The paper's central discovery is that text-to-image models exhibit stereotype-reinforcing disparities across a wide set of human-centric dimensions. For example, high-income roles are depicted as White (70% vs 43% for low-income in SD1.5), negative attributes and actions are overwhelmingly associated with males (91% and 87% respectively), positive place descriptions are 99% Western while negative place descriptions shift toward Africa and the Middle East, and neutral 'place of worship' prompts produce Christian imagery 83% of the time. Flux-1 shows even stronger skews, generating almost exclusively White individuals across most prompt groups. These patterns hold in both a UNet-based and a DiT-based model, suggesting the bias does not depend on one architecture.
Load-bearing premise
The statistics depend on human labelers categorizing each generated image into fixed demographic buckets by looking at it; the paper reports no measure of how often labelers agree, so if different labelers put the same image in different buckets, the percentages in the tables would shift.
Editorial extensions
If this is right
- If biases persist across architectures, future models trained on AI-generated web content could inherit and amplify these skews as training data.
- Users relying on text-to-image models for professional or creative work would unknowingly reproduce stereotyped imagery, reinforcing representational harms.
- The search-engine comparison suggests the models are more extreme than web image distributions in some categories, so the problem is not merely a reflection of general internet content.
- Simple prompt engineering without explicit demographic specification does not remove the biases; neutral prompts still produce skewed outputs.
- The observed biases intersect (e.g., race and gender combine, as in 'sushi maker' being Asian female), so mitigation must be intersectional.
Reading between the lines
- A concrete next test would be to measure inter-annotator agreement on a sample of the generated images; if agreement is low for race or age categories, the reported percentages may be unstable.
- If the prompt battery is released as claimed, others could run the same prompts on newer models (e.g., SDXL, DALL-E 3) to track whether biases recede or shift with scale and fine-tuning.
- The 'a person eating watermelon' result hints at cultural-context effects; a follow-up could systematically vary the food item and nationality to map how culinary prompts trigger gendered and racialized defaults.
- The study's reliance on binary gender and coarse race bins may obscure non-binary and multiracial representations; an extension could use open-ended descriptions or continuous skin-tone scales.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an empirical audit of social biases in two text-to-image models, Stable Diffusion 1.5 and Flux-1, and compares them with Google Image Search results. The authors generated over 16,000 images from 160 prompt topics spanning occupations, attributes, actions, ideologies, emotions, family descriptions, place descriptions, religion, and life events, then manually labeled the images for gender, race, age, somatotype, and place/religion categories. They report percentage breakdowns for positive versus negative prompt groups and conclude that the models exhibit significant disparities that often reinforce harmful stereotypes.
Significance. If the quantitative results were fully substantiated, the paper would be a useful broadening of bias evaluation in T2I models beyond the usual occupation/gender/skin-tone focus, covering actions, ideologies, emotions, family structures, place descriptions, and life events. The study has notable strengths: two architecturally distinct model families (UNet-based SD1.5 and DiT-based Flux-1), consistent generation settings, a comparison corpus from Google Image Search, and qualitative examples that illustrate the phenomena. However, the central quantitative claim of 'significant disparities' is currently unsupported by the reported evidence because the annotation process, sample sizes, and statistical measures are not disclosed.
major comments (4)
- [Section 3.1/3.2, Tables 1-3] The central quantitative claim rests on percentage breakdowns that are not auditable. The paper does not report the number of images per prompt group or per demographic cell, the number of annotators, or any inter-annotator reliability statistic (e.g., Cohen's kappa). Section 3.1 states only that images 'have been labeled by multiple human operators,' and Section 3.2 acknowledges the subjectivity of somatotype labeling. Without per-cell counts and reliability data, the percentages in Tables 1-3 could be dominated by a small number of topics or by annotator disagreement. The authors should report sample sizes, an agreement metric, and ideally release the annotations and prompts to support the headline claim of significant disparities.
- [Section 3.1] The exclusion of 'distorted, unclear, abstract, or nonsensical' images is described with no criteria, no counts, and no per-model or per-prompt-group exclusion rates. Post hoc filtering can differentially remove images by demographic content, which would directly bias every percentage in Tables 1-3. The filtering protocol and exclusion statistics must be reported for the results to be interpretable.
- [Abstract and Section 4] The word 'significant' is used throughout without any statistical support. For example, Section 4 states that 'racial bias was insignificant' for attribute prompts and describes percentage gaps as 'striking' and 'significant,' but no confidence intervals, hypothesis tests, or multiple-comparison corrections are reported for any table. The authors should compute appropriate uncertainty measures (e.g., bootstrap confidence intervals or chi-square tests) for the differences they highlight; otherwise the abstract's claim of 'significant disparities' is not justified.
- [Section 3.2] The grouping of prompts into 'positive' and 'negative' categories is a normative assumption that drives many of the paper's headline comparisons (e.g., positive vs. negative attribute, action, family, and emotion prompts). No validation of this valence classification is provided, and there is no neutral-prompt baseline or sensitivity analysis. Since a misclassification of even a few prompts could alter the reported percentages, the authors should justify the classification (e.g., with independent valence ratings) or show that the results are robust to alternative groupings.
minor comments (6)
- [Section 4, 'Bias in Roles'] There is a typo: 'the rations were close' should read 'the ratios were close,' and later 'SD.15' should be 'SD1.5'.
- [Conclusion] The sentence 'The observed biases in political views, emotions, family structures, places, religious depictions, and life events' is a sentence fragment and should be completed or merged with the following sentence.
- [References] Reference [11] is misattributed: the U.S. Census Bureau is not authored by Black Forest Labs. Please correct the author/organization and URL.
- [Figures and Tables] The text refers to 'Figure 4' and 'Tables 1, 2, 3,' but the figures are not numbered in the captions, and Table 2's column headers do not align with the rows. Also, the table numbering should be checked for consistency with the in-text citations.
- [Section 3.1] The arithmetic is unclear: 160 topics at 'over 50 images per topic' would yield about 8,000 images, not 16,000. The relationship between the two models, the number of images per prompt, and the total count should be stated explicitly.
- [Footnote 1] The promised release of the prompt benchmark has no URL or timeline; for reproducibility, the prompts and annotations should be made available with the paper.
Circularity Check
No significant circularity: the paper is an empirical audit whose reported disparities are measured outcomes, not consequences of fitted parameters or self-citations.
full rationale
This paper is an empirical measurement study, not a derivation. The pipeline is: manually curated prompts → generate images with fixed checkpoints → human annotators label demographic categories → compare percentage distributions. The central claim (that T2I models exhibit disparities across gender, race, age, and somatotype) is a reported observation about generated image sets, not a quantity derived from an equation or a fitted parameter. The authors never fit a model to a subset of data and then "predict" a closely related quantity; the percentages in Tables 1–3 are descriptive statistics of human-assigned labels. The prompt categories ("positive" vs. "negative", high-income vs. low-income occupations) are author-chosen inputs, but they do not by construction force the reported demographic outcomes: for example, labeling a prompt as "negative attribute" does not mathematically imply that 91% of SD1.5 images will be labeled Male, or that negative-place prompts will yield 37% African-coded images. Those are contingent empirical findings. The paper's demographic label definitions (Section 3.2) are operational definitions for annotation, not self-referential derivations of the conclusion. There are no load-bearing self-citations: the references are to prior external bias audits and technical model papers, and no uniqueness theorem or prior result by the same authors is invoked to forbid alternatives. The acknowledged subjectivity in somatotype labeling and the absence of inter-annotator reliability statistics are methodological validity/robustness concerns, not circularity. For this reason, the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Prompt wording per topic
- Image filtering threshold
- Annotation label set
assumptions (4)
- domain assumption Gender, race, age, and somatotype can be reliably inferred from visual cues in generated images
- domain assumption Google Image Search provides a valid external reference for comparison
- domain assumption The selected 160 prompt topics are representative of real-world T2I usage
- domain assumption Stable Diffusion 1.5 and Flux-1 are representative enough of current T2I models to support general conclusions
Cite this review
Pith. "Pith review of Hidden Bias in the Machine: Stereotypes in Text-to-Image Models." pith.science (2026). https://pith.science/paper/DB4U3THO
@misc{pith2026250613780,
author = {Pith},
title = {Pith review of: Hidden Bias in the Machine: Stereotypes in Text-to-Image Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/DB4U3THO}},
note = {Machine review of arXiv:2506.13780}
}
read the original abstract
Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To investigate these issues, we curated a diverse set of prompts spanning thematic categories such as occupations, traits, actions, ideologies, emotions, family roles, place descriptions, spirituality, and life events. For each of the 160 unique topics, we crafted multiple prompt variations to reflect a wide range of meanings and perspectives. Using Stable Diffusion 1.5 (UNet-based) and Flux-1 (DiT-based) models with original checkpoints, we generated over 16,000 images under consistent settings. Additionally, we collected 8,000 comparison images from Google Image Search. All outputs were filtered to exclude abstract, distorted, or nonsensical results. Our analysis reveals significant disparities in the representation of gender, race, age, somatotype, and other human-centric factors across generated images. These disparities often mirror and reinforce harmful stereotypes embedded in societal narratives. We discuss the implications of these findings and emphasize the need for more inclusive datasets and development practices to foster fairness in generative visual systems.
Figures
Reference graph
Works this paper leans on
-
[1]
DALL-E-3: Improving Image Genera- tion with Better Captions.Computer Science, 2023
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, and etal. DALL-E-3: Improving Image Genera- tion with Better Captions.Computer Science, 2023. 1
work page 2023
-
[2]
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In IJCNLP, 2021. 2
work page 2021
-
[3]
Junming Chen, Zichun Shao, and Bin Hu. Generating In- terior Design from Text: A New Diffusion Model-Based Method for Efficient Creative Design.Buildings, 2023. 1
work page 2023
-
[4]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Syn- thesis.arXiv:2310.00426, 2023. 1
-
[5]
Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, and Matthew Turk. Tibet: Identifying and Evaluating Biases in Text-to-Image Genera- (rounded %) White Black E.Asian Latino Middle E Other Male Female Other Underweight Average Overweight Young Adult Senior Attributes Positive 98 0 0 2 0 0 68 32 0 20 66 14 6 70 24 A...
work page 2024
-
[6]
Dall-eval: Probing the Reasoning Skills and Social Biases of Text-to- Image Generation Models
Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the Reasoning Skills and Social Biases of Text-to- Image Generation Models. InCVPR, pages 3043–3054,
-
[7]
Publications Office of the European Union, 2024
Europol Innovation Lab.Facing Reality? – Law Enforce- ment and the Challenge of Deepfakes – An observatory re- port. Publications Office of the European Union, 2024. 1
work page 2024
-
[8]
Beat Biden.https://www.youtube.com/ watch?v=kLMMxgtxQ1Y&t=32s, 2023
GOP. Beat Biden.https://www.youtube.com/ watch?v=kLMMxgtxQ1Y&t=32s, 2023. 1
work page 2023
Show all 29 references
-
[9]
Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models.CVPR,
Jiuxiang Gu, Jianfei Cai, Shafiq Joty, Li Niu, and Gang Wang. Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models.CVPR,
-
[10]
Exploring text- to-image generation models: Applications and cloud re- source utilization.Elsevier Computers and Electrical En- gineering, 123, 2025
Sahani Jaiprakash and Choudhary Prakash. Exploring text- to-image generation models: Applications and cloud re- source utilization.Elsevier Computers and Electrical En- gineering, 123, 2025. 1
2025
-
[11]
Black Forest Labs. U.s. census bureau.https://www. census.gov, 2020. 3
2020
-
[12]
Flux.https://github.com/ black-forest-labs/flux, 2024
Black Forest Labs. Flux.https://github.com/ black-forest-labs/flux, 2024. 1
2024
-
[13]
StoryGAN: A Sequential Conditional GAN for Story Visualization .CVPR, 2019
Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. StoryGAN: A Sequential Conditional GAN for Story Visualization .CVPR, 2019. 1
2019
-
[14]
Stable Bias: Analyzing So- cietal Representations in Diffusion Models.NeurIPS, 2023
Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. Stable Bias: Analyzing So- cietal Representations in Diffusion Models.NeurIPS, 2023. 1, 2
2023
-
[15]
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
Nila Masrourisaadat, Nazanin Sedaghatkish, Fatemeh Sar- shartehrani, and Edward A Fox. Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models. arXiv:2407.00138, 2024. 1, 2
2024 arXiv
-
[16]
Social Biases Through the Text-to-Image Generation Lens
Ranjita Naik and Besmira Nushi. Social Biases Through the Text-to-Image Generation Lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 786–808, 2023. 1, 2
2023
-
[17]
Multilingual Diversity Improves Vision- Language Representations.arXiv:2405.16915, 2024
Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei- Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei Koh, and Ranjay Krishna. Multilingual Diversity Improves Vision- Language Representations.arXiv:2405.16915, 2024. 2
2024
-
[18]
Modelling agency — deep agency.https: //www.deepagency.com/, 2024
Danny Postma. Modelling agency — deep agency.https: //www.deepagency.com/, 2024. 1
2024
-
[19]
Hierarchical Text-Conditional Image Gen- eration with CLIP Latents.arXiv:2204.06125, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical Text-Conditional Image Gen- eration with CLIP Latents.arXiv:2204.06125, 2022. 1
2022 arXiv
-
[20]
High-Resolution Image Synthesis with Latent Diffusion Models .CVPR, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models .CVPR, 2022. 1
2022
-
[21]
Fast High- Resolution Image Synthesis with Latent Adversarial Diffu- sion Distillation.arXiv:2403.12015, 2024
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast High- Resolution Image Synthesis with Latent Adversarial Diffu- sion Distillation.arXiv:2403.12015, 2024. 1
2024 arXiv
-
[22]
LAION- 400M: Open Dataset of CLIP-filtered 400 Million Image- Text Pairs.arXiv:2111.02114, 2021
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION- 400M: Open Dataset of CLIP-filtered 400 Million Image- Text Pairs.arXiv:2111.02114, 2021. 1
2021 arXiv
-
[23]
Andrew Shaw, Andre Ye, Ranjay Krishna, and Amy X. Zhang. Unsettling the Hegemony of Intention: Agonistic Image Generation.arXiv:2502.15242, 2025. 1, 2
2025 arXiv
-
[24]
ReStGAN: A step towards visually guided shopper experi- ence via text-to-image synthesis.WACV, 2020
Shiv Surya, Amrith Setlur, Arijit Biswas, and Sumit Negi. ReStGAN: A step towards visually guided shopper experi- ence via text-to-image synthesis.WACV, 2020. 1
2020
-
[25]
Survey of Bias In Text- to-Image Generation: Definition, Evaluation, and Mitiga- tion.arXiv:2404.01030, 2024
Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Re- becca Pattichis, and Kai-Wei Chang. Survey of Bias In Text- to-Image Generation: Definition, Evaluation, and Mitiga- tion.arXiv:2404.01030, 2024. 1, 2
2024 arXiv
-
[26]
Fleet, Radu Soricut, Jason Baldridge, Mo- hammad Norouzi, Peter Anderson, and William Chan
Su Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont- Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mo- hammad Norouzi, Peter Anderson, and William Chan. Ima- gen Editor and EditBench: Advancing and Evaluati...
2023
-
[27]
Hwang, Amy X
Andre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang, and Ranjay Krishna. Cultural and Linguistic Diversity Im- proves Visual Representations.arXiv:2310.14356, 2024. 2
2024 arXiv
-
[28]
Text-to-Image Synthesis: A Decade Survey.arXiv:2411.16164, 2024
Nonghai Zhang and Hao Tang. Text-to-Image Synthesis: A Decade Survey.arXiv:2411.16164, 2024. 1
2024 arXiv
-
[29]
SINE: SINgle Image Editing with Text-to-Image Diffusion Models.CVPR, 2022
Zhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris Metaxas, and Jian Ren. SINE: SINgle Image Editing with Text-to-Image Diffusion Models.CVPR, 2022. 1
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.