REVIEW 3 major objections 5 minor 95 references
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Existing automated gender-bias detectors for text-to-image models do not reproduce human-labeled bias, with one overestimating it by 26.95%; a combined face-detection and CLIP detector closes the gap.
desk verdict An overdue head-to-head validation of gender-bias detectors for T2I models; the recall insight is real and useful, but the quantitative claims need error bars and the proposed enhancement is evaluated too optimistically. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's analytic machinery decomposes a gender-bias detector into three stages, filtering low-quality images, classifying gender, and computing bias scores, and then isolates detector error to the filtering and classification stages. The load-bearing metrics are the model bias score, the average per-prompt absolute male-female difference divided by the total, and the prompt bias score; the proposed CLIP-Enhance detector combines a face-detection model, YOLOv8-based multi-person filtering and cropping, and CLIP zero-shot gender classification. The face-detection stage is what lets CLIP-Enhance filter out 82.91% of low-quality images while keeping a recall of 97.54% on clear images, which is the property that existing vision-language-only detectors lack.
What would settle it
A re-run of the seven detectors on a fresh human-labeled sample from the same three text-to-image models, using the original authors' code or exact configurations, that shows deviations far smaller than 26.95% for CLIP-Prob would falsify the claim that the published detectors mis-measure bias, pointing instead to an implementation artifact in this study's setup.
Extended reading notes
Core claim
The central discovery is that widely used automated gender-bias detectors do not accurately capture the bias that human annotators see in text-to-image model outputs, and the mismatch is driven mostly by the filtering step rather than the gender-classification step. On a manually labeled set of 6,000 images, the ground-truth model bias scores are 0.752 for SDXL, 0.730 for SD3, and 0.631 for Dreamlike; the seven evaluated detectors deviate from these scores by up to 26.95%, and detectors with classification accuracy above 95% (CLIP-Prob, BLIP-2) still mis-measure bias because they discard clear images or keep low-quality ones. The paper identifies the cause and a remedy: a face-detection model filters images without clear faces, YOLOv8 removes and crops multi-person images, and CLIP then classifies gender. This CLIP-Enhance pipeline reports model-bias scores within 0.47% to 1.23% of the human labels and the lowest prompt-level error across all three models.
Load-bearing premise
The study assumes that the seven detector implementations it builds, for example CLIP-Prob with a face detector and a 90% confidence cutoff, faithfully match the detectors as originally proposed, so that the measured deviations reflect the original tools rather than artifacts of this study's reimplementation.
Editorial extensions
If this is right
- Published bias numbers for text-to-image models that rely on CLIP, CLIP-Prob, CLIP-Uncertain, or BLIP-2 are likely to overestimate or underestimate the true bias, so model rankings based on those detectors may be wrong; the paper shows CLIP-Prob would rank Dreamlike as the second most biased model when it is actually the least biased.
- A detector with very high gender-classification accuracy can still fail at bias measurement if its filtering stage has low recall, meaning future bias studies should report filtering performance, not just classification accuracy.
- Face-detection-based filtering combined with a vision-language classifier is a reliable recipe: MiVOLO, FairFace, and CLIP-Enhance all stay within a few percentage points of the human-labeled bias, while detectors lacking face-based filtering do not.
- The proposed CLIP-Enhance detector provides a concrete measuring stick for future text-to-image fairness claims, with model-bias scores within 0.47% to 1.23% and prompt-level errors of 0.065, 0.048, and 0.073 on the three tested models.
- The findings imply that fairness testing for generative models should treat low-quality-image filtering as a first-class evaluation step, since a 12.48% average rate of low-quality images is large enough to distort any downstream gender distribution.
Reading between the lines
- If the detector discrepancy extends beyond the three open-source models tested, earlier conclusions about which text-to-image model is most biased may need re-examination, because the bias ranking itself can flip depending on the detector used.
- The same filtering failure likely affects bias measurements for race, age, and other demographic attributes, since any low-quality image corrupts the downstream classifier regardless of the attribute being measured.
- A natural testable extension is to run CLIP-Enhance on closed models such as DALL-E 3 or Imagen with a human-labeled sample, to see whether the 0.47% to 1.23% accuracy holds outside the open-source model family.
- The 12.48% low-quality-image rate may itself carry bias-relevant information: a model that generates more unreadable images for certain prompts could be hiding a stereotype rather than correcting it, so filtering these images out might understate the bias a user actually experiences.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper asks whether existing automated detectors can accurately measure gender bias in text-to-image models. It constructs a dataset of 6,000 images from SDXL, SD3, and Dreamlike Photoreal 2.0 using 100 gender-neutral prompts, manually labels each image as male, female, or low-quality (Cohen's Kappa 0.86), and compares the model bias scores and prompt bias scores computed from seven detectors from prior work with the human-labeled ground truth. The paper reports that all three models generate more male than female images, that profession prompts are the most biased category, that none of the seven detectors reproduces the human bias scores (CLIP-Prob overestimates Dreamlike's model bias score by 26.95%), and that vision-language model-based detectors are poor at filtering low-quality images. Based on these findings, the authors propose CLIP-Enhance, which combines dlib face detection, YOLOv8 multi-person filtering, and CLIP classification, and claim it achieves model-bias-score differences of only 0.47%-1.23% and filters 82.91% of low-quality images. The dataset and code are publicly available.
Significance. If the findings hold, this paper provides a useful, sobering benchmark for T2I bias testing: it would show that common automated detectors can deviate considerably from human judgments and that the filtering step is a major source of error. The strengths are the release of a human-labeled dataset with high inter-annotator agreement, the decomposition of detector errors into filtering and classification, and the reproducible study design. However, the bold quantitative claims are currently not backed by statistical uncertainty analysis, and the fidelity of the seven detector implementations to the originally proposed methods is not established; both are fixable and are addressed in the major comments.
major comments (3)
- [§3.1 and Appendix A] The RQ2 headline deviations, including the 26.95% overestimate for CLIP-Prob on Dreamlike (Table 1), are meaningful only if the seven detector implementations faithfully correspond to the original proposals. The manuscript provides no fidelity check: it does not compare against the original authors' code or published outputs, and for CLIP-Prob it is not demonstrated that Seshadri et al. [73] use MediaPipe face detection with a 90% CLIP-similarity cutoff as implemented here (Appendix A). If the face detector or threshold differs, the reported 80.9% filtering of clear images and the resulting bias deviation could be artifacts of this study's setup rather than properties of the original detectors. Please add a fidelity analysis (e.g., reproducing a sample of the original papers' reported results or a sensitivity analysis over face detectors and confidence thresholds) and report the provenance of each implementation.
- [§4.2, Tables 1 and 2] All detector comparisons are point estimates with no confidence intervals or significance tests. Because each prompt-model cell contains only 20 images and T2I generation is stochastic, the differences among the better detectors (e.g., CLIP-Enhance 0.53% vs. MiVOLO 0.93% for SDXL in Table 1) may lie within sampling noise, and the same may hold for the prompt bias score differences in Table 2. Please provide standard errors, bootstrap confidence intervals, or a per-prompt paired test to support the ranking 'CLIP-Enhance is the most accurate detector' and the claim that existing detectors are inaccurate.
- [§5.1] CLIP-Enhance is designed by inspecting the failure modes of the same dataset on which it is evaluated, and its key free parameter—the 50% second-person bounding-box-area ratio for multi-person filtering—is hand-chosen without any held-out data or sensitivity analysis. Measuring the detector on the data used for its design can overstate the reported 0.47%-1.23% model-bias-score accuracy and 82.91% filter rate. Please evaluate CLIP-Enhance on a held-out set of prompts or models, or at minimum present a sensitivity analysis over the 50% threshold.
minor comments (5)
- [Appendix A and B] The CLIP-Uncertain prompt contains a typo: 'a phot of a person' should be 'a photo of a person'.
- [§3.3 and Table 5] The text reports average male/female percentages of 63.57%/24.18%, while Table 5 gives 63.50%/24.02%; please reconcile the numbers.
- [§4.2] The text gives FairFace prompt bias score differences as 0.093, 0.055, and 0.109, but Table 2 reports 0.092, 0.053, and 0.108; the values are inconsistent.
- [§1 and §4.2] The statement that CLIP's deviation is 'seven times more' than FairFace holds only for SD3 (3.97%/0.55% = 7.2); on Dreamlike CLIP is actually closer than FairFace (4.91% vs 5.55%). Please qualify the claim per model.
- [§3.3] The image generation step does not report random seeds or sampling parameters; adding these details would improve reproducibility.
Circularity Check
No significant circularity: the RQ2 comparison is anchored to independent human labels, and the paper's self-citations are non-load-bearing; the in-sample design of CLIP-Enhance and the unvalidated CLIP-Prob reimplementation are validity caveats, not circular reductions.
full rationale
The paper's derivation chain is: (i) build a 6,000-image dataset from three T2I models; (ii) obtain ground-truth gender labels from two human annotators (Cohen's kappa 0.86, Section 3.3); (iii) run seven detectors on the same images (Appendix A); (iv) compute model and prompt bias scores (Eqs. 1 and 2) from each labeling source; and (v) measure the discrepancy via Eq. 3. At no point is any detector's output defined in terms of the human labels, nor is the human ground truth derived from a detector. The central RQ2 finding — 'None of the detectors can accurately capture the gender bias in T2I models, with some overestimating bias by as much as 26.95%' — is therefore a comparison against an independent external benchmark, not a self-fulfilling construction. Eqs. 1-3 are definitional metrics but they do not encode the result; the 26.95% figure is an empirical consequence of which images CLIP-Prob's 90%-confidence filter discards. The self-citations present (e.g., Lyu et al. 2024 and 2023 for kappa-threshold and metric conventions; Yang et al./Lo-group fairness papers in Related Work) are non-load-bearing: none is invoked to justify a detector's design or to rule out alternatives. The only circularity-adjacent step is Section 5.1's CLIP-Enhance, whose components (dlib face detection, YOLOv8 multi-person filter with a 50% area threshold, cropping) are 'based on the empirical findings' obtained on the very same 6,000 images, and which is then evaluated on that same set with claimed deviations of 0.47%-1.23%. This is in-sample model selection rather than a derivation that reduces to its inputs: the 50% threshold is an un-fitted hand choice, and the reported accuracy is a measurement that could plausibly have been worse, so the result is not forced by construction. It does, however, mean the 0.47%-1.23% figure is not an out-of-sample validation. A separate caveat, outside circularity: the paper does not verify that its Appendix A reimplementation of CLIP-Prob (MediaPipe face detector plus 90% threshold) matches Seshadri et al.'s original, so the 26.95% overestimate may be an artifact of this study's configuration; that is a construct-validity threat, not circularity. The paper's own stated limitations (Section 7: binary gender framing, prompt/T2I model coverage) likewise bound generality without introducing a circular step.
Assumptions & free parameters
free parameters (1)
- Multi-person filtering ratio (CLIP-Enhance) =
second-largest YOLOv8 bounding box area > 50% of largest triggers filtering
assumptions (4)
- domain assumption Gender is treated as binary (male/female); non-binary identities are excluded from measurement
- domain assumption A fair T2I model should generate equal numbers of male and female images for gender-neutral prompts
- domain assumption Human labels from two annotators are the ground truth for both gender and image quality
- domain assumption Gender bias is computed only on 'clear' images; low-quality images are excluded from the denominator
Cite this review
Pith. "Pith review of Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?." pith.science (2026). https://pith.science/paper/XBJN6FBP
@misc{pith2026250115775,
author = {Pith},
title = {Pith review of: Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?},
year = {2026},
howpublished = {\url{https://pith.science/paper/XBJN6FBP}},
note = {Machine review of arXiv:2501.15775}
}
read the original abstract
Text-to-Image (T2I) models have recently gained significant attention due to their ability to generate high-quality images and are consequently used in a wide range of applications. However, there are concerns about the gender bias of these models. Previous studies have shown that T2I models can perpetuate or even amplify gender stereotypes when provided with neutral text prompts. Researchers have proposed automated gender bias uncovering detectors for T2I models, but a crucial gap exists: no existing work comprehensively compares the various detectors and understands how the gender bias detected by them deviates from the actual situation. This study addresses this gap by validating previous gender bias detectors using a manually labeled dataset and comparing how the bias identified by various detectors deviates from the actual bias in T2I models, as verified by manual confirmation. We create a dataset consisting of 6,000 images generated from three cutting-edge T2I models: Stable Diffusion XL, Stable Diffusion 3, and Dreamlike Photoreal 2.0. During the human-labeling process, we find that all three T2I models generate a portion (12.48% on average) of low-quality images (e.g., generate images with no face present), where human annotators cannot determine the gender of the person. Our analysis reveals that all three T2I models show a preference for generating male images, with SDXL being the most biased. Additionally, images generated using prompts containing professional descriptions (e.g., lawyer or doctor) show the most bias. We evaluate seven gender bias detectors and find that none fully capture the actual level of bias in T2I models, with some detectors overestimating bias by up to 26.95%. We further investigate the causes of inaccurate estimations, highlighting the limitations of detectors in dealing with low-quality images. Based on our findings, we propose an enhanced detector...
Figures
Reference graph
Works this paper leans on
-
[73]
Preethi Seshadri, Sameer Singh, and Yanai Elazar. 2023. The bias amplification paradox in text-to-image generation. arXiv preprint arXiv:2308.00755 (2023)
arXiv 2023
-
[1]
Stability AI. 2024. Stable Diffusion 3 Medium. https://huggingface.co/stabilityai/ stable-diffusion-3-medium
2024
-
[2]
Stability AI. 2024. Stable Diffusion XL Base 1.0. https://huggingface.co/stabilityai/ stable-diffusion-xl-base-1.0
2024
-
[3]
Ana Kessler. 2023. Breakdown of Coca-Cola Commercial Made with Stable Diffusion Revealed. https://80.lv/articles/breakdown-of-coca-cola-commercial- made-with-stable-diffusion-revealed/
2023
-
[4]
T2IReplication Anonymous. 2024. T2IReplication-ISSTA25. (10 2024). https: //doi.org/10.6084/m9.figshare.27377649.v1
-
[5]
Muhammad Hilmi Asyrofi, Zhou Yang, Imam Nur Bani Yusuf, Hong Jin Kang, Ferdian Thung, and David Lo. 2021. Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems. IEEE Transactions on Software Engineering 48, 12 (2021), 5087–5101
2021
-
[6]
Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. 2022. How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 1358–1370
2022
-
[7]
Solon Barocas and Andrew D Selbst. 2016. Big data’s disparate impact. Calif. L. Rev. 104 (2016), 671
2016
Show all 95 references
-
[8]
Richard Berk, Hoda Heidari, Shahin Jabbari, Michael Kearns, and Aaron Roth
-
[9]
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al . 2023. Improving im- age generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2, 3 (2023), 8
2023
-
[10]
Yuriy Brun and Alexandra Meliou. 2018. Software fairness. In Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering . 754–759
2018
-
[11]
Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accu- racy disparities in commercial gender classification. In Conference on fairness, accountability and transparency. PMLR, 77–91
2018
-
[12]
Alessandro Castelnovo, Riccardo Crupi, Greta Greco, Daniele Regoli, Ilaria Giuseppina Penco, and Andrea Claudio Cosentini. 2022. A clarification of the nuances in the fairness metrics landscape. Scientific Reports 12, 1 (2022), 4209
2022
-
[13]
Zhang, Max Hort, Mark Harman, and Federica Sarro
Zhenpeng Chen, Jie M. Zhang, Max Hort, Mark Harman, and Federica Sarro
-
[14]
Zhenpeng Chen, Jie M Zhang, Federica Sarro, and Mark Harman. 2023. A com- prehensive empirical study of bias mitigation methods for machine learning classifiers. ACM transactions on software engineering and methodology 32, 4 (2023), 1–30
2023
-
[15]
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3043–3054
2023
-
[16]
Tom Davenport. 2023. Cuebric: Generative AI Comes to Hollywood. (2023). https://www.forbes.com/sites/tomdavenport/2023/03/13/cuebric- generative-ai-comes-to-hollywood/?sh=19b07abb174b
2023
-
[17]
Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle
-
[18]
Dlib. 2017. High quality face recognition. http://dlib.net/dnn_face_recognition_ ex.cpp.html. Accessed: 2024-07-02
2017
-
[19]
dreamlike art. 2023. Dreamlike Photoreal 2.0. https://huggingface.co/dreamlike- art/dreamlike-photoreal-2.0
2023
-
[20]
Eran Eidinger, Roee Enbar, and Tal Hassner. 2014. Age and gender estimation of unfiltered faces. IEEE Transactions on information forensics and security 9, 12 (2014), 2170–2179
2014
-
[21]
Piero Esposito, Parmida Atighehchian, Anastasis Germanidis, and Deepti Ghadi- yaram. 2023. Mitigating stereotypical biases in text to image generative systems. arXiv preprint arXiv:2310.06904 (2023)
2023 arXiv
-
[22]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis. arXiv preprint arXiv:2403.03206 (2024)
2024 arXiv
-
[23]
Fairness analysis
Anthony Finkelstein, Mark Harman, S Afshin Mansouri, Jian Ren, and Yuanyuan Zhang. 2008. “Fairness analysis” in requirements assignments. In 2008 16th IEEE International Requirements Engineering Conference . IEEE, 115–124
2008
-
[24]
Kathleen C Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023. Diversity is not a one-way street: Pilot study on ethical interventions for racial bias in text-to-image systems. ICCV, accepted (2023)
2023
-
[25]
Kathleen C Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023. A friendly face: Do text-to-image systems rely on stereotypes when the input is under- specified? arXiv preprint arXiv:2302.07159 (2023)
2023 arXiv
-
[26]
Felix Friedrich, Manuel Brack, Lukas Struppek, Dominik Hintersdorf, Patrick Schramowski, Sasha Luccioni, and Kristian Kersting. 2023. Fair diffusion: Instruct- ing text-to-image generation models on fairness. arXiv preprint arXiv:2302.10893 (2023)
2023 arXiv
-
[27]
Felix Friedrich, Katharina Hämmerl, Patrick Schramowski, Jindrich Libovicky, Kristian Kersting, and Alexander Fraser. 2024. Multilingual Text-to-Image Gen- eration Magnifies Gender Stereotypes and Prompt Engineering May Not Help You. arXiv preprint arXiv:2401.16092 (2024)
2024 arXiv
-
[28]
gofundme. 2023. GoFundMe | Help Changes Everything. https://www.youtube. com/watch?v=NqdC0WX-f6o
2023
-
[29]
Google. 2024. MediaPipe Face Detector - Python API. https://ai.google.dev/ edge/mediapipe/solutions/vision/face_detector/python Accessed: April 10, 2025
2024
-
[30]
GOP. 2023. Beat Biden. https://www.youtube.com/watch?v=kLMMxgtxQ1Y&t= 32s
2023
-
[31]
Nina Grgic-Hlaca, Muhammad Bilal Zafar, Krishna P Gummadi, and Adrian Weller. 2016. The case for process fairness in learning: Feature selection for fair decision making. In NIPS symposium on machine learning and the law , Vol. 1. Barcelona, Spain, 11
2016
-
[32]
Huizhong Guo, Jinfeng Li, Jingyi Wang, Xiangyu Liu, Dongxia Wang, Zehong Hu, Rong Zhang, and Hui Xue. 2023. FairRec: Fairness testing for deep recommender systems. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 310–321
2023
-
[33]
Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems 29 (2016)
2016
-
[34]
Deborah Hellman. 2020. Measuring algorithmic fairness. Virginia Law Review 106, 4 (2020), 811–866
2020
-
[35]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840–6851
2020
-
[36]
Glenn Jocher, Ayush Chaurasia, and Jing Qiu. 2023. Ultralytics YOLO. https: //github.com/ultralytics/ultralytics
2023
-
[37]
Kimmo Karkkainen and Jungseock Joo. 2021. Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1548–1558
2021
-
[38]
Os Keyes, Chandler May, and Annabelle Carrell. 2021. You keep using that word: Ways of thinking about gender in computing research. Proceedings of the ACM on human-computer interaction 5, CSCW1 (2021), 1–23
2021
-
[39]
Eunji Kim, Siwon Kim, Chaehun Shin, and Sungroh Yoon. 2023. De-stereotyping text-to-image models through prompt tuning. (2023)
2023
-
[40]
Maksim Kuprashevich and Irina Tolstykh. 2023. Mivolo: Multi-input transformer for age and gender estimation. arXiv preprint arXiv:2307.04616 (2023)
2023 arXiv
-
[41]
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017. Counterfac- tual fairness. Advances in neural information processing systems 30 (2017)
2017
-
[42]
J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159–174
1977
-
[43]
Julia Kaiwen Lau, Kelvin Kai Wen Kong, Julian Hao Yong, Per Hoong Tan, Zhou Yang, Zi Qian Yong, Joshua Chern Wey Low, Chun Yong Chong, Mei Kuan Lim, and David Lo. 2023. Synthesizing Speech Test Cases with Text-to-Speech? An Empirical Study on the False Alarms in Automated Spee...
2023
-
[44]
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, et al. 2024. Holistic evaluation of text-to-image models. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[45]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742
2023
-
[46]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900
2022
-
[47]
Alexander Lin, Lucas Monteiro Paes, Sree Harsha Tanneru, Suraj Srinivas, and Himabindu Lakkaraju. 2023. Word-Level Explanations for Analyzing Bias in Text-to-Image Models. arXiv preprint arXiv:2306.05500 (2023)
2023 arXiv
-
[48]
Yiming Lin, Jie Shen, Yujiang Wang, and Maja Pantic. 2022. Fp-age: Leveraging face parsing attention for facial age estimation in the wild. IEEE Transactions on Image Processing (2022). 9 MM ’25, October 27–31, 2025, Dublin, Ireland Lyu et al
2022
-
[49]
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang
-
[50]
Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2024. Stable bias: Evaluating societal representations in diffusion models. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[51]
Yunbo Lyu, Hong Jin Kang, Ratnadira Widyasari, Julia Lawall, and David Lo
-
[52]
Le, Ming Li, and David Lo
Yunbo Lyu, Thanh Le-Cong, Hong Jin Kang, Ratnadira Widyasari, Zhipeng Zhao, Xuan-Bach D. Le, Ming Li, and David Lo. 2023. Chronos: Time-Aware Zero-Shot Identification of Libraries from Vulnerability Reports. In Proceedings of the 45th International Conference on Software Engin...
2023
-
[53]
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. 2015. Generating images from captions with attention. arXiv preprint arXiv:1511.02793 (2015)
2015 arXiv
-
[54]
Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276–282
2012
-
[55]
Megvii. 2024. Face++ Cognitive Services. https://www.faceplusplus.com/
2024
-
[56]
IEEE Trans
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel. IEEE Trans. Softw. Eng. 50, 9 (Sept. 2024), 2219–2239. https://doi.org/10.1109/ TSE.2024.3406718
2024
-
[57]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM com- puting surveys (CSUR) 54, 6 (2021), 1–35
2021
-
[58]
Shira Mitchell, Eric Potash, Solon Barocas, Alexander D’Amour, and Kristian Lum. 2021. Algorithmic fairness: Choices, assumptions, and definitions. Annual review of statistics and its application 8, 1 (2021), 141–163
2021
-
[59]
Ranjita Naik and Besmira Nushi. 2023. Social biases through the text-to-image generation lens. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. 786–808
2023
-
[60]
Author’s Full Name. 2024. 90% of Online Content Could Be Gen- erated by AI by 2025, Expert Says. Yahoo Finance (2024). https: //finance.yahoo.com/news/90-of-online-content-could-be-generated-by- ai-by-2025-expert-says-201023872.html Accessed: 2024-08-22
2024
-
[61]
Ninareh Mehrabi, Thamme Gowda, Fred Morstatter, Nanyun Peng, and Aram Galstyan. 2020. Man is to person as woman is to location: Measuring gender bias in named entity recognition. In Proceedings of the 31st ACM conference on Hypertext and Social Media . 231–232
2020
-
[62]
Evgeny Obedkov. 2023. How AI-Assisted RPG Tales of Syn Utilizes Stable Diffu- sion and ChatGPT to Create Assets and Dialogues. https://gameworldobserver. com/2023/03/06/tales-of-syn-ai-rpg-stable-diffusion-chatgpt-game Accessed: 2024-08-22
2023
-
[63]
OpenAI. 2022. DALL-E 2. https://openai.com/index/dall-e-2/
2022
-
[64]
OpenAI. 2023. Dall ·E 3 System Card. https://openai.com/research/dall-e-3- system-card
2023
-
[65]
Hadas Orgad, Bahjat Kawar, and Yonatan Belinkov. 2023. Editing implicit as- sumptions in text-to-image diffusion models. arXiv preprint arXiv:2303.08084 (2023)
2023 arXiv
-
[66]
Leonardo Nicoletti and Dina Bass. 2023. Humans are biased. Generative AI is even worse. https://www.bloomberg.com/graphics/2023-generative-ai-bias/
2023
-
[67]
Danny Postma. [n. d.]. AI Modelling Agency — Deep Agency. https://www. deepagency.com/
-
[68]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[69]
Sai Sathiesh Rajan, Sakshi Udeshi, and Sudipta Chattopadhyay. 2022. Aequevox: Automated fairness testing of speech recognition systems. InInternational Confer- ence on Fundamental Approaches to Software Engineering . Springer International Publishing Cham, 245–267
2022
-
[70]
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In Inter- national conference on machine learning . PMLR, 1060–1069
2016
-
[71]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023)
2023 arXiv
-
[72]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Infor...
2022
-
[74]
Ezekiel Soremekun, Sakshi Udeshi, and Sudipta Chattopadhyay. 2022. Astraea: Grammar-based fairness testing. IEEE Transactions on Software Engineering 48, 12 (2022), 5188–5211
2022
-
[75]
Lin Sze Khoo, Jia Qi Bay, Ming Lee Kimberly Yap, Mei Kuan Lim, Chun Yong Chong, Zhou Yang, and David Lo. 2023. Exploring and Repairing Gen- der Fairness Violations in Word Embedding-based Sentiment Analysis Model through Adversarial Patches. In 2023 IEEE International Conferen...
2023
-
[76]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
2022
-
[77]
Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. 2024. Sur- vey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation. arXiv preprint arXiv:2404.01030 (2024)
2024 arXiv
-
[78]
Wenxuan Wang, Haonan Bai, Jen-tse Huang, Yuxuan Wan, Youliang Yuan, Haoyi Qiu, Nanyun Peng, and Michael R Lyu. 2024. New Job, New Gender? Measuring the Social Bias in Image Generation Models. arXiv preprint arXiv:2401.00763 (2024)
2024 arXiv
-
[79]
Zichong Wang, Yang Zhou, Meikang Qiu, Israat Haque, Laura Brown, Yi He, Jianwu Wang, David Lo, and Wenbin Zhang. 2023. Towards fair machine learn- ing software: Understanding and addressing model bias through counterfactual thinking. arXiv preprint arXiv:2302.08018 (2023)
2023
-
[80]
Yisong Xiao, Aishan Liu, Tianlin Li, and Xianglong Liu. 2023. Latent imitator: Generating natural individual discriminatory instances for black-box fairness testing. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis . 829–841
2023
-
[81]
Yixin Wan and Kai-Wei Chang. 2024. The Male CEO and the Female Assistant: Probing Gender Biases in Text-To-Image Models Through Paired Stereotype Test. arXiv preprint arXiv:2402.11089 (2024)
2024 arXiv
-
[82]
Zhou Yang, Muhammad Hilmi Asyrofi, and David Lo. 2021. Biasrv: Uncovering biased sentiment predictions at runtime. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1540–1544
2021
-
[83]
Zhou Yang, Harshit Jain, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2021. Biasheal: On-the-fly black-box healing of bias in sentiment analysis systems. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 644–648
2021
-
[84]
Zhou Yang, Zhensu Sun, Terry Zhuo Yue, Premkumar Devanbu, and David Lo
-
[85]
Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De la Torre. 2023. Iti-gen: Inclusive text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3969– 3980
2023
-
[86]
Junjie Yang, Jiajun Jiang, Zeyu Sun, and Junjie Chen. 2024. A Large-Scale Empir- ical Study on Improving the Fairness of Image Classification Models. In Proceed- ings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 210–222
2024
-
[87]
Lingfeng Zhang, Yueling Zhang, and Min Zhang. 2021. Efficient white-box fairness testing through gradient search. In Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis . 103–114
2021
-
[88]
Peixin Zhang, Jingyi Wang, Jun Sun, and Xinyu Wang. 2021. Fairness testing of deep image classification with adequacy metrics. arXiv preprint arXiv:2111.08856 (2021)
2021 arXiv
-
[89]
a photo of a cat
Zhifei Zhang, Yang Song, and Hairong Qi. 2017. Age progression/regression by conditional adversarial autoencoder. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5810–5818. 10 Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Im...
2017
-
[90]
arXiv preprint arXiv:2403.07506 (2024)
Robustness, security, privacy, explainability, efficiency, and usability of large language models for code. arXiv preprint arXiv:2403.07506 (2024)
2024 arXiv
-
[92]
Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu. 2020. Machine learning testing: Survey, landscapes and horizons. IEEE Transactions on Software Engineering 48, 1 (2020), 1–36
2020
-
[2018]
In Proceedings of the 2018 chi conference on human factors in computing systems
Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 chi conference on human factors in computing systems . 1–14
2018
-
[2019]
arXiv preprint arXiv:1910.10486 (2019)
Does gender matter? towards fairness in dialogue systems. arXiv preprint arXiv:1910.10486 (2019)
2019 arXiv
-
[2021]
Fairness in criminal justice risk assessments: The state of the art.Sociological Methods & Research 50, 1 (2021), 3–44
2021
-
[2024]
ACM Trans
Fairness Testing: A Comprehensive Survey and Analysis of Trends. ACM Trans. Softw. Eng. Methodol. 33, 5, Article 137 (June 2024), 59 pages. https: //doi.org/10.1145/3652155
2024 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.