{"id":"5a1b9d55-89e2-4cba-9cf5-da934b414d47","arxiv_id":"2501.01973","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new evaluation framework reports that nine popular text-to-image models mostly fail equal-representation fairness criteria, with skintone alignment errors far larger than gender errors.","lead":"This paper introduces INFELM, a fairness benchmark for text-to-image AI models, adding a new skintone classifier and metrics for demographic representation and prompt alignment. Testing nine widely used models, it concludes that most fail a simple 80 percent fairness rule and that skintone errors dominate gender errors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All fairness conclusions in Table 7 rest on INFELM's skintone classifier, but its accuracy on images from the nine evaluated models is never validated; a style-shift or group-correlated error would change every headline result.","rationale":"I read the paper in good faith. The proposed classifier improvement on WBB and High-Aes is plausible and the reported gains are large, but the headline fairness evaluation cannot be separated from how the classifier behaves on the actual generated images. The reader's weakest_assumption focuses on the uniform reference distribution (Section 4.3.2). That is a real limitation, but it is explicit and normative: fairness definitions are value-laden, and one can recompute bias scores under alternative references. The classifier generalization issue is more load-bearing because it is an unexamined empirical claim. Every number in Table 7—bias scores, alignment errors, the four-fifth-rule verdicts—is filtered through the automatic classifier. If the classifier's error rate is different on DALL-E 3 or FLUX outputs than on the synthetic High-Aes test set, or if errors are correlated with skintone group, the paper's central conclusions could be artifacts of measurement. The paper provides no human spot-check, no per-model accuracy, and no analysis of error by skintone scale. Because the reader's rationale does mention the lack of human validation on test images, my agreement is partial rather than full. The concrete human-annotation test would settle whether the concern lands. If the classifier holds up on real model outputs, the conditional acceptance is justified; if it does not, the benchmark and the headline fairness findings would need substantial revision. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":12270,"tokens_out":6200,"duration_ms":65542,"concrete_test":"Independently sample 200 generated images per model (1,800 total), stratified across prompts and predicted groups, and have at least three trained annotators assign Monk skintone scales. Compare INFELM's predictions against the annotator majority with ±1 tolerance, computing per-model precision and normalized MSE as in Eq. (6). If average precision is more than about 10 points below Table 6's High-Aes value (0.9033), or if per-scale confusion correlates with model or true skintone, then the Table 7 fairness metrics and the 'most models fail' conclusion are not supported. Publish the annotated sample and the per-model accuracy table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is twofold: (i) INFELM's skintone classifier substantially improves over prior methods, and (ii) applying it to nine text-to-image models shows that most fail the four-fifth rule, with representation bias larger than alignment error and skintone alignment error much larger than gender error. Claim (ii) is computed entirely from automatic gender and skintone labels, so it is only as strong as classifier accuracy on the actual generated images. The classifier is validated only on WBB and High-Aes (Table 6); High-Aes is an internal synthetic-image dataset, and the topology module is trained on RealisticVision-generated faces (Section 4.2.1). The nine evaluation models—DALL-E 3, FLUX, SDXL Lightning, and others—produce image styles and lighting that differ from that training distribution. The paper asserts robustness to image style (C4), but it provides no per-model accuracy, no human validation on real model outputs, and no error analysis broken down by model or skintone group. If the classifier systematically confuses adjacent Monk scales, or if error rates correlate with skintone group (e.g., darker skins under stylized renderings), then b_s and e_s in Table 7 shift in unknown directions. The stated finding that skintone alignment error is more than 60% higher than gender error could then be an artifact of classifier miscalibration rather than a property of the generators. This is distinct from the uniform reference-distribution assumption (Section 4.3.2): the uniform baseline is explicit and normative, whereas classifier generalization is an unexamined empirical claim that affects every reported fairness number.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents INFELM, a fairness evaluation framework for large text-to-image (T2I) models. It proposes a skintone classifier that fuses latent facial-topology features with a dominant-skin-pixel distribution, and evaluates content alignment error and representation bias across gender and five skintone groups. Using 246 prompts in six socially sensitive domains and 100 generated images per prompt, the authors compare nine models (Stable Diffusion v1.4/1.5/2.1, Openjourney v4, DALL-E 3, RealisticVision v5.1/6.0, SDXL Lightning, FLUX.1-schnell). They report that the classifier improves skintone precision by at least 16.04 percentage points over HEIM and VIT baselines, and that most evaluated models fail the four-fifth rule, with representation bias generally larger than alignment error and skintone alignment error substantially larger than gender error.","tokens_in":12606,"tokens_out":8093,"duration_ms":74249,"significance":"If the measurement assumptions hold, the paper would provide a useful reusable fairness benchmark and a convincing demonstration that pixel-only skintone detection is insufficient. The large-scale comparison across nine models and six domains, the explicit treatment of the Monk scale's uneven color distribution, and the inclusion of both representation and alignment metrics are genuine strengths. The classifier results on the public WBB dataset are encouraging. However, the headline fairness conclusions rest on an unvalidated generalization of the skintone classifier to the evaluated models, an untested uniform reference distribution, and an underspecified alignment-error mapping; these issues must be resolved before the empirical claims can be taken as established.","major_comments":[{"comment":"All skintone metrics for the nine evaluated models are computed using the INFELM classifier, but the classifier is never validated on images produced by these models. The only evaluations are on WBB and High-Aes (Table 6), and the topology module is trained on synthetic faces generated by RealisticVision v5.1 (Section 4.2.1), which is itself one of the nine evaluated models. Since style shifts or group-correlated misclassification (e.g., adjacent Monk-scale confusion under stylized lighting) would change every value of b_s and e_s in Table 7, the headline conclusions in the Abstract and Takeaway 1 are not supported without per-model human-annotated validation or an explicit style-robustness analysis.","section":"Section 5.2, Table 7"},{"comment":"The representation bias metric defines the reference distribution by the statement 'we assume that the demographic groups are equally distributed.' No justification or sensitivity analysis is provided, even though the text immediately acknowledges that social statistics from authoritative sources could serve as the groundtruth. Because the four-fifth-rule checks in Section 5.3 compare every model to this uniform baseline, all bias scores in Table 7 and the conclusion that 'most models do not meet the criteria of fairness' are conditional on an untested normative choice. The authors should either test alternative reference distributions or explicitly reframe the results as an equal-representation benchmark rather than a fairness verdict.","section":"Section 4.3.2, Eq. (4)"},{"comment":"The mismatch ratio used to compute e_g and e_s is not defined formally, and no mapping is given from the demographic phrases in the prompts (e.g., 'a black CEO') to the 10 Monk-scale labels produced by the classifier. Without this mapping, the large skintone alignment errors in Table 7 are uninterpretable; the claim that skintone alignment error is 'significantly higher' than gender error (Takeaway 2) could be an artifact of an inconsistent or overly strict mapping rather than a property of the generators. The paper should specify the mapping, give examples, and ideally validate the alignment labels by human raters.","section":"Section 4.3.2, alignment error definition"},{"comment":"The High-Aes dataset is described only as an internal synthetic dataset of 11,000 images with human annotations. It is not stated whether these images are generated with the same RealisticVision v5.1 pipeline used to create the topology training data in Section 4.2.1, nor are class balance, annotation protocol, or the test split described. Without this information, the reported 0.9033 precision on High-Aes and the claimed 'at least 16.04%' improvement over HEIM cannot be independently assessed, and the comparison may not be a fair out-of-distribution test.","section":"Section 5.1, Table 6"},{"comment":"All model comparisons are point estimates with no confidence intervals, standard errors, or significance tests, despite the finite sample sizes (100 images per prompt, 246 prompts). For example, differences as small as 0.001 in e_g between SDXL Lightning and FLUX.1-schnell are treated as meaningful without uncertainty quantification. This undermines the comparative claims in Takeaway 4; bootstrap confidence intervals or per-prompt standard errors should be reported.","section":"Section 5.2, Table 7"}],"minor_comments":[{"comment":"The typos 'evaluaiton' and 'exsting' should be corrected to 'evaluation' and 'existing'; the Related Work section should be copy-edited throughout.","section":"Section 2"},{"comment":"'Ordinary classification problem' should read 'ordinal classification problem,' and the sentence explaining the relationship between tolerance and ordinal classification should be rewritten for clarity.","section":"Section 4.2.3"},{"comment":"The sentence 'compare them with the fairness reference computed using the empirical four-fifth rule, .' contains an empty clause and appears to be missing a phrase or cross-reference to Figure 6.","section":"Section 5.3"},{"comment":"The symbol alpha is described as 'the attention weights' but the equation uses it as a scalar balancing coefficient; clarify whether alpha is a learned attention vector or a fixed scalar loss weight.","section":"Equation (3)"},{"comment":"'Off alltext-to-image models' should read 'Of all text-to-image models'; the text also alternates between 'DALL-E 3' and 'Dall-E 3' and should be made consistent.","section":"Section 5.2"},{"comment":"The paper does not provide a link to code, the full prompt set, or generated-image samples; releasing these artifacts would substantially improve reproducibility.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"One point for the editor's attention: one of the evaluated models, SDXL Lightning, is developed by ByteDance, which is affiliated with the authors' employer (TikTok Inc.); the manuscript does not disclose this relationship. In addition, the High-Aes dataset and the trained classifier appear to be proprietary, so the public-benchmark claim is hard to verify independently. These points do not change my technical assessment but may be relevant to transparency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the skintone classifier is the real contribution, and it looks legit, but every fairness number in Table 7 depends on that classifier's accuracy on images from the nine models, and that accuracy is never measured. The stress-test note is right: this is an unexamined empirical claim, not a normative quibble.\n\nWhat's new and what works: the classifier combines facial topology with dominant skin-pixel distributions and beats HEIM and VIT on WBB and High-Aes by large margins (16% precision, 4% MSE). That is credible evidence that the method is better than pixel-only baselines. The prompt set across six bias-sensitive domains and the explicit separation between representation bias and content alignment go beyond HEIM and DALL-Eval. The large-scale evaluation of nine models is also useful, and the finding that fine-tuned RealisticVision models are more biased than base Stable Diffusion is a plausible result with practical implications.\n\nWhere it gets soft: the biggest problem is the lack of validation on the actual generated images. The topology module is trained on RealisticVision v5.1, which is one of the nine evaluated models, so there is a direct circularity risk. Validation on WBB helps, but WBB contains real photos, not stylized FLUX or DALL-E outputs. If the classifier systematically confuses adjacent Monk scales or darkens lighter skins under non-photorealistic rendering, the headline claim that skintone alignment error is roughly twenty times gender error could be an artifact of classifier miscalibration. The paper needs per-model accuracy on a human-labeled subset, or at least a calibration plot across image styles. Without that, the absolute bias numbers in Table 7 are conditional on the classifier being right.\n\nThe uniform reference distribution is a second issue, though less damaging because the paper states it explicitly. Using equal group sizes is a defensible normative baseline, but the results should be labeled as \"under a uniform reference\" and tested against alternatives from social statistics. Minor problems: no error bars on any of the point estimates, no released artifacts, and the gender classifier is also unvalidated on generated images.\n\nWho this is for: people auditing text-to-image models in industry, and researchers working on skintone classification. The classifier alone is worth a serious look. My recommendation: send it to peer review. It is a solid, useful benchmark with a real technical contribution, but revision needs to add human validation on the nine models' outputs and sensitivity checks on the uniform assumption. If the authors cannot validate the classifier on model outputs, the fairness conclusions should be scaled back.","headline":"The skintone classifier is a genuine step forward, but every headline fairness number in Table 7 depends on that classifier's unmeasured accuracy on the very models being audited.","tokens_in":13149,"tokens_out":3535,"would_cite":true,"duration_ms":36946,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"INFELM claims a facial-topology skintone classifier exposes that nine text-to-image models fail a common fairness rule when bias is measured accurately.","keywords":["fairness evaluation","text-to-image models","skintone classification","representation bias","content alignment","four-fifth rule","facial topology"],"falsifier":"Take the same 246 prompts and generated images, but compute representation bias using real-world social statistics for each domain (e.g., occupation gender shares from labor statistics, skintone proportions from census-like data) instead of uniform p_g; if the four-fifth rule failures disappear or reorder, the audit's conclusion depends on that uniform assumption rather than on the models' outputs.","tokens_in":12099,"feed_emoji":"⚖️","tokens_out":4552,"duration_ms":40551,"temperature":0.7,"pith_summary":"This paper argues that existing text-to-image fairness audits are unreliable because their skintone classifiers rely on detecting dominant pixel colors and mapping to the nearest Monk scale, which misclassifies images due to uneven color spaces and lighting. It proposes INFELM, a fairness evaluation framework with a new skintone classifier that combines facial topological features with a distribution of dominant skin pixels, and reports that this classifier beats pixel-only baselines by at least 16.04% in precision. Using this measurement on nine text-to-image models across six bias-sensitive domains, the paper finds that most models fail an empirical four-fifth rule fairness criterion, with representation bias more pronounced than content alignment errors and skintone alignment errors substantially higher than gender alignment errors. The upshot is a reusable benchmark and a warning that current models are not meeting basic demographic fairness standards when measured more accurately.","feed_headline":"Most text-to-image models fail a four-fifth rule fairness audit","feed_subtitle":"Finer skintone measurement shows representation bias and skintone alignment errors dwarf gender errors in nine models.","key_machinery":"The central mechanism is the INFELM skintone classifier. It first trains a small CNN on 12,000 synthetically generated faces (using RealisticVision v5.1 across six geographic-origin groups, with varying environment lighting) to extract latent facial topological features, then extracts the K = 15 most dominant skin pixels in the facial region via Otsu's thresholding in YCbCr and HSV color spaces, representing them as a bin-weighted distribution over Monk scales. These two feature types are fused through a self-attention module and trained with a loss L = alpha*L_ft + (1-alpha)*L_st as an ordinal classification problem with ±1 scale tolerance. Fairness is then scored with representation bias b = (1/Z)*sum_g |n_g/N - p_g| normalized to the maximum possible bias, and content alignment error as the prompt-classifier mismatch ratio, both evaluated against a uniform reference distribution and a four-fifth rule threshold of 0.2.","core_discovery":"The central discovery is that a skintone classifier which fuses learned facial topological features with an ordinal distribution of dominant skin pixels (rather than a single mean color) reduces skintone misclassification, and that this measurement change affects the assessment of large text-to-image models: none of the nine models evaluated consistently satisfies the four-fifth rule threshold b < 0.2 for representation bias. Across all models, average gender representation bias is 0.639 and skintone bias 0.593, while average gender alignment error is only 0.029 versus 0.668 for skintone, so the dominant failure is inaccurate skintone alignment rather than gender misalignment. The paper also finds that fine-tuned photorealistic models (RealisticVision variants) show more polarized demographics than their base Stable Diffusion models, and that DALL-E 3 is closest to the fairness criteria but still fails on skintone.","pith_inferences":["The uniform reference distribution is a policy choice, not a fact; the same audit pipeline would yield different bias scores under a reference distribution drawn from population statistics, so the numerical thresholds should be read as relative rather than absolute.","The claim that fine-tuning amplifies bias is based on only two RealisticVision variants; a broader sample of fine-tuned models would be needed to confirm the trend.","The synthetic training images for the topology classifier come from a single text-to-image model (RealisticVision v5.1); if that generator has its own style bias, the latent topological features may inherit it, which is testable by retraining on a different generator.","The content alignment error metric counts any mismatch between prompt and classifier as an error, but some mismatches may be legitimate (e.g., a prompt with no demographic attribute); refining the metric could change the error magnitudes."],"forward_implications":["If the classifier is adopted, downstream fairness audits become more trustworthy because skintone measurements are no longer skewed by uneven Monk-scale color spans or lighting.","Most evaluated models fail the four-fifth rule; practitioners should not assume a single model is fair across both gender and skintone.","Fine-tuned photorealistic models tend to show more polarized demographic outputs, suggesting fine-tuning on curated datasets can amplify bias.","Skintone alignment errors exceed gender errors by a large margin, pointing to missing or noisy skintone captions in training data as a lever for improvement."],"supporting_citations":[{"why":"Supplies the main pixel-based skintone baseline (HEIM) and the evaluation framework that INFELM extends.","marker":"[12]"},{"why":"Defines the 10-point Monk skintone scale used for classification and for labeling the internal dataset.","marker":"[17]"},{"why":"The text-to-image model used to generate the synthetic facial images for training the topology classifier.","marker":"[3]"},{"why":"Another pixel-based fairness evaluation baseline (DALL-Eval) that accounts for illumination, compared in Table 1.","marker":"[7]"},{"why":"Stable Diffusion models form the base family for several evaluated models and are the core generation pipeline for comparison.","marker":"[22]"},{"why":"DALL-E 3 is the best-performing evaluated model and the closest to the fairness criteria in the experiments.","marker":"[18]"},{"why":"The VIT-based gender classification model used to label gender expression in generated images.","marker":"[1]"}],"fun_headline_variants":["AI image generators fail four-fifth rule fairness audit","Skintone bias, not gender, is the big fairness gap in AI images","None of nine text-to-image models meet fairness standards","New fairness benchmark INFELM shows AI image models are biased","DALL-E 3 still fails on skintone fairness in INFELM audit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All bias scores assume demographic groups should be equally represented (p_g uniform); if the true reference population is non-uniform, every bias number changes.","fun_headline_variants_meta":{"raw":{"variants":["AI image generators fail four-fifth rule fairness audit","Skintone bias, not gender, is the big fairness gap in AI images","None of nine text-to-image models meet fairness standards","New fairness benchmark INFELM shows AI image models are biased","DALL-E 3 still fails on skintone fairness in INFELM audit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1664,"prompt_tokens":1002,"completion_tokens":662,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":572}},"tokens_in":618,"tokens_out":662,"duration_ms":6664,"temperature":1.0,"reasoning_tokens":572,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:42:54.460619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 246 prompts and generated images, but compute representation bias using real-world social statistics for each domain (e.g., occupation gender shares from labor statistics, skintone proportions from census-like data) instead of uniform p_g; if the four-fifth rule failures disappear or reorder, the audit's conclusion depends on that uniform assumption rather than on the models' outputs.","supporting_citations":[{"cited_title":"Holistic evaluation of text-to-image models","cited_arxiv_id":null,"evidence_quote":"Supplies the main pixel-based skintone baseline (HEIM) and the evaluation framework that INFELM extends."},{"cited_title":"The monk skin tone scale","cited_arxiv_id":null,"evidence_quote":"Defines the 10-point Monk skintone scale used for classification and for labeling the internal dataset."},{"cited_title":"https://huggingface.co/SG161222/Realistic_ Vision_V5.1_noVAE","cited_arxiv_id":null,"evidence_quote":"The text-to-image model used to generate the synthetic facial images for training the topology classifier."},{"cited_title":"Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models","cited_arxiv_id":null,"evidence_quote":"Another pixel-based fairness evaluation baseline (DALL-Eval) that accounts for illumination, compared in Table 1."},{"cited_title":"High-resolution image synthesis with latent diffu- sion models","cited_arxiv_id":null,"evidence_quote":"Stable Diffusion models form the base family for several evaluated models and are the core generation pipeline for comparison."},{"cited_title":"Dall-e 3: The latest in text-to-image generation","cited_arxiv_id":null,"evidence_quote":"DALL-E 3 is the best-performing evaluated model and the closest to the fairness criteria in the experiments."},{"cited_title":"https://huggingface.co/touchtech/fashion-images- gender-age-vit-large-patch16-224-in21k-v3","cited_arxiv_id":null,"evidence_quote":"The VIT-based gender classification model used to label gender expression in generated images."}],"review_version":1}