{"id":"7f146798-4ccf-4f61-9a62-ab0cfdbc588c","arxiv_id":"2602.00032","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Emotion words in text-to-image prompts act as demographic selectors: negative emotions shift outputs toward White, middle-aged, male-coded faces, and young Black women are nearly absent across all models.","lead":"Researchers made 56,000 AI-generated faces with and without emotion words like 'angry' or 'happy' and measured the gender, race, age, and attractiveness of each face. They found that negative emotion prompts push models toward White, middle-aged, male-looking faces, while young Black women are almost never generated.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FairFace's emotion-invariance on synthetic faces is unvalidated; the central emotion-driven demographic shifts in §4.3 may be classifier artifacts rather than generator behavior.","rationale":"The reader's weakest_assumption correctly identifies the most load-bearing concern: the entire emotion-shift analysis depends on FairFace labels, yet the paper does not validate that FairFace's demographic estimates are invariant to facial expression. If this assumption fails, the central claim—that emotion prompts shift demographic distributions—could be an artifact of the classifier rather than a property of the generators. The existing validation in Appendix A.2 only checks synthetic faces under neutral prompts, so it does not rule out the confound. This concern is more fundamental than the absence of error bars or significance tests, because a systematic classifier bias could reverse the direction of the claimed effect; sampling noise would only add uncertainty. I therefore agree with the reader that the paper requires additional validation before the central claim can be fully accepted. Since the reader's verdict is already CONDITIONAL, my assessment does not change the recommendation; the paper should be revised to address this specific confound, along with the uncertainty quantification and artifact release also noted by the reader.","tokens_in":21285,"tokens_out":5566,"duration_ms":58382,"concrete_test":"Extend the Appendix A.2 validation to emotion-conditioned prompts. For each demographic specification (e.g., 'young Asian female', 'middle-aged Black male', etc.) and each of the six emotion words, generate 100 synthetic faces per audited model (or at least 2 Western + 2 Chinese models), then apply FairFace. Compute the demographic label distributions per emotion for each fixed demographic prompt. If, for a fixed specified demographic, FairFace's predicted gender/race/age distributions shift significantly across emotion conditions (e.g., >5 percentage points), the emotion-driven shifts in §4.3 are confounded by classifier expression sensitivity. If the labels are stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that negatively valenced emotion prompts shift generated faces toward White, male, and older demographics (abstract, §5.3)—rests entirely on demographic labels produced by FairFace (Section 3.4). The validation in Appendix A.2 tests FairFace on synthetic faces generated with explicit demographic prompts (e.g., 'young Asian female') but omits emotion variation: the prompts are neutral. The benchmark validation in Appendix A.1 reports overall accuracy on CFD, FACES, APPA-REAL, and FGNET but does not stratify by facial expression. If FairFace's age, gender, or race estimates are expression-dependent—e.g., anger increasing perceived age or masculinity—the emotion-induced distribution shifts in §4.3 and Figure 5 would reflect classifier bias rather than generator behavior. The paper's conclusion acknowledges 'classification errors' generically but does not address this specific confound. This is not a minor caveat: it undermines the causal interpretation that 'emotion prompts act as demographic selectors.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper audits eight text-to-image (T2I) models — four from Western organizations and four from Chinese institutions — by generating 56,000 faces under a neutral prompt and six emotion prompts. Demographic attributes (gender, race, age) are estimated with FairFace and attractiveness with a dedicated model. The paper quantifies bias relative to UN-derived global population statistics and, for emotion effects, relative to each model's neutral-prompt baseline, using KL, JS, and TVD divergences at the marginal and intersectional level. The central claim is that adding an emotion to a prompt does more than change facial expression: negatively valenced emotions (sadness, anger, fear, disgust) consistently shift generated faces toward White, middle-aged, male-coded appearances, while happiness yields the smallest shift from the neutral baseline. The authors also report a broad Western/Chinese homogenization of demographic bias and propose that emotion-conditioned, intersectional, and multilingual audits become standard practice.","tokens_in":21485,"tokens_out":5278,"duration_ms":61831,"significance":"If the central claim holds, this is an important and non-obvious contribution: emotional prompts would act as demographic selectors in T2I models, a dimension that prior audits, focusing on neutral prompts, have missed. The study is strong in scale and design—eight models, 56,000 images, separate validation of FairFace on benchmark and on controlled synthetic faces, and a multi-metric information-theoretic framework. The explicit comparison of Western and Chinese model families is also valuable. However, the main causal interpretation currently rests on an unvalidated assumption—that FairFace's demographic estimates are invariant to facial expression—and the headline quantitative claims lack uncertainty quantification. These issues are addressable with additional validation and re-analysis, but they are load-bearing for the paper's central assertion.","major_comments":[{"comment":"The central emotion-shift claim rests entirely on FairFace demographic labels, but the validation in Appendix A.2 tests synthetic faces generated with explicit demographic prompts and neutral expression only; the benchmark validation in Appendix A.1 does not stratify accuracy by facial expression. If FairFace's age, gender, or race predictions are expression-dependent—e.g., angry faces being classified as older or more male-coded—then the emotion-induced shifts reported in §4.3 and Figure 5 would be classifier artifacts rather than generator behavior. The limitation paragraph acknowledges 'classification errors' generically but does not address this specific differential-bias confound. The paper should either demonstrate emotion-invariance on expression-labelled benchmark or synthetic sets (e.g., stratifying FACES/CFD by expression, or generating emotion-conditioned faces with explicit d","section":"§3.4, §A.2, §4.3"},{"comment":"The phrase 'statistically significant' is used (e.g., §5.3, bullet 3) to describe the emotion-driven increase in male and White faces, but no significance test, confidence interval, or bootstrap procedure is reported anywhere. Tables 1 and 2 report scalar KL/JS/TVD values without uncertainty, despite each condition being only 1,000 samples. The directional claims ('consistently shift,' 'statistically significant') require at least bootstrap CIs or permutation tests over the per-model/per-emotion distributions, particularly because several reported differences are small in absolute magnitude.","section":"§5.3 and §4.3"},{"comment":"The intersectional KL/JS metrics are computed over 2 × 4 × 3 = 24 cells with only 1,000 samples per model and emotion. The paper does not state how zero cells are handled; if P(d|e0)=0 while P(d|e)>0, the KL term in Eq. (4) is undefined, and with near-erasure of combinations such as young×female×Black, such zeros are plausible. Without explicit smoothing (e.g., additive epsilon) or reporting of cell counts, the values in Table 2 — and the ranking of emotions by intersectional shift — may be dominated by sparse-cell noise. The authors should report the smoothing procedure or use a sparse-robust alternative (e.g., smoothed JSD).","section":"§4.3.1, Table 2, Eqs. (4)–(5)"}],"minor_comments":[{"comment":"The reference [Huber et al.(2023)] contains a placeholder DOI/arXiv id ('arXiv:2304.XXXX') and cannot be verified; [Doh et al.([n. d.])] lacks a year and a stable publication venue. Please complete these citations before publication.","section":"References"},{"comment":"The text says 'The joint demographic distributions reported in Table 2 further suggest...' but Table 2 reports emotion-conditioned KL/JS divergences, not Western/Chinese joint demographic distributions. The intended pointer is probably Table 1 or Figure 3.","section":"§4.2"},{"comment":"The heatmaps show 'mean ΔP' of facial attributes, but the color scale and the precise definition of ΔP (e.g., P(emotion) − P(neutral) averaged over models) are not stated. Please add a legend and explicit definition.","section":"Figure 5"},{"comment":"Equation (7) is not a KL divergence but a pointwise weighted log-ratio for a single category c. Renaming it (e.g., 'per-category shift contribution') would avoid confusion with the proper divergences in Eqs. (1)–(5).","section":"Eq. (7)"},{"comment":"The global 'White ≈ 11%' reference value is a constructed quantity from a country-to-race assignment. The construction is described in Appendix C, but the paper could more explicitly flag that this reference is a modeling choice and show sensitivity to alternative assignments, since RQ1's absolute claims depend on it.","section":"§3.5 / Appendix C"},{"comment":"The paper does not mention releasing code, prompts, or generated image data. Given the audit's value as a benchmark, sharing the exact prompt templates, generation seeds (if deterministic), and the evaluation pipeline would substantially strengthen reproducibility.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful part of this paper is the scope: eight models, 56,000 faces, neutral plus six emotions, intersectional analysis, and a Western-vs-Chinese comparison. If you work on T2I fairness, this is the first audit that systematically combines emotion conditioning with intersectional demographics. The finding that negative emotions shift outputs toward White, middle-aged, male-coded faces is new and, if it holds, important. It has a load-bearing soft spot, though: everything rests on FairFace labels, and the paper never validates that FairFace's demographic predictions are invariant to facial expression on synthetic faces. The validation in Appendix A.2 covers only neutral prompts with explicit demographics. Anger, fear, and disgust plausibly make faces look older or read as more masculine, so part of the emotion shift could be classifier artifact rather than generator behavior. That's not a minor caveat.\n\nWhat the paper does well: the design is sound, the benchmark validation is careful (CFD, FACES, APPA-REAL, FGNET), and the synthetic-face validation with explicit demographic prompts is more than most audits do. The intersectional near-erasure of young x female x Black faces is clear and important regardless of the emotion question. The Chinese-prompt appendix and the sad-vs-unhappy robustness check are thoughtful additions.\n\nSoft spots beyond the confound: no error bars or significance tests, yet Section 5.3 calls the shifts 'statistically significant.' With 1,000 images per cell, bootstrap CIs are cheap. The intersectional KL/JS values are computed over sparse cells without smoothing, which can inflate divergences when the neutral baseline has near-zero counts. And no code or data is released, which is a real problem for an audit paper.\n\nVerdict: send it to peer review. It deserves referee time. The authors should be asked to add the emotion-conditioned classifier validation, uncertainty metrics, and public artifacts. I would cite the neutral-prompt intersectional findings with moderate confidence now; I'd hold off citing the emotion-shift claim until the confound is addressed.","headline":"First cross-ecosystem, emotion-conditioned audit of face-generation bias; solid scope, but the headline emotion-shift claim may partly be a classifier artifact.","tokens_in":21950,"tokens_out":3121,"would_cite":true,"duration_ms":34154,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding an emotion to a text-to-image prompt shifts the demographics of the generated faces — negative emotions bias outputs toward White, middle-aged, male-coded faces.","keywords":["text-to-image models","demographic bias","emotion conditioning","intersectionality","synthetic faces","fairness audit","KL divergence","cross-cultural comparison"],"falsifier":"Take a set of neutral faces generated by the same models, digitally re-render each with different expressions (happy, angry, sad) while preserving identity, and check whether the automated classifier's demographic predictions shift toward White, middle-aged, male for negative expressions. If they do, the emotion-driven demographic shift is at least partly an artifact of the measurement tool. Alternatively, have human annotators label a random sample of the emotion-conditioned faces and see whether human-perceived demographics follow the same valence-driven pattern.","tokens_in":21175,"feed_emoji":"😠","tokens_out":5114,"duration_ms":45542,"temperature":0.7,"pith_summary":"This paper tries to establish that emotion words in text-to-image prompts do not merely change a face's expression; they quietly change who is depicted. Auditing eight text-to-image models (four from Western developers, four from Chinese developers) with 56,000 generated faces, the authors find that negatively valenced emotions (sadness, anger, fear, disgust) consistently shift outputs toward White-coded, middle-aged, male-coded faces, while happiness produces the smallest demographic shift from the neutral baseline. The paper also documents strong overrepresentation of young faces and near-erasure of specific intersections such as young female Black faces across all models. If correct, emotion conditioning is a demographic selector in these systems, and audits that only use neutral prompts systematically miss a dimension of representational bias.","feed_headline":"Audit: sad and angry prompts skew AI faces White and male","feed_subtitle":"Across 8 Western and Chinese models, negative emotions shift generated faces toward older White men; happiness barely moves the default.","key_machinery":"The argument rests on an emotion-conditioned prompt template — 'A photorealistic portrait of a [emotion] person, front-facing' — combined with information-theoretic divergences (Kullback-Leibler, Jensen-Shannon, total variation distance) that compare each emotion's output distribution against the neutral baseline and against global population statistics. The intersectional analysis further computes the joint distribution over gender, race, and age, exposing compounded underrepresentation that single-attribute metrics miss. The demographic labels come from an automated face-attribute classifier, and perceived attractiveness from a separate model.","core_discovery":"On the paper's own terms, the central discovery is a valence-driven demographic mapping: instructing a model to generate a sad, angry, fearful, or disgusted face moves the output distribution toward White, middle-aged, male appearance, relative to the model's neutral default, and away from Asian and young faces. Happy prompts barely move the distribution, indicating that the neutral default already sits close to the 'happy young woman' prototype in latent space. The direction of these shifts is consistent across all eight models, Western and Chinese alike, and is accompanied by a drop in perceived attractiveness for negative emotions.","pith_inferences":["If the valence-driven demographic shift is real, it suggests the model's latent space ties emotional valence to demographic prototypes (grumpy = old White man, happy = young woman); a direct probe would be to generate 'an angry young Asian woman' and measure whether the model resists the combination.","The result could be partly an artifact of the attribute classifier: if the classifier reads angry expressions as older and male, the emotion-shift findings would be inflated. Testing same-identity faces with altered expressions would settle this.","The cross-ecosystem homogenization implies that mitigation efforts on one ecosystem (e.g., dataset documentation for one model family) may propagate globally, and that regional regulation could have outsized effects.","A practical extension: explicitly grounding demographic terms in emotion prompts (e.g., 'an angry Black woman') could counteract the default shift, providing a user-level mitigation — but it also risks over-correcting or stereotyping."],"forward_implications":["Emotion prompts are demographic selectors: content creators who routinely generate 'angry' or 'fearful' imagery will, without intending it, populate their visuals disproportionately with White middle-aged men.","Bias audits that test only neutral prompts understate the representational skew of a model; emotion-conditioned auditing should be part of pre-deployment evaluation.","Western and Chinese models converge on the same biases, implying shared training data and evaluation standards; regional origin alone does not diversify outputs.","Negative emotions lower perceived attractiveness of generated faces, reinforcing an association between negative affect and unattractiveness that mirrors the attractiveness halo effect.","Intersectional near-erasure (e.g., young female Black faces) shows that fixing single-attribute parity would not fix representational diversity."],"fun_headline_variants":["AI faces go White and male when told to be sad or angry","Negative emotions turn AI faces White and male, study finds","Emotion prompts bias AI faces: sad means White male","Sad, angry prompts shift AI faces to White, middle-aged men"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The automated classifier's demographic labels (gender, race, age) are accurate on synthetic faces and are not swayed by the facial expression in the image; if angry faces are systematically misread as older and male, the paper's central emotion-shift finding could be a classifier artifact rather than a generator behavior.","fun_headline_variants_meta":{"raw":{"variants":["AI faces go White and male when told to be sad or angry","Negative emotions turn AI faces White and male, study finds","Emotion prompts bias AI faces: sad means White male","Sad, angry prompts shift AI faces to White, middle-aged men"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000727,"raw_usage":{"total_tokens":3128,"prompt_tokens":810,"completion_tokens":2318,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2247}},"tokens_in":554,"tokens_out":2318,"duration_ms":16591,"temperature":1.0,"reasoning_tokens":2247,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:47:23.614155+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of neutral faces generated by the same models, digitally re-render each with different expressions (happy, angry, sad) while preserving identity, and check whether the automated classifier's demographic predictions shift toward White, middle-aged, male for negative expressions. If they do, the emotion-driven demographic shift is at least partly an artifact of the measurement tool. Alternatively, have human annotators label a random sample of the emotion-conditioned faces and see whether human-perceived demographics follow the same valence-driven pattern.","supporting_citations":[],"review_version":1}