{"id":"68cbb657-5134-4d91-82e8-1165d73cc0fe","arxiv_id":"2507.10755","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"An audit of AffectNet and RAF-DB finds many posed images and reports that two FER models disproportionately predict negative emotions for non-white and darker-skinned smiling faces, but the bias evidence lacks controls and validation.","lead":"Two popular facial-expression datasets contain many posed photos, despite being described as in-the-wild. Two emotion-recognition models also more often predict negative emotions for non-white and darker-skinned faces that are smiling, though the evidence for bias is weakened by missing baselines and unvalidated labels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bias claim rests on unverified assumption: without race-specific smile base rates and validated smile/neutral labels, higher P(smiling | negative prediction) for non-white groups does not establish model bias; base-rate confounding alone can produce the reported pattern.","rationale":"The reader's weakest_assumption identifies the same load-bearing gap: the bias conclusion assumes that smiling and neutral faces are correct ground-truth labels and that group differences in the proportion of smiling faces among negative predictions reflect model bias rather than base-rate differences. My stress-test agrees. The paper's reported chi-square tests compare each non-white group against the white group's proportion of smiling faces in the negative prediction set; for this to test bias, smile prevalence in FairFace must be race-invariant, which is never established. The manual smile/neutral labeling is also unvalidated and potentially biased by knowledge of the hypothesis. Both issues are internal to the argument rather than merely disagreements with consensus, so the REJECT verdict stands. The posed-expression proportion is also a concern, but the racial-bias claim is the more consequential and load-bearing assertion in the abstract and conclusion. No additional independent support (released code, validated labels, inter-rater reliability) is present to offset these methodological gaps.","tokens_in":11352,"tokens_out":3111,"duration_ms":42855,"concrete_test":"Compute race-specific smile base rates on a gold-standard annotated subset: take 300 random FairFace images per observed race (2,100 total), have at least two independent coders label smile/neutral (smile defined as AU12+AU25; report Cohen's kappa), then re-run the model analysis using P(negative prediction | smile, race) and P(negative prediction | non-smile, race), or equivalently adjust P(smile | negative prediction, race) by Bayes' rule using P(smile | race). If the adjusted error rates are statistically equal across races, or if the reported pattern disappears after conditioning on smile prevalence, the Section IV-B bias conclusion does not hold. As a secondary check, recompute the paper's chi-square tests using a null expectation of P(smile | race) rather than the white proportion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central bias claim in Section IV-B compares, across observed races, the proportion of smiling faces among images that the FER models labeled with a negative emotion. The paper takes the white proportion as the expected value under the null hypothesis of no bias (Section IV-B, Tables III-IV). This is valid only if the prevalence of smiling in FairFace is identical across observed races. FairFace is a naturalistic face dataset; no evidence is given that smile prevalence is race-balanced. If Black or Southeast Asian faces are more often smiling in FairFace, then even a race-blind model that cannot read expression would show more smiling faces in its negative prediction bins for those groups. The missing control is P(smile | race) over the full evaluation set, or preferably the error rates P(negative prediction | smile, race) and P(negative prediction | non-smile, race), rather than the reverse conditional used in the paper. The smile/neutral labels are also assigned by the authors with no inter-rater reliability, no FACS verification, and no reported blinding to model prediction or observed race; because the hypothesis is known, this creates systematic mislabeling risk. The posed-expression result (Section IV-A) is similarly built on an unvalidated heuristic (AU6 absence, plain backgrounds, actors, lighting), but the headline bias claim already fails on the base-rate issue alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript audits two facial expression recognition (FER) datasets, AffectNet and RAF-DB, and two models trained on them (RUL and MENet). It makes two central claims: first, that a large fraction of supposedly in-the-wild images in these datasets are actually posed (46.5% in AffectNet and 35.3% in RAF-DB, based on a custom visual heuristic); second, that the audited FER models are racially biased because, when the models predict a negative emotion, the proportion of smiling or neutral faces is higher for people observed as non-white or with darker skin than for people observed as white. The paper proposes this as the first audit to verify the presence of posed expressions in FER datasets and to examine fairness across more than two observed-race groups.","tokens_in":11610,"tokens_out":6352,"duration_ms":78223,"significance":"If the claims were established, this would be a valuable contribution to the algorithmic-auditing literature: the topic is timely, the use of FairFace with multiple observed-race categories is a step beyond binary race audits, and the paper makes falsifiable, quantitative claims rather than only qualitative arguments. The paper also usefully foregrounds the connection between data-collection practices and downstream model behavior. However, as presented, the evidence does not support the central claims. The bias analysis compares a reverse conditional—P(smiling | negative prediction, race)—without controlling for how often each group smiles in the evaluation set, so a race-blind model could produce the same pattern solely from group differences in smile base rates. The smile/neutral labels used as ground truth are unvalidated and appear to be assigned without inter-rater reliability or blinding. The posed-expression heuristic is similarly unvalidated and partly circular, and the reported statistical intervals are inconsistent with the stratified sampling design.","major_comments":[{"comment":"The bias analysis compares P(smiling | negative prediction, observed race) across groups and treats the white group's value as the no-bias null. This is valid only if smile prevalence in FairFace is equal across observed-race groups, which the paper does not establish and which is unlikely to hold in a naturalistic dataset. A race-blind classifier whose errors are independent of race would still produce the reported pattern if, for example, Black or Southeast Asian faces are more often smiling in FairFace, because more smiling faces would fall into any negative-prediction bin. The authors should report P(smile | race) on the full evaluation set and the conditional error rates P(negative prediction | smile, race) and P(negative prediction | non-smile, race) with confidence intervals; the reverse conditional alone cannot support the bias conclusion. The same concern applies to the neutral-face analysis and to the skin-tone analysis in Figure 3. Additionally, because inference outputs were sampled by predicted emotion and observed race, the reported marginal proportions are not representative of the models' behavior on FairFace as a whole.","section":"Section IV-B, Tables III–IV, Figures 2–3"},{"comment":"The smile and neutral labels used as ground truth for the bias claim are assigned by the authors with no inter-rater reliability, no FACS verification, and no reported blinding to model predictions or observed race. Because the audit hypothesis is known, this creates a risk of systematic mislabeling that could generate the reported differences. The paper also assumes, without support, that a smiling or neutral face cannot validly receive a negative emotion prediction; this ignores display rules, social context, and the possibility of genuine smiles accompanying negative affect. At minimum, the authors need an independent label validation (for example, multiple coders with agreement statistics) and a sensitivity analysis using only high-confidence labels.","section":"Section III-F"},{"comment":"The posed-expression methodology is not validated. The FACS-based smile criterion applies only to smiles; for non-smiling faces, the paper relies on visual heuristics (recognizable actors, plain monochrome backgrounds, direct gaze, artificial lighting) that are asserted without evidence or annotation reliability. Because these same heuristics partly define what the auditors count as posed, the claim that 46.5% of AffectNet and 35.3% of RAF-DB images are posed is to a large extent a restatement of the labeling rule rather than an empirical discovery. The null hypothesis of \"zero posed images in a completely wild setting\" is also not an appropriate benchmark for in-the-wild datasets, and the chi-square test against it does not validate the heuristic. The authors should validate their protocol on a dataset with known spontaneous and posed labels and report per-criterion precision and recall.","section":"Section III-B and Section IV-A"},{"comment":"The statistical reporting for the posed-image proportions is inconsistent with the sampling design. The sample size is initially computed under simple random sampling with a 95% confidence level and 5% margin of error, but the authors then state that they \"sample equal numbers of each emotion label,\" which is a stratified design. The confidence intervals and the combined chi-square test with N = 761 ignore this stratification and are therefore not valid as reported. Additionally, the caption of Figure 1 gives the combined proportion as 40.9% with a 95% CI of 41.3%–51.6%, which is impossible because the point estimate lies outside the interval; the interval appears to have been copied from the AffectNet row. The authors should compute weighted estimates and report strata-specific results.","section":"Section III-A and Figure 1"}],"minor_comments":[{"comment":"The manuscript contains frequent typos, including \"interations\" (Introduction), \"exppressions\" (Section II-B), \"Relative Uncertainity\" (Table I and Section IV-B), \"Southest Asian\" (Table II), and \"osbserved\" (Table II caption).","section":"Throughout"},{"comment":"The paper says it aimed to select 50 samples for each permutation of observed race and predicted emotion (Section III-E), but the row totals in Table II are not multiples of 50 for some groups; please clarify the actual sampling and inclusion criteria.","section":"Section III-E and Table II"},{"comment":"The bar charts report proportions without error bars or confidence intervals, so the reader cannot assess the precision of the between-group differences; adding these would help interpretation.","section":"Figures 2 and 3"},{"comment":"The multiple chi-square tests are not corrected for multiple comparisons; with 12 tests, some significant results would be expected by chance even under the null hypothesis.","section":"Tables III and IV"},{"comment":"The statement that \"only 4% of the sampled images were of the darkest skin tone\" is not tied to any reported count or table; please provide exact numbers and clarify which sample it refers to.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper addresses a timely and important topic, and the authors are right to push beyond binary race categories. However, the stress-test concern about base-rate confounding lands directly: the bias statistic in Section IV-B cannot distinguish model bias from group differences in smile prevalence, and this alone invalidates the paper's headline claim as currently stated. The posed-expression analysis also needs external validation before the proportions can be taken at face value. These are substantial but in principle fixable: the bias analysis could be redone with proper base-rate controls and conditional error rates, and the labeling protocol could be validated with independent coders. I therefore recommend a major revision rather than outright rejection, with the understanding that the revised paper must reanalyze the data, not merely soften the language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one before you read it. First, the central bias claim is built on a comparison that doesn't control for how often each observed-race group smiles in FairFace. The paper computes the share of smiling faces among negative predictions and treats the white share as the null expectation. That is only valid if smile prevalence is race-balanced, and no evidence of that is given. A race-blind model would produce exactly this pattern if Black or Southeast Asian faces in FairFace are more often smiling. Second, the posed-expression percentages are computed from a sample the authors stratified by emotion label, but they report confidence intervals as if the sample were a simple random sample. They never weight the strata back to the dataset's label distribution, so the 46.5% figure is not a valid estimate of anything in the population.\n\nThe paper does have real merits. The question of how many posed images are inside two widely used FER datasets is a good and under-studied one. The fairness audit spans seven observed-race categories, which is more than the binary audits that have dominated the literature. The authors are also transparent about the difficulty of defining race and about the limitations of their own labels. That honesty matters.\n\nThe soft spots are serious, though. The manual smile/neutral labels come with no inter-rater reliability, no blinding, and no independent check. The posed/spontaneous heuristic (plain backgrounds, actors, direct gaze, AU6 absence) is unvalidated and circular in places: the same cues used to infer posing are then used as evidence that the images are posed. The chi-square tests in Tables III and IV use white as the expected value without justifying that the model error rate should be race-invariant after conditioning on expression. None of this is fatal to the idea of the audit, but it is fatal to the current conclusions. The paper's own discussion of stereotypes in human emotion perception is interesting but does not do the statistical work.\n\nWho is this for? Fairness auditors and FER researchers might want to cite it as an example of the pitfalls of audit design, but the specific findings should not be used as evidence of model bias. I would send it to peer review because the topic is important and the flaws are correctable, but I would expect major revisions: add base-rate controls, report error rates per group given ground-truth labels, validate the manual annotations, and redo the sampling math. As it stands, I would not cite it as a reliable source.","headline":"The racial-bias claim doesn't survive contact with the missing base-rate control, and the posed-image percentages come from a stratified sample that isn't weighted back; the question is good, but the evidence as presented doesn't support the headline.","tokens_in":12108,"tokens_out":1602,"would_cite":false,"duration_ms":22732,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Auditing AffectNet and RAF-DB, this paper finds that a substantial share of supposedly in-the-wild images are posed, and that models trained on these datasets are more likely to predict negative emotions for smiling faces observed as…","keywords":["facial expression recognition","algorithmic audit","racial bias","posed expressions","spontaneous expressions","AffectNet","RAF-DB","FairFace"],"falsifier":"Re-run the bias audit on a subset of test images where the base rate of smiling and neutral faces is matched across observed-race groups, or where smile/neutral labels are independently verified by multiple annotators; if the gap in negative predictions on smiling faces between White and non-White groups disappears or reverses, the claimed bias pattern would be shown to reflect test-set demographics rather than systematic model bias.","tokens_in":11141,"feed_emoji":"🎭","tokens_out":8611,"duration_ms":91853,"temperature":0.7,"pith_summary":"This paper tries to establish that two widely used facial expression recognition (FER) datasets, AffectNet and RAF-DB, contain a substantial number of posed images even though they are presented as capturing spontaneous 'in-the-wild' expressions, and that models trained on them are biased against people perceived as non-white or as having darker skin. Using a proposed method for distinguishing posed from spontaneous smiles, the audit finds 46.5% of sampled AffectNet images and 35.3% of sampled RAF-DB images are posed. Running two top-performing models on the FairFace test set, the paper reports that faces observed as Black, East Asian, Southeast Asian, or Indian are significantly more likely than White faces to be predicted as angry, sad, disgusted, or contemptuous while smiling or neutral, with the effect increasing for darker skin tones. If true, these findings matter because real-world uses of emotion recognition—automated interviews, security, and hiring—could both lose accuracy on genuine expressions and perpetuate harmful misreadings of people's emotions based on appearance.","feed_headline":"Audit: FER models read smiling darker faces as angry","feed_subtitle":"Two flagship emotion datasets are ~40% posed; negative predictions skew against non-white smiles.","key_machinery":"The central object is the smile classifier based on the Facial Action Coding System (FACS): a genuine smile is indicated by Action Units 6 (cheek raising), 12 (lip corner pull), and 25 (parted lips), while a posed smile lacks AU6. For non-smiling images, the audit treats recognizable actors, plain monochrome backgrounds, and bright artificial lighting combined with direct camera gaze as markers of posing. On the fairness side, the load-bearing metric is the proportion of smiling or neutral faces inside the set of negative emotion predictions, stratified by observed race from the FairFace test set and by a six-point skin-tone scale grouped into three bins, with White faces serving as the reference category for chi-square tests. This metric carries the bias conclusion: if models were unbiased, the smiling rate among negative predictions should be the same across observed races.","core_discovery":"The paper's core assertion is that the audited FER systems are racially biased and their training data are not genuinely spontaneous. On the data side, the authors propose a FACS-inspired rule for smiles (presence of AU6/AU12/AU25 for genuine, absence of AU6 for posed) plus background and gaze heuristics for non-smiles, and apply it to random samples of AffectNet and RAF-DB, concluding that 46.5% of AffectNet and 35.3% of RAF-DB images are posed, a result statistically significant against a null hypothesis of zero posed images. On the model side, the paper runs the Relative Uncertainty Learning model and the Multi-task EfficientNet-B2 model on FairFace, then compares the rate of negative emotion predictions (anger, sadness, disgust, contempt) among faces the auditors label as smiling or neutral. Across both models, 23.4% of negative predictions on White faces are smiling, versus 33% on Black faces and higher on several other non-White groups, with significant chi-square differences for Black, East Asian, Southeast Asian, and Indian observed races; the neutral-face and skin-tone analyses show a similar or stronger pattern. The paper interprets these gaps as evidence that the models over-assign negative emotions to non-White and darker-skinned faces, arguing that a smiling or neutral face should not receive a negative label.","pith_inferences":["A stratified re-analysis that controls for how often each observed-race group smiles in FairFace could separate true model bias from base-rate effects, something the paper does not report.","Because a large share of training images are posed, models trained on these datasets may be tuned to exaggerated, actor-like expressions; evaluating on a corpus of genuinely spontaneous interactions would likely show a larger accuracy drop for non-White faces.","The posed-image detection method, built for smiles, could be extended to video FER datasets and to non-smile emotions, where FACS markers are less established, to test whether posing affects other expressions similarly.","A stronger causal test of skin-tone sensitivity would manipulate perceived skin tone in images while holding expression and identity fixed, isolating the model's reliance on skin cues from the demographic composition of training data."],"forward_implications":["Benchmark accuracy of these models on 'in-the-wild' data overstates real-world performance, because a large portion of their training images are posed rather than spontaneous expressions.","Deployed emotion-recognition systems risk misclassifying smiling or neutral faces of non-White and darker-skinned people as angry, sad, disgusted, or contemptuous, affecting automated interviews, security, and hiring.","Underrepresentation of dark skin tones in the training samples (about 4% in the darkest group) likely contributes to the bias, though the sample is too small to test annotation bias.","The proposed posed-image detection methodology can be reused by other auditors to check spontaneity in FER image datasets."],"supporting_citations":[{"why":"Supplies AffectNet, the first dataset audited for posed images and the training set for one of the biased models.","marker":"[34]"},{"why":"Supplies RAF-DB, the second dataset audited for posed images and the training set for the other biased model.","marker":"[30]"},{"why":"Provides FairFace, the observed-race and skin-tone test set used for the model bias audit.","marker":"[25]"},{"why":"Provides the FACS-based smile identification rule (AU6/AU12/AU25) used to label smiles as genuine or posed.","marker":"[5]"},{"why":"Documents the lack of an established methodology for distinguishing spontaneous from posed expressions, motivating the paper's proposed method.","marker":"[20]"},{"why":"Prior audit showing facial analysis performs worse on darker-skinned faces, motivating the fairness audit.","marker":"[7]"},{"why":"The Relative Uncertainty Learning model trained on RAF-DB whose predictions are audited.","marker":"[51]"},{"why":"The Multi-task EfficientNet-B2 model trained on AffectNet whose predictions are audited.","marker":"[42]"}],"fun_headline_variants":["Emotion AI misreads non-White smiles as anger","Audited emotion datasets are full of posed faces","Racial bias in emotion AI: smiles seen as anger","Emotion recognition models penalize darker skin","The hidden posed faces in emotion AI datasets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The bias claim hinges on treating a smiling or neutral face as a ground-truth label that cannot legitimately receive a negative emotion prediction, and on trusting the auditors' manual smile/neutral labels; if smiling simply occurs at different frequencies across observed-race groups in the test set, the reported gaps could be an artifact of those base-rate differences rather than model bias.","fun_headline_variants_meta":{"raw":{"variants":["Emotion AI misreads non-White smiles as anger","Audited emotion datasets are full of posed faces","Racial bias in emotion AI: smiles seen as anger","Emotion recognition models penalize darker skin","The hidden posed faces in emotion AI datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3335,"prompt_tokens":1084,"completion_tokens":2251,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":700,"completion_tokens_details":{"reasoning_tokens":2178}},"tokens_in":700,"tokens_out":2251,"duration_ms":19792,"temperature":1.0,"reasoning_tokens":2178,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:26:35.229391+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the bias audit on a subset of test images where the base rate of smiling and neutral faces is matched across observed-race groups, or where smile/neutral labels are independently verified by multiple annotators; if the gap in negative predictions on smiling faces between White and non-White groups disappears or reverses, the claimed bias pattern would be shown to reflect test-set demographics rather than systematic model bias.","supporting_citations":[{"cited_title":"AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild","cited_arxiv_id":"1708.03985","evidence_quote":"Supplies AffectNet, the first dataset audited for posed images and the training set for one of the biased models."},{"cited_title":"Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild","cited_arxiv_id":null,"evidence_quote":"Supplies RAF-DB, the second dataset audited for posed images and the training set for the other biased model."},{"cited_title":"FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age for Bias Measurement and Mitigation","cited_arxiv_id":null,"evidence_quote":"Provides FairFace, the observed-race and skin-tone test set used for the model bias audit."},{"cited_title":"Proximity Begins with a Smile, But Which One? Associating Non-duchenne Smiles with Higher Psy- chological Distance","cited_arxiv_id":null,"evidence_quote":"Provides the FACS-based smile identification rule (AU6/AU12/AU25) used to label smiles as genuine or posed."},{"cited_title":"J., AND LI, X","cited_arxiv_id":null,"evidence_quote":"Documents the lack of an established methodology for distinguishing spontaneous from posed expressions, motivating the paper's proposed method."},{"cited_title":"Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification","cited_arxiv_id":null,"evidence_quote":"Prior audit showing facial analysis performs worse on darker-skinned faces, motivating the fairness audit."},{"cited_title":"Relative Uncertainty Learning for Facial Expression Recognition","cited_arxiv_id":null,"evidence_quote":"The Relative Uncertainty Learning model trained on RAF-DB whose predictions are audited."},{"cited_title":"V., S AVCHENKO , L","cited_arxiv_id":null,"evidence_quote":"The Multi-task EfficientNet-B2 model trained on AffectNet whose predictions are audited."}],"review_version":1}