{"id":"5288f007-b955-4edc-a379-7765b8167a26","arxiv_id":"2504.16104","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Multimodal LLMs systematically answer forced-choice social questions differently for beautified versus original faces, assigning positive traits to the beautified versions in most scenarios.","lead":"Seven open-source multimodal AI models changed their judgments on the same faces in 84 to 86 percent of tested social scenarios when a beauty filter was applied. The authors report that, like humans, these models attach positive traits such as confidence and trustworthiness to more attractive faces, and that the effect interacts with gender, age, and race.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing risk is the filter-confound assumption: beauty filters alter low-level visual features beyond attractiveness, so the reported shifts may partly reflect filter artifacts rather than attractiveness perception.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing premise: that the beauty filter changes only perceived attractiveness. My reading of the paper confirms this is the key causal assumption. The paper relies on human attractiveness ratings to validate the manipulation, but those ratings cannot certify that low-level visual features are invariant; they only certify that the filter increases perceived attractiveness. Since MLLMs are known to exploit superficial image statistics, the observed response shifts could reflect filter artifacts or an 'edited photo' cue rather than attractiveness as a social construct. This is the single most important threat to the central claim because it determines what the headline percentages mean. The paper is otherwise strong: seven models, 91 scenarios, order averaging, multiple seeds, and direction-consistent halo results provide real evidence that something about beautified faces changes MLLM decisions. The concern does not warrant rejection, but it does warrant a control experiment before the numbers are taken as measuring attractiveness specifically. The reader's CONDITIONAL verdict already captures this, so no verdict change is needed; the concrete sham-filter test would sharpen the condition and either validate or reinterpret the headline claim.","tokens_in":39099,"tokens_out":2711,"duration_ms":33497,"concrete_test":"Run the same 91-scenario pipeline with a sham-filter control: apply a matched edit to each original face that reproduces the beauty filter's low-level perturbation magnitude (e.g., same skin-smoothing kernel, contrast adjustment, and output resolution) while not increasing human-rated attractiveness. If MLLM answer distributions between original and sham-edited faces shift as much as they do with the real beauty filter, the reported effect is not attributable to attractiveness. If the sham edit produces no significant shift, the filter-confound concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MLLMs exhibit an attractiveness bias because they respond differently to the same individual's original versus beautified face. This inference requires the assumption, stated in Section 3.2, that \"the only difference between them is attractiveness, with minimal confounds.\" The paper validates the manipulation by citing human ratings showing beautified faces are perceived as more attractive, but that does not establish that all decision-relevant visual properties are otherwise invariant. Beauty filters also change skin texture/smoothing, apparent age, lighting, local contrast, and introduce a common filter signature. MLLMs trained on internet images may associate that signature with positive traits or with 'edited/idealized' appearance, independent of attractiveness as humans perceive it. If so, the 83.8%-86.2% scenario-level shift and the 92.6% halo-effect figure would not isolate attractiveness. The paper's own Limitations section acknowledges that inputs are curated to \"minimize confounding variables\" but does not test the invariance claim directly. The human-ratings validation from [Gulati et al.(2024)] addresses the perception of attractiveness, not the absence of other visual changes. This is a correctness risk rather than an internal inconsistency: the statistical result is plausible, but its interpretation as 'attractiveness bias' depends on a control that is asserted rather than demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an empirical study of whether multimodal large language models (MLLMs) exhibit an attractiveness bias. The authors use a dataset of 462 individuals, each with an original face image and a beautified version, and 91 forced-choice social judgment scenarios spanning stereotyped jobs, traits, and conditions. Seven open-source MLLMs are queried with each image under four option orderings and three seeds, yielding over seven million prompts. The authors define an attractiveness bias as a statistically significant Kruskal-Wallis difference between the distributions of stereotype-consistency scores for original versus beautified faces, and an attractiveness halo effect as the same bias in sentiment scenarios where the beautified face is more likely to be assigned the positive trait. The paper reports that attractiveness affects decisions in the large majority of scenarios (83.8% per Table 2; 86.2% in the abstract and Section 5), that the halo effect appears in 92.6% or 94.8% of relevant scenarios depending on the section, and that gender, age, and race biases interact with attractiveness, with female faces showing a stronger attractiveness bias and beautified images amplifying gender bias in most models. The paper concludes that MLLMs replicate the human attractiveness halo effect and calls for intersectional bias mitigation.","tokens_in":39359,"tokens_out":3773,"duration_ms":40719,"significance":"If the reported effects are real, this is a meaningful and timely contribution: it provides the first large-scale evidence that open-source MLLMs alter forced-choice judgments about the same person when facial attractiveness is manipulated, and it connects this behavior to a well-documented human cognitive bias. The study has noteworthy strengths: it covers seven diverse open-source models, uses paired original/beautified images of the same identities, averages over option orderings and random seeds, reports per-scenario statistics in the appendices, and makes the experimental design transparent enough to be replicated. The main risk is interpretive: the claim that the observed differences are specifically due to attractiveness, rather than to other visual changes introduced by beauty filters, rests on an invariance assumption that is asserted but not directly tested. The inconsistency among the headline percentages (83.8% vs 86.2%; 92.6% vs 94.8%) also needs to be resolved before the central quantitative claims can be taken at face value.","major_comments":[{"comment":"The central inference, that differences between original and beautified faces isolate attractiveness, depends on the claim in Section 3.2 that \"the only difference between them is attractiveness, with minimal confounds.\" The cited human-rating validation from Gulati et al. (2024) shows that beautified faces are perceived as more attractive, but it does not establish that other decision-relevant visual properties are invariant. Beauty filters typically alter skin texture, apparent age, lighting, local contrast, and introduce a common filter signature, and MLLMs may exploit any of these cues. Because the paper's RQ1 and RQ2 conclusions are precisely about attractiveness, the authors should provide direct evidence on this invariance, for example by measuring low-level image similarity between paired images, testing a filter-detection classifier, or evaluating models on control tasks where attractiveness should be irrelevant (e.g., identity, expression, or age judgment). Without such a control, the reported shifts in phi_i could be driven partly by filter artifacts rather than by attractiveness as humans perceive it.","section":"Section 3.2; Section 5; Figure 1"},{"comment":"The headline numbers are inconsistent across the manuscript. The abstract states that attractiveness impacts decisions in 86.2% of scenarios and that the halo effect appears in 94.8% of relevant scenarios, while Table 2 reports an average of 83.8% for attractiveness bias and Section 4 reports 92.6% for the halo effect in the 33 sentiment scenarios. Section 5 repeats 86.2% and 92.6%. The paper does not explain whether these are different aggregation methods, different scenario subsets, or errors. Since these percentages are the paper's central quantitative claims, the authors must reconcile the numbers and specify the exact computation underlying each reported figure.","section":"Abstract; Section 4; Section 5; Table 2"},{"comment":"The attractiveness-bias test compares the distributions of phi_i over original faces and their paired beautified faces using a Kruskal-Wallis test, which treats the two samples as independent groups. Because each beautified image is derived from a specific original image of the same individual, the observations are paired. A paired test (e.g., Wilcoxon signed-rank test) or a mixed-effects model with identity as a random effect would use the pairing structure and is more directly aligned with the paper's claim of comparing \"the same individuals\" before and after beautification. The authors should either justify the unpaired test or report paired analyses; this is not necessarily fatal given the large effects, but it is a methodological mismatch that affects the statistical framing.","section":"Section 3.5; Section 4"}],"minor_comments":[{"comment":"The captions say \"Out of 19 scenarios\" for the race-stereotyped jobs, but there are 12 such scenarios and each table has 12 rows. This appears to be a copy-paste error from the gender-stereotyped jobs appendix and should be corrected.","section":"Appendix G, Tables 24-30"},{"comment":"The sentence \"In all scenarios and for all models (except for 3 out of 31 scenarios for DeepSeek and 1 out of 30 for Qwen2)\" is confusing because Tables 12 and 14 report different totals (28 out of 33 and 28 out of 33 significant scenarios, with 3 and 1 opposite-direction cases respectively). The text should state the denominator and direction conventions clearly.","section":"Section 4, RQ2"},{"comment":"The table formatting uses colored shading with a legend that may not survive black-and-white printing; the authors should ensure all information is also encoded textually or with patterns.","section":"Table 2 and Tables 31-33"},{"comment":"The definition of the Bias section uses H with subscripts interchangeably for the Kruskal-Wallis statistic and the hypothesis label; for clarity, the paper should consistently distinguish the test statistic from the hypothesis being tested.","section":"Section 3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is an empirical measurement, and the direction of the finding is consistent across models and scenarios, so the qualitative conclusion is credible. The load-bearing issue is the filter-confound assumption: the authors need to show that the observed changes are specifically due to attractiveness rather than to correlated low-level image changes. The inconsistency in headline percentages is also a blocking issue that must be fixed before acceptance. The dependence on a dataset and human ratings from a prior paper by the same first author is not inherently circular, but the authors should make the dataset's availability and the precise rating protocol explicit so that reviewers and readers can assess the manipulation check independently."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first controlled paired-image study I've seen that directly tests attractiveness bias in MLLMs using natural faces. The protocol is clean—same individual, original vs. beautified, 462 faces, 91 scenarios, seven open models, order-invariant scoring, three seeds. That is a real step beyond the synthetic-image benchmarks in the prior work, and the empirical pattern is consistent: beautified faces are significantly more likely to be assigned positive traits, and the effect shows up across all seven models. The large Kruskal-Wallis statistics in the appendix support the qualitative conclusion, and the gender intersection result (stronger attractiveness bias for female faces) aligns with the human literature. I find the core finding credible and worth taking seriously.\n\nWhere the paper is softer: first, the abstract numbers don't match the results tables. The abstract says 86.2% of scenarios and 94.8% for the halo effect; Table 2 and the text say 83.8% and 92.6%. That's a minor but embarrassing inconsistency that must be fixed. Second, the load-bearing assumption that beauty filters change 'only attractiveness' is asserted, not demonstrated. The paper's own discussion admits filters make people look younger, and they also smooth skin, alter lighting, and leave a filter signature. So the measured shifts could partly reflect filter artifacts or an \"edited/idealized\" cue rather than attractiveness per se. The human-rating validation from the prior study shows the manipulation works on attractiveness perception, but it doesn't show other decision-relevant features are invariant. This is the main interpretive caveat, and it means the specific percentages are probably upper bounds on the pure attractiveness effect. Third, the per-scenario KW tests are run at p<0.01 with no multiple-comparison correction across 91 scenarios, so some false positives are likely. The effect sizes are large enough that the main conclusion survives, but the reported proportions are optimistic. Fourth, code and data are \"available upon request\"—for a paper whose contribution is a dataset and protocol, that is weak; they should ship them.\n\nWho is this for? Anyone working on fairness evaluation of vision-language models, and people studying how cognitive biases transfer to AI. It deserves a serious referee: the design is thoughtful, the finding is important if it holds, and the weaknesses are addressable rather than fatal. My recommendation: send it to review, and ask the authors to reconcile the numbers, release the artifacts, and directly test the filter-confound concern—e.g., by measuring which low-level features change or by including a control condition with a non-attractiveness filter.","headline":"A credible, large-scale measurement that beauty filters shift MLLM judgments in a stereotype-consistent direction, but the 'only difference is attractiveness' claim is overstated and the abstract numbers need a sanity check.","tokens_in":39857,"tokens_out":1759,"would_cite":true,"duration_ms":21317,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that seven open-source multimodal large language models give systematically different forced-choice answers about the same person when their face is beautified, associating beautified faces with positive traits and…","keywords":["attractiveness bias","attractiveness halo effect","multimodal large language models","beauty filters","intersectional bias","algorithmic fairness","face perception","cognitive bias in AI"],"falsifier":"Re-run the 91 scenarios with an attractiveness manipulation that does not touch low-level image properties — for example, morph each face by a fixed amount toward a consensus-attractive prototype while holding age, skin texture, lighting, expression, and identity constant. If the original and manipulated versions no longer produce significantly different answer distributions at $p<0.01$, the reported bias is an artifact of the beauty filter rather than an attractiveness halo effect in the models.","tokens_in":38933,"feed_emoji":"✨","tokens_out":10440,"duration_ms":95789,"temperature":0.7,"pith_summary":"This paper tries to establish that a well-documented human judgment bias — the \"attractiveness halo effect,\" in which attractive people are assumed to have positive traits — is also present in multimodal large language models (MLLMs), systems that answer text prompts about images. To test this, the authors show 462 faces to seven open-source MLLMs both in original form and after a commercial beauty filter, and ask 91 forced-choice questions about jobs, personality traits, and social conditions. The paper reports that attractiveness changed model answers in 86.2% of scenarios on average, and that beautified faces were significantly more often assigned positive traits, such as confidence and trustworthiness, in most of the sentiment scenarios, which is direct evidence of the halo effect in machines. The result matters because MLLMs are beginning to inform real decisions in hiring, education, and professional evaluation, where physical attractiveness should be irrelevant.","feed_headline":"Beauty filters flip AI judgments in 86% of test scenarios","feed_subtitle":"Seven vision-language models judge the same face more positively once it is beautified.","key_machinery":"The load-bearing device is a paired-image counterfactual: every one of the 462 identities appears twice, once as the original photograph and once passed through a beauty filter, so the input pair is intended to differ only in perceived attractiveness. For each of the 91 scenarios, the MLLM must choose between two options; responses are converted to a Stereotype Consistency Score (SCS), the fraction of times an image is assigned to the stereotyped choice, averaged over the four possible orderings of the two options and three random seeds to remove position and sampling effects. An attractiveness bias is declared when a Kruskal-Wallis test at $p<0.01$ finds that the score distributions for originals and beautified versions differ; the halo effect is the directional case in which beautified faces are more often matched to the positive trait. Wilcoxon paired rank tests across the 91 scenarios, with Bonferroni correction, then quantify how gender, age, and race change the strength of the attractiveness bias.","core_discovery":"The central claim is stated directly: physical attractiveness biases the decisions made by MLLMs. Using the same individual's face in original and beautified versions, the paper observes statistically significant differences in forced-choice answers across a large majority of 91 scenarios, with the effect present in every one of the seven models tested. The attractiveness halo effect appears in the sentiment trait scenarios, where beautified faces are more likely than original faces to be labelled confident, trustworthy, happy, kind, or competent; the paper reports this in 94.8% of the relevant scenarios. The bias is not uniform: it is stronger for female faces than male faces, beauty filters amplify gender bias in most models, and the effect of race and age is attenuated by beautification in several models. The authors conclude that attractiveness is an \"invisible\" decision cue for MLLMs, operating implicitly and intersecting with demographic stereotypes.","pith_inferences":["A beauty filter also changes skin texture, apparent age, and lighting; until attractiveness is manipulated independently of those cues, part of the measured bias could come from filter artifacts rather than from attractiveness as humans perceive it.","The forced-choice format may understate the bias in free-form conversation, where a model can hedge or explain; asking the same models open-ended questions about the same paired faces is a direct next test.","If the halo effect is encoded in pretraining associations, then text-only language models may show a similar effect when faces are described verbally, which would make the bias about language statistics rather than vision.","A practical extension would be to use this same original-versus-beautified contrast as a routine fairness probe during MLLM release testing, complementing demographic parity checks."],"forward_implications":["If the result holds, any vision-based screening or assessment system built on an MLLM can silently favor more attractive-looking people even in settings where looks carry no information.","The human attractiveness halo effect transfers to machines: the same person is judged more trustworthy, confident, or competent when their face is beautified, so appearance editing can change model-based evaluations.","Attractiveness bias is intersectional: it is strongest for female faces, and beautification amplifies gender bias in most of the seven models, while reducing age and race biases in some.","Fairness audits that check only gender, race, or age will miss a bias that operates on appearance and modulates those demographic biases; attractiveness should be an explicit evaluation dimension.","The paired original/beautified image protocol provides a reusable probe for appearance-based bias in open MLLMs, including for models not covered in this study."],"supporting_citations":[{"why":"Supplies the paired original/beautified face dataset and human attractiveness ratings used to validate the beauty manipulation.","marker":"[Gulati et al.(2024)]"},{"why":"Supplies race-diverse face images from one of the two source databases.","marker":"[Ma et al.(2015)]"},{"why":"Supplies age-diverse face images from the other source database.","marker":"[Ebner et al.(2010)]"},{"why":"Defines the attractiveness halo effect that the paper tests in MLLMs.","marker":"[Dion et al.(1972)]"},{"why":"Supplies the So-B-IT trait taxonomy from which sentiment trait pairs were selected.","marker":"[Hamidieh et al.(2024)]"},{"why":"Supplies VLStereoSet stereotyped trait and condition pairs used in scenarios.","marker":"[Zhou et al.(2022)]"},{"why":"Supplies ModSCAN gender-stereotyped hobby pairs used in trait scenarios.","marker":"[Jiang et al.(2024)]"},{"why":"Supplies gender-stereotyped job pairs drawn from labor-force data.","marker":"[Fraser and Kiritchenko(2024)]"},{"why":"Supplies additional gender-stereotyped job pairs used in job scenarios.","marker":"[Xiao et al.(2024)]"}],"fun_headline_variants":["AI models show beauty bias: attractive faces get favorable traits","In 86% of scenarios, AI judgments flip with beauty filters","Attractiveness halo effect: AI gives pretty faces an edge","AI's beauty bias: attractive faces seen as smarter","Pretty privilege in AI: beauty filters tilt model decisions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire measurement rests on the assumption that the beauty filter changes only perceived attractiveness and leaves every other decision-relevant property of the face unchanged; if the filter also shifts apparent age, skin texture, lighting, or expression, the answer differences could be caused by those changes rather than by attractiveness.","fun_headline_variants_meta":{"raw":{"variants":["AI models show beauty bias: attractive faces get favorable traits","In 86% of scenarios, AI judgments flip with beauty filters","Attractiveness halo effect: AI gives pretty faces an edge","AI's beauty bias: attractive faces seen as smarter","Pretty privilege in AI: beauty filters tilt model decisions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000668,"raw_usage":{"total_tokens":3085,"prompt_tokens":1019,"completion_tokens":2066,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":1984}},"tokens_in":635,"tokens_out":2066,"duration_ms":15094,"temperature":1.0,"reasoning_tokens":1984,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:35:17.657193+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 91 scenarios with an attractiveness manipulation that does not touch low-level image properties — for example, morph each face by a fixed amount toward a consensus-attractive prototype while holding age, skin texture, lighting, expression, and identity constant. If the original and manipulated versions no longer produce significantly different answer distributions at $p<0.01$, the reported bias is an artifact of the beauty filter rather than an attractiveness halo effect in the models.","supporting_citations":[],"review_version":1}