Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Generative image models systematically default to US-style, modern-leaning imagery, and iterative image editing quietly erodes cultural fidelity even when standard metrics improve — a sign that culture-sensitive generation and editing remai

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 08:34 UTC pith:LYRKKIQR

load-bearing objection A serious cultural-bias audit with a valuable new I2I-editing angle and a clean T2I default finding, but the headline erosion claim is undercut by a quality/culture confound in the human metric and an inconsistent rating protocol. the 3 major comments →

arxiv 2510.20042 v3 pith:LYRKKIQR submitted 2025-10-22 cs.CV

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

classification cs.CV
keywords cultural biastext-to-image generationimage-to-image editingcultural fidelityevaluation metricshuman evaluationera-aware promptsgenerative image models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that current generative image models carry a systematic cultural bias that standard evaluation metrics miss. Using a balanced six-country, 36-subcategory protocol with era-aware prompts, it finds that country-agnostic prompts collapse to a US-like, modern-leaning default. It then shows that iterative image-to-image editing erodes cultural fidelity across edit steps while conventional automated alignment scores stay flat or improve, and that editors often substitute superficial cues such as palette shifts for era-consistent cultural detail. The authors argue that culture-sensitive generation and editing are therefore unreliable today, and that auditing cultural bias requires culture-aware metrics and human judgments, not generic quality scores.

Core claim

The paper's central claim is that current open generative image models are systematically culturally biased in three ways. First, when prompts do not specify a country, generated images default to a US-like, modern aesthetic, flattening distinctions among countries. Second, in iterative image-to-image editing, repeated instructions to align an image with a target culture progressively erode culturally specific cues even as conventional automated alignment scores remain flat or improve; native-expert ratings and a culture-aware retrieval-augmented VQA metric both register the decline. Third, editors tend to apply superficial markers — palette shifts, generic props — rather than era-consistent

What carries the argument

The load-bearing machinery is a standardized, three-layer evaluation stack: a granular prompt schema (six countries × eight categories × 36 subcategories × traditional/modern/era-agnostic prompt modes), a dual protocol that first generates a T2I base corpus and then runs three I2I editing experiments (multi-loop edits, attribute addition, cross-country restylization), and a comparison of three evaluation channels — conventional automatic alignment metrics, a culture-aware retrieval-augmented visual question-answering metric, and expert ratings by native reviewers. The argument turns on the divergence between the conventional metrics and the human/culture-aware channels across iterative edit

Load-bearing premise

The paper's central editing-erosion finding depends on its combined human quality score — the average of image-quality and cultural-representation ratings — actually measuring cultural fidelity; the strong correlation with aesthetic score means the stepwise decline could be driven mainly by generic visual quality rather than cultural loss.

What would settle it

Re-rate the released image corpus using only the cultural-representation rating, without averaging in image quality, and compare step-1 to step-5 trajectories per model-country pair; if cultural-only scores stay flat or rise while combined human quality scores fall, the claim that iterative editing erodes cultural fidelity is refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Country-agnostic prompting cannot be treated as culturally neutral; without an explicit country, models produce US-like modern imagery, so specifying culture is necessary for diversity.
  • Standard prompt-alignment metrics are inadequate for auditing cultural bias: they can report improvement while human-perceived cultural quality falls sharply.
  • Iterative editing amplifies rather than corrects cultural bias, so relying on repeated edits to 'fix' an image is unsafe without culture-specific monitoring.
  • Current editors signal culture through surface cues such as palette and props instead of context-consistent transformations, and they often preserve source identity when targeting Global-South countries, leaving correction burden on users.
  • A culture-aware automated metric can track human best/worst judgments at high agreement rates, making scalable auditing of gross cultural failure plausible.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If the metric-blind erosion generalizes, edit-history UI features — such as a cultural-fidelity score per step or auto-stop after the first edit — would be more useful than raw quality thumbnails for culture-sensitive products.
  • Editorial inference: The country-level unit obscures subnational and diaspora variation; a finer-grained version of this protocol would likely reveal even larger default biases within countries like the US and China.
  • Editorial inference: The strong agreement on 'worst' selections suggests automated culture-aware auditing works best as a failure detector; detecting subtle or mid-range cultural distortion may still need humans.
  • Editorial inference: A directly testable extension is an early-stopping rule: applying only one or two edit steps may preserve cultural fidelity better than five, a hypothesis the stepwise HQS decline data suggests.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a unified, reproducible evaluation of cultural bias in both text-to-image (T2I) and image-to-image (I2I) models, covering six countries, an 8-category/36-subcategory schema, era-aware prompts, and three I2I editing protocols. It combines standard automatic metrics (CLIPScore, DreamSim, Aesthetic Score) with a retrieval-augmented VQA-based culture-aware metric and expert human judgments from native raters. The main claims are: (1) country-agnostic prompts default to a U.S.-like, modern-leaning style; (2) iterative I2I editing erodes cultural fidelity even when conventional metrics stay flat or improve, while human ratings and the culture-aware metric register the degradation; and (3) I2I models rely on superficial cues rather than context-consistent cultural changes. The paper also reports occupational gender and skin-tone skews. The authors release their image corpus, prompts, and configurations.

Significance. If the central claims hold, the paper provides a valuable benchmark and methodology for auditing cultural bias in generative image models, especially for the underexplored I2I setting. The release of reproducible artifacts, the use of cross-model cluster analysis with permutation tests and FDR control for the T2I default finding, and the triangulation of automatic, culture-aware, and human judgments are concrete strengths. The country-native expert protocol is also a positive feature. However, the key I2I erosion claim depends on a human quality measure whose composition is ambiguous and potentially confounded with generic image quality, so the significance of finding (2) is not yet established as stated.

major comments (3)
  1. [§3.4 vs Appendix B] There is an internal inconsistency in the definition of the primary human reference. Section 3.4 states that raters assign two scores—Image Quality and Cultural Representation—and that HQS is their average. Appendix B states that raters assign three 1–5 Likert scores—Image Quality, Prompt Alignment, and Cultural Representation—and §4.2.2 also says three dimensions. This is not a cosmetic issue: HQS is the reference against which the I2I erosion claim is made. The paper must specify exactly which components entered HQS and whether Prompt Alignment was included, and report component-wise results.
  2. [§3.4, Appendix D.3, Appendix E.2] The central claim that iterative I2I editing erodes cultural fidelity, while conventional metrics stay flat or improve, is not directly supported because HQS is the average of Image Quality and Cultural Representation and no component-wise trajectories are reported. Appendix E.2 shows that Aesthetic Score correlates r=0.78 with HQS and also declines across edit steps; Appendix D.3 reports HQS declines in percentages. If the HQS decline is driven mainly by the Image Quality component, then the finding reduces to a generic quality-collapse effect, not cultural erosion. The high agreement of the culture-aware metric with human Best/Worst selections (73.8%/83.7%) is also plausibly explained by shared sensitivity to quality. The authors should report the two (or three) HQS components separately across edit steps, and ideally condition the culture-aware metric agreement on the Cultural Represe
  3. [Eq. (5), §4.1.2] The traditional–modern leaning score in Eq. (5) relies on category-specific prototypes μ_trad and μ_mod, but their construction is not specified. If these prototypes are mean embeddings of the model's own traditional/modern generations, then the resulting 'modern lean' and 'traditional lean' labels are partly circular: they measure agreement with the model's own modes rather than with an external, culturally grounded standard. The paper should state how μ_trad and μ_mod are obtained, whether they are fixed independent anchors, and, if they are derived from the models, what robustness checks were performed.
minor comments (5)
  1. [§4.1, Eq. (1)] The number of clusters Km in the k-means analysis is never reported. Since the distributional-proximity results depend on Km, a value and a sensitivity analysis (e.g., Km∈{5,10,20}) would strengthen the US-default claim.
  2. [Abstract, §4.2.1, Appendix E.2] The abstract and Section 4.2.1 state that conventional metrics 'remain flat or improve,' but the paper's own Appendix E.2 reports that Aesthetic Score declines across edit steps. Please clarify which metrics are covered by the claim, or qualify it as applying to CLIPScore in particular.
  3. [Appendix D.2.1] The text says CLIPScore changes range from -5.1% to +5.1%, but Table 5 lists several values close to -5.0% and +5.1%; please reconcile the exact numbers and report the mean change consistently.
  4. [Appendix D.3] The subsection is titled 'Cultural Bias Analysis' but the reported values are HQS changes. Since HQS includes Image Quality, the title overstates what is measured. Renaming or separating the component analyses would avoid confusion.
  5. [§6, Limitations] The limitations paragraph appropriately acknowledges the coarseness of country-level labels, but the paper could also note that the occupation bias audit uses automated classification whose error rates are not reported; a brief validation would help.

Circularity Check

0 steps flagged

No significant circularity: central claims rest on external human ratings and a culture-aware VQA independently validated against them, not on fitted targets or self-citation chains.

full rationale

The paper's central claims are empirical: (1) country-agnostic prompts default to U.S.-like/modern depictions, (2) iterative I2I editing erodes cultural fidelity while CLIPScore stays flat or improves, and (3) I2I models rely on superficial cues. The load-bearing evidence for these claims is triangulated from three independent sources: standard automatic metrics (CLIPScore, DreamSim, Aesthetic Score), a culture-aware retrieval-augmented VQA metric, and expert human ratings from country-native reviewers. The culture-aware metric is not fitted to the human ratings it is later compared against; the paper states it is built by retrieving Wikipedia context via FAISS, generating yes/no questions with Qwen2.5-0.5B, and answering with Qwen2-VL-7B, and only after construction is it benchmarked against human Best/Worst selections. That is validation, not circular fitting. The Human Quality Score (HQS) is defined as the average of Image Quality and Cultural Representation, and is a human reference rather than a derived prediction of the model's own outputs. The concern that HQS may confound cultural fidelity with image quality is a construct-validity issue, not a circularity: no equation in the paper reduces the reported 'cultural erosion' to the fitted value of any parameter. The manuscript does contain an internal inconsistency, in that Section 3.4 says raters assign two scores while Appendix B says they assign three (adding Prompt Alignment); this is a reporting flaw that weakens the precision of the HQS composition but does not by itself make the derivation circular. The traditional-modern leaning score (Eq. 5) uses category-specific prototypes whose construction is not fully specified in the main text, and if those prototypes were mean embeddings of the models' own traditional/modern outputs, the 'modern default' finding would partly measure the model against itself. However, the paper does not state that construction, and the main country-level results are corroborated by the cluster-proportion analysis (Eqs. 1-4), which does not depend on such prototypes. Self-citations to the authors' prior CCUB and SCoFT work appear only as related-work motivation and as extensions of prior benchmarking efforts, not as the justification for any prediction or uniqueness claim. No load-bearing step reduces to its own input by construction, and no fitted parameter is renamed as a prediction. Therefore the appropriate circularity score is 0.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper's quantitative claims rely on several ungrounded choices: the number of clusters, the definition of traditional-modern prototypes, and the weighting of human rating components. Domain assumptions about country-as-culture, CLIP semantics, and VQA validity are acknowledged or implicit.

free parameters (3)
  • Cluster count Km (k-means) = not reported
    Per-model cluster count Km in Eq. 1/4 changes country proximity estimates; no criterion or value is given, so the number of latent modes is a free choice.
  • Traditional/modern prototypes μtrad, μmod = not specified externally
    Eq. 5's leaning score depends on category-specific prototypes; unless these are externally curated anchors, they are internal means from the model's own generations, making the score self-referential.
  • HQS equal-weight averaging = 0.5 image quality + 0.5 cultural representation
    Human Quality Score is defined as the average of Image Quality and Cultural Representation; equal weighting is an arbitrary modeling choice that conflates aesthetic and cultural components.
axioms (4)
  • domain assumption Country-level labels are an adequate proxy for culture in this evaluation
    All country-level comparisons and conclusions take sovereign states as the unit of culture; the paper acknowledges this obscures subnational, minority, and diaspora heterogeneity (Limitations).
  • domain assumption Native/expert raters' judgments are the ground truth for cultural fidelity
    HQS and Best/Worst selections are treated as the primary reference for cultural correctness; no external ground-truth annotation set is used.
  • domain assumption CLIP embeddings encode cultural and traditional-modern semantics
    Eq. 5 and cluster proximity rely on CLIP embedding space; if CLIP lacks cultural distinctions, the measured proximities and leaning scores may be artifacts.
  • domain assumption Qwen2-VL-7B with Wikipedia retrieval is a valid culture-aware evaluator
    The proposed culture-aware metric depends on the VQA model's answers; the paper validates it against human Best/Worst choices but not against human ratings, and VQA biases are acknowledged as a limitation.

pith-pipeline@v1.3.0-alltime-deepseek · 11 in / 13072 out tokens · 298388 ms · 2026-08-04T08:34:24.843284+00:00 · methodology

0 comments
read the original abstract

Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image (I2I) editors underexplored. We bridge this gap with a unified evaluation across six countries, an 8-category/36-subcategory schema, and era-aware prompts, auditing both T2I generation and I2I editing under a standardized protocol that yields comparable diagnostics. Using open models with fixed settings, we derive cross-country, cross-era, and cross-category evaluations. Our framework combines standard automatic metrics, a culture-aware retrieval-augmented VQA, and expert human judgments collected from native reviewers. To enable reproducibility, we release the complete image corpus, prompts, and configurations. Our study reveals three findings: (1) under country-agnostic prompts, models default to Global-North, modern-leaning depictions that flatten cross-country distinctions; (2) iterative I2I editing erodes cultural fidelity even when conventional metrics remain flat or improve; and (3) I2I models apply superficial cues (palette shifts, generic props) rather than era-consistent, context-aware changes, often retaining source identity for Global-South targets. These results highlight that culture-sensitive edits remain unreliable in current systems. By releasing standardized data, prompts, and human evaluation protocols, we provide a reproducible, culture-centered benchmark for diagnosing and tracking cultural bias in generative image models. Project page: https://seochan99.github.io/ECB

Figures

Figures reproduced from arXiv: 2510.20042 by Huichan Seo, Jean Oh, Jihie Kim, Junseo Kim, Lukman Ismaila, Mehul Agarwal, Minki Hong, Naome Etori, Sieun Choi, Yi Zhou, Zhixuan Liu.

Figure 1
Figure 1. Figure 1: Representative cultural biases in T2I generations across six countries. Examples include [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overall framework overview. (a) Schema inputs: six countries, eight categories, and three [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparative samples for U.S. (top) and country-agnostic (bottom) prompts across two [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Divergence between Automated and Human Judgment in Iterative Editing. (a) CLIPScore [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Alignment of the Culture-aware Metric with Human Judgment (All Models Averaged). (a) [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Representative stepwise attribute additions for Korea and the U.S. Columns progress from [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Representative cross-country restyling results for China and Nigeria; columns are ordered [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Occupation-level demographic distributions in T2I outputs. We evaluate 12 occupations [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: ECB Human Survey UI snapshots. (a) consent/IRB gating; (b) participant dashboard; (c) [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Alignment of the human selection with our culture-aware metric by model and country. [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Average CLIPScore Progression Across Iterative I2I Steps by Country and Model. [PITH_FULL_IMAGE:figures/full_fig_p019_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Average HQS Progression Across Iterative I2I Steps by Country and Model. [PITH_FULL_IMAGE:figures/full_fig_p020_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: CLIPScore analysis. (a) Average CLIPScore by model reveals minimal variation (3.1–3.4 [PITH_FULL_IMAGE:figures/full_fig_p023_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Aesthetic Score analysis. (a) Model-wise comparision: SD3.5 most stable (3.3–3.9), [PITH_FULL_IMAGE:figures/full_fig_p024_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: DreamSim delta analysis. (a) Model-wise: FLUX.1 highest total change (0.38–0.54), [PITH_FULL_IMAGE:figures/full_fig_p025_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Multi-loop I2I editing across countries using Qwen-Image-Edit (best-performing editor). [PITH_FULL_IMAGE:figures/full_fig_p027_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Occupational demographic bias examples using HiDream-I1 (best-performing generator). [PITH_FULL_IMAGE:figures/full_fig_p028_17.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

    cs.CY 2026-05 unverdicted novelty 6.0

    A participatory red-teaming project in the Global South created the PLACES dataset of 26k T2I failure examples that reveal unique cultural and linguistic harms missed by existing safety frameworks.

  2. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

Reference graph

Works this paper leans on

42 extracted references · 12 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Photorealistic text-to- image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to- image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

  2. [2]

    Scaling rectified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. InForty-first international conference on machine learning, 2024

  3. [3]

    A review on generative ai for text-to-image and image-to-image generation and implications to scientific images.arXiv preprint arXiv:2502.21151, 2025

    Zineb Sordo, Eric Chagnon, and Daniela Ushizima. A review on generative ai for text-to-image and image-to-image generation and implications to scientific images.arXiv preprint arXiv:2502.21151, 2025

  4. [4]

    A survey of text-to-image diffusion models in generative ai

    Siddharth Kandwal and Vibha Nehra. A survey of text-to-image diffusion models in generative ai. In2024 14th International Conference on Cloud Computing, Data Science & Engineering (Confluence), pages 73–78. IEEE, 2024

  5. [5]

    Survey of bias in text-to-image generation: Definition, evaluation, and mitigation.arXiv preprint arXiv:2404.01030, 2024

    Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. Survey of bias in text-to-image generation: Definition, evaluation, and mitigation.arXiv preprint arXiv:2404.01030, 2024

  6. [6]

    An image speaks a thousand words, but can everyone listen? on image transcreation for cultural relevance

    Simran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, and Graham Neubig. An image speaks a thousand words, but can everyone listen? on image transcreation for cultural relevance. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10258–10279, 2024

  7. [7]

    On the cultural gap in text-to-image generation.arXiv preprint arXiv:2307.02971, 2023

    Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang, Jinsong Su, Shuming Shi, and Zhaopeng Tu. On the cultural gap in text-to-image generation.arXiv preprint arXiv:2307.02971, 2023

  8. [8]

    Beyond aesthetics: cultural competence in text-to-image models

    Nithish Kannen, Arif Ahmad, Marco Andreetto, Vinodkumar Prabhakaran, Utsav Prabhu, Adji Bousso Dieng, Pushpak Bhattacharyya, and Shachi Dave. Beyond aesthetics: cultural competence in text-to-image models. InProceedings of the 38th International Conference on Neural Information Processing Systems, pages 13716–13747, 2024

  9. [9]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

  10. [10]

    Clipscore: A reference-free evaluation metric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7514–7528, 2021

  11. [11]

    Diffusion models through a global lens: Are they culturally inclusive?arXiv preprint arXiv:2502.08914, 2025

    Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, and Alice Oh. Diffusion models through a global lens: Are they culturally inclusive?arXiv preprint arXiv:2502.08914, 2025

  12. [12]

    Culturalframes: Assessing cultural expectation alignment in text- to-image models and evaluation metrics.arXiv preprint arXiv:2506.08835, 2025

    Shravan Nayak, Mehar Bhatia, Xiaofeng Zhang, Verena Rieser, Lisa Anne Hendricks, Sjoerd van Steenkiste, Yash Goyal, Aishwarya Agrawal, et al. Culturalframes: Assessing cultural expectation alignment in text- to-image models and evaluation metrics.arXiv preprint arXiv:2506.08835, 2025

  13. [13]

    Dreamsim: learning new dimensions of human visual similarity using synthetic data

    Stephanie Fu, Netanel Y Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dreamsim: learning new dimensions of human visual similarity using synthetic data. InProceedings of the 37th International Conference on Neural Information Processing Systems, pages 50742–50768, 2023

  14. [14]

    Predicting scores of various aesthetic attribute sets by learning from overall score labels

    Heng Huang, Xin Jin, Yaqi Liu, Hao Lou, Chaoen Xiao, Shuai Cui, Xining Li, and Dongqing Zou. Predicting scores of various aesthetic attribute sets by learning from overall score labels. InProceedings of the 2nd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice, pages 63–71, 2024. 12

  15. [15]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  16. [16]

    Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, et al. Flux. 1 kontext: Flow matching for in-context image generation and editing in latent space.arXiv preprint arXiv:2506.15742, 2025

  17. [17]

    Hidream-i1: A high-efficient image generative foundation model with sparse diffusion transformer.arXiv preprint arXiv:2505.22705, 2025

    Qi Cai, Jingwen Chen, Yang Chen, Yehao Li, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Yiheng Zhang, Fengbin Gao, Peihan Xu, et al. Hidream-i1: A high-efficient image generative foundation model with sparse diffusion transformer.arXiv preprint arXiv:2505.22705, 2025

  18. [18]

    Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025

    Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng-ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al. Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025

  19. [19]

    Nextstep-1: Toward autoregressive image generation with continuous tokens at scale.arXiv preprint arXiv:2508.10711, 2025

    NextStep Team, Chunrui Han, Guopeng Li, Jingwei Wu, Quan Sun, Yan Cai, Yuang Peng, Zheng Ge, Deyu Zhou, Haomiao Tang, et al. Nextstep-1: Toward autoregressive image generation with continuous tokens at scale.arXiv preprint arXiv:2508.10711, 2025

  20. [20]

    Easily accessible text-to-image generation amplifies demographic stereotypes at large scale

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. InProceedings of the 2023 ACM conference on fairness, accountability, and transparency, pages 1493–1504, 2023

  21. [21]

    Social biases through the text-to-image generation lens

    Ranjita Naik and Besmira Nushi. Social biases through the text-to-image generation lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 786–808, 2023

  22. [22]

    Tibet: Identifying and evaluating biases in text-to-image generative models

    Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, and Matthew Turk. Tibet: Identifying and evaluating biases in text-to-image generative models. InEuropean Conference on Computer Vision, pages 429–446. Springer, 2024

  23. [23]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. InEuropean conference on computer vision, pages 740–755. Springer, 2014

  24. [24]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009

  25. [25]

    Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

  26. [26]

    The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world.Advances in Neural Information Processing Systems, 35:12979–12990, 2022

    William Gaviria Rojas, Sudnya Diamos, Keertan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman. The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world.Advances in Neural Information Processing Systems, 35:12979–12990, 2022

  27. [27]

    Wit: Wikipedia- based image text dataset for multimodal multilingual machine learning

    Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. Wit: Wikipedia- based image text dataset for multimodal multilingual machine learning. InProceedings of the 44th international ACM SIGIR conference on research and development in information retrieval, pages 2443– 2449, 2021

  28. [28]

    Towards equitable representation in text-to-image synthesis models with the cross-cultural understanding benchmark (ccub) dataset.arXiv preprint arXiv:2301.12073, 2023

    Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu, Youngsik Yun, Lia Coleman, Peter Schaldenbrand, Jihie Kim, and Jean Oh. Towards equitable representation in text-to-image synthesis models with the cross-cultural understanding benchmark (ccub) dataset.arXiv preprint arXiv:2301.12073, 2023

  29. [29]

    Scoft: Self-contrastive fine-tuning for equitable image generation

    Zhixuan Liu, Peter Schaldenbrand, Beverley-Claire Okogwu, Wenxuan Peng, Youngsik Yun, Andrew Hundt, Jihie Kim, and Jean Oh. Scoft: Self-contrastive fine-tuning for equitable image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10822– 10832, 2024

  30. [30]

    Benchmarking vision language models for cultural understanding

    Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Steenkiste, Lisa Hendricks, Karolina Stanczak, and Aishwarya Agrawal. Benchmarking vision language models for cultural understanding. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5769–5790, 2024. 13

  31. [31]

    Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22500–22510, 2023

  32. [32]

    Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis

    Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. CoRR, 2023

  33. [33]

    Evaluating text-to-visual generation with image-to-text generation.CoRR, 2024

    Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. Evaluating text-to-visual generation with image-to-text generation.CoRR, 2024

  34. [34]

    Deconstructing bias: A multifaceted framework for diagnosing cultural and compositional inequities in text-to-image generative models.arXiv preprint arXiv:2505.01430, 2025

    Muna Numan Said, Aarib Zaidi, Rabia Usman, Sonia Okon, Praneeth Medepalli, Kevin Zhu, Vasu Sharma, and Sean O’Brien. Deconstructing bias: A multifaceted framework for diagnosing cultural and compositional inequities in text-to-image generative models.arXiv preprint arXiv:2505.01430, 2025

  35. [35]

    Cvqa: culturally- diverse multilingual visual question answering benchmark

    David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo, Teresa Lynn, Injy Hamed, Aditya Nanda Kishore, Aishik Mandal, Alina Dragonetti, Artem Abzaliev, Atnafu Lambebo Tonja, et al. Cvqa: culturally- diverse multilingual visual question answering benchmark. InProceedings of the 38th International Conference on Neural Information Processing Systems, pages 1147...

  36. [36]

    Retrieval-augmented generation for knowledge- intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge- intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

  37. [37]

    Ravenea: A benchmark for multimodal retrieval-augmented visual culture understanding.arXiv preprint arXiv:2505.14462, 2025

    Jiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi, Daniel Hershcovich, Anders Søgaard, Ivan Vuli´c, Wenxuan Zhang, Paul Pu Liang, Yang Deng, et al. Ravenea: A benchmark for multimodal retrieval-augmented visual culture understanding.arXiv preprint arXiv:2505.14462, 2025

  38. [38]

    Qwen2.5 technical report, 2025

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...

  39. [39]

    Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2:3, 2024

    Qwen Team et al. Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2:3, 2024

  40. [40]

    Balancing preservation and modification: A region and semantic aware metric for instruction-based image editing.arXiv preprint arXiv:2506.13827, 2025

    Zhuoying Li, Zhu Xu, Yuxin Peng, and Yang Liu. Balancing preservation and modification: A region and semantic aware metric for instruction-based image editing.arXiv preprint arXiv:2506.13827, 2025

  41. [41]

    Gender bias in coreference resolution: Evaluation and debiasing methods

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Gender bias in coreference resolution: Evaluation and debiasing methods. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, 2018. 14 Appendix A Model...

  42. [42]

    convergence

    When comparing models, SD3.5 proved the most stable (averaging 3.3–3.9), whereas NextStep showed the steepest decline (averaging 2.7–3.5). Analyzing countries, Kenya and the United States registered the highest overall scores (averaging 3.2–3.8), while Nigeria recorded the lowest (averaging 2.7–3.5). Crucially, the Aesthetic Score demonstrates a strong co...