Pith. sign in

REVIEW 6 cited by

BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15240 v6 pith:3FBZODQS submitted 2024-07-21 cs.CV

classification cs.CV
keywords biasesbigbenchmodelsbiasanalysisbenchmarkbigbench2024different
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particularly in human image generation. Sociological research has established systematic classifications of bias. Yet, existing studies on bias in T2I models largely conflate different types of bias, impeding methodological progress. In this paper, we introduce BIGbench, a unified benchmark for Biases of Image Generation, featuring a carefully designed dataset. Unlike existing benchmarks, BIGbench classifies and evaluates biases across four dimensions to enable a more granular evaluation and deeper analysis. Furthermore, BIGbench applies advanced multi-modal large language models to achieve fully automated and highly accurate evaluations. We apply BIGbench to evaluate eight representative T2I models and three debiasing methods. Our human evaluation results by trained evaluators from different races underscore BIGbench's effectiveness in aligning images and identifying various biases. Moreover, our study also reveals new research directions about biases with insightful analysis of our results. Our work is openly accessible at https://github.com/BIGbench2024/BIGbench2024/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

    cs.SE 2026-07 conditional novelty 6.5 of 10

    Across 12 checkpoints of Yi, Qwen, Mistral and Gemma, mean absolute adjacent-generation trust-score drift is 8.00 pp—3.6× an independence-based no-drift null—and remains elevated under leave-one-out and strict-scoring checks.

  2. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

  3. Multi-Group Proportional Representation for Text-to-Image Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors apply the MPR metric (an integral probability metric) to text-to-image generation, derive tractable forms for linear and decision-tree function classes, and use it as a fine-tuning objective that reduces i...

  4. Evaluating the Sensitivity of LLMs to Prior Context

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Prior conversational context, especially from a different knowledge domain, can sharply reduce LLM multiple-choice accuracy, and repeating the task near the query mitigates the drop.

  5. Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings

    stat.AP 2026-06 conditional novelty 5.0 of 10

    In switched MDD follow-up data, causal forest outperformed seven baselines at next-visit counterfactual prediction, and adjusted treatment effects were modest (about 0.3–0.9 HAMD-17 points).

  6. BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining

    cs.CL 2025-06 reject novelty 3.0 of 10

    A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.

Pith tools