REVIEW 6 cited by
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particularly in human image generation. Sociological research has established systematic classifications of bias. Yet, existing studies on bias in T2I models largely conflate different types of bias, impeding methodological progress. In this paper, we introduce BIGbench, a unified benchmark for Biases of Image Generation, featuring a carefully designed dataset. Unlike existing benchmarks, BIGbench classifies and evaluates biases across four dimensions to enable a more granular evaluation and deeper analysis. Furthermore, BIGbench applies advanced multi-modal large language models to achieve fully automated and highly accurate evaluations. We apply BIGbench to evaluate eight representative T2I models and three debiasing methods. Our human evaluation results by trained evaluators from different races underscore BIGbench's effectiveness in aligning images and identifying various biases. Moreover, our study also reveals new research directions about biases with insightful analysis of our results. Our work is openly accessible at https://github.com/BIGbench2024/BIGbench2024/.
Forward citations
Cited by 6 Pith papers
-
The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs
Across 12 checkpoints of Yi, Qwen, Mistral and Gemma, mean absolute adjacent-generation trust-score drift is 8.00 pp—3.6× an independence-based no-drift null—and remains elevated under leave-one-out and strict-scoring checks.
-
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.
-
Multi-Group Proportional Representation for Text-to-Image Models
The authors apply the MPR metric (an integral probability metric) to text-to-image generation, derive tractable forms for linear and decision-tree function classes, and use it as a fine-tuning objective that reduces i...
-
Evaluating the Sensitivity of LLMs to Prior Context
Prior conversational context, especially from a different knowledge domain, can sharply reduce LLM multiple-choice accuracy, and repeating the task near the query mitigates the drop.
-
Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings
In switched MDD follow-up data, causal forest outperformed seven baselines at next-visit counterfactual prediction, and adjusted treatment effects were modest (about 0.3–0.9 HAMD-17 points).
-
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.
Discussion (0). Continue with ORCID to comment.