REVIEW 6 cited by
Consistency-diversity-realism Pareto fronts of conditional image generative models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Building world models that accurately and comprehensively represent the real world is the utmost aspiration for conditional image generative models as it would enable their use as world simulators. For these models to be successful world models, they should not only excel at image quality and prompt-image consistency but also ensure high representation diversity. However, current research in generative models mostly focuses on creative applications that are predominantly concerned with human preferences of image quality and aesthetics. We note that generative models have inference time mechanisms - or knobs - that allow the control of generation consistency, quality, and diversity. In this paper, we use state-of-the-art text-to-image and image-and-text-to-image models and their knobs to draw consistency-diversity-realism Pareto fronts that provide a holistic view on consistency-diversity-realism multi-objective. Our experiments suggest that realism and consistency can both be improved simultaneously; however there exists a clear tradeoff between realism/consistency and diversity. By looking at Pareto optimal points, we note that earlier models are better at representation diversity and worse in consistency/realism, and more recent models excel in consistency/realism while decreasing significantly the representation diversity. By computing Pareto fronts on a geodiverse dataset, we find that the first version of latent diffusion models tends to perform better than more recent models in all axes of evaluation, and there exist pronounced consistency-diversity-realism disparities between geographical regions. Overall, our analysis clearly shows that there is no best model and the choice of model should be determined by the downstream application. With this analysis, we invite the research community to consider Pareto fronts as an analytical tool to measure progress towards world models.
Forward citations
Cited by 6 Pith papers
-
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
GASS enhances fixed-prompt diversity in T2I models by expanding CLIP embedding spread along the text direction and a computed orthogonal background direction.
-
Scaling Group Inference for Diverse and High-Quality Generation
Groups of generated images become more diverse while staying high-quality when K outputs are chosen from M candidates via a quadratic integer program with progressive pruning.
-
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
IConMark adds preselected, human-readable objects to AI images via prompt engineering and detects them with a vision-language model, achieving higher AUROC than noise-based watermarks on tested augmentations.
-
Adultification Bias in LLMs and Text-to-Image Models
Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.
-
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
A per-category best-of-K RL reward, multi-axis max@K, shifts SD3.5-M perceived-appearance distributions toward uniform coverage (Fairness Score +0.23 to +0.36) without quality loss.
-
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
EMAG replaces selected attention maps with their exponential moving average during diffusion sampling, reporting +0.46 HPS over CFG on SD3 and composing with APG/CADS.
Discussion (0). Continue with ORCID to comment.