Pith. sign in

REVIEW 5 cited by

ImagenHub: Standardizing the evaluation of conditional image generation models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01596 v4 pith:NZB3PPXU submitted 2023-10-02 cs.CV cs.GRcs.MM

classification cs.CVcs.GRcs.MM
keywords generationimagemodelsevaluationconditionalevaluateinferencemetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, a myriad of conditional image generation and editing models have been developed to serve different downstream tasks, including text-to-image generation, text-guided image editing, subject-driven image generation, control-guided image generation, etc. However, we observe huge inconsistencies in experimental conditions: datasets, inference, and evaluation metrics - render fair comparisons difficult. This paper proposes ImagenHub, which is a one-stop library to standardize the inference and evaluation of all the conditional image generation models. Firstly, we define seven prominent tasks and curate high-quality evaluation datasets for them. Secondly, we built a unified inference pipeline to ensure fair comparison. Thirdly, we design two human evaluation scores, i.e. Semantic Consistency and Perceptual Quality, along with comprehensive guidelines to evaluate generated images. We train expert raters to evaluate the model outputs based on the proposed metrics. Our human evaluation achieves a high inter-worker agreement of Krippendorff's alpha on 76% models with a value higher than 0.4. We comprehensively evaluated a total of around 30 models and observed three key takeaways: (1) the existing models' performance is generally unsatisfying except for Text-guided Image Generation and Subject-driven Image Generation, with 74% models achieving an overall score lower than 0.5. (2) we examined the claims from published papers and found 83% of them hold with a few exceptions. (3) None of the existing automatic metrics has a Spearman's correlation higher than 0.2 except subject-driven image generation. Moving forward, we will continue our efforts to evaluate newly published models and update our leaderboard to keep track of the progress in conditional image generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IDEA-Bench: How Far are Generative Models from Professional Designing?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    IDEA-Bench measures generative models on 100 professional design tasks and finds the best tested system scores only 22.48 out of 100.

  2. Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Inpainting unusual objects into street scenes reveals that open-vocabulary detectors miss objects based on image location rather than object semantics.

  3. MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image encoder maps any face to a 'name embedding' that, when prepended to a text prompt, makes an SDXL model generate consistent identities for arbitrary people without fine-tuning.

  4. Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.

  5. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

Pith tools