Pith. sign in

REVIEW 9 cited by

Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16820 v4 pith:CXO43YPF submitted 2024-04-25 cs.CV

classification cs.CV
keywords humantemplatesmetricsmodelspromptratingsacrossarise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While previous work has evaluated T2I alignment by proposing metrics, benchmarks, and templates for collecting human judgements, the quality of these components is not systematically measured. Human-rated prompt sets are generally small and the reliability of the ratings -- and thereby the prompt set used to compare models -- is not evaluated. We address this gap by performing an extensive study evaluating auto-eval metrics and human templates. We provide three main contributions: (1) We introduce a comprehensive skills-based benchmark that can discriminate models across different human templates. This skills-based benchmark categorises prompts into sub-skills, allowing a practitioner to pinpoint not only which skills are challenging, but at what level of complexity a skill becomes challenging. (2) We gather human ratings across four templates and four T2I models for a total of >100K annotations. This allows us to understand where differences arise due to inherent ambiguity in the prompt and where they arise due to differences in metric and model quality. (3) Finally, we introduce a new QA-based auto-eval metric that is better correlated with human ratings than existing metrics for our new dataset, across different human templates, and on TIFA160.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Selective search plus generator-reasoner co-training improves knowledge-grounded image generation, but the reported gains are scored by the same VLM judge used to train the system.

  2. A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Layer-normalized averaging of all decoder-only LLM hidden states, rather than last-layer embeddings, improves text-to-image compositional alignment and beats T5 on GenAI-Bench.

  3. OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OneIG-Bench introduces a 2,440-prompt, six-dimension benchmark with automated metrics for text-to-image models, covering alignment, text, reasoning, style, and diversity in English and Chinese.

  4. MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MMIG-Bench is a unified benchmark of 4,850 prompts and 1,750 reference images with a three-level evaluation suite, including the VQA-based Aspect Matching Score that correlates with human ratings.

  5. Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.

  6. DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Introduces DetailMaster, a 4,116-prompt benchmark with fine-grained evaluation of long-prompt text-to-image generation, finding that state-of-the-art models achieve only about 50% accuracy on attribute binding and spa...

  7. A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A survey and framework that categorizes generative model unlearning by point-wise versus concept-wise objectives, parameter-based versus non-parametric methods, and completeness/utility/efficiency evaluation.

  8. Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Current image-text alignment metrics, including CLIPScore and DSGScore, produce unstable model rankings under random seeds and are highly sensitive to tiny image perturbations.

  9. NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

    cs.CV 2025-05 conditional novelty 4.0 of 10

    The NTIRE 2025 challenge report compares 20 methods for fine-grained text-to-image quality assessment, introduces the EvalMuse-Structure dataset, and finds every participating team outperformed the baselines.

Pith tools