Pith. sign in

REVIEW 8 cited by

TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11897 v3 pith:PGNSOGP2 submitted 2023-03-21 cs.CV

classification cs.CV
keywords text-to-imagetifamodelsevaluationfaithfulnesstextansweringexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite thousands of researchers, engineers, and artists actively working on improving text-to-image generation models, systems often fail to produce images that accurately align with the text inputs. We introduce TIFA (Text-to-Image Faithfulness evaluation with question Answering), an automatic evaluation metric that measures the faithfulness of a generated image to its text input via visual question answering (VQA). Specifically, given a text input, we automatically generate several question-answer pairs using a language model. We calculate image faithfulness by checking whether existing VQA models can answer these questions using the generated image. TIFA is a reference-free metric that allows for fine-grained and interpretable evaluations of generated images. TIFA also has better correlations with human judgments than existing metrics. Based on this approach, we introduce TIFA v1.0, a benchmark consisting of 4K diverse text inputs and 25K questions across 12 categories (object, counting, etc.). We present a comprehensive evaluation of existing text-to-image models using TIFA v1.0 and highlight the limitations and challenges of current models. For instance, we find that current text-to-image models, despite doing well on color and material, still struggle in counting, spatial relations, and composing multiple objects. We hope our benchmark will help carefully measure the research progress in text-to-image synthesis and provide valuable insights for further research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Poplar-9K is a curated human-centric image dataset generated through a reproducible attribute sampling, diffusion rendering, and vision-language inspection pipeline, with 9,401 accepted pairs from 11,765 candidates.

  2. Discovering Divergent Representations between Text-to-Image Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An evolutionary algorithm discovers visual attributes that appear in one text-to-image model's outputs but not another's, and identifies the prompt concepts that trigger them.

  3. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  4. Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    With a carefully selected noise schedule, diffusion models trained with as few as 32 latent states, or composed from single-state models, match 1,000-state training and converge 4-6x faster.

  5. DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new evaluation framework, DIMCIM, measures default-mode diversity and prompted generalization in text-to-image models, finding a scale trade-off and a 0.85 correlation between default diversity and training data diversity.

  6. A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Reference-free caption quality is scored by the downstream vision-language accuracy of a caption-conditioned reconstructed image, via a new CTTD benchmark.

  7. AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.

  8. Towards Evaluating Robustness of Prompt Adherence in Text to Image Models

    cs.CV 2025-07 conditional novelty 4.0 of 10

    New benchmark results show that Stable Diffusion 3.x and Janus Pro models struggle to place simple geometric shapes in the correct image quadrant, with best F1 scores around 0.41 to 0.5.

Pith tools