A new benchmark with 5,000 varied prompts and a VLM-based fine-grained evaluation protocol aims to measure and rank text-to-image models' instruction-following ability.
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis, December 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
A new benchmark with 5,000 varied prompts and a VLM-based fine-grained evaluation protocol aims to measure and rank text-to-image models' instruction-following ability.