A new benchmark with 5,000 varied prompts and a VLM-based fine-grained evaluation protocol aims to measure and rank text-to-image models' instruction-following ability.
PixArt- Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation, March 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
A new benchmark with 5,000 varied prompts and a VLM-based fine-grained evaluation protocol aims to measure and rank text-to-image models' instruction-following ability.