TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation

Shintaro Ozaki , Tomoyuki Jinno , Kazuki Hayashi , Yusuke Sakai , Jingun Kwon , Hidetaka Kamigaito , Katsuhiko Hayashi , Manabu Okumura

show 1 more author

Taro Watanabe

Authors on Pith no claims yet

classification 💻 cs.CL cs.CV

keywords generationentitiestexttigerdescriptionsimageperformancepromptprompts

0 comments

read the original abstract

When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is almost countless, and new entities emerge; memorizing all of them completely is not realistic. To bridge this gap, our work proposes Text-based Intelligent Generation with Entity Prompt Refinement (TextTIGER). TextTIGER strengthens knowledge about entities that appear in the prompt by augmenting external information and then summarizes the expanded descriptions with large language models, preventing performance degradation that arises from excessively long inputs. To evaluate our method, we construct a new dataset consisting of captions, images, detailed descriptions, and lists of entities. Experiments with multiple image generation models show that TextTIGER improves image generation performance on widely used evaluation metrics compared with prompts that use captions alone. In addition, using Multimodal LLM (MLLM)-as-a-judge, which shows a strong correlation with human evaluation, we demonstrate that our method consistently achieves higher scores, which underscores its effectiveness. These results show that strengthening entity-related descriptions, summarizing them, and refining prompts to an appropriate length leads to substantial improvements in image generation performance. We will release the created dataset and code upon acceptance.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

SCMAPR: Self-Correcting Multi-Agent Prompt Refinement for Complex-Scenario Text-to-Video Generation
cs.AI 2026-04 unverdicted novelty 6.0

SCMAPR is a self-correcting multi-agent prompt refinement framework that boosts text-to-video alignment and quality in complex scenarios, with reported gains on VBench, EvalCrafter, and a new T2V-Complexity benchmark.