REVIEW 12 cited by
Improving Text-to-Image Consistency via Automatic Prompt Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Impressive advances in text-to-image (T2I) generative models have yielded a plethora of high performing models which are able to generate aesthetically appealing, photorealistic images. Despite the progress, these models still struggle to produce images that are consistent with the input prompt, oftentimes failing to capture object quantities, relations and attributes properly. Existing solutions to improve prompt-image consistency suffer from the following challenges: (1) they oftentimes require model fine-tuning, (2) they only focus on nearby prompt samples, and (3) they are affected by unfavorable trade-offs among image quality, representation diversity, and prompt-image consistency. In this paper, we address these challenges and introduce a T2I optimization-by-prompting framework, OPT2I, which leverages a large language model (LLM) to improve prompt-image consistency in T2I models. Our framework starts from a user prompt and iteratively generates revised prompts with the goal of maximizing a consistency score. Our extensive validation on two datasets, MSCOCO and PartiPrompts, shows that OPT2I can boost the initial consistency score by up to 24.9% in terms of DSG score while preserving the FID and increasing the recall between generated and real data. Our work paves the way toward building more reliable and robust T2I systems by harnessing the power of LLMs.
Forward citations
Cited by 12 Pith papers
-
DuET: Dual Expert Trajectories for Diffusion Image Editing
Switching a diffusion editor from image-conditioned to caption-only mode for a mid-trajectory interval and back improves edit fidelity and naturalness on FLUX2-Klein and BAGEL, while predictably reducing source-image ...
-
Visual Persuasion: What Influences Decisions of Vision-Language Models?
Naturalistic edits to background, lighting, and staging systematically shift the choices of frontier vision-language models, and a new optimization-plus-interpretation framework (CVPO) discovers and explains the visua...
-
Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
A multi-agent prompt-refinement system using pairwise AI judging and targeted edit signals outperforms prior automated methods on complex text-to-image tasks.
-
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
A noise hypernetwork learns a reward-tilted initial noise distribution for frozen distilled diffusion generators, recovering roughly half of test-time noise-optimization gains at a fraction of the compute.
-
WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image
A training-free method, WAVE, improves multi-view consistency in single-image novel view synthesis by using 3D-warped views to guide diffusion attention and initial noise.
-
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
A 7B reasoning model trained with supervised fine-tuning and reinforcement learning generates ComfyUI workflows from text instructions and outperforms GPT-4o and Claude-based baselines on the authors' tests.
-
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
RePrompt uses RL-trained reasoning traces to enhance text-to-image prompts, boosting spatial composition and counting scores across FLUX, SD3, and PixArt-Σ while keeping image generators fixed.
-
PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
A single VLM learns reusable T2I prompt rewriting by diagnosing its own generated images and optimizing with a hybrid ideal-point/Chebyshev multi-objective reward.
-
GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.
-
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
An LLM-powered multi-agent Copilot retrieves and constructs ComfyUI workflows, reporting at least 88.5% recall on its own test set and 85.9% online acceptance of proposed workflows.
-
LLMs can see and hear without any training
MILS uses an LLM plus a pretrained scorer in an iterative generate-score-refine loop to produce competitive zero-shot captions and improved text-to-image prompts without training.
-
Seamless and Efficient Interactions within a Mixed-Dimensional Information Space
A thesis that three design strategies, multimodal AI, context-aware placement, and combined 2D/3D views, make mixed-dimensional information spaces seamless and efficient, demonstrated with three systems.
Discussion (0). Continue with ORCID to comment.