Pith. sign in

REVIEW 12 cited by

Improving Text-to-Image Consistency via Automatic Prompt Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17804 v1 pith:UDTQA4TG submitted 2024-03-26 cs.CV cs.CL

classification cs.CVcs.CL
keywords consistencymodelspromptprompt-imagescoretheychallengesframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Impressive advances in text-to-image (T2I) generative models have yielded a plethora of high performing models which are able to generate aesthetically appealing, photorealistic images. Despite the progress, these models still struggle to produce images that are consistent with the input prompt, oftentimes failing to capture object quantities, relations and attributes properly. Existing solutions to improve prompt-image consistency suffer from the following challenges: (1) they oftentimes require model fine-tuning, (2) they only focus on nearby prompt samples, and (3) they are affected by unfavorable trade-offs among image quality, representation diversity, and prompt-image consistency. In this paper, we address these challenges and introduce a T2I optimization-by-prompting framework, OPT2I, which leverages a large language model (LLM) to improve prompt-image consistency in T2I models. Our framework starts from a user prompt and iteratively generates revised prompts with the goal of maximizing a consistency score. Our extensive validation on two datasets, MSCOCO and PartiPrompts, shows that OPT2I can boost the initial consistency score by up to 24.9% in terms of DSG score while preserving the FID and increasing the recall between generated and real data. Our work paves the way toward building more reliable and robust T2I systems by harnessing the power of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DuET: Dual Expert Trajectories for Diffusion Image Editing

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Switching a diffusion editor from image-conditioned to caption-only mode for a mid-trajectory interval and back improves edit fidelity and naturalness on FLUX2-Klein and BAGEL, while predictably reducing source-image ...

  2. Visual Persuasion: What Influences Decisions of Vision-Language Models?

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Naturalistic edits to background, lighting, and staging systematically shift the choices of frontier vision-language models, and a new optimization-plus-interpretation framework (CVPO) discovers and explains the visua...

  3. Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A multi-agent prompt-refinement system using pairwise AI judging and targeted edit signals outperforms prior automated methods on complex text-to-image tasks.

  4. Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A noise hypernetwork learns a reward-tilted initial noise distribution for frozen distilled diffusion generators, recovering roughly half of test-time noise-optimization gains at a fraction of the compute.

  5. WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free method, WAVE, improves multi-view consistency in single-image novel view synthesis by using 3D-warped views to guide diffusion attention and initial noise.

  6. ComfyUI-R1: Exploring Reasoning Models for Workflow Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 7B reasoning model trained with supervised fine-tuning and reinforcement learning generates ComfyUI workflows from text instructions and outperforms GPT-4o and Claude-based baselines on the authors' tests.

  7. RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RePrompt uses RL-trained reasoning traces to enhance text-to-image prompts, boosting spatial composition and counting scores across FLUX, SD3, and PixArt-Σ while keeping image generators fixed.

  8. PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single VLM learns reusable T2I prompt rewriting by diagnosing its own generated images and optimizing with a hybrid ideal-point/Chebyshev multi-objective reward.

  9. GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design

    cs.HC 2025-08 conditional novelty 5.0 of 10

    GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.

  10. ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An LLM-powered multi-agent Copilot retrieves and constructs ComfyUI workflows, reporting at least 88.5% recall on its own test set and 85.9% online acceptance of proposed workflows.

  11. LLMs can see and hear without any training

    cs.CV 2025-01 conditional novelty 5.0 of 10

    MILS uses an LLM plus a pretrained scorer in an iterative generate-score-refine loop to produce competitive zero-shot captions and improved text-to-image prompts without training.

  12. Seamless and Efficient Interactions within a Mixed-Dimensional Information Space

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A thesis that three design strategies, multimodal AI, context-aware placement, and combined 2D/3D views, make mixed-dimensional information spaces seamless and efficient, demonstrated with three systems.

Pith tools