Pith. sign in

REVIEW 6 cited by

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11708 v3 pith:FTXZOSBU submitted 2024-01-22 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffusiongenerationtext-to-imageeditingmodelsmultipleabilitycomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have exhibit exceptional performance in text-to-image generation and editing. However, existing methods often face challenges when handling complex text prompts that involve multiple objects with multiple attributes and relationships. In this paper, we propose a brand new training-free text-to-image generation/editing framework, namely Recaption, Plan and Generate (RPG), harnessing the powerful chain-of-thought reasoning ability of multimodal LLMs to enhance the compositionality of text-to-image diffusion models. Our approach employs the MLLM as a global planner to decompose the process of generating complex images into multiple simpler generation tasks within subregions. We propose complementary regional diffusion to enable region-wise compositional generation. Furthermore, we integrate text-guided image generation and editing within the proposed RPG in a closed-loop fashion, thereby enhancing generalization ability. Extensive experiments demonstrate our RPG outperforms state-of-the-art text-to-image diffusion models, including DALL-E 3 and SDXL, particularly in multi-category object composition and text-image semantic alignment. Notably, our RPG framework exhibits wide compatibility with various MLLM architectures (e.g., MiniGPT-4) and diffusion backbones (e.g., ControlNet). Our code is available at: https://github.com/YangLing0818/RPG-DiffusionMaster

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

    cs.CV 2025-05 conditional novelty 7.0 of 10

    PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.

  4. TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

    cs.AI 2026-05 conditional novelty 5.0 of 10

    A training-free, test-time guidance rule that tilts a diffusion model's samples toward regions where every concept in a prompt is jointly present; it improves several T2ICompBench categories over prior correctors and ...

  5. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  6. Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

    cs.CV 2025-05 reject novelty 4.0 of 10

    Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.

Pith tools