Pith. sign in

REVIEW 2 cited by

Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.13858 v1 pith:GLIYVMCS submitted 2024-08-25 cs.CV cs.LG

classification cs.CVcs.LG
keywords complexdiffusionpaintingretouchingscenecompositiondefinitiongeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-to-image diffusion models have demonstrated impressive capabilities in image quality. However, complex scene generation remains relatively unexplored, and even the definition of `complex scene' itself remains unclear. In this paper, we address this gap by providing a precise definition of complex scenes and introducing a set of Complex Decomposition Criteria (CDC) based on this definition. Inspired by the artists painting process, we propose a training-free diffusion framework called Complex Diffusion (CxD), which divides the process into three stages: composition, painting, and retouching. Our method leverages the powerful chain-of-thought capabilities of large language models (LLMs) to decompose complex prompts based on CDC and to manage composition and layout. We then develop an attention modulation method that guides simple prompts to specific regions to complete the complex scene painting. Finally, we inject the detailed output of the LLM into a retouching model to enhance the image details, thus implementing the retouching stage. Extensive experiments demonstrate that our method outperforms previous SOTA approaches, significantly improving the generation of high-quality, semantically consistent, and visually diverse images for complex scenes, even with intricate prompts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Integrating Large Language Models into Text Animation: An Intelligent Editing System with Inline and Chat Interaction

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A text-animation editor with inline and chat LLM agents was rated usable (SUS 75) by 11 non-professional testers.

  2. CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step

    cs.CV 2025-07 conditional novelty 4.0 of 10

    CoT-Diff couples a multimodal LLM's step-by-step 3D layout reasoning into the diffusion denoising loop, claiming large gains in spatial alignment for text-to-image generation.

Pith tools