Pith. sign in

REVIEW 20 cited by

Composer: Creative and Controllable Image Synthesis with Composable Conditions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.09778 v2 pith:YQHMN7NL submitted 2023-02-20 cs.CV cs.GR

classification cs.CVcs.GR
keywords composerconditionsfactorsimagecomposablecontrollabilitygenerativemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability. This work offers a new generation paradigm that allows flexible control of the output image, such as spatial layout and palette, while maintaining the synthesis quality and model creativity. With compositionality as the core idea, we first decompose an image into representative factors, and then train a diffusion model with all these factors as the conditions to recompose the input. At the inference stage, the rich intermediate representations work as composable elements, leading to a huge design space (i.e., exponentially proportional to the number of decomposed factors) for customizable content creation. It is noteworthy that our approach, which we call Composer, supports various levels of conditions, such as text description as the global information, depth map and sketch as the local guidance, color histogram for low-level details, etc. Besides improving controllability, we confirm that Composer serves as a general framework and facilitates a wide range of classical generative tasks without retraining. Code and models will be made available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenSpace: Benchmarking Spatially-Aware Image Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.

  2. Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling

    eess.IV 2026-07 conditional novelty 6.5 of 10

    Next-dense-stride prediction enables coarse-to-fine autoregressive image generation on a single-scale grid and unifies multi-contrast MRI translation, generation, and segmentation in one model.

  3. WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.

  4. RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A closed-loop, training-free controller uses CLIP similarity feedback and bidirectional IP-Adapter scales to keep rare attributes and base objects balanced throughout the diffusion trajectory.

  5. DanceOPD: On-Policy Generative Field Distillation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.

  6. ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.

  7. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.

  8. Palette Aligned Image Diffusion

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Palette-Adapter conditions text-to-image diffusion on a sparse color palette treated as a histogram, with entropy and distance controls and a negative-color guidance mechanism.

  9. CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CompSlider learns to synthesize image-conditioning latents from multiple attribute sliders at once, aiming for more disentangled and structure-preserving multi-attribute control in text-to-image generation.

  10. CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free extension of InteractDiffusion that uses LLM-mined relations, action feature offsets, and entity masks to improve entity and interaction control in generated images.

  11. See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fine-tuning LVLMs on original/altered image pairs with targeted visual instructions reduces object and attribute hallucinations on POPE, LLaVA-Bench, and MMHal-Bench.

  12. Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints

    cs.LG 2025-07 conditional novelty 6.0 of 10

    CP-Composer trains a geometric diffusion model on linear peptides and imposes cyclization constraints at generation time, achieving 38-84% constraint satisfaction across four cyclization strategies.

  13. Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback

    cs.CV 2025-07 conditional novelty 6.0 of 10

    InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.

  14. UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    UniGlyph replaces pre-rendered glyph conditions with segmentation-derived masks in a ControlNet diffusion model, reporting gains on visual text rendering benchmarks.

  15. DreamLight: Towards Harmonious and Consistent Image Relighting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.

  16. Image Editing As Programs with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.

  17. ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer

    eess.IV 2025-05 conditional novelty 6.0 of 10

    An anatomy- and clinical-conditioned diffusion model that synthesizes post-treatment CT and uses its features to improve immunotherapy response prediction in NSCLC.

  18. Compositional Scene Understanding through Inverse Generative Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.

  19. Introspective Attention Modulation for Safe Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Inference-time attention modulation suppresses unsafe content in diffusion-transformer T2I models without retraining and beats concept-erasure baselines in the paper's benchmarks.

  20. LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A multimodal-LLM planner, polygon-based layout masks, and SVD-based structure injection combine to improve spatial and textual control of pre-trained text-to-image diffusion models.

Pith tools