REVIEW 20 cited by
Composer: Creative and Controllable Image Synthesis with Composable Conditions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability. This work offers a new generation paradigm that allows flexible control of the output image, such as spatial layout and palette, while maintaining the synthesis quality and model creativity. With compositionality as the core idea, we first decompose an image into representative factors, and then train a diffusion model with all these factors as the conditions to recompose the input. At the inference stage, the rich intermediate representations work as composable elements, leading to a huge design space (i.e., exponentially proportional to the number of decomposed factors) for customizable content creation. It is noteworthy that our approach, which we call Composer, supports various levels of conditions, such as text description as the global information, depth map and sketch as the local guidance, color histogram for low-level details, etc. Besides improving controllability, we confirm that Composer serves as a general framework and facilitates a wide range of classical generative tasks without retraining. Code and models will be made available.
Forward citations
Cited by 20 Pith papers
-
GenSpace: Benchmarking Spatially-Aware Image Generation
GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.
-
Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling
Next-dense-stride prediction enables coarse-to-fine autoregressive image generation on a single-scale grid and unifies multi-contrast MRI translation, generation, and segmentation in one model.
-
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.
-
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement
A closed-loop, training-free controller uses CLIP similarity feedback and bidirectional IP-Adapter scales to keep rare attributes and base objects balanced throughout the diffusion trajectory.
-
DanceOPD: On-Policy Generative Field Distillation
Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.
-
ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
ReSAGE-PAR adapts diffusion models with LoRA, scores generated images via vision-language prompts, and applies Bayesian classification to produce pseudo-labels, yielding up to 8.7% gains when used to expand PAR datasets.
-
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.
-
Palette Aligned Image Diffusion
Palette-Adapter conditions text-to-image diffusion on a sparse color palette treated as a histogram, with entropy and distance controls and a negative-color guidance mechanism.
-
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
CompSlider learns to synthesize image-conditioning latents from multiple attribute sliders at once, aiming for more disentangled and structure-preserving multi-attribute control in text-to-image generation.
-
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
A training-free extension of InteractDiffusion that uses LLM-mined relations, action feature offsets, and entity masks to improve entity and interaction control in generated images.
-
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
Fine-tuning LVLMs on original/altered image pairs with targeted visual instructions reduces object and attribute hallucinations on POPE, LLaVA-Bench, and MMHal-Bench.
-
Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints
CP-Composer trains a geometric diffusion model on linear peptides and imposes cyclization constraints at generation time, achieving 38-84% constraint satisfaction across four cyclization strategies.
-
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.
-
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
UniGlyph replaces pre-rendered glyph conditions with segmentation-derived masks in a ControlNet diffusion model, reporting gains on visual text rendering benchmarks.
-
DreamLight: Towards Harmonious and Consistent Image Relighting
A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.
-
Image Editing As Programs with Diffusion Models
IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.
-
ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer
An anatomy- and clinical-conditioned diffusion model that synthesizes post-treatment CT and uses its features to improve immunotherapy response prediction in NSCLC.
-
Compositional Scene Understanding through Inverse Generative Modeling
Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.
-
Introspective Attention Modulation for Safe Text-to-Image Generation
Inference-time attention modulation suppresses unsafe content in diffusion-transformer T2I models without retraining and beats concept-erasure baselines in the paper's benchmarks.
-
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
A multimodal-LLM planner, polygon-based layout masks, and SVD-based structure injection combine to improve spatial and textual control of pre-trained text-to-image diffusion models.
Discussion (0). Sign in to comment.