Pith. sign in

REVIEW 2 cited by

Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16990 v1 pith:TXTJP6RJ submitted 2024-03-25 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords subjectsattentionboundedgenerationleakagemultiplecomplexdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introduced to improve user control, aiming to localize subjects represented by specific tokens. Yet, these methods often produce semantically inaccurate images, especially when dealing with multiple semantically or visually similar subjects. In this work, we study and analyze the causes of these limitations. Our exploration reveals that the primary issue stems from inadvertent semantic leakage between subjects in the denoising process. This leakage is attributed to the diffusion model's attention layers, which tend to blend the visual features of different subjects. To address these issues, we introduce Bounded Attention, a training-free method for bounding the information flow in the sampling process. Bounded Attention prevents detrimental leakage among subjects and enables guiding the generation to promote each subject's individuality, even with complex multi-subject conditioning. Through extensive experimentation, we demonstrate that our method empowers the generation of multiple subjects that better align with given prompts and layouts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.

  2. Multitwine: Multi-Object Compositing with Text and Layout Control

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single diffusion model simultaneously composites multiple objects into a scene with text and layout control, outperforming sequential insertion on interacting cases.

Pith tools