Pith. sign in

REVIEW 2 cited by

Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08877 v3 pith:7IJETWBZ submitted 2023-06-15 cs.CL cs.CV

classification cs.CLcs.CV
keywords entitiesbindingimagelinguisticmodifiersattentionduringflamingo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-conditioned image generation models often generate incorrect associations between entities and their visual attributes. This reflects an impaired mapping between linguistic binding of entities and modifiers in the prompt and visual binding of the corresponding elements in the generated image. As one notable example, a query like "a pink sunflower and a yellow flamingo" may incorrectly produce an image of a yellow sunflower and a pink flamingo. To remedy this issue, we propose SynGen, an approach which first syntactically analyses the prompt to identify entities and their modifiers, and then uses a novel loss function that encourages the cross-attention maps to agree with the linguistic binding reflected by the syntax. Specifically, we encourage large overlap between attention maps of entities and their modifiers, and small overlap with other entities and modifier words. The loss is optimized during inference, without retraining or fine-tuning the model. Human evaluation on three datasets, including one new and challenging set, demonstrate significant improvements of SynGen compared with current state of the art methods. This work highlights how making use of sentence structure during inference can efficiently and substantially improve the faithfulness of text-to-image generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.

  2. Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Padding tokens in text-to-image models can carry semantic information or act as diffusion-time registers, depending on training and attention architecture.

Pith tools