Pith. sign in

REVIEW 5 cited by

TokenCompose: Text-to-Image Diffusion with Token-level Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03626 v2 pith:5QWGC4CG submitted 2023-12-06 cs.CV

classification cs.CV
keywords diffusiontokencomposeconsistencymodelpromptstextcompositionenhanced
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standard denoising process in the Latent Diffusion Model takes text prompts as conditions only, absent explicit constraint for the consistency between the text prompts and the image contents, leading to unsatisfactory results for composing multiple object categories. TokenCompose aims to improve multi-category instance composition by introducing the token-wise consistency terms between the image content and object segmentation maps in the finetuning stage. TokenCompose can be applied directly to the existing training pipeline of text-conditioned diffusion models without extra human labeling information. By finetuning Stable Diffusion, the model exhibits significant improvements in multi-category instance composition and enhanced photorealism for its generated images. Project link: https://mlpc-ucsd.github.io/TokenCompose

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Aligning cross-attention map similarity to text self-attention maps at test time improves semantic alignment in Stable Diffusion for prompts with multiple objects and attributes.

  2. LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers

    cs.LG 2024-12 conditional novelty 5.0 of 10

    LazyDiT learns small gates that decide when to reuse cached layer outputs, cutting diffusion transformer compute by up to half while matching or beating DDIM quality.

  3. VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VersaGen enables users to control text-to-image diffusion models with partial sketch inputs at object and scene levels, with automatic localization and adaptive control strength.

  4. High-Order Matching for One-Step Shortcut Diffusion Models

    cs.CV 2025-02 reject novelty 4.0 of 10

    HOMO extends shortcut diffusion with acceleration and jerk supervision, but the proof of superior approximation is not supported and experiments lack error bars.

  5. Text-to-Image Synthesis: A Decade Survey

    cs.CV 2024-11 conditional novelty 1.0 of 10

    A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.

Pith tools