REVIEW 3 major objections 5 minor 85 references
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read CoMPaSS claims that spatial failures in text-to-image diffusion models are fixable by pairing curated spatial data with token-order reinjection, achieving up to +131% relative gains on GenEval Position across four open-weight models.
desk verdict A well-executed recipe for improving spatial compliance on COCO-style benchmarks, but the headline gains track the training distribution more than they prove a general spatial understanding. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two components carry the argument. SCOP (Spatial Constraints-Oriented Pairing) is a data engine that enumerates object pairs in an image and keeps only those passing five geometric constraints—visual significance, semantic distinction, spatial clarity, minimal overlap, and size balance—then decodes the surviving pairs into image crops paired with templated spatial captions. TENOR (Token ENcoding ORdering) is a parameter-free module that adds sinusoidal positional encodings to the key vectors in UNet cross-attention and to the text query/key vectors in MMDiT blocks, making token order visible at every attention step so structurally different prompts produce different conditioning signals.
What would settle it
Evaluate CoMPaSS against its base models on a held-out benchmark built from non-COCO object categories (for example, 'a wrench to the left of a screwdriver') or object-centric spatial language ('to the child's right hand side'); if accuracy returns to baseline levels, the gains are distribution matching and the claim of enhanced general spatial understanding fails.
Extended reading notes
Core claim
CoMPaSS establishes that injecting token-order information into the text-image attention of diffusion models, in combination with training on a small set of spatially unambiguous image-text pairs, makes both UNet-based and MMDiT-based text-to-image models substantially better at rendering left, right, above, and below relations. The paper's central claim is that the two interventions are complementary: SCOP supplies clean spatial supervision that was missing from web-scale training data, and TENOR provides the structural signal that lets the model tell 'A left of B' from 'B left of A', which standard encoders fail to preserve. On FLUX.1 the combination lifts VISOR from 37.96 to 75.17, T2I-CompBench Spatial from 0.18 to 0.30, and GenEval Position from 0.26 to 0.60, with no trainable parameters added at inference time.
Load-bearing premise
The reported gains generalize beyond the exact training distribution: SCOP pairs come only from COCO object categories with eight spatial tags, and the benchmarks test the same categories and the same binary left/right/above/below relations, so the improvements could reflect distribution matching rather than general spatial understanding.
Editorial extensions
If this is right
- Any existing UNet- or MMDiT-based text-to-image model can be upgraded for spatial accuracy with a short fine-tuning phase that adds no parameters at inference and only about 3% latency.
- A random 500-image subset of SCOP already lifts GenEval Position from 0.26 to 0.56 on FLUX.1, so the recipe is data-efficient enough for settings without access to web-scale datasets.
- The model trained only on two-object pairs improves three-object spatial accuracy (e.g., FLUX.1 'any' accuracy from 30.12 to 52.44), indicating the token-order signal transfers beyond the training template.
- The improvements are not confined to spatial metrics: overall GenEval, DPG-Bench, FID, and CMMD all improve, suggesting that cleaning spatial supervision also helps general prompt following.
- The ablations assign distinct roles to the two components: SCOP alone raises spatial accuracy substantially, and TENOR adds generalization to unseen prompt structures.
Reading between the lines
- Every reported benchmark shares SCOP's own COCO vocabulary and binary relation set, so the true test of general spatial understanding would be an out-of-distribution probe with non-COCO objects or context-dependent spatial language; the paper does not provide one.
- The same token-order blindness that scrambles left/right also plausibly degrades attribute binding and other order-sensitive compositions, so TENOR may transfer to color, size, and count tasks—an untested implication of the paper's analysis.
- The authors' listed limitations (extreme size disparities, object-centric frames) suggest concrete next experiments: building SCOP-style pairs that include size-contrast or object-centric annotations should extend the method toward fuller Qualitative Spatial Relations coverage, and their Fig. 7 shows a preliminary positive result for size.
- The 85.2% human-agreement check validates the SCOP captions, but the paper does not decompose how much of the benchmark gain comes from the crop-and-template decoding versus the geometric filtering itself.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper identifies two causes of poor spatial-relation generation in text-to-image (T2I) diffusion models: ambiguous spatial captions in existing image-text datasets and loss of token-ordering information in text encoders. It proposes CoMPaSS, composed of SCOP, a constraint-based data engine applied to the COCO training split that extracts object pairs satisfying visual significance, semantic distinction, spatial clarity, minimal overlap, and size balance, and TENOR, a parameter-free module that injects positional encodings into the text-image attention keys and queries. Experiments on SD1.4, SD1.5, SD2.1, and FLUX.1 report large gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%), together with improved overall scores on DPG-Bench and fidelity metrics, at low computational overhead and with promising data efficiency.
Significance. If the results hold, CoMPaSS is a practical and lightweight recipe for improving spatial compliance of open-weight T2I models: it adds no trainable parameters, requires only a brief fine-tuning phase, works across UNet and MMDiT architectures, and its data engine is simple and reproducible. The strengths of the paper are the systematic threshold-based curation pipeline, the sensible diagnostic in Table 1 showing that text encoders fail to rank logically equivalent spatial paraphrases as most similar, and the ablations in Tables 5 and 6 showing that both components contribute and that performance scales with training data. However, because SCOP and the headline benchmarks share the COCO category vocabulary and a small set of binary viewer-centric spatial relations, the evidence as presented is stronger for distribution-matched spatial compliance than for a general spatial-understanding capability. The broad claim of generalization needs an out-of-distribution evaluation and error bars before the conclusion is fully supported.
major comments (3)
- [Sec. 3.1, Sec. A.1, Tab. 2] The evidence for the paper's main claim that CoMPaSS enhances spatial understanding generally is currently confined to the same distribution used to build SCOP. SCOP curates pairs from the COCO training split and encodes only eight spatial tokens (<left>, <right>, <above>, <below>, and four diagonals), while the three headline benchmarks evaluate simple binary viewer-centric relations over essentially the same object vocabulary. The paper's own Sec. 5 lists context-dependent spatial language and object-centric frames as unsupported. Without an evaluation on categories and relation types outside this closed vocabulary, the large relative gains are equally consistent with distribution matching. Please add such an out-of-distribution test, or revise the conclusion to claim improved spatial compliance on this benchmark distribution rather than general spatial understanding.
- [Sec. 4.3, Tab. 5] The SCOP thresholds tau_v, tau_u, tau_o, and tau_s are selected by grid search on the same benchmarks that produce the headline SOTA numbers, and no repeated-seed or bootstrap intervals are reported for any accuracy result in Tabs. 2, 4, 5, 6, or A8-A11. This makes it impossible to quantify how much of the reported margin is selection bias and leaves the true improvement over baselines uncertain. Please provide confidence intervals and a validation split that is not used for threshold selection.
- [Tab. A11 vs. abstract/Sec. 4.2] The abstract and Sec. 4.2 state that gains are achieved 'without compromising general generation capabilities,' but the per-task breakdown in Tab. A11 shows several non-spatial tasks degrading: SD2.1+CoMPaSS drops GenEval Color from 0.85 to 0.71 and Count from 0.44 to 0.20, and SD1.5+CoMPaSS drops DPG-Bench Other from 67.81 to 60.80. Since overall scores can mask these trade-offs, the no-compromise claim should be made conditional on aggregate metrics or accompanied by a per-task analysis of which capabilities are preserved and which are not.
minor comments (5)
- [Sec. 3.1, Eqs. (1)-(5)] The thresholds are introduced as 'principled constraints' but are free parameters; please justify the chosen values or soften the terminology, and state how the resulting dataset size varies with each threshold.
- [Tab. 1] The proxy task tests nearest-neighbor ranking among four prompt variations; please clarify how ties are handled and report per-relation results, since 'above'/'below' may behave differently from 'left'/'right'.
- [Sec. 3.1] The human validation reports an 85.2% agreement rate but does not state the number of annotators, the number of items judged, or the exact instructions given; please add this information for reproducibility.
- [Sec. 5, Fig. 7] The size-disparity fine-tuning experiment is described only with one qualitative example; provide the training protocol and quantitative results, or label it explicitly as preliminary.
- [Appendix B] The latency overhead table reports mean +/- SD but not the number of measurement repetitions or the hardware conditions; please state the measurement protocol so the overhead numbers can be reproduced.
Circularity Check
No equation-level circularity and no load-bearing self-citation; the main residual circular element is that SCOP's thresholds are grid-searched on the GenEval Position benchmark that is then reported as a headline +131% 'prediction'.
-
fitted input called prediction
[Sec. 4.3 (Ablation Studies), Table 5; headline results in Sec. 1 and Table 2]
"The SCOP data engine has four tunable hyperparameters. We empirically determine the optimal values to be {τv, τu, τo, τs} = {0.2, 2.0, 0.3, 0.5} via grid search. In Tab. 5, we report the model’s sensitivity to each of these four hyperparameters by evaluating performance on nearby values. While our chosen hyperparameters yield optimal results, the model’s performance remains high across a range of nearby values."
Table 5's 'Ours' values are 0.54 for SD1.5 and 0.60 for FLUX.1, exactly the GenEval Position scores reported for SD1.5+CoMPaSS and FLUX.1+CoMPaSS in Table 2. Thus the SCOP threshold tuple was selected by maximizing the GenEval Position benchmark, and the same number is later presented as an independent +131% spatial-understanding prediction. The GenEval Position claim is therefore partly a fitted quantity rather than an out-of-sample result. This does not make the entire method circular: the VISOR and T2I-CompBench Spatial gains, and the TENOR ablation in Table 6, provide independent evidence. The issue is localized to one headline number and is a tuning-on-the-test-set loop, not an identity of equations.
full rationale
The paper's derivation chain is otherwise self-contained. SCOP is a data-curation engine that filters COCO pairs by explicit geometric constraints; TENOR adds absolute positional encodings to cross-attention key/query vectors, which is a parameter-free architectural intervention tested against the original models in Table 6. No claim is derived from a fitted parameter in the sense of a formula reducing to its inputs, and no load-bearing self-citation or imported uniqueness theorem appears. The main circularity concern is the SCOP hyperparameter grid search: the optimal thresholds in Table 5 coincide with the GenEval Position numbers later reported as SOTA, indicating selection on that benchmark. The reported VISOR (+98%) and T2I-CompBench Spatial (+67%) gains are not tied to that tuning loop, and the ablation shows large gains even for nearby threshold values, so the central claim retains substantial independent content. The broader worry that SCOP's COCO/8-token distribution matches the evaluation benchmarks is a generalization/overfitting concern rather than circularity and is therefore not scored as a circular step.
Assumptions & free parameters
free parameters (4)
- tau_v (visual significance threshold) =
0.2
- tau_u (spatial clarity threshold) =
2.0
- tau_o (minimal overlap threshold) =
0.3
- tau_s (size balance threshold) =
0.5
assumptions (4)
- domain assumption Bounding-box geometry and category labels suffice to determine the spatial relation between two objects.
- domain assumption The proxy task in Table 1 (highest similarity to rephrased variation) measures how well a text encoder preserves spatial semantics.
- domain assumption Benchmark prompts (VISOR, GenEval Position, T2I-CompBench Spatial) are unseen relative to SCOP training data and indicative of general spatial ability.
- domain assumption Fine-tuning on 28k spatial pairs does not degrade other generative abilities.
Cite this review
Pith. "Pith review of CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models." pith.science (2026). https://pith.science/paper/2V7MGPF5
@misc{pith2026241213195,
author = {Pith},
title = {Pith review of: CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2V7MGPF5}},
note = {Machine review of arXiv:2412.13195}
}
read the original abstract
Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models
Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elhoseiny. Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models. In IEEE/CVF International Confer- ence on Computer Vision, ICCV 2023, Paris, France, Octo- ber 1-6, 2023, pages 19984–19996. IEEE, 2023. 2
2023
-
[2]
Improving Image Genera- tion with Better Captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving Image Genera- tion with Better Captions. 1
-
[3]
Training diffusion models with reinforce- ment learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforce- ment learning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. 2
2024
-
[4]
FLUX.1-dev
Black Forest Labs. FLUX.1-dev. https : / / huggingface . co / black - forest - labs / FLUX . 1-dev, 2024. Accessed: 2025-07-16. 1, 5
2024
-
[5]
Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts. InIEEE Con- ference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 3558–3568. Com- puter Vision Foundation / IEEE, 2021. 1, 2, 3, 4
2021
-
[6]
Getting it right: Improving spatial consis- tency in text-to-image models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, and Yezhou Yang. Getting it right: Improving spatial consis- tency in text-to-image models. In Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 20...
work page 2024
-
[7]
Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models. ACM Trans. Graph., 42(4):148:1–148:10, 2023. 2, 5
work page 2023
-
[8]
Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu
Juan Chen, Anthony G. Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu. A survey of qual- itative spatial representations. pages 106–136, 2015. 8
work page 2015
Show all 85 references
-
[9]
Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu
Juan Chen, Anthony G. Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu. A survey of qual- itative spatial representations. Knowl. Eng. Rev., 30(1):106– 136, 2015. 8
2015
-
[10]
Pixart- Σ: Weak-to-strong training of dif- fusion transformer for 4k text-to-image generation
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- Σ: Weak-to-strong training of dif- fusion transformer for 4k text-to-image generation. In Com- puter Vision - ECCV 2024 - 18th European Conference...
2024
-
[11]
Pixart- δ: Fast and controllable image generation with latent consistency mod- els
Junsong Chen, Yue Wu, Simian Luo, Enze Xie, Sayak Paul, Ping Luo, Hang Zhao, and Zhenguo Li. Pixart- δ: Fast and controllable image generation with latent consistency mod- els. CoRR, abs/2401.05252, 2024
2024 arXiv
-
[12]
Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- alpha: Fast training of diffu- sion transformer for photorealistic text-to-image synthesis. In The Twelfth International Conference on Lear...
2024
-
[13]
Training-free layout control with cross-attention guidance
Minghao Chen, Iro Laina, and Andrea Vedaldi. Training-free layout control with cross-attention guidance. In IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2024, Waikoloa, HI, USA, January 3-8, 2024 , pages 5331–5341. IEEE, 2024. 2, 5, 6, 7
2024
-
[14]
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuh- mann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In IEEE/CVF Conference on Computer Vision and Pattern Reco...
2023
-
[15]
Visual pro- gramming for step-by-step text-to-image generation and evaluation
Jaemin Cho, Abhay Zala, and Mohit Bansal. Visual pro- gramming for step-by-step text-to-image generation and evaluation. In Advances in Neural Information Process- ing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, U...
2023
-
[16]
Coventry, Merc `e Prat-Sala, and Lynn Richards
Kenny R. Coventry, Merc `e Prat-Sala, and Lynn Richards. The interplay between geometry and function in the compre- hension of over, under, above, and below.Journal of Memory and Language, 44(3):376–398, 2001. 8
2001
-
[17]
Dall·e mini
Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Ph ´uc L ˆe Khac, Luke Melas, and Rito- brata Ghosh. Dall·e mini. https://github.com/ borisdayma/dalle-mini, 2021. 8
2021
-
[18]
Diffu- sion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffu- sion models beat gans on image synthesis. In Advances in Neural Information Processing Systems 34: Annual Con- ference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , pages 8780– 8...
2021
-
[19]
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. In Advances in Neural Informa- tion Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New O...
2022
-
[20]
Scaling rec- tified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rec- tified flow transformers for high-resolution image syn...
2024
-
[21]
Re- inforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Moham- mad Ghavamzadeh, Kangwook Lee, and Kimin Lee. Re- inforcement learning for fine-tuning text-to-image diffusion models. In Advances in Neural Information Processing Sys- tems 36:...
2023
-
[22]
Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang
Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Ar- jun R. Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured dif- fusion guidance for compositional text-to-image synthesis. In The Eleventh International Conference on Learn...
2023
-
[23]
Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Ar- jun R. Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang. Layoutgpt: Compositional visual plan- ning and generation with large language models. In Ad- vances in Neural Information Processing Systems 36: ...
2023
-
[24]
Tc-bench: Benchmark- ing temporal compositionality in text-to-video and image-to- video generation
Weixi Feng, Jiachen Li, Michael Saxon, Tsu-Jui Fu, Wenhu Chen, and William Yang Wang. Tc-bench: Benchmark- ing temporal compositionality in text-to-video and image-to- video generation. CoRR, abs/2406.08656, 2024. 2
2024 arXiv
-
[25]
Geneval: An object-focused framework for evaluating text- to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text- to-image alignment. In Advances in Neural Information Pro- cessing Systems 36: Annual Conference on Neural Informa- tion Processing Systems 2023, NeurIPS 2023, New ...
2023
-
[26]
Benchmarking spatial relationships in text-to-image generation
Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vi- neet, Eric Horvitz, Ece Kamar, Chitta Baral, and Yezhou Yang. Benchmarking spatial relationships in text-to-image generation. CoRR, abs/2212.10015, 2022. 2, 6, 7, 8
2022 arXiv
-
[27]
Diffusion-rpo: Aligning diffusion mod- els through relative preference optimization
Yi Gu, Zhendong Wang, Yueqin Yin, Yujia Xie, and Mingyuan Zhou. Diffusion-rpo: Aligning diffusion mod- els through relative preference optimization. CoRR, abs/2406.06382, 2024. 2
2024 arXiv
-
[28]
Yagmur G ¨uc ¸l¨ut¨urk, Umut G ¨uc ¸l¨u, Rob van Lier, and Mar- cel A. J. van Gerven. Convolutional sketch inversion. In Computer Vision - ECCV 2016 Workshops - Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I, pages 810–824. Springer, 2016. 2
2016
-
[29]
Ganspace: Discovering interpretable GAN controls
Erik H ¨ark¨onen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Discovering interpretable GAN controls. In Advances in Neural Information Processing Sys- tems 33: Annual Conference on Neural Information Process- ing Systems 2020, NeurIPS 2020, December 6-12, 2...
2020
-
[30]
Prompt-to-prompt image editing with cross-attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross-attention control. InThe Eleventh Interna- tional Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, ...
2023
-
[31]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Proc...
2017
-
[32]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. In Advances in Neural Informa- tion Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, De- cember 6-12, 2020, virtual, 2020. 2
2020
-
[33]
ELLA: equip diffusion models with LLM for en- hanced semantic alignment
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. ELLA: equip diffusion models with LLM for en- hanced semantic alignment. CoRR, abs/2403.05135, 2024. 2, 5, 6, 7, 8
2024 arXiv
-
[34]
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Os- tendorf, Ranjay Krishna, and Noah A. Smith. TIFA: accurate and interpretable text-to-image faithfulness evaluation with question answering. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, Fran...
2023
-
[35]
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xi- hui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. In Ad- vances in Neural Information Processing Systems 36: An- nual Conference on Neural Information Processing Syste...
2023
-
[36]
Re- thinking FID: towards a better evaluation metric for image generation
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Re- thinking FID: towards a better evaluation metric for image generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, ...
2024
-
[37]
Comat: Aligning text-to-image diffusion model with image- to-text concept matching
Dongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang, Dazhong Shen, Zhuofan Zong, Yu Liu, and Hongsheng Li. Comat: Aligning text-to-image diffusion model with image- to-text concept matching. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural ...
2024
-
[38]
Scalable ranked preference optimization for text-to-image generation
Shyamgopal Karthik, Huseyin Coskun, Zeynep Akata, Sergey Tulyakov, Jian Ren, and Anil Kag. Scalable ranked preference optimization for text-to-image generation. CoRR, abs/2410.18013, 2024. 2
2024 arXiv
-
[39]
Evaluating and improving composi- tional text-to-visual generation
Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Xide Xia, Pengchuan Zhang, Graham Neubig, and Deva Ramanan. Evaluating and improving composi- tional text-to-visual generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - W...
2024
-
[40]
Photomaker: Customizing realistic human photos via stacked ID embedding
Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming- Ming Cheng, and Ying Shan. Photomaker: Customizing realistic human photos via stacked ID embedding. CoRR, abs/2312.04461, 2023. 2
2023 arXiv
-
[41]
Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models
Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. Trans. Mach. Learn. Res., 2024, 2024. 2, 6
2024
-
[42]
Collins, Yiwen Luo, Yang Li, Kai J
Youwei Liang, Junfeng He, Gang Li, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang, Junjie Ke, Krishnamurthy Dj Dvijotham, Katherine M. Collins, Yiwen Luo, Yang Li, Kai J. Kohlhoff, Deepak Ramachandran, and Vidhya Naval- pak...
2024
-
[43]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In Computer Vision - ECCV 2014 - 13th Eu- ropean Conference, Zurich, Switzerland, September 6-12, 2014, ...
2014
-
[44]
Tenenbaum
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B. Tenenbaum. Compositional visual generation with composable diffusion models. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, Octo- ber 23-27, 2022, Proceedings, Part XVII , pages 423–439...
2022
-
[45]
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andr´es Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022...
2022
-
[46]
Pick-and-draw: Training-free semantic guidance for text-to-image person- alization
Henglei Lv, Jiayu Xiao, and Liang Li. Pick-and-draw: Training-free semantic guidance for text-to-image person- alization. In Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Aus- tralia, 28 October 2024 - 1 November 2024 , pages 1053...
2024
-
[47]
MidJourney
Inc. MidJourney. Midjourney: Ai-powered image genera- tion. https://www.midjourney.com/ , 2023. Ac- cessed: 2024-04-27. 1
2023
-
[48]
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Thirty-Eighth AAAI Conference on Ar- tificial Intelligence, AAAI 2024, Thirty-Sixt...
2024
-
[49]
GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, I...
2022
-
[50]
Drag your GAN: interactive point-based manipulation on the generative image manifold
Xingang Pan, Ayush Tewari, Thomas Leimk”uhler, Lingjie Liu, Abhimitra Meka, and Christian Theobalt. Drag your GAN: interactive point-based manipulation on the generative image manifold. In ACM SIGGRAPH 2023 Conference Pro- ceedings, SIGGRAPH 2023, Los Angeles, CA, USA, August ...
2023
-
[51]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1- 6, 2023, pages 4172–4182. IEEE, 2023. 2, 5
2023
-
[52]
Grounded text-to-image synthesis with attention refocusing
Quynh Phung, Songwei Ge, and Jia-Bin Huang. Grounded text-to-image synthesis with attention refocusing. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 7932–7942. IEEE, 2024. 2, 6, 7
2024
-
[53]
SDXL: improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M”uller, Joe Penna, and Robin Rombach. SDXL: improving latent diffusion models for high-resolution image synthesis. In The Twelfth Interna- tional Conference on Learning Representations, ICLR 2024,...
2024
-
[54]
Barron, and Ben Milden- hall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representa- tions, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenRe- view.net, 2023. 1
2023
-
[55]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[56]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21: 140:1–140:67, 2020. 2, 5
2020
-
[57]
Hierarchical text-conditional image gener- ation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with CLIP latents. CoRR, abs/2204.06125, 2022. 1, 2, 8
2022 arXiv
-
[58]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pages 10674–...
2022
-
[59]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, CVPR 2023, Vancouver, BC, Ca...
2023
-
[60]
Runway AI
Inc. Runway AI. Runwayml: Creative ai tools for content creation. https://runwayml.com/, 2023. Accessed: 2024-04-27. 1
2023
-
[61]
Dual caption preference optimization for diffusion models
Amir Saeidi, Yiran Luo, Agneet Chatterjee, Shamanthak Hegde, Bimsara Pathiraja, Yezhou Yang, and Chitta Baral. Dual caption preference optimization for diffusion models. CoRR, abs/2502.06023, 2025. 2
2025
-
[62]
Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Moham- mad Norouzi. Photorealistic text-to-image diffusion mod- els wit...
2022
-
[63]
Scribbler: Controlling deep image synthesis with sketch and color
Patsorn Sangkloy, Jingwan Lu, Chen Fang, Fisher Yu, and James Hays. Scribbler: Controlling deep image synthesis with sketch and color. In 2017 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2017, Hon- olulu, HI, USA, July 21-26, 2017 , pages 6836–6845. IEEE...
2017
-
[64]
LAION- 400M: open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION- 400M: open dataset of clip-filtered 400 million image-text pairs. CoRR, abs/2111.02114, 2021. 1, 2, 3, 4
2021 arXiv
-
[65]
LAION-5B: an open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAI...
2022
-
[66]
A picture is worth a thousand words: Principled recaptioning improves image generation
Eyal Segalis, Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. A picture is worth a thousand words: Principled recaptioning improves image generation. CoRR, abs/2310.16656, 2023. 2
2023 arXiv
-
[67]
Box it to bind it: Unified layout control and attribute binding in t2i diffusion models
Ashkan Taghipour, Morteza Ghahremani, Mohammed Ben- namoun, Aref Miri Rekavandi, Hamid Laga, and Farid Bous- said. Box it to bind it: Unified layout control and attribute binding in t2i diffusion models. CoRR, abs/2402.17910,
-
[68]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30, pages 5998–6008, 2017. 5
2017
-
[69]
Diffusion model align- ment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model align- ment using direct preference optimization. In IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2024
-
[70]
Instantid: Zero-shot identity-preserving gener- ation in seconds
Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and An- thony Chen. Instantid: Zero-shot identity-preserving gener- ation in seconds. CoRR, abs/2401.07519, 2024. 2
2024 arXiv
-
[71]
Tokencompose: Text-to-image diffusion with token-level supervision
Zirui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang, and Zhuowen Tu. Tokencompose: Text-to-image diffusion with token-level supervision. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 8553–8564. IEEE, 2024. 2, 1
2024
-
[72]
Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau
Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguis-...
2023
-
[73]
Seesr: Towards semantics-aware real-world image super-resolution
Rongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang, Shuai Li, and Lei Zhang. Seesr: Towards semantics-aware real-world image super-resolution. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , pages 25456–25467...
2024
-
[74]
Paragraph-to-image gener- ation with information-enriched diffusion model
Weijia Wu, Zhuang Li, Yefei He, Mike Zheng Shou, Chunhua Shen, Lele Cheng, Yan Li, Tingting Gao, Di Zhang, and Zhongyuan Wang. Paragraph-to-image gener- ation with information-enriched diffusion model. CoRR, abs/2311.14284, 2023. 2, 5
2023 arXiv
-
[75]
Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis
Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. CoRR, abs/2306.09341, 2023. 2
2023 arXiv
-
[76]
Human preference score: Better aligning text-to- image models with human preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hong- sheng Li. Human preference score: Better aligning text-to- image models with human preference. In IEEE/CVF Inter- national Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 , pages 2096–2105. IEEE, 2023. 2
2023
-
[77]
Stylespace analysis: Disentangled controls for stylegan image genera- tion
Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled controls for stylegan image genera- tion. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 12863–12872. Computer Vision Foundation / IEEE...
2021
-
[78]
Freeman, Fr ´edo Durand, and Song Han
Guangxuan Xiao, Tianwei Yin, William T. Freeman, Fr ´edo Durand, and Song Han. Fastcomposer: Tuning-free multi- subject image generation with localized attention. Int. J. Comput. Vis., 133(3):1175–1194, 2025. 2
2025
-
[79]
R&b: Region and boundary aware zero-shot grounded text-to-image generation
Jiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang, and Qing- ming Huang. R&b: Region and boundary aware zero-shot grounded text-to-image generation. In The Twelfth Interna- tional Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2...
2024
-
[80]
Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagere- ward: Learning and evaluating human preferences for text- to-image generation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Infor- m...
2023
-
[81]
Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. CoRR, abs/2308.06721, 2023. 2
2023 arXiv
-
[82]
Scaling up to excellence: Practicing model scaling for photo- realistic image restoration in the wild
Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen He, Yu Qiao, and Chao Dong. Scaling up to excellence: Practicing model scaling for photo- realistic image restoration in the wild. In IEEE/CVF Confer- ence on Computer Vision and Pattern Recognit...
2024
-
[83]
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content...
2022
-
[84]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 3813–
2023
-
[85]
A horse to the left of a bottle
Bo Zhao, Lili Meng, Weidong Yin, and Leonid Sigal. Image generation from layout. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , pages 8584–8593. Computer Vision Foundation / IEEE, 2019. 2 Appendix A. Additional...
2019
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.