REVIEW 3 major objections 5 minor 2 cited by
Image Editing with Diffusion Models: A Survey
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This survey argues that all diffusion-based image editing can be organized by one taxonomy: visual content versus visual expression, four operations, and three method families.
desk verdict A useful but over-claimed taxonomy survey: broad, readable, and worth engaging, but the three-way method split is cleaner on paper than on the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the taxonomy itself: a decomposition of an image into visual content and visual expression, a four-way action set of add, delete, change, and combine, and a three-way method split determined by whether the base model is untouched, fine-tuned, or extended with an adapter (a trained add-on module that injects processed prompts into a frozen model). The taxonomy does the work of assigning every task a coordinate, every method a bucket, and every benchmark or dataset a rationale; all four survey sections are organized around it.
What would settle it
Take the methods from a recent major computer-vision venue and try to place each one into exactly one of the three categories. A single method whose contribution depends equally on fine-tuning base-model parameters and training an adapter—with ablations showing both are necessary—would disprove exclusivity, and a recognized editing task that changes neither visual content nor visual expression and is not an add, delete, change, or combine would disprove completeness.
Extended reading notes
Core claim
The paper's central claim is organizational: image editing can be defined as modifying an existing image to satisfy a user intention, and a systematic taxonomy falls out of that definition. Images are first separated into feature images (keypoints, edges, depth) and natural images; natural images are then divided into visual content (objects, backgrounds) and visual expression (structure, style, light, texture). User operations reduce to four actions—add, delete, change, combine—and instructions can arrive as text, as a feature image, or as in-context image pairs. Methods are grouped by how much of the base model they modify: inversion-based methods steer noise latents or attention maps without touching parameters, fine-tuning-based methods retrain all or part of the network at training or test time, and adapter-based methods keep the base model frozen and inject processed multimodal prompts through trained adapters. The survey claims this classification is clearer and simpler than earlier schemes, and it uses the same distinctions to organize evaluation, benchmarks, and dataset construction.
Load-bearing premise
The load-bearing premise is that every editing task fits the content/expression grid with the four actions and every method fits exactly one of the three buckets by its dominant mechanism; the paper itself notes that some methods combine adapter training with base-model fine-tuning, so the exclusivity of the buckets is the assumption most likely to give way.
Editorial extensions
If this is right
- A new editing method can be located by asking three questions: which image component it targets, which of the four actions it performs, and whether it leaves the base model untouched, fine-tunes it, or trains an adapter.
- The task grid gives benchmark builders a checklist of task cells to cover; the survey reviews EditBench, EditVal, TEdBench, I2EBench, and EditEval as variants of one pipeline.
- Dataset construction splits into two pipelines—extraction-based (feature maps, subject sets, video frames) and generation-based (LLM-written instructions plus consistent image pairs)—so data builders can choose by task type.
- The survey's challenges follow from the taxonomy: inversion methods lose detail to stochasticity, fine-tuning needs data and compute, adapters handle multi-image inputs poorly, and current metrics miss semantic coherence; the stated future directions are unified, multimodal, multi-turn editing models.
Reading between the lines
- An implication the authors leave implicit is that the task grid is generative: crossing every content/expression component with every action defines a design space, and cells such as 'combine applied to lighting' or 'delete applied to texture' remain largely unexplored.
- Because the paper concedes that some methods train adapters while also fine-tuning base parameters, the three method families are best read as labels for the dominant mechanism rather than exclusive classes; marking hybrid methods explicitly would make the taxonomy more robust.
- If MLLM scoring plus human validation continues to displace pure compute metrics, editing evaluation will increasingly need to measure instruction-intent understanding, not just fidelity—a direction the survey signals but does not develop.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of image editing with diffusion models. It proposes a task taxonomy in which an image is decomposed into visual content and visual expression, editing operations are characterized as add/delete/change/combine, and instructions are classified as text, feature-image, or in-context. It then organizes editing methods into inversion-based, fine-tuning-based, and adapter-based categories, and reviews evaluation metrics, benchmarks, and dataset construction approaches. The final sections discuss open challenges and future directions. The paper states that its main distinction from prior surveys is a clearer and simpler classification.
Significance. The survey's four-part organization (tasks, methods, evaluation, datasets) is sensible, and the detailed account of dataset construction pipelines in Section 5, including generation-based construction and data filtering, is a practical contribution that goes beyond many earlier surveys. The paper covers a broad range of recent methods and provides helpful illustrative figures. The main caveat is that the central method taxonomy is not yet internally consistent: the manuscript itself contains examples that do not fit the trichotomy and it resolves a hybrid case with a subjective tie-breaker. This weakens the survey's central claim, but the issue is addressable by revision. The contribution is organizational; the paper provides no code or machine-checked artifacts.
major comments (3)
- [Section 3, Section 3.3.2] The central claim that every editing method falls into exactly one of the inversion-based, fine-tuning-based, or adapter-based categories is not supported by the manuscript's own text. Section 3 defines the categories by 'the extent of modifications applied to the base model' and states that adapter-based methods require no modification of base-model parameters. Section 3.3.2 then says that 'some recent methods simultaneously use adapters to process multimodal prompts and fine-tune the original model parameters,' and assigns these methods to the adapter category because their 'main innovation and contribution' lies in prompt processing. This is a subjective priority rule rather than a classification criterion, so the proposed categories are not mutually exclusive as claimed. I recommend either defining an explicit primary-modification rule or presenting the three categories as overlapping families and discussing hybrid methods.
- [Section 3.2.2] Textual Inversion is described as a testing-time fine-tuning method, but under the definitions in Section 3 it does not modify base-model parameters, and it does not manipulate noise latents or attention maps in the way the inversion-based category is described. It optimizes a learnable text embedding from a few images. This is an internal counterexample to the claimed partition: the method belongs to none of the three categories as defined. The taxonomy needs either a fourth category, such as embedding or prompt optimization, or a broader definition of fine-tuning that is stated explicitly and applied consistently.
- [Section 5.1.1] The dataset MultiGen-20M is attributed to 'UniControl (Zhao et al., 2024a),' but MultiGen-20M was introduced by Qin et al. (2023) with UniControl, while Zhao et al. (2024a) is Uni-ControlNet. The inconsistency is evident because Section 3.3.1 correctly cites UniControl as Qin et al. (2023). This attribution error should be corrected, since Section 5 is one of the survey's main contributions and accuracy of dataset provenance is essential in a survey.
minor comments (5)
- [Section 5.2] The paragraph introducing generation-based dataset construction contains a duplicated fragment: '...the strategies they employ to ensure the quality and consistency of the generated datasets. have approached these steps and the strategies they employ...' The second fragment should be deleted.
- [Section 3.2.2] The description of Imagic says that it 'fixes the network parameters' and then states that 'the fine-tuning process involves adjusting the entire denoising network.' Please clarify that these statements refer to the two stages of Imagic, because as written the description is self-contradictory.
- [Section 4.2.1] The text cites 'Reason-Edit (Huang et al., 2024b),' but the bibliography entry for Huang et al. (2024b) is titled SmartEdit; please align the in-text name with the cited work.
- [Sections 3.2.1, 5.2.1] The paper uses inconsistent spelling for OmniControl: 'Omini-control' in Section 3.2.1 and 'OminiControl' in Section 5.2.1 versus 'OmniControl' elsewhere.
- [Figure 5, Section 4.2.1] There are several typos: 'beah' in the Fig. 5 caption should be 'beach'; 'reasonina' in Section 4.2.1 should be 'reasoning'; and 'OMniEdit-Filtered-1.2M' in Section 5.2.3 has unusual capitalization.
Circularity Check
No significant circularity: the survey's taxonomy and summaries describe external literature rather than deriving predictions from fitted inputs.
full rationale
This is a survey paper whose proposed contributions are organizational: a task taxonomy (visual content vs. visual expression; add/delete/change/combine), a method trichotomy (inversion-based, fine-tuning-based, adapter-based), a metrics summary, and a dataset-construction overview. None of these are derived from a fitted parameter that is then renamed as a prediction, and no equation or theorem in the paper is shown to reduce to its own inputs. The method classification is justified by the stated criterion of 'the extent of modifications applied to the base model' (Section 3), with each category illustrated by representative external methods; this is an interpretive taxonomy, not a self-referential derivation. The paper explicitly acknowledges that some recent methods combine adapter processing with fine-tuning (Section 3.3.2), and the placement of borderline methods such as Textual Inversion under testing-time fine-tuning is a classification judgment, not a circular argument. The stated distinctiveness relative to prior surveys is 'clearer, simpler classifications,' which is a claim about expository utility rather than an empirically derived result. No load-bearing self-citations, uniqueness theorems, or fitted-input-called-prediction steps are present. Any weakness about taxonomy completeness or overlap is a correctness/evidence concern, not circularity. The survey is therefore self-contained as a literature organization, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The selected references are representative of the image editing field and are described accurately.
- ad hoc to paper Each editing method can be assigned to exactly one of the three categories: inversion-based, fine-tuning-based, or adapter-based.
Cite this review
Pith. "Pith review of Image Editing with Diffusion Models: A Survey." pith.science (2026). https://pith.science/paper/YBTSNVFV
@misc{pith2026250413226,
author = {Pith},
title = {Pith review of: Image Editing with Diffusion Models: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/YBTSNVFV}},
note = {Machine review of arXiv:2504.13226}
}
read the original abstract
With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images continues to improve, so does the demand for further application like image editing. In recent years, many remarkable works are realizing a wide variety of editing effects. However, the wide variety of editing types and diverse editing approaches have made it difficult for researchers to establish a comprehensive view of the development of this field. In this survey, we summarize the image editing field from four aspects: tasks definition, methods classification, results evaluation and editing datasets. First, we provide a definition of image editing, which in turn leads to a variety of editing task forms from the perspective of operation parts and manipulation actions. Subsequently, we categorize and summary methods for implementing editing into three categories: inversion-based, fine-tuning-based and adapter-based. In addition, we organize the currently used metrics, available datasets and corresponding construction methods. At the end, we present some visions for the future development of the image editing field based on the previous summaries.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 2 Pith papers
-
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.
-
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
MultiFakeVerse provides 845,286 person-centric images edited through VLM-generated instructions; state-of-the-art deepfake detectors and human observers misclassify a large fraction of them.
Reference graph
Works this paper leans on
-
[1]
(2023) Gpt-4 technical report
Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, Almeida D, Altenschmidt J, Altman S, Anadkat S, et al. (2023) Gpt-4 technical report. arXiv preprint arXiv:230308774
2023
-
[2]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 18208--18218
Avrahami O, Lischinski D, Fried O (2022) Blended diffusion for text-driven editing of natural images. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 18208--18218
2022
-
[3]
ACM transactions on graphics (TOG) 42(4):1--11
Avrahami O, Fried O, Lischinski D (2023) Blended latent diffusion. ACM transactions on graphics (TOG) 42(4):1--11
2023
-
[4]
arXiv preprint arXiv:230812966
Bai J, Bai S, Yang S, Wang S, Tan S, Wang P, Lin J, Zhou C, Zhou J (2023) Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:230812966
2023
-
[5]
Advances in Neural Information Processing Systems 36
Bansal A, Borgnia E, Chu HM, Li J, Kazemi H, Huang F, Goldblum M, Geiping J, Goldstein T (2024) Cold diffusion: Inverting arbitrary image transforms without noise. Advances in Neural Information Processing Systems 36
2024
-
[6]
arXiv preprint arXiv:220106503
Bao F, Li C, Zhu J, Zhang B (2022) Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. arXiv preprint arXiv:220106503
2022
-
[7]
In: European conference on computer vision, Springer, pp 707--723
Bar-Tal O, Ofri-Amar D, Fridman R, Kasten Y, Dekel T (2022) Text2live: Text-driven layered image and video editing. In: European conference on computer vision, Springer, pp 707--723
2022
-
[8]
arXiv preprint arXiv:231002426
Basu S, Saberi M, Bhardwaj S, Chegini AM, Massiceti D, Sanjabi M, Hu SX, Feizi S (2023) Editval: Benchmarking diffusion based text-guided image editing methods. arXiv preprint arXiv:231002426
2023
Show all 176 references
-
[9]
arXiv preprint arXiv:180101401
Bi \'n kowski M, Sutherland DJ, Arbel M, Gretton A (2018) Demystifying mmd gans. arXiv preprint arXiv:180101401
2018
-
[10]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 22563--22575
Blattmann A, Rombach R, Ling H, Dockhorn T, Kim SW, Fidler S, Kreis K (2023) Align your latents: High-resolution video synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 22563--22575
2023
-
[11]
arXiv preprint arXiv:200410934
Bochkovskiy A, Wang CY, Liao HYM (2020) Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:200410934
2020
-
[12]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8861--8870
Brack M, Friedrich F, Kornmeier K, Tsaban L, Schramowski P, Kersting K, Passos A (2024) Ledits++: Limitless image editing using text-to-image models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8861--8870
2024
-
[13]
arXiv preprint arXiv:180911096
Brock A (2018) Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:180911096
2018
-
[14]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18392--18402
Brooks T, Holynski A, Efros AA (2023) Instructpix2pix: Learning to follow image editing instructions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18392--18402
2023
-
[15]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1209--1218
Caesar H, Uijlings J, Ferrari V (2018) Coco-stuff: Thing and stuff classes in context. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1209--1218
2018
-
[16]
IEEE Transactions on pattern analysis and machine intelligence (6):679--698
Canny J (1986) A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelligence (6):679--698
1986
-
[17]
IEEE Transactions on Knowledge and Data Engineering
Cao H, Tan C, Gao Z, Xu Y, Chen G, Heng PA, Li SZ (2024 a ) A survey on generative diffusion models. IEEE Transactions on Knowledge and Data Engineering
2024
-
[18]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 22560--22570
Cao M, Wang X, Qi Z, Shan Y, Qie X, Zheng Y (2023) Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 22560--22570
2023
-
[19]
arXiv preprint arXiv:240304279
Cao P, Zhou F, Song Q, Yang L (2024 b ) Controllable generation with text-to-image diffusion models: A survey. arXiv preprint arXiv:240304279
2024
-
[20]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7291--7299
Cao Z, Simon T, Wei SE, Sheikh Y (2017) Realtime multi-person 2d pose estimation using part affinity fields. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7291--7299
2017
-
[21]
In: Proceedings of the IEEE/CVF international conference on computer vision, pp 9650--9660
Caron M, Touvron H, Misra I, J \'e gou H, Mairal J, Bojanowski P, Joulin A (2021) Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 9650--9660
2021
-
[22]
arXiv preprint arXiv:231019145
Chakrabarty T, Singh K, Saakyan A, Muresan S (2023) Learning to follow object-centric image editing instructions faithfully. arXiv preprint arXiv:231019145
2023
-
[23]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6593--6602
Chen X, Huang L, Liu Y, Shen Y, Zhao D, Zhao H (2024 a ) Anydoor: Zero-shot object-level image customization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6593--6602
2024
-
[24]
(2024 b ) Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen Z, Wu J, Wang W, Su W, Chen G, Xing S, Zhong M, Zhang Q, Zhu X, Lu L, et al. (2024 b ) Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp...
2024
-
[25]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 14131--14140
Choi S, Park S, Lee M, Choo J (2021) Viton-hd: High-resolution virtual try-on via misalignment-aware normalization. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 14131--14140
2021
-
[26]
arXiv preprint arXiv:240305139
Choi Y, Kwak S, Lee K, Choi H, Shin J (2024) Improving diffusion models for virtual try-on. arXiv preprint arXiv:240305139
2024
-
[27]
(2024) Mobilevlm v2: Faster and stronger baseline for vision language model
Chu X, Qiao L, Zhang X, Xu S, Wei F, Yang Y, Sun X, Hu Y, Lin X, Zhang B, et al. (2024) Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:240203766
2024
-
[28]
Contributors M (2020) Openmmlab pose estimation toolbox and benchmark
2020
-
[29]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45(9):10850--10869
Croitoru FA, Hondru V, Ionescu RT, Shah M (2023) Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(9):10850--10869
2023
-
[30]
arXiv preprint arXiv:160603798
DeTone D, Malisiewicz T, Rabinovich A (2016) Deep image homography estimation. arXiv preprint arXiv:160603798
2016
-
[31]
Advances in neural information processing systems 34:8780--8794
Dhariwal P, Nichol A (2021) Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34:8780--8794
2021
-
[32]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7430--7440
Dong W, Xue S, Duan X, Han S (2023) Prompt tuning inversion for text-driven image editing using diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7430--7440
2023
-
[33]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 12873--12883
Esser P, Rombach R, Ommer B (2021) Taming transformers for high-resolution image synthesis. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 12873--12883
2021
-
[34]
arXiv preprint arXiv:230917102
Fu TJ, Hu W, Du X, Wang WY, Yang Y, Gan Z (2023) Guiding instruction-based image editing via multimodal large language models. arXiv preprint arXiv:230917102
2023
-
[35]
arXiv preprint arXiv:220801618
Gal R, Alaluf Y, Atzmon Y, Patashnik O, Bermano AH, Chechik G, Cohen-Or D (2022) An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:220801618
2022
-
[36]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10021--10030
Gao S, Liu X, Zeng B, Xu S, Li Y, Luo X, Liu J, Zhen X, Zhang B (2023) Implicit diffusion models for continuous super-resolution. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10021--10030
2023
-
[37]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 22930--22941
Ge S, Nah S, Liu G, Poon T, Tao A, Catanzaro B, Jacobs D, Huang JB, Liu MY, Balaji Y (2023) Preserve your own correlation: A noise prior for video diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 22930--22941
2023
-
[38]
(2024) Instructdiffusion: A generalist modeling interface for vision tasks
Geng Z, Yang B, Hang T, Li C, Gu S, Zhang T, Bao J, Zhang Z, Li H, Hu H, et al. (2024) Instructdiffusion: A generalist modeling interface for vision tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12709--12720
2024
-
[39]
arXiv preprint arXiv:221008933
Gong S, Li M, Feng J, Wu Z, Kong L (2022) Diffuseq: Sequence to sequence text generation with diffusion models. arXiv preprint arXiv:221008933
2022
-
[40]
Advances in neural information processing systems 27
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. Advances in neural information processing systems 27
2014
-
[41]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 36, pp 726--734
Gu G, Ko B, Go S, Lee SH, Lee J, Shin M (2022) Towards light-weight and real-time line segment detection. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 36, pp 726--734
2022
-
[42]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14049--14058
Guo L, Wang C, Yang W, Huang S, Wang Y, Pfister H, Wen B (2023 a ) Shadowdiffusion: When degradation prior meets diffusion model for shadow removal. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14049--14058
2023
-
[43]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 12097--12107
Guo Y, Xiao X, Chang Y, Deng S, Yan L (2023 b ) From sky to the ground: A large-scale benchmark and simple baseline towards real rain removal. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 12097--12107
2023
-
[44]
(2024) Proxedit: Improving tuning-free real image editing with proximal guidance
Han L, Wen S, Chen Q, Zhang Z, Song K, Ren M, Gao R, Stathopoulos A, He X, Chen Y, et al. (2024) Proxedit: Improving tuning-free real image editing with proximal guidance. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp 4291--4301
2024
-
[45]
arXiv preprint arXiv:240918071
He R, Ma K, Huang L, Huang S, Gao J, Wei X, Dai J, Han J, Liu S (2024) Freeedit: Mask-free reference-based image editing with multi-modal instruction. arXiv preprint arXiv:240918071
2024
-
[46]
arXiv preprint arXiv:220801626
Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D (2022) Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:220801626
2022
-
[47]
arXiv preprint arXiv:210408718
Hessel J, Holtzman A, Forbes M, Bras RL, Choi Y (2021) Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:210408718
2021
-
[48]
Advances in neural information processing systems 30
Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30
2017
-
[49]
ICLR (Poster) 3
Higgins I, Matthey L, Pal A, Burgess CP, Glorot X, Botvinick MM, Mohamed S, Lerchner A (2017) beta-vae: Learning basic visual concepts with a constrained variational framework. ICLR (Poster) 3
2017
-
[50]
arXiv preprint arXiv:220712598
Ho J, Salimans T (2022) Classifier-free diffusion guidance. arXiv preprint arXiv:220712598
2022
-
[51]
Advances in neural information processing systems 33:6840--6851
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Advances in neural information processing systems 33:6840--6851
2020
-
[52]
(2022 a ) Imagen video: High definition video generation with diffusion models
Ho J, Chan W, Saharia C, Whang J, Gao R, Gritsenko A, Kingma DP, Poole B, Norouzi M, Fleet DJ, et al. (2022 a ) Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:221002303
2022
-
[53]
Journal of Machine Learning Research 23(47):1--33
Ho J, Saharia C, Chan W, Fleet DJ, Norouzi M, Salimans T (2022 b ) Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research 23(47):1--33
2022
-
[54]
Advances in Neural Information Processing Systems 35:8633--8646
Ho J, Salimans T, Gritsenko A, Chan W, Norouzi M, Fleet DJ (2022 c ) Video diffusion models. Advances in Neural Information Processing Systems 35:8633--8646
2022
-
[55]
arXiv preprint arXiv:240217525
Huang Y, Huang J, Liu Y, Yan M, Lv J, Liu J, Xiong W, Zhang H, Chen S, Cao L (2024 a ) Diffusion model-based image editing: A survey. arXiv preprint arXiv:240217525
2024
-
[56]
(2024 b ) Smartedit: Exploring complex instruction-based image editing with multimodal large language models
Huang Y, Xie L, Wang X, Yuan Z, Cun X, Ge Y, Zhou J, Dong C, Huang R, Zhang R, et al. (2024 b ) Smartedit: Exploring complex instruction-based image editing with multimodal large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[57]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12469--12478
Huberman-Spiegelglas I, Kulikov V, Michaeli T (2024) An edit friendly ddpm noise space: Inversion and manipulations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 12469--12478
2024
-
[58]
arXiv preprint arXiv:240409990
Hui M, Yang S, Zhao B, Shi Y, Wang H, Wang P, Zhou Y, Xie C (2024) Hq-edit: A high-quality dataset for instruction-based image editing. arXiv preprint arXiv:240409990
2024
-
[59]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1125--1134
Isola P, Zhu JY, Zhou T, Efros AA (2017) Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1125--1134
2017
-
[60]
arXiv preprint arXiv:241011831
Karaev N, Makarov I, Wang J, Neverova N, Vedaldi A, Rupprecht C (2024) Cotracker3: Simpler and better point tracking by pseudo-labelling real videos. arXiv preprint arXiv:241011831
2024
-
[61]
arXiv preprint arXiv:181204948
Karras T (2019) A style-based generator architecture for generative adversarial networks. arXiv preprint arXiv:181204948
2019
-
[62]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6007--6017
Kawar B, Zada S, Lang O, Tov O, Chang H, Dekel T, Mosseri I, Irani M (2023) Imagic: Text-based real image editing with diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6007--6017
2023
-
[63]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 2426--2435
Kim G, Kwon T, Ye JC (2022) Diffusionclip: Text-guided diffusion models for robust image manipulation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 2426--2435
2022
-
[64]
arXiv preprint arXiv:13126114
Kingma DP (2013) Auto-encoding variational bayes. arXiv preprint arXiv:13126114
2013
-
[65]
(2023) Segment anything
Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY, et al. (2023) Segment anything. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 4015--4026
2023
-
[66]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10051--10060
Kolkin N, Salavon J, Shakhnarovich G (2019) Style transfer by relaxed optimal transport and self-similarity. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10051--10060
2019
-
[67]
arXiv preprint arXiv:200909761
Kong Z, Ping W, Huang J, Zhao K, Catanzaro B (2020) Diffwave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:200909761
2020
-
[68]
IEEE Transactions on Intelligent Transportation Systems 23(8):13498--13511
Kreiss S, Bertoni L, Alahi A (2021) Openpifpaf: Composite fields for semantic keypoint detection and spatio-temporal association. IEEE Transactions on Intelligent Transportation Systems 23(8):13498--13511
2021
-
[69]
arXiv preprint arXiv:190701341
Lasinger K, Ranftl R, Schindler K, Koltun V (2019) Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. arXiv preprint arXiv:190701341
2019
-
[70]
arXiv preprint arXiv:240309055
Lee J, Jung DS, Lee K, Lee KM (2024) Semanticdraw: towards real-time interactive content creation from image diffusion models. arXiv preprint arXiv:240309055
2024
-
[71]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 13906--13915
Lee Y, Park J (2020) Centermask: Real-time anchor-free instance segmentation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 13906--13915
2020
-
[72]
arXiv preprint arXiv:230600950
Levin E, Fried O (2023) Differential diffusion: Giving each pixel its strength. arXiv preprint arXiv:230600950
2023
-
[73]
Advances in Neural Information Processing Systems 36:30146--30166
Li D, Li J, Hoi S (2023 a ) Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. Advances in Neural Information Processing Systems 36:30146--30166
2023
-
[74]
In: International conference on machine learning, PMLR, pp 12888--12900
Li J, Li D, Xiong C, Hoi S (2022) Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International conference on machine learning, PMLR, pp 12888--12900
2022
-
[75]
In: European Conference on Computer Vision, Springer, pp 129--147
Li M, Yang T, Kuang H, Wu J, Wang Z, Xiao X, Chen C (2025) Controlnet ++ : Improving conditional controls with efficient consistency feedback. In: European Conference on Computer Vision, Springer, pp 129--147
2025
-
[76]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 2039--2048
Li N, Liu Q, Singh KK, Wang Y, Zhang J, Plummer BA, Lin Z (2024) Unihuman: A unified model for editing human images in the wild. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 2039--2048
2024
-
[77]
arXiv preprint arXiv:230904372
Li S, Chen C, Lu H (2023 b ) Moecontroller: Instruction-based arbitrary image manipulation with mixture-of-expert controllers. arXiv preprint arXiv:230904372
2023
-
[78]
arXiv preprint arXiv:231206738
Li S, Singh H, Grover A (2023 c ) Instructany2pix: Flexible visual editing via multimodal instruction following. arXiv preprint arXiv:231206738
2023
-
[79]
In: Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer, pp 740--755
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll \'a r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer, pp 740--755
2014
-
[80]
(2024) Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Liu D, Zhang R, Qiu L, Huang S, Lin W, Zhao S, Geng S, Lin Z, Jin P, Zhang K, et al. (2024) Sphinx-x: Scaling data and parameters for a family of multi-modal large language models. arXiv preprint arXiv:240205935
2024
-
[81]
arXiv preprint arXiv:230112503
Liu H, Chen Z, Yuan Y, Mei X, Liu X, Mandic D, Wang W, Plumbley MD (2023) Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:230112503
2023
-
[82]
In: Proceedings of the 29th ACM international conference on multimedia, pp 50--58
Liu Y, Zhu L, Pei S, Fu H, Qin J, Zhang Q, Wan L, Feng W (2021) From synthetic to real: Image dehazing collaborating with unlabeled real data. In: Proceedings of the 29th ACM international conference on multimedia, pp 50--58
2021
-
[83]
In: Proceedings of the 30th ACM International Conference on Multimedia, pp 638--647
Ma Y, Xu G, Sun X, Yan M, Zhang J, Ji R (2022) X-clip: End-to-end multi-grained contrastive learning for video-text retrieval. In: Proceedings of the 30th ACM International Conference on Multimedia, pp 638--647
2022
-
[84]
arXiv preprint arXiv:240814180
Ma Y, Ji J, Ye K, Lin W, Wang Z, Zheng Y, Zhou Q, Sun X, Ji R (2024) I2ebench: A comprehensive benchmark for instruction-based image editing. arXiv preprint arXiv:240814180
2024
-
[85]
arXiv preprint arXiv:210801073
Meng C, He Y, Song Y, Song J, Wu J, Zhu JY, Ermon S (2021) Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:210801073
2021
-
[86]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14297--14306
Meng C, Rombach R, Gao R, Kingma D, Ermon S, Ho J, Salimans T (2023) On distillation of guided diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 14297--14306
2023
-
[87]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 21033--21043
Miao J, Wang X, Wu Y, Li W, Zhang X, Wei Y, Yang Y (2022) Large-scale video panoptic segmentation in the wild: A benchmark. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 21033--21043
2022
-
[88]
arXiv preprint arXiv:14111784
Mirza M (2014) Conditional generative adversarial nets. arXiv preprint arXiv:14111784
2014
-
[89]
arXiv preprint arXiv:230516807
Miyake D, Iohara A, Saito Y, Tanaka T (2023) Negative-prompt inversion: Fast image inversion for editing with text-guided diffusion models. arXiv preprint arXiv:230516807
2023
-
[90]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6038--6047
Mokady R, Hertz A, Aberman K, Pritch Y, Cohen-Or D (2023) Null-text inversion for editing real images using guided diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 6038--6047
2023
-
[91]
In: Proceedings of the 31st ACM International Conference on Multimedia, pp 8580--8589
Morelli D, Baldrati A, Cartella G, Cornia M, Bertini M, Cucchiara R (2023) Ladi-vton: Latent diffusion textual-inversion enhanced virtual try-on. In: Proceedings of the 31st ACM International Conference on Multimedia, pp 8580--8589
2023
-
[92]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 38, pp 4296--4304
Mou C, Wang X, Xie L, Wu Y, Zhang J, Qi Z, Shan Y (2024) T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 38, pp 4296--4304
2024
-
[93]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3883--3891
Nah S, Hyun Kim T, Mu Lee K (2017) Deep multi-scale convolutional neural network for dynamic scene deblurring. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3883--3891
2017
-
[94]
arXiv preprint arXiv:211210741
Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, Sutskever I, Chen M (2021) Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:211210741
2021
-
[95]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45(8):10346--10357
\"O zdenizci O, Legenstein R (2023) Restoring vision in adverse weather conditions with patch-based denoising diffusion models. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(8):10346--10357
2023
-
[96]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 15912--15921
Pan Z, Gherardi R, Xie X, Huang S (2023) Effective real image editing with accelerated iterative diffusion inversion. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 15912--15921
2023
-
[97]
Advances in neural information processing systems 30
Papamakarios G, Pavlakou T, Murray I (2017) Masked autoregressive flow for density estimation. Advances in neural information processing systems 30
2017
-
[98]
Advances in Neural Information Processing Systems 33:7198--7211
Park T, Zhu JY, Wang O, Lu J, Shechtman E, Efros A, Zhang R (2020) Swapping autoencoder for deep image manipulation. Advances in Neural Information Processing Systems 33:7198--7211
2020
-
[99]
In: ACM SIGGRAPH 2023 Conference Proceedings, pp 1--11
Parmar G, Kumar Singh K, Zhang R, Li Y, Lu J, Zhu JY (2023) Zero-shot image-to-image translation. In: ACM SIGGRAPH 2023 Conference Proceedings, pp 1--11
2023
-
[100]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10199--10208
Phung H, Dao Q, Tran A (2023) Wavelet diffusion models are fast and scalable image generators. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10199--10208
2023
-
[101]
(2024) State of the art on diffusion models for visual computing
Po R, Yifan W, Golyanik V, Aberman K, Barron JT, Bermano A, Chan E, Dekel T, Holynski A, Kanazawa A, et al. (2024) State of the art on diffusion models for visual computing. In: Computer Graphics Forum, Wiley Online Library, vol 43, p e15063
2024
-
[102]
arXiv preprint arXiv:230701952
Podell D, English Z, Lacey K, Blattmann A, Dockhorn T, M \"u ller J, Penna J, Rombach R (2023) Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:230701952
2023
-
[103]
(2023) Unicontrol: A unified diffusion model for controllable visual generation in the wild
Qin C, Zhang S, Yu N, Feng Y, Yang X, Zhou Y, Wang H, Niebles JC, Xiong C, Savarese S, et al. (2023) Unicontrol: A unified diffusion model for controllable visual generation in the wild. arXiv preprint arXiv:230511147
2023
-
[104]
(2021) Learning transferable visual models from natural language supervision
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al. (2021) Learning transferable visual models from natural language supervision. In: International conference on machine learning, PMLR, pp 8748--8763
2021
-
[105]
In: International conference on machine learning, Pmlr, pp 8821--8831
Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, Chen M, Sutskever I (2021) Zero-shot text-to-image generation. In: International conference on machine learning, Pmlr, pp 8821--8831
2021
-
[106]
arXiv preprint arXiv:220406125 1(2):3
Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:220406125 1(2):3
2022
-
[107]
IEEE transactions on pattern analysis and machine intelligence 44(3):1623--1637
Ranftl R, Lasinger K, Hafner D, Schindler K, Koltun V (2020) Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence 44(3):1623--1637
2020
-
[108]
Advances in neural information processing systems 32
Razavi A, Van den Oord A, Vinyals O (2019) Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems 32
2019
-
[109]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 658--666
Rezatofighi H, Tsoi N, Gwak J, Sadeghian A, Reid I, Savarese S (2019) Generalized intersection over union: A metric and a loss for bounding box regression. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 658--666
2019
-
[110]
In: International conference on machine learning, PMLR, pp 1530--1538
Rezende D, Mohamed S (2015) Variational inference with normalizing flows. In: International conference on machine learning, PMLR, pp 1530--1538
2015
-
[111]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10684--10695
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10684--10695
2022
-
[112]
arXiv preprint arXiv:241010792
Rout L, Chen Y, Ruiz N, Caramanis C, Shakkottai S, Chu WS (2024) Semantic image inversion and editing using rectified stochastic differential equations. arXiv preprint arXiv:241010792
2024
-
[113]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 10219--10228
Ruan L, Ma Y, Yang H, He H, Liu B, Fu J, Yuan NJ, Jin Q, Guo B (2023) Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 10219--10228
2023
-
[114]
arXiv preprint arXiv:240109084
Ruan L, Tian L, Huang C, Zhang X, Xiao X (2024) Univg: Towards unified-modal video generation. arXiv preprint arXiv:240109084
2024
-
[115]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 22500--22510
Ruiz N, Li Y, Jampani V, Pritch Y, Rubinstein M, Aberman K (2023) Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 22500--22510
2023
-
[116]
(2022 a ) Photorealistic text-to-image diffusion models with deep language understanding
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton EL, Ghasemipour K, Gontijo Lopes R, Karagol Ayan B, Salimans T, et al. (2022 a ) Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35:36479--36494
2022
-
[117]
IEEE transactions on pattern analysis and machine intelligence 45(4):4713--4726
Saharia C, Ho J, Chan W, Salimans T, Fleet DJ, Norouzi M (2022 b ) Image super-resolution via iterative refinement. IEEE transactions on pattern analysis and machine intelligence 45(4):4713--4726
2022
-
[118]
(2022) Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann C, Beaumont R, Vencu R, Gordon C, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M, et al. (2022) Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems 35:25278--25294
2022
-
[119]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 38, pp 8975--8983
Shang S, Shan Z, Liu G, Wang L, Wang X, Zhang Z, Zhang J (2024) Resdiff: Combining cnn and diffusion model for image super-resolution. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 38, pp 8975--8983
2024
-
[120]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8871--8879
Sheynin S, Polyak A, Singer U, Kirstain Y, Zohar A, Ashual O, Parikh D, Taigman Y (2024) Emu edit: Precise image editing via recognition and generation tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8871--8879
2024
-
[121]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8839--8849
Shi Y, Xue C, Liew JH, Pan J, Yan H, Zhang W, Tan VY, Bai S (2024) Dragdiffusion: Harnessing diffusion models for interactive point-based image editing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8839--8849
2024
-
[122]
arXiv preprint arXiv:240614555
Shuai X, Ding H, Ma X, Tu R, Jiang YG, Tao D (2024) A survey of multimodal-guided image editing with text-to-image diffusion models. arXiv preprint arXiv:240614555
2024
-
[123]
(2022) Make-a-video: Text-to-video generation without text-video data
Singer U, Polyak A, Hayes T, Yin X, An J, Zhang S, Hu Q, Yang H, Ashual O, Gafni O, et al. (2022) Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:220914792
2022
-
[124]
In: International conference on machine learning, PMLR, pp 2256--2265
Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In: International conference on machine learning, PMLR, pp 2256--2265
2015
-
[125]
arXiv preprint arXiv:201002502
Song J, Meng C, Ermon S (2020 a ) Denoising diffusion implicit models. arXiv preprint arXiv:201002502
2020
-
[126]
Advances in neural information processing systems 32
Song Y, Ermon S (2019) Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems 32
2019
-
[127]
Advances in neural information processing systems 33:12438--12448
Song Y, Ermon S (2020) Improved techniques for training score-based generative models. Advances in neural information processing systems 33:12438--12448
2020
-
[128]
arXiv preprint arXiv:201113456
Song Y, Sohl-Dickstein J, Kingma DP, Kumar A, Ermon S, Poole B (2020 b ) Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:201113456
2020
-
[129]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18310--18319
Song Y, Zhang Z, Lin Z, Cohen S, Price B, Zhang J, Kim SY, Aliaga D (2023) Objectstitch: Object compositing with diffusion model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18310--18319
2023
-
[130]
arXiv preprint arXiv:220308382
Su X, Song J, Meng C, Ermon S (2022) Dual diffusion implicit bridges for image-to-image translation. arXiv preprint arXiv:220308382
2022
-
[131]
In: Proceedings of the IEEE/CVF international conference on computer vision, pp 5117--5127
Su Z, Liu W, Yu Z, Hu D, Liao Q, Tian Q, Pietik \"a inen M, Liu L (2021) Pixel difference networks for efficient edge detection. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 5117--5127
2021
-
[132]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1--9
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1--9
2015
-
[133]
arXiv preprint arXiv:241115098 3
Tan Z, Liu S, Yang X, Xue Q, Wang X (2024) Ominicontrol: Minimal and universal control for diffusion transformer. arXiv preprint arXiv:241115098 3
2024
-
[134]
(2024) Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team G, Georgiev P, Lei VI, Burnell R, Bai L, Gulati A, Tanzer G, Vincent D, Pan Z, Wang S, et al. (2024) Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:240305530
2024
-
[135]
arXiv preprint arXiv:250221291
Tian X, Li W, Xu B, Yuan Y, Wang Y, Shen H (2025) Mige: A unified framework for multimodal instruction-based image generation and editing. arXiv preprint arXiv:250221291
2025
-
[136]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1921--1930
Tumanyan N, Geyer M, Bagon S, Dekel T (2023) Plug-and-play diffusion features for text-driven image-to-image translation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1921--1930
2023
-
[137]
In: International conference on machine learning, PMLR, pp 1747--1756
Van Den Oord A, Kalchbrenner N, Kavukcuoglu K (2016) Pixel recurrent neural networks. In: International conference on machine learning, PMLR, pp 1747--1756
2016
-
[138]
(2017) Neural discrete representation learning
Van Den Oord A, Vinyals O, et al. (2017) Neural discrete representation learning. Advances in neural information processing systems 30
2017
-
[139]
(2019) Diode: A dense indoor and outdoor depth dataset
Vasiljevic I, Kolkin N, Zhang S, Luo R, Wang H, Dai FZ, Daniele AF, Mostajabi M, Basart S, Walter MR, et al. (2019) Diode: A dense indoor and outdoor depth dataset. arXiv preprint arXiv:190800463
2019
-
[140]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 4146--4156
Vinker Y, Alaluf Y, Cohen-Or D, Shamir A (2023) Clipascene: Scene sketching with different types and levels of abstraction. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 4146--4156
2023
-
[141]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 22532--22541
Wallace B, Gokul A, Naik N (2023) Edict: Exact diffusion inversion via coupled transformations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 22532--22541
2023
-
[142]
arXiv preprint arXiv:230518047
Wang Q, Zhang B, Birsak M, Wonka P (2023 a ) Instructedit: Improving automatic masks for diffusion-based image editing with user instructions. arXiv preprint arXiv:230518047
2023
-
[143]
(2023 b ) Imagen editor and editbench: Advancing and evaluating text-guided image inpainting
Wang S, Saharia C, Montgomery C, Pont-Tuset J, Noy S, Pellegrini S, Onoe Y, Laszlo S, Fleet DJ, Soricut R, et al. (2023 b ) Imagen editor and editbench: Advancing and evaluating text-guided image inpainting. In: Proceedings of the IEEE/CVF conference on computer vision and pat...
2023
-
[144]
In: Proceedings of the IEEE/CVF international conference on computer vision, pp 10776--10785
Wang W, Feiszli M, Wang H, Tran D (2021) Unidentified video objects: A benchmark for dense, open-world segmentation. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 10776--10785
2021
-
[145]
In: European Conference on Computer Vision, Springer, pp 36--54
Wang Y, Lipson L, Deng J (2024) Sea-raft: Simple, efficient, accurate raft for optical flow. In: European Conference on Computer Vision, Springer, pp 36--54
2024
-
[146]
In: The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Ieee, vol 2, pp 1398--1402
Wang Z, Simoncelli EP, Bovik AC (2003) Multiscale structural similarity for image quality assessment. In: The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, Ieee, vol 2, pp 1398--1402
2003
-
[147]
In: The Thirteenth International Conference on Learning Representations
Wei C, Xiong Z, Ren W, Du X, Zhang G, Chen W (2024) Omniedit: Building image editing generalist models through specialist supervision. In: The Thirteenth International Conference on Learning Representations
2024
-
[148]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7378--7387
Wu CH, De la Torre F (2023) A latent space of stochastic diffusion models for zero-shot image editing and guidance. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7378--7387
2023
-
[149]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7623--7633
Wu JZ, Ge Y, Wang X, Lei SW, Gu Y, Shi Y, Hsu W, Shan Y, Qie X, Shou MZ (2023) Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 7623--7633
2023
-
[150]
arXiv preprint arXiv:250402160
Wu S, Huang M, Wu W, Cheng Y, Ding F, He Q (2025) Less-to-more generalization: Unlocking more controllability by in-context generation. arXiv preprint arXiv:250402160
2025
-
[151]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 13095--13105
Xia B, Zhang Y, Wang S, Wang Y, Wu X, Tian Y, Yang W, Van Gool L (2023) Diffir: Efficient diffusion model for image restoration. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 13095--13105
2023
-
[152]
arXiv preprint arXiv:241217098
Xia B, Zhang Y, Li J, Wang C, Wang Y, Wu X, Yu B, Jia J (2024) Dreamomni: Unified image generation and editing. arXiv preprint arXiv:241217098
2024
-
[153]
arXiv preprint arXiv:250306419
Xia T, Zhang Y, Zhang TLL (2025) Consistent image layout editing with diffusion models. arXiv preprint arXiv:250306419
2025
-
[154]
arXiv preprint arXiv:240911340
Xiao S, Wang Y, Zhou J, Yuan H, Xing X, Yan R, Wang S, Huang T, Liu Z (2024) Omnigen: Unified image generation. arXiv preprint arXiv:240911340
2024
-
[155]
In: Proceedings of the IEEE international conference on computer vision, pp 1395--1403
Xie S, Tu Z (2015) Holistically-nested edge detection. In: Proceedings of the IEEE international conference on computer vision, pp 1395--1403
2015
-
[156]
ACM Computing Surveys 57(2):1--42
Xing Z, Feng Q, Chen H, Dai Q, Hu H, Xu H, Wu Z, Jiang YG (2024) A survey on video diffusion models. ACM Computing Surveys 57(2):1--42
2024
-
[157]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18381--18391
Yang B, Gu S, Zhang B, Zhang T, Chen X, Sun X, Chen D, Wen F (2023 a ) Paint by example: Exemplar-based image editing with diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18381--18391
2023
-
[158]
ACM Computing Surveys 56(4):1--39
Yang L, Zhang Z, Song Y, Hong S, Xu R, Zhao Y, Zhang W, Cui B, Yang MH (2023 b ) Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys 56(4):1--39
2023
-
[159]
arXiv preprint arXiv:230806721
Ye H, Zhang J, Liu S, Han X, Yang W (2023) Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:230806721
2023
-
[160]
arXiv preprint arXiv:241115738
Yu Q, Chow W, Yue Z, Pan K, Wu Y, Wan X, Li J, Tang S, Zhang H, Zhuang Y (2024) Anyedit: Mastering unified high-quality image editing for any idea. arXiv preprint arXiv:241115738
2024
-
[161]
(2023) Mvimgnet: A large-scale dataset of multi-view images
Yu X, Xu M, Zhang Y, Liu H, Ye C, Wu Y, Yan Z, Zhu C, Xiong Z, Liang T, et al. (2023) Mvimgnet: A large-scale dataset of multi-view images. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 9150--9161
2023
-
[162]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8372--8382
Zeng J, Song D, Nie W, Tian H, Wang T, Liu AA (2024) Cat-dm: Controllable accelerated virtual try-on with diffusion model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 8372--8382
2024
-
[163]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1881--1889
Zhan X, Pan X, Liu Z, Lin D, Loy CC (2019) Self-supervised learning via conditional motion propagation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 1881--1889
2019
-
[164]
arXiv preprint arXiv:240919365
Zhan Z, Chen D, Mei JP, Zhao Z, Chen J, Chen C, Lyu S, Wang C (2024) Conditional image synthesis with diffusion models: A survey. arXiv preprint arXiv:240919365
2024
-
[165]
Advances in Neural Information Processing Systems 36
Zhang K, Mo L, Chen W, Sun H, Su Y (2024 a ) Magicbrush: A manually annotated dataset for instruction-guided image editing. Advances in Neural Information Processing Systems 36
2024
-
[166]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 3836--3847
Zhang L, Rao A, Agrawala M (2023) Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 3836--3847
2023
-
[167]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 586--595
Zhang R, Isola P, Efros AA, Shechtman E, Wang O (2018) The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 586--595
2018
-
[168]
(2024 b ) Hive: Harnessing human feedback for instructional visual editing
Zhang S, Yang X, Feng Y, Qin C, Chen CC, Yu N, Chen Z, Wang H, Savarese S, Ermon S, et al. (2024 b ) Hive: Harnessing human feedback for instructional visual editing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 9026--9036
2024
-
[169]
arXiv preprint arXiv:250108225
Zhang Y, Zhou X, Zeng Y, Xu H, Li H, Zuo W (2025) Framepainter: Endowing interactive image editing with video diffusion priors. arXiv preprint arXiv:250108225
2025
-
[170]
arXiv preprint arXiv:160903126
Zhao J (2016) Energy-based generative adversarial network. arXiv preprint arXiv:160903126
2016
-
[171]
Advances in Neural Information Processing Systems 36
Zhao S, Chen D, Chen YC, Bao J, Hao S, Yuan L, Wong KYK (2024 a ) Uni-controlnet: All-in-one control to text-to-image diffusion models. Advances in Neural Information Processing Systems 36
2024
-
[172]
arXiv preprint arXiv:240515769
Zhao X, Guan J, Fan C, Xu D, Lin Y, Pan H, Feng P (2024 b ) Fastdrag: Manipulate anything in one step. arXiv preprint arXiv:240515769
2024
-
[173]
IEEE transactions on pattern analysis and machine intelligence 40(6):1452--1464
Zhou B, Lapedriza A, Khosla A, Oliva A, Torralba A (2017 a ) Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence 40(6):1452--1464
2017
-
[174]
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 633--641
Zhou B, Zhao H, Puig X, Fidler S, Barriuso A, Torralba A (2017 b ) Scene parsing through ade20k dataset. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 633--641
2017
-
[175]
arXiv preprint arXiv:221111018
Zhou D, Wang W, Yan H, Lv W, Zhu Y, Feng J (2022) Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:221111018
2022
-
[176]
In: Proceedings of the IEEE international conference on computer vision, pp 2223--2232
Zhu JY, Park T, Isola P, Efros AA (2017) Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proceedings of the IEEE international conference on computer vision, pp 2223--2232
2017
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.