REVIEW 4 major objections 5 minor 132 references
Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper claims that unpaired image-to-image translation tasks should be chosen by how much source content must survive, and that this choice is best made before picking a model.
desk verdict A useful review taxonomy and a modest Sim2Real benchmark, but the cross-paper FID/KID rankings in Section 6 are not commensurable and need to be reframed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the three-way content-preservation categorization: Fully Content preserving, Partially Content preserving, and Non-Content preserving, defined by how much of the source image must be reproduced in the output. Around this axis the paper organizes an architectural taxonomy of GAN-based, VAE-based, diffusion, flow-based, and transformer models, and a new Sim2Real benchmark built from simulated vehicle images and real images from MIO, Visdrone, and VeRi, with FID, KID, IS, NDB, JSD, and LPIPS as evaluation tools. The taxonomy supplies the structure of the review, and the benchmark supplies a reusable testbed for the content-preserving label.
What would settle it
Re-run a handful of the compared models end-to-end on the same train and validation split of the Sim2Real benchmark and of one Fully Content preserving task such as GTA to Cityscape, computing FID and KID with a single script. If the best model by the copied numbers is not the best model in the unified run, the paper's comparative conclusions collapse.
Extended reading notes
Core claim
On its own terms, the central claim is that content preservation is not a side effect but a design axis. The paper shows that tasks and datasets can be sorted into Fully Content preserving, Partially Content preserving, and Non-Content preserving, and that different models perform best in different categories. For example, DRIT++ achieves the lowest FID for GTA to Cityscape, CUT achieves the best KID on the same task, UNSB achieves the lowest FID for Horse to Zebra, and DCLGAN performs best on Label to Cityscapes. The paper also introduces a Sim2Real benchmark and reports evaluations of CycleGAN, StyleFlow, GcGAN, VSAIT, SRUNIT, and UNSB on it, showing that hyperparameter choices such as batch size and loss weights visibly change content preservation and image quality.
Load-bearing premise
The evaluation tables compare FID and KID scores copied from each model's original paper, so the comparison assumes those numbers were produced under fairly comparable training conditions; if they were not, the rankings and the conclusions drawn from them are not reliable.
Editorial extensions
If this is right
- If the categorization is right, practitioners can narrow a model search by first deciding whether a task is fully, partially, or non-content preserving.
- The Sim2Real benchmark gives later researchers a shared vehicle-focused dataset on which content-preserving methods can be compared directly.
- The paper's tables show that no single model dominates all three categories, so reporting results per content-preservation category could become a useful convention.
- The hyperparameter studies on CycleGAN, GcGAN, and StyleFlow suggest that content preservation in a given model is tunable through batch size and loss weights, not fixed by architecture alone.
Reading between the lines
- A testable extension would measure content preservation directly, for instance by comparing segmentation or keypoint consistency between source and translated images, rather than relying only on FID and KID.
- The taxonomy could plausibly extend to newer diffusion and transformer models that were not fully benchmarked here, since those architectures also face the same content-versus-style tradeoff.
- The benchmark could be reused for partially content-preserving tasks by adding tasks that deliberately change object class or layout, which would stress-test where the FCP/PCP boundary actually lies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a review of unpaired image-to-image translation organized around the degree to which the source image content must be preserved. It introduces a three-way task taxonomy (Fully Content Preserving, Partially Content Preserving, Non-Content Preserving), surveys around 70 models grouped by architecture, summarizes datasets and evaluation metrics, compiles FID and KID results from the original papers, and presents a new Sim2Real benchmark for content-preserving simulation-to-real translation. The paper concludes that the degree of content preservation should be considered when selecting an I2I model for a given application.
Significance. The proposed FCP/PCP/NCP taxonomy addresses a real and practically useful distinction that is often implicit in I2I papers, and the broad survey of models, datasets, and metrics provides a useful entry point for practitioners. The Sim2Real benchmark, combining VisDrone, MIO, and VeRi data with a simulated vehicle-image domain, is a potentially reusable evaluation resource, and the paper is explicit about the hyperparameters and metrics used in its own experiments. However, the central empirical claim that the taxonomy supports model selection rests on Section 6, where FID/KID values are copied from heterogeneous original papers without a shared evaluation protocol, and on the Sim2Real section, which has internal inconsistencies (the stated eight models are not the reported six) and no released data, code, or error bars. Because these issues affect the main practical conclusion, the paper needs revision before the claim can be accepted. I see no circular reasoning or fabricated entities; the weaknesses are methodological and presentation-related rather than conceptual.
major comments (4)
- [§6, Tables 4–5] The cross-model rankings in Section 6 are not commensurable because the FID and KID values are taken from the models' own articles under different training protocols, data splits, resolutions, and metric implementations, as the text itself states: 'These results are taken from the models' own articles.' Consequently, statements such as 'DRIT++ achieve the lowest FID' (Table 4, GTA-to-Cityscape) and 'CUT performs better than other methods' (Table 5, GTA-to-Cityscape) are not supported as general claims about model quality. The authors should either re-run all compared methods on a common split with a common metric implementation, or explicitly rephrase these statements as reports of the original papers' numbers rather than as findings of this review.
- [§7.2, Table 12] The evaluation section states that eight models were selected (CycleGAN, DRIT, GcGAN, StyleFlow, SRUNIT, VSAIT, UNSB), but the text lists only seven model names and Table 12 reports results for only six models, with no DRIT row and no explanation for its absence. This discrepancy makes the benchmark evaluation incomplete and raises the question of whether the missing model's results were omitted selectively. The authors should specify the exact set of models, justify any exclusions, and either add the missing results or remove the claim of eight models.
- [§7.1, §7.2] The Sim2Real benchmark lacks a stated train/test split, repeated runs with seeds, error bars, and a released data/code artifact. For example, in §7.2.1 the authors say CycleGAN was 'trained and tested on Sim2Real dataset' for 20 epochs, but no held-out evaluation set is described, so the reported FID, IS, NDB, JSD, and LPIPS numbers cannot be interpreted as generalization measures. The conclusion that 'changing hyperparameter λ from 10 to 5 do not increase the performance' is therefore not verifiable. The authors should describe the exact split, report multiple seeds with variance, and release the dataset construction code and evaluation scripts.
- [§4, Table 3] The FCP/PCP/NCP assignment is not operationalized: the paper gives qualitative descriptions of the three categories but no quantitative or procedural rule by which a reader can assign a new task to a category. Some assignments in Table 3 also appear inconsistent with the definitions, for instance Label2Cityscape is listed as PCP while Keypoint2Photo is listed as NCP, although both involve structured source information being transformed into a target image. Without an explicit annotation protocol or inter-annotator agreement, the taxonomy cannot be applied reproducibly by future users, which weakens the paper's central contribution.
minor comments (5)
- [§7.2] The sentence listing selected models ends with 'an UNSB [38]'; this should read 'and UNSB [38]'.
- [Table 11] The column header 'V an' contains a stray space and should be 'Van'.
- [Throughout] The text contains several typos and grammatical errors, for example 'benckmark' in §7.2.1, 'metircs' in the Table 12 caption, and 'do not increase the performance results' in §7.2.1; a careful proofread is needed.
- [Figure 10 caption and text] The caption says the first row is simulation images and the second row is real images, but the text in §7.1 says 'images of 5 classes in the VeRi data set have been randomly sampled and demonstrated'; the figure and text should be reconciled to avoid confusion about which domain is shown.
- [Table 9] Several KID values in Table 9 are on the order of 50–100, which is far outside the typical range for KID on standard benchmarks (usually below 0.1); the authors should state whether these values are scaled (e.g., multiplied by 100) or whether a different implementation was used.
Circularity Check
No significant circularity: the taxonomy and benchmark are organizational and empirical contributions, with no fitted parameter renamed as a prediction and no load-bearing self-citation.
full rationale
The paper is a review plus an empirical benchmark. The FCP/PCP/NCP categories are descriptive labels assigned to tasks according to the stated degree of content preservation; they are not derived from the benchmark results, and no equation defines one category in terms of another. Section 6 rankings are drawn from the original model papers, as the text acknowledges ('These results are taken from the models' own articles'), which creates a comparability and validity limitation but not circularity: the numbers are external evidence, not outputs of a model fitted to the taxonomy. Section 7's Sim2Real evaluation reports measured FID, IS, NDB, JSD, and LPIPS values for eight named models; the choice of models and hyperparameters is a design decision rather than a fitted parameter later renamed as a prediction. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation. The concluding recommendation that content-preservation extent should inform model choice restates the organizing principle of the taxonomy, but it is a practical guideline rather than a derived result claimed to be proven by the benchmark. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption I2I tasks can be cleanly partitioned into Fully, Partially, and Non-Content preserving categories based on source and target domains.
- domain assumption FID and KID scores reported in different original papers are comparable enough to rank methods in the survey tables.
- domain assumption The combination of Visdrone, MIO, VeRi, and the simulated vehicle dataset constitutes a valid Sim2Real benchmark for content-preserving translation.
Cite this review
Pith. "Pith review of Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review." pith.science (2026). https://pith.science/paper/AD4C5HGT
@misc{pith2026250208667,
author = {Pith},
title = {Pith review of: Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/AD4C5HGT}},
note = {Machine review of arXiv:2502.08667}
}
read the original abstract
Image-to-image translation (I2I) transforms an image from a source domain to a target domain while preserving source content. Most computer vision applications are in the field of image-to-image translation, such as style transfer, image segmentation, and photo enhancement. The degree of preservation of the content of the source images in the translation process can be different according to the problem and the intended application. From this point of view, in this paper, we divide the different tasks in the field of image-to-image translation into three categories: Fully Content preserving, Partially Content preserving, and Non-Content preserving. We present different tasks, datasets, methods, results of methods for these three categories in this paper. We make a categorization for I2I methods based on the architecture of different models and study each category separately. In addition, we introduce well-known evaluation criteria in the I2I translation field. Specifically, nearly 70 different I2I models were analyzed, and more than 10 quantitative evaluation metrics and 30 distinct tasks and datasets relevant to the I2I translation problem were both introduced and assessed. Translating from simulation to real images could be well viewed as an application of fully content preserving or partially content preserving unsupervised image-to-image translation methods. So, we provide a benchmark for Sim-to-Real translation, which can be used to evaluate different methods. In general, we conclude that because of the different extent of the obligation to preserving content in various applications, it is better to consider this issue in choosing a suitable I2I model for a specific application.
Figures
Figures from the paper (15 more)
Reference graph
Works this paper leans on
-
[1]
& Efros, A
Zhu, J., Park, T., Isola, P. & Efros, A. Unpaired image-to-image trans- lation using cycle-consistent adversarial networks. Proceedings Of The IEEE International Conference On Computer Vision . pp. 2223-2232 (2017)
2017
-
[2]
& Efros, A
Isola, P., Zhu, J., Zhou, T. & Efros, A. Image-to-image translation with conditional adversarial networks. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 1125-1134 (2017)
2017
-
[3]
& Cucchiara, R
Tomei, M., Cornia, M., Baraldi, L. & Cucchiara, R. Art2real: Un- folding the reality of artworks via semantically-aware image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 5849-5859 (2019)
2019
-
[4]
& Chuang, Y
Chang, H., Wang, Z. & Chuang, Y. Domain-specific mappings for gen- erative adversarial style transfer. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16 . pp. 573-589 (2020)
2020
-
[5]
& Gupta, V
Kumar, P. & Gupta, V. Unpaired Image-to-Image Translation Based Artwork Restoration Using Generative Adversarial Networks. Interna- tional Conference On Intelligent Manufacturing And Energy Sustain- ability. pp. 581-591 (2023) 58
2023
-
[6]
& Kim, A
Cho, Y., Malav, R., Pandey, G. & Kim, A. DehazeGAN: underwa- ter haze image restoration using unpaired image-to-image translation. IF AC-PapersOnLine. 52, 82-85 (2019)
2019
-
[7]
& Huang, J
Guo, X., Wang, Z., Yang, Q., Lv, W., Liu, X., Wu, Q. & Huang, J. Gan-based virtual-to-real image translation for urban scene semantic segmentation. Neurocomputing. 394 pp. 127-135 (2020)
2020
-
[8]
& Wong, H
Li, R., Cao, W., Jiao, Q., Wu, S. & Wong, H. Simplified unsuper- vised image translation for semantic segmentation adaptation. Pattern Recognition. 105 pp. 107343 (2020)
2020
Show all 132 references
-
[9]
& Kim, K
Murez, Z., Kolouri, S., Kriegman, D., Ramamoorthi, R. & Kim, K. Image to image translation for domain adaptation. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 4500-4509 (2018)
2018
-
[10]
& Shen, H
Li, J., Lu, K., Huang, Z., Zhu, L. & Shen, H. Heterogeneous domain adaptation through progressive alignment.IEEE Transactions On Neu- ral Networks And Learning Systems . 30, 1381-1391 (2018)
2018
-
[11]
& Cook, D
Wilson, G. & Cook, D. A survey of unsupervised deep domain adap- tation. ACM Transactions On Intelligent Systems And Technology (TIST). 11, 1-46 (2020)
2020
-
[12]
& Yang, M
Lee, H., Tseng, H., Huang, J., Singh, M. & Yang, M. Diverse image-to- image translation via disentangled representations. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 35-51 (2018)
2018
-
[13]
& Gong, M
Yi, Z., Zhang, H., Tan, P. & Gong, M. Dualgan: Unsupervised dual learning for image-to-image translation. Proceedings Of The IEEE In- ternational Conference On Computer Vision . pp. 2849-2857 (2017)
2017
-
[14]
& Kautz, J
Huang, X., Liu, M., Belongie, S. & Kautz, J. Multimodal unsupervised image-to-image translation. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 172-189 (2018)
2018
-
[15]
& Chen, Z
Pang, Y., Lin, J., Qin, T. & Chen, Z. Image-to-image translation: Methods and applications. IEEE Transactions On Multimedia . 24 pp. 3859-3881 (2021) 59
2021
-
[16]
& Choo, J
Cho, W., Choi, S., Park, D., Shin, I. & Choo, J. Image-to-image trans- lation via group-wise deep whitening-and-coloring transformation. Pro- ceedings Of The IEEE/CVF Conference On Computer Vision And Pat- tern Recognition. pp. 10639-10647 (2019)
2019
-
[17]
& Prakash, A
Theiss, J., Leverett, J., Kim, D. & Prakash, A. Unpaired image transla- tion via vector symbolic architectures. European Conference On Com- puter Vision . pp. 17-32 (2022)
2022
-
[18]
Fan, W., Chen, J., Ma, J., Hou, J. & Yi, S. Styleflow for content-fixed image to image translation. ArXiv Preprint arXiv:2207.01909 . (2022)
2022 arXiv
-
[19]
& Tao, D
Fu, H., Gong, M., Wang, C., Batmanghelich, K., Zhang, K. & Tao, D. Geometry-consistent generative adversarial networks for one-sided unsupervised domain mapping. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2427-2436 (2019)
2019
-
[20]
& Fang, B
Chen, R., Huang, W., Huang, B., Sun, F. & Fang, B. Reusing discrimi- nators for encoding: Towards unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 8168-8177 (2020)
2020
-
[21]
& Zhang, K
Xie, S., Gong, M., Xu, Y. & Zhang, K. Unaligned image-to-image translation by learning to reweight. Proceedings Of The IEEE/CVF International Conference On Computer Vision. pp. 14174-14184 (2021)
2021
-
[22]
& Abbeel, P
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W. & Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. 2017 IEEE/RSJ International Conference On Intelligent Robots And Systems (IROS) . pp. 23-30 (2017)
2017
-
[23]
& Bengio, Y
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. & Bengio, Y. Generative adversarial nets. Advances In Neural Information Processing Systems . 27 (2014)
2014
-
[24]
& Metaxas, D
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X. & Metaxas, D. Stackgan++: Realistic image synthesis with stacked generative ad- versarial networks. IEEE Transactions On Pattern Analysis And Ma- chine Intelligence . 41, 1947-1962 (2018) 60
2018
-
[25]
& Papa, J
De Rosa, G. & Papa, J. A survey on text generation using generative adversarial networks. Pattern Recognition. 119 pp. 108098 (2021)
2021
-
[26]
& Liang, Z
Liu, Z., Wang, J. & Liang, Z. Catgan: Category-aware generative ad- versarial networks with hierarchical evolutionary learning for category text generation. Proceedings Of The AAAI Conference On Artificial Intelligence. 34, 8425-8432 (2020)
2020
-
[27]
& Mohammadi, G
Aldausari, N., Sowmya, A., Marcus, N. & Mohammadi, G. Video generative adversarial networks: a review. ACM Computing Surveys (CSUR). 55, 1-25 (2022)
2022
-
[28]
& Tan, M
Chen, Q., Wu, Q., Chen, J., Wu, Q., Hengel, A. & Tan, M. Scripted video generation with a bottom-up generative adversarial network. IEEE Transactions On Image Processing . 29 pp. 7454-7467 (2020)
2020
-
[29]
Conditional generative adversarial nets
Mirza, M. Conditional generative adversarial nets. ArXiv Preprint arXiv:1411.1784. (2014)
2014 arXiv
-
[30]
Conditional generative adversarial nets for convolutional face generation
Gauthier, J. Conditional generative adversarial nets for convolutional face generation. Class Project For Stanford CS231N: Convolutional Neural Networks For Visual Recognition, Winter Semester . 2014, 2 (2014)
2014
-
[31]
Auto-encoding variational bayes
Kingma, D. Auto-encoding variational bayes. ArXiv Preprint arXiv:1312.6114. (2013)
2013 arXiv
-
[32]
& Ferrari, V
Pumarola, A., Popov, S., Moreno-Noguer, F. & Ferrari, V. C-flow: Conditional generative flow models for images and 3d point clouds. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 7949-7958 (2020)
2020
-
[33]
& Mohamed, S
Rezende, D. & Mohamed, S. Variational inference with normalizing flows. International Conference On Machine Learning . pp. 1530-1538 (2015)
2015
-
[34]
& Abbeel, P
Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. Advances In Neural Information Processing Systems. 33 pp. 6840-6851 (2020) 61
2020
-
[35]
& Breckon, T
Sasaki, H., Willcocks, C. & Breckon, T. Unit-ddpm: Unpaired im- age translation with denoising diffusion probabilistic models. ArXiv Preprint arXiv:2104.05358. (2021)
2021 arXiv
-
[36]
& Zhu, J
Zhao, M., Bao, F., Li, C. & Zhu, J. Egsde: Unpaired image-to- image translation via energy-guided stochastic differential equations. Advances In Neural Information Processing Systems. 35 pp. 3609-3623 (2022)
2022
-
[37]
& Ermon, S
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J. & Ermon, S. SDEdit: Guided Image Synthesis and Editing with Stochastic Differ- ential Equations. (2022)
2022
-
[38]
Kim, B., Kwon, G., Kim, K. & Ye, J. Unpaired Image-to- Image Translation via Neural Schrodinger Bridge. ArXiv Preprint arXiv:2305.15086. (2023)
2023 arXiv
-
[39]
Attention is all you need
Vaswani, A. Attention is all you need. Advances In Neural Information Processing Systems. (2017)
2017
-
[41]
& Koltun, V
Ranftl, R., Bochkovskiy, A. & Koltun, V. Vision transformers for dense prediction. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 12179-12188 (2021)
2021
-
[42]
& Others A survey on vision transformer
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y. & Others A survey on vision transformer. IEEE Transactions On Pattern Analysis And Machine Intelligence. 45, 87-110 (2022)
2022
-
[43]
& Shah, M
Khan, S., Naseer, M., Hayat, M., Zamir, S., Khan, F. & Shah, M. Transformers in vision: A survey. ACM Computing Surveys (CSUR) . 54, 1-41 (2022)
2022
-
[44]
& Others An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S. & Others An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv Preprint arXiv:2010.11929 . (2020) 62
2020 arXiv
-
[45]
Flow-based Deep Generative Models
Weng, L. Flow-based Deep Generative Models. (2018), https://lilianweng.github.io/posts/2018-10-13-flow-models/, Accessed: 2024-08-27
2018
-
[46]
& Kim, J
Kim, T., Cha, M., Kim, H., Lee, J. & Kim, J. Learning to discover cross-domain relations with generative adversarial networks. Interna- tional Conference On Machine Learning . pp. 1857-1865 (2017)
2017
-
[47]
& Wolf, L
Benaim, S. & Wolf, L. One-sided unsupervised domain mapping. Ad- vances In Neural Information Processing Systems . 30 (2017)
2017
-
[48]
& Bottou, L
Arjovsky, M., Chintala, S. & Bottou, L. Wasserstein generative adver- sarial networks. International Conference On Machine Learning . pp. 214-223 (2017)
2017
-
[49]
& Van Gool, L
Ma, L., Jia, X., Georgoulis, S., Tuytelaars, T. & Van Gool, L. Ex- emplar guided unsupervised image-to-image translation with semantic consistency. ArXiv Preprint arXiv:1805.11145 . (2018)
2018 arXiv
-
[50]
& Zhu, J
Park, T., Liu, M., Wang, T. & Zhu, J. Semantic image synthesis with spatially-adaptive normalization. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2337-2346 (2019)
2019
-
[51]
Jia, Z., Yuan, B., Wang, K., Wu, H., Clifford, D., Yuan, Z. & Su, H. Semantically robust unpaired image translation for data with un- matched semantics statistics. Proceedings Of The IEEE/CVF Interna- tional Conference On Computer Vision . pp. 14273-14283 (2021)
2021
-
[52]
& Cai, J
Zheng, C., Cham, T. & Cai, J. The spatially-correlative loss for various image translation tasks. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 16407-16417 (2021)
2021
-
[53]
& Chinnam, R
Emami, H., Aliabadi, M., Dong, M. & Chinnam, R. SPA-GAN: Spatial attention GAN for image-to-image translation. IEEE Transactions On Multimedia. 23 pp. 391-401 (2020)
2020
-
[54]
& Dong, H
Zhao, Y., Wu, R. & Dong, H. Unpaired image-to-image translation using adversarial consistency loss. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16 . pp. 800-815 (2020) 63
2020
-
[55]
Liu, Y., Wang, H., Yue, Y. & Lu, F. Separating content and style for unsupervised image-to-image translation. ArXiv Preprint arXiv:2110.14404. (2021)
2021 arXiv
-
[56]
& Belongie, S
Huang, X. & Belongie, S. Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings Of The IEEE Interna- tional Conference On Computer Vision . pp. 1501-1510 (2017)
2017
-
[57]
& Wen, F
Zhou, X., Zhang, B., Zhang, T., Zhang, P., Bao, J., Chen, D., Zhang, Z. & Wen, F. Cocosnet v2: Full-resolution correspondence learning for image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 11465-11475 (2021)
2021
-
[58]
& Zhu, J
Park, T., Efros, A., Zhang, R. & Zhu, J. Contrastive learning for unpaired image-to-image translation. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Pro- ceedings, Part IX 16 . pp. 319-345 (2020)
2020
-
[59]
& Armin, M
Han, J., Shoeiby, M., Petersson, L. & Armin, M. Dual contrastive learning for unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recog- nition. pp. 746-755 (2021)
2021
-
[60]
Wang, W., Zhou, W., Bao, J., Chen, D. & Li, H. Instance-wise hard negative example generation for contrastive learning in unpaired image- to-image translation. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 14020-14029 (2021)
2021
-
[61]
& Charette, R
Pizzati, F., Lalonde, J. & Charette, R. Manifest: Manifold deforma- tion for few-shot image translation. European Conference On Computer Vision. pp. 440-456 (2022)
2022
-
[62]
Lin, J., Wang, Y., Chen, Z. & He, T. Learning to transfer: unsuper- vised domain translation via meta-learning. Proceedings Of The AAAI Conference On Artificial Intelligence . 34, 11507-11514 (2020)
2020
-
[63]
Koksal, A. & Lu, S. Rf-gan: A light and reconfigurable network for unpaired image-to-image translation. Proceedings Of The Asian Con- ference On Computer Vision . (2020) 64
2020
-
[64]
Ye, K., Ye, Y., Yang, M. & Hu, B. Independent encoder for deep hierarchical unsupervised image-to-image translation. ArXiv Preprint arXiv:2107.02494. (2021)
2021 arXiv
-
[65]
& Chen, Q
Zhao, J., Lee, F., Hu, C., Yu, H. & Chen, Q. LDA-GAN: Lightweight domain-attention GAN for unpaired image-to-image translation. Neu- rocomputing. 506 pp. 355-368 (2022)
2022
-
[66]
& Wang, Z
Deng, H., Wu, Q., Huang, H., Yang, X. & Wang, Z. Involution- GAN: lightweight GAN with involution for unsupervised image-to- image translation. Neural Computing And Applications . 35, 16593- 16605 (2023)
2023
-
[67]
& Huang, H
Ganjdanesh, A., Gao, S., Alipanah, H. & Huang, H. Compressing image-to-image translation gans using local density structures on their learned manifold. Proceedings Of The AAAI Conference On Artificial Intelligence. 38, 12118-12126 (2024)
2024
-
[68]
& Krishnaswamy, S
Amodio, M. & Krishnaswamy, S. Travelgan: Image-to-image transla- tion by transformation vector learning. Proceedings Of The Ieee/cvf Conference On Computer Vision And Pattern Recognition . pp. 8983- 8992 (2019)
2019
-
[69]
Siamese neural networks: An overview
Chicco, D. Siamese neural networks: An overview. Artificial Neural Networks. pp. 73-94 (2021)
2021
-
[71]
& Chang, X
Cao, Y., Yao, L., Pan, L., Sheng, Q. & Chang, X. Guided Image-to- Image Translation by Discriminator-Generator Communication. IEEE Transactions On Multimedia. (2023)
2023
-
[72]
& Choo, J
Choi, Y., Choi, M., Kim, M., Ha, J., Kim, S. & Choo, J. Stargan: Uni- fied generative adversarial networks for multi-domain image-to-image translation. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 8789-8797 (2018)
2018
-
[73]
Choi, Y., Uh, Y., Yoo, J. & Ha, J. Stargan v2: Diverse image synthesis for multiple domains. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 8188-8197 (2020) 65
2020
-
[74]
Yu, X., Chen, Y., Liu, S., Li, T. & Li, G. Multi-mapping image-to- image translation via learning disentanglement. Advances In Neural Information Processing Systems. 32 (2019)
2019
-
[75]
Liu, R., Ge, Y., Choi, C., Wang, X. & Li, H. Divco: Diverse condi- tional image synthesis via contrastive generative adversarial network. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 16377-16386 (2021)
2021
-
[76]
& Yang, M
Mao, Q., Lee, H., Tseng, H., Ma, S. & Yang, M. Mode seeking gener- ative adversarial networks for diverse image synthesis. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recog- nition. pp. 1429-1437 (2019)
2019
-
[77]
& Yan, Y
Tang, H., Xu, D., Sebe, N., Wang, Y., Corso, J. & Yan, Y. Multi- channel attention selection gan with cascaded semantic guidance for cross-view image translation. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2417-2426 (2019)
2019
-
[78]
& Yan, Y
Tang, H., Xu, D., Sebe, N. & Yan, Y. Attention-Guided Generative Adversarial Networks for Unsupervised Image-to-Image Translation. CoRR. abs/1903.12296 (2019), http://arxiv.org/abs/1903.12296
2019 arXiv
-
[79]
& Sebe, N
Tang, H., Liu, H., Xu, D., Torr, P. & Sebe, N. Attentiongan: Unpaired image-to-image translation using attention-guided generative adversar- ial networks. IEEE Transactions On Neural Networks And Learning Systems. 34, 1972-1987 (2021)
2021
-
[80]
& Kim, K
Mejjati, Y., Richardt, C., Tompkin, J., Cosker, D. & Kim, K. Unsu- pervised Attention-guided Image to Image Translation. (2018)
2018
-
[81]
U-gat-it: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation
Kim, J. U-gat-it: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation. ArXiv Preprint arXiv:1907.10830 . (2019)
2019 arXiv
-
[82]
& Miao, C
Zhan, F., Yu, Y., Wu, R., Zhang, J., Cui, K., Xiao, A., Lu, S. & Miao, C. Bi-level feature alignment for versatile image translation and manipulation. European Conference On Computer Vision. pp. 224-241 (2022) 66
2022
-
[83]
& De Guevara, M
Cazenavette, G. & De Guevara, M. MixerGAN: An MLP-based ar- chitecture for unpaired image-to-image translation. ArXiv Preprint arXiv:2105.14110. (2021)
2021 arXiv
-
[84]
& Oth- ers Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Un- terthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J. & Oth- ers Mlp-mixer: An all-mlp architecture for vision. Advances In Neural Information Processing Systems. 34 pp. 24261-24272 (2021)
2021
-
[85]
& Hao, Y
Lai, X., Bai, X. & Hao, Y. Unsupervised generative adversarial net- works with cross-model weight transfer mechanism for image-to-image translation. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 1814-1822 (2021)
2021
-
[86]
& Kautz, J
Liu, M., Breuel, T. & Kautz, J. Unsupervised image-to-image trans- lation networks. Advances In Neural Information Processing Systems . 30 (2017)
2017
-
[87]
& Loy, C
Wu, W., Cao, K., Li, C., Qian, C. & Loy, C. Transgaga: Geometry- aware unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 8012-8021 (2019)
2019
-
[88]
& Liu, M
Saito, K., Saenko, K. & Liu, M. Coco-funit: Few-shot unsupervised image translation with a content conditioned style encoder. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 . pp. 382-398 (2020)
2020
-
[89]
& Han, S
Han, G., Min, J. & Han, S. EM-LAST: Effective Multidimensional Latent Space Transport for an Unpaired Image-to-Image Translation With an Energy-Based Model. IEEE Access. 10 pp. 72839-72849 (2022)
2022
-
[90]
& Dhariwal, P
Kingma, D. & Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. Advances In Neural Information Processing Systems . 31 (2018)
2018
-
[91]
Very deep convolutional networks for large-scale image recognition
Simonyan, K. Very deep convolutional networks for large-scale image recognition. ArXiv Preprint arXiv:1409.1556 . (2014)
2014 arXiv
-
[92]
& Liu, Z
Fan, W., Chen, J. & Liu, Z. Hierarchy Flow For High-Fidelity Image- to-Image Translation. ArXiv Preprint arXiv:2308.06909 . (2023) 67
2023 arXiv
-
[93]
& Singh, V
Sun, H., Mehta, R., Zhou, H., Huang, Z., Johnson, S., Prabhakaran, V. & Singh, V. Dual-glow: Conditional flow-based generative model for modality transfer. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 10611-10620 (2019)
2019
-
[94]
& Wang, Z
Zheng, W., Li, Q., Zhang, G., Wan, P. & Wang, Z. Ittr: Un- paired image-to-image translation with transformers. ArXiv Preprint arXiv:2203.16015. (2022)
2022 arXiv
-
[95]
& Kim, S
Kim, S., Baek, J., Park, J., Kim, G. & Kim, S. InstaFormer: Instance- aware image-to-image translation with transformer.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 18321-18331 (2022)
2022
-
[96]
& Ren, Y
Torbunov, D., Huang, Y., Yu, H., Huang, J., Yoo, S., Lin, M., Viren, B. & Ren, Y. Uvcgan: Unet vision transformer cycle-consistent gan for unpaired image-to-image translation. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Vision . pp. 702-712 (2023)
2023
-
[97]
& Schiele, B
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benen- son, R., Franke, U., Roth, S. & Schiele, B. The cityscapes dataset for semantic urban scene understanding. Proceedings Of The IEEE Con- ference On Computer Vision And Pattern Recognition . pp. 3213-3223 (2016)
2016
-
[98]
& Torralba, A
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A. & Torralba, A. Scene parsing through ade20k dataset. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 633- 641 (2017)
2017
-
[99]
& Wyeth, G
Milford, M. & Wyeth, G. SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights.2012 IEEE International Conference On Robotics And Automation . pp. 1643-1649 (2012)
2012
-
[100]
& Darrell, T
Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V. & Darrell, T. Bdd100k: a diverse driving video database with scalable annotation tooling. 2018. ArXiv Preprint arXiv:1805.04687 . (1805) 68
2018 arXiv
-
[101]
& Tang, X
Liu, Z., Luo, P., Wang, X. & Tang, X. Deep learning face attributes in the wild. Proceedings Of The IEEE International Conference On Computer Vision . pp. 3730-3738 (2015)
2015
-
[102]
& Tang, X
Liu, Z., Luo, P., Qiu, S., Wang, X. & Tang, X. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. Pro- ceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 1096-1104 (2016)
2016
-
[103]
& Belongie, S
Poursaeed, O., Matera, T. & Belongie, S. Vision-based real estate price estimation. Machine Vision And Applications . 29, 667-676 (2018)
2018
-
[104]
&ˇS´ ara, R
Tyleˇ cek, R. &ˇS´ ara, R. Spatial pattern templates for recognition of ob- jects with regular structure. Pattern Recognition: 35th German Con- ference, GCPR 2013, Saarbr¨ ucken, Germany, September 3-6, 2013. Proceedings 35. pp. 364-374 (2013)
2013
-
[105]
& Aila, T
Karras, T., Laine, S. & Aila, T. A style-based generator architecture for generative adversarial networks. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 4401- 4410 (2019)
2019
-
[106]
& Lazebnik, S
Plummer, B., Wang, L., Cervantes, C., Caicedo, J., Hockenmaier, J. & Lazebnik, S. Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models. Proceedings Of The IEEE International Conference On Computer Vision . pp. 2641-2649 (2015)
2015
-
[107]
& Koltun, V
Richter, S., Vineet, V., Roth, S. & Koltun, V. Playing for data: Ground truth from computer games.Computer Vision–ECCV 2016: 14th Euro- pean Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14. pp. 102-118 (2016)
2016
-
[108]
& Fei-Fei, L
Deng, J., Dong, W., Socher, R., Li, L., Li, K. & Fei-Fei, L. Imagenet: A large-scale hierarchical image database. 2009 IEEE Conference On Computer Vision And Pattern Recognition . pp. 248-255 (2009)
2009
-
[109]
& Urtasun, R
Geiger, A., Lenz, P. & Urtasun, R. Are we ready for autonomous driv- ing? the kitti vision benchmark suite. 2012 IEEE Conference On Com- puter Vision And Pattern Recognition . pp. 3354-3361 (2012) 69
2012
-
[110]
& Fergus, R
Silberman, N., Hoiem, D., Kohli, P. & Fergus, R. Indoor segmenta- tion and support inference from rgbd images. Computer Vision–ECCV 2012: 12th European Conference On Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12 . pp. 746-760 (2012)
2012
-
[111]
& Sheikh, Y
Cao, Z., Hidalgo, G., Simon, T., Wei, S. & Sheikh, Y. Openpose: real- time multi-person 2d pose estimation using part affinity fields (2018). ArXiv Preprint arXiv:1812.08008 . 6 (1812)
2018 arXiv
-
[112]
& Ra- mamoorthi, R
Bi, S., Sunkavalli, K., Perazzi, F., Shechtman, E., Kim, V. & Ra- mamoorthi, R. Deep cg2real: Synthetic-to-real translation via image disentanglement. Proceedings Of The IEEE/CVF International Con- ference On Computer Vision . pp. 2730-2739 (2019)
2019
-
[113]
& Grauman, K
Yu, A. & Grauman, K. Fine-grained visual comparisons with local learning. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 192-199 (2014)
2014
-
[114]
& Borth, D
Kalkowski, S., Schulze, C., Dengel, A. & Borth, D. Real-time analysis and visualization of the yfcc100m dataset. Proceedings Of The 2015 Workshop On Community-organized Multimodal Mining: Opportuni- ties For Novel Solutions . pp. 25-30 (2015)
2015
-
[115]
& Salzmann, M
Ozaydin, B., Zhang, T., S¨ usstrunk, S. & Salzmann, M. DSI2I: Dense Style for Unpaired Image-to-Image Translation. ArXiv Preprint arXiv:2212.13253. (2022)
2022 arXiv
-
[116]
& Chen, C
Zhao, Y., Li, C., Yu, P., Gao, J. & Chen, C. Feature quantization improves gan training. ArXiv Preprint arXiv:2004.02088 . (2020)
2020 arXiv
-
[117]
& Aslam, M
Lee, H., Li, Y., Lee, T. & Aslam, M. Progressively unsupervised gener- ative attentional networks with adaptive layer-instance normalization for image-to-image translation. Sensors. 23, 6858 (2023)
2023
-
[118]
& Luo, J
Lin, J., Chen, Z., Xia, Y., Liu, S., Qin, T. & Luo, J. Exploring ex- plicit domain supervision for latent space disentanglement in unpaired image-to-image translation. IEEE Transactions On Pattern Analysis And Machine Intelligence . 43, 1254-1266 (2019) 70
2019
-
[119]
& Shi, J
Zheng, Z., Wu, Y., Han, X. & Shi, J. Forkgan: Seeing into the rainy night. Computer Vision–ECCV 2020: 16th European Conference, Glas- gow, UK, August 23–28, 2020, Proceedings, Part III 16 . pp. 155-170 (2020)
2020
-
[120]
& Sun, Z
Cao, J., Huang, H., Li, Y., He, R. & Sun, Z. Informative sample mining network for multi-domain image-to-image translation. European Con- ference On Computer Vision . pp. 404-419 (2020)
2020
-
[121]
& Cohen-Or, D
Katzir, O., Lischinski, D. & Cohen-Or, D. Cross-domain cascaded deep feature translation. ArXiv Preprint arXiv:1906.01526 . (2019)
2019 arXiv
-
[122]
& Robertson, N
Shubhra Ghosh, S., Hua, Y., Subhra Mukherjee, S. & Robertson, N. IEGAN: Multi-purpose Perceptual Quality Image Enhancement Us- ing Generative Adversarial Network. ArXiv E-prints . pp. arXiv-1811 (2018)
2018
-
[123]
& Chen, X
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A. & Chen, X. Improved techniques for training gans. Advances In Neural Information Processing Systems. 29 (2016)
2016
-
[124]
& Rabinovich, A
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V. & Rabinovich, A. Going deeper with con- volutions. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 1-9 (2015)
2015
-
[125]
& Michaeli, T
Shaham, T., Dekel, T. & Michaeli, T. Singan: Learning a generative model from a single natural image. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 4570-4580 (2019)
2019
-
[126]
& Gretton, A
Bi´ nkowski, M., Sutherland, D., Arbel, M. & Gretton, A. Demystifying mmd gans. ArXiv Preprint arXiv:1801.01401 . (2018)
2018 arXiv
-
[127]
& Khan, L
Lin, Y., Wang, Y., Li, Y., Gao, Y., Wang, Z. & Khan, L. Attention- based spatial guidance for image-to-image translation. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Vi- sion. pp. 816-825 (2021)
2021
-
[128]
& Wang, O
Zhang, R., Isola, P., Efros, A., Shechtman, E. & Wang, O. The un- reasonable effectiveness of deep features as a perceptual metric. Pro- ceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 586-595 (2018) 71
2018
-
[129]
& Weiss, Y
Richardson, E. & Weiss, Y. On gans and gmms. Advances In Neural Information Processing Systems. 31 (2018)
2018
-
[130]
& Zhu, J
Su, S., Song, J., Gao, L. & Zhu, J. Towards Unsupervised Deformable- Instances Image-to-Image Translation.. IJCAI. pp. 1004-1010 (2021)
2021
-
[131]
& Simoncelli, E
Wang, Z., Bovik, A., Sheikh, H. & Simoncelli, E. Image quality assess- ment: from error visibility to structural similarity. IEEE Transactions On Image Processing. 13, 600-612 (2004)
2004
-
[132]
& Burnaev, E
Korotin, A., Selikhanovych, D. & Burnaev, E. Neural Optimal Trans- port. (2023)
2023
-
[133]
& Wen, F
Zhang, P., Zhang, B., Chen, D., Yuan, L. & Wen, F. Cross-domain correspondence learning for exemplar-based image translation.Proceed- ings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 5143-5153 (2020)
2020
-
[134]
& Chen, C
Zhao, Y. & Chen, C. Unpaired image-to-image translation via latent energy transport. Proceedings Of The IEEE/CVF Conference On Com- puter Vision And Pattern Recognition . pp. 16418-16427 (2021) 72
2021
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.