Pith. sign in

REVIEW 4 major objections 5 minor 132 references

Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims that unpaired image-to-image translation tasks should be chosen by how much source content must survive, and that this choice is best made before picking a model.

desk verdict A useful review taxonomy and a modest Sim2Real benchmark, but the cross-paper FID/KID rankings in Section 6 are not commensurable and need to be reframed. read the letter →

arxiv 2502.08667 v2 pith:AD4C5HGT submitted 2025-02-11 eess.IV cs.CV

classification eess.IVcs.CV
keywords Image-to-imagetranslationUnpairedContentpreservationSim-to-realbenchmarkGenerativeadversarialnetworksDiffusionmodelsFIDKIDevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that unpaired image-to-image translation tasks differ in how much of the source image must survive, and that this difference should drive model choice. It divides tasks into Fully, Partially, and Non-Content preserving, sorts roughly seventy models by architecture, and collects FID and KID numbers across thirty tasks. It also builds a new Sim2Real benchmark combining simulated vehicle images with real images from three datasets and evaluates eight models on it. The practical payoff would be a checklist: decide how much content must be preserved, then pick a model family known to respect that constraint.

What carries the argument

The carrying object is the three-way content-preservation categorization: Fully Content preserving, Partially Content preserving, and Non-Content preserving, defined by how much of the source image must be reproduced in the output. Around this axis the paper organizes an architectural taxonomy of GAN-based, VAE-based, diffusion, flow-based, and transformer models, and a new Sim2Real benchmark built from simulated vehicle images and real images from MIO, Visdrone, and VeRi, with FID, KID, IS, NDB, JSD, and LPIPS as evaluation tools. The taxonomy supplies the structure of the review, and the benchmark supplies a reusable testbed for the content-preserving label.

What would settle it

Re-run a handful of the compared models end-to-end on the same train and validation split of the Sim2Real benchmark and of one Fully Content preserving task such as GTA to Cityscape, computing FID and KID with a single script. If the best model by the copied numbers is not the best model in the unified run, the paper's comparative conclusions collapse.

Watch

Extended reading notes

Core claim

On its own terms, the central claim is that content preservation is not a side effect but a design axis. The paper shows that tasks and datasets can be sorted into Fully Content preserving, Partially Content preserving, and Non-Content preserving, and that different models perform best in different categories. For example, DRIT++ achieves the lowest FID for GTA to Cityscape, CUT achieves the best KID on the same task, UNSB achieves the lowest FID for Horse to Zebra, and DCLGAN performs best on Label to Cityscapes. The paper also introduces a Sim2Real benchmark and reports evaluations of CycleGAN, StyleFlow, GcGAN, VSAIT, SRUNIT, and UNSB on it, showing that hyperparameter choices such as batch size and loss weights visibly change content preservation and image quality.

Load-bearing premise

The evaluation tables compare FID and KID scores copied from each model's original paper, so the comparison assumes those numbers were produced under fairly comparable training conditions; if they were not, the rankings and the conclusions drawn from them are not reliable.

Editorial extensions

If this is right

  • If the categorization is right, practitioners can narrow a model search by first deciding whether a task is fully, partially, or non-content preserving.
  • The Sim2Real benchmark gives later researchers a shared vehicle-focused dataset on which content-preserving methods can be compared directly.
  • The paper's tables show that no single model dominates all three categories, so reporting results per content-preservation category could become a useful convention.
  • The hyperparameter studies on CycleGAN, GcGAN, and StyleFlow suggest that content preservation in a given model is tunable through batch size and loss weights, not fixed by architecture alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would measure content preservation directly, for instance by comparing segmentation or keypoint consistency between source and translated images, rather than relying only on FID and KID.
  • The taxonomy could plausibly extend to newer diffusion and transformer models that were not fully benchmarked here, since those architectures also face the same content-versus-style tradeoff.
  • The benchmark could be reused for partially content-preserving tasks by adding tasks that deliberately change object class or layout, which would stress-test where the FCP/PCP boundary actually lies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a review of unpaired image-to-image translation organized around the degree to which the source image content must be preserved. It introduces a three-way task taxonomy (Fully Content Preserving, Partially Content Preserving, Non-Content Preserving), surveys around 70 models grouped by architecture, summarizes datasets and evaluation metrics, compiles FID and KID results from the original papers, and presents a new Sim2Real benchmark for content-preserving simulation-to-real translation. The paper concludes that the degree of content preservation should be considered when selecting an I2I model for a given application.

Significance. The proposed FCP/PCP/NCP taxonomy addresses a real and practically useful distinction that is often implicit in I2I papers, and the broad survey of models, datasets, and metrics provides a useful entry point for practitioners. The Sim2Real benchmark, combining VisDrone, MIO, and VeRi data with a simulated vehicle-image domain, is a potentially reusable evaluation resource, and the paper is explicit about the hyperparameters and metrics used in its own experiments. However, the central empirical claim that the taxonomy supports model selection rests on Section 6, where FID/KID values are copied from heterogeneous original papers without a shared evaluation protocol, and on the Sim2Real section, which has internal inconsistencies (the stated eight models are not the reported six) and no released data, code, or error bars. Because these issues affect the main practical conclusion, the paper needs revision before the claim can be accepted. I see no circular reasoning or fabricated entities; the weaknesses are methodological and presentation-related rather than conceptual.

major comments (4)
  1. [§6, Tables 4–5] The cross-model rankings in Section 6 are not commensurable because the FID and KID values are taken from the models' own articles under different training protocols, data splits, resolutions, and metric implementations, as the text itself states: 'These results are taken from the models' own articles.' Consequently, statements such as 'DRIT++ achieve the lowest FID' (Table 4, GTA-to-Cityscape) and 'CUT performs better than other methods' (Table 5, GTA-to-Cityscape) are not supported as general claims about model quality. The authors should either re-run all compared methods on a common split with a common metric implementation, or explicitly rephrase these statements as reports of the original papers' numbers rather than as findings of this review.
  2. [§7.2, Table 12] The evaluation section states that eight models were selected (CycleGAN, DRIT, GcGAN, StyleFlow, SRUNIT, VSAIT, UNSB), but the text lists only seven model names and Table 12 reports results for only six models, with no DRIT row and no explanation for its absence. This discrepancy makes the benchmark evaluation incomplete and raises the question of whether the missing model's results were omitted selectively. The authors should specify the exact set of models, justify any exclusions, and either add the missing results or remove the claim of eight models.
  3. [§7.1, §7.2] The Sim2Real benchmark lacks a stated train/test split, repeated runs with seeds, error bars, and a released data/code artifact. For example, in §7.2.1 the authors say CycleGAN was 'trained and tested on Sim2Real dataset' for 20 epochs, but no held-out evaluation set is described, so the reported FID, IS, NDB, JSD, and LPIPS numbers cannot be interpreted as generalization measures. The conclusion that 'changing hyperparameter λ from 10 to 5 do not increase the performance' is therefore not verifiable. The authors should describe the exact split, report multiple seeds with variance, and release the dataset construction code and evaluation scripts.
  4. [§4, Table 3] The FCP/PCP/NCP assignment is not operationalized: the paper gives qualitative descriptions of the three categories but no quantitative or procedural rule by which a reader can assign a new task to a category. Some assignments in Table 3 also appear inconsistent with the definitions, for instance Label2Cityscape is listed as PCP while Keypoint2Photo is listed as NCP, although both involve structured source information being transformed into a target image. Without an explicit annotation protocol or inter-annotator agreement, the taxonomy cannot be applied reproducibly by future users, which weakens the paper's central contribution.
minor comments (5)
  1. [§7.2] The sentence listing selected models ends with 'an UNSB [38]'; this should read 'and UNSB [38]'.
  2. [Table 11] The column header 'V an' contains a stray space and should be 'Van'.
  3. [Throughout] The text contains several typos and grammatical errors, for example 'benckmark' in §7.2.1, 'metircs' in the Table 12 caption, and 'do not increase the performance results' in §7.2.1; a careful proofread is needed.
  4. [Figure 10 caption and text] The caption says the first row is simulation images and the second row is real images, but the text in §7.1 says 'images of 5 classes in the VeRi data set have been randomly sampled and demonstrated'; the figure and text should be reconciled to avoid confusion about which domain is shown.
  5. [Table 9] Several KID values in Table 9 are on the order of 50–100, which is far outside the typical range for KID on standard benchmarks (usually below 0.1); the authors should state whether these values are scaled (e.g., multiplied by 100) or whether a different implementation was used.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the taxonomy and benchmark are organizational and empirical contributions, with no fitted parameter renamed as a prediction and no load-bearing self-citation.

full rationale

The paper is a review plus an empirical benchmark. The FCP/PCP/NCP categories are descriptive labels assigned to tasks according to the stated degree of content preservation; they are not derived from the benchmark results, and no equation defines one category in terms of another. Section 6 rankings are drawn from the original model papers, as the text acknowledges ('These results are taken from the models' own articles'), which creates a comparability and validity limitation but not circularity: the numbers are external evidence, not outputs of a model fitted to the taxonomy. Section 7's Sim2Real evaluation reports measured FID, IS, NDB, JSD, and LPIPS values for eight named models; the choice of models and hyperparameters is a design decision rather than a fitted parameter later renamed as a prediction. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation. The concluding recommendation that content-preservation extent should inform model choice restates the organizing principle of the taxonomy, but it is a practical guideline rather than a derived result claimed to be proven by the benchmark. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a review paper with an empirical benchmark, not a mathematical derivation. There are no fitted free parameters in a modeling sense: the hyperparameter choices in Section 7.2 are experimental configurations, not parameters fitted to make a formula work. The paper relies on three domain assumptions: the meaningfulness of the FCP/PCP/NCP split, the comparability of cross-paper metrics, and the validity of the assembled Sim2Real dataset. No new physical or conceptual entities are introduced.

assumptions (3)
  • domain assumption I2I tasks can be cleanly partitioned into Fully, Partially, and Non-Content preserving categories based on source and target domains.
    Section 4 introduces the three categories and Table 3 assigns specific tasks to them, but no measurable or operational criterion is given; the boundaries rely on the authors' judgment.
  • domain assumption FID and KID scores reported in different original papers are comparable enough to rank methods in the survey tables.
    Section 6 assembles results taken directly from each model's article without re-running models under a common protocol. This assumes cross-paper comparability, which is fragile given differing training setups and data splits.
  • domain assumption The combination of Visdrone, MIO, VeRi, and the simulated vehicle dataset constitutes a valid Sim2Real benchmark for content-preserving translation.
    Section 7.1 describes the data combination and a pixel threshold for splitting low/high resolution, but does not validate domain gap, label consistency, or usefulness of the benchmark beyond the authors' own runs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review." pith.science (2026). https://pith.science/paper/AD4C5HGT

@misc{pith2026250208667,
  author       = {Pith},
  title        = {Pith review of: Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AD4C5HGT}},
  note         = {Machine review of arXiv:2502.08667}
}
read the original abstract

Image-to-image translation (I2I) transforms an image from a source domain to a target domain while preserving source content. Most computer vision applications are in the field of image-to-image translation, such as style transfer, image segmentation, and photo enhancement. The degree of preservation of the content of the source images in the translation process can be different according to the problem and the intended application. From this point of view, in this paper, we divide the different tasks in the field of image-to-image translation into three categories: Fully Content preserving, Partially Content preserving, and Non-Content preserving. We present different tasks, datasets, methods, results of methods for these three categories in this paper. We make a categorization for I2I methods based on the architecture of different models and study each category separately. In addition, we introduce well-known evaluation criteria in the I2I translation field. Specifically, nearly 70 different I2I models were analyzed, and more than 10 quantitative evaluation metrics and 30 distinct tasks and datasets relevant to the I2I translation problem were both introduced and assessed. Translating from simulation to real images could be well viewed as an application of fully content preserving or partially content preserving unsupervised image-to-image translation methods. So, we provide a benchmark for Sim-to-Real translation, which can be used to evaluate different methods. In general, we conclude that because of the different extent of the obligation to preserving content in various applications, it is better to consider this issue in choosing a suitable I2I model for a specific application.

Figures

Figures reproduced from arXiv: 2502.08667 by the authors.

Figure 1
Figure 1. Several image-to-image translation problems. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Visual comparisons of results of different models (CycleGAN [1], MUNIT [14], [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. An overview of image-to-image translation methods. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: GAN-based image-to-image translation methods [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Conditional Generative Adversarial Network [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: (a) and (b) are paired datasets. (c) and (d) are unpaired datasets. [PITH_FULL_IMAGE:figures/full_fig_p035_6.png]
Figure 7
Figure 7. Figure 7: (a) and (b) are Fully Content preserving . (c) and (d) are Partially Content [PITH_FULL_IMAGE:figures/full_fig_p037_7.png]
Figure 8
Figure 8. Figure 8: Visual comparison of I2I methods for GTA [PITH_FULL_IMAGE:figures/full_fig_p043_8.png]
Figure 9
Figure 9. Figure 9: Visual results for some PCP and NCP tasks. [PITH_FULL_IMAGE:figures/full_fig_p044_9.png]
Figure 10
Figure 10. Figure 10: Simulation sample images (first row) and Real sample images (second row). [PITH_FULL_IMAGE:figures/full_fig_p050_10.png]
Figure 11
Figure 11. Figure 11: Visual results of experimenting CycleGAN on Sim2Real benckmark dataset. [PITH_FULL_IMAGE:figures/full_fig_p051_11.png]
Figure 12
Figure 12. Figure 12: StyleFlow visual results for hyperparameter combination B1, B4, B8 and B16 [PITH_FULL_IMAGE:figures/full_fig_p053_12.png]
Figure 13
Figure 13. Figure 13: StyleFlow visual results for a fixed configuration and different input style [PITH_FULL_IMAGE:figures/full_fig_p053_13.png]
Figure 14
Figure 14. Figure 14: StyleFlow visual results for hyperparameter combination P1 and P3 explained [PITH_FULL_IMAGE:figures/full_fig_p054_14.png]
Figure 15
Figure 15. Figure 15: Effect of different values for GcGAN hyperparameters [PITH_FULL_IMAGE:figures/full_fig_p055_15.png]
Figure 16
Figure 16. Figure 16: Visual results of experimenting VSAIT model on Sim2Real benckmark dataset [PITH_FULL_IMAGE:figures/full_fig_p056_16.png]
Figure 17
Figure 17. Figure 17: Visual results of experimenting with the SRUNIT model on the Sim2Real [PITH_FULL_IMAGE:figures/full_fig_p057_17.png]
Figure 18
Figure 18. Figure 18: UNSB model results. 57 [PITH_FULL_IMAGE:figures/full_fig_p057_18.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

132 extracted references · 69 canonical work pages

  1. [1]

    & Efros, A

    Zhu, J., Park, T., Isola, P. & Efros, A. Unpaired image-to-image trans- lation using cycle-consistent adversarial networks. Proceedings Of The IEEE International Conference On Computer Vision . pp. 2223-2232 (2017)

  2. [2]

    & Efros, A

    Isola, P., Zhu, J., Zhou, T. & Efros, A. Image-to-image translation with conditional adversarial networks. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 1125-1134 (2017)

  3. [3]

    & Cucchiara, R

    Tomei, M., Cornia, M., Baraldi, L. & Cucchiara, R. Art2real: Un- folding the reality of artworks via semantically-aware image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 5849-5859 (2019)

  4. [4]

    & Chuang, Y

    Chang, H., Wang, Z. & Chuang, Y. Domain-specific mappings for gen- erative adversarial style transfer. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16 . pp. 573-589 (2020)

  5. [5]

    & Gupta, V

    Kumar, P. & Gupta, V. Unpaired Image-to-Image Translation Based Artwork Restoration Using Generative Adversarial Networks. Interna- tional Conference On Intelligent Manufacturing And Energy Sustain- ability. pp. 581-591 (2023) 58

  6. [6]

    & Kim, A

    Cho, Y., Malav, R., Pandey, G. & Kim, A. DehazeGAN: underwa- ter haze image restoration using unpaired image-to-image translation. IF AC-PapersOnLine. 52, 82-85 (2019)

  7. [7]

    & Huang, J

    Guo, X., Wang, Z., Yang, Q., Lv, W., Liu, X., Wu, Q. & Huang, J. Gan-based virtual-to-real image translation for urban scene semantic segmentation. Neurocomputing. 394 pp. 127-135 (2020)

  8. [8]

    & Wong, H

    Li, R., Cao, W., Jiao, Q., Wu, S. & Wong, H. Simplified unsuper- vised image translation for semantic segmentation adaptation. Pattern Recognition. 105 pp. 107343 (2020)

Show all 132 references
  1. [9]

    & Kim, K

    Murez, Z., Kolouri, S., Kriegman, D., Ramamoorthi, R. & Kim, K. Image to image translation for domain adaptation. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 4500-4509 (2018)

  2. [10]

    & Shen, H

    Li, J., Lu, K., Huang, Z., Zhu, L. & Shen, H. Heterogeneous domain adaptation through progressive alignment.IEEE Transactions On Neu- ral Networks And Learning Systems . 30, 1381-1391 (2018)

  3. [11]

    & Cook, D

    Wilson, G. & Cook, D. A survey of unsupervised deep domain adap- tation. ACM Transactions On Intelligent Systems And Technology (TIST). 11, 1-46 (2020)

  4. [12]

    & Yang, M

    Lee, H., Tseng, H., Huang, J., Singh, M. & Yang, M. Diverse image-to- image translation via disentangled representations. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 35-51 (2018)

  5. [13]

    & Gong, M

    Yi, Z., Zhang, H., Tan, P. & Gong, M. Dualgan: Unsupervised dual learning for image-to-image translation. Proceedings Of The IEEE In- ternational Conference On Computer Vision . pp. 2849-2857 (2017)

  6. [14]

    & Kautz, J

    Huang, X., Liu, M., Belongie, S. & Kautz, J. Multimodal unsupervised image-to-image translation. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 172-189 (2018)

  7. [15]

    & Chen, Z

    Pang, Y., Lin, J., Qin, T. & Chen, Z. Image-to-image translation: Methods and applications. IEEE Transactions On Multimedia . 24 pp. 3859-3881 (2021) 59

  8. [16]

    & Choo, J

    Cho, W., Choi, S., Park, D., Shin, I. & Choo, J. Image-to-image trans- lation via group-wise deep whitening-and-coloring transformation. Pro- ceedings Of The IEEE/CVF Conference On Computer Vision And Pat- tern Recognition. pp. 10639-10647 (2019)

  9. [17]

    & Prakash, A

    Theiss, J., Leverett, J., Kim, D. & Prakash, A. Unpaired image transla- tion via vector symbolic architectures. European Conference On Com- puter Vision . pp. 17-32 (2022)

  10. [18]

    Fan, W., Chen, J., Ma, J., Hou, J. & Yi, S. Styleflow for content-fixed image to image translation. ArXiv Preprint arXiv:2207.01909 . (2022)

  11. [19]

    & Tao, D

    Fu, H., Gong, M., Wang, C., Batmanghelich, K., Zhang, K. & Tao, D. Geometry-consistent generative adversarial networks for one-sided unsupervised domain mapping. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2427-2436 (2019)

  12. [20]

    & Fang, B

    Chen, R., Huang, W., Huang, B., Sun, F. & Fang, B. Reusing discrimi- nators for encoding: Towards unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 8168-8177 (2020)

  13. [21]

    & Zhang, K

    Xie, S., Gong, M., Xu, Y. & Zhang, K. Unaligned image-to-image translation by learning to reweight. Proceedings Of The IEEE/CVF International Conference On Computer Vision. pp. 14174-14184 (2021)

  14. [22]

    & Abbeel, P

    Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W. & Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. 2017 IEEE/RSJ International Conference On Intelligent Robots And Systems (IROS) . pp. 23-30 (2017)

  15. [23]

    & Bengio, Y

    Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. & Bengio, Y. Generative adversarial nets. Advances In Neural Information Processing Systems . 27 (2014)

  16. [24]

    & Metaxas, D

    Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X. & Metaxas, D. Stackgan++: Realistic image synthesis with stacked generative ad- versarial networks. IEEE Transactions On Pattern Analysis And Ma- chine Intelligence . 41, 1947-1962 (2018) 60

  17. [25]

    & Papa, J

    De Rosa, G. & Papa, J. A survey on text generation using generative adversarial networks. Pattern Recognition. 119 pp. 108098 (2021)

  18. [26]

    & Liang, Z

    Liu, Z., Wang, J. & Liang, Z. Catgan: Category-aware generative ad- versarial networks with hierarchical evolutionary learning for category text generation. Proceedings Of The AAAI Conference On Artificial Intelligence. 34, 8425-8432 (2020)

  19. [27]

    & Mohammadi, G

    Aldausari, N., Sowmya, A., Marcus, N. & Mohammadi, G. Video generative adversarial networks: a review. ACM Computing Surveys (CSUR). 55, 1-25 (2022)

  20. [28]

    & Tan, M

    Chen, Q., Wu, Q., Chen, J., Wu, Q., Hengel, A. & Tan, M. Scripted video generation with a bottom-up generative adversarial network. IEEE Transactions On Image Processing . 29 pp. 7454-7467 (2020)

  21. [29]

    Conditional generative adversarial nets

    Mirza, M. Conditional generative adversarial nets. ArXiv Preprint arXiv:1411.1784. (2014)

  22. [30]

    Conditional generative adversarial nets for convolutional face generation

    Gauthier, J. Conditional generative adversarial nets for convolutional face generation. Class Project For Stanford CS231N: Convolutional Neural Networks For Visual Recognition, Winter Semester . 2014, 2 (2014)

  23. [31]

    Auto-encoding variational bayes

    Kingma, D. Auto-encoding variational bayes. ArXiv Preprint arXiv:1312.6114. (2013)

  24. [32]

    & Ferrari, V

    Pumarola, A., Popov, S., Moreno-Noguer, F. & Ferrari, V. C-flow: Conditional generative flow models for images and 3d point clouds. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 7949-7958 (2020)

  25. [33]

    & Mohamed, S

    Rezende, D. & Mohamed, S. Variational inference with normalizing flows. International Conference On Machine Learning . pp. 1530-1538 (2015)

  26. [34]

    & Abbeel, P

    Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. Advances In Neural Information Processing Systems. 33 pp. 6840-6851 (2020) 61

  27. [35]

    & Breckon, T

    Sasaki, H., Willcocks, C. & Breckon, T. Unit-ddpm: Unpaired im- age translation with denoising diffusion probabilistic models. ArXiv Preprint arXiv:2104.05358. (2021)

  28. [36]

    & Zhu, J

    Zhao, M., Bao, F., Li, C. & Zhu, J. Egsde: Unpaired image-to- image translation via energy-guided stochastic differential equations. Advances In Neural Information Processing Systems. 35 pp. 3609-3623 (2022)

  29. [37]

    & Ermon, S

    Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J. & Ermon, S. SDEdit: Guided Image Synthesis and Editing with Stochastic Differ- ential Equations. (2022)

  30. [38]

    Kim, B., Kwon, G., Kim, K. & Ye, J. Unpaired Image-to- Image Translation via Neural Schrodinger Bridge. ArXiv Preprint arXiv:2305.15086. (2023)

  31. [39]

    Attention is all you need

    Vaswani, A. Attention is all you need. Advances In Neural Information Processing Systems. (2017)

  32. [41]

    & Koltun, V

    Ranftl, R., Bochkovskiy, A. & Koltun, V. Vision transformers for dense prediction. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 12179-12188 (2021)

  33. [42]

    & Others A survey on vision transformer

    Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y. & Others A survey on vision transformer. IEEE Transactions On Pattern Analysis And Machine Intelligence. 45, 87-110 (2022)

  34. [43]

    & Shah, M

    Khan, S., Naseer, M., Hayat, M., Zamir, S., Khan, F. & Shah, M. Transformers in vision: A survey. ACM Computing Surveys (CSUR) . 54, 1-41 (2022)

  35. [44]

    & Others An image is worth 16x16 words: Transformers for image recognition at scale

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S. & Others An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv Preprint arXiv:2010.11929 . (2020) 62

  36. [45]

    Flow-based Deep Generative Models

    Weng, L. Flow-based Deep Generative Models. (2018), https://lilianweng.github.io/posts/2018-10-13-flow-models/, Accessed: 2024-08-27

  37. [46]

    & Kim, J

    Kim, T., Cha, M., Kim, H., Lee, J. & Kim, J. Learning to discover cross-domain relations with generative adversarial networks. Interna- tional Conference On Machine Learning . pp. 1857-1865 (2017)

  38. [47]

    & Wolf, L

    Benaim, S. & Wolf, L. One-sided unsupervised domain mapping. Ad- vances In Neural Information Processing Systems . 30 (2017)

  39. [48]

    & Bottou, L

    Arjovsky, M., Chintala, S. & Bottou, L. Wasserstein generative adver- sarial networks. International Conference On Machine Learning . pp. 214-223 (2017)

  40. [49]

    & Van Gool, L

    Ma, L., Jia, X., Georgoulis, S., Tuytelaars, T. & Van Gool, L. Ex- emplar guided unsupervised image-to-image translation with semantic consistency. ArXiv Preprint arXiv:1805.11145 . (2018)

  41. [50]

    & Zhu, J

    Park, T., Liu, M., Wang, T. & Zhu, J. Semantic image synthesis with spatially-adaptive normalization. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2337-2346 (2019)

  42. [51]

    Jia, Z., Yuan, B., Wang, K., Wu, H., Clifford, D., Yuan, Z. & Su, H. Semantically robust unpaired image translation for data with un- matched semantics statistics. Proceedings Of The IEEE/CVF Interna- tional Conference On Computer Vision . pp. 14273-14283 (2021)

  43. [52]

    & Cai, J

    Zheng, C., Cham, T. & Cai, J. The spatially-correlative loss for various image translation tasks. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 16407-16417 (2021)

  44. [53]

    & Chinnam, R

    Emami, H., Aliabadi, M., Dong, M. & Chinnam, R. SPA-GAN: Spatial attention GAN for image-to-image translation. IEEE Transactions On Multimedia. 23 pp. 391-401 (2020)

  45. [54]

    & Dong, H

    Zhao, Y., Wu, R. & Dong, H. Unpaired image-to-image translation using adversarial consistency loss. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16 . pp. 800-815 (2020) 63

  46. [55]

    Liu, Y., Wang, H., Yue, Y. & Lu, F. Separating content and style for unsupervised image-to-image translation. ArXiv Preprint arXiv:2110.14404. (2021)

  47. [56]

    & Belongie, S

    Huang, X. & Belongie, S. Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings Of The IEEE Interna- tional Conference On Computer Vision . pp. 1501-1510 (2017)

  48. [57]

    & Wen, F

    Zhou, X., Zhang, B., Zhang, T., Zhang, P., Bao, J., Chen, D., Zhang, Z. & Wen, F. Cocosnet v2: Full-resolution correspondence learning for image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 11465-11475 (2021)

  49. [58]

    & Zhu, J

    Park, T., Efros, A., Zhang, R. & Zhu, J. Contrastive learning for unpaired image-to-image translation. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Pro- ceedings, Part IX 16 . pp. 319-345 (2020)

  50. [59]

    & Armin, M

    Han, J., Shoeiby, M., Petersson, L. & Armin, M. Dual contrastive learning for unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recog- nition. pp. 746-755 (2021)

  51. [60]

    Wang, W., Zhou, W., Bao, J., Chen, D. & Li, H. Instance-wise hard negative example generation for contrastive learning in unpaired image- to-image translation. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 14020-14029 (2021)

  52. [61]

    & Charette, R

    Pizzati, F., Lalonde, J. & Charette, R. Manifest: Manifold deforma- tion for few-shot image translation. European Conference On Computer Vision. pp. 440-456 (2022)

  53. [62]

    Lin, J., Wang, Y., Chen, Z. & He, T. Learning to transfer: unsuper- vised domain translation via meta-learning. Proceedings Of The AAAI Conference On Artificial Intelligence . 34, 11507-11514 (2020)

  54. [63]

    Koksal, A. & Lu, S. Rf-gan: A light and reconfigurable network for unpaired image-to-image translation. Proceedings Of The Asian Con- ference On Computer Vision . (2020) 64

  55. [64]

    Ye, K., Ye, Y., Yang, M. & Hu, B. Independent encoder for deep hierarchical unsupervised image-to-image translation. ArXiv Preprint arXiv:2107.02494. (2021)

  56. [65]

    & Chen, Q

    Zhao, J., Lee, F., Hu, C., Yu, H. & Chen, Q. LDA-GAN: Lightweight domain-attention GAN for unpaired image-to-image translation. Neu- rocomputing. 506 pp. 355-368 (2022)

  57. [66]

    & Wang, Z

    Deng, H., Wu, Q., Huang, H., Yang, X. & Wang, Z. Involution- GAN: lightweight GAN with involution for unsupervised image-to- image translation. Neural Computing And Applications . 35, 16593- 16605 (2023)

  58. [67]

    & Huang, H

    Ganjdanesh, A., Gao, S., Alipanah, H. & Huang, H. Compressing image-to-image translation gans using local density structures on their learned manifold. Proceedings Of The AAAI Conference On Artificial Intelligence. 38, 12118-12126 (2024)

  59. [68]

    & Krishnaswamy, S

    Amodio, M. & Krishnaswamy, S. Travelgan: Image-to-image transla- tion by transformation vector learning. Proceedings Of The Ieee/cvf Conference On Computer Vision And Pattern Recognition . pp. 8983- 8992 (2019)

  60. [69]

    Siamese neural networks: An overview

    Chicco, D. Siamese neural networks: An overview. Artificial Neural Networks. pp. 73-94 (2021)

  61. [71]

    & Chang, X

    Cao, Y., Yao, L., Pan, L., Sheng, Q. & Chang, X. Guided Image-to- Image Translation by Discriminator-Generator Communication. IEEE Transactions On Multimedia. (2023)

  62. [72]

    & Choo, J

    Choi, Y., Choi, M., Kim, M., Ha, J., Kim, S. & Choo, J. Stargan: Uni- fied generative adversarial networks for multi-domain image-to-image translation. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 8789-8797 (2018)

  63. [73]

    Choi, Y., Uh, Y., Yoo, J. & Ha, J. Stargan v2: Diverse image synthesis for multiple domains. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 8188-8197 (2020) 65

  64. [74]

    Yu, X., Chen, Y., Liu, S., Li, T. & Li, G. Multi-mapping image-to- image translation via learning disentanglement. Advances In Neural Information Processing Systems. 32 (2019)

  65. [75]

    Liu, R., Ge, Y., Choi, C., Wang, X. & Li, H. Divco: Diverse condi- tional image synthesis via contrastive generative adversarial network. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 16377-16386 (2021)

  66. [76]

    & Yang, M

    Mao, Q., Lee, H., Tseng, H., Ma, S. & Yang, M. Mode seeking gener- ative adversarial networks for diverse image synthesis. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recog- nition. pp. 1429-1437 (2019)

  67. [77]

    & Yan, Y

    Tang, H., Xu, D., Sebe, N., Wang, Y., Corso, J. & Yan, Y. Multi- channel attention selection gan with cascaded semantic guidance for cross-view image translation. Proceedings Of The IEEE/CVF Con- ference On Computer Vision And Pattern Recognition . pp. 2417-2426 (2019)

  68. [78]

    & Yan, Y

    Tang, H., Xu, D., Sebe, N. & Yan, Y. Attention-Guided Generative Adversarial Networks for Unsupervised Image-to-Image Translation. CoRR. abs/1903.12296 (2019), http://arxiv.org/abs/1903.12296

  69. [79]

    & Sebe, N

    Tang, H., Liu, H., Xu, D., Torr, P. & Sebe, N. Attentiongan: Unpaired image-to-image translation using attention-guided generative adversar- ial networks. IEEE Transactions On Neural Networks And Learning Systems. 34, 1972-1987 (2021)

  70. [80]

    & Kim, K

    Mejjati, Y., Richardt, C., Tompkin, J., Cosker, D. & Kim, K. Unsu- pervised Attention-guided Image to Image Translation. (2018)

  71. [81]

    U-gat-it: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation

    Kim, J. U-gat-it: unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation. ArXiv Preprint arXiv:1907.10830 . (2019)

  72. [82]

    & Miao, C

    Zhan, F., Yu, Y., Wu, R., Zhang, J., Cui, K., Xiao, A., Lu, S. & Miao, C. Bi-level feature alignment for versatile image translation and manipulation. European Conference On Computer Vision. pp. 224-241 (2022) 66

  73. [83]

    & De Guevara, M

    Cazenavette, G. & De Guevara, M. MixerGAN: An MLP-based ar- chitecture for unpaired image-to-image translation. ArXiv Preprint arXiv:2105.14110. (2021)

  74. [84]

    & Oth- ers Mlp-mixer: An all-mlp architecture for vision

    Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Un- terthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J. & Oth- ers Mlp-mixer: An all-mlp architecture for vision. Advances In Neural Information Processing Systems. 34 pp. 24261-24272 (2021)

  75. [85]

    & Hao, Y

    Lai, X., Bai, X. & Hao, Y. Unsupervised generative adversarial net- works with cross-model weight transfer mechanism for image-to-image translation. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 1814-1822 (2021)

  76. [86]

    & Kautz, J

    Liu, M., Breuel, T. & Kautz, J. Unsupervised image-to-image trans- lation networks. Advances In Neural Information Processing Systems . 30 (2017)

  77. [87]

    & Loy, C

    Wu, W., Cao, K., Li, C., Qian, C. & Loy, C. Transgaga: Geometry- aware unsupervised image-to-image translation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 8012-8021 (2019)

  78. [88]

    & Liu, M

    Saito, K., Saenko, K. & Liu, M. Coco-funit: Few-shot unsupervised image translation with a content conditioned style encoder. Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 . pp. 382-398 (2020)

  79. [89]

    & Han, S

    Han, G., Min, J. & Han, S. EM-LAST: Effective Multidimensional Latent Space Transport for an Unpaired Image-to-Image Translation With an Energy-Based Model. IEEE Access. 10 pp. 72839-72849 (2022)

  80. [90]

    & Dhariwal, P

    Kingma, D. & Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. Advances In Neural Information Processing Systems . 31 (2018)

  81. [91]

    Very deep convolutional networks for large-scale image recognition

    Simonyan, K. Very deep convolutional networks for large-scale image recognition. ArXiv Preprint arXiv:1409.1556 . (2014)

  82. [92]

    & Liu, Z

    Fan, W., Chen, J. & Liu, Z. Hierarchy Flow For High-Fidelity Image- to-Image Translation. ArXiv Preprint arXiv:2308.06909 . (2023) 67

  83. [93]

    & Singh, V

    Sun, H., Mehta, R., Zhou, H., Huang, Z., Johnson, S., Prabhakaran, V. & Singh, V. Dual-glow: Conditional flow-based generative model for modality transfer. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 10611-10620 (2019)

  84. [94]

    & Wang, Z

    Zheng, W., Li, Q., Zhang, G., Wan, P. & Wang, Z. Ittr: Un- paired image-to-image translation with transformers. ArXiv Preprint arXiv:2203.16015. (2022)

  85. [95]

    & Kim, S

    Kim, S., Baek, J., Park, J., Kim, G. & Kim, S. InstaFormer: Instance- aware image-to-image translation with transformer.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 18321-18331 (2022)

  86. [96]

    & Ren, Y

    Torbunov, D., Huang, Y., Yu, H., Huang, J., Yoo, S., Lin, M., Viren, B. & Ren, Y. Uvcgan: Unet vision transformer cycle-consistent gan for unpaired image-to-image translation. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Vision . pp. 702-712 (2023)

  87. [97]

    & Schiele, B

    Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benen- son, R., Franke, U., Roth, S. & Schiele, B. The cityscapes dataset for semantic urban scene understanding. Proceedings Of The IEEE Con- ference On Computer Vision And Pattern Recognition . pp. 3213-3223 (2016)

  88. [98]

    & Torralba, A

    Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A. & Torralba, A. Scene parsing through ade20k dataset. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 633- 641 (2017)

  89. [99]

    & Wyeth, G

    Milford, M. & Wyeth, G. SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights.2012 IEEE International Conference On Robotics And Automation . pp. 1643-1649 (2012)

  90. [100]

    & Darrell, T

    Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V. & Darrell, T. Bdd100k: a diverse driving video database with scalable annotation tooling. 2018. ArXiv Preprint arXiv:1805.04687 . (1805) 68

  91. [101]

    & Tang, X

    Liu, Z., Luo, P., Wang, X. & Tang, X. Deep learning face attributes in the wild. Proceedings Of The IEEE International Conference On Computer Vision . pp. 3730-3738 (2015)

  92. [102]

    & Tang, X

    Liu, Z., Luo, P., Qiu, S., Wang, X. & Tang, X. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. Pro- ceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 1096-1104 (2016)

  93. [103]

    & Belongie, S

    Poursaeed, O., Matera, T. & Belongie, S. Vision-based real estate price estimation. Machine Vision And Applications . 29, 667-676 (2018)

  94. [104]

    &ˇS´ ara, R

    Tyleˇ cek, R. &ˇS´ ara, R. Spatial pattern templates for recognition of ob- jects with regular structure. Pattern Recognition: 35th German Con- ference, GCPR 2013, Saarbr¨ ucken, Germany, September 3-6, 2013. Proceedings 35. pp. 364-374 (2013)

  95. [105]

    & Aila, T

    Karras, T., Laine, S. & Aila, T. A style-based generator architecture for generative adversarial networks. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 4401- 4410 (2019)

  96. [106]

    & Lazebnik, S

    Plummer, B., Wang, L., Cervantes, C., Caicedo, J., Hockenmaier, J. & Lazebnik, S. Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models. Proceedings Of The IEEE International Conference On Computer Vision . pp. 2641-2649 (2015)

  97. [107]

    & Koltun, V

    Richter, S., Vineet, V., Roth, S. & Koltun, V. Playing for data: Ground truth from computer games.Computer Vision–ECCV 2016: 14th Euro- pean Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14. pp. 102-118 (2016)

  98. [108]

    & Fei-Fei, L

    Deng, J., Dong, W., Socher, R., Li, L., Li, K. & Fei-Fei, L. Imagenet: A large-scale hierarchical image database. 2009 IEEE Conference On Computer Vision And Pattern Recognition . pp. 248-255 (2009)

  99. [109]

    & Urtasun, R

    Geiger, A., Lenz, P. & Urtasun, R. Are we ready for autonomous driv- ing? the kitti vision benchmark suite. 2012 IEEE Conference On Com- puter Vision And Pattern Recognition . pp. 3354-3361 (2012) 69

  100. [110]

    & Fergus, R

    Silberman, N., Hoiem, D., Kohli, P. & Fergus, R. Indoor segmenta- tion and support inference from rgbd images. Computer Vision–ECCV 2012: 12th European Conference On Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12 . pp. 746-760 (2012)

  101. [111]

    & Sheikh, Y

    Cao, Z., Hidalgo, G., Simon, T., Wei, S. & Sheikh, Y. Openpose: real- time multi-person 2d pose estimation using part affinity fields (2018). ArXiv Preprint arXiv:1812.08008 . 6 (1812)

  102. [112]

    & Ra- mamoorthi, R

    Bi, S., Sunkavalli, K., Perazzi, F., Shechtman, E., Kim, V. & Ra- mamoorthi, R. Deep cg2real: Synthetic-to-real translation via image disentanglement. Proceedings Of The IEEE/CVF International Con- ference On Computer Vision . pp. 2730-2739 (2019)

  103. [113]

    & Grauman, K

    Yu, A. & Grauman, K. Fine-grained visual comparisons with local learning. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 192-199 (2014)

  104. [114]

    & Borth, D

    Kalkowski, S., Schulze, C., Dengel, A. & Borth, D. Real-time analysis and visualization of the yfcc100m dataset. Proceedings Of The 2015 Workshop On Community-organized Multimodal Mining: Opportuni- ties For Novel Solutions . pp. 25-30 (2015)

  105. [115]

    & Salzmann, M

    Ozaydin, B., Zhang, T., S¨ usstrunk, S. & Salzmann, M. DSI2I: Dense Style for Unpaired Image-to-Image Translation. ArXiv Preprint arXiv:2212.13253. (2022)

  106. [116]

    & Chen, C

    Zhao, Y., Li, C., Yu, P., Gao, J. & Chen, C. Feature quantization improves gan training. ArXiv Preprint arXiv:2004.02088 . (2020)

  107. [117]

    & Aslam, M

    Lee, H., Li, Y., Lee, T. & Aslam, M. Progressively unsupervised gener- ative attentional networks with adaptive layer-instance normalization for image-to-image translation. Sensors. 23, 6858 (2023)

  108. [118]

    & Luo, J

    Lin, J., Chen, Z., Xia, Y., Liu, S., Qin, T. & Luo, J. Exploring ex- plicit domain supervision for latent space disentanglement in unpaired image-to-image translation. IEEE Transactions On Pattern Analysis And Machine Intelligence . 43, 1254-1266 (2019) 70

  109. [119]

    & Shi, J

    Zheng, Z., Wu, Y., Han, X. & Shi, J. Forkgan: Seeing into the rainy night. Computer Vision–ECCV 2020: 16th European Conference, Glas- gow, UK, August 23–28, 2020, Proceedings, Part III 16 . pp. 155-170 (2020)

  110. [120]

    & Sun, Z

    Cao, J., Huang, H., Li, Y., He, R. & Sun, Z. Informative sample mining network for multi-domain image-to-image translation. European Con- ference On Computer Vision . pp. 404-419 (2020)

  111. [121]

    & Cohen-Or, D

    Katzir, O., Lischinski, D. & Cohen-Or, D. Cross-domain cascaded deep feature translation. ArXiv Preprint arXiv:1906.01526 . (2019)

  112. [122]

    & Robertson, N

    Shubhra Ghosh, S., Hua, Y., Subhra Mukherjee, S. & Robertson, N. IEGAN: Multi-purpose Perceptual Quality Image Enhancement Us- ing Generative Adversarial Network. ArXiv E-prints . pp. arXiv-1811 (2018)

  113. [123]

    & Chen, X

    Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A. & Chen, X. Improved techniques for training gans. Advances In Neural Information Processing Systems. 29 (2016)

  114. [124]

    & Rabinovich, A

    Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V. & Rabinovich, A. Going deeper with con- volutions. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 1-9 (2015)

  115. [125]

    & Michaeli, T

    Shaham, T., Dekel, T. & Michaeli, T. Singan: Learning a generative model from a single natural image. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 4570-4580 (2019)

  116. [126]

    & Gretton, A

    Bi´ nkowski, M., Sutherland, D., Arbel, M. & Gretton, A. Demystifying mmd gans. ArXiv Preprint arXiv:1801.01401 . (2018)

  117. [127]

    & Khan, L

    Lin, Y., Wang, Y., Li, Y., Gao, Y., Wang, Z. & Khan, L. Attention- based spatial guidance for image-to-image translation. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Vi- sion. pp. 816-825 (2021)

  118. [128]

    & Wang, O

    Zhang, R., Isola, P., Efros, A., Shechtman, E. & Wang, O. The un- reasonable effectiveness of deep features as a perceptual metric. Pro- ceedings Of The IEEE Conference On Computer Vision And Pattern Recognition. pp. 586-595 (2018) 71

  119. [129]

    & Weiss, Y

    Richardson, E. & Weiss, Y. On gans and gmms. Advances In Neural Information Processing Systems. 31 (2018)

  120. [130]

    & Zhu, J

    Su, S., Song, J., Gao, L. & Zhu, J. Towards Unsupervised Deformable- Instances Image-to-Image Translation.. IJCAI. pp. 1004-1010 (2021)

  121. [131]

    & Simoncelli, E

    Wang, Z., Bovik, A., Sheikh, H. & Simoncelli, E. Image quality assess- ment: from error visibility to structural similarity. IEEE Transactions On Image Processing. 13, 600-612 (2004)

  122. [132]

    & Burnaev, E

    Korotin, A., Selikhanovych, D. & Burnaev, E. Neural Optimal Trans- port. (2023)

  123. [133]

    & Wen, F

    Zhang, P., Zhang, B., Chen, D., Yuan, L. & Wen, F. Cross-domain correspondence learning for exemplar-based image translation.Proceed- ings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 5143-5153 (2020)

  124. [134]

    & Chen, C

    Zhao, Y. & Chen, C. Unpaired image-to-image translation via latent energy transport. Proceedings Of The IEEE/CVF Conference On Com- puter Vision And Pattern Recognition . pp. 16418-16427 (2021) 72

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.