Pith. sign in

REVIEW 3 cited by

Contrastive Learning for Unpaired Image-to-Image Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.15651 v3 pith:VH7GF47E submitted 2020-07-30 cs.CV cs.LG

classification cs.CVcs.LG
keywords translationcontrastiveimageimage-to-imagelearningmethodsettingcorresponding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In image-to-image translation, each patch in the output should reflect the content of the corresponding patch in the input, independent of domain. We propose a straightforward method for doing so -- maximizing mutual information between the two, using a framework based on contrastive learning. The method encourages two elements (corresponding patches) to map to a similar point in a learned feature space, relative to other elements (other patches) in the dataset, referred to as negatives. We explore several critical design choices for making contrastive learning effective in the image synthesis setting. Notably, we use a multilayer, patch-based approach, rather than operate on entire images. Furthermore, we draw negatives from within the input image itself, rather than from the rest of the dataset. We demonstrate that our framework enables one-sided translation in the unpaired image-to-image translation setting, while improving quality and reducing training time. In addition, our method can even be extended to the training setting where each "domain" is only a single image.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing

    q-bio.QM 2025-06 conditional novelty 7.0 of 10

    DisProtEdit learns disentangled protein representations from separate structural and functional text descriptions, enabling controllable single- and multi-attribute protein editing via latent interpolation.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A lightweight hybrid-patch GAN enhances synthetic game images toward photorealism in real time, beating prior lightweight paired translators on speed, KID, and semantic consistency.

Pith tools