Pith. sign in

REVIEW 5 cited by

Zero-shot Image-to-Image Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03027 v1 pith:37TFQU37 submitted 2023-02-06 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords editingimagecontentexistinginputmethodmodelschanges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale text-to-image generative models have shown their remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First, it is hard for users to come up with a perfect text prompt that accurately describes every visual detail in the input image. Second, while existing models can introduce desirable changes in certain regions, they often dramatically alter the input content and introduce unexpected changes in unwanted regions. In this work, we propose pix2pix-zero, an image-to-image translation method that can preserve the content of the original image without manual prompting. We first automatically discover editing directions that reflect desired edits in the text embedding space. To preserve the general content structure after editing, we further propose cross-attention guidance, which aims to retain the cross-attention maps of the input image throughout the diffusion process. In addition, our method does not need additional training for these edits and can directly use the existing pre-trained text-to-image diffusion model. We conduct extensive experiments and show that our method outperforms existing and concurrent works for both real and synthetic image editing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.

  3. Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting editing pipeline that combines 2D diffusion localization, depth-based point seeding, and sequential view refinement to achieve up to 4x faster local edits.

  4. Generative AI for Industrial Contour Detection: A Language-Guided Vision System

    cs.CV 2025-08 reject novelty 4.0 of 10

    A GAN-plus-VLM pipeline improves industrial remnant contour extraction, with GPT-image-1 outperforming Gemini 2.0 Flash on SSIM, LPIPS, and Hausdorff distance.

  5. SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework

    cs.CV 2025-01 reject novelty 4.0 of 10

    SmartSpatial combines depth injection and attention guidance in Stable Diffusion with a new VLM-based spatial metric, but reported improvements are not statistically significant per the paper's own p-value statement.

Pith tools