Pith. sign in

REVIEW 2 cited by

Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15798 v1 pith:KELARHA6 submitted 2024-12-20 cs.CV

classification cs.CV
keywords guidanceimagetargetdiffusionmodelsourceimage-to-imagelatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery Localization

    cs.CV 2025-06 conditional novelty 4.0 of 10

    M2SFormer, a Transformer encoder architecture with multi-spectral/multi-scale skip-connection attention and curvature-based difficulty guidance, achieves state-of-the-art cross-domain image forgery localization on sev...

  2. Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization

    cs.CV 2025-05 conditional novelty 4.0 of 10

    CASO learns per-attribute continuous embeddings via classifier gradients, steering Stable Diffusion for text-free, disentangled image editing.

Pith tools