Pith. sign in

REVIEW 19 cited by

One-Step Image Translation with Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12036 v1 pith:MOXKA3VO submitted 2024-03-18 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords modeldiffusionmodelssingle-stepexistingimageinferencelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we address two limitations of existing conditional diffusion models: their slow inference speed due to the iterative denoising process and their reliance on paired data for model fine-tuning. To tackle these issues, we introduce a general method for adapting a single-step diffusion model to new tasks and domains through adversarial learning objectives. Specifically, we consolidate various modules of the vanilla latent diffusion model into a single end-to-end generator network with small trainable weights, enhancing its ability to preserve the input image structure while reducing overfitting. We demonstrate that, for unpaired settings, our model CycleGAN-Turbo outperforms existing GAN-based and diffusion-based methods for various scene translation tasks, such as day-to-night conversion and adding/removing weather effects like fog, snow, and rain. We extend our method to paired settings, where our model pix2pix-Turbo is on par with recent works like Control-Net for Sketch2Photo and Edge2Image, but with a single-step inference. This work suggests that single-step diffusion models can serve as strong backbones for a range of GAN learning objectives. Our code and models are available at https://github.com/GaParmar/img2img-turbo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting

    cs.GR 2026-04 unverdicted novelty 7.0 of 10

    FieryGS integrates LLM-based material reasoning, volumetric combustion simulation, and a unified renderer with 3D Gaussian Splatting to generate physically plausible and user-controllable fire in in-the-wild scenes.

  2. PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    PRISM is a flow-matching method for unpaired image translation that uses a learned per-feature gate, derived from target-distribution distance, to control both initialization and transport timing, improving realism an...

  3. GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GSRAIN fuses measured-raindrop-calibrated Gaussian streaks/haze with a diffusion-based rainy-appearance transfer in 3D Gaussian Splatting scenes, enabling 0–13 mm/h rainfall control for autonomous-driving testing.

  4. GMODiff: One-Step Gain Map Refinement with Diffusion Priors for HDR Reconstruction

    cs.CV 2025-12 conditional novelty 6.0 of 10

    HDR reconstruction is reformulated as one-step gain map refinement, enabling a pre-trained latent diffusion model plus regression priors to produce high-quality HDR at a fraction of the cost of prior diffusion approaches.

  5. Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Elastic3D converts monocular video to stereo by directly synthesizing the right-eye view with a one-step latent diffusion model conditioned on a user-set median disparity, using a guided decoder to preserve left-view details.

  6. Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A multi-stage diffusion-based framework that generates labeled synthetic aerial images from weak image-level labels improves cross-domain vehicle detection AP50 over prior adaptation methods.

  7. UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model

    eess.IV 2025-07 conditional novelty 6.0 of 10

    UniSegDiff uses staged training and inference with alternating mask/noise prediction targets plus STAPLE fusion of multiple samples to reach state-of-the-art lesion segmentation across six datasets and modalities.

  8. UNICE: Training A Universal Image Contrast Enhancer

    cs.CV 2025-07 conditional novelty 6.0 of 10

    UNICE trains a two-stage model to generate and fuse a pseudo multi-exposure sequence from one image, generalizing across four contrast-enhancement tasks without human labels.

  9. StableCodec: Taming One-Step Diffusion for Extreme Image Compression

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.

  10. The Devil is in the Darkness: Diffusion-Based Nighttime Dehazing Anchored in Brightness Perception

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DiffND combines a depth- and sky-guided data synthesis pipeline with a diffusion model gated by a brightness perception network to achieve nighttime dehazing with day-level brightness.

  11. FastFace: Tuning Identity Preservation in Distilled Diffusion via Guidance and Attention

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An inference-time framework of decoupled classifier-free guidance and attention manipulation improves identity preservation and prompt alignment when pretrained face ID adapters are used with few-step distilled diffus...

  12. WeatherEdit: Controllable Weather Editing with 4D Gaussian Field

    cs.CV 2025-05 conditional novelty 6.0 of 10

    WeatherEdit generates controllable, temporally consistent weather effects (snow, rain, fog) in driving scenes by editing 2D backgrounds with a multi-style adapter and rendering dynamic particles with a 4D Gaussian field.

  13. BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

    cs.CV 2026-07 conditional novelty 5.0 of 10

    One latent diffusion model, with token-level cross-modal attention, performs calibration-free visible-guided infrared super-resolution and infrared-visible fusion as two outputs of the same process.

  14. Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single latent-diffusion model, trained with cycle consistency, self-distillation, and CLIP guidance on unpaired driving data, edits fog, rain, and snow in driving scenes and modestly improves downstream perception.

  15. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  16. From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An unsupervised diffusion-based enhancer with caption, reflectance, and cycle-attention consistency losses improves zero-shot classification, face detection, and night segmentation on low-light images.

  17. CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    CycleVAR adapts a pretrained visual autoregressive model to unpaired image translation using softmax-relaxed quantization and source-token prefixes, achieving FID scores competitive with CycleGAN-Turbo.

  18. ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ProSplat combines a feed-forward 3D Gaussian Splatting generator with a one-step diffusion improvement model and epipolar-aware attention, reporting about 1 dB PSNR gain over SOTA on wide-baseline sparse-view novel vi...

  19. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

Pith tools