REVIEW 19 cited by
One-Step Image Translation with Text-to-Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we address two limitations of existing conditional diffusion models: their slow inference speed due to the iterative denoising process and their reliance on paired data for model fine-tuning. To tackle these issues, we introduce a general method for adapting a single-step diffusion model to new tasks and domains through adversarial learning objectives. Specifically, we consolidate various modules of the vanilla latent diffusion model into a single end-to-end generator network with small trainable weights, enhancing its ability to preserve the input image structure while reducing overfitting. We demonstrate that, for unpaired settings, our model CycleGAN-Turbo outperforms existing GAN-based and diffusion-based methods for various scene translation tasks, such as day-to-night conversion and adding/removing weather effects like fog, snow, and rain. We extend our method to paired settings, where our model pix2pix-Turbo is on par with recent works like Control-Net for Sketch2Photo and Edge2Image, but with a single-step inference. This work suggests that single-step diffusion models can serve as strong backbones for a range of GAN learning objectives. Our code and models are available at https://github.com/GaParmar/img2img-turbo.
Forward citations
Cited by 19 Pith papers
-
FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting
FieryGS integrates LLM-based material reasoning, volumetric combustion simulation, and a unified renderer with 3D Gaussian Splatting to generate physically plausible and user-controllable fire in in-the-wild scenes.
-
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
PRISM is a flow-matching method for unpaired image translation that uses a learned per-feature gate, derived from target-distribution distance, to control both initialization and transport timing, improving realism an...
-
GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes
GSRAIN fuses measured-raindrop-calibrated Gaussian streaks/haze with a diffusion-based rainy-appearance transfer in 3D Gaussian Splatting scenes, enabling 0–13 mm/h rainfall control for autonomous-driving testing.
-
GMODiff: One-Step Gain Map Refinement with Diffusion Priors for HDR Reconstruction
HDR reconstruction is reformulated as one-step gain map refinement, enabling a pre-trained latent diffusion model plus regression priors to produce high-quality HDR at a fraction of the cost of prior diffusion approaches.
-
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
Elastic3D converts monocular video to stereo by directly synthesizing the right-eye view with a one-step latent diffusion model conditioned on a user-set median disparity, using a guided decoder to preserve left-view details.
-
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
A multi-stage diffusion-based framework that generates labeled synthetic aerial images from weak image-level labels improves cross-domain vehicle detection AP50 over prior adaptation methods.
-
UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model
UniSegDiff uses staged training and inference with alternating mask/noise prediction targets plus STAPLE fusion of multiple samples to reach state-of-the-art lesion segmentation across six datasets and modalities.
-
UNICE: Training A Universal Image Contrast Enhancer
UNICE trains a two-stage model to generate and fuse a pseudo multi-exposure sequence from one image, generalizing across four contrast-enhancement tasks without human labels.
-
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.
-
The Devil is in the Darkness: Diffusion-Based Nighttime Dehazing Anchored in Brightness Perception
DiffND combines a depth- and sky-guided data synthesis pipeline with a diffusion model gated by a brightness perception network to achieve nighttime dehazing with day-level brightness.
-
FastFace: Tuning Identity Preservation in Distilled Diffusion via Guidance and Attention
An inference-time framework of decoupled classifier-free guidance and attention manipulation improves identity preservation and prompt alignment when pretrained face ID adapters are used with few-step distilled diffus...
-
WeatherEdit: Controllable Weather Editing with 4D Gaussian Field
WeatherEdit generates controllable, temporally consistent weather effects (snow, rain, fog) in driving scenes by editing 2D backgrounds with a multi-style adapter and rendering dynamic particles with a 4D Gaussian field.
-
BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion
One latent diffusion model, with token-level cross-modal attention, performs calibration-free visible-guided infrared super-resolution and infrared-visible fusion as two outputs of the same process.
-
Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data
A single latent-diffusion model, trained with cycle consistency, self-distillation, and CLIP guidance on unpaired driving data, edits fog, rain, and snow in driving scenes and modestly improves downstream perception.
-
Translationese as a Rational Response to Translation Task Difficulty
Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.
-
From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
An unsupervised diffusion-based enhancer with caption, reflectance, and cycle-attention consistency losses improves zero-shot classification, face detection, and night segmentation on low-light images.
-
CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
CycleVAR adapts a pretrained visual autoregressive model to unpaired image translation using softmax-relaxed quantization and source-token prefixes, achieving FID scores competitive with CycleGAN-Turbo.
-
ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views
ProSplat combines a feed-forward 3D Gaussian Splatting generator with a one-step diffusion improvement model and epipolar-aware attention, reporting about 1 dB PSNR gain over SOTA on wide-baseline sparse-view novel vi...
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
Discussion (0). Continue with ORCID to comment.