Pith. sign in

REVIEW 12 cited by

ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02938 v2 pith:OFUO5GV7 submitted 2021-08-06 cs.CV

classification cs.CV
keywords ddpmimagegenerationimagesmethodgenerativeilvrprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Denoising diffusion probabilistic models (DDPM) have shown remarkable performance in unconditional image generation. However, due to the stochasticity of the generative process in DDPM, it is challenging to generate images with the desired semantics. In this work, we propose Iterative Latent Variable Refinement (ILVR), a method to guide the generative process in DDPM to generate high-quality images based on a given reference image. Here, the refinement of the generative process in DDPM enables a single DDPM to sample images from various sets directed by the reference image. The proposed ILVR method generates high-quality images while controlling the generation. The controllability of our method allows adaptation of a single DDPM without any additional learning in various image generation tasks, such as generation from various downsampling factors, multi-domain image translation, paint-to-image, and editing with scribbles.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MENO restores multi-scale structure in neural-operator PDE surrogates via one-step improved MeanFlow, claiming up to 2× better power-spectrum accuracy and up to 14× faster inference than DDIM enhancement.

  2. RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Using an LLM to generate prompt-specific visual rubrics and grade each criterion independently gives a more interpretable reward that improves text-to-image model alignment beyond composite and learned scalar rewards.

  3. Score-based Diffusion Model for Unpaired Virtual Histology Staining

    eess.IV 2025-06 conditional novelty 6.0 of 10

    An unpaired, mutual-information-guided diffusion model translates H&E histology images into IHC images with improved structural and staining fidelity.

  4. FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

    cs.CV 2026-08 conditional novelty 5.0 of 10

    FlowForm is a diffusion model that adds shallow-water-equation penalties and terrain-conditioned adapters to synthesize flood satellite images, and reports top scores on a new 10,000-pair dataset.

  5. Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    cs.LG 2026-07 reject novelty 5.0 of 10

    A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.

  6. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  7. Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade

    cs.LG 2025-12 conditional novelty 5.0 of 10

    A cascade of a functional autoencoder (coarse structure) and a residual conditional diffusion model (fine details), with mask-cascade training and manifold-constrained gradients, reconstructs sparse-sensed physical fields.

  8. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  9. Time-variant Image Inpainting via Interactive Distribution Transition Estimation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    The authors introduce time-variant image inpainting (TAMP), a benchmark (TAMP-Street), and InDiTE-Diff, a diffusion-based method with a semantic complementation module that outperforms prior reference-guided inpaintin...

  10. Solving Inverse Problems via Diffusion-Based Priors: An Approximation-Free Ensemble Sampling Approach

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A weighted-particle sampler evolves the posterior through the diffusion model's reverse dynamics, with theoretical error bounds and improved image reconstructions.

  11. Restoring Real-World Images with an Internal Detail Enhancement Diffusion Model

    cs.CV 2025-05 conditional novelty 5.0 of 10

    This paper fine-tunes ControlNet on a frozen Stable Diffusion prior using a self-regularization objective that conditions on both the degraded input and a DDIM-estimated clean version, improving perceptual quality and...

  12. Projection-Based Correction for Enhancing Deep Inverse Networks

    cs.LG 2025-05 conditional novelty 2.0 of 10

    A projection step that forces a deep network's reconstruction to satisfy y = Ax gives small PSNR gains in low-noise imaging tests, but the supporting theory is a restatement of the definition of a well-trained network.

Pith tools