Pith. sign in

REVIEW 7 cited by

Palette: Image-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.05826 v2 pith:TPXFMCON submitted 2021-11-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusionimage-to-imageevaluationmodelstranslationarchitectureframeworkloss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper develops a unified framework for image-to-image translation based on conditional diffusion models and evaluates this framework on four challenging image-to-image translation tasks, namely colorization, inpainting, uncropping, and JPEG restoration. Our simple implementation of image-to-image diffusion models outperforms strong GAN and regression baselines on all tasks, without task-specific hyper-parameter tuning, architecture customization, or any auxiliary loss or sophisticated new techniques needed. We uncover the impact of an L2 vs. L1 loss in the denoising diffusion objective on sample diversity, and demonstrate the importance of self-attention in the neural architecture through empirical studies. Importantly, we advocate a unified evaluation protocol based on ImageNet, with human evaluation and sample quality scores (FID, Inception Score, Classification Accuracy of a pre-trained ResNet-50, and Perceptual Distance against original images). We expect this standardized evaluation protocol to play a role in advancing image-to-image translation research. Finally, we show that a generalist, multi-task diffusion model performs as well or better than task-specific specialist counterparts. Check out https://diffusion-palette.github.io for an overview of the results.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Strong Gravitational Lensing Posterior Sampling in Pixel-Space Using Diffusion Models and Recurrent Inference Machines

    astro-ph.IM 2026-07 conditional novelty 7.0 of 10

    DiRIM uses a diffusion model with recurrent score refinement to sample pixel-space joint posteriors of the lensed source and foreground mass map, reproducing mock strong-lens observations to the noise level.

  2. Detangled: A Framework for Creating, Editing, and Inferencing Feature Rich Hair Strands

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A 5D texture parameterization plus centerline-based canonical space and supervised diffusion enables generation and texture transfer of feature-rich hair strands independent of style.

  3. SHFormer: Dynamic Spectral Filtering Convolutional Neural Network and High-pass Kernel Generation Transformer for Adaptive MRI Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SHFormer, a hybrid spectral-filtering CNN and high-pass-kernel transformer, improves accelerated MRI reconstruction and generalization to unseen contrasts and acceleration factors.

  4. Semantic Color Naturalness Breaker: Preventing Illegitimate Colorization via Content-Aware Color Priors

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A tiny, invisible perturbation added to a published grayscale image can force AI colorizers to produce content-wrong colors (blue apples), an effect this paper measures and optimizes with a new semantic color-plausibi...

  5. AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content

    cs.CV 2025-08 conditional novelty 6.0 of 10

    AU-IQA is a new 4,800-image benchmark showing existing quality models, mainly those trained on ordinary user content, only partially predict human ratings of AI-enhanced photos.

  6. Pinterest Canvas: Large-Scale Image Generation at Pinterest

    cs.CV 2026-03 conditional novelty 4.0 of 10

    A FLUX-style base diffusion model plus task-specific fine-tunes and product-preserving pipelines yields double-digit Pinterest ads engagement lifts and higher no-defect rates than GPT-Image, FLUX Kontext, and Nano Banana.

  7. DiffPR: Diffusion-Based Phase Reconstruction via Frequency-Decoupled Learning

    eess.IV 2025-06 reject novelty 4.0 of 10

    DiffPR couples a quarter-resolution U-Net phase predictor with an unconditional diffusion refiner and reports solid but modest gains over U-Net baselines on four QPI datasets, with the claimed spectral-bias mechanism ...

Pith tools