REVIEW 9 cited by
One-Step Effective Diffusion Network for Real-World Image Super-Resolution
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The pre-trained text-to-image diffusion models have been increasingly employed to tackle the real-world image super-resolution (Real-ISR) problem due to their powerful generative image priors. Most of the existing methods start from random noise to reconstruct the high-quality (HQ) image under the guidance of the given low-quality (LQ) image. While promising results have been achieved, such Real-ISR methods require multiple diffusion steps to reproduce the HQ image, increasing the computational cost. Meanwhile, the random noise introduces uncertainty in the output, which is unfriendly to image restoration tasks. To address these issues, we propose a one-step effective diffusion network, namely OSEDiff, for the Real-ISR problem. We argue that the LQ image contains rich information to restore its HQ counterpart, and hence the given LQ image can be directly taken as the starting point for diffusion, eliminating the uncertainty introduced by random noise sampling. We finetune the pre-trained diffusion network with trainable layers to adapt it to complex image degradations. To ensure that the one-step diffusion model could yield HQ Real-ISR output, we apply variational score distillation in the latent space to conduct KL-divergence regularization. As a result, our OSEDiff model can efficiently and effectively generate HQ images in just one diffusion step. Our experiments demonstrate that OSEDiff achieves comparable or even better Real-ISR results, in terms of both objective metrics and subjective evaluations, than previous diffusion model-based Real-ISR methods that require dozens or hundreds of steps. The source codes are released at https://github.com/cswry/OSEDiff.
Forward citations
Cited by 9 Pith papers
-
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.
-
The Devil is in the Dark Pixels: Toward Brightness Bias-Robust Denoising
BBRD, a drop-in replacement for MSE that normalizes per-brightness-band errors by their own noise variance and upweights the lagging band via softmax Group-DRO, improves dark, bright, and aggregate PSNR simultaneously...
-
Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
A transfer training scheme converts Stable Diffusion's 8x VAE into a 4x VAE that stays compatible with the pretrained UNet, improving fine-structure preservation in real-world super-resolution at lower FLOPs.
-
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
PBR-SR super-resolves PBR texture maps (albedo, roughness, metallic, normal) in a zero-shot way by optimizing textures so differentiable renderings match super-resolved multi-view renderings from a pretrained image SR model.
-
One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
One-step diffusion super-resolution via CLIP semantic alignment and DWT high-frequency perception distillation improves no-reference perceptual quality scores.
-
Compression-Aware One-Step Diffusion Model for JPEG Artifact Removal
CODiff adds a compression-aware embedder to a one-step diffusion model and reports state-of-the-art perceptual metrics for JPEG artifact removal in a single sampling step.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
-
Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models
VLM-based selection and ensembling of diffusion super-resolution samples is proposed, together with a hybrid Trustworthiness Score, but the score is partially fitted and its claimed human-preference correlation is not...
-
ControlMambaIR: Conditional Controls with State-Space Model for Image Restoration
A diffusion image restoration model with a Mamba condition network reports low LPIPS/FID on several benchmarks, but the PSNR losses and internal inconsistencies undermine the stated performance claims.
Discussion (0). Continue with ORCID to comment.