Pith. sign in

REVIEW 10 cited by

Exploiting Diffusion Prior for Real-World Image Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07015 v4 pith:SKB2Q2ID submitted 2023-05-11 cs.CV

classification cs.CV
keywords diffusionmodelspre-trainedpriorfidelityreal-worldsuper-resolutionachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a novel approach to leverage prior knowledge encapsulated in pre-trained text-to-image diffusion models for blind super-resolution (SR). Specifically, by employing our time-aware encoder, we can achieve promising restoration results without altering the pre-trained synthesis model, thereby preserving the generative prior and minimizing training cost. To remedy the loss of fidelity caused by the inherent stochasticity of diffusion models, we employ a controllable feature wrapping module that allows users to balance quality and fidelity by simply adjusting a scalar value during the inference process. Moreover, we develop a progressive aggregation sampling strategy to overcome the fixed-size constraints of pre-trained diffusion models, enabling adaptation to resolutions of any size. A comprehensive evaluation of our method using both synthetic and real-world benchmarks demonstrates its superiority over current state-of-the-art approaches. Code and models are available at https://github.com/IceClear/StableSR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  2. Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    In VAR-based super-resolution, K2N predicts the first three coarse scales in parallel from the low-resolution input and generates only the remaining fine scales autoregressively, reducing hallucination while staying c...

  3. Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A transfer training scheme converts Stable Diffusion's 8x VAE into a 4x VAE that stays compatible with the pretrained UNet, improving fine-structure preservation in real-world super-resolution at lower FLOPs.

  4. Robust ID-Specific Face Restoration via Alignment Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RIDFR injects a reference person's identity into diffusion-based face restoration and uses Alignment Learning across multiple same-identity references to suppress pose, expression, and makeup interference.

  5. MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A cascaded, segmentation-conditioned, per-instance diffusion method synthesizes globally coherent gigapixel microscopic detail from a phone photo and sparse microscope references at up to 350×.

  6. DECAF: De-Clustering for Adaptive Representational Unlearning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    DECAF is a forget-only unlearning method that adds input noise, suppresses the forget-class probability, and diversifies outputs, achieving 0.10% forget accuracy and 79.4% retain accuracy on CIFAR-10/ResNet-18 while d...

  7. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  8. Efficient Burst Super-Resolution with One-step Diffusion

    cs.CV 2025-07 reject novelty 5.0 of 10

    E-BSRD applies EDM sampling and consistency-model distillation to burst super-resolution, achieving one-step diffusion at 0.44 s/image while roughly matching or slightly underperforming the multi-step BSRD baseline on...

  9. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  10. Incorporating Uncertainty-Guided and Top-k Codebook Matching for Real-World Blind Image Super-Resolution

    cs.CV 2025-06 conditional novelty 4.0 of 10

    UGTSR improves codebook-based blind super-resolution by combining uncertainty-guided loss weighting, top-3 codebook matching, and an align-attention module for fusing low- and high-quality features.

Pith tools