Pith. sign in

REVIEW 7 cited by

Pixel-Aware Stable Diffusion for Realistic Image Super-resolution and Personalized Stylization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14469 v4 pith:43XOT6BW submitted 2023-08-28 cs.CV

classification cs.CV
keywords imagediffusionpasdstylizationmodelspixel-awarestabletasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have demonstrated impressive performance in various image generation, editing, enhancement and translation tasks. In particular, the pre-trained text-to-image stable diffusion models provide a potential solution to the challenging realistic image super-resolution (Real-ISR) and image stylization problems with their strong generative priors. However, the existing methods along this line often fail to keep faithful pixel-wise image structures. If extra skip connections between the encoder and the decoder of a VAE are used to reproduce details, additional training in image space will be required, limiting the application to tasks in latent space such as image stylization. In this work, we propose a pixel-aware stable diffusion (PASD) network to achieve robust Real-ISR and personalized image stylization. Specifically, a pixel-aware cross attention module is introduced to enable diffusion models perceiving image local structures in pixel-wise level, while a degradation removal module is used to extract degradation insensitive features to guide the diffusion process together with image high level information. An adjustable noise schedule is introduced to further improve the image restoration results. By simply replacing the base diffusion model with a stylized one, PASD can generate diverse stylized images without collecting pairwise training data, and by shifting the base model with an aesthetic one, PASD can bring old photos back to life. Extensive experiments in a variety of image enhancement and stylization tasks demonstrate the effectiveness of our proposed PASD approach. Our source codes are available at \url{https://github.com/yangxy/PASD/}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  2. Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A transfer training scheme converts Stable Diffusion's 8x VAE into a 4x VAE that stays compatible with the pretrained UNet, improving fine-structure preservation in real-world super-resolution at lower FLOPs.

  3. TextSR: Diffusion Super-Resolution with Multilingual OCR Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TextSR super-resolves multilingual scene text by conditioning a diffusion model on UTF-8-encoded OCR characters, achieving top OCR accuracy on TextZoom and on small/medium text in a self-defined TextVQA evaluation.

  4. MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A cascaded, segmentation-conditioned, per-instance diffusion method synthesizes globally coherent gigapixel microscopic detail from a phone photo and sparse microscope references at up to 350×.

  5. Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SCST reports the best perceptual quality (LPIPS/DISTS) on four synthetic benchmarks and the best no-reference quality scores on the real-world VideoLQ benchmark by adding spatio-temporal Mamba and contrastive ControlN...

  6. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  7. Incorporating Uncertainty-Guided and Top-k Codebook Matching for Real-World Blind Image Super-Resolution

    cs.CV 2025-06 conditional novelty 4.0 of 10

    UGTSR improves codebook-based blind super-resolution by combining uncertainty-guided loss weighting, top-3 codebook matching, and an align-attention module for fusing low- and high-quality features.

Pith tools