Pith. sign in

REVIEW 3 cited by

Activating More Pixels in Image Super-Resolution Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.04437 v3 pith:TXXDMUE6 submitted 2022-05-09 eess.IV cs.CV

classification eess.IVcs.CV
keywords transformerattentionbetterfurtherimageinformationinputmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is still not fully exploited in existing networks. In order to activate more input pixels for better reconstruction, we propose a novel Hybrid Attention Transformer (HAT). It combines both channel attention and window-based self-attention schemes, thus making use of their complementary advantages of being able to utilize global statistics and strong local fitting capability. Moreover, to better aggregate the cross-window information, we introduce an overlapping cross-attention module to enhance the interaction between neighboring window features. In the training stage, we additionally adopt a same-task pre-training strategy to exploit the potential of the model for further improvement. Extensive experiments show the effectiveness of the proposed modules, and we further scale up the model to demonstrate that the performance of this task can be greatly improved. Our overall method significantly outperforms the state-of-the-art methods by more than 1dB. Codes and models are available at https://github.com/XPixelGroup/HAT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study

    cs.CV 2025-04 conditional novelty 5.0 of 10

    KwaiSR is a new image super-resolution benchmark made from short-form user-generated content, and existing SR models and quality metrics struggle on it.

  2. NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results

    eess.IV 2025-04 conditional novelty 4.0 of 10

    The NTIRE 2025 challenge report presents leaderboards for efficient video quality assessment on KVQ and diffusion-based image super-resolution on the new KwaiSR dataset, with user-study results showing subjective and ...

  3. $\text{S}^{3}$Mamba: Arbitrary-Scale Super-Resolution via Scaleable State Space Model

    cs.CV 2024-11 reject novelty 4.0 of 10

    S3Mamba applies scale-modulated state space models to arbitrary-scale super-resolution, reporting marginal PSNR gains over prior INR-based methods.

Pith tools