Pith. sign in

REVIEW 3 cited by

Local Distortion Aware Efficient Transformer Adaptation for Image Quality Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12001 v1 pith:7DNPPL6B submitted 2023-08-23 cs.CV

classification cs.CV
keywords distortionlocalpretrainedfeaturesimagefoundationlarge-scalemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image Quality Assessment (IQA) constitutes a fundamental task within the field of computer vision, yet it remains an unresolved challenge, owing to the intricate distortion conditions, diverse image contents, and limited availability of data. Recently, the community has witnessed the emergence of numerous large-scale pretrained foundation models, which greatly benefit from dramatically increased data and parameter capacities. However, it remains an open problem whether the scaling law in high-level tasks is also applicable to IQA task which is closely related to low-level clues. In this paper, we demonstrate that with proper injection of local distortion features, a larger pretrained and fixed foundation model performs better in IQA tasks. Specifically, for the lack of local distortion structure and inductive bias of vision transformer (ViT), alongside the large-scale pretrained ViT, we use another pretrained convolution neural network (CNN), which is well known for capturing the local structure, to extract multi-scale image features. Further, we propose a local distortion extractor to obtain local distortion features from the pretrained CNN and a local distortion injector to inject the local distortion features into ViT. By only training the extractor and injector, our method can benefit from the rich knowledge in the powerful foundation models and achieve state-of-the-art performance on popular IQA datasets, indicating that IQA is not only a low-level problem but also benefits from stronger high-level features drawn from large-scale pretrained models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MS-IQA: A Multi-Scale Feature Fusion Network for PET/CT Image Quality Assessment

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A multi-scale ResNet and Swin Transformer fusion network with channel attention predicts PET/CT image quality scores from a new 2,700-image radiologist-labeled dataset and outperforms existing IQA methods.

  2. Visual-Language Model Knowledge Distillation Method for Image Quality Assessment

    cs.CV 2025-07 reject novelty 5.0 of 10

    A CLIP teacher with quality-grade prompts distills knowledge into lightweight student encoders via cosine-annealed soft and hard label losses, reporting SOTA IQA scores at 5-28M parameters.

  3. Toward Total Recall: Enhancing FAIRness through AI-Driven Metadata Standardization

    cs.IR 2025-02 conditional novelty 5.0 of 10

    Combining GPT-4 with CEDAR metadata templates to standardize biomedical metadata substantially improves exact-match retrieval recall over raw metadata.

Pith tools