Pith. sign in

REVIEW 4 major objections 6 minor 107 references

NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This challenge report claims that for diffusion-based super-resolution of short-form user-generated content, the ranking preferred by human experts does not match the ranking given by standard objective quality metrics.

desk verdict A competent NTIRE challenge report whose only scientific claim — subjective/objective inconsistency in Track 2 — rests on a five-expert user study with no reliability statistics; the paper is worth refereeing but that conclusion needs scaling back or supporting. read the letter →

arxiv 2504.13131 v1 pith:7BZAGTZ6 submitted 2025-04-17 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords short-formUGCvideoqualityassessmentefficientVQAdiffusion-basedsuper-resolutionKwaiSRdatasetsubjectivepreferenceobjectiveperceptualmetricsbenchmarkchallenge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This challenge report claims that two things can be done for short-form user-generated content (S-UGC): video quality can be assessed accurately by lightweight models, and diffusion-based super-resolution can improve the subjective quality of in-the-wild images. The paper's central finding, however, is that the two goals do not line up: in the super-resolution track, the ranking produced by standard objective metrics disagreed with the ranking produced by five expert viewers. A method that won the objective leaderboard was rated fourth by experts, while the experts' top choice ranked sixth objectively. If the user study is representative, leaderboard positions on PSNR, LPIPS, or MUSIQ should not be read as statements about which enhanced image people actually prefer.

What carries the argument

The argument is carried by the challenge's evaluation protocol and its new dataset. Track 1 uses the KVQ database with a coarse-to-fine scoring scheme and a hard limit of 120 GFLOPs, forcing contestants to trade accuracy against compute. Track 2 introduces the KwaiSR dataset (1,800 synthetic paired images and 1,900 real-world low-quality images, split 8:1:1) and evaluates the six objectively best submissions through a five-expert user study in which each expert spent about eight hours choosing the most realistic result. That human-preference step is the load-bearing mechanism: it produces the subjective rankings against which the objective metrics are compared.

What would settle it

Run the same pairwise user study with a larger and more diverse panel, say 50 or more non-expert viewers, and compute a majority ranking with agreement statistics. If the majority ranking matches the objective ordering, or if some standard metric such as LPIPS or MUSIQ correlates with the larger panel's preferences, then the claimed inconsistency is an artifact of the five-expert panel rather than a general property of current metrics.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes a benchmark result and a negative result. The benchmark result is that efficient video quality assessment under a 120 GFLOPs limit is viable: the winning Track 1 model scored 0.922 on the combined metric while using 47.39 GFLOPs and 33.01M parameters, with the top five teams all above 0.90. The negative result is that, for diffusion-based super-resolution of short-form UGC images, subjective preference and objective quality metrics diverge. In the user study of the six teams shortlisted by objective performance, TACO SR had the highest expert winning rates (0.2775 on synthetic and 0.3529 on wild images) yet ranked sixth by objective measures, while the objectively top-ranked team, SYSU-FVL-Team, ranked fourth in expert preference. The paper states this as evidence that current perceptual metrics may not reliably reflect perceived quality in generative-model-based S-UGC super-resolution.

Load-bearing premise

The whole subjective-versus-objective conclusion rests on the assumption that five professional image-processing experts, each spending about eight hours, represent how viewers in general would rank the six super-resolution outputs; the paper reports winning rates but no inter-rater agreement, confidence intervals, or statistical test.

Editorial extensions

If this is right

  • Generative super-resolution leaderboards that rely on PSNR, SSIM, LPIPS, or MUSIQ may reward the wrong teams; future challenge reports should include a human preference stage or a metric calibrated to it.
  • Lightweight VQA is deployable at platform scale: the best Track 1 result was achieved at 47.39 GFLOPs with 33.01M parameters, well under the 120 GFLOPs ceiling.
  • The KwaiSR dataset gives the research community a shared test bed with both synthetic pairs and real low-quality images for short-form UGC super-resolution.
  • The top Track 1 teams used teacher-student pseudo-labeling or hybrid Mamba-attention designs to stay under the compute budget, so efficiency can come from training strategy rather than from shrinking a single network alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the expert votes may be tracking a realism axis, namely texture plausibility and absence of generative artifacts, that fidelity-oriented metrics like PSNR and SSIM are structurally blind to; a metric that scores realism rather than reconstruction error would likely close part of the gap.
  • Editorial inference: if the inconsistency generalizes beyond five experts, the same divergence should appear in other generative restoration tasks such as face restoration, denoising, and video enhancement, where objective benchmarks similarly dominate.
  • Editorial inference: the six final outputs plus the expert preference judgments could be reused as training data for a lightweight reward model or ranking-based no-reference metric; whether such a metric beats MUSIQ on held-out expert choices is a direct, testable next step.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This challenge report describes the NTIRE 2025 competition on short-form UGC video quality assessment and enhancement. Track 1 evaluates efficient no-reference VQA models on the KVQ dataset under a 120 GFLOPs constraint, and Track 2 introduces the KwaiSR dataset for diffusion-based single-image super-resolution, reporting objective quality metrics and a user study of the top six teams. The paper presents the leaderboard for Track 1, summarizes the architecture, training, and testing details of each of the 18 participating teams, and concludes that subjective preferences and objective metrics are noticeably inconsistent for generative SR on this content. The challenge repository is made publicly available, and the paper serves as a record of methods and results for the community.

Significance. If the reported results are reliable, the paper offers a useful community resource: a new S-UGC super-resolution dataset (KwaiSR, detailed in a companion paper), a public challenge repository, and a structured overview of 18 submitted methods under a strict efficiency budget. The Track 1 leaderboard, showing strong performance within 120 GFLOPs, is of practical interest for lightweight VQA deployment. The claimed subjective/objective inconsistency in Track 2, if supported by proper statistical evidence, would be an important caution against reading PSNR, LPIPS, or MUSIQ as proxies for user preference when comparing generative SR outputs. The paper is transparent about per-team training protocols and makes participant fact sheets available, which is commendable. However, the evidence for the central interpretive claim is currently thin, and the Track 1 out-of-sample comparison is compromised by at least one team explicitly using test-set pseudo-labels in training.

major comments (4)
  1. [Section 4.3 (ZX-AIE-Vector)] The training procedure described for ZX-AIE-Vector includes generating pseudo-labels for the KVQ test set, merging them with the training and validation data, and then fine-tuning the lightweight model on the refined pseudo-labeled test data using two training phases. This is a transductive use of the test set and means the reported test performance is not a clean out-of-sample evaluation. The paper does not state whether this practice was permitted by the challenge rules. Because Table 1 is the central deliverable of Track 1, the authors must disclose the rule and, if the practice was not allowed, recompute the leaderboard without this team. At minimum, a sensitivity analysis showing the ranking with and without test-set adaptation is needed.
  2. [Section 3 and Table 2] The claim that "current perceptual metrics may not reliably reflect perceptual quality" rests entirely on a user study described in a single sentence: five professional image-processing experts spent about eight hours each comparing the top six teams' outputs. No number of test images, no stimulus presentation protocol, no aggregation rule, no inter-rater agreement measure, no confidence intervals, and no significance test are reported. With only five raters and six teams, the differences between adjacent winning rates are small (e.g., 0.2775 vs. 0.2640 on synthetic; 0.1540 vs. 0.0947 on wild), so a single expert's preference flips several rank orders. Moreover, the raters were asked to select "the most visually convincing and realistic result," an expert-centric criterion that does not necessarily represent the general viewer population of short-form UGC platforms. The authors should either provide the full user-study protocol with statistical analysis (e.g., Fleiss' kappa, confidence intervals, or a test against chance) or substantially weaken the conclusion to a hypothesis rather than a finding.
  3. [Section 2 and Table 1] The Track 1 leaderboard is presented as a "Final Score" combining SROCC, PLCC, Rank1, and Rank2, but the paper never defines how this score is computed. The challenge description mentions coarse-grained quality scoring and fine-grained rankings for difficult samples, yet no formula, weighting, or aggregation rule is given. Without this definition, readers cannot interpret the rankings or reproduce the leaderboard. Please add the exact computation of Final Score, including how the fine-grained Rank1 and Rank2 components enter the score.
  4. [Table 2] The column header "User Study Score (objective)" is ambiguous: the first numeric column (e.g., 0.2775/0.3529) appears to be user-study winning rates, while the following columns are conventional objective metrics. The "Ranking (Objective)" column is not tied to any stated aggregation of PSNR, SSIM, LPIPS, MUSIQ, ManIQA, or CLIPIQA, and the text says the top six teams were shortlisted by objective metrics, yet objective scores are listed for all nine teams. Please restructure the table to clearly separate user-study results from objective metrics, define the objective ranking criterion, and state how the six teams were selected for the user study.
minor comments (6)
  1. [Section 2] The text states that KVQ contains "nine primary content scenarios" but then lists eight categories (landscape, crowd, person, food, portrait, computer graphic, caption, and stage). Please correct the count or add the missing category.
  2. [Section 1] The phrase "might inevitability suffer" should be "might inevitably suffer."
  3. [Abstract and Section 1] The repository URL "https://github.com/lixinustc/KVQE- Challenge-CVPR-NTIRE2025" contains a space after the hyphen; please provide a single, correct URL.
  4. [Appendix A and B] The organizer team blocks are titled "NTIRE2024 Organizers" although this is the NTIRE 2025 challenge; the year should be corrected.
  5. [Equation (1), Section 5.1] The loss definitions contain spacing and font inconsistencies, e.g., "Iper" appears both italic and non-italic, and the L1 term is written as "L1(f(Ipsnr,I per),I GT)". Please unify the notation.
  6. [Section 3] The sentence "The teams with top performances including SharpMind, ZQE, ZX-AIE-Vector, ECNU-SJTU VQA Team, and TenVQA achieved excellent results..." is informal; please rephrase to a measured description of the numerical results.

Circularity Check

1 steps flagged · score 4.0 of 10

One team's Track 1 test score is partly built from the test set itself via pseudo-label retraining; the paper's main subjective/objective inconsistency claim is empirical, not circular.

  1. fitted input called prediction [Section 4.3, ZX-AIE-Vector, Training details]
    "Next, they fine-tune a large-scale model on the KVQ dataset [51] (merging the official training and validation sets), and use it to generate pseudo-labels for the test set. These pseudo-labels are then merged with the original KVQ training and validation data to construct an augmented dataset. ... In the second-phase fine-tuning, they use the full KVQ training and validation sets along with the refined pseudo-labeled test data to retrain their lightweight model for an additional 30 epochs. ... During testing, they evaluate their final lightweight model on the KVQ test set."

    The team trains on the test set itself (via pseudo-labels) and then reports a 'test' score. The final evaluation on the KVQ test set is therefore not an out-of-sample prediction: the test inputs and their pseudo-labels were already seen during training, so the reported rank-3 final score (0.912) is partly constructed from the test data rather than independently predicted. The paper discloses this strategy, but the leaderboard entry is still a fitted input renamed as a prediction.

full rationale

The paper is primarily a competition report, so most results are measured benchmark outcomes rather than derived claims. The central interpretive statement in Section 3 and Table 2—that there is a noticeable inconsistency between subjective preferences and objective metrics—is an empirical comparison between an independent five-expert user study and objective scores; it is not forced by definition, and no equation makes the subjective ranking equal to the objective metrics. Its weakness is statistical (five experts, no inter-rater agreement or significance testing), not circular. Self-citations to KVQ and KwaiSR are dataset-provenance citations, and the datasets are public artifacts rather than load-bearing derivations. The one genuine circularity is ZX-AIE-Vector's use of pseudo-labeled test data in training followed by reporting of test performance, which makes that specific prediction reduce to a fitted input. Because this affects one team's leaderboard row rather than the paper's main scientific conclusion, the overall circularity score is moderate rather than high.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no mathematical model, fitted parameters, or new postulated entities. Its claims depend on the validity of the KVQ labels and the small user study, both listed as domain assumptions.

assumptions (2)
  • domain assumption KVQ ground-truth quality scores, annotated by professional researchers, are valid and representative for short-form UGC.
    All Track 1 leaderboard numbers (SROCC, PLCC) are correlations against these labels; if the labels are biased, the rankings are biased. Used throughout Section 3.
  • domain assumption The five-expert user study provides a reliable perceptual quality ranking.
    Track 2 subjective ranking and the inconsistency conclusion in Section 3 depend on this small panel being representative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results." pith.science (2026). https://pith.science/paper/7BZAGTZ6

@misc{pith2026250413131,
  author       = {Pith},
  title        = {Pith review of: NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7BZAGTZ6}},
  note         = {Machine review of arXiv:2504.13131}
}
read the original abstract

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image Super-Resolution (KwaiSR). Track 1 aims to advance the development of lightweight and efficient video quality assessment (VQA) models, with an emphasis on eliminating reliance on model ensembles, redundant weights, and other computationally expensive components in the previous IQA/VQA competitions. Track 2 introduces a new short-form UGC dataset tailored for single image super-resolution, i.e., the KwaiSR dataset. It consists of 1,800 synthetically generated S-UGC image pairs and 1,900 real-world S-UGC images, which are split into training, validation, and test sets using a ratio of 8:1:1. The primary objective of the challenge is to drive research that benefits the user experience of short-form UGC platforms such as Kwai and TikTok. This challenge attracted 266 participants and received 18 valid final submissions with corresponding fact sheets, significantly contributing to the progress of short-form UGC VQA and image superresolution. The project is publicly available at https://github.com/lixinustc/KVQE- ChallengeCVPR-NTIRE2025.

Figures

Figures reproduced from arXiv: 2504.13131 by the authors.

Figure 1
Figure 1. Enhancement results by Kwai-LPM (Large Processing Model), which is a diffusion-based SR method. 3. Challenge Results The challenge results are presented in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A comparison of the subjective quality between six teams on the synthetic dataset part: Example 1. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. A comparison of the subjective quality between six teams on the synthetic dataset part: Example 2. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: A comparison of the subjective quality between six teams on the wild dataset part: Example 1. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: A comparison of the subjective quality between six teams on the wild dataset part: Example 2. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: The overall framework of Team ZQE. they generate 2,000 originals from the web-sourced videos and compress them using H.264 at 6 levels [51]. RQ-VQA is used to generate pseudo-labels for all videos. Finally, they pretrain E-VQA on this dataset using fidelity loss [75] a…
Figure 10
Figure 10. Figure 10: The overall framework of Team GoldenChef. [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 9
Figure 9. Figure 9: The overall framework of Team ECNU-SJTU. [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 11
Figure 11. Figure 11: The overall framework of Team DAIQAM. from each frame, a projector module is applied to reduce the feature dimensionality. The features from all frames are then concatenated and passed through a multilayer percep￾tron (MLP) for quality regression. Training details Dur…
Figure 13
Figure 13. Figure 13: The overall framework of Team PiNAFusion-SR. [PITH_FULL_IMAGE:figures/full_fig_p011_13.png]
Figure 12
Figure 12. Figure 12: The overall framework of Team Nourayn. Their model combines spatial feature extraction with a pre￾trained ResNet-50 backbone and Faster-RNN, alongside temporal modeling using a Bidirectional LSTM (BiLSTM). The overall design aims to predict the Mean Opinion Score (MOS…
Figure 14
Figure 14. Figure 14: The framework proposed by Team RealismDiff. [PITH_FULL_IMAGE:figures/full_fig_p012_14.png]
Figure 15
Figure 15. Figure 15: Overall Pipeline of the solution of Team SRlab. [PITH_FULL_IMAGE:figures/full_fig_p012_15.png]
Figure 17
Figure 17. Figure 17: The framework of Team ZigZagSeeSR [PITH_FULL_IMAGE:figures/full_fig_p014_17.png]
Figure 16
Figure 16. Figure 16: The overall pipeline of the method proposed by Team [PITH_FULL_IMAGE:figures/full_fig_p014_16.png]
Figure 18
Figure 18. Figure 18: The framework proposed by Team BVIVSR. encoder Eφ is responsible for extracting deep latent fea￾tures from the input low-resolution image. Dρ consists of a multi-scale hierarchical encoding module, multiple multi￾head linear attention blocks, and MLPs. Its hierarchica…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

107 extracted references · 45 canonical work pages

  1. [38]

    NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Kwaisr dataset and study

    Xin Li, Xijun Wang, Bingchen Li, Kun Yuan, Yizhen Shao, Suhang Yao, Ming Sun, Chao Zhou, Radu Timofte, and Zhibo Chen. NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Kwaisr dataset and study. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2025. 2

  2. [1]

    Ntire 2017 challenge on single image super-resolution: Dataset and study

    Eirikur Agustsson and Radu Timofte. Ntire 2017 challenge on single image super-resolution: Dataset and study. In CVPR workshops, pages 126–135, 2017. 12, 13, 15

  3. [2]

    Zigzag diffusion sam- pling: The path to success is zigzag

    Lichen Bai, Shitong Shao, Zikai Zhou, Zipeng Qi, Zhiqiang Xu, Haoyi Xiong, and Zeke Xie. Zigzag diffusion sam- pling: The path to success is zigzag. arXiv preprint arXiv:2412.10891, 2024. 14

  4. [3]

    Activat- ing more pixels in image super-resolution transformer

    X Chen, X Wang, J Zhou, and C Dong. Activat- ing more pixels in image super-resolution transformer. arXiv:2205.04437. 2

  5. [4]

    Activating more pixels in image super- resolution transformer

    Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, and Chao Dong. Activating more pixels in image super- resolution transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 22367–22377, 2023. 2

  6. [5]

    NTIRE 2025 challenge on image super-resolution (×4): Methods and results

    Zheng Chen, Kai Liu, Jue Gong, Jingkai Wang, Lei Sun, Zongwei Wu, Radu Timofte, Yulun Zhang, et al. NTIRE 2025 challenge on image super-resolution (×4): Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2025. 2

  7. [6]

    NTIRE 2025 challenge on real-world face restoration: Methods and results

    Zheng Chen, Jingkai Wang, Kai Liu, Jue Gong, Lei Sun, Zongwei Wu, Radu Timofte, Yulun Zhang, et al. NTIRE 2025 challenge on real-world face restoration: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2025. 2

  8. [7]

    Nafssr: Stereo image super-resolution using nafnet

    Xiaojie Chu, Liangyu Chen, and Wenqing Yu. Nafssr: Stereo image super-resolution using nafnet. InCVPR, pages 1239–1248, 2022. 11

Show all 107 references
  1. [8]

    NTIRE 2025 challenge on raw image restoration and super-resolution

    Marcos Conde, Radu Timofte, et al. NTIRE 2025 challenge on raw image restoration and super-resolution. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  2. [9]

    Raw image reconstruc- tion from RGB on smartphones

    Marcos Conde, Radu Timofte, et al. Raw image reconstruc- tion from RGB on smartphones. NTIRE 2025 challenge re- port. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) Workshops ,

  3. [10]

    Imagenet: A large-scale hierarchical im- age database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical im- age database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5

  4. [11]

    NTIRE 2025 challenge on night photography rendering

    Egor Ershov, Sergey Korchagin, Alexei Khalin, Artyom Panshin, Arseniy Terekhin, Ekaterina Zaychenkova, Georgiy Lobarev, Vsevolod Plokhotnyuk, Denis Abramov, Elisey Zhdanov, Sofia Dorogova, Yasin Mamedov, Nikola Banic, Georgii Perevozchikov, Radu Timofte, et al. NTIRE 2025 chal...

  5. [12]

    X3d: Expanding architectures for efficient video recognition

    Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 203–213, 2020. 7

  6. [13]

    Slowfast networks for video recognition

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019. 6

  7. [14]

    NTIRE 2025 challenge on cross-domain few-shot object detection: Methods and results

    Yuqian Fu, Xingyu Qiu, Bin Ren Yanwei Fu, Radu Timofte, Nicu Sebe, Ming-Hsuan Yang, Luc Van Gool, et al. NTIRE 2025 challenge on cross-domain few-shot object detection: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...

  8. [15]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 6

  9. [16]

    Internvqa: Advancing compressed video qual- ityassessment with distilling large foundation model

    Fengbin Guan, Zihao Yu, Yiting Lu, Xin Li, and Zhibo Chen. Internvqa: Advancing compressed video qual- ityassessment with distilling large foundation model. arXiv preprint arXiv:2502.19026, 2025. 2

  10. [17]

    Mam- bairv2: Attentive state space restoration

    Hang Guo, Yong Guo, Yaohua Zha, Yulun Zhang, Wenbo Li, Tao Dai, Shu-Tao Xia, and Yawei Li. Mam- bairv2: Attentive state space restoration. arXiv preprint arXiv:2411.15269, 2024. 14

  11. [18]

    Mambair: A simple baseline for image restoration with state-space model

    Hang Guo, Jinmin Li, Tao Dai, Zhihao Ouyang, Xudong Ren, and Shu-Tao Xia. Mambair: A simple baseline for image restoration with state-space model. InEuropean con- ference on computer vision, pages 222–241. Springer, 2024. 2

  12. [19]

    NTIRE 2025 challenge on text to image generation model qual- ity assessment

    Shuhao Han, Haotian Fan, Fangyuan Kong, Wenjie Liao, Chunle Guo, Chongyi Li, Radu Timofte, et al. NTIRE 2025 challenge on text to image generation model qual- ity assessment. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Workshop...

  13. [20]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 13

  14. [21]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 10

  15. [22]

    NTIRE 2025 challenge on video quality enhancement for video conferencing: Datasets, methods and results

    Varun Jain, Zongwei Wu, Quan Zou, Louis Florentin, Henrik Turbell, Sandeep Siddhartha, Radu Timofte, et al. NTIRE 2025 challenge on video quality enhancement for video conferencing: Datasets, methods and results. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision a...

  16. [23]

    Hiif: Hi- erarchical encoding based implicit image function for con- tinuous super-resolution

    Yuxuan Jiang, Ho Man Kwan, Tianhao Peng, Ge Gao, Fan Zhang, Xiaoqing Zhu, Joel Sole, and David Bull. Hiif: Hi- erarchical encoding based implicit image function for con- tinuous super-resolution. arXiv preprint arXiv:2412.03748,

  17. [24]

    C2d-isr: Optimizing attention-based image super-resolution from continuous to discrete scales

    Yuxuan Jiang, Chengxi Zeng, Siyue Teng, Fan Zhang, Xi- aoqing Zhu, Joel Sole, and David Bull. C2d-isr: Optimizing attention-based image super-resolution from continuous to discrete scales. arXiv preprint arXiv:2503.13740, 2025. 15

  18. [25]

    A style-based generator architecture for generative adversarial networks

    Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, pages 4401–4410, 2019. 13

  19. [26]

    One millisecond face alignment with an ensemble of regression trees

    Vahid Kazemi and Josephine Sullivan. One millisecond face alignment with an ensemble of regression trees. In CVPR, pages 1867–1874, 2014. 12

  20. [27]

    Accu- rate image super-resolution using very deep convolutional networks

    Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Accu- rate image super-resolution using very deep convolutional networks. In Proc. IEEE Conf. Comput. Vis. Pattern Recog- nit. workshops, pages 1646–1654, 2016. 2

  21. [28]

    Adam: A method for stochastic optimization

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ,

  22. [29]

    NTIRE 2025 challenge on efficient burst hdr and restoration: Datasets, methods, and results

    Sangmin Lee, Eunpil Park, Angel Canelo, Hyunhee Park, Youngjo Kim, Hyungju Chun, Xin Jin, Chongyi Li, Chun- Le Guo, Radu Timofte, et al. NTIRE 2025 challenge on efficient burst hdr and restoration: Datasets, methods, and results. In Proceedings of the IEEE/CVF Conference on Co...

  23. [30]

    Hst: Hierarchical swin transformer for com- pressed image super-resolution

    Bingchen Li, Xin Li, Yiting Lu, Sen Liu, Ruoyu Feng, and Zhibo Chen. Hst: Hierarchical swin transformer for com- pressed image super-resolution. In European conference on computer vision, pages 651–668. Springer, 2022. 2

  24. [31]

    Sed: Semantic-aware discriminator for image super-resolution

    Bingchen Li, Xin Li, Hanxin Zhu, Yeying Jin, Ruoyu Feng, Zhizheng Zhang, and Zhibo Chen. Sed: Semantic-aware discriminator for image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 25784–25795, 2024. 2

  25. [32]

    Videomamba: State space model for efficient video understanding

    Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. Videomamba: State space model for efficient video understanding. In European Conference on Computer Vision, pages 237–255. Springer,

  26. [33]

    A close look at few-shot real image super-resolution from the distortion relation perspective

    Xin Li, Xin Jin, Jun Fu, Xiaoyuan Yu, Bei Tong, and Zhibo Chen. A close look at few-shot real image super-resolution from the distortion relation perspective. arXiv preprint arXiv:2111.13078, 2021. 2

  27. [34]

    Learning distortion invariant representation for im- age restoration from a causality perspective

    Xin Li, Bingchen Li, Xin Jin, Cuiling Lan, and Zhibo Chen. Learning distortion invariant representation for im- age restoration from a causality perspective. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 1714–1724, 2023. 2

  28. [35]

    Ucip: A universal frame- work for compressed image super-resolution using dynamic prompt

    Xin Li, Bingchen Li, Yeying Jin, Cuiling Lan, Hanxin Zhu, Yulin Ren, and Zhibo Chen. Ucip: A universal frame- work for compressed image super-resolution using dynamic prompt. In European Conference on Computer Vision , pages 107–125. Springer, 2024. 2

  29. [36]

    Ntire 2024 challenge on short-form ugc video qual- ity assessment: Methods and results

    Xin Li, Kun Yuan, Yajing Pei, Yiting Lu, Ming Sun, Chao Zhou, Zhibo Chen, Radu Timofte, Wei Sun, Haoning Wu, et al. Ntire 2024 challenge on short-form ugc video qual- ity assessment: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  30. [37]

    NTIRE 2025 challenge on day and night raindrop removal for dual-focused images: Methods and results

    Xin Li, Yeying Jin, Xin Jin, Zongwei Wu, Bingchen Li, Yufei Wang, Wenhan Yang, Yu Li, Zhibo Chen, Bihan Wen, Robby Tan, Radu Timofte, et al. NTIRE 2025 challenge on day and night raindrop removal for dual-focused images: Methods and results. In Proceedings of the IEEE/CVF Conf...

  31. [39]

    NTIRE 2025 challenge on short-form ugc video qual- ity assessment and enhancement: Methods and results

    Xin Li, Kun Yuan, Bingchen Li, Fengbin Guan, Yizhen Shao, Zihao Yu, Xijun Wang, Yiting Lu, Wei Luo, Suhang Yao, Ming Sun, Chao Zhou, Zhibo Chen, Radu Timofte, et al. NTIRE 2025 challenge on short-form ugc video qual- ity assessment and enhancement: Methods and results. In Proc...

  32. [40]

    Lsdir: A large scale dataset for image restora- tion

    Yawei Li, Kai Zhang, Jingyun Liang, Jiezhang Cao, Ce Liu, Rui Gong, Yulun Zhang, Hao Tang, Yun Liu, Denis Deman- dolx, et al. Lsdir: A large scale dataset for image restora- tion. In Proc. IEEE Conf. Comput. Vis. Pattern Recog. , pages 1775–1787, 2023. 12, 13, 15

  33. [41]

    Swinir: Image restoration using swin transformer

    Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer. In Proceedings of the IEEE/CVF international conference on computer vision , pages 1833– 1844, 2021. 2

  34. [42]

    NTIRE 2025 the 2nd restore any image model (RAIM) in the wild challenge

    Jie Liang, Radu Timofte, Qiaosi Yi, Zhengqiang Zhang, Shuaizheng Liu, Lingchen Sun, Rongyuan Wu, Xindong Zhang, Hui Zeng, Lei Zhang, et al. NTIRE 2025 the 2nd restore any image model (RAIM) in the wild challenge. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion a...

  35. [43]

    Enhanced deep residual networks for single image super-resolution

    Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion workshops, pages 136–144, 2017. 2

  36. [44]

    Enhanced deep residual networks for sin- gle image super-resolution

    Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for sin- gle image super-resolution. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. workshops, pages 136–144, 2017. 2

  37. [45]

    Enhanced deep residual networks for single image super-resolution

    Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In CVPR workshops, pages 136–144, 2017. 12, 15

  38. [46]

    Ada- dqa: Adaptive diverse quality-aware feature acquisition for video quality assessment

    Hongbo Liu, Mingda Wu, Kun Yuan, Ming Sun, Yansong Tang, Chuanchuan Zheng, Xing Wen, and Xiu Li. Ada- dqa: Adaptive diverse quality-aware feature acquisition for video quality assessment. In ACM Multimedia, pages 6695–6704. ACM, 2023. 2

  39. [47]

    NTIRE 2025 XGC quality assessment challenge: Methods and results

    Xiaohong Liu, Xiongkuo Min, Qiang Hu, Xiaoyun Zhang, Jie Guo, et al. NTIRE 2025 XGC quality assessment challenge: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  40. [48]

    NTIRE 2025 challenge on low light image enhancement: Methods and results

    Xiaoning Liu, Zongwei Wu, Florin-Alexandru Vasluianu, Hailong Yan, Bin Ren, Yulun Zhang, Shuhang Gu, Le Zhang, Ce Zhu, Radu Timofte, et al. NTIRE 2025 challenge on low light image enhancement: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion ...

  41. [49]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 7

  42. [50]

    Swin transformer v2: Scaling up capacity and resolu- tion

    Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolu- tion. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pages 12009–12019, 2022. 2

  43. [51]

    Kvq: Kwai video quality assessment for short-form videos

    Yiting Lu, Xin Li, Yajing Pei, Kun Yuan, Qizhi Xie, Yun- peng Qu, Ming Sun, Chao Zhou, and Zhibo Chen. Kvq: Kwai video quality assessment for short-form videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 25963–25973, 2024. 2,...

  44. [52]

    Q-adapt: Adapting lmm for visual qual- ity assessment with progressive instruction tuning

    Yiting Lu, Xin Li, Haoning Wu, Bingchen Li, Weisi Lin, and Zhibo Chen. Q-adapt: Adapting lmm for visual qual- ity assessment with progressive instruction tuning. arXiv preprint arXiv:2504.01655, 2025. 2

  45. [53]

    Efficient transformer for single image super-resolution

    Zhisheng Lu, Hong Liu, Juncheng Li, and Linlin Zhang. Efficient transformer for single image super-resolution. arXiv:2108.11084, 2021. 2

  46. [54]

    Transformer for single image super-resolution

    Zhisheng Lu, Juncheng Li, Hong Liu, Chaoyan Huang, Lin- lin Zhang, and Tieyong Zeng. Transformer for single image super-resolution. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pages 457–466, 2022. 2

  47. [55]

    Reduced- reference video quality assessment of compressed video se- quences

    Lin Ma, Songnan Li, and King Ngi Ngan. Reduced- reference video quality assessment of compressed video se- quences. IEEE Transactions on circuits and systems for video technology, 22(10):1441–1456, 2012. 2

  48. [56]

    Bvi-aom: A new training dataset for deep video compression optimization

    Jakub Nawała, Yuxuan Jiang, Fan Zhang, Xiaoqing Zhu, Joel Sole, and David Bull. Bvi-aom: A new training dataset for deep video compression optimization. In VCIP, pages 1–5. IEEE, 2024. 15

  49. [57]

    Scalable diffusion mod- els with transformers

    William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. In ICCV, pages 4195–4205, 2023. 13

  50. [58]

    XPSR: cross-modal priors for diffusion-based image super-resolution

    Yunpeng Qu, Kun Yuan, Kai Zhao, Qizhi Xie, Jinhua Hao, Ming Sun, and Chao Zhou. XPSR: cross-modal priors for diffusion-based image super-resolution. In ECCV (11), pages 285–303. Springer, 2024. 2

  51. [59]

    KVQ: boosting video quality as- sessment via saliency-guided local perception

    Yunpeng Qu, Kun Yuan, Qizhi Xie, Ming Sun, Chao Zhou, and Jian Wang. KVQ: boosting video quality as- sessment via saliency-guided local perception. CoRR, abs/2503.10259, 2025. 2

  52. [60]

    Sam 2: Segment anything in images and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Ro- man R ¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 12

  53. [61]

    The tenth NTIRE 2025 efficient super-resolution challenge report

    Bin Ren, Hang Guo, Lei Sun, Zongwei Wu, Radu Tim- ofte, Yawei Li, et al. The tenth NTIRE 2025 efficient super-resolution challenge report. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  54. [62]

    Mambacsr: Dual-interleaved scanning for compressed image super-resolution with ssms

    Yulin Ren, Xin Li, Mengxi Guo, Bingchen Li, Shijie Zhao, and Zhibo Chen. Mambacsr: Dual-interleaved scanning for compressed image super-resolution with ssms. arXiv preprint arXiv:2408.11758, 2024. 2

  55. [63]

    NTIRE 2025 challenge on UGC video enhancement: Meth- ods and results

    Nickolay Safonov, Alexey Bryntsev, Andrey Moskalenko, Dmitry Kulikov, Dmitriy Vatolin, Radu Timofte, et al. NTIRE 2025 challenge on UGC video enhancement: Meth- ods and results. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Works...

  56. [64]

    Laion-5b: An open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, 35: 25278–25294...

  57. [65]

    Video quality assessment by reduced reference spatio-temporal entropic differencing

    Rajiv Soundararajan and Alan C Bovik. Video quality assessment by reduced reference spatio-temporal entropic differencing. IEEE Transactions on Circuits and Systems for Video Technology, 23(4):684–694, 2012. 2

  58. [66]

    Pixel-level and semantic- level adjustable super-resolution: A dual-lora approach

    Lingchen Sun, Rongyuan Wu, Zhiyuan Ma, Shuaizheng Liu, Qiaosi Yi, and Lei Zhang. Pixel-level and semantic- level adjustable super-resolution: A dual-lora approach. arXiv preprint arXiv:2412.03017, 2024. 11, 13

  59. [67]

    NTIRE 2025 challenge on event-based image deblurring: Methods and results

    Lei Sun, Andrea Alfarano, Peiqi Duan, Shaolin Su, Kaiwei Wang, Boxin Shi, Radu Timofte, Danda Pani Paudel, Luc Van Gool, et al. NTIRE 2025 challenge on event-based image deblurring: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  60. [68]

    The tenth ntire 2025 image denoising challenge report

    Lei Sun, Hang Guo, Bin Ren, Luc Van Gool, Radu Timo- fte, Yawei Li, et al. The tenth ntire 2025 image denoising challenge report. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  61. [69]

    Deep learning based full-reference and no- reference quality assessment models for compressed ugc videos

    Wei Sun, Tao Wang, Xiongkuo Min, Fuwang Yi, and Guangtao Zhai. Deep learning based full-reference and no- reference quality assessment models for compressed ugc videos. In 2021 IEEE International Conference on Mul- timedia & Expo Workshops (ICMEW) , pages 1–6. IEEE,

  62. [70]

    A deep learning based no-reference quality assessment model for ugc videos

    Wei Sun, Xiongkuo Min, Wei Lu, and Guangtao Zhai. A deep learning based no-reference quality assessment model for ugc videos. In Proceedings of the 30th ACM Interna- tional Conference on Multimedia, pages 856–865, 2022. 2, 7

  63. [71]

    A deep learning based no-reference quality assessment model for ugc videos

    Wei Sun, Xiongkuo Min, Wei Lu, and Guangtao Zhai. A deep learning based no-reference quality assessment model for ugc videos. In Proceedings of the 30th ACM Interna- tional Conference on Multimedia, pages 856–865, 2022. 2

  64. [72]

    Analysis of video quality datasets via design of minimalistic video quality models

    Wei Sun, Wen Wen, Xiongkuo Min, Long Lan, Guangtao Zhai, and Kede Ma. Analysis of video quality datasets via design of minimalistic video quality models. IEEE Transactions on Pattern Analysis and Machine Intelligence,

  65. [73]

    Enhancing blind video quality assessment with rich quality-aware features

    Wei Sun, Haoning Wu, Zicheng Zhang, Jun Jia, Zhichao Zhang, Linhan Cao, Qiubo Chen, Xiongkuo Min, Weisi Lin, and Guangtao Zhai. Enhancing blind video quality assessment with rich quality-aware features. arXiv preprint arXiv:2405.08745, 2024. 2, 3, 8

  66. [74]

    An empirical study for efficient video quality assessment

    Wei Sun, Kang Fu, Linhan Cao, Dandan Zhu, Kai- wei Zhang, Yucheng Zhu, Zicheng Zhang, Menghan Hu, Xiongkuo Min, and Guangtao Zhai. An empirical study for efficient video quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...

  67. [75]

    Frank: a ranking method with fidelity loss

    Ming-Feng Tsai, Tie-Yan Liu, Tao Qin, Hsin-Hsi Chen, and Wei-Ying Ma. Frank: a ranking method with fidelity loss. In Proceedings of the 30th annual international ACM SI- GIR conference on Research and development in informa- tion retrieval, pages 383–390, 2007. 8

  68. [76]

    Ugc-vqa: Benchmarking blind video quality assessment for user generated content

    Zhengzhong Tu, Yilin Wang, Neil Birkbeck, Balu Adsumilli, and Alan C Bovik. Ugc-vqa: Benchmarking blind video quality assessment for user generated content. IEEE Transactions on Image Processing , 30:4449–4464,

  69. [77]

    Maxim: Multi-axis mlp for image processing

    Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxim: Multi-axis mlp for image processing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5769–5780, 2022. 2

  70. [78]

    NTIRE 2025 image shadow removal challenge report

    Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou, Cailian Chen, Zongwei Wu, Radu Timofte, et al. NTIRE 2025 image shadow removal challenge report. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  71. [79]

    NTIRE 2025 ambi- ent lighting normalization challenge

    Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou, Zongwei Wu, Radu Timofte, et al. NTIRE 2025 ambi- ent lighting normalization challenge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025. 2

  72. [80]

    haiqiang Wang, Gary Li, Shan Liu, and C.-C. Jay Kuo. Icme 2021 ugc-vqa challenge. In Available: http://ugcvqa.com/. 2

  73. [81]

    Lit: Delving into a simplified linear diffusion transformer for image generation

    Jiahao Wang, Ning Kang, Lewei Yao, Mengzhao Chen, Chengyue Wu, Songyang Zhang, Shuchen Xue, Yong Liu, Taiqiang Wu, Xihui Liu, et al. Lit: Delving into a simplified linear diffusion transformer for image generation. arXiv preprint arXiv:2501.12976, 2025. 13

  74. [82]

    Recovering realistic texture in image super-resolution by deep spatial feature transform

    Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Recovering realistic texture in image super-resolution by deep spatial feature transform. In CVPR, pages 606–615,

  75. [83]

    Real-esrgan: Training real-world blind super-resolution with pure synthetic data

    Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In ICCV, pages 1905–1914, 2021. 2, 13, 14

  76. [84]

    Sinsr: diffusion-based image super- resolution in a single step

    Yufei Wang, Wenhan Yang, Xinyuan Chen, Yaohui Wang, Lanqing Guo, Lap-Pui Chau, Ziwei Liu, Yu Qiao, Alex C Kot, and Bihan Wen. Sinsr: diffusion-based image super- resolution in a single step. In CVPR, pages 25796–25805,

  77. [85]

    Enhanced semantic extraction and guid- ance for ugc image super resolution

    Yiwen Wang, Ying Liang, Yuxuan Zhang, Xinning Chai, Zhengxue Cheng, Yinsheng Qin, Yucai Yang, rong Xie, and Li Song. Enhanced semantic extraction and guid- ance for ugc image super resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  78. [86]

    NTIRE 2025 challenge on light field image super-resolution: Methods and results

    Yingqian Wang, Zhengyu Liang, Fengyuan Zhang, Lvli Tian, Longguang Wang, Juncheng Li, Jungang Yang, Radu Timofte, Yulan Guo, et al. NTIRE 2025 challenge on light field image super-resolution: Methods and results. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision a...

  79. [87]

    Fast- vqa: Efficient end-to-end video quality assessment with fragment sampling

    Haoning Wu, Chaofeng Chen, Jingwen Hou, Liang Liao, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Fast- vqa: Efficient end-to-end video quality assessment with fragment sampling. In European conference on computer vision, pages 538–554. Springer, 2022. 2, 7

  80. [88]

    Exploring video quality assessment on user gener- ated contents from aesthetic and technical perspectives

    Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jing- wen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring video quality assessment on user gener- ated contents from aesthetic and technical perspectives. In Proceedings of the IEEE/CVF International Conferenc...

  81. [89]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al. Q-align: Teaching lmms for visual scoring via discrete text-defined levels. arXiv preprint arXiv:2312.17090, 2023. 2

  82. [90]

    Q-instruct: Improving low-level visual abilities for multi-modality foundation models

    Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Kaixin Xu, Chunyi Li, Jingwen Hou, Guangtao Zhai, et al. Q-instruct: Improving low-level visual abilities for multi-modality foundation models. In Proceedings of the IEEE/CVF conference on computer v...

  83. [91]

    Seesr: Towards semantics- aware real-world image super-resolution

    Rongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang, Shuai Li, and Lei Zhang. Seesr: Towards semantics- aware real-world image super-resolution. In CVPR, pages 25456–25467, 2024. 14

  84. [92]

    QPT-V2: masked image mod- eling advances visual scoring

    Qizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu, Ming Sun, Chao Zhou, and Jihong Zhu. QPT-V2: masked image mod- eling advances visual scoring. In ACM Multimedia, pages 2709–2718. ACM, 2024. 2

  85. [93]

    NTIRE 2025 challenge on single image reflection removal in the wild: Datasets, methods and results

    Kangning Yang, Jie Cai, Ling Ouyang, Florin-Alexandru Vasluianu, Radu Timofte, Jiaming Ding, Huiming Sun, Lan Fu, Jinlong Li, Chiu Man Ho, Zibo Meng, et al. NTIRE 2025 challenge on single image reflection removal in the wild: Datasets, methods and results. In Proceedings of th...

  86. [94]

    Aim 2022 challenge on super-resolution of compressed image and video: Dataset, methods and results

    Ren Yang, Radu Timofte, Xin Li, Qi Zhang, Lin Zhang, Fanglong Liu, Dongliang He, Fu Li, He Zheng, Weihang Yuan, et al. Aim 2022 challenge on super-resolution of compressed image and video: Dataset, methods and results. In European Conference on Computer Vision , pages 174–

  87. [95]

    Wider face: A face detection benchmark

    Shuo Yang, Ping Luo, Chen-Change Loy, and Xiaoou Tang. Wider face: A face detection benchmark. In CVPR, pages 5525–5533, 2016. 13

  88. [96]

    Patch-vq:’patching up’the video qual- ity problem

    Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, and Alan Bovik. Patch-vq:’patching up’the video qual- ity problem. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14019– 14029, 2021. 6

  89. [97]

    Teaching large language models to regress accu- rate image quality scores using score distribution

    Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, and Chao Dong. Teaching large language models to regress accu- rate image quality scores using score distribution. arXiv preprint arXiv:2501.11561, 2025. 3

  90. [98]

    Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild

    Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiang- tao Kong, Xintao Wang, Jingwen He, Yu Qiao, and Chao Dong. Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild. In CVPR, pages 25669–25680, 2024. 12

  91. [99]

    Sf-iqa: Quality and similarity integration for ai generated image quality assessment

    Zihao Yu, Fengbin Guan, Yiting Lu, Xin Li, and Zhibo Chen. Sf-iqa: Quality and similarity integration for ai generated image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6692–6701, 2024. 2

  92. [100]

    Video quality assessment based on swin trans- formerv2 and coarse to fine strategy

    Zihao Yu, Fengbin Guan, Yiting Lu, Xin Li, and Zhibo Chen. Video quality assessment based on swin trans- formerv2 and coarse to fine strategy. arXiv preprint arXiv:2401.08522, 2024. 2

  93. [101]

    Ef- ficient diffusion model for image restoration by residual shifting

    Zongsheng Yue, Jianyi Wang, and Chen Change Loy. Ef- ficient diffusion model for image restoration by residual shifting. PAMI, 2024. 13, 14

  94. [102]

    NTIRE 2025 challenge on hr depth from images of specular and transparent surfaces

    Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, Samuele Salti, Stefano Mattoccia, et al. NTIRE 2025 challenge on hr depth from images of specular and transparent surfaces. In Proceedings of the IEEE/CVF Conference on Computer ...

  95. [103]

    A spatial–temporal video quality as- sessment method via comprehensive hvs simulation

    Ao-Xiang Zhang, Yuan-Gen Wang, Weixuan Tang, Leida Li, and Sam Kwong. A spatial–temporal video quality as- sessment method via comprehensive hvs simulation. IEEE Transactions on Cybernetics, 54(8):4749–4762, 2023. 4

  96. [104]

    Swinfir: Revisiting the swinir with fast fourier convolution and improved training for image super- resolution

    Dafeng Zhang, Feiyu Huang, Shizhuo Liu, Xiaobing Wang, and Zhezhu Jin. Swinfir: Revisiting the swinir with fast fourier convolution and improved training for image super- resolution. arXiv:2208.11247, 2022. 2

  97. [105]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. workshops, pages 586–595,

  98. [106]

    Blind image quality assessment via vision- language correspondence: A multitask learning perspec- tive

    Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. Blind image quality assessment via vision- language correspondence: A multitask learning perspec- tive. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 14071–14081,

  99. [107]

    Quality-aware pre-trained models for blind image quality assessment

    Kai Zhao, Kun Yuan, Ming Sun, Mading Li, and Xing Wen. Quality-aware pre-trained models for blind image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22302– 22313, 2023. 2

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.