Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

Video Quality Assessment: A Comprehensive Survey

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Video quality assessment has shifted from handcrafted statistical metrics to deep and multimodal models, and this survey benchmarks that shift on user-generated and AI-generated video, finding the learned models ahead except on temporal…

desk verdict Useful, current VQA survey, but the benchmark tables mix protocols and need provenance before the comparative claims can be trusted. read the letter →

arxiv 2412.04508 v2 pith:GLB73QR6 submitted 2024-12-04 eess.IV cs.CV

classification eess.IVcs.CV
keywords videoqualityassessmentsubjectivestudiesfull-referenceVQAno-referencedeeplearninglargemultimodalmodelsuser-generatedcontentAI-generated
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish a reliable map of the current video quality assessment field: how subjective databases are built, how full-reference and no-reference algorithms have evolved, and which models actually predict human judgments best on emerging content. Its central comparative claim is that deep learning and large multimodal models now outperform traditional handcrafted statistical metrics on user-generated and AI-generated video, while several handcrafted full-reference models remain competitive on temporal distortions such as frame-rate variation. If the map is accurate, it tells streaming platforms and codec designers where to place their bets: learned no-reference models for real-world UGC, multimodal models for AIGC, and bespoke temporal models for high-frame-rate and frame-interpolated content. The survey also argues that data scarcity and the cost of subjective studies are the main limits on further progress.

What carries the argument

The device that carries the survey's argument is a double taxonomy plus a comparative benchmark. The taxonomy separates subjective studies and databases from objective algorithms, and within objective algorithms separates full-reference from no-reference and knowledge-driven from deep learning-based; within deep models it separates temporal pooling, 3D CNNs, transformers, and large multimodality models. The benchmark then maps representative models onto databases with legacy, UGC, and AIGC content, using SROCC and PLCC, two rank and linear correlation measures against human mean opinion scores. This machinery lets the authors convert a literature review into a performance map, and it is also the point where the survey's assumptions are concentrated.

What would settle it

Re-run every model in Tables V and VI on the same databases under one fixed protocol: identical train and test splits, identical preprocessing and frame sampling, identical nonlinear mapping to MOS, and identical subject subsets. If the resulting SROCC and PLCC rankings differ materially from the tables, the survey's comparative insights and design recommendations are not supported.

Watch

Extended reading notes

Core claim

In the paper's own framing, the field of video quality assessment has shifted from measuring predefined distortion properties to learning quality from human-labeled video, and its benchmark section is the evidence. The survey classifies objective models into knowledge-driven versus deep learning-based categories, then reports SROCC and PLCC correlations with human scores for representative full-reference and no-reference models on five FR databases and five NR databases, including user-generated and AI-generated content. On those tables, transformer-based and large-multimodality models such as FAST-VQA, DOVER, COVER, MaxVQA, Q-Align, and LMM-VQA show the highest correlations on UGC and AIGC, while knowledge-driven models such as VMAF and ST-GREED remain strong on traditional and temporal distortions. The paper's stated conclusion is that effective temporal modeling and the integration of multimodal priors are the two most productive directions for future VQA.

Load-bearing premise

The paper's comparative conclusions rest on the premise that the SROCC and PLCC numbers in its benchmark tables are directly comparable across models, even though they are taken from different papers with unstated differences in training splits and preprocessing.

Editorial extensions

If this is right

  • For user-generated content, the survey implies practical blind quality monitoring should use deep or transformer-based no-reference models rather than frame-level image metrics, since the benchmark shows knowledge-driven IQA models performing poorly on video.
  • For AI-generated video, multimodal models and prompt-based scoring are the only tested family that tracks human judgment across both fidelity and text alignment, so AIGC evaluation is likely to be built around them.
  • For high-frame-rate and frame-interpolated video, general-purpose metrics are insufficient; the survey's tables show bespoke temporal models like ST-GREED and FloLPIPS winning those databases, so codec and interpolation evaluation should use task-specific metrics.
  • Temporal aggregation is not optional: the survey's comparison indicates that simple frame score averaging fails, and memory-aware pooling or recurrent or temporal modules are needed to match human perception.
  • If these rankings hold, adopting the leading learned metrics as loss functions in coding and enhancement pipelines should improve perceptual optimization beyond what SSIM and VMAF-based losses achieve today.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The benchmark's cross-paper numbers leave an open question: because SROCC and PLCC values are drawn from original papers where training splits and preprocessing differ, the exact ordering of models is less certain than the broad split between handcrafted and learned approaches; a unified re-evaluation could shift positions within each family.
  • A natural testable extension would be an AIGC benchmark that holds the generator, prompt set, and frame rate fixed while varying only the VQA model, to isolate how much of the large-multimodal-model advantage comes from text alignment versus raw fidelity scoring.
  • The survey's emphasis on temporal distortion suggests that future gains on UGC and AIGC may come from explicit motion and memory modules rather than from larger spatial backbones, a hypothesis one could test by ablating temporal modules while holding backbone size fixed.
  • The dual demands of AIGC quality, perceptual fidelity and prompt-video alignment, may require VQA to split into two scores rather than one mean opinion score, extending the aesthetic and technical decomposition the survey documents in DOVER.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript is a survey of video quality assessment (VQA), covering subjective study methodology and VQA databases, full-reference and no-reference objective models, deep learning architectures, loss functions, benchmark comparisons on legacy/UGC/AIGC content, and applications and challenges. Its stated central claim is that it provides a comprehensive survey of recent progress in VQA algorithms and of the benchmarking studies and databases that support them. The only novel empirical contribution is the performance comparison in Tables V and VI, which reports SROCC/PLCC for representative FR and NR image/video quality models across five FR databases and five NR databases.

Significance. If the taxonomy and the benchmark comparison are reliable, this would be a useful reference for the VQA community: the paper covers a very broad literature, including recent transformer-based and large-multimodality-model methods, summarizes a large number of subjective databases including AIGC databases, and provides design-oriented observations. It also makes a public GitHub repository available. However, the survey's reliability is currently weakened by two load-bearing problems: the evaluation protocol underlying Tables V and VI is not specified, and the citation numbering is internally inconsistent in ways that make it difficult to trace claims to sources. The paper does not ship machine-checked proofs or reproducible evaluation code for the benchmark, so the comparison must be assessed on the strength of the protocol description, which is currently insufficient.

major comments (3)
  1. [Section V-B, Tables V and VI] The only empirical contribution of the paper is the cross-model comparison in Tables V and VI, and the conclusions in Sections V-C and V-D (e.g., that transformer-based and LMM-based models "demonstrate the most excellent performance") depend directly on the comparability of the SROCC/PLCC values. Section V-B states only that "if available, performance data was taken from the original papers, otherwise we conducted the evaluation," without specifying train/test splits, whether image models were used off-the-shelf or fine-tuned, frame sampling and temporal pooling, preprocessing, or the logistic fitting used for PLCC. Some entries in the tables are themselves implausible under any single protocol: for example, Table V reports DeepQA at 0.0815 SROCC on LIVE-YT-HFR while LPIPS reports 0.6920, and Table VI reports SAMA at 0.0136 SROCC on T2VQA-DB while DOVER reports 0.7609. These gaps are far more likely to reflect different evaluation protocols than model quality. The authors should provide a per-entry provenance for every number in Tables V and VI, release the evaluation code and checkpoints, and either rerun all models under one protocol or clearly mark the provenance and add caveats where numbers are not comparable.
  2. [Throughout; e.g., Sections II-A, II-B, IV-A and Fig. 7] The citation numbering is internally inconsistent, which is load-bearing for a survey whose purpose is to let readers trace claims. SSIM is cited as [13] in Section II-A but as [164] in Section IV-A; MS-SSIM is cited as [13] in Section II-A and as [14] in Fig. 7; VMAF is cited as [43] in Section II-B, as [171] in Section IV-A Type iv, and as [173] in the same subsection; ST-GREED is cited as [137] in Table V and as [138] in Section IV-A; MC-SSIM is cited as [166] in the text and as [171] in Fig. 7; 3D-SSIM is cited as [167] in the text and as [172] in Fig. 7. A reader cannot reliably identify which reference a number or claim belongs to, so the "comprehensive survey" claim is weakened. The manuscript needs a systematic reference audit before it can be considered publication-ready.
  3. [Section V-C and V-D] The design insights drawn from Tables V and VI go beyond what the data can support given the protocol described. For example, the observation that "deep learning-based IQA models perform reasonably well on general distortion datasets" and the advice to prefer temporal NN modules are based on a small, non-random selection of models and on numbers whose protocol is opaque. Even if the protocol problem were fixed, the paper should temper these statements by noting the small number of databases, the content overlap among them, and the fact that many entries are copied from papers with different training regimes.
minor comments (4)
  1. [Section IV-C, Eq. (5)] Equation (5) writes the monotonicity loss as a double sum over i only (both summation indices are i), whereas the pair term L_ij^rank in Eq. (3) requires two distinct indices; the second index should be j.
  2. [Section V-A, Eq. (12)] The logistic function for nonlinear regression appears malformed: the expression exp{(-x + beta_3/|beta_4|}) has mismatched parentheses and the placement of the fraction is unclear. Please rewrite it in standard form, e.g., f(x) = beta_2 + (beta_1 - beta_2)/(1 + exp((x - beta_3)/|beta_4|)).
  3. [Throughout] There are numerous typographical errors that should be corrected in a final pass, including "percieved" (Introduction), "seventies" for "severities" (Section III-A1), "sucecesses" (Section II-D), "correspondance" (Section II-E), "insterest" (Section IV-B1), "histgram" (Fig. 7), "Pre-processig" (Table I), "A VC" for "AVC" (Section III-B3), and "spaital" for "spatial" (Table III). These do not block the scientific content but they do reduce the professionalism of the manuscript.
  4. [Section V-B] Please define the meaning of "-/-" entries in Tables V and VI (not evaluated, not reported in the original paper, or not applicable) and add a legend explaining that italic and orthographic fonts distinguish IQA from VQA models, since the table captions currently rely on the reader inferring this convention.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey synthesizes external literature and its only novel contribution, the benchmark tables, is explicitly grounded in external reported results rather than in any self-derived or self-fitted quantity.

full rationale

This paper is a comprehensive survey of video quality assessment algorithms, databases, and loss functions. It introduces no predictive model, no fitted parameter, and no derived quantity of its own. Its organizational claims (taxonomies in Figures 2 and 7, method descriptions in Section IV, loss formulations in Section IV-C) are restatements or summaries of published external work, with standard equation definitions (L1, L2, ranking loss, SROCC/PLCC formulas) that are textbook definitions rather than derivation outcomes. The central comparative contribution, Tables V and VI in Section V, is explicitly described as collecting performance data from original papers with the authors conducting evaluation only when numbers were not available ('If available, performance data was taken from the original papers, otherwise we conducted the evaluation'). Gathering and tabulating reported SROCC/PLCC values is not a derivation, and no claim reduces by construction to an input. The paper's conclusions (e.g., transformer-based and LMM-based models perform strongly on UGC and AIGC content) are empirical summaries of those tabulated external numbers, not predictions manufactured from the paper's own definitions. Authors do cite their own prior works (RAPIQUE, VIQE, STS-QA, COVER, temporal pooling studies), but these citations are contextual references to the literature being surveyed and are not load-bearing justifications of any derivation or uniqueness claim; the survey's structure does not depend on any single self-citation being accepted. The skeptical concern about protocol comparability of mixed benchmark numbers is a validity or reproducibility issue, not a circularity issue, since the numbers are not produced by the surveyed paper's own fitted parameters. Accordingly, there is no circular step to quote and no step where an equation reduces to its own input.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

This is a review paper, so there are no free parameters or invented entities. The main implicit assumption is that the external results cited, and the authors' own recomputations, are correct and directly comparable across different studies.

assumptions (1)
  • domain assumption The cited publications' reported results and the authors' own benchmark evaluations are accurate and comparable.
    The survey's central value as a reference depends on the trustworthiness of the numbers it compiles in Tables V and VI. If these numbers are inaccurate or not comparable, the survey's guidance is misleading.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Video Quality Assessment: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/GLB73QR6

@misc{pith2026241204508,
  author       = {Pith},
  title        = {Pith review of: Video Quality Assessment: A Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GLB73QR6}},
  note         = {Machine review of arXiv:2412.04508}
}
read the original abstract

Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video statistics, which are inspired both by models of projected images of the real world and by dual models of the human visual system, deliver only limited prediction performances on real-world user-generated content (UGC), as exemplified in recent large-scale VQA databases containing large numbers of diverse video contents crawled from the web. Fortunately, recent advances in deep neural networks and Large Multimodality Models (LMMs) have enabled significant progress in solving this problem, yielding better results than prior handcrafted models. Numerous deep learning-based VQA models have been developed, with progress in this direction driven by the creation of content-diverse, large-scale human-labeled databases that supply ground truth psychometric video quality data. Here, we present a comprehensive survey of recent progress in the development of VQA algorithms and the benchmarking studies and databases that make them possible. We also analyze open research directions on study design and VQA algorithm architectures. Github link: https://github.com/taco-group/Video-Quality-Assessment-A-Comprehensive-Survey.

Figures

Figures reproduced from arXiv: 2412.04508 by the authors.

Figure 1
Figure 1. Number of publications on image and video quality [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of existing subjective and objective video quality assessment methods. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Example of a visual interface used when playing video. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 6
Figure 6. Figure 6: Samples of video contents selected from legacy datasets ((a)-(f)), UGC datasets ((g)-(l)), and AIGC datasets ((m)-(r)). [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Classification and evolution of objective quality assessment models. The left figure presents framework diagrams for [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Three types of IQA framework based on how they extract and employ multi-scale features: the parallel, bottom-up and [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: The framework of LPIPS [80]. to weight similarities across space and channels. Unlike the above bottom-up paradigm, TOPIQ [178] adopts a top-down approach for both FR and NR image quality assessment. Multi-scale features from the first five layers of ResNet-50 [193] ar…
Figure 10
Figure 10. Figure 10: The framework of CLIP-IQA [100]. quality-aware contrastive loss. Type vi: Large multimodality model-based IQA. CLIP￾IQA [100] assesses image quality by leveraging the vision-language prior in CLIP [8]. It uses cosine similarity between image embeddings and text embedd…
Figure 11
Figure 11. Figure 11: The framework of VSFA [74]. video-level quality predictions. • Varga et al. [73] employs AlexNet [189], Inception￾V3 [268], and Inception-ResNet-V2 [269] pre-trained on KoNViD-10K for transfer learning. Frame-level features are extracted and processed through a two-la…
Figure 12
Figure 12. Figure 12: The framework of FAST-VQA [95]. between labeled source videos and unlabeled target videos. • Shen et al. [262] captures multi-scale motion through a spa￾tiotemporal feature pyramid, with pyramid attention modules for channel and spatial selective sensitivity. Type iv:…
Figure 13
Figure 13. Figure 13: Application overview of video quality assessment. [PITH_FULL_IMAGE:figures/full_fig_p025_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  2. Towards Standardized Light Field Quality Assessment: Hybrid Subjective Benchmarking and Objective Metric Evaluation

    cs.CV 2026-07 conditional novelty 5.5 of 10

    A hybrid DSCS+PC light-field QA framework and public dataset show objective metrics drop when view-synthesis/3DGS distortions join coding artifacts, and view pooling affects agreement.

Reference graph

Works this paper leans on

294 extracted references · 45 canonical work pages · cited by 2 Pith papers

  1. [13]

    Multiscale structural sim- ilarity for image quality assessment,

    Z. Wang, E. P. Simoncelli, and A. C. Bovik, “Multiscale structural sim- ilarity for image quality assessment,” in The Thirty-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 , vol. 2, 2003, pp. 1398–1402

  2. [164]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE Transactions on Image Processing , vol. 13, no. 4, pp. 600–612, 2004

  3. [171]

    Wasserstein generative ad- versarial networks,

    M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative ad- versarial networks,” in International Conference on Machine Learning, 2017, pp. 214–223

  4. [173]

    Toward a practical perceptual video quality metric,

    Z. Li, A. Aaron, I. Katsavounidis, A. Moorthy, and M. Manohara, “Toward a practical perceptual video quality metric,” The Netflix Tech Blog, vol. 6, no. 2, 2016

  5. [14]

    Complex wavelet structural similarity: A new image similarity index,

    M. P. Sampat, Z. Wang, S. Gupta, A. C. Bovik, and M. K. Markey, “Complex wavelet structural similarity: A new image similarity index,” IEEE Transactions on Image Processing , vol. 18, no. 11, pp. 2385– 2401, 2009

  6. [43]

    Scale mixtures of gaussians and the statistics of natural images,

    M. J. Wainwright and E. Simoncelli, “Scale mixtures of gaussians and the statistics of natural images,” in Advances in Neural Information Processing Systems , S. Solla, T. Leen, and K. M ¨uller, Eds., vol. 12. MIT Press, 1999. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 1999/file/6a5dfac4be1502501489fc0f5a24b667-Paper.pdf

  7. [137]

    ST-GREED: Space-time generalized entropic differences for frame rate dependent video quality prediction,

    P. C. Madhusudana, N. Birkbeck, Y . Wang, B. Adsumilli, and A. C. Bovik, “ST-GREED: Space-time generalized entropic differences for frame rate dependent video quality prediction,” IEEE Trans. on Image Processing, vol. 30, pp. 7446–7457, 2021

  8. [138]

    Making video quality assessment models robust to bit depth,

    J. P. Ebenezer, Z. Shang, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Making video quality assessment models robust to bit depth,” IEEE Signal Processing Letters , vol. 30, pp. 488–492, 2023

  9. [166]

    Efficient video quality assessment along temporal trajectories,

    A. K. Moorthy and A. C. Bovik, “Efficient video quality assessment along temporal trajectories,” IEEE Transactions on Circuits and Sys- tems for Video Technology, vol. 20, no. 11, pp. 1653–1658, 2010

  10. [167]

    3d-ssim for video quality assessment,

    K. Zeng and Z. Wang, “3d-ssim for video quality assessment,” in 2012 19th IEEE International Conference on Image Processing , 2012, pp. 621–624

  11. [172]

    Making Video Quality Assessment Models Sensitive to Frame Rate Distortions,

    P. C. Madhusudana, N. Birkbeck, Y . Wang, B. Adsumilli, and A. C. Bovik, “Making Video Quality Assessment Models Sensitive to Frame Rate Distortions,” IEEE Signal Processing Letters , vol. 29, pp. 897– 901, 2022

Show all 294 references
  1. [1]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017

  2. [2]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255

  3. [3]

    Deep convolutional neural models for picture-quality prediction: Chal- lenges and solutions to data-driven image quality assessment,

    J. Kim, H. Zeng, D. Ghadiyaram, S. Lee, L. Zhang, and A. C. Bovik, “Deep convolutional neural models for picture-quality prediction: Chal- lenges and solutions to data-driven image quality assessment,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 130–141, 2017

  4. [4]

    Patch- VQ:’patching up’ the video quality problem,

    Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch- VQ:’patching up’ the video quality problem,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 019–14 029

  5. [5]

    YouTube UGC dataset for video compression research,

    Y . Wang, S. Inguva, and B. Adsumilli, “YouTube UGC dataset for video compression research,” in IEEE International Workshop on Multimedia Signal Processing (MMSP) , 2019, pp. 1–5

  6. [6]

    Study of subjective and objective quality assessment of audio-visual signals,

    X. Min, G. Zhai, J. Zhou, M. C. Farias, and A. C. Bovik, “Study of subjective and objective quality assessment of audio-visual signals,” IEEE Transactions on Image Processing, vol. 29, pp. 6054–6068, 2020

  7. [7]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollar, and R. Girshick, “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 4015–4026

  8. [8]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine L...

  9. [9]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 34 892–34 916. [Online]. Available: ...

  10. [10]

    mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

    Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang, “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. ...

  11. [11]

    Q-instruct: Improving low-level visual abilities for multi-modality foundation models,

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, G. Xue, W. Sun, Q. Yan, and W. Lin, “Q-instruct: Improving low-level visual abilities for multi-modality foundation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...

  12. [15]

    Information content weighting for perceptual image quality assessment,

    Z. Wang and Q. Li, “Information content weighting for perceptual image quality assessment,” IEEE Transactions on Image Processing , vol. 20, no. 5, pp. 1185–1198, 2010

  13. [16]

    Edge strength similarity for image quality assessment,

    X. Zhang, X. Feng, W. Wang, and W. Xue, “Edge strength similarity for image quality assessment,” IEEE Signal Processing Letters, vol. 20, no. 4, pp. 319–322, 2013

  14. [17]

    Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,

    W. Xue, L. Zhang, X. Mou, and A. C. Bovik, “Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,” IEEE Transactions on Image Processing , vol. 23, no. 2, pp. 684–695, 2013

  15. [18]

    Image quality assessment based on gradient similarity,

    A. Liu, W. Lin, and M. Narwaria, “Image quality assessment based on gradient similarity,” IEEE Transactions on Image Processing , vol. 21, no. 4, pp. 1500–1512, 2011

  16. [19]

    FSIM: A feature similarity index for image quality assessment,

    L. Zhang, L. Zhang, X. Mou, and D. Zhang, “FSIM: A feature similarity index for image quality assessment,” IEEE Transactions on Image Processing, vol. 20, no. 8, pp. 2378–2386, 2011

  17. [20]

    VSI: A visual saliency-induced index for perceptual image quality assessment,

    L. Zhang, Y . Shen, and H. Li, “VSI: A visual saliency-induced index for perceptual image quality assessment,” IEEE Transactions on Image Processing, vol. 23, no. 10, pp. 4270–4281, 2014

  18. [21]

    A perceptual image quality index based on global and double-random window similarity,

    Z. Shi, K. Chen, K. Pang, J. Zhang, and Q. Cao, “A perceptual image quality index based on global and double-random window similarity,” Digital Signal Processing , vol. 60, pp. 277–286, 2017

  19. [22]

    Video quality assessment based on structural distortion measurement,

    Z. Wang, L. Lu, and A. C. Bovik, “Video quality assessment based on structural distortion measurement,” Signal processing: Image commu- nication, vol. 19, no. 2, pp. 121–132, 2004

  20. [23]

    A structural similarity metric for video based on motion models,

    K. Seshadrinathan and A. C. Bovik, “A structural similarity metric for video based on motion models,” in IEEE International Conference on Acoustics, Speech and Signal Processing , vol. 1, 2007, pp. I–869

  21. [24]

    An optical flow-based full reference video quality assessment algorithm,

    K. Manasa and S. S. Channappayya, “An optical flow-based full reference video quality assessment algorithm,” IEEE Transactions on Image Processing, vol. 25, no. 6, pp. 2480–2492, 2016

  22. [25]

    Image quality assess- ment: Unifying structure and texture similarity,

    K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assess- ment: Unifying structure and texture similarity,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 5, pp. 2567 – 2581, 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 29

  23. [26]

    Locally adaptive structure and texture similarity for image quality assessment,

    K. Ding, Y . Liu, X. Zou, S. Wang, and K. Ma, “Locally adaptive structure and texture similarity for image quality assessment,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 2483–2491

  24. [27]

    Deep learning based full- reference and no-reference quality assessment models for compressed UGC videos,

    W. Sun, T. Wang, X. Min, F. Yi, and G. Zhai, “Deep learning based full- reference and no-reference quality assessment models for compressed UGC videos,” in IEEE International Conference on Multimedia & Expo Workshops (ICMEW), 2021, pp. 1–6

  25. [28]

    End-to-end Optimized Image Compression,

    J. Ball ´e, V . Laparra, and E. P. Simoncelli, “End-to-end Optimized Image Compression,” 2017. [Online]. Available: https://arxiv.org/abs/ 1611.01704

  26. [29]

    Learning to compress videos without computing motion,

    M. Chen, T. Goodall, A. Patney, and A. C. Bovik, “Learning to compress videos without computing motion,” Signal Processing: Image Communication, vol. 103, p. 116633, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0923596522000029

  27. [30]

    Metamers of the ventral stream,

    J. Freeman and E. P. Simoncelli, “Metamers of the ventral stream,” Nature neuroscience, vol. 14, no. 9, pp. 1195–1201, 2011

  28. [31]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017

  29. [32]

    Design of linear equalizers optimized for the structural similarity index,

    S. S. Channappayya, A. C. Bovik, C. Caramanis, and R. W. Heath, “Design of linear equalizers optimized for the structural similarity index,” IEEE Transactions on Image Processing , vol. 17, no. 6, pp. 857–872, 2008

  30. [33]

    Mcvd - masked conditional video diffusion for prediction, generation, and interpolation,

    V . V oleti, A. Jolicoeur-Martineau, and C. Pal, “Mcvd - masked conditional video diffusion for prediction, generation, and interpolation,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. ...

  31. [34]

    High-fidelity generative image compression,

    F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson, “High-fidelity generative image compression,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 11 9...

  32. [35]

    Lossy image compression with conditional dif- fusion models,

    R. Yang and S. Mandt, “Lossy image compression with conditional dif- fusion models,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 64 971–64 995. [Onl...

  33. [36]

    The level weighted structural similarity loss: A step away from mse,

    Y . Lu, “The level weighted structural similarity loss: A step away from mse,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, pp. 9989–9990, Jul. 2019. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/5131

  34. [37]

    Deep generative adversarial compression artifact removal,

    L. Galteri, L. Seidenari, M. Bertini, and A. Del Bimbo, “Deep generative adversarial compression artifact removal,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , Oct 2017

  35. [38]

    Loss functions for pose guided person image generation,

    H. Shi, L. Wang, N. Zheng, G. Hua, and W. Tang, “Loss functions for pose guided person image generation,” Pattern Recognition, vol. 122, p. 108351, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320321005318

  36. [39]

    The statistics of natural images,

    D. L. Ruderman, “The statistics of natural images,” Netw.: Comput. Neural Syst., vol. 5, no. 4, pp. 517–548, 1994

  37. [40]

    Image information and visual quality,

    H. R. Sheikh and A. C. Bovik, “Image information and visual quality,” IEEE Transactions on Image Processing , vol. 15, no. 2, pp. 430–444, 2006

  38. [41]

    Distributions of the Two-Dimensional DCT Coefficients for Images,

    R. Reininger and J. Gibson, “Distributions of the Two-Dimensional DCT Coefficients for Images,” IEEE Transactions on Communications, vol. 31, no. 6, pp. 835–839, 1983

  39. [42]

    A theory for multiresolution signal decomposition: the wavelet representation,

    S. Mallat, “A theory for multiresolution signal decomposition: the wavelet representation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 11, no. 7, pp. 674–693, 1989

  40. [44]

    VMAF: The journey continues,

    Z. Li, C. Bampis, J. Novak, A. Aaron, K. Swanson, A. Moorthy, and J. Cock, “VMAF: The journey continues,” Netflix Technology Blog , vol. 25, no. 1, 2018

  41. [45]

    RAPIQUE: Rapid and accurate video quality prediction of user generated content,

    Z. Tu, X. Yu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “RAPIQUE: Rapid and accurate video quality prediction of user generated content,” IEEE Open Journal of Signal Processing , vol. 2, pp. 425–440, 2021

  42. [46]

    No-reference quality assessment of variable frame-rate videos using temporal band- pass statistics,

    Q. Zheng, Z. Tu, Y . Fan, X. Zeng, and A. C. Bovik, “No-reference quality assessment of variable frame-rate videos using temporal band- pass statistics,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 1795–1799

  43. [47]

    Making a “completely blind

    A. Mittal, R. Soundararajan, and A. C. Bovik, “Making a “completely blind” image quality analyzer,”IEEE Signal Processing Letters, vol. 20, no. 3, pp. 209–212, 2013

  44. [48]

    No-reference image quality assessment in the spatial domain,

    A. Mittal, A. K. Moorthy, and A. C. Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Trans. Image Process., vol. 21, no. 12, pp. 4695–4708, 2012

  45. [49]

    Blind image quality assessment using joint statistics of gradient magnitude and laplacian features,

    W. Xue, X. Mou, L. Zhang, A. C. Bovik, and X. Feng, “Blind image quality assessment using joint statistics of gradient magnitude and laplacian features,” IEEE Trans. Image Process. , vol. 23, no. 11, pp. 4850–4862, 2014

  46. [50]

    A completely blind video quality evaluator,

    Q. Zheng, Z. Tu, X. Zeng, A. C. Bovik, and Y . Fan, “A completely blind video quality evaluator,” IEEE Signal Processing Letters, vol. 29, pp. 2228–2232, 2022

  47. [51]

    No- reference quality assessment of tone-mapped HDR pictures,

    D. Kundu, D. Ghadiyaram, A. C. Bovik, and B. L. Evans, “No- reference quality assessment of tone-mapped HDR pictures,” IEEE Trans. Image Process., vol. 26, no. 6, pp. 2957–2971, 2017

  48. [52]

    Perceptual quality prediction on authentically dis- torted images using a bag of features approach,

    D. Ghadiyaram, “Perceptual quality prediction on authentically dis- torted images using a bag of features approach,” Journal of Vision , vol. 17(1), no. 32, pp. 1–25, 2017

  49. [53]

    Blind video quality assessment via space-time slice statistics,

    Q. Zheng, Z. Tu, Z. Hao, X. Zeng, A. C. Bovik, and Y . Fan, “Blind video quality assessment via space-time slice statistics,” in 2022 IEEE International Conference on Image Processing (ICIP) , 2022, pp. 451– 455

  50. [54]

    Blind prediction of natural video quality,

    M. A. Saad, A. C. Bovik, and C. Charrier, “Blind prediction of natural video quality,” IEEE Trans. Image Process. , vol. 23, no. 3, pp. 1352– 1365, 2014

  51. [55]

    Spatiotemporal statistics for video quality assessment,

    X. Li, Q. Guo, and X. Lu, “Spatiotemporal statistics for video quality assessment,” IEEE Trans. Image Process. , vol. 25, no. 7, pp. 3329– 3342, 2016

  52. [56]

    Hastie, R

    T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman, The Ele- ments of Statistical Learning: Data Mining, Inference, and Prediction . Springer, 2009, vol. 2

  53. [57]

    Deep video quality asses- sor: From spatio-temporal visual sensitivity to a convolutional neural aggregation network,

    W. Kim, J. Kim, S. Ahn, J. Kim, and S. Lee, “Deep video quality asses- sor: From spatio-temporal visual sensitivity to a convolutional neural aggregation network,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 219–234

  54. [58]

    C3dvqa: Full-reference video quality assessment with 3D convolutional neural network,

    M. Xu, J. Chen, H. Wang, S. Liu, G. Li, and Z. Bai, “C3dvqa: Full-reference video quality assessment with 3D convolutional neural network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 4447–4451

  55. [59]

    Deep VQA based on a novel hybrid training methodology,

    C. Feng, F. Zhang, and D. R. Bull, “Deep VQA based on a novel hybrid training methodology,” arXiv preprint arXiv:2202.08595, 2022

  56. [60]

    No-reference video quality evaluation by a deep transfer CNN architecture,

    R. Hou, Y . Zhao, Y . Hu, and H. Liu, “No-reference video quality evaluation by a deep transfer CNN architecture,” Signal Processing: Image Communication, vol. 83, p. 115782, 2020

  57. [61]

    A deep learning based no- reference quality assessment model for UGC videos,

    W. Sun, X. Min, W. Lu, and G. Zhai, “A deep learning based no- reference quality assessment model for UGC videos,” arXiv preprint arXiv:2204.14047, 2022

  58. [62]

    Deep neural networks for no-reference video quality assessment,

    J. You and J. Korhonen, “Deep neural networks for no-reference video quality assessment,” in IEEE International Conference on Image Processing (ICIP), 2019, pp. 2349–2353

  59. [63]

    No-reference video quality assessment with 3d shearlet transform and convolutional neural networks,

    Y . Li, L.-M. Po, C.-H. Cheung, X. Xu, L. Feng, F. Yuan, and K.- W. Cheung, “No-reference video quality assessment with 3d shearlet transform and convolutional neural networks,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 26, no. 6, pp. 1044– 1057, 2015

  60. [64]

    Blind video quality assessment with weakly supervised learning and resampling strategy,

    Y . Zhang, X. Gao, L. He, W. Lu, and R. He, “Blind video quality assessment with weakly supervised learning and resampling strategy,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 29, no. 8, pp. 2244–2255, 2018

  61. [66]

    A strong baseline for image and video quality assessment,

    S. Wen and J. Wang, “A strong baseline for image and video quality assessment,” arXiv preprint arXiv:2111.07104 , 2021

  62. [67]

    End-to-end blind quality assessment of compressed videos using deep neural networks,

    W. Liu, Z. Duanmu, and Z. Wang, “End-to-end blind quality assessment of compressed videos using deep neural networks,” ACM Multimedia, pp. 546–554, 2018. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 30

  63. [68]

    No-reference video quality assessment based on the tempo- ral pooling of deep features,

    D. Varga, “No-reference video quality assessment based on the tempo- ral pooling of deep features,” Neural Processing Letters, vol. 50, no. 3, pp. 2595–2608, 2019

  64. [69]

    Semantic information oriented no- reference video quality assessment,

    W. Wu, Q. Li, Z. Chen, and S. Liu, “Semantic information oriented no- reference video quality assessment,” IEEE Signal Processing Letters , vol. 28, pp. 204–208, 2021

  65. [70]

    No-reference video quality assessment using multi-pooled, saliency weighted deep features and decision fusion,

    D. Varga, “No-reference video quality assessment using multi-pooled, saliency weighted deep features and decision fusion,” Sensors, vol. 22, no. 6, p. 2209, 2022

  66. [71]

    Blind natural video quality prediction via statistical temporal features and deep spatial features,

    J. Korhonen, Y . Su, and J. You, “Blind natural video quality prediction via statistical temporal features and deep spatial features,” in Proceed- ings of the ACM International Conference on Multimedia , 2020, pp. 3311–3319

  67. [72]

    No-reference video quality assessment using multi-level spatially pooled features,

    F. G ¨otz-Hahn, V . Hosu, H. Lin, and D. Saupe, “No-reference video quality assessment using multi-level spatially pooled features,” arXiv preprint arXiv:1912.07966, 2019

  68. [73]

    No-reference video quality assessment via pretrained cnn and lstm networks,

    D. Varga and T. Szir ´anyi, “No-reference video quality assessment via pretrained cnn and lstm networks,”Signal, Image and Video Processing, vol. 13, no. 8, pp. 1569–1576, 2019

  69. [74]

    Quality assessment of in-the-wild videos,

    D. Li, T. Jiang, and M. Jiang, “Quality assessment of in-the-wild videos,” in Proceedings of the 27th ACM International Conference on Multimedia, 2019, pp. 2351–2359

  70. [75]

    Rirnet: Recurrent-in- recurrent network for video quality assessment,

    P. Chen, L. Li, L. Ma, J. Wu, and G. Shi, “Rirnet: Recurrent-in- recurrent network for video quality assessment,” in Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 834–842

  71. [76]

    Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,

    B. Li, W. Zhang, M. Tian, G. Zhai, and X. Wang, “Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 9, pp. 5944 – 5958, 2022

  72. [77]

    Study on no-reference video quality assessment method incorporating dual deep learning networks,

    J. Li and X. Li, “Study on no-reference video quality assessment method incorporating dual deep learning networks,” Multimedia Tools and Applications, pp. 1–20, 2022

  73. [78]

    Deep neural networks for end-to-end spatiotemporal video quality prediction and aggregation,

    J. Chen, H. Wang, M. Xu, G. Li, and S. Liu, “Deep neural networks for end-to-end spatiotemporal video quality prediction and aggregation,” in IEEE International Conference on Multimedia and Expo (ICME), 2021, pp. 1–6

  74. [80]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595

  75. [81]

    Why are deep repre- sentations good perceptual quality features?

    T. Tariq, O. T. Tursun, M. Kim, and P. Didyk, “Why are deep repre- sentations good perceptual quality features?” in European Conference on Computer Vision . Springer, 2020, pp. 445–461

  76. [82]

    FlOLPIPS: A bespoke video quality metric for frame interpoation,

    D. Danier, F. Zhang, and D. Bull, “FlOLPIPS: A bespoke video quality metric for frame interpoation,” arXiv preprint arXiv:2207.08119, 2022

  77. [83]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997

  78. [84]

    Gate-variants of gated recurrent unit (GRU) neural networks,

    R. Dey and F. M. Salem, “Gate-variants of gated recurrent unit (GRU) neural networks,” in IEEE 60th International Midwest Symposium on Circuits and Systems (MWSCAS) , 2017, pp. 1597–1600

  79. [85]

    2BiVQA: Double BI-LSTM based Video Quality Assessment of UGC Videos,

    A. Telili, S. A. Fezza, W. Hamidouche, and H. F. Meftah, “2BiVQA: Double BI-LSTM based Video Quality Assessment of UGC Videos,” arXiv preprint arXiv:2208.14774 , 2022

  80. [86]

    Cbam: Convolutional block attention module,

    S. Woo, J. Park, J.-Y . Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 3–19

  81. [87]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141

  82. [88]

    Attention based network for no-reference ugc video quality assessment,

    F. Yi, M. Chen, W. Sun, X. Min, Y . Tian, and G. Zhai, “Attention based network for no-reference ugc video quality assessment,” in IEEE International Conference on Image Processing (ICIP), 2021, pp. 1414– 1418

  83. [89]

    Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,

    B. Chen, L. Zhu, G. Li, F. Lu, H. Fan, and S. Wang, “Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1903–1916, 2022

  84. [90]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020

  85. [91]

    Swin transformer: Hierarchical vision transformer using shifted win- dows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted win- dows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022

  86. [92]

    Long short-term convolutional transformer for no-reference video quality assessment,

    J. You, “Long short-term convolutional transformer for no-reference video quality assessment,” in ACM International Conference on Mul- timedia, 2021, pp. 2112–2120

  87. [93]

    StarVQA: Space-time attention for video quality assessment,

    F. Xing, Y .-G. Wang, H. Wang, L. Li, and G. Zhu, “StarVQA: Space-time attention for video quality assessment,” arXiv preprint arXiv:2108.09635, 2021

  88. [94]

    DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment,

    H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, and W. Lin, “DisCoVQA: Temporal Distortion-Content Transformers for Video Quality Assessment,” arXiv preprint arXiv:2206.09853 , 2022

  89. [95]

    Fast-VQA: Efficient End-to-end Video Quality Assessment with Fragment Sampling,

    H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-VQA: Efficient End-to-end Video Quality Assessment with Fragment Sampling,” arXiv preprint arXiv:2207.02595 , 2022

  90. [96]

    SAMscore: A semantic structural similarity metric for image translation evaluation,

    Y . Li, M. Chen, W. Yang, K. Wang, J. Ma, A. C. Bovik, and Y . Zhang, “SAMscore: A semantic structural similarity metric for image translation evaluation,” arXiv preprint arXiv:2305.15367 , 2023

  91. [97]

    Sam-iqa: Can segment anything boost image quality assessment?

    X. Li, T. Jiang, H. Fan, and S. Liu, “Sam-iqa: Can segment anything boost image quality assessment?” arXiv preprint arXiv:2307.04455 , 2023

  92. [98]

    Cover: A comprehensive video quality evaluator,

    C. He, Q. Zheng, R. Zhu, X. Zeng, Y . Fan, and Z. Tu, “Cover: A comprehensive video quality evaluator,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2024, pp. 5799–5809

  93. [99]

    Ptm-vqa: Efficient video quality assessment leveraging diverse pretrained models from the wild,

    K. Yuan, H. Liu, M. Li, M. Sun, M. Sun, J. Gong, J. Hao, C. Zhou, and Y . Tang, “Ptm-vqa: Efficient video quality assessment leveraging diverse pretrained models from the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2...

  94. [100]

    Exploring clip for assessing the look and feel of images,

    J. Wang, K. C. Chan, and C. C. Loy, “Exploring clip for assessing the look and feel of images,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 2, pp. 2555–2563, Jun. 2023. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/ 25353

  95. [101]

    Exploring opinion-unaware video quality assessment with semantic affinity criterion,

    H. Wu, L. Liao, J. Hou, C. Chen, E. Zhang, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring opinion-unaware video quality assessment with semantic affinity criterion,” 2023. [Online]. Available: https://arxiv.org/abs/2302.13269

  96. [102]

    Towards explainable in-the-wild video quality assessment: a database and a language-prompted approach,

    H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Towards explainable in-the-wild video quality assessment: a database and a language-prompted approach,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 1045– 1054

  97. [103]

    Q-bench: A benchmark for general-purpose foundation models on low-level vision,

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai et al. , “Q-bench: A benchmark for general-purpose foundation models on low-level vision,” arXiv preprint arXiv:2309.14181, 2023

  98. [104]

    Lmm-vqa: Advancing video quality assessment with large multimodal models,

    Q. Ge, W. Sun, Y . Zhang, Y . Li, Z. Ji, F. Sun, S. Jui, X. Min, and G. Zhai, “Lmm-vqa: Advancing video quality assessment with large multimodal models,” arXiv preprint arXiv:2408.14008 , 2024

  99. [106]

    Video quality assessment on mobile devices: Subjective, behavioral and objective studies,

    A. K. Moorthy, L. K. Choi, A. C. Bovik, and G. De Veciana, “Video quality assessment on mobile devices: Subjective, behavioral and objective studies,” IEEE Journal of Selected Topics on Signal Processing, vol. 6, no. 6, pp. 652–671, 2012

  100. [108]

    Video quality assessment accounting for temporal visual masking of local flicker,

    L. K. Choi and A. C. Bovik, “Video quality assessment accounting for temporal visual masking of local flicker,” Signal Processing: image communication, vol. 67, pp. 182–198, 2018

  101. [109]

    A subjective and objective study of space-time subsampled video quality,

    D. Y . Lee, S. Paul, C. G. Bampis, H. Ko, J. Kim, S. Y . Jeong, B. Homan, and A. C. Bovik, “A subjective and objective study of space-time subsampled video quality,” IEEE Transactions on Image Processing , vol. 31, pp. 934–948, 2021

  102. [110]

    HDR or SDR? a subjective and objective study of scaled and compressed videos,

    J. P. Ebenezer, Z. Shang, Y . Chen, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “HDR or SDR? a subjective and objective study of scaled and compressed videos,” IEEE Transactions on Image Processing , 2024

  103. [111]

    Large-scale study of perceptual video quality,

    Z. Sinno and A. C. Bovik, “Large-scale study of perceptual video quality,” IEEE Transactions on Image Processing , vol. 28, no. 2, pp. 612–627, feb 2019. [Online]. Available: https://doi.org/10.1109% 2Ftip.2018.2869673 JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 31

  104. [112]

    From patches to pictures (PaQ-2-PiQ): Mapping the perceptual space of picture quality,

    Z. Ying, H. Niu, P. Gupta, D. Mahajan, D. Ghadiyaram, and A. Bovik, “From patches to pictures (PaQ-2-PiQ): Mapping the perceptual space of picture quality,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 3575–3585

  105. [113]

    Telepresence video quality assessment,

    Z. Ying, D. Ghadiyaram, and A. Bovik, “Telepresence video quality assessment,” European Conference on Computer Vision, Tel Aviv , pp. 327–347, Oct 2022

  106. [114]

    MCL-V: A streaming video quality assessment database,

    J. Y . Lin, R. Song, C.-H. Wu, T. Liu, H. Wang, and C.-C. J. Kuo, “MCL-V: A streaming video quality assessment database,” Journal of Visual Communication and Image Representation , vol. 30, pp. 1–9, 2015

  107. [115]

    CVD2014—a database for evaluating no-reference video quality assessment algorithms,

    M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen, and J. H ¨akkinen, “CVD2014—a database for evaluating no-reference video quality assessment algorithms,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 3073–3086, 2016

  108. [116]

    A study of high frame rate video formats,

    A. Mackin, F. Zhang, and D. R. Bull, “A study of high frame rate video formats,” IEEE Trans. Multimedia. , vol. 21, no. 6, pp. 1499– 1512, 2018

  109. [117]

    Subjective and objective quality assessment of high frame rate videos,

    P. C. Madhusudana, X. Yu, N. Birkbeck, Y . Wang, B. Adsumilli, and A. C. Bovik, “Subjective and objective quality assessment of high frame rate videos,” IEEE Access, vol. 9, pp. 108 069–108 082, 2021

  110. [118]

    Bvi-vfi: A video quality database for video frame interpolation,

    D. Danier, F. Zhang, and D. R. Bull, “Bvi-vfi: A video quality database for video frame interpolation,” IEEE Transactions on Image Processing, vol. 32, pp. 6004–6019, 2023

  111. [119]

    The konstanz natural video database (konvid-1k),

    V . Hosu, F. Hahn, M. Jenadeleh, H. Lin, H. Men, T. Szir ´anyi, S. Li, and D. Saupe, “The konstanz natural video database (konvid-1k),” in International Conference on Quality of Multimedia Experience (QoMEX), 2017, pp. 1–6

  112. [120]

    Perceptual quality assessment of internet videos,

    J. Xu, J. Li, X. Zhou, W. Zhou, B. Wang, and Z. Chen, “Perceptual quality assessment of internet videos,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 1248–1257

  113. [121]

    PUGCQ: A Large Scale Dataset for Quality Assessment of Professional User-Generated Content,

    G. Li, B. Chen, L. Zhu, Q. He, H. Fan, and S. Wang, “PUGCQ: A Large Scale Dataset for Quality Assessment of Professional User-Generated Content,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 3728–3736

  114. [122]

    Rich features for perceptual quality assess- ment of UGC videos,

    Y . Wang, J. Ke, H. Talebi, J. G. Yim, N. Birkbeck, B. Adsumilli, P. Milanfar, and F. Yang, “Rich features for perceptual quality assess- ment of UGC videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 13 435–13 444

  115. [123]

    Subjective and objective analysis of streamed gaming videos,

    X. Yu, Z. Ying, N. Birkbeck, Y . Wang, B. Adsumilli, and A. C. Bovik, “Subjective and objective analysis of streamed gaming videos,” IEEE Transactions on Games , vol. 16, no. 2, pp. 445 – 458, 2022

  116. [124]

    Towards explainable in-the-wild video quality assessment: A database and a language-prompted approach,

    H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Towards explainable in-the-wild video quality assessment: A database and a language-prompted approach,” in Proceedings of the 31st ACM International Conference on Multimedia, ser. MM ’23. New York...

  117. [125]

    Exploring video quality assessment on user generated con- tents from aesthetic and technical perspectives,

    ——, “Exploring video quality assessment on user generated con- tents from aesthetic and technical perspectives,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 20 144–20 154

  118. [126]

    Md-vqa: Multi-dimensional quality assessment for ugc live videos,

    Z. Zhang, W. Wu, W. Sun, D. Tu, W. Lu, X. Min, Y . Chen, and G. Zhai, “Md-vqa: Multi-dimensional quality assessment for ugc live videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 1746–1755

  119. [127]

    Kvq: Kwai video quality assessment for short-form videos,

    Y . Lu, X. Li, Y . Pei, K. Yuan, Q. Xie, Y . Qu, M. Sun, C. Zhou, and Z. Chen, “Kvq: Kwai video quality assessment for short-form videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 25 963–25 973

  120. [128]

    Methodology for the subjective assessment of the quality of television pictures,

    B. Series, “Methodology for the subjective assessment of the quality of television pictures,” Recommendation ITU-R BT , vol. 500, no. 13, 2012

  121. [129]

    Recover subjective quality scores from noisy measurements,

    Z. Li and C. G. Bampis, “Recover subjective quality scores from noisy measurements,” in Data compression conference (DCC), 2017, pp. 52– 61

  122. [130]

    Study of subjective and objective quality assessment of video,

    K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack, “Study of subjective and objective quality assessment of video,” IEEE Transactions on Image Processing, vol. 19, no. 6, pp. 1427–1441, 2010

  123. [131]

    A subjective and objective study of space-time subsampled video quality,

    D. Y . Lee, S. Paul, C. G. Bampis, H. Ko, J. Kim, S. Y . Jeong, B. Homan, and A. C. Bovik, “A subjective and objective study of space-time subsampled video quality,” IEEE Transactions on Image Processing , vol. 31, pp. 934–948, 2022

  124. [132]

    Comparing subjective video quality testing methodologies,

    M. H. Pinson and S. Wolf, “Comparing subjective video quality testing methodologies,” in Visual Communications and Image Processing 2003, T. Ebrahimi and T. Sikora, Eds., vol. 5150, International Society for Optics and Photonics. SPIE, 2003, pp. 573 – 582. [Online]. Available:...

  125. [133]

    Quality asessment of coded images using numerical category scaling,

    A. M. van Dijk, J.-B. Martens, and A. B. Watson, “Quality asessment of coded images using numerical category scaling,” in Advanced Image and Video Communications and Storage Technologies, N. Ohta, H. U. Lemke, and J. C. Lehureau, Eds., vol. 2451, International Society for Opti...

  126. [134]

    Study of the subjective and objective quality of high mo- tion live streaming videos,

    Z. Shang, J. P. Ebenezer, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Study of the subjective and objective quality of high mo- tion live streaming videos,” IEEE Transactions on Image Processing , vol. 31, pp. 1027–1041, 2021

  127. [135]

    A study of subjective and objec- tive quality assessment of hdr videos,

    Z. Shang, J. P. Ebenezer, A. K. Venkataramanan, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “A study of subjective and objec- tive quality assessment of hdr videos,” IEEE Transactions on Image Processing, vol. 33, pp. 42–57, 2024

  128. [136]

    Chipqa: No-reference video quality prediction via space-time chips,

    J. P. Ebenezer, Z. Shang, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Chipqa: No-reference video quality prediction via space-time chips,” IEEE Transactions on Image Processing , vol. 30, pp. 8059– 8074, 2021

  129. [139]

    Massive online crowdsourced study of subjective and objective picture quality,

    D. Ghadiyaram and A. C. Bovik, “Massive online crowdsourced study of subjective and objective picture quality,” IEEE Transactions on Image Processing, vol. 25, no. 1, pp. 372–387, 2015

  130. [140]

    Yfcc100m: The new data in multimedia research,

    B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,” Communications of the ACM , vol. 59, no. 2, pp. 64–73, 2016

  131. [141]

    Measuring the quality of text-to-video model outputs: Metrics and dataset,

    I. Chivileva, P. Lynch, T. E. Ward, and A. F. Smeaton, “Measuring the quality of text-to-video model outputs: Metrics and dataset,” 2023. [Online]. Available: https://arxiv.org/abs/2309.08009

  132. [142]

    Evalcrafter: Benchmarking and evaluating large video generation models,

    Y . Liu, X. Cun, X. Liu, X. Wang, Y . Zhang, H. Chen, Y . Liu, T. Zeng, R. Chan, and Y . Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2024, pp....

  133. [143]

    Fetv: A benchmark for fine-grained evaluation of open-domain text- to-video generation,

    Y . Liu, L. Li, S. Ren, R. Gao, S. Li, S. Chen, X. Sun, and L. Hou, “Fetv: A benchmark for fine-grained evaluation of open-domain text- to-video generation,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levi...

  134. [144]

    Vbench: Comprehensive benchmark suite for video generative models,

    Z. Huang, Y . He, J. Yu, F. Zhang, C. Si, Y . Jiang, Y . Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y . Wang, X. Chen, L. Wang, D. Lin, Y . Qiao, and Z. Liu, “Vbench: Comprehensive benchmark suite for video generative models,” in Proceedings of the IEEE/CVF Conference on Computer Vi...

  135. [145]

    Subjective-aligned dataset and metric for text- to-video quality assessment,

    T. Kou, X. Liu, Z. Zhang, C. Li, H. Wu, X. Min, G. Zhai, and N. Liu, “Subjective-aligned dataset and metric for text- to-video quality assessment,” 2024. [Online]. Available: https: //arxiv.org/abs/2403.11956

  136. [146]

    Gaia: Rethinking action quality assessment for ai-generated videos,

    Z. Chen, W. Sun, Y . Tian, J. Jia, Z. Zhang, J. Wang, R. Huang, X. Min, G. Zhai, and W. Zhang, “Gaia: Rethinking action quality assessment for ai-generated videos,” 2024. [Online]. Available: https://arxiv.org/abs/2406.06087

  137. [147]

    Benchmarking aigc video quality assessment: A dataset and unified model,

    Z. Zhang, X. Li, W. Sun, J. Jia, X. Min, Z. Zhang, C. Li, Z. Chen, P. Wang, Z. Ji, F. Sun, S. Jui, and G. Zhai, “Benchmarking aigc video quality assessment: A dataset and unified model,” 2024. [Online]. Available: https://arxiv.org/abs/2407.21408

  138. [148]

    Yfcc100m: the new data in multimedia research,

    B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: the new data in multimedia research,” Commun. ACM , vol. 59, no. 2, p. 64–73, jan

  139. [149]

    The kinetics human action video dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman, “The kinetics human action video dataset,” CoRR, vol. abs/1705.06950, 2017. [Online]. Available: http://arxiv.org/abs/1705.06950 ...

  140. [150]

    TaoLive

    Taobao Alibaba, Inc., “TaoLive.” [Online]. Available: https://taolive. taobao.com

  141. [151]

    Available: https: //www.kwai.com/foryou

    Kuaishou Technology, Inc., “Kwai.” [Online]. Available: https: //www.kwai.com/foryou

  142. [152]

    Imagereward: Learning and evaluating human preferences for text-to-image generation,

    J. Xu, X. Liu, Y . Wu, Y . Tong, Q. Li, M. Ding, J. Tang, and Y . Dong, “Imagereward: Learning and evaluating human preferences for text-to-image generation,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Le...

  143. [153]

    Pick-a-pic: An open dataset of user preferences for text-to-image generation,

    Y . Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy, “Pick-a-pic: An open dataset of user preferences for text-to-image generation,” in Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Ed...

  144. [154]

    A perceptual quality assessment exploration for aigc images,

    Z. Zhang, C. Li, W. Sun, X. Liu, X. Min, and G. Zhai, “A perceptual quality assessment exploration for aigc images,” in 2023 IEEE Inter- national Conference on Multimedia and Expo Workshops (ICMEW) , 2023, pp. 440–445

  145. [155]

    Aigciqa2023: A large-scale image quality assessment database for ai generated images: From the perspectives of quality, authenticity and correspon- dence,

    J. Wang, H. Duan, J. Liu, S. Chen, X. Min, and G. Zhai, “Aigciqa2023: A large-scale image quality assessment database for ai generated images: From the perspectives of quality, authenticity and correspon- dence,” in Artificial Intelligence, L. Fang, J. Pei, G. Zhai, and R. Wan...

  146. [156]

    Exploring the naturalness of ai-generated images,

    Z. Chen, W. Sun, H. Wu, Z. Zhang, J. Jia, Z. Ji, F. Sun, S. Jui, X. Min, G. Zhai, and W. Zhang, “Exploring the naturalness of ai-generated images,” 2024. [Online]. Available: https://arxiv.org/abs/2312.05476

  147. [157]

    Agiqa-3k: An open database for ai-generated image quality assessment,

    C. Li, Z. Zhang, H. Wu, W. Sun, X. Min, X. Liu, G. Zhai, and W. Lin, “Agiqa-3k: An open database for ai-generated image quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 8, pp. 6833–6846, 2024

  148. [158]

    AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment,

    C. Li, T. Kou, Y . Gao, Y . Cao, W. Sun, Z. Zhang, Y . Zhou, Z. Zhang, H. Wu, W. Zhang, X. Liu, X. Min, and Z. Guangtao, “AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  149. [159]

    Cmc-bench: Towards a new paradigm of visual signal compression,

    C. Li, X. Wu, H. Wu, D. Feng, Z. Zhang, G. Lu, X. Min, X. Liu, G. Zhai, and W. Lin, “Cmc-bench: Towards a new paradigm of visual signal compression,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09356

  150. [160]

    GPT-4 technical report,

    OpenAI, “GPT-4 technical report,” https://openai.com/research/gpt-4, 2023

  151. [161]

    A frame rate dependent video quality metric based on temporal wavelet decomposition and spatiotemporal pooling,

    F. Zhang, A. Mackin, and D. R. Bull, “A frame rate dependent video quality metric based on temporal wavelet decomposition and spatiotemporal pooling,” in Proc. IEEE Int. Conf. Image Process., 2017, pp. 300–304

  152. [162]

    A comparative evaluation of temporal pooling methods for blind video quality assessment,

    Z. Tu, C.-J. Chen, L.-H. Chen, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “A comparative evaluation of temporal pooling methods for blind video quality assessment,” in IEEE International Conference on Image Processing (ICIP) , 2020, pp. 141–145

  153. [163]

    A statistical evaluation of recent full reference image quality assess- ment algorithms,

    “A statistical evaluation of recent full reference image quality assess- ment algorithms,” IEEE Transactions on Image Processing , vol. 15, no. 11, pp. 3440–3451, 2006

  154. [165]

    Video quality assessment using a statistical model of human visual speed perception,

    Z. Wang and Q. Li, “Video quality assessment using a statistical model of human visual speed perception,” JOSA A, vol. 24, no. 12, pp. B61– B69, 2007

  155. [168]

    An information fidelity criterion for image quality assessment using natural scene statistics,

    H. R. Sheikh, A. C. Bovik, and G. De Veciana, “An information fidelity criterion for image quality assessment using natural scene statistics,” IEEE Transactions on Image Processing , vol. 14, no. 12, pp. 2117– 2128, 2005

  156. [169]

    Most apparent distortion: full- reference image quality assessment and the role of strategy,

    E. C. Larson and D. M. Chandler, “Most apparent distortion: full- reference image quality assessment and the role of strategy,” Journal of Electronic Imaging , vol. 19, no. 1, p. 011006, 2010

  157. [170]

    Video Quality Assessment by Reduced Reference Spatio-Temporal Entropic Differencing,

    R. Soundararajan and A. C. Bovik, “Video Quality Assessment by Reduced Reference Spatio-Temporal Entropic Differencing,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 23, no. 4, pp. 684–694, 2013

  158. [174]

    Image Quality Assessment by Separately Evaluating Detail Losses and Additive Impairments,

    S. Li, F. Zhang, L. Ma, and K. N. Ngan, “Image Quality Assessment by Separately Evaluating Detail Losses and Additive Impairments,” IEEE Transactions on Multimedia , vol. 13, no. 5, pp. 935–949, 2011

  159. [175]

    Spatiotemporal feature inte- gration and model fusion for full reference video quality assessment,

    C. G. Bampis, Z. Li, and A. C. Bovik, “Spatiotemporal feature inte- gration and model fusion for full reference video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 29, no. 8, pp. 2256–2270, 2018

  160. [176]

    Funque: Fusion of unified quality evaluators,

    A. K. Venkataramanan, C. Stejerean, and A. C. Bovik, “Funque: Fusion of unified quality evaluators,” in 2022 IEEE International Conference on Image Processing (ICIP) , 2022, pp. 2147–2151

  161. [177]

    One transform to compute them all: Efficient fusion-based full-reference video quality assessment,

    A. K. Venkataramanan, C. Stejerean, I. Katsavounidis, and A. C. Bovik, “One transform to compute them all: Efficient fusion-based full-reference video quality assessment,” IEEE Transactions on Image Processing, vol. 33, pp. 509–524, 2024

  162. [178]

    Topiq: A top-down approach from semantics to distortions for image quality assessment,

    C. Chen, J. Mo, J. Hou, H. Wu, L. Liao, W. Sun, Q. Yan, and W. Lin, “Topiq: A top-down approach from semantics to distortions for image quality assessment,” IEEE Transactions on Image Processing , vol. 33, pp. 2404–2418, 2024

  163. [179]

    Motion tuned spatio-temporal quality assessment of natural videos,

    K. Seshadrinathan and A. C. Bovik, “Motion tuned spatio-temporal quality assessment of natural videos,” IEEE Transactions on Image Processing, vol. 19, no. 2, pp. 335–350, 2009

  164. [180]

    A spatiotemporal most- apparent-distortion model for video quality assessment,

    P. V . Vu, C. T. Vu, and D. M. Chandler, “A spatiotemporal most- apparent-distortion model for video quality assessment,” in IEEE International Conference on Image Processing , 2011, pp. 2505–2508

  165. [181]

    Attention driven foveated video quality assessment,

    J. You, T. Ebrahimi, and A. Perkis, “Attention driven foveated video quality assessment,” IEEE Transactions on Image Processing , vol. 23, no. 1, pp. 200–213, 2013

  166. [182]

    Temporal video quality model accounting for variable frame delay distortions,

    M. H. Pinson, L. K. Choi, and A. C. Bovik, “Temporal video quality model accounting for variable frame delay distortions,” IEEE Transac- tions on Broadcasting , vol. 60, no. 4, pp. 637–649, 2014

  167. [183]

    A perception-based hybrid model for video quality assessment,

    F. Zhang and D. R. Bull, “A perception-based hybrid model for video quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 26, no. 6, pp. 1017–1028, 2015

  168. [184]

    Deep neural networks for no-reference and full-reference image quality assessment,

    S. Bosse, D. Maniry, K.-R. M ¨uller, T. Wiegand, and W. Samek, “Deep neural networks for no-reference and full-reference image quality assessment,” IEEE Transactions on Image Processing , vol. 27, no. 1, pp. 206–219, 2017

  169. [185]

    Deep learning of human visual sensitivity in image quality assessment framework,

    J. Kim and S. Lee, “Deep learning of human visual sensitivity in image quality assessment framework,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 1676–1684

  170. [186]

    Deep learning-based distortion sensitivity prediction for full-reference image quality assessment,

    S. Ahn, Y . Choi, and K. Yoon, “Deep learning-based distortion sensitivity prediction for full-reference image quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 344–353

  171. [187]

    U-Net: Convolutional net- works for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional net- works for biomedical image segmentation,” inInternational Conference on Medical Image Computing and Computer-assisted Intervention . Springer, 2015, pp. 234–241

  172. [188]

    Squeezenet: Alexnet-level accuracy with 50x fewer parameters and less than 0.5 mb model size,

    F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and less than 0.5 mb model size,” arXiv preprint arXiv:1602.07360, 2016

  173. [189]

    One weird trick for parallelizing convolutional neural networks,

    A. Krizhevsky, “One weird trick for parallelizing convolutional neural networks,” arXiv preprint arXiv:1404.5997 , 2014

  174. [190]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556 , 2014

  175. [191]

    E-LPIPS: robust per- ceptual image similarity via random transformation ensembles,

    M. Kettunen, E. H ¨ark¨onen, and J. Lehtinen, “E-LPIPS: robust per- ceptual image similarity via random transformation ensembles,” arXiv preprint arXiv:1906.03973, 2019

  176. [192]

    Deepwsd: Projecting degradations in perceptual space to wasserstein distance in deep feature space,

    X. Liao, B. Chen, H. Zhu, S. Wang, M. Zhou, and S. Kwong, “Deepwsd: Projecting degradations in perceptual space to wasserstein distance in deep feature space,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 970–978. JOURNAL OF LATEX CLASS FIL...

  177. [193]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778

  178. [194]

    Video quality assessment for spatio-temporal resolution adaptive coding,

    H. Zhu, B. Chen, L. Zhu, P. Chen, L. Song, and S. Wang, “Video quality assessment for spatio-temporal resolution adaptive coding,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 7, pp. 6403–6415, 2024

  179. [195]

    A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,

    N. D. Narvekar and L. J. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in 2009 International Workshop on Quality of Multimedia Experience , 2009, pp. 87–91

  180. [196]

    Image sharpness assess- ment based on local phase coherence,

    R. Hassen, Z. Wang, and M. M. A. Salama, “Image sharpness assess- ment based on local phase coherence,” IEEE Trans. Image Process. , vol. 22, no. 7, pp. 2798–2810, 2013

  181. [197]

    No-reference perceptual quality assessment of JPEG compressed images,

    Z. Wang, H. R. Sheikh, and A. C. Bovik, “No-reference perceptual quality assessment of JPEG compressed images,” in Proc. IEEE Int. Conf. Image Process., vol. 1, 2002, pp. I–I

  182. [198]

    Blind measurement of blocking artifacts in images,

    Z. Wang, A. C. Bovik, and B. L. Evan, “Blind measurement of blocking artifacts in images,” in Proc. IEEE Int. Conf. Image Process. , vol. 3, 2000, pp. 981–984

  183. [199]

    No-reference quality as- sessment of jpeg images via a quality relevance map,

    S. A. Golestaneh and D. M. Chandler, “No-reference quality as- sessment of jpeg images via a quality relevance map,” IEEE Signal Processing Letters, vol. 21, no. 2, pp. 155–158, 2014

  184. [200]

    Two-level approach for no-reference consumer video quality assessment,

    J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,” IEEE Trans. Image Process. , vol. 28, no. 12, pp. 5923–5938, 2019

  185. [201]

    Unsupervised feature learning framework for no-reference image quality assessment,

    P. Ye, J. Kumar, L. Kang, and D. Doermann, “Unsupervised feature learning framework for no-reference image quality assessment,” in IEEE Conference on Computer Vision and Pattern Recognition , 2012, pp. 1098–1105

  186. [202]

    Blind image quality assessment based on high order statistics aggregation,

    J. Xu, P. Ye, Q. Li, H. Du, Y . Liu, and D. Doermann, “Blind image quality assessment based on high order statistics aggregation,” IEEE Transactions on Image Processing, vol. 25, no. 9, pp. 4444–4457, 2016

  187. [203]

    A completely blind video integrity oracle,

    A. Mittal, M. A. Saad, and A. C. Bovik, “A completely blind video integrity oracle,” IEEE Trans. Image Process., vol. 25, no. 1, pp. 289– 300, 2015

  188. [204]

    UGC- VQA: Benchmarking blind video quality assessment for user generated content,

    Z. Tu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “UGC- VQA: Benchmarking blind video quality assessment for user generated content,” IEEE Trans. Image Process. , vol. 30, pp. 4449–4464, 2021

  189. [205]

    FA VER: Blind Quality Prediction of Variable Frame Rate Videos,

    Q. Zheng, Z. Tu, P. C. Madhusudana, X. Zeng, A. C. Bovik, and Y . Fan, “FA VER: Blind Quality Prediction of Variable Frame Rate Videos,” Signal Processing: Image Communication , vol. 122, 2024

  190. [206]

    A feature-enriched completely blind image quality evaluator,

    L. Zhang, L. Zhang, and A. C. Bovik, “A feature-enriched completely blind image quality evaluator,” IEEE Trans. Image Process. , vol. 24, no. 8, pp. 2579–2591, 2015

  191. [207]

    A no-reference video quality predictor for compression and scaling artifacts,

    D. Ghadiyaram, C. Chen, S. Inguva, and A. Kokaram, “A no-reference video quality predictor for compression and scaling artifacts,” in IEEE International Conference on Image Processing , 2017, pp. 3445–3449

  192. [208]

    Completely blind quality assessment of user generated video content,

    P. Kancharla and S. S. Channappayya, “Completely blind quality assessment of user generated video content,” IEEE Trans. Image Process., vol. 31, pp. 263–274, 2021

  193. [209]

    Blind image quality assessment by natural scene statistics and perceptual characteristics,

    Y . Liu, K. Gu, X. Li, and Y . Zhang, “Blind image quality assessment by natural scene statistics and perceptual characteristics,” ACM Trans- actions on Multimedia Computing, Communications, and Applications (TOMM), vol. 16, no. 3, pp. 1–91, 2020

  194. [210]

    Unsu- pervised blind image quality evaluation via statistical measurements of structure, naturalness, and perception,

    Y . Liu, K. Gu, Y . Zhang, X. Li, G. Zhai, D. Zhao, and W. Gao, “Unsu- pervised blind image quality evaluation via statistical measurements of structure, naturalness, and perception,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 4, pp. 929–943, 2019

  195. [211]

    Perceptual straighten- ing of natural videos,

    O. J. H ´enaff, R. L. Goris, and E. P. Simoncelli, “Perceptual straighten- ing of natural videos,” Nature Neuroscience, vol. 22, no. 6, pp. 984– 991, 2019

  196. [213]

    Blind quality assessment based on pseudo-reference image,

    X. Min, K. Gu, G. Zhai, J. Liu, X. Yang, and C. W. Chen, “Blind quality assessment based on pseudo-reference image,” IEEE Transactions on Multimedia, vol. 20, no. 8, pp. 2049–2062, 2018

  197. [214]

    Convolutional neural networks for no-reference image quality assessment,

    L. Kang, P. Ye, Y . Li, and D. Doermann, “Convolutional neural networks for no-reference image quality assessment,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2014, pp. 1733–1740

  198. [215]

    A deep neural net- work for image quality assessment,

    S. Bosse, D. Maniry, T. Wiegand, and W. Samek, “A deep neural net- work for image quality assessment,” in IEEE International Conference on Image Processing (ICIP) , 2016, pp. 3773–3777

  199. [216]

    End-to- end blind image quality assessment using deep neural networks,

    K. Ma, W. Liu, K. Zhang, Z. Duanmu, Z. Wang, and W. Zuo, “End-to- end blind image quality assessment using deep neural networks,” IEEE Transactions on Image Processing, vol. 27, no. 3, pp. 1202–1213, 2017

  200. [217]

    Nima: Neural image assessment,

    H. Talebi and P. Milanfar, “Nima: Neural image assessment,” IEEE Transactions on Image Processing, vol. 27, no. 8, pp. 3998–4011, 2018

  201. [218]

    A probabilistic quality represen- tation approach to deep blind image quality prediction,

    H. Zeng, L. Zhang, and A. C. Bovik, “A probabilistic quality represen- tation approach to deep blind image quality prediction,” arXiv preprint arXiv:1708.08190, 2017

  202. [219]

    Blind image quality assessment using a deep bilinear convolutional neural network,

    W. Zhang, K. Ma, J. Yan, D. Deng, and Z. Wang, “Blind image quality assessment using a deep bilinear convolutional neural network,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 1, pp. 36–47, 2018

  203. [220]

    Fast R-CNN,

    R. Girshick, “Fast R-CNN,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1440–1448

  204. [221]

    Blind quality assessment for in-the-wild images via hierarchical feature fusion and iterative mixed database training,

    W. Sun, X. Min, D. Tu, S. Ma, and G. Zhai, “Blind quality assessment for in-the-wild images via hierarchical feature fusion and iterative mixed database training,” IEEE Journal of Selected Topics in Signal Processing, vol. 17, no. 6, pp. 1178–1192, 2023

  205. [222]

    Graphiqa: Learning distortion graph representations for blind image quality assessment,

    S. Sun, T. Yu, J. Xu, W. Zhou, and Z. Chen, “Graphiqa: Learning distortion graph representations for blind image quality assessment,” IEEE Transactions on Multimedia , vol. 25, pp. 2912–2925, 2023

  206. [223]

    Image quality assessment: From mean opinion score to opinion score distribution,

    Y . Gao, X. Min, Y . Zhu, J. Li, X.-P. Zhang, and G. Zhai, “Image quality assessment: From mean opinion score to opinion score distribution,” in Proceedings of the 30th ACM International Conference on Multimedia , ser. MM ’22. New York, NY , USA: Association for Computing Mach...

  207. [224]

    Hallucinated-iqa: No-reference image quality assessment via adversarial learning,

    K.-Y . Lin and G. Wang, “Hallucinated-iqa: No-reference image quality assessment via adversarial learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018

  208. [225]

    No-reference image quality assessment by hallucinating pristine features,

    B. Chen, L. Zhu, C. Kong, H. Zhu, S. Wang, and Z. Li, “No-reference image quality assessment by hallucinating pristine features,” IEEE Transactions on Image Processing , vol. 31, pp. 6139–6151, 2022

  209. [226]

    Blind image quality assessment based on geometric order learning,

    N.-H. Shin, S.-H. Lee, and C.-S. Kim, “Blind image quality assessment based on geometric order learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 12 799–12 808

  210. [227]

    Fully deep blind image quality predictor,

    J. Kim and S. Lee, “Fully deep blind image quality predictor,” IEEE Journal of Selected Topics on Signal Processing , vol. 11, no. 1, pp. 206–220, 2016

  211. [228]

    Video swin transformer,

    Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 3202–3211

  212. [229]

    Maxvit: Multi-axis vision transformer,

    Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y . Li, “Maxvit: Multi-axis vision transformer,” in European conference on computer vision. Springer, 2022, pp. 459–479

  213. [230]

    Cswin transformer: A general vision transformer backbone with cross-shaped windows,

    X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, “Cswin transformer: A general vision transformer backbone with cross-shaped windows,” in Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition , 2022, pp. 12 124– 12 134

  214. [231]

    MUSIQ: Multi- scale image quality transformer,

    J. Ke, Q. Wang, Y . Wang, P. Milanfar, and F. Yang, “MUSIQ: Multi- scale image quality transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5148–5157

  215. [232]

    Transformer for image quality assessment,

    J. You and J. Korhonen, “Transformer for image quality assessment,” in IEEE International Conference on Image Processing (ICIP) , 2021, pp. 1389–1393

  216. [233]

    Data-efficient image quality assessment with attention-panel decoder,

    G. Qin, R. Hu, Y . Liu, X. Zheng, H. Liu, X. Li, and Y . Zhang, “Data-efficient image quality assessment with attention-panel decoder,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 2, pp. 2091–2100, Jun. 2023

  217. [234]

    No-reference image quality assessment via transformers, relative ranking, and self- consistency,

    S. A. Golestaneh, S. Dadsetan, and K. M. Kitani, “No-reference image quality assessment via transformers, relative ranking, and self- consistency,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , January 2022, pp. 1220– 1230

  218. [235]

    Maniqa: Multi-dimension attention network for no-reference image quality assessment,

    S. Yang, T. Wu, S. Shi, S. Lao, Y . Gong, M. Cao, J. Wang, and Y . Yang, “Maniqa: Multi-dimension attention network for no-reference image quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , June 2022, pp...

  219. [236]

    Continual learning for blind image quality assessment,

    W. Zhang, D. Li, C. Ma, G. Zhai, X. Yang, and K. Ma, “Continual learning for blind image quality assessment,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022

  220. [237]

    Metaiqa: Deep meta- learning for no-reference image quality assessment,

    H. Zhu, L. Li, J. Wu, W. Dong, and G. Shi, “Metaiqa: Deep meta- learning for no-reference image quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2020, pp. 14 143–14 152

  221. [238]

    From distortion manifold to perceptual quality: a data efficient blind image quality assessment approach,

    S. Su, Q. Yan, Y . Zhu, J. Sun, and Y . Zhang, “From distortion manifold to perceptual quality: a data efficient blind image quality assessment approach,” Pattern Recognition, vol. 133, p. 109047, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0...

  222. [239]

    Forgetting to remember: A scalable incremental learning framework for cross-task blind image quality assessment,

    R. Ma, Q. Wu, K. N. Ngan, H. Li, F. Meng, and L. Xu, “Forgetting to remember: A scalable incremental learning framework for cross-task blind image quality assessment,” IEEE Transactions on Multimedia , vol. 25, pp. 8817–8827, 2023

  223. [240]

    Continual learn- ing of blind image quality assessment with channel modulation kernel,

    H. Li, L. Liao, C. Chen, X. Fan, W. Zuo, and W. Lin, “Continual learn- ing of blind image quality assessment with channel modulation kernel,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2024

  224. [241]

    Deep blind image quality assessment powered by online hard example mining,

    Z. Wang, Q. Jiang, S. Zhao, W. Feng, and W. Lin, “Deep blind image quality assessment powered by online hard example mining,” IEEE Transactions on Multimedia , vol. 25, pp. 4774–4784, 2023

  225. [242]

    Task-specific normalization for continual learning of blind image quality models,

    W. Zhang, K. Ma, G. Zhai, and X. Yang, “Task-specific normalization for continual learning of blind image quality models,” IEEE Transac- tions on Image Processing , vol. 33, pp. 1898–1910, 2024

  226. [243]

    Image quality assessment using contrastive learning,

    P. C. Madhusudana, N. Birkbeck, Y . Wang, B. Adsumilli, and A. C. Bovik, “Image quality assessment using contrastive learning,” IEEE Transactions on Image Processing , vol. 31, pp. 4149–4161, 2022

  227. [244]

    Opinion unaware image quality assessment via adversarial convolutional variational autoencoder,

    A. Shukla, A. Upadhyay, S. Bhugra, and M. Sharma, “Opinion unaware image quality assessment via adversarial convolutional variational autoencoder,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , January 2024, pp. 2153– 2163

  228. [245]

    Re-iqa: Unsupervised learning for image quality assessment in the wild,

    A. Saha, S. Mishra, and A. C. Bovik, “Re-iqa: Unsupervised learning for image quality assessment in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 5846–5855

  229. [246]

    No reference opinion unaware quality assessment of authentically distorted images,

    N. C. Babu, V . Kannan, and R. Soundararajan, “No reference opinion unaware quality assessment of authentically distorted images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2023, pp. 2459–2468

  230. [247]

    Arniqa: Learning distortion manifold for image quality assessment,

    L. Agnolucci, L. Galteri, M. Bertini, and A. Del Bimbo, “Arniqa: Learning distortion manifold for image quality assessment,” in Pro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024, pp. 189–198

  231. [248]

    Quality-aware pre- trained models for blind image quality assessment,

    K. Zhao, K. Yuan, M. Sun, M. Li, and X. Wen, “Quality-aware pre- trained models for blind image quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 22 302–22 313

  232. [249]

    Blind image quality assessment via vision-language correspondence: A multitask learning perspective,

    W. Zhang, G. Zhai, Y . Wei, X. Yang, and K. Ma, “Blind image quality assessment via vision-language correspondence: A multitask learning perspective,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 14 071–14 081

  233. [250]

    Towards transparent deep image aesthetics assessment with tag-based content descriptors,

    J. Hou, W. Lin, Y . Fang, H. Wu, C. Chen, L. Liao, and W. Liu, “Towards transparent deep image aesthetics assessment with tag-based content descriptors,” IEEE Transactions on Image Processing, pp. 1–1, 2023

  234. [251]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,”

  235. [252]

    Minigpt- 4: Enhancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt- 4: Enhancing vision-language understanding with advanced large language models,” 2023. [Online]. Available: https://arxiv.org/abs/ 2304.10592

  236. [253]

    Instructblip: Towards general-purpose vision-language models with instruction tuning,

    W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023. [Online]. Available: https://arxiv.org/abs/2305.06500

  237. [254]

    Otter: A multi-modal model with in-context instruction tuning,

    B. Li, Y . Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu, “Otter: A multi-modal model with in-context instruction tuning,” 2023. [Online]. Available: https://arxiv.org/abs/2305.03726

  238. [255]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels,

    H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y . Gao, A. Wang, E. Zhang, W. Sun et al., “Q-align: Teaching lmms for visual scoring via discrete text-defined levels,” arXiv preprint arXiv:2312.17090 , 2023

  239. [256]

    Towards open-ended visual quality comparison,

    H. Wu, H. Zhu, Z. Zhang, E. Zhang, C. Chen, L. Liao, C. Li, A. Wang, W. Sun, Q. Yanet al., “Towards open-ended visual quality comparison,” arXiv preprint arXiv:2402.16641 , 2024

  240. [257]

    A no-reference au- toencoder video quality metric,

    H. B. Martinez, M. C. Farias, and A. Hines, “A no-reference au- toencoder video quality metric,” in IEEE International Conference on Image Processing (ICIP) , 2019, pp. 1755–1759

  241. [258]

    No- reference vmaf: A deep neural network-based approach to blind video quality assessment,

    A. De Decker, J. De Cock, P. Lambert, and G. Van Wallendael, “No- reference vmaf: A deep neural network-based approach to blind video quality assessment,” IEEE Transactions on Broadcasting , pp. 1–0, 2024

  242. [259]

    Video quality assessment for online processing: From spatial to temporal sampling,

    J. Yan, L. Wu, Y . Fang, X. Liu, X. Xia, and W. Liu, “Video quality assessment for online processing: From spatial to temporal sampling,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2024

  243. [260]

    Unified quality assessment of in-the- wild videos with mixed datasets training,

    D. Li, T. Jiang, and M. Jiang, “Unified quality assessment of in-the- wild videos with mixed datasets training,” International Journal of Computer Vision, vol. 129, no. 4, pp. 1238–1257, 2021

  244. [261]

    Unsupervised curriculum domain adaptation for no-reference video quality assessment,

    P. Chen, L. Li, J. Wu, W. Dong, and G. Shi, “Unsupervised curriculum domain adaptation for no-reference video quality assessment,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 5178–5187

  245. [262]

    A blind video quality assessment method via spatiotemporal pyramid attention,

    W. Shen, M. Zhou, X. Wei, H. Wang, B. Fang, C. Ji, X. Zhuang, J. Wang, J. Luo, H. Pu, X. Huang, S. Wang, H. Cao, Y . Feng, T. Xiang, and Z. Shang, “A blind video quality assessment method via spatiotemporal pyramid attention,”IEEE Transactions on Broadcasting, vol. 70, no. 1, ...

  246. [263]

    Neighbourhood representative sampling for efficient end-to- end video quality assessment,

    H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, J. Gu, and W. Lin, “Neighbourhood representative sampling for efficient end-to- end video quality assessment,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 12, pp. 15 185–15 202, 2023

  247. [264]

    Scaling and masking: A new paradigm of data sampling for image and video quality assessment,

    Y . Liu, Y . Quan, G. Xiao, A. Li, and J. Wu, “Scaling and masking: A new paradigm of data sampling for image and video quality assessment,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, pp. 3792–3801, Mar. 2024. [Online]. Available: https://oj...

  248. [265]

    Knowledge guided semi-supervised learning for quality assessment of user generated videos,

    S. Mitra and R. Soundararajan, “Knowledge guided semi-supervised learning for quality assessment of user generated videos,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, pp. 4251–4260, Mar. 2024

  249. [266]

    Modular blind video quality assessment,

    W. Wen, M. Li, Y . Zhang, Y . Liao, J. Li, L. Zhang, and K. Ma, “Modular blind video quality assessment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 2763–2772

  250. [267]

    Ze-fesg: A zero-shot feature extrac- tion method based on semantic guidance for no-reference video quality assessment,

    Y . Mi, Y . Li, Y . Shu, and S. Liu, “Ze-fesg: A zero-shot feature extrac- tion method based on semantic guidance for no-reference video quality assessment,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 3640– 3644

  251. [268]

    Rethink- ing the inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethink- ing the inception architecture for computer vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2818–2826

  252. [269]

    Inception-v4, Inception-Resnet and the impact of residual connections on learning,

    C. Szegedy, S. Ioffe, V . Vanhoucke, and A. A. Alemi, “Inception-v4, Inception-Resnet and the impact of residual connections on learning,” in Thirty-first AAAI Conference on Artificial Intelligence , 2017

  253. [270]

    Spatial pyramid pooling in deep convolutional networks for visual recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 37, no. 9, pp. 1904– 1916, 2015

  254. [271]

    Inceptiontime: Finding alexnet for time series classification,

    H. Ismail Fawaz, B. Lucas, G. Forestier, C. Pelletier, D. F. Schmidt, J. Weber, G. I. Webb, L. Idoumghar, P.-A. Muller, and F. Petitjean, “Inceptiontime: Finding alexnet for time series classification,” Data Mining and Knowledge Discovery, vol. 34, no. 6, pp. 1936–1962, 2020

  255. [272]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2019

  256. [273]

    Learning spatiotemporal features with 3d convolutional networks,

    D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceed- ings of the IEEE International Conference on Computer Vision (ICCV), December 2015

  257. [274]

    Video swin transformer,

    Z. Liu, J. Ning, Y . Cao, Y . Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 3202–3211

  258. [275]

    A convnet for the 2020s,

    Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 11 966–11 976

  259. [276]

    Blind image quality assessment with a probabilistic quality representation,

    H. Zeng, L. Zhang, and A. C. Bovik, “Blind image quality assessment with a probabilistic quality representation,” in IEEE International Conference on Image Processing (ICIP) , 2018, pp. 609–613. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 35

  260. [277]

    Regression or classification? new methods to evaluate no-reference picture and video quality models,

    Z. Tu, C.-J. Chen, L.-H. Chen, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Regression or classification? new methods to evaluate no-reference picture and video quality models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 20...

  261. [278]

    Fast differen- tiable sorting and ranking,

    M. Blondel, O. Teboul, Q. Berthet, and J. Djolonga, “Fast differen- tiable sorting and ranking,” in International Conference on Machine Learning, 2020, pp. 950–959

  262. [279]

    Final report from the video quality experts group on the validation of objective models of video quality assessment,

    V . Q. E. Group et al., “Final report from the video quality experts group on the validation of objective models of video quality assessment,” in VQEG meeting, Ottawa, Canada, March, 2000 , 2000

  263. [280]

    Squared earth mover’s distance- based loss for training deep neural networks,

    L. Hou, C.-P. Yu, and D. Samaras, “Squared earth mover’s distance- based loss for training deep neural networks,” arXiv preprint arXiv:1611.05916, 2016

  264. [281]

    Predicting the quality of compressed videos with pre-existing distortions,

    X. Yu, N. Birkbeck, Y . Wang, C. G. Bampis, B. Adsumilli, and A. C. Bovik, “Predicting the quality of compressed videos with pre-existing distortions,” IEEE Transactions on Image Processing , vol. 30, pp. 7511–7526, 2021

  265. [282]

    Aim 2024 challenge on compressed video quality assessment: Methods and results,

    M. Smirnov, A. Gushchin, A. Antsiferova, D. Vatolin, R. Timofte, Z. Jia, Z. Zhang, W. Sun, J. Qian, Y . Cao et al., “Aim 2024 challenge on compressed video quality assessment: Methods and results,” arXiv preprint arXiv:2408.11982, 2024

  266. [283]

    Ais 2024 challenge on video quality assessment of user-generated content: Methods and results,

    M. V . Conde, S. Zadtootaghaj, N. Barman, R. Timofte, C. He, Q. Zheng, R. Zhu, Z. Tu, H. Wang, X. Chen et al., “Ais 2024 challenge on video quality assessment of user-generated content: Methods and results,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Patt...

  267. [284]

    Assessing quality of images or videos using a two-stage quality assessment,

    A. C.Bovik, “Assessing quality of images or videos using a two-stage quality assessment,” Jan. 7 2020, US Patent 10,529,066

  268. [285]

    On the use of SSIM in HEVC,

    T. Zhao, K. Zeng, A. Rehman, and Z. Wang, “On the use of SSIM in HEVC,” in Asilomar Conference on Signals, Systems and Computers . IEEE, 2013, pp. 1107–1111

  269. [286]

    VMAF based rate-distortion optimization for video coding,

    S. Deng, J. Han, and Y . Xu, “VMAF based rate-distortion optimization for video coding,” in IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP) , 2020, pp. 1–6

  270. [287]

    Deep perceptual preprocessing for video coding,

    A. Chadha and Y . Andreopoulos, “Deep perceptual preprocessing for video coding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 852–14 861

  271. [288]

    Video pre-processing with JND-based Gaussian filtering of superpix- els, author=Ding, Lei and Li, Ge and Wang, Ronggang and Wang, Wenmin,

    “Video pre-processing with JND-based Gaussian filtering of superpix- els, author=Ding, Lei and Li, Ge and Wang, Ronggang and Wang, Wenmin,” in SPIE Visual Information Processing and Communication VI, vol. 9410, 2015, pp. 20–25

  272. [289]

    Quality-constant per-shot encoding by two-pass learning-based rate factor prediction,

    C. Cai, Y . Wang, X. Li, and T. Ye, “Quality-constant per-shot encoding by two-pass learning-based rate factor prediction,” arXiv preprint arXiv:2208.10739, 2022

  273. [290]

    Predicting rate control target through a learning based content adaptive model,

    H. Xing, Z. Zhou, J. Wang, H. Shen, D. He, and F. Li, “Predicting rate control target through a learning based content adaptive model,” in Picture Coding Symposium (PCS) . IEEE, 2019, pp. 1–5

  274. [291]

    Dynamic optimizer-A perceptual video encoding optimization framework,

    I. Katsavounidis, “Dynamic optimizer-A perceptual video encoding optimization framework,” The NETFLIX tech blog , 2018

  275. [292]

    SSIM Motivated Quality Control for Versatile Video Coding,

    M. Wang, S. Wang, J. Li, L. Zhang, Y . Wang, and S. Ma, “SSIM Motivated Quality Control for Versatile Video Coding,” in Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 2020, pp. 1122–1127

  276. [293]

    SSIM-based perceptual rate control for video coding,

    T.-S. Ou, Y .-H. Huang, and H. H. Chen, “SSIM-based perceptual rate control for video coding,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 21, no. 5, pp. 682–691, 2011

  277. [294]

    Rate-distortion optimization for video compression,

    G. Sullivan and T. Wiegand, “Rate-distortion optimization for video compression,” IEEE Signal Processing Magazine , vol. 15, no. 6, pp. 74–90, 1998

  278. [295]

    Rate-SSIM optimization for video coding,

    S. Wang, A. Rehman, Z. Wang, S. Ma, and W. Gao, “Rate-SSIM optimization for video coding,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2011, pp. 833–836

  279. [296]

    SSIM-motivated rate-distortion optimization for video coding,

    ——, “SSIM-motivated rate-distortion optimization for video coding,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 22, no. 4, pp. 516–529, 2011

  280. [297]

    Elic: Efficient learned image compression with unevenly grouped space- channel contextual adaptive coding,

    D. He, Z. Yang, W. Peng, R. Ma, H. Qin, and Y . Wang, “Elic: Efficient learned image compression with unevenly grouped space- channel contextual adaptive coding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 5718–5727

  281. [298]

    Multi-modality deep network for extreme learned image compression,

    X. Jiang, W. Tan, T. Tan, B. Yan, and L. Shen, “Multi-modality deep network for extreme learned image compression,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 1, pp. 1033– 1041, Jun. 2023

  282. [2016]

    Available: https://doi.org/10.1145/2812802

    [Online]. Available: https://doi.org/10.1145/2812802

  283. [2023]

    Available: https://arxiv.org/abs/2304.08485

    [Online]. Available: https://arxiv.org/abs/2304.08485

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.