{"id":"db1dde86-de40-45ff-9088-ec7b61e68594","arxiv_id":"2412.04508","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey of video quality assessment methods and databases, with benchmark comparisons of full-reference and no-reference models on UGC and AIGC datasets.","lead":"This paper surveys video quality assessment (VQA), reviewing subjective databases and objective algorithms from handcrafted methods to deep learning and large multimodal models. It also compares representative models on user-generated and AI-generated video datasets, providing a taxonomy and benchmark overview.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark tables V and VI mix numbers from different evaluation protocols, undermining the paper's comparative claims.","rationale":"I reviewed the paper in good faith. Its central claim is to provide a comprehensive, reliable survey and an effectiveness/adaptability comparison of FR and NR VQA models. The taxonomy and literature coverage appear broad, and the survey's descriptive content has independent value. However, the benchmark tables are the only place where the paper makes a novel empirical statement about model performance. Those tables compare numbers drawn from heterogeneous sources: some from original papers (which themselves use different training and test protocols) and some from the authors' own runs (with no specified protocol). The reader's weakest assumption correctly identifies this as the core vulnerability. My proposed test is deliberately concrete: take two open-source models already in Table VI, evaluate them under a single explicit protocol, and check whether the reported ordering and values reproduce. If they do not, the design insights (e.g., 'Transformer-based and LMM-based models demonstrate the most excellent performance') lose their evidential basis and the paper must be revised to either supply a unified protocol or soften the comparative claims. Other issues—citation numbering errors, typos, and a missing methodology for Fig. 1—are secondary and fixable; they do not threaten the central claim as directly as the unverifiable benchmark comparison. I therefore find no reason to change the reader's CONDITIONAL verdict.","tokens_in":52914,"tokens_out":6780,"duration_ms":62493,"concrete_test":"Run DOVER and FAST-VQA (both open-source) on T2VQA-DB and GAIA with an identical, documented protocol: 32 uniformly sampled frames, resize to each model's native input resolution, use official pre-trained weights without fine-tuning, and compute SROCC/PLCC after a single 4-parameter logistic fit. If the resulting values or their cross-model ordering deviate by more than 0.05 SROCC from Table VI, the benchmark comparison is protocol-dependent and the comparative insights in Section V are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V's Tables V and VI are the paper's only novel empirical contribution—a 'comprehensive comparison' that grounds design insights (e.g., Transformers and LMMs outperform other families). This comparison hinges on SROCC/PLCC values being comparable across models, yet Section V-B states only 'If available, performance data was taken from the original papers, otherwise we conducted the evaluation,' with no protocol details (train/test splits, fine-tuning, frame sampling, preprocessing, temporal pooling, logistic fitting). The numbers show suspicious cross-model gaps: DeepQA scores 0.0815 SROCC on LIVE-YT-HFR while LPIPS scores 0.6920, and SAMA scores 0.0136 on T2VQA-DB while DOVER scores 0.7609. Such gaps are more plausibly explained by differing evaluation protocols than by model quality alone, so the comparative conclusions currently lack a reproducible foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a survey of video quality assessment (VQA), covering subjective study methodology and VQA databases, full-reference and no-reference objective models, deep learning architectures, loss functions, benchmark comparisons on legacy/UGC/AIGC content, and applications and challenges. Its stated central claim is that it provides a comprehensive survey of recent progress in VQA algorithms and of the benchmarking studies and databases that support them. The only novel empirical contribution is the performance comparison in Tables V and VI, which reports SROCC/PLCC for representative FR and NR image/video quality models across five FR databases and five NR databases.","tokens_in":53031,"tokens_out":4585,"duration_ms":46891,"significance":"If the taxonomy and the benchmark comparison are reliable, this would be a useful reference for the VQA community: the paper covers a very broad literature, including recent transformer-based and large-multimodality-model methods, summarizes a large number of subjective databases including AIGC databases, and provides design-oriented observations. It also makes a public GitHub repository available. However, the survey's reliability is currently weakened by two load-bearing problems: the evaluation protocol underlying Tables V and VI is not specified, and the citation numbering is internally inconsistent in ways that make it difficult to trace claims to sources. The paper does not ship machine-checked proofs or reproducible evaluation code for the benchmark, so the comparison must be assessed on the strength of the protocol description, which is currently insufficient.","major_comments":[{"comment":"The only empirical contribution of the paper is the cross-model comparison in Tables V and VI, and the conclusions in Sections V-C and V-D (e.g., that transformer-based and LMM-based models \"demonstrate the most excellent performance\") depend directly on the comparability of the SROCC/PLCC values. Section V-B states only that \"if available, performance data was taken from the original papers, otherwise we conducted the evaluation,\" without specifying train/test splits, whether image models were used off-the-shelf or fine-tuned, frame sampling and temporal pooling, preprocessing, or the logistic fitting used for PLCC. Some entries in the tables are themselves implausible under any single protocol: for example, Table V reports DeepQA at 0.0815 SROCC on LIVE-YT-HFR while LPIPS reports 0.6920, and Table VI reports SAMA at 0.0136 SROCC on T2VQA-DB while DOVER reports 0.7609. These gaps are far more likely to reflect different evaluation protocols than model quality. The authors should provide a per-entry provenance for every number in Tables V and VI, release the evaluation code and checkpoints, and either rerun all models under one protocol or clearly mark the provenance and add caveats where numbers are not comparable.","section":"Section V-B, Tables V and VI"},{"comment":"The citation numbering is internally inconsistent, which is load-bearing for a survey whose purpose is to let readers trace claims. SSIM is cited as [13] in Section II-A but as [164] in Section IV-A; MS-SSIM is cited as [13] in Section II-A and as [14] in Fig. 7; VMAF is cited as [43] in Section II-B, as [171] in Section IV-A Type iv, and as [173] in the same subsection; ST-GREED is cited as [137] in Table V and as [138] in Section IV-A; MC-SSIM is cited as [166] in the text and as [171] in Fig. 7; 3D-SSIM is cited as [167] in the text and as [172] in Fig. 7. A reader cannot reliably identify which reference a number or claim belongs to, so the \"comprehensive survey\" claim is weakened. The manuscript needs a systematic reference audit before it can be considered publication-ready.","section":"Throughout; e.g., Sections II-A, II-B, IV-A and Fig. 7"},{"comment":"The design insights drawn from Tables V and VI go beyond what the data can support given the protocol described. For example, the observation that \"deep learning-based IQA models perform reasonably well on general distortion datasets\" and the advice to prefer temporal NN modules are based on a small, non-random selection of models and on numbers whose protocol is opaque. Even if the protocol problem were fixed, the paper should temper these statements by noting the small number of databases, the content overlap among them, and the fact that many entries are copied from papers with different training regimes.","section":"Section V-C and V-D"}],"minor_comments":[{"comment":"Equation (5) writes the monotonicity loss as a double sum over i only (both summation indices are i), whereas the pair term L_ij^rank in Eq. (3) requires two distinct indices; the second index should be j.","section":"Section IV-C, Eq. (5)"},{"comment":"The logistic function for nonlinear regression appears malformed: the expression exp{(-x + beta_3/|beta_4|}) has mismatched parentheses and the placement of the fraction is unclear. Please rewrite it in standard form, e.g., f(x) = beta_2 + (beta_1 - beta_2)/(1 + exp((x - beta_3)/|beta_4|)).","section":"Section V-A, Eq. (12)"},{"comment":"There are numerous typographical errors that should be corrected in a final pass, including \"percieved\" (Introduction), \"seventies\" for \"severities\" (Section III-A1), \"sucecesses\" (Section II-D), \"correspondance\" (Section II-E), \"insterest\" (Section IV-B1), \"histgram\" (Fig. 7), \"Pre-processig\" (Table I), \"A VC\" for \"AVC\" (Section III-B3), and \"spaital\" for \"spatial\" (Table III). These do not block the scientific content but they do reduce the professionalism of the manuscript.","section":"Throughout"},{"comment":"Please define the meaning of \"-/-\" entries in Tables V and VI (not evaluated, not reported in the original paper, or not applicable) and add a legend explaining that italic and orthographic fonts distinguish IQA from VQA models, since the table captions currently rely on the reader inferring this convention.","section":"Section V-B"}],"recommendation":"major_revision","confidential_remarks":"The core survey content is potentially valuable, but the paper cannot be accepted while the benchmark comparison lacks a reproducible protocol and while the citation numbering is inconsistent. These are fixable within the scope of a revision, so I do not recommend rejection. I would also suggest asking the authors to justify the selection of models and databases in Tables V and VI, since the \"comprehensive\" claim in the title and abstract depends on that selection being representative rather than convenient."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — here is my take.\n\nThe paper is a genuinely useful survey. Its scope is current and broad: it covers subjective databases across legacy, UGC, and AIGC content, organizes FR and NR algorithms by architecture and temporal-modeling strategy, includes recent LMM-based methods, and discusses applications like transcoding and live streaming. The taxonomy diagrams and dataset tables are a convenient entry point. For a survey, the absence of a new algorithm or dataset is not a flaw; its stated contribution is synthesis, and the synthesis is mostly coherent.\n\nThe soft spots are real, and I agree with the stress-test that the benchmark section is the weakest link. Tables V and VI are the only new empirical content, yet the protocol is under-specified. The paper says performance data was taken from original papers when available and otherwise computed by the authors, but it does not say which cells came from where, nor report train/test splits, fine-tuning, frame sampling, temporal pooling, or the logistic fit used for PLCC. With numbers mixed across protocols, the comparative gaps (e.g., DeepQA at 0.0815 SROCC on LIVE-YT-HFR while LPIPS scores 0.6920; SAMA at 0.0136 on T2VQA-DB while DOVER scores 0.7609) cannot be read as pure model-quality differences. That does not make the benchmark fraudulent, but it does make the design insights in Section V less reliable than the paper implies. This is fixable with disclosed provenance and released evaluation code.\n\nThere are also smaller reliability problems. Citation numbering is inconsistent (SSIM appears as [13] and [164], VMAF as [171] and [173]), and typos are scattered through the text. For a reference survey, that erodes trust.\n\nWhat holds up: the taxonomy is coherent, the dataset catalog is informative, and I found no sign of fabricated results. The overall direction of the comparison—transformers and LMMs lead on UGC/AIGC, knowledge-driven FR models remain competitive on temporal distortions—is plausible, just not yet reproducible.\n\nI would send this to peer review. The right referee ask would be to disclose the provenance of every number in Tables V and VI and to clean up the reference numbering. After that, I would be happy to cite it as a survey and hand it to a student.","headline":"Useful, current VQA survey, but the benchmark tables mix protocols and need provenance before the comparative claims can be trusted.","tokens_in":53590,"tokens_out":3745,"would_cite":false,"duration_ms":38822,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Video quality assessment has shifted from handcrafted statistical metrics to deep and multimodal models, and this survey benchmarks that shift on user-generated and AI-generated video, finding the learned models ahead except on temporal…","keywords":["video quality assessment","subjective quality studies","full-reference VQA","no-reference VQA","deep learning","large multimodal models","user-generated content","AI-generated content"],"falsifier":"Re-run every model in Tables V and VI on the same databases under one fixed protocol: identical train and test splits, identical preprocessing and frame sampling, identical nonlinear mapping to MOS, and identical subject subsets. If the resulting SROCC and PLCC rankings differ materially from the tables, the survey's comparative insights and design recommendations are not supported.","tokens_in":52716,"feed_emoji":"🎬","tokens_out":10203,"duration_ms":92277,"temperature":0.7,"pith_summary":"This survey tries to establish a reliable map of the current video quality assessment field: how subjective databases are built, how full-reference and no-reference algorithms have evolved, and which models actually predict human judgments best on emerging content. Its central comparative claim is that deep learning and large multimodal models now outperform traditional handcrafted statistical metrics on user-generated and AI-generated video, while several handcrafted full-reference models remain competitive on temporal distortions such as frame-rate variation. If the map is accurate, it tells streaming platforms and codec designers where to place their bets: learned no-reference models for real-world UGC, multimodal models for AIGC, and bespoke temporal models for high-frame-rate and frame-interpolated content. The survey also argues that data scarcity and the cost of subjective studies are the main limits on further progress.","feed_headline":"Deep and multimodal models lead video quality prediction","feed_subtitle":"The survey maps which metrics to trust for streaming, user-generated, and AI-generated video.","key_machinery":"The device that carries the survey's argument is a double taxonomy plus a comparative benchmark. The taxonomy separates subjective studies and databases from objective algorithms, and within objective algorithms separates full-reference from no-reference and knowledge-driven from deep learning-based; within deep models it separates temporal pooling, 3D CNNs, transformers, and large multimodality models. The benchmark then maps representative models onto databases with legacy, UGC, and AIGC content, using SROCC and PLCC, two rank and linear correlation measures against human mean opinion scores. This machinery lets the authors convert a literature review into a performance map, and it is also the point where the survey's assumptions are concentrated.","core_discovery":"In the paper's own framing, the field of video quality assessment has shifted from measuring predefined distortion properties to learning quality from human-labeled video, and its benchmark section is the evidence. The survey classifies objective models into knowledge-driven versus deep learning-based categories, then reports SROCC and PLCC correlations with human scores for representative full-reference and no-reference models on five FR databases and five NR databases, including user-generated and AI-generated content. On those tables, transformer-based and large-multimodality models such as FAST-VQA, DOVER, COVER, MaxVQA, Q-Align, and LMM-VQA show the highest correlations on UGC and AIGC, while knowledge-driven models such as VMAF and ST-GREED remain strong on traditional and temporal distortions. The paper's stated conclusion is that effective temporal modeling and the integration of multimodal priors are the two most productive directions for future VQA.","pith_inferences":["The benchmark's cross-paper numbers leave an open question: because SROCC and PLCC values are drawn from original papers where training splits and preprocessing differ, the exact ordering of models is less certain than the broad split between handcrafted and learned approaches; a unified re-evaluation could shift positions within each family.","A natural testable extension would be an AIGC benchmark that holds the generator, prompt set, and frame rate fixed while varying only the VQA model, to isolate how much of the large-multimodal-model advantage comes from text alignment versus raw fidelity scoring.","The survey's emphasis on temporal distortion suggests that future gains on UGC and AIGC may come from explicit motion and memory modules rather than from larger spatial backbones, a hypothesis one could test by ablating temporal modules while holding backbone size fixed.","The dual demands of AIGC quality, perceptual fidelity and prompt-video alignment, may require VQA to split into two scores rather than one mean opinion score, extending the aesthetic and technical decomposition the survey documents in DOVER."],"forward_implications":["For user-generated content, the survey implies practical blind quality monitoring should use deep or transformer-based no-reference models rather than frame-level image metrics, since the benchmark shows knowledge-driven IQA models performing poorly on video.","For AI-generated video, multimodal models and prompt-based scoring are the only tested family that tracks human judgment across both fidelity and text alignment, so AIGC evaluation is likely to be built around them.","For high-frame-rate and frame-interpolated video, general-purpose metrics are insufficient; the survey's tables show bespoke temporal models like ST-GREED and FloLPIPS winning those databases, so codec and interpolation evaluation should use task-specific metrics.","Temporal aggregation is not optional: the survey's comparison indicates that simple frame score averaging fails, and memory-aware pooling or recurrent or temporal modules are needed to match human perception.","If these rankings hold, adopting the leading learned metrics as loss functions in coding and enhancement pipelines should improve perceptual optimization beyond what SSIM and VMAF-based losses achieve today."],"supporting_citations":[{"why":"Supplies the LIVE-VQA database used as the full-reference benchmark for legacy compression and transmission distortions.","marker":"[105]"},{"why":"Supplies the LIVE-YT-HFR database used to benchmark models on combined compression and frame-rate variation.","marker":"[117]"},{"why":"Supplies the BVI-VFI database used to benchmark models on frame-interpolation distortions.","marker":"[118]"},{"why":"Supplies the LIVE-VQC database used as the no-reference benchmark for large-scale in-the-wild UGC.","marker":"[111]"},{"why":"Supplies the KoNViD-1k database used to compare no-reference models on diverse authentic natural distortions.","marker":"[119]"},{"why":"Supplies the T2VQA-DB database used as the AIGC benchmark for text-to-video fidelity and alignment.","marker":"[145]"},{"why":"Supplies the GAIA database used as the AIGC benchmark for action quality in generated video.","marker":"[146]"},{"why":"ST-GREED is the frame-rate-aware full-reference model whose high results on HFR databases support the temporal-distortion finding.","marker":"[137]"},{"why":"VMAF is the deployed fusion-based full-reference model that sets the traditional baseline in the FR benchmark and in coding applications.","marker":"[173]"},{"why":"FAST-VQA is the fragment-sampling transformer model whose top UGC results support the transformer-based design insight.","marker":"[95]"}],"fun_headline_variants":["Survey: deep learning and LMMs top video quality prediction","Video quality survey: why AI models beat handcrafted on UGC","Deep learning and multimodal models lead video quality metrics","How AI models predict video quality better than traditional metrics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's comparative conclusions rest on the premise that the SROCC and PLCC numbers in its benchmark tables are directly comparable across models, even though they are taken from different papers with unstated differences in training splits and preprocessing.","fun_headline_variants_meta":{"raw":{"variants":["Survey: deep learning and LMMs top video quality prediction","Video quality survey: why AI models beat handcrafted on UGC","Deep learning and multimodal models lead video quality metrics","How AI models predict video quality better than traditional metrics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1331,"prompt_tokens":935,"completion_tokens":396,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":328}},"tokens_in":551,"tokens_out":396,"duration_ms":4446,"temperature":1.0,"reasoning_tokens":328,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:48:39.956650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run every model in Tables V and VI on the same databases under one fixed protocol: identical train and test splits, identical preprocessing and frame sampling, identical nonlinear mapping to MOS, and identical subject subsets. If the resulting SROCC and PLCC rankings differ materially from the tables, the survey's comparative insights and design recommendations are not supported.","supporting_citations":[],"review_version":1}