{"id":"6c59253c-9ecc-4563-86b4-55f871af94c3","arxiv_id":"2504.13131","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The NTIRE 2025 challenge report presents leaderboards for efficient video quality assessment on KVQ and diffusion-based image super-resolution on the new KwaiSR dataset, with user-study results showing subjective and objective quality rankings diverge.","lead":"A computer-vision competition report that ranks 18 algorithms for scoring short-form video quality and for super-resolving social-media images, and describes a new dataset of short-form image pairs. The report documents which methods won and notes that objective quality metrics disagree with human preference for AI-enhanced images.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subjective/objective inconsistency claim rests on an unreported five-expert user study; the paper supplies no inter-rater agreement, confidence intervals, or significance tests for the Table 2 win rates.","rationale":"The reader's weakest assumption identified the same five-expert user study as the crux, and my reading agrees fully. The full text confirms no inter-rater agreement, no confidence intervals, no test count, and no significance test for the subjective rankings. The paper's own sentences frame the raters as five professional experts asked to find the 'most visually convincing and realistic result,' so the claim about perceptual metrics failing to reflect subjective quality is a general inference from a narrow expert sample. The objective-based shortlisting step further reinforces the need for reliability statistics. I considered other potential concerns: the pseudo-labeling by ZX-AIE-Vector on the test set is disclosed in the team's own method section and affects only that team's reported Track 1 numbers, not the paper's main scientific conclusion; the Track 1 leaderboard is otherwise a straightforward report of competition outcomes, and the public dataset and code availability support the factual portions. Therefore the verdict should remain CONDITIONAL, with the condition that the Track 2 user-study conclusion needs supporting statistics (agreement, confidence intervals, or a larger externally validated study) before the headline inference is treated as robust. The reader's proposed conditional verdict is correct; no retreat to REJECT or UNVERDICTED is warranted because the challenge-report facts are credible and the limitation is explicitly locatable in the paper text.","tokens_in":23975,"tokens_out":2923,"duration_ms":22575,"concrete_test":"Request or reconstruct the raw user-study data for Track 2: per-expert, per-image preference judgments for the six shortlisted teams on both synthetic and wild subsets. Compute Fleiss' kappa or Krippendorff's alpha across the five experts, and bootstrap 95% confidence intervals for each team's win rate. Then re-test the paper's inconsistency claim: if the confidence interval around SYSU-FVL-Team's subjective win rate overlaps with the intervals around TACO SR and RealismDiff, or if inter-rater agreement is below 0.3, the subjective/objective inconsistency is not statistically supported by this study.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central interpretive claim, in Section 3 and Table 2, is that \"the results highlight a noticeable inconsistency between subjective preferences and objective metrics, suggesting that current perceptual metrics may not reliably reflect perceptual quality in generative model-based S-UGC image super-resolution.\" This is load-bearing because it is the only scientific conclusion drawn beyond reporting leaderboard ranks. The entire subjective side rests on a user study described in one sentence: five professional image-processing experts spent about eight hours each comparing the top six teams' outputs. No number of test images, no question format, no aggregation rule, no inter-rater agreement (e.g., Fleiss' kappa), no confidence interval, and no significance test is reported. With five raters and six teams, the win-rate differences between adjacent ranks are small (e.g., 0.2775 vs 0.2640 on synthetic; 0.1540 vs 0.0947 on wild); a single expert's preference flips several of these differences. The paper's own text flags the limitation: the raters are professional experts asked to find \"the most visually convincing and realistic result,\" which is a narrow, expert-centric criterion, not a general-population perceptual preference. The conclusion that current perceptual metrics \"may not reliably reflect perceptual quality\" requires generalizing from five experts to the broad viewer population of short-form UGC platforms, and the paper provides no evidence for that generalization. Finally, the user study was conducted after shortlisting teams on objective metrics, so the shortlist and the reported objective rankings are correlated with the study design. This does not make the inconsistency claim false, but it makes Table 2's subjective rankings insufficient to support the conclusion without additional reliability evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This challenge report describes the NTIRE 2025 competition on short-form UGC video quality assessment and enhancement. Track 1 evaluates efficient no-reference VQA models on the KVQ dataset under a 120 GFLOPs constraint, and Track 2 introduces the KwaiSR dataset for diffusion-based single-image super-resolution, reporting objective quality metrics and a user study of the top six teams. The paper presents the leaderboard for Track 1, summarizes the architecture, training, and testing details of each of the 18 participating teams, and concludes that subjective preferences and objective metrics are noticeably inconsistent for generative SR on this content. The challenge repository is made publicly available, and the paper serves as a record of methods and results for the community.","tokens_in":24322,"tokens_out":6593,"duration_ms":59978,"significance":"If the reported results are reliable, the paper offers a useful community resource: a new S-UGC super-resolution dataset (KwaiSR, detailed in a companion paper), a public challenge repository, and a structured overview of 18 submitted methods under a strict efficiency budget. The Track 1 leaderboard, showing strong performance within 120 GFLOPs, is of practical interest for lightweight VQA deployment. The claimed subjective/objective inconsistency in Track 2, if supported by proper statistical evidence, would be an important caution against reading PSNR, LPIPS, or MUSIQ as proxies for user preference when comparing generative SR outputs. The paper is transparent about per-team training protocols and makes participant fact sheets available, which is commendable. However, the evidence for the central interpretive claim is currently thin, and the Track 1 out-of-sample comparison is compromised by at least one team explicitly using test-set pseudo-labels in training.","major_comments":[{"comment":"The training procedure described for ZX-AIE-Vector includes generating pseudo-labels for the KVQ test set, merging them with the training and validation data, and then fine-tuning the lightweight model on the refined pseudo-labeled test data using two training phases. This is a transductive use of the test set and means the reported test performance is not a clean out-of-sample evaluation. The paper does not state whether this practice was permitted by the challenge rules. Because Table 1 is the central deliverable of Track 1, the authors must disclose the rule and, if the practice was not allowed, recompute the leaderboard without this team. At minimum, a sensitivity analysis showing the ranking with and without test-set adaptation is needed.","section":"Section 4.3 (ZX-AIE-Vector)"},{"comment":"The claim that \"current perceptual metrics may not reliably reflect perceptual quality\" rests entirely on a user study described in a single sentence: five professional image-processing experts spent about eight hours each comparing the top six teams' outputs. No number of test images, no stimulus presentation protocol, no aggregation rule, no inter-rater agreement measure, no confidence intervals, and no significance test are reported. With only five raters and six teams, the differences between adjacent winning rates are small (e.g., 0.2775 vs. 0.2640 on synthetic; 0.1540 vs. 0.0947 on wild), so a single expert's preference flips several rank orders. Moreover, the raters were asked to select \"the most visually convincing and realistic result,\" an expert-centric criterion that does not necessarily represent the general viewer population of short-form UGC platforms. The authors should either provide the full user-study protocol with statistical analysis (e.g., Fleiss' kappa, confidence intervals, or a test against chance) or substantially weaken the conclusion to a hypothesis rather than a finding.","section":"Section 3 and Table 2"},{"comment":"The Track 1 leaderboard is presented as a \"Final Score\" combining SROCC, PLCC, Rank1, and Rank2, but the paper never defines how this score is computed. The challenge description mentions coarse-grained quality scoring and fine-grained rankings for difficult samples, yet no formula, weighting, or aggregation rule is given. Without this definition, readers cannot interpret the rankings or reproduce the leaderboard. Please add the exact computation of Final Score, including how the fine-grained Rank1 and Rank2 components enter the score.","section":"Section 2 and Table 1"},{"comment":"The column header \"User Study Score (objective)\" is ambiguous: the first numeric column (e.g., 0.2775/0.3529) appears to be user-study winning rates, while the following columns are conventional objective metrics. The \"Ranking (Objective)\" column is not tied to any stated aggregation of PSNR, SSIM, LPIPS, MUSIQ, ManIQA, or CLIPIQA, and the text says the top six teams were shortlisted by objective metrics, yet objective scores are listed for all nine teams. Please restructure the table to clearly separate user-study results from objective metrics, define the objective ranking criterion, and state how the six teams were selected for the user study.","section":"Table 2"}],"minor_comments":[{"comment":"The text states that KVQ contains \"nine primary content scenarios\" but then lists eight categories (landscape, crowd, person, food, portrait, computer graphic, caption, and stage). Please correct the count or add the missing category.","section":"Section 2"},{"comment":"The phrase \"might inevitability suffer\" should be \"might inevitably suffer.\"","section":"Section 1"},{"comment":"The repository URL \"https://github.com/lixinustc/KVQE- Challenge-CVPR-NTIRE2025\" contains a space after the hyphen; please provide a single, correct URL.","section":"Abstract and Section 1"},{"comment":"The organizer team blocks are titled \"NTIRE2024 Organizers\" although this is the NTIRE 2025 challenge; the year should be corrected.","section":"Appendix A and B"},{"comment":"The loss definitions contain spacing and font inconsistencies, e.g., \"Iper\" appears both italic and non-italic, and the L1 term is written as \"L1(f(Ipsnr,I per),I GT)\". Please unify the notation.","section":"Equation (1), Section 5.1"},{"comment":"The sentence \"The teams with top performances including SharpMind, ZQE, ZX-AIE-Vector, ECNU-SJTU VQA Team, and TenVQA achieved excellent results...\" is informal; please rephrase to a measured description of the numerical results.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The test-set pseudo-labeling practice in Section 4.3 is the most serious issue: if it was disallowed by the challenge rules, the Track 1 leaderboard is invalid and should be recomputed; if it was allowed, the paper must state that explicitly and discuss the implications for out-of-sample generalization. The user study is too thinly documented to support the paper's main scientific claim; the authors should be pushed to either provide the full protocol and statistics or to clearly mark the inconsistency as an observation from a small expert study rather than a general conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the NTIRE 2025 short-form UGC VQA/SR challenge report. It is exactly what it says on the tin: a clean, well-organized summary of two competition tracks, leaderboard tables, and short method descriptions for each submitted team. If you want to know who won, with what architecture and at what compute cost, this paper answers that. The organization is honestly done: 266 participants, 18 valid final submissions with fact sheets, computational limits specified, and the KwaiSR dataset is released through the companion paper [38]. That kind of reproducible infrastructure is real value even though the dataset details live elsewhere.\n\nWhat is actually new here is modest. The KwaiSR dataset is only summarized; the full description is the companion challenge paper. The leaderboard numbers are competition outcomes, not a new scientific result. The one interpretive claim — that subjective preferences and objective metrics are noticeably inconsistent for generative SR, so current perceptual metrics may not reflect what people prefer — is genuinely interesting, but it is also the load-bearing conclusion and it is under-built.\n\nThe subjective side of Table 2 comes from five professional image-processing experts, each spending about eight hours choosing the most visually convincing result among the top six teams. The paper reports winning rates but no number of test images, no aggregation rule, no inter-rater agreement, no confidence intervals, no significance tests. The adjacent win-rate gaps on the synthetic set (0.2775 vs 0.2640) are small enough that one expert flipping preference changes the ranking. Five experts whose job is to spot artifacts are also not a proxy for general short-form UGC viewers. So the \"current perceptual metrics may not reliably reflect perceptual quality\" sentence should be softened to something like \"these perceptual metrics did not match the preferences of our five expert raters in this particular comparison.\" That weaker claim is fine and probably true.\n\nTwo smaller issues. One, the ZX-AIE-Vector team explicitly used test-set pseudo-labels to train, so their leaderboard score is not strictly out-of-sample. The paper discloses this in the team section but does not flag its implication for the comparison. Two, the challenge is organized by the KVQ/KwaiSR creators and the winning teams lean heavily on organizer-group priors; that is a mild self-referential context, not a fatal flaw, but it is worth keeping in mind when citing the benchmark.\n\nBottom line: as a challenge report it is solid, useful, and deserves a serious referee. The leaderboard and dataset release will be cited by people working in efficient VQA and S-UGC SR. The subjective/objective inconsistency claim needs either more reliability evidence or a scaled-back statement before it should be treated as established. I would send it to review with a request to fix exactly that.\n\nRecommendation: accept for peer review, conditional on the authors either reporting agreement/error statistics for the user study or explicitly reframing the conclusion as a case study of five experts rather than a general finding.","headline":"A competent NTIRE challenge report whose only scientific claim — subjective/objective inconsistency in Track 2 — rests on a five-expert user study with no reliability statistics; the paper is worth refereeing but that conclusion needs scaling back or supporting.","tokens_in":25263,"tokens_out":734,"would_cite":true,"duration_ms":9586,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This challenge report claims that for diffusion-based super-resolution of short-form user-generated content, the ranking preferred by human experts does not match the ranking given by standard objective quality metrics.","keywords":["short-form UGC","video quality assessment","efficient VQA","diffusion-based super-resolution","KwaiSR dataset","subjective preference","objective perceptual metrics","benchmark challenge"],"falsifier":"Run the same pairwise user study with a larger and more diverse panel, say 50 or more non-expert viewers, and compute a majority ranking with agreement statistics. If the majority ranking matches the objective ordering, or if some standard metric such as LPIPS or MUSIQ correlates with the larger panel's preferences, then the claimed inconsistency is an artifact of the five-expert panel rather than a general property of current metrics.","tokens_in":23761,"feed_emoji":"🎞️","tokens_out":6478,"duration_ms":56862,"temperature":0.7,"pith_summary":"This challenge report claims that two things can be done for short-form user-generated content (S-UGC): video quality can be assessed accurately by lightweight models, and diffusion-based super-resolution can improve the subjective quality of in-the-wild images. The paper's central finding, however, is that the two goals do not line up: in the super-resolution track, the ranking produced by standard objective metrics disagreed with the ranking produced by five expert viewers. A method that won the objective leaderboard was rated fourth by experts, while the experts' top choice ranked sixth objectively. If the user study is representative, leaderboard positions on PSNR, LPIPS, or MUSIQ should not be read as statements about which enhanced image people actually prefer.","feed_headline":"Experts and objective metrics pick different AI upscaling winners","feed_subtitle":"A five-expert study ranked the objective leader's output fourth, signaling a gap in PSNR, LPIPS, and MUSIQ.","key_machinery":"The argument is carried by the challenge's evaluation protocol and its new dataset. Track 1 uses the KVQ database with a coarse-to-fine scoring scheme and a hard limit of 120 GFLOPs, forcing contestants to trade accuracy against compute. Track 2 introduces the KwaiSR dataset (1,800 synthetic paired images and 1,900 real-world low-quality images, split 8:1:1) and evaluates the six objectively best submissions through a five-expert user study in which each expert spent about eight hours choosing the most realistic result. That human-preference step is the load-bearing mechanism: it produces the subjective rankings against which the objective metrics are compared.","core_discovery":"On its own terms, the paper establishes a benchmark result and a negative result. The benchmark result is that efficient video quality assessment under a 120 GFLOPs limit is viable: the winning Track 1 model scored 0.922 on the combined metric while using 47.39 GFLOPs and 33.01M parameters, with the top five teams all above 0.90. The negative result is that, for diffusion-based super-resolution of short-form UGC images, subjective preference and objective quality metrics diverge. In the user study of the six teams shortlisted by objective performance, TACO SR had the highest expert winning rates (0.2775 on synthetic and 0.3529 on wild images) yet ranked sixth by objective measures, while the objectively top-ranked team, SYSU-FVL-Team, ranked fourth in expert preference. The paper states this as evidence that current perceptual metrics may not reliably reflect perceived quality in generative-model-based S-UGC super-resolution.","pith_inferences":["Editorial inference: the expert votes may be tracking a realism axis, namely texture plausibility and absence of generative artifacts, that fidelity-oriented metrics like PSNR and SSIM are structurally blind to; a metric that scores realism rather than reconstruction error would likely close part of the gap.","Editorial inference: if the inconsistency generalizes beyond five experts, the same divergence should appear in other generative restoration tasks such as face restoration, denoising, and video enhancement, where objective benchmarks similarly dominate.","Editorial inference: the six final outputs plus the expert preference judgments could be reused as training data for a lightweight reward model or ranking-based no-reference metric; whether such a metric beats MUSIQ on held-out expert choices is a direct, testable next step."],"forward_implications":["Generative super-resolution leaderboards that rely on PSNR, SSIM, LPIPS, or MUSIQ may reward the wrong teams; future challenge reports should include a human preference stage or a metric calibrated to it.","Lightweight VQA is deployable at platform scale: the best Track 1 result was achieved at 47.39 GFLOPs with 33.01M parameters, well under the 120 GFLOPs ceiling.","The KwaiSR dataset gives the research community a shared test bed with both synthetic pairs and real low-quality images for short-form UGC super-resolution.","The top Track 1 teams used teacher-student pseudo-labeling or hybrid Mamba-attention designs to stay under the compute budget, so efficiency can come from training strategy rather than from shrinking a single network alone."],"supporting_citations":[{"why":"Provides the KVQ video database, the training and evaluation benchmark for Track 1 and the content source for the challenge.","marker":"[51]"},{"why":"Introduces the KwaiSR dataset and task framing that Track 2 rests on.","marker":"[38]"},{"why":"Supplies the quality-aware features and teacher model used for pseudo-labeling and distillation by Track 1 teams.","marker":"[73]"},{"why":"Defines the fragment-sampling strategy used for efficient spatial-temporal modeling by several Track 1 teams.","marker":"[87]"},{"why":"The dual-branch aesthetic and technical base model that ZQE adapts with attention-based fusion and a video semantic extractor.","marker":"[88]"},{"why":"The large-scale video-quality pretraining set used by ZX-AIE-Vector before fine-tuning on KVQ.","marker":"[96]"},{"why":"The real-world degradation pipeline several Track 2 teams reuse or refine to synthesize low-quality training images.","marker":"[83]"},{"why":"The adjustable pixel and semantic super-resolution scheme that TACO SR and SYSU-FVL-Team build on to control fidelity versus realism.","marker":"[66]"},{"why":"The large diffusion restoration model used as the second stage of RealismDiff's two-stage pipeline.","marker":"[98]"},{"why":"The efficient Markov-chain diffusion scheduler NetLab adopts for fast 15-step sampling.","marker":"[101]"}],"fun_headline_variants":["Experts and metrics disagree on best AI upscaler","Top AI upscaler by metrics ranks fourth in expert preference","Objective metrics don't reflect perceived quality in UGC upscaling","Human preference and objective scores diverge in UGC SR","Perceptual metrics fail to predict user taste in AI upscaling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole subjective-versus-objective conclusion rests on the assumption that five professional image-processing experts, each spending about eight hours, represent how viewers in general would rank the six super-resolution outputs; the paper reports winning rates but no inter-rater agreement, confidence intervals, or statistical test.","fun_headline_variants_meta":{"raw":{"variants":["Experts and metrics disagree on best AI upscaler","Top AI upscaler by metrics ranks fourth in expert preference","Objective metrics don't reflect perceived quality in UGC upscaling","Human preference and objective scores diverge in UGC SR","Perceptual metrics fail to predict user taste in AI upscaling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000611,"raw_usage":{"total_tokens":2869,"prompt_tokens":997,"completion_tokens":1872,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":1789}},"tokens_in":613,"tokens_out":1872,"duration_ms":12657,"temperature":1.0,"reasoning_tokens":1789,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:13:14.604804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pairwise user study with a larger and more diverse panel, say 50 or more non-expert viewers, and compute a majority ranking with agreement statistics. If the majority ranking matches the objective ordering, or if some standard metric such as LPIPS or MUSIQ correlates with the larger panel's preferences, then the claimed inconsistency is an artifact of the five-expert panel rather than a general property of current metrics.","supporting_citations":[{"cited_title":"Kvq: Kwai video quality assessment for short-form videos","cited_arxiv_id":null,"evidence_quote":"Provides the KVQ video database, the training and evaluation benchmark for Track 1 and the content source for the challenge."},{"cited_title":"NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Kwaisr dataset and study","cited_arxiv_id":null,"evidence_quote":"Introduces the KwaiSR dataset and task framing that Track 2 rests on."},{"cited_title":"Fast- vqa: Efficient end-to-end video quality assessment with fragment sampling","cited_arxiv_id":null,"evidence_quote":"Defines the fragment-sampling strategy used for efficient spatial-temporal modeling by several Track 1 teams."},{"cited_title":"Patch-vq:’patching up’the video qual- ity problem","cited_arxiv_id":null,"evidence_quote":"The large-scale video-quality pretraining set used by ZX-AIE-Vector before fine-tuning on KVQ."},{"cited_title":"Real-esrgan: Training real-world blind super-resolution with pure synthetic data","cited_arxiv_id":null,"evidence_quote":"The real-world degradation pipeline several Track 2 teams reuse or refine to synthesize low-quality training images."},{"cited_title":"Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild","cited_arxiv_id":null,"evidence_quote":"The large diffusion restoration model used as the second stage of RealismDiff's two-stage pipeline."},{"cited_title":"Ef- ficient diffusion model for image restoration by residual shifting","cited_arxiv_id":null,"evidence_quote":"The efficient Markov-chain diffusion scheduler NetLab adopts for fast 15-step sampling."}],"review_version":1}