{"id":"f799fe47-216d-49be-8a31-5cf604789edf","arxiv_id":"2504.15003","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"KwaiSR is a new image super-resolution benchmark made from short-form user-generated content, and existing SR models and quality metrics struggle on it.","lead":"The paper introduces KwaiSR, a benchmark dataset of 1,800 synthetic and 1,900 real low-quality images from the short-video platform Kwai, plus results from running existing super-resolution models on it. A generalist might read it for the finding that common image-quality metrics rank results differently from human viewers on short-form UGC content.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The synthetic degradation model grounding the 'in the wild' claim is never specified or validated; without it, the benchmark's transferability to real UGC images is unestablished.","rationale":"The reader's weakest_assumption correctly identifies the missing degradation model as a primary fragility, and my analysis agrees that this is the single most load-bearing issue. I downweight the KVQ-filtering concern as secondary because the wild subset's role is mainly to provide real low-quality images for qualitative and no-reference evaluation; even if the KVQ threshold is imperfect, the synthetic subset is the part carrying the quantitative claims. The internal size inconsistencies (1,800 vs 1,900; 360/380 vs 1,800/1,900) are distracting but likely reflect train/validation/test split descriptions rather than a deeper flaw. Since the reader already assigned CONDITIONAL, my recommendation is UNCHANGED: the paper should be accepted only if the authors supply the degradation model, validate it against the wild distribution, and correct the size reporting. The concern does not force a stronger rejection because the dataset, once released, can be independently tested for distribution match.","tokens_in":13958,"tokens_out":2792,"duration_ms":27279,"concrete_test":"Ask the authors to release the full degradation pipeline (blur kernel model and range, noise type and level, compression settings, any stochastic parameters) and then compute interpretable degradation descriptors (estimated blur kernel width, noise variance, blocking/banding artifact strength, frequency falloff) for both the synthetic LR training set and the wild LR set. Run a two-sample MMD test on these descriptors; if the distributions are significantly different, or if a model trained on the synthetic train split performs markedly worse on the wild test split than on the synthetic test split, the claim that the synthetic benchmark represents real short-form UGC degradation is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that KwaiSR is the first benchmark for short-form UGC image super-resolution 'in the wild', with the synthetic subset 'produced by simulating the degradation following the distribution of real-world low-quality short-form UGC images' (Sec. 1). However, the paper never states the degradation model, its parameter values, how the distribution was estimated, or how it was validated against the wild images. The synthetic subset is the only part with ground truth and is used for training and for objective comparisons in Tables 1–2; therefore the conclusion that existing SR methods are 'pretty challenging' for short-form UGC depends entirely on this unstated simulation matching the real UGC degradation distribution. Figure 5 shows only KVQ scores of wild images, not a comparison of synthetic LR degradation characteristics to wild LR characteristics. Without such a comparison, or a cross-domain experiment (e.g., training on the synthetic split and evaluating on the wild split), the 'in the wild' generalization claim is an unsupported assumption. Section 4.1's '360 synthetic image pairs and 380 low-quality wild images' also conflicts with the stated 1,800 pairs and 1,900 images elsewhere, adding protocol ambiguity even if it likely refers to the validation-plus-test splits.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces KwaiSR, a benchmark dataset for short-form UGC image super-resolution, collected from the Kwai platform and comprising 1,800 synthetic LR-HR pairs (or 1,900, per the abstract) and 1,900 wild low-quality images without ground truth. The authors describe the semantic category distribution, report KVQ-based quality analysis, and present the results of the NTIRE 2025 challenge track in which nine teams submitted results evaluated on synthetic and wild splits using fidelity and no-reference perceptual metrics. The central claim is that KwaiSR is the first benchmark dataset for short-form UGC image super-resolution 'in the wild' and that existing SR methods find it challenging.","tokens_in":14161,"tokens_out":4030,"duration_ms":34652,"significance":"If the dataset is appropriately validated and documented, KwaiSR fills a concrete gap: existing SR benchmarks are either synthetic with simple degradations or real-world without ground truth, and no dataset targets the specific characteristics of short-form UGC platforms. The paper's strengths include the real-world collection from the Kwai platform, the inclusion of eleven semantic categories, the public challenge organization with nine participating teams, and the quantitative comparison of nine recent SR methods on both synthetic and wild splits. The dataset release is likely to be useful to the community regardless of the outcome of the challenge. However, the significance is currently tempered by a lack of specification of the synthetic degradation pipeline and by the partly circular quality analysis, both of which bear directly on the 'in the wild' claim.","major_comments":[{"comment":"The dataset size is stated inconsistently. The abstract says the synthetic dataset includes '1,900 image pairs', while Section 1 says 'resulting in 1800 image pairs'; Section 4.1 then says the benchmark includes '360 synthetic image pairs and 380 low-quality wild images'. Please clarify which numbers are correct and whether Section 4.1 refers only to the validation and test subsets, since the figures in Section 3 show 1,440/180/180 synthetic and 1,520/190/190 wild splits.","section":"Abstract and Section 1"},{"comment":"The central claim that the synthetic subset is 'produced by simulating the degradation following the distribution of real-world low-quality short-form UGC images' is never substantiated. The degradation model, its parameter values or ranges, and the procedure used to estimate the distribution from real UGC images are not described, and no comparison is made between the degradation characteristics of the synthetic LR images and the wild LR images. Since the synthetic pairs are the only ground-truth data used for training and for the objective comparisons in Tables 1 and 2, the 'in the wild' validity of the benchmark currently rests on an unstated assumption. Please provide the degradation specification and a direct comparison or a cross-domain experiment (e.g., training on the synthetic split and evaluating on the wild split).","section":"Section 1 and Section 3.2"},{"comment":"The quality analysis is partly circular. The wild images were selected by filtering with KVQ, and Figure 5 then reports the KVQ score distribution of those selected images; this confirms the filter's behavior but does not validate that the selected images reflect the quality distribution of short-form UGC images on the platform or that KVQ scores correspond to human perception. The claim that 'the overall distribution of images aligns with that observed on the Kwai Platform' is made without supporting evidence. Please provide human ratings on a sample of the selected and unselected images, or a comparison with the distribution before filtering.","section":"Section 3.2 and Figure 5"},{"comment":"The conclusion that 'existing metrics are not accurate to measure the human perception quality in this dataset' is supported only by a brief mention of a user study with no described protocol. The number of participants, rating scale, stimulus presentation method, and significance testing are not reported, and no statistics are given for the claim that the top-ranked objective team 'did not deliver the best user experience'. As written, this is an unsupported claim that nevertheless appears as a headline conclusion of the challenge analysis.","section":"Section 3.3"}],"minor_comments":[{"comment":"The abstract states the synthetic dataset includes '1,900 image pairs' while the body of the paper consistently uses 1,800; please make the counts consistent throughout.","section":"Abstract"},{"comment":"Please state explicitly that the '360 synthetic image pairs and 380 low-quality wild images' are the validation and test subsets only, to avoid confusion with the full dataset sizes of 1,800 and 1,900.","section":"Section 4.1"},{"comment":"In the Qualitative Comparison paragraph, the text says 'in Figures 1 and 2' but the referenced qualitative figures are Figures 6 and 7; please correct the cross-reference.","section":"Section 4.2"},{"comment":"The caption labels two panels as '(f) OSEDiff'; the sixth panel should be labeled '(g)'.","section":"Figure 6"},{"comment":"The description of HAT as one that 'outperforms state-of-the-art methods in multiple tasks' is vague; please specify the comparison or remove the unspecific claim.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a challenge report, and some omitted details (e.g., degradation settings, user study protocol) may exist in the companion challenge report [29]. However, the manuscript should stand alone as a dataset description, especially for a dataset that is being released for community use. I recommend requesting the degradation specification and a resolution of the size inconsistencies before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Name],\n\nShort version: KwaiSR is the first benchmark aimed specifically at short-form UGC super-resolution, and that alone makes it worth paying attention to. The paper also runs a challenge and evaluates a dozen methods, which is useful legwork. But the load-bearing piece — the synthetic degradation pipeline that supposedly matches real UGC degradation — is never actually described, and there are some sloppy number inconsistencies. It's a conditional accept as a dataset, not as a finished study.\n\nWhat's genuinely new: a domain-specific dataset with both synthetic pairs and wild low-quality images, split into train/val/test, and a challenge track built around it. The observation that objective metrics disagree with the user study is also a useful data point for the field.\n\nWhat's solid: the evaluation tables are reproducible in principle (they used official code), and the dataset is hosted openly. The semantic category balance is decently thought out.\n\nWhere it gets soft: the biggest issue is that the synthetic subset is the only part with ground truth, and the paper says it was produced by \"simulating the degradation following the distribution of real-world low-quality short-form UGC images\" but never gives the degradation model, parameters, or any validation that the simulation actually matches the wild images. Without that, the \"in the wild\" claim is an assumption. A cross-domain experiment (train on synthetic, test on wild) would have been a quick check, and they didn't do it.\n\nSecond, the numbers don't line up: the abstract says 1,900 pairs then 1,800; Section 4.1 says the benchmark includes 360 synthetic pairs and 380 wild images, which is only the val+test split. That kind of sloppiness doesn't sink a dataset paper, but it erodes trust.\n\nThird, the wild set is filtered by KVQ from the authors' own prior work, with no threshold or human validation. That's acceptable as a choice, but it should be transparent.\n\nFourth, the user study result — that the objective top team didn't win subjectively — is asserted without data, just a pointer to the challenge report. For this paper to stand alone, that needs more substance.\n\nBottom line: the dataset is a real contribution and the paper deserves a serious referee, but it needs major revision: fully disclose the degradation model and its validation, fix the numbers, and either add the user study data or clearly mark it as out of scope. I'd send it back for a major revision, not desk reject it.","headline":"KwaiSR is a genuinely useful new dataset for short-form UGC super-resolution, but the paper's main validity gap is an undescribed synthetic degradation pipeline that needs to be fixed before the 'in the wild' claim holds.","tokens_in":14757,"tokens_out":3461,"would_cite":true,"duration_ms":29683,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"KwaiSR: the first benchmark for super-resolving short-form UGC images, and current methods fail it","keywords":["image super-resolution","short-form UGC","benchmark dataset","KwaiSR","no-reference image quality assessment","diffusion models","NTIRE challenge","real-world degradation"],"falsifier":"Estimate the degradation parameters (blur kernel, noise level, compression strength) from the wild KwaiSR images and compare them with the synthetic low-resolution images; if the estimated distributions differ substantially, the synthetic pairs do not represent real short-form UGC degradation and results on them may not transfer to the wild setting.","tokens_in":13739,"feed_emoji":"🎬","tokens_out":3904,"duration_ms":36621,"temperature":0.7,"pith_summary":"The paper introduces KwaiSR, the first benchmark dataset built specifically for image super-resolution on short-form user-generated content (UGC) from a real video platform. It contains 1,800 synthetic low-resolution/high-resolution image pairs designed to mimic real short-form UGC degradation, plus 1,900 wild low-quality images directly collected from the Kwai platform and filtered by the KVQ quality assessment method. The dataset is used to run the NTIRE 2025 challenge track on short-form UGC image enhancement, and the challenge results show that existing image super-resolution methods, including diffusion-based ones, do not reliably deliver good perceptual quality on this content. The paper further reports that objective metrics disagree with human preference on this data. If correct, KwaiSR gives the community a shared resource for developing and measuring super-resolution algorithms tailored to the distortions found on short-form video platforms.","feed_headline":"First UGC super-resolution benchmark stumps existing methods","feed_subtitle":"KwaiSR pairs 1,800 synthetic and 1,900 wild images to push image super-resolution toward real short-form content.","key_machinery":"The central object is the KwaiSR dataset itself: 1,800 synthetic low-resolution/high-resolution pairs plus 1,900 wild low-quality images from the Kwai platform, spanning eleven semantic categories (mountain, night, water, field, food, caption, person, portrait, crowd, CG, and stage). The synthetic pairs are intended to supply paired ground truth for training and objective comparison, while the wild images provide an unpaired, in-the-wild test set; both are split 8:1:1 into training, validation, and testing. The dataset is the load-bearing mechanism because every experimental finding in the paper—that existing methods struggle, that diffusion methods trade fidelity for realism, and that objective metrics misjudge subjective quality—is derived from running existing super-resolution models on these images.","core_discovery":"The paper's central claim is that KwaiSR is the first benchmark dataset for short-form UGC image super-resolution in the wild, and that this dataset is genuinely hard for current super-resolution methods. The synthetic subset pairs 1,800 low-resolution images with ground-truth 1920×1080 high-resolution images using a simulated degradation that is said to follow the real distribution of low-quality short-form UGC images; the wild subset contains 1,900 low-quality images filtered by the KVQ quality metric. On the synthetic subset, methods must perform 4x super-resolution, while wild images are evaluated at native resolution. Results from the NTIRE 2025 challenge on this dataset show three things: a realism-versus-perceptual-quality trade-off that no method resolves cleanly, failure of no-reference metrics such as MUSIQ, CLIPIQA, and MANIQA to match user experience, and limited effectiveness of existing diffusion-based restoration methods on this content.","pith_inferences":["The paper never specifies the synthetic degradation model or its parameter values, so a reader cannot yet judge whether the synthetic pairs really match the degradation distribution of wild short-form UGC images; a validation study comparing estimated degradation parameters or human paired comparisons between synthetic and wild low-resolution images would test this directly.","Because the challenge's objective metric ranking disagreed with the user study, the reported team rankings may themselves be an artifact of the chosen composite score; re-ranking the teams by human preference alone could give a different picture of which methods actually work.","The dataset could plausibly serve as a frame-level resource for short-form UGC video super-resolution, but the paper does not address temporal consistency, so treating it as a video benchmark would be an extension rather than a claim of the paper.","The night and stage categories are underrepresented in the synthetic subset, which may bias results toward the more common categories; future dataset versions could deliberately balance or augment these difficult scenarios."],"forward_implications":["Existing image super-resolution methods trained on conventional datasets perform noticeably worse on KwaiSR, so new methods will need to handle the specific degradations of short-form UGC content.","The challenge results show a persistent trade-off between realism and perceptual quality: methods that improve fidelity tend to deform faces, while methods that boost perceptual scores generate unreal, AI-style textures.","No-reference quality metrics such as MUSIQ, CLIPIQA, and MANIQA do not align with human preference on this dataset, indicating that progress on this task will require a better quality assessment method.","The synthetic/wild split lets researchers evaluate both paired fidelity (on synthetic data) and generalization to genuine in-the-wild images (on wild data), a combination that previous SR datasets did not offer.","One-step diffusion models such as OSEDiff achieve strong perceptual scores on the wild subset despite lower PSNR and SSIM, suggesting that sampling efficiency does not necessarily hurt perceptual quality on this domain."],"supporting_citations":[{"why":"KVQ is the quality assessment method used to filter low-quality wild images, making it the basis for the wild subset's construction.","marker":"[36]"},{"why":"The companion NTIRE 2025 challenge report provides the user study and team results that ground the paper's claim that KwaiSR is challenging and that metrics fail.","marker":"[29]"},{"why":"StableSR is one of the diffusion-based baselines evaluated on both synthetic and wild subsets, supporting the claim about existing method limitations.","marker":"[48]"},{"why":"DiffBIR is a diffusion-based baseline whose perceptual scores on KwaiSR help demonstrate the realism-versus-perceptual-quality trade-off.","marker":"[33]"},{"why":"OSEDiff is the one-step diffusion baseline that performs best on wild perceptual metrics, supporting the discussion of single-step sampling.","marker":"[56]"},{"why":"DiffIR achieves the highest PSNR and SSIM on synthetic test and validation sets, serving as the fidelity-focused baseline in the comparison.","marker":"[57]"},{"why":"PASD is a semantic-guided diffusion baseline used in the wild and synthetic comparisons, contributing to the evidence on detail preservation and failure modes.","marker":"[61]"},{"why":"InvSR is a one-step diffusion method evaluated on synthetic data, providing a contrast in reconstruction versus perceptual quality.","marker":"[67]"},{"why":"ResShift is a diffusion baseline whose qualitative results on text and color help illustrate the visual failure patterns on KwaiSR.","marker":"[66]"},{"why":"SwinIR is a transformer-based baseline that performs better on text in qualitative comparisons, showing that generative priors are not always advantageous on this data.","marker":"[31]"}],"fun_headline_variants":["First UGC super-res benchmark stumps existing AI methods","KwaiSR: short-form UGC SR dataset too tough for current methods","NTIRE 2025 challenge reveals SR methods falter on UGC content","KwaiSR benchmark: existing super-resolution can't match UGC reality","New UGC SR dataset highlights realism-quality tradeoff"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The synthetic low-resolution images are assumed to reproduce the real degradation distribution of short-form UGC images, but the paper never states the degradation model, its parameters, or how that distribution was estimated or validated against the wild images.","fun_headline_variants_meta":{"raw":{"variants":["First UGC super-res benchmark stumps existing AI methods","KwaiSR: short-form UGC SR dataset too tough for current methods","NTIRE 2025 challenge reveals SR methods falter on UGC content","KwaiSR benchmark: existing super-resolution can't match UGC reality","New UGC SR dataset highlights realism-quality tradeoff"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00026,"raw_usage":{"total_tokens":1633,"prompt_tokens":1030,"completion_tokens":603,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":512}},"tokens_in":646,"tokens_out":603,"duration_ms":6019,"temperature":1.0,"reasoning_tokens":512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:35:31.260192+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the degradation parameters (blur kernel, noise level, compression strength) from the wild KwaiSR images and compare them with the synthetic low-resolution images; if the estimated distributions differ substantially, the synthetic pairs do not represent real short-form UGC degradation and results on them may not transfer to the wild setting.","supporting_citations":[{"cited_title":"Kvq: Kwai video quality assessment for short-form videos","cited_arxiv_id":null,"evidence_quote":"KVQ is the quality assessment method used to filter low-quality wild images, making it the basis for the wild subset's construction."},{"cited_title":"NTIRE 2025 challenge on short-form ugc video quality assessment and enhancement: Methods and results","cited_arxiv_id":null,"evidence_quote":"The companion NTIRE 2025 challenge report provides the user study and team results that ground the paper's claim that KwaiSR is challenging and that metrics fail."},{"cited_title":"Exploiting diffusion prior for real-world image super-resolution","cited_arxiv_id":null,"evidence_quote":"StableSR is one of the diffusion-based baselines evaluated on both synthetic and wild subsets, supporting the claim about existing method limitations."},{"cited_title":"Diff- bir: Toward blind image restoration with generative diffusion prior","cited_arxiv_id":null,"evidence_quote":"DiffBIR is a diffusion-based baseline whose perceptual scores on KwaiSR help demonstrate the realism-versus-perceptual-quality trade-off."},{"cited_title":"One-step effective diffusion network for real-world image super-resolution","cited_arxiv_id":null,"evidence_quote":"OSEDiff is the one-step diffusion baseline that performs best on wild perceptual metrics, supporting the discussion of single-step sampling."},{"cited_title":"Diffir: Efficient diffusion model for image restoration","cited_arxiv_id":null,"evidence_quote":"DiffIR achieves the highest PSNR and SSIM on synthetic test and validation sets, serving as the fidelity-focused baseline in the comparison."},{"cited_title":"Resshift: Efficient diffusion model for image super- resolution by residual shifting","cited_arxiv_id":null,"evidence_quote":"ResShift is a diffusion baseline whose qualitative results on text and color help illustrate the visual failure patterns on KwaiSR."}],"review_version":1}