{"id":"cee25576-6696-4a24-afad-95965c79af30","arxiv_id":"1908.11517","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This paper introduces UGC-VIDEO, a 550-clip TikTok-based video quality database with subjective scores, and shows existing quality metrics correlate only moderately with human perception.","lead":"Researchers built a video quality database from 50 TikTok clips, recompressed them with two codecs at five quality levels, and collected ratings from 30 people. The database lets companies and researchers test how well automated video quality scores work on ordinary user-shot video.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No public dataset or score files are provided, so the central claim of a new subjective UGC database cannot be independently verified; the non-pristine references further complicate DMOS interpretation.","rationale":"The reader's conditional verdict is appropriate: the paper is a dataset-and-benchmark contribution whose core artifact is not accessible, so the central claims cannot be checked. The reader's weakest_assumption focuses on the non-pristine nature of the TikTok source videos used as references. That is a legitimate technical concern and it does affect how DMOS should be interpreted, but the paper explicitly motivates the database from the absence of pristine sources in UGC, making the choice less of a hidden flaw and more of a design premise. The more load-bearing problem is that no data or scores are released at all; without them, even the existence of the 550 sequences and the subjective ratings is unverifiable. I agree with the reader's overall verdict (CONDITIONAL) and therefore recommend no change. The disagreement is partial because I would rank the missing data artifact above the non-pristine reference issue in deciding the conditional status; the non-pristine references are a secondary complication that would be addressable if the scores were public.","tokens_in":8326,"tokens_out":6897,"duration_ms":70148,"concrete_test":"Search the paper, arXiv metadata, and the web for a 'UGC-VIDEO' download link. If none exists, contact the corresponding author to request the full database: 50 source videos, 500 compressed videos, per-subject ratings, MOS, DMOS, and the source-quality category labels from Section IV.A. Upon receipt, independently recompute the SROCC, PLCC, and RMSE values in Tables I and II using the logistic mapping of Eq. 5, and verify the reported VMAF values (0.8726 on DMOS over all data, 0.8141 on MOS over all data) and the category-1/2 results. If the data are public and the numbers reproduce, the concern is resolved; if the data cannot be obtained or the numbers differ, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the creation of the UGC-VIDEO database and the benchmark showing that existing quality measures correlate only moderately with human opinion. For that claim to hold, the 50 source videos, 500 compressed versions, and the subjective scores (per-subject ratings, MOS, DMOS) must exist and be correctly computed. The least secure condition is verifiability: the paper provides no URL, repository, or supplementary material with any of these artifacts. Without the data, none of the tables can be audited; a transposed column, a mislabeled QP, or a different logistic fitting procedure could materially shift the headline SROCC values (e.g., VMAF 0.8726 on DMOS, 0.8141 on MOS). This is not a cosmetic issue for a dataset paper: the resource itself is the evidence, and its absence makes the core contribution an unreviewable assertion. The reader's weakest_assumption about non-pristine references is real but partly acknowledged in the paper's motivation: UGC uploads have no pristine original, so recompressing already-compressed TikTok videos is an intentional design choice. The DMOS computation in Section III.B is non-standard because the reference is not pristine, and the category-1/2 split in Section IV.A conditions on reference quality, potentially introducing range-restriction artifacts. These complications reinforce the need for the actual scores to be released so that alternative analyses (e.g., MOS-only benchmarks or partial correlations controlling for source quality) can be run. The decisive gap, however, is the missing data artifact itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UGC-VIDEO, a new subjectively annotated database for user-generated video quality assessment, consisting of 50 source videos collected from TikTok (covering selfie, indoor, outdoor, and screen content), each further compressed with H.264 and H.265 at five quantization levels, yielding 550 sequences including the sources. Subjective ratings were collected from 30 subjects using the ACR-HR paradigm, and MOS and DMOS scores were computed. The authors benchmark twelve objective quality metrics using SROCC, PLCC, and RMSE, and analyze performance by reference quality (category 1 vs. 2) and by content category. The main reported finding is that existing metrics, including VMAF, correlate only moderately with human opinion on this database, with VMAF achieving the best overall SROCC of 0.8415 on DMOS and 0.8141 on MOS.","tokens_in":8591,"tokens_out":4750,"duration_ms":44528,"significance":"If the database and scores are made available, this would be a useful resource for the VQA community, since existing databases mostly rely on pristine sources or synthetic distortions, whereas UGC-VIDEO reflects the realistic multi-stage compression scenario on a hosting platform. The content-aware sampling strategy based on SI, TI, and blur (Section II.C) is a methodological strength, as is the use of standard subjective-testing procedures (ITU-R BT.500-13). The benchmark results provide a baseline for future UGC VQA research and align with the growing interest in no-reference and reduced-reference quality assessment. However, as a dataset paper, the absence of any data availability statement is a serious limitation: the core contribution—the subjective scores and the video stimuli—cannot be independently verified or reused.","major_comments":[{"comment":"The manuscript provides no URL, repository, or supplementary material for the UGC-VIDEO database, the subjective scores (per-subject ratings, MOS, DMOS), or the source/processed video files. Because the central claim of the paper is the creation of a new subjective database and the benchmark results derived from it, the absence of public access makes Tables I and II unverifiable. A dataset paper must make the data available, or at minimum release the MOS/DMOS scores and a description of how to obtain the videos; without this, the contribution is an assertion rather than a resource. This should be fixed by providing a stable download link and data format description.","section":"Sections II and IV (overall)"},{"comment":"The paper claims a \"significant performance degradation\" on low-quality reference videos (category 1) and draws conclusions about the relative performance of algorithms (e.g., VMAF being best), but no confidence intervals, bootstrap estimates, or significance tests are reported. With 500 compressed videos and correlations in the range 0.7–0.9, differences such as SROCC 0.8415 (VMAF) vs. 0.8443 (ViS3) may not be statistically meaningful. The authors should add significance testing (e.g., bootstrap on subjects or a Steiger test for correlated correlations) to support the benchmark claims.","section":"Section IV.A, Tables I and II"},{"comment":"The DMOS is computed as the difference between the source and distorted video scores, but the source videos are already-compressed TikTok uploads with widely varying quality (as acknowledged in Section IV.A and Fig. 3). This makes the reference non-pristine, so DMOS conflates the recompression effect with the intrinsic quality of the source. The subsequent split into category 1 (low-quality source) and category 2 (higher-quality source) conditions on source MOS, which can introduce range-restriction artifacts: lower correlations in category 1 might reflect a narrower DMOS range or a ceiling/floor effect rather than a genuine failure of the objective metrics. The authors partially acknowledge the issue, but they do not control for it in the benchmark. Please report MOS-based results as the primary analysis or include partial correlations controlling for source MOS, and discuss how the DMOS benchmark should be interpreted given the non-pristine reference.","section":"Section III.B and Section IV.A"},{"comment":"The text states that \"110 compressed videos with low quality source are classified into the category 1\" and the remaining 390 into category 2, based on a 20th percentile split of source MOS. Since there are 50 source videos and 10 compressed versions each, a 20th percentile split should select 10 sources and hence 100 compressed videos, not 110 (which would correspond to 11 sources). The paper should clarify the exact computation of the percentile threshold and how ties were handled; this is needed for reproducibility of the category-level results in Table I.","section":"Section IV.A"}],"minor_comments":[{"comment":"The sentence \"Considering the fact that our primary goal of investigating the quality assessment of UGC videos for improving the video coding/transcoding performance\" is grammatically incomplete; please rephrase.","section":"Section II.D"},{"comment":"The column heading \"Ourdoor\" should be corrected to \"Outdoor\".","section":"Table II"},{"comment":"There is a typo: \"dummpy presentations\" should be \"dummy presentations\".","section":"Section III.A"},{"comment":"There is a typo: \"As suah\" should be \"As such\".","section":"Section IV.B"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope. The most pressing issue is data availability: this is fundamentally a dataset paper, and without access to the scores and stimuli, the contribution cannot be validated. The authors should be required to release the data as part of the revision. The statistical and DMOS-validity concerns are secondary but should also be addressed, as they affect the strength of the benchmark claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the UGC-VIDEO database is a well-intentioned addition to the UGC VQA line, and the benchmark numbers are plausible, but as submitted the paper's central artifact—the database and the subjective scores—is not available, so the whole thing rests on trust. The paper does have real strengths: the selection procedure uses a published optimal-sampling method to spread 50 sources across SI/TI/blur in each of four content categories; the encoding setup (two codecs, five QPs) is simple and reproducible in principle; and the subjective testing follows ITU-R BT.500-13 with screening, dummy presentations, and a single calibrated CRT. The benchmark tables are internally consistent and the finding that VMAF tops out around 0.87 SROCC on DMOS and 0.81 on MOS is the kind of number people will quote.\n\nThe soft spots are not subtle. First, there is no URL, repository, or supplementary material. For a dataset paper, data availability is the evidence. Without the per-subject scores and the actual encoded videos, none of the tables can be audited; a transposed column or a different logistic fit would change the headline numbers. That alone makes the submission conditional. Second, the 'source' videos are themselves TikTok uploads that already went through at least one compression round. The paper acknowledges this in the introduction—no pristine source exists in UGC—but then computes DMOS by subtracting scores for these non-pristine references. That is a defensible design choice, but the paper should present MOS as the primary target and treat DMOS with more caution. The category-1/category-2 split (by the 20th percentile of reference MOS) creates a potential range-restriction artifact, and the authors should show whether the VMAF drop in category 1 persists after controlling for reference quality or using MOS directly. Third, the paper reports no confidence intervals or significance tests on differences between correlation coefficients. With 550 clips, some apparent differences between algorithms are noise; a bootstrap or a Williams test would help.\n\nThe citation pattern is fine. The sampling method is borrowed from [13] and cited. No invented entities, no fitted model used to draw conclusions. The central argument—that existing quality measures leave room for improvement on UGC—is consistent with prior work (KoNViD-1k, LIVE-VQC), so if the data see the light, the paper's conclusions will probably survive.\n\nThis paper deserves a serious referee, but only after the authors commit to releasing the data and scores, and to strengthening the statistical reporting. Anyone working on UGC VQA benchmarks will want this database; until it is downloadable, it is just a description.","headline":"A plausible UGC video quality database with a solid setup, but the missing data and scores make the central claim unverifiable as submitted.","tokens_in":9133,"tokens_out":2110,"would_cite":false,"duration_ms":18927,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper creates UGC-VIDEO, a subjectively rated database of 50 TikTok source videos recompressed into 550 clips, and argues that existing quality metrics, especially no-reference ones, correlate only moderately with human opinion on…","keywords":["video quality assessment","user-generated content","subjective quality database","TikTok","H.264","H.265/HEVC","VMAF","no-reference quality metrics"],"falsifier":"Take the same 50 clips, re-encode them from their original pre-upload masters, repeat the subjective test, and check whether VMAF's SROCC of 0.8726 on DMOS and the category-1 degradation pattern reproduce; if the correlations shift materially, the reported gap is partly an artifact of treating pre-compressed uploads as pristine references.","tokens_in":8145,"feed_emoji":"🎬","tokens_out":6398,"duration_ms":57908,"temperature":0.7,"pith_summary":"The paper sets out to measure how well existing video quality algorithms perform on user-generated videos, which lack pristine originals and typically undergo multiple compression stages before being viewed. To do this, it builds UGC-VIDEO, a database of 50 TikTok source clips spanning selfie, indoor, outdoor, and screen-content categories, re-encodes each with H.264 and H.265 at five quantization levels, and collects subjective ratings from 30 viewers. Benchmarking shows that full-reference metrics correlate only moderately with human opinion: VMAF reaches the best SROCC of 0.8726 against DMOS and 0.8141 against MOS, while no-reference metrics such as NIQE, BRISQUE, VIIDEO, and BLIINDS lag well behind. The authors conclude that current objective measures leave substantial room for improvement on UGC content.","feed_headline":"New TikTok video dataset exposes a UGC quality scoring gap","feed_subtitle":"550 recompressed clips rated by humans show the best metric, VMAF, caps at 0.87 SROCC.","key_machinery":"The carrying instrument is the UGC-VIDEO database: 50 TikTok videos sampled uniformly across spatial information, temporal information, and blur using a dataset-shaping formulation, then re-encoded into 10 distorted versions per source with two codecs and five quantization levels. Subjective scores came from single-stimulus ACR-HR sessions with 30 subjects, screened according to a standard ITU-R subject-screening protocol, yielding both MOS and DMOS. The database is what carries the benchmark: eight full-reference or reduced-reference metrics and four no-reference metrics are evaluated against these scores, split by reference-quality level and by content category.","core_discovery":"The central claim is that a realistic UGC video database, built from already-uploaded TikTok videos rather than pristine studio content, exposes a clear gap between current objective quality measures and human perception. The database itself is the discovery instrument: 50 source videos were selected to be nearly uniform in spatial information, temporal information, and blur, then each was recompressed with x264 and x265 at QPs 22, 27, 32, 37, and 42, producing 550 rated sequences. The authors report that full-reference algorithms agree with DMOS only moderately, that performance degrades sharply when the reference video is itself low quality, and that screen-content videos are the hardest category for most algorithms. They also observe that some recompressed clips receive negative DMOS, meaning compression can slightly improve perceived quality by smoothing noise in already-imperfect UGC sources.","pith_inferences":["If the source videos' prior compression histories vary across clips, DMOS blends the reference's own quality with the effect of the new recompression; a follow-up analysis could regress out source MOS and bitrate to isolate the incremental quality loss.","The negative-DMOS observation suggests a testable extension: a UGC quality metric should be allowed to be non-monotonic in bitrate, because re-encoding can act as denoising on already-noisy uploads.","The same uniform-sampling design could be reused to build larger UGC databases with per-clip provenance metadata, letting the field separate codec-induced artifacts from capture and editing artifacts rather than treating every uploaded video as a pristine source.","Benchmarking against both MOS and DMOS on the same data, as this paper does, is a practice worth standardizing, since an algorithm optimized for DMOS may not optimize for absolute perceived quality."],"forward_implications":["VMAF should be treated as the strongest existing reference model on UGC content, so future UGC quality metrics should be compared against VMAF rather than PSNR or SSIM alone.","Screen-content videos are the most difficult category for existing algorithms, so UGC quality assessment likely needs content-aware modeling rather than a single universal metric.","Negative DMOS cases show that recompression can perceptually improve noisy UGC videos, so objective models that assume distortion only degrades quality will misjudge these clips.","Low-quality references degrade the correlation of full-reference metrics, meaning databases built on pristine sources will tend to overstate how well those metrics will work on real user-generated content.","No-reference metrics trained on natural images are poorly suited to diverse UGC, and the measured gap motivates developing blind metrics that account for multi-stage compression and special effects."],"supporting_citations":[{"why":"Supplies the dataset-shaping formulation used to select 50 source videos that are nearly uniform across spatial information, temporal information, and blur features.","marker":"[13]"},{"why":"Supplies the x264 encoder used to generate the H.264 compressed versions of every source video at five QPs.","marker":"[14]"},{"why":"Supplies the x265 encoder used to generate the H.265 compressed versions of every source video at five QPs.","marker":"[15]"},{"why":"Provides the subject-screening and rating methodology that the subjective test follows when computing MOS and rejecting unreliable viewers.","marker":"[16]"},{"why":"Defines the spatial and temporal information features and the single-stimulus ACR-HR protocol used to collect opinion scores.","marker":"[11]"},{"why":"Describes VMAF, the full-reference model whose scores produce the best benchmark correlations in the paper.","marker":"[23]"},{"why":"Supplies the closest prior subjectively annotated UGC-style video database that this work extends by simulating the upload-and-recompress chain.","marker":"[7]"},{"why":"Provides the logistic regression function used to compute PLCC and RMSE values from the subjective scores.","marker":"[24]"}],"fun_headline_variants":["TikTok recompression dataset shows UGC quality metrics fall short","550 human-rated TikTok clips expose gap in video quality metrics","UGC video scoring blind spot: metrics miss real-world compression effects","Negative DMOS: compression can clean noisy user videos, new data show","New benchmark: existing video metrics lag on TikTok-sourced UGC"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the already-compressed TikTok videos treated as source clips can serve as valid references, so the MOS and DMOS differences recorded after recompression measure the added compression's effect rather than each source's unknown prior encoding and editing history.","fun_headline_variants_meta":{"raw":{"variants":["TikTok recompression dataset shows UGC quality metrics fall short","550 human-rated TikTok clips expose gap in video quality metrics","UGC video scoring blind spot: metrics miss real-world compression effects","Negative DMOS: compression can clean noisy user videos, new data show","New benchmark: existing video metrics lag on TikTok-sourced UGC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000851,"raw_usage":{"total_tokens":3667,"prompt_tokens":879,"completion_tokens":2788,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":2697}},"tokens_in":495,"tokens_out":2788,"duration_ms":18094,"temperature":1.0,"reasoning_tokens":2697,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:12:35.144053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 50 clips, re-encode them from their original pre-upload masters, repeat the subjective test, and check whether VMAF's SROCC of 0.8726 on DMOS and the category-1 degradation pattern reproduce; if the correlations shift materially, the reported gap is partly an artifact of treating pre-compressed uploads as pristine references.","supporting_citations":[{"cited_title":"Shaping datasets: Optimal data selection for speciﬁc target distributions across dimensions,","cited_arxiv_id":null,"evidence_quote":"Supplies the dataset-shaping formulation used to select 50 source videos that are nearly uniform across spatial information, temporal information, and blur features."},{"cited_title":"x264: A high performance H.264/A VC encoder,","cited_arxiv_id":null,"evidence_quote":"Supplies the x264 encoder used to generate the H.264 compressed versions of every source video at five QPs."},{"cited_title":"x265 HEVC Encoder/H.265 Video Codec,","cited_arxiv_id":null,"evidence_quote":"Supplies the x265 encoder used to generate the H.265 compressed versions of every source video at five QPs."},{"cited_title":"Methodology for the subjective assessment of the quality of television pictures,","cited_arxiv_id":null,"evidence_quote":"Provides the subject-screening and rating methodology that the subjective test follows when computing MOS and rejecting unreliable viewers."},{"cited_title":"P. 910: Subjective video quality assessment methods for multimedia applications,","cited_arxiv_id":null,"evidence_quote":"Defines the spatial and temporal information features and the single-stimulus ACR-HR protocol used to collect opinion scores."},{"cited_title":"Challenges in cloud based ingest and encoding for high quality streaming media,","cited_arxiv_id":null,"evidence_quote":"Describes VMAF, the full-reference model whose scores produce the best benchmark correlations in the paper."},{"cited_title":"The Konstanz natural video database (KoNViD-1k),","cited_arxiv_id":null,"evidence_quote":"Supplies the closest prior subjectively annotated UGC-style video database that this work extends by simulating the upload-and-recompress chain."},{"cited_title":"A statistical evaluation of recent full reference image quality assessment algorithms,","cited_arxiv_id":null,"evidence_quote":"Provides the logistic regression function used to compute PLCC and RMSE values from the subjective scores."}],"review_version":1}