Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

UGC-VIDEO: perceptual quality assessment of user-generated videos

T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper creates UGC-VIDEO, a subjectively rated database of 50 TikTok source videos recompressed into 550 clips, and argues that existing quality metrics, especially no-reference ones, correlate only moderately with human opinion on…

desk verdict A plausible UGC video quality database with a solid setup, but the missing data and scores make the central claim unverifiable as submitted. read the letter →

arxiv 1908.11517 v2 pith:ITIWY3QE submitted 2019-08-30 cs.MM eess.IV

classification cs.MMeess.IV
keywords videoqualityassessmentuser-generatedcontentsubjectivedatabaseTikTokH.264H.265/HEVCVMAFno-referencemetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure how well existing video quality algorithms perform on user-generated videos, which lack pristine originals and typically undergo multiple compression stages before being viewed. To do this, it builds UGC-VIDEO, a database of 50 TikTok source clips spanning selfie, indoor, outdoor, and screen-content categories, re-encodes each with H.264 and H.265 at five quantization levels, and collects subjective ratings from 30 viewers. Benchmarking shows that full-reference metrics correlate only moderately with human opinion: VMAF reaches the best SROCC of 0.8726 against DMOS and 0.8141 against MOS, while no-reference metrics such as NIQE, BRISQUE, VIIDEO, and BLIINDS lag well behind. The authors conclude that current objective measures leave substantial room for improvement on UGC content.

What carries the argument

The carrying instrument is the UGC-VIDEO database: 50 TikTok videos sampled uniformly across spatial information, temporal information, and blur using a dataset-shaping formulation, then re-encoded into 10 distorted versions per source with two codecs and five quantization levels. Subjective scores came from single-stimulus ACR-HR sessions with 30 subjects, screened according to a standard ITU-R subject-screening protocol, yielding both MOS and DMOS. The database is what carries the benchmark: eight full-reference or reduced-reference metrics and four no-reference metrics are evaluated against these scores, split by reference-quality level and by content category.

What would settle it

Take the same 50 clips, re-encode them from their original pre-upload masters, repeat the subjective test, and check whether VMAF's SROCC of 0.8726 on DMOS and the category-1 degradation pattern reproduce; if the correlations shift materially, the reported gap is partly an artifact of treating pre-compressed uploads as pristine references.

Watch

Extended reading notes

Core claim

The central claim is that a realistic UGC video database, built from already-uploaded TikTok videos rather than pristine studio content, exposes a clear gap between current objective quality measures and human perception. The database itself is the discovery instrument: 50 source videos were selected to be nearly uniform in spatial information, temporal information, and blur, then each was recompressed with x264 and x265 at QPs 22, 27, 32, 37, and 42, producing 550 rated sequences. The authors report that full-reference algorithms agree with DMOS only moderately, that performance degrades sharply when the reference video is itself low quality, and that screen-content videos are the hardest category for most algorithms. They also observe that some recompressed clips receive negative DMOS, meaning compression can slightly improve perceived quality by smoothing noise in already-imperfect UGC sources.

Load-bearing premise

The load-bearing premise is that the already-compressed TikTok videos treated as source clips can serve as valid references, so the MOS and DMOS differences recorded after recompression measure the added compression's effect rather than each source's unknown prior encoding and editing history.

Editorial extensions

If this is right

  • VMAF should be treated as the strongest existing reference model on UGC content, so future UGC quality metrics should be compared against VMAF rather than PSNR or SSIM alone.
  • Screen-content videos are the most difficult category for existing algorithms, so UGC quality assessment likely needs content-aware modeling rather than a single universal metric.
  • Negative DMOS cases show that recompression can perceptually improve noisy UGC videos, so objective models that assume distortion only degrades quality will misjudge these clips.
  • Low-quality references degrade the correlation of full-reference metrics, meaning databases built on pristine sources will tend to overstate how well those metrics will work on real user-generated content.
  • No-reference metrics trained on natural images are poorly suited to diverse UGC, and the measured gap motivates developing blind metrics that account for multi-stage compression and special effects.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the source videos' prior compression histories vary across clips, DMOS blends the reference's own quality with the effect of the new recompression; a follow-up analysis could regress out source MOS and bitrate to isolate the incremental quality loss.
  • The negative-DMOS observation suggests a testable extension: a UGC quality metric should be allowed to be non-monotonic in bitrate, because re-encoding can act as denoising on already-noisy uploads.
  • The same uniform-sampling design could be reused to build larger UGC databases with per-clip provenance metadata, letting the field separate codec-induced artifacts from capture and editing artifacts rather than treating every uploaded video as a pristine source.
  • Benchmarking against both MOS and DMOS on the same data, as this paper does, is a practice worth standardizing, since an algorithm optimized for DMOS may not optimize for absolute perceived quality.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces UGC-VIDEO, a new subjectively annotated database for user-generated video quality assessment, consisting of 50 source videos collected from TikTok (covering selfie, indoor, outdoor, and screen content), each further compressed with H.264 and H.265 at five quantization levels, yielding 550 sequences including the sources. Subjective ratings were collected from 30 subjects using the ACR-HR paradigm, and MOS and DMOS scores were computed. The authors benchmark twelve objective quality metrics using SROCC, PLCC, and RMSE, and analyze performance by reference quality (category 1 vs. 2) and by content category. The main reported finding is that existing metrics, including VMAF, correlate only moderately with human opinion on this database, with VMAF achieving the best overall SROCC of 0.8415 on DMOS and 0.8141 on MOS.

Significance. If the database and scores are made available, this would be a useful resource for the VQA community, since existing databases mostly rely on pristine sources or synthetic distortions, whereas UGC-VIDEO reflects the realistic multi-stage compression scenario on a hosting platform. The content-aware sampling strategy based on SI, TI, and blur (Section II.C) is a methodological strength, as is the use of standard subjective-testing procedures (ITU-R BT.500-13). The benchmark results provide a baseline for future UGC VQA research and align with the growing interest in no-reference and reduced-reference quality assessment. However, as a dataset paper, the absence of any data availability statement is a serious limitation: the core contribution—the subjective scores and the video stimuli—cannot be independently verified or reused.

major comments (4)
  1. [Sections II and IV (overall)] The manuscript provides no URL, repository, or supplementary material for the UGC-VIDEO database, the subjective scores (per-subject ratings, MOS, DMOS), or the source/processed video files. Because the central claim of the paper is the creation of a new subjective database and the benchmark results derived from it, the absence of public access makes Tables I and II unverifiable. A dataset paper must make the data available, or at minimum release the MOS/DMOS scores and a description of how to obtain the videos; without this, the contribution is an assertion rather than a resource. This should be fixed by providing a stable download link and data format description.
  2. [Section IV.A, Tables I and II] The paper claims a "significant performance degradation" on low-quality reference videos (category 1) and draws conclusions about the relative performance of algorithms (e.g., VMAF being best), but no confidence intervals, bootstrap estimates, or significance tests are reported. With 500 compressed videos and correlations in the range 0.7–0.9, differences such as SROCC 0.8415 (VMAF) vs. 0.8443 (ViS3) may not be statistically meaningful. The authors should add significance testing (e.g., bootstrap on subjects or a Steiger test for correlated correlations) to support the benchmark claims.
  3. [Section III.B and Section IV.A] The DMOS is computed as the difference between the source and distorted video scores, but the source videos are already-compressed TikTok uploads with widely varying quality (as acknowledged in Section IV.A and Fig. 3). This makes the reference non-pristine, so DMOS conflates the recompression effect with the intrinsic quality of the source. The subsequent split into category 1 (low-quality source) and category 2 (higher-quality source) conditions on source MOS, which can introduce range-restriction artifacts: lower correlations in category 1 might reflect a narrower DMOS range or a ceiling/floor effect rather than a genuine failure of the objective metrics. The authors partially acknowledge the issue, but they do not control for it in the benchmark. Please report MOS-based results as the primary analysis or include partial correlations controlling for source MOS, and discuss how the DMOS benchmark should be interpreted given the non-pristine reference.
  4. [Section IV.A] The text states that "110 compressed videos with low quality source are classified into the category 1" and the remaining 390 into category 2, based on a 20th percentile split of source MOS. Since there are 50 source videos and 10 compressed versions each, a 20th percentile split should select 10 sources and hence 100 compressed videos, not 110 (which would correspond to 11 sources). The paper should clarify the exact computation of the percentile threshold and how ties were handled; this is needed for reproducibility of the category-level results in Table I.
minor comments (4)
  1. [Section II.D] The sentence "Considering the fact that our primary goal of investigating the quality assessment of UGC videos for improving the video coding/transcoding performance" is grammatically incomplete; please rephrase.
  2. [Table II] The column heading "Ourdoor" should be corrected to "Outdoor".
  3. [Section III.A] There is a typo: "dummpy presentations" should be "dummy presentations".
  4. [Section IV.B] There is a typo: "As suah" should be "As such".

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's claims are a new subjective database plus a benchmark of existing quality measures, neither of which is derived from its own inputs by construction.

full rationale

The paper's central claims are (1) construction of the UGC-VIDEO database with subjective scores and (2) a benchmark evaluation of existing full-, reduced-, and no-reference quality measures against MOS and DMOS. No author-derived metric or fitted model is used to generate a prediction; each objective algorithm (PSNR, SSIM, VMAF, BRISQUE, etc.) is an external, pre-existing measure whose scores are correlated with the subjectively collected ratings. The logistic regression in Eq. (5) is a standard monotonic mapping used only to compute PLCC and RMSE, and it does not enter the SROCC results or the comparative conclusions; correlating any external metric with human labels is not a derivation of that metric from the labels. DMOS is computed as a difference between source and distorted ratings, and the source videos are themselves previously compressed TikTok uploads, which complicates interpretation of the reference quality and of the category-1/2 split in Section IV.A; however, this is a validity and design weakness, not circularity, because the benchmark conclusion does not assume the sources are pristine and does not reduce to the subjective scores by construction. The paper also provides no public release link for the database or scores, which is a verifiability limitation rather than evidence of circular derivation. I find no step in which a claimed result is equivalent to its input by definition or by self-citation.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The conclusions rest on several clearly stated domain assumptions about representativeness and reference validity, none of which are independently validated. No free parameters are fitted into the central benchmark claim; the logistic mapping in Eq. 5 is standard evaluation preprocessing and is not reported as a scientific result.

assumptions (4)
  • domain assumption TikTok videos, after the described filtering, are representative of user-generated video content at large.
    The database is claimed to represent typical UGC, but the source pool comes from a single platform with unknown download and re-encoding history. This enters in Section II.A.
  • domain assumption Single-stimulus ACR-HR ratings from 30 subjects on one CRT monitor yield valid MOS and DMOS values.
    The procedure follows ITU-R BT.500-13, but no inter-laboratory validation, display diversity, or subject-retest reliability is reported. This enters in Section III.A.
  • domain assumption Already-compressed TikTok clips can serve as a valid reference for differential (DMOS) quality assessment.
    DMOS is computed as the difference from these references even though they are not pristine and have unknown prior compression histories. This enters in Section III.B and Table I.
  • domain assumption Uniform coverage of the SI, TI, and blur feature space (via Eq. 3) yields a representative content set for perceptual quality.
    The minimization in Eq. 3 balances three hand-chosen low-level features; the link between that feature coverage and perceptual-quality diversity is assumed. This enters in Section II.C.

how reviews work

0 comments
Cite this review

Pith. "Pith review of UGC-VIDEO: perceptual quality assessment of user-generated videos." pith.science (2026). https://pith.science/paper/ITIWY3QE

@misc{pith2026190811517,
  author       = {Pith},
  title        = {Pith review of: UGC-VIDEO: perceptual quality assessment of user-generated videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ITIWY3QE}},
  note         = {Machine review of arXiv:1908.11517}
}
read the original abstract

Recent years have witnessed an ever-expandingvolume of user-generated content (UGC) videos available on the Internet. Nevertheless, progress on perceptual quality assessmentof UGC videos still remains quite limited. There are many distinguished characteristics of UGC videos in the complete video production and delivery chain, and one important property closely relevant to video quality is that there does not exist the pristine source after they are uploaded to the hosting platform,such that they often undergo multiple compression stages before ultimately viewed. To facilitate the UGC video quality assessment,we created a UGC video perceptual quality assessment database. It contains 50 source videos collected from TikTok with diverse content, along with multiple distortion versions generated bythe compression with different quantization levels and coding standards. Subjective quality assessment was conducted to evaluate the video quality. Furthermore, we benchmark the database using existing quality assessment algorithms, and potential roomis observed to future improve the accuracy of UGC video quality measures.

Figures

Figures reproduced from arXiv: 1908.11517 by the authors.

Figure 2
Figure 2. Selected source video sequences. The four images from left to right are [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. The relationship between DMOS and ∆R for each category. assessment BRISQUE [25], [26], NIQE [27], VIIDEO [28] and BLIINDS [29] were also evaluated. The SROCC, PLCC and RMSE results of four separate categories and the whole database are shown in Table II. We can find that the existing algorithms may not provide reasonably accurate predictions on the UGC videos. For most algorithms, they perform the worst on screen co… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LEHA-CVQAD: Dataset To Enable Generalized Video Quality Assessment of Compression Artifacts

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LEHA-CVQAD is a 6,240-clip compressed video dataset with fused MOS and pairwise labels, a hidden test set, and a new Rate-Distortion Alignment Error metric.

Reference graph

Works this paper leans on

29 extracted references · 26 canonical work pages · cited by 1 Pith paper

  1. [13]

    Shaping datasets: Optimal data selection for specific target distributions across dimensions,

    Vassilios V onikakis, Ramanathan Subramanian, and Stefan Winkler, “Shaping datasets: Optimal data selection for specific target distributions across dimensions,” in 2016 IEEE International Conference on Image Processing (ICIP). IEEE, 2016, pp. 3753–3757

  2. [1]

    Study of subjective and objective quality assessment of video,

    Kalpana Seshadrinathan, Rajiv Soundararajan, Alan Conrad Bovik, and Lawrence K Cormack, “Study of subjective and objective quality assessment of video,” IEEE transactions on Image Processing , vol. 19, no. 6, pp. 1427–1441, 2010

  3. [2]

    A subjective study to evaluate video quality assessment algorithms,

    Kalpana Seshadrinathan, Rajiv Soundararajan, Alan C Bovik, and Lawrence K Cormack, “A subjective study to evaluate video quality assessment algorithms,” in Human vision and electronic imaging XV . International Society for Optics and Photonics, 2010, vol. 7527, p. 75270H

  4. [3]

    Video quality assessment on mobile devices: Subjective, behavioral and objective studies,

    Anush Krishna Moorthy, Lark Kwon Choi, Alan Conrad Bovik, and Gustavo De Veciana, “Video quality assessment on mobile devices: Subjective, behavioral and objective studies,” IEEE Journal of Selected Topics in Signal Processing , vol. 6, no. 6, pp. 652–671, 2012

  5. [4]

    Subjective analysis of video quality on mobile devices,

    Anush K Moorthy, Lark K Choi, Gustavo De Veciana, and Alan C Bovik, “Subjective analysis of video quality on mobile devices,” in Sixth International Workshop on Video Processing and Quality Metrics for Consumer Electronics (VPQM), Scottsdale, Arizona . Citeseer, 2012

  6. [5]

    MCL-JCV: a JND-based H.264/A VC video quality as- sessment dataset,

    Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavounidis, Anne Aaron, and C-C Jay Kuo, “MCL-JCV: a JND-based H.264/A VC video quality as- sessment dataset,” in Image Processing (ICIP), 2016 IEEE International Conference on. IEEE, 2016, pp. 1509–1513

  7. [6]

    In-capture mobile video distor- tions: A study of subjective behavior and objective algorithms,

    Deepti Ghadiyaram, Janice Pan, Alan C Bovik, Anush Krishna Moorthy, Prasanjit Panda, and Kai-Chieh Yang, “In-capture mobile video distor- tions: A study of subjective behavior and objective algorithms,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 9, pp. 2061–2077, 2017

  8. [7]

    The Konstanz natural video database (KoNViD-1k),

    Vlad Hosu, Franz Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tam ´as Szir´anyi, Shujun Li, and Dietmar Saupe, “The Konstanz natural video database (KoNViD-1k),” in 2017 Ninth International Conference on Quality of Multimedia Experience (QoMEX) . IEEE, 2017, pp. 1–6

Show all 29 references
  1. [8]

    Quality assessment of images undergoing multiple distortion stages,

    Shahrukh Athar, Abdul Rehman, and Zhou Wang, “Quality assessment of images undergoing multiple distortion stages,” in Image Processing (ICIP), 2017 IEEE International Conference on . IEEE, 2017, pp. 3175– 3179

  2. [9]

    Overview of the H.264/A VC video coding standard,

    Thomas Wiegand, Gary J Sullivan, Gisle Bjontegaard, and Ajay Luthra, “Overview of the H.264/A VC video coding standard,” IEEE Transac- tions on circuits and systems for video technology , vol. 13, no. 7, pp. 560–576, 2003

  3. [10]

    Overview of the high efficiency video coding (HEVC) standard,

    Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand, “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on circuits and systems for video technology , vol. 22, no. 12, pp. 1649–1668, 2012

  4. [11]

    P. 910: Subjective video quality assessment methods for multimedia applications,

    ITUT Rec, “P. 910: Subjective video quality assessment methods for multimedia applications,” International Telecommunication Union, Geneva, vol. 2, 2008

  5. [12]

    A no-reference image blur metric based on the cumulative probability of blur detection (CPBD),

    Niranjan D Narvekar and Lina J Karam, “A no-reference image blur metric based on the cumulative probability of blur detection (CPBD),” IEEE Transactions on Image Processing, vol. 20, no. 9, pp. 2678–2683, 2011

  6. [14]

    x264: A high performance H.264/A VC encoder,

    Loren Merritt and Rahul Vanam, “x264: A high performance H.264/A VC encoder,” [online] http://neuron2. net/library/avc/overview x264 v8 5. pdf, 2006

  7. [15]

    x265 HEVC Encoder/H.265 Video Codec,

    “x265 HEVC Encoder/H.265 Video Codec,” http://x265.org/

  8. [16]

    Methodology for the subjective assessment of the quality of television pictures,

    ITU-R Rec. BT. 500-13, “Methodology for the subjective assessment of the quality of television pictures,” 2012

  9. [17]

    Subjective video quality assessment methods for multimedia applications,

    P ITU-T RECOMMENDATION, “Subjective video quality assessment methods for multimedia applications,” International telecommunication union, 1999

  10. [18]

    Image quality assessment: from error visibility to structural similarity,

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004

  11. [19]

    Image information and visual quality,

    Hamid R Sheikh and Alan C Bovik, “Image information and visual quality,” in 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing . IEEE, 2004, vol. 3, pp. iii–709

  12. [20]

    Multiscale structural similarity for image quality assessment,

    Zhou Wang, Eero P Simoncelli, and Alan C Bovik, “Multiscale structural similarity for image quality assessment,” inThe Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003 . Ieee, 2003, vol. 2, pp. 1398–1402

  13. [21]

    SpEED-QA: Spatial efficient entropic differencing for image and video quality,

    Christos G Bampis, Praful Gupta, Rajiv Soundararajan, and Alan C Bovik, “SpEED-QA: Spatial efficient entropic differencing for image and video quality,” IEEE Signal Processing Letters , vol. 24, no. 9, pp. 1333–1337, 2017

  14. [22]

    ViS3: an algorithm for video quality assessment via analysis of spatial and spatiotemporal slices,

    Phong V Vu and Damon M Chandler, “ViS3: an algorithm for video quality assessment via analysis of spatial and spatiotemporal slices,” Journal of Electronic Imaging , vol. 23, no. 1, pp. 013016, 2014

  15. [23]

    Challenges in cloud based ingest and encoding for high quality streaming media,

    Anne Aaron, Zhi Li, Megha Manohara, Joe Yuchieh Lin, Eddy Chi- Hao Wu, and C-C Jay Kuo, “Challenges in cloud based ingest and encoding for high quality streaming media,” in 2015 IEEE International Conference on Image Processing (ICIP) . IEEE, 2015, pp. 1732–1736

  16. [24]

    A statistical evaluation of recent full reference image quality assessment algorithms,

    Hamid R Sheikh, Muhammad F Sabir, and Alan C Bovik, “A statistical evaluation of recent full reference image quality assessment algorithms,” IEEE Transactions on image processing, vol. 15, no. 11, pp. 3440–3451, 2006

  17. [25]

    Referenceless image spatial quality evaluation engine,

    A Mittal, AK Moorthy, and AC Bovik, “Referenceless image spatial quality evaluation engine,” in 45th Asilomar Conference on Signals, Systems and Computers , 2011, vol. 38, pp. 53–54

  18. [26]

    No- reference image quality assessment in the spatial domain,

    Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik, “No- reference image quality assessment in the spatial domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012

  19. [27]

    Making a “Completely Blind

    Anish Mittal, Rajiv Soundararajan, and Alan C Bovik, “Making a “Completely Blind” Image Quality Analyzer,” IEEE Signal Process. Lett., vol. 20, no. 3, pp. 209–212, 2013

  20. [28]

    A completely blind video integrity oracle,

    Anish Mittal, Michele A Saad, and Alan C Bovik, “A completely blind video integrity oracle,” IEEE Transactions on Image Processing , vol. 25, no. 1, pp. 289–300, 2015

  21. [29]

    Blind prediction of natural video quality,

    Michele A Saad, Alan C Bovik, and Christophe Charrier, “Blind prediction of natural video quality,” IEEE Transactions on Image Processing, vol. 23, no. 3, pp. 1352–1365, 2014

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.