REVIEW 3 major objections 5 minor 2 cited by
HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper introduces a large-scale paired HDR/SDR video quality dataset and shows HDR's benefit depends on content brightness, texture, and motion.
desk verdict A genuinely useful new HDR/SDR pairwise dataset, but the 'general HDR TV' claim needs per-TV analysis before the headline conclusions can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the HDRSDR-VQA dataset itself: 54 source clips (with measured spatial, temporal, and colorfulness statistics plus luminance distributions) rendered in HDR10 and SDR, each at nine distortion levels, judged by pairwise comparison. The measurement engine is active sampling (ASAP) to pick informative pairs, followed by maximum-likelihood scaling (pwcmp) into Just-Objectionable-Differences, where one JOD means 75 percent of observers prefer one version; the JOD scale is what lets HDR and SDR be placed on one quality axis.
What would settle it
Compute the HDR-minus-SDR JOD difference separately for each of the six televisions. If the crossover pattern disappears on the brightest TV, or if the dimmest TVs reverse the bright-content HDR advantage, then the pooled conclusion that HDR wins on bright scenes is an artifact of averaging different displays rather than a general property.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a content-dependent preference law: HDR10 outperforms SDR when the source has highlight pixels beyond the SDR range, saturated colors outside the sRGB gamut, and fine texture that higher bit depth can preserve, but this advantage disappears or inverts in scenes with little brightness or color range and in high-motion content, where SDR's lower bit-depth demand is more robust to compression. The direction and size of the preference also depend on bitrate and resolution: HDR's gap widens at high bitrates and can turn negative at low bitrates.
Load-bearing premise
The pooled JOD analysis assumes that viewer preferences from six TVs spanning 252 to 2,564 cd/m² peak brightness can be combined into one scale that represents the general HDR viewing experience.
Editorial extensions
If this is right
- Because all videos have paired HDR and SDR versions on the same JOD scale, objective VQA models can be tested against a direct format-preference signal rather than separate quality scores.
- Content-adaptive streaming can use content statistics such as brightness, gamut, texture, and motion to decide when HDR is worth the bitrate, since the dataset maps where HDR's advantage appears.
- Rate-distortion behavior differs by format: at low bitrates HDR can fall behind SDR on unfavorable content, so a fixed HDR-first ladder is not optimal.
- The anchor contents link the new scale to an earlier HDR/SDR comparison database, making scores from both studies comparable.
- The public subset of the dataset lets other labs reproduce the scaling and extend the analysis to new models.
Reading between the lines
- If the per-TV data were released, one could test whether peak brightness moderates the preference: HDR's bright-scene advantage might shrink or reverse on the dimmest displays, which would qualify the pooled conclusion.
- The SDR baseline depends on the chosen conversion and grading, so the comparison is one sample of SDR rather than SDR-in-general; a different SDR master could shift the bitrates where the crossover occurs.
- The reported content statistics (SI, TI, CF, and luminance extremes) could train a simple predictor of format preference, generalizing the four-case analysis to unseen content.
- Because active sampling prioritizes ranking information, the resulting scale may under-represent rare but strong disagreements; a re-analysis by viewer or device group could reveal whether the HDR preference on bright scenes is universal or driven by a subset.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HDRSDR-VQA, a subjective video quality dataset built for direct HDR-versus-SDR comparison. The dataset contains 960 videos from 54 source sequences, each source rendered in HDR10 and SDR at nine quality levels (one reference and eight distorted variants), and the authors report a subjective study with 145 participants on six consumer HDR-capable televisions, collecting over 22,000 pairwise comparisons that are scaled to JOD scores using the pwcmp algorithm. The paper presents four content-level case studies and concludes that HDR generally provides perceptual advantages for content with rich textures and high brightness or color gamut, but that these advantages diminish or reverse in low-dynamic-range or motion-intensive scenes. The authors position the dataset as a reusable benchmark for VQA, adaptive streaming, and perceptual model development.
Significance. If the dataset is sound, it fills a real gap: prior VQA datasets typically evaluate only one dynamic-range format, and the few HDR/SDR comparison studies use limited display conditions or absolute category rating rather than pairwise comparison. The scale of the study (960 videos, six TVs, 145 participants) is substantial, and the concrete protocol details, including the use of ASAP active sampling and pwcmp scaling, are credible. The count arithmetic is internally consistent: 51 non-anchor sources times 18 versions plus 3 anchors times 14 versions gives 960 videos, and 145 participants times 160 comparisons equals 23,200 comparisons, which is consistent with the claimed 'over 22,000'. The main scientific value would be in enabling display-aware study of when HDR is preferred, assuming the pooled JOD analysis is not hiding strong display dependence. The public release of scores and stimuli would also make this a practical resource for the community. The key weakness is external validity: the paper claims results representative of a 'general HDR TV' without reporting or modeling per-display variation.
major comments (3)
- [IV.A, Table II] The pooled JOD analysis combines comparisons from six televisions whose HDR peak brightness ranges from 252 to 2564 cd/m2 and whose Rec.2020 coverage ranges from 23.8% to 54.8%. The text in Section IV.A explicitly aims for results 'representative of a general HDR TV', but the paper reports no per-TV JOD scales, no display-as-covariate model, and no interaction analysis. If HDR versus SDR preference reverses on low-brightness or low-gamut displays such as the CU8000 or Vizio M6, the aggregate conclusions in Section VI about when HDR wins would be an artifact of the particular mix of TVs. The central claim of the paper therefore needs a per-TV analysis or a statistical model treating the display as a random effect, at minimum to show that the HDR-SDR JOD differences do not change sign or significance across the six displays.
- [V, Fig. 5] The concluding generalization that HDR advantages 'diminish or even reverse in low dynamic or motion-intensive scenes' is supported only by four selected content examples. No aggregate statistics are reported across the 54 contents, such as the distribution of HDR-minus-SDR JOD differences at each bitrate, the proportion of contents for which the difference is statistically distinguishable from zero, or correlations with SI, TI, colorfulness, and luminance metrics. For instance, the text describes the '28 Swan' HDR-SDR difference as negligible without reporting confidence intervals, so the reader cannot tell whether the claimed reversals are real effects or noise. The paper should provide confidence intervals for the JOD estimates and a content-level analysis linking content attributes to HDR-SDR preference.
- [Abstract and III.A-Anchor Contents] The three anchor contents are taken from the authors' own LIVE-HDRvsSDR database and are used to calibrate that database with the new dataset, but the paper does not report any result of this calibration, such as consistency checks between overlapping conditions, nor does it state how the anchor contents' 14-variation structure affects the pooled scaling. Since anchor contents have fewer versions than the other 51 sequences, the effective coverage is imbalanced, and the JOD scale may be less precisely determined for those contents. Please provide the anchor-calibration outcome and explain whether the pooled JOD scale is robust to excluding or reweighting the anchor contents.
minor comments (5)
- [III.A] There are typos and spacing issues in the text, including 'sourcd' instead of 'sourced' and inconsistent 'V oD' spacing; these should be corrected.
- [I and Table I] The text says 'eight distinct levels of distortion' and then Table I is described as detailing 'the specific categories of distortions', but Table I only lists resolutions and bitrates; the table caption and surrounding text should be made consistent.
- [IV.B] The paper reports 145 participants but the listed gender counts (53 female, 91 male, one undisclosed) sum to 145, while the per-TV counts (19, 26, 21, 26, 26, 27) also sum to 145; this is internally consistent, but the 'over 22,000 comparisons' statement could be made exact (160 per participant times 145 participants gives 23,200).
- [IV.B] One color-deficient participant was retained in the study 'in line with our practice of accommodating diverse participants'; since colorfulness is discussed as a factor in HDR preference, please clarify whether this participant's data were removed or separately analyzed in the scaling.
- [III.C] The conversion to SDR is described only as using NBCU LUTs for non-VoD content and professional mastering for VoD content; please state whether the same LUT was applied to all source types and whether any tone mapping or gamut mapping was applied to the HDR versions for display on each TV, since this affects interpretation of the format comparison.
Circularity Check
No circularity: the dataset construction and JOD scaling are self-contained; the prior-database anchors are calibration, not a load-bearing derivation.
full rationale
This is a dataset-construction paper, not a derivation. The 960 videos are generated from 54 source sequences with explicitly specified encodings, and the 22,000+ pairwise comparisons are collected from 145 participants on six televisions. The transformation from pairwise choices to JOD scores uses the external pwcmp maximum-likelihood tool [15]; no parameter is fitted to a target and then renamed as a prediction. The conclusions in Section V about when HDR is preferred are descriptive analyses of the collected JOD values, not outputs of a model whose inputs already contain those conclusions. The only self-reference is the inclusion of three anchor contents from the authors' earlier LIVE-HDRvsSDR database [9], described as 'calibrate and integrate data'; this is a transparency measure for scale alignment, and subjective judgments for those anchors are still collected in the present study. The cited prior datasets, including LIVE-HDR [8], are independent external resources; none is invoked as a uniqueness theorem or as a premise that forces the reported preferences. The skeptical concern about pooling across televisions with 252-2564 cd/m² peak brightness is an external-validity limitation, not a circularity, because the JOD scale is not defined in terms of the television mix in a way that makes the conclusions true by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Pairwise comparison data scaled by pwcmp yields reliable interval-scale JOD scores.
- domain assumption Scores from six TVs with different peak brightness can be pooled into one scale.
- domain assumption NBCU LUT conversion for non-VOD content and Amazon Studios grading for VOD produce comparable SDR versions.
- domain assumption The ASAP active sampling strategy produces unbiased pairwise comparisons.
Cite this review
Pith. "Pith review of HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation." pith.science (2026). https://pith.science/paper/LZHRM2JW
@misc{pith2026250521831,
author = {Pith},
title = {Pith review of: HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/LZHRM2JW}},
note = {Machine review of arXiv:2505.21831}
}
read the original abstract
We introduce HDRSDR-VQA, a large-scale video quality assessment dataset designed to facilitate comparative analysis between High Dynamic Range (HDR) and Standard Dynamic Range (SDR) content under realistic viewing conditions. The dataset comprises 960 videos generated from 54 diverse source sequences, each presented in both HDR and SDR formats across nine distortion levels. To obtain reliable perceptual quality scores, we conducted a comprehensive subjective study involving 145 participants and six consumer-grade HDR-capable televisions. A total of over 22,000 pairwise comparisons were collected and scaled into Just-Objectionable-Difference (JOD) scores. Unlike prior datasets that focus on a single dynamic range format or use limited evaluation protocols, HDRSDR-VQA enables direct content-level comparison between HDR and SDR versions, supporting detailed investigations into when and why one format is preferred over the other. The open-sourced part of the dataset is publicly available to support further research in video quality assessment, content-adaptive streaming, and perceptual model development.
Figures
Forward citations
Cited by 2 Pith papers
-
Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions
A 300+ device crowd-sourced VQA dataset plus Blade-Chest aggregation and a condition-adaptation MLP let standard metrics predict quality orderings under real mobile viewing conditions far better than unadapted baselines.
-
Learning Quality from Complexity and Structure: A Feature-Fused XGBoost Model for Video Quality Assessment
A concatenation of VCA complexity residuals and SSIM fed into an XGBoost regressor reaches PLCC 0.787 on the VQA Grand Challenge test set, though the method is not truly reduced-reference.
Reference graph
Works this paper leans on
-
[1]
Study of subjective and objective quality assessment of video,
K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack, “Study of subjective and objective quality assessment of video,”IEEE transactions on Image Processing, vol. 19, no. 6, pp. 1427–1441, 2010
work page 2010
-
[2]
Video quality assessment on mobile devices: Subjective, behavioral and ob- jective studies,
A. K. Moorthy, L. K. Choi, A. C. Bovik, and G. De Veciana, “Video quality assessment on mobile devices: Subjective, behavioral and ob- jective studies,”IEEE Journal of Selected Topics in Signal Processing, vol. 6, no. 6, pp. 652–671, 2012
work page 2012
-
[3]
M. Azimi, A. Banitalebi-Dehkordi, Y . Dong, M. T. Pourazad, and P. Nasiopoulos, “Evaluating the performance of existing full-reference quality metrics on high dynamic range (hdr) video content,”arXiv preprint arXiv:1803.04815, 2018
work page Pith review arXiv 2018
-
[4]
Hdr video quality assessment: Perceptual evaluation of compressed hdr video,
X. Pan, J. Zhang, S. Wang, S. Wang, Y . Zhou, W. Ding, and Y . Yang, “Hdr video quality assessment: Perceptual evaluation of compressed hdr video,”Journal of Visual Communication and Image Representation, vol. 57, pp. 76–83, 2018
work page 2018
-
[5]
Verification test report for hdr/wcg video coding using hevc main 10 profile,
V . Baroncini, K. Andersson, A. Ramasubramonian, and G. Sullivan, “Verification test report for hdr/wcg video coding using hevc main 10 profile,” inProc. JCTVC-X1018 24th JCT-VC Meeting, 2016, pp. 293– 303
work page 2016
-
[6]
Subjective and objective evaluation of hdr video compression,
M. Rerabek, P. Hanhart, P. Korshunov, and T. Ebrahimi, “Subjective and objective evaluation of hdr video compression,” in9th International Workshop on Video Processing and Quality Metrics for Consumer Electronics (VPQM), 2015
work page 2015
-
[7]
Perceptual quality assessment of uhd-hdr-wcg videos,
S. Athar, T. Costa, K. Zeng, and Z. Wang, “Perceptual quality assessment of uhd-hdr-wcg videos,” in2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019, pp. 1740–1744
work page 2019
-
[8]
A study of subjective and objec- tive quality assessment of hdr videos,
Z. Shang, J. P. Ebenezer, A. K. Venkataramanan, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “A study of subjective and objec- tive quality assessment of hdr videos,”IEEE Transactions on Image Processing, vol. 33, pp. 42–57, 2023
work page 2023
Show all 15 references
-
[9]
Hdr or sdr? a subjective and objective study of scaled and compressed videos,
J. P. Ebenezer, Z. Shang, Y . Chen, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Hdr or sdr? a subjective and objective study of scaled and compressed videos,”IEEE Transactions on Image Processing, 2024
2024
-
[10]
Avt-vqdb- uhd-2-hdr: An open 8k hdr source dataset for video quality research,
D. Keller, T. Goebel, V . Siebenkees, J. Prenzel, and A. Raake, “Avt-vqdb- uhd-2-hdr: An open 8k hdr source dataset for video quality research,” in2024 16th International Conference on Quality of Multimedia Expe- rience (QoMEX), 2024, pp. 186–192
2024
-
[11]
Ultra-high-definition television (rec. itu-r bt.2020): A generational leap in the evolution of television [standards in a nutshell],
M. Sugawara, S.-Y . Choi, and D. Wood, “Ultra-high-definition television (rec. itu-r bt.2020): A generational leap in the evolution of television [standards in a nutshell],”IEEE Signal Processing Magazine, vol. 31, no. 3, pp. 170–174, 2014
2020
-
[12]
Perceptual signal coding for more efficient usage of bit codes,
S. Miller, M. Nezamabadi, and S. Daly, “Perceptual signal coding for more efficient usage of bit codes,”SMPTE Motion Imaging Journal, vol. 122, pp. 52–59, 05 2013
2013
-
[13]
Methodology for the subjective assessment of the quality of television pictures,
B. Series, “Methodology for the subjective assessment of the quality of television pictures,”Recommendation ITU-R BT, vol. 500, no. 13, 2012
2012
-
[14]
Active sampling for pairwise comparisons via approximate message passing and information gain maximization,
A. Mikhailiuk, C. Wilmot, M. Perez-Ortiz, D. Yue, and R. K. Mantiuk, “Active sampling for pairwise comparisons via approximate message passing and information gain maximization,” in2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021, pp. 2559– 2566
2021
-
[15]
A practical guide and soft- ware for analysing pairwise comparison experiments,
M. Perez-Ortiz and R. K. Mantiuk, “A practical guide and soft- ware for analysing pairwise comparison experiments,”arXiv preprint arXiv:1712.03686, 2017
2017 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.