Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper introduces a large-scale paired HDR/SDR video quality dataset and shows HDR's benefit depends on content brightness, texture, and motion.

desk verdict A genuinely useful new HDR/SDR pairwise dataset, but the 'general HDR TV' claim needs per-TV analysis before the headline conclusions can be trusted. read the letter →

arxiv 2505.21831 v1 pith:LZHRM2JW submitted 2025-05-27 cs.CV

classification cs.CV
keywords videoqualityassessmentHDRSDRsubjectivedatasetpairwisecomparisonjust-objectionabledifferenceadaptivestreamingofexperience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds a benchmark for comparing video quality between High Dynamic Range (HDR10) and Standard Dynamic Range (SDR) versions of the same content. It collects 960 videos from 54 sources, each encoded at nine quality levels, and has 145 viewers make over 22,000 pairwise choices on six consumer televisions, converted into a single Just-Objectionable-Difference scale. The central finding is that HDR is not uniformly better: it wins on bright, textured, wide-gamut content, while its advantage shrinks and sometimes reverses on dark, low-color, or motion-heavy scenes and at low bitrates. If the dataset holds up, it gives streaming engineers and quality metrics a direct measurement of when a format switch actually pays off.

What carries the argument

The load-bearing object is the HDRSDR-VQA dataset itself: 54 source clips (with measured spatial, temporal, and colorfulness statistics plus luminance distributions) rendered in HDR10 and SDR, each at nine distortion levels, judged by pairwise comparison. The measurement engine is active sampling (ASAP) to pick informative pairs, followed by maximum-likelihood scaling (pwcmp) into Just-Objectionable-Differences, where one JOD means 75 percent of observers prefer one version; the JOD scale is what lets HDR and SDR be placed on one quality axis.

What would settle it

Compute the HDR-minus-SDR JOD difference separately for each of the six televisions. If the crossover pattern disappears on the brightest TV, or if the dimmest TVs reverse the bright-content HDR advantage, then the pooled conclusion that HDR wins on bright scenes is an artifact of averaging different displays rather than a general property.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a content-dependent preference law: HDR10 outperforms SDR when the source has highlight pixels beyond the SDR range, saturated colors outside the sRGB gamut, and fine texture that higher bit depth can preserve, but this advantage disappears or inverts in scenes with little brightness or color range and in high-motion content, where SDR's lower bit-depth demand is more robust to compression. The direction and size of the preference also depend on bitrate and resolution: HDR's gap widens at high bitrates and can turn negative at low bitrates.

Load-bearing premise

The pooled JOD analysis assumes that viewer preferences from six TVs spanning 252 to 2,564 cd/m² peak brightness can be combined into one scale that represents the general HDR viewing experience.

Editorial extensions

If this is right

  • Because all videos have paired HDR and SDR versions on the same JOD scale, objective VQA models can be tested against a direct format-preference signal rather than separate quality scores.
  • Content-adaptive streaming can use content statistics such as brightness, gamut, texture, and motion to decide when HDR is worth the bitrate, since the dataset maps where HDR's advantage appears.
  • Rate-distortion behavior differs by format: at low bitrates HDR can fall behind SDR on unfavorable content, so a fixed HDR-first ladder is not optimal.
  • The anchor contents link the new scale to an earlier HDR/SDR comparison database, making scores from both studies comparable.
  • The public subset of the dataset lets other labs reproduce the scaling and extend the analysis to new models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the per-TV data were released, one could test whether peak brightness moderates the preference: HDR's bright-scene advantage might shrink or reverse on the dimmest displays, which would qualify the pooled conclusion.
  • The SDR baseline depends on the chosen conversion and grading, so the comparison is one sample of SDR rather than SDR-in-general; a different SDR master could shift the bitrates where the crossover occurs.
  • The reported content statistics (SI, TI, CF, and luminance extremes) could train a simple predictor of format preference, generalizing the four-case analysis to unseen content.
  • Because active sampling prioritizes ranking information, the resulting scale may under-represent rare but strong disagreements; a re-analysis by viewer or device group could reveal whether the HDR preference on bright scenes is universal or driven by a subset.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces HDRSDR-VQA, a subjective video quality dataset built for direct HDR-versus-SDR comparison. The dataset contains 960 videos from 54 source sequences, each source rendered in HDR10 and SDR at nine quality levels (one reference and eight distorted variants), and the authors report a subjective study with 145 participants on six consumer HDR-capable televisions, collecting over 22,000 pairwise comparisons that are scaled to JOD scores using the pwcmp algorithm. The paper presents four content-level case studies and concludes that HDR generally provides perceptual advantages for content with rich textures and high brightness or color gamut, but that these advantages diminish or reverse in low-dynamic-range or motion-intensive scenes. The authors position the dataset as a reusable benchmark for VQA, adaptive streaming, and perceptual model development.

Significance. If the dataset is sound, it fills a real gap: prior VQA datasets typically evaluate only one dynamic-range format, and the few HDR/SDR comparison studies use limited display conditions or absolute category rating rather than pairwise comparison. The scale of the study (960 videos, six TVs, 145 participants) is substantial, and the concrete protocol details, including the use of ASAP active sampling and pwcmp scaling, are credible. The count arithmetic is internally consistent: 51 non-anchor sources times 18 versions plus 3 anchors times 14 versions gives 960 videos, and 145 participants times 160 comparisons equals 23,200 comparisons, which is consistent with the claimed 'over 22,000'. The main scientific value would be in enabling display-aware study of when HDR is preferred, assuming the pooled JOD analysis is not hiding strong display dependence. The public release of scores and stimuli would also make this a practical resource for the community. The key weakness is external validity: the paper claims results representative of a 'general HDR TV' without reporting or modeling per-display variation.

major comments (3)
  1. [IV.A, Table II] The pooled JOD analysis combines comparisons from six televisions whose HDR peak brightness ranges from 252 to 2564 cd/m2 and whose Rec.2020 coverage ranges from 23.8% to 54.8%. The text in Section IV.A explicitly aims for results 'representative of a general HDR TV', but the paper reports no per-TV JOD scales, no display-as-covariate model, and no interaction analysis. If HDR versus SDR preference reverses on low-brightness or low-gamut displays such as the CU8000 or Vizio M6, the aggregate conclusions in Section VI about when HDR wins would be an artifact of the particular mix of TVs. The central claim of the paper therefore needs a per-TV analysis or a statistical model treating the display as a random effect, at minimum to show that the HDR-SDR JOD differences do not change sign or significance across the six displays.
  2. [V, Fig. 5] The concluding generalization that HDR advantages 'diminish or even reverse in low dynamic or motion-intensive scenes' is supported only by four selected content examples. No aggregate statistics are reported across the 54 contents, such as the distribution of HDR-minus-SDR JOD differences at each bitrate, the proportion of contents for which the difference is statistically distinguishable from zero, or correlations with SI, TI, colorfulness, and luminance metrics. For instance, the text describes the '28 Swan' HDR-SDR difference as negligible without reporting confidence intervals, so the reader cannot tell whether the claimed reversals are real effects or noise. The paper should provide confidence intervals for the JOD estimates and a content-level analysis linking content attributes to HDR-SDR preference.
  3. [Abstract and III.A-Anchor Contents] The three anchor contents are taken from the authors' own LIVE-HDRvsSDR database and are used to calibrate that database with the new dataset, but the paper does not report any result of this calibration, such as consistency checks between overlapping conditions, nor does it state how the anchor contents' 14-variation structure affects the pooled scaling. Since anchor contents have fewer versions than the other 51 sequences, the effective coverage is imbalanced, and the JOD scale may be less precisely determined for those contents. Please provide the anchor-calibration outcome and explain whether the pooled JOD scale is robust to excluding or reweighting the anchor contents.
minor comments (5)
  1. [III.A] There are typos and spacing issues in the text, including 'sourcd' instead of 'sourced' and inconsistent 'V oD' spacing; these should be corrected.
  2. [I and Table I] The text says 'eight distinct levels of distortion' and then Table I is described as detailing 'the specific categories of distortions', but Table I only lists resolutions and bitrates; the table caption and surrounding text should be made consistent.
  3. [IV.B] The paper reports 145 participants but the listed gender counts (53 female, 91 male, one undisclosed) sum to 145, while the per-TV counts (19, 26, 21, 26, 26, 27) also sum to 145; this is internally consistent, but the 'over 22,000 comparisons' statement could be made exact (160 per participant times 145 participants gives 23,200).
  4. [IV.B] One color-deficient participant was retained in the study 'in line with our practice of accommodating diverse participants'; since colorfulness is discussed as a factor in HDR preference, please clarify whether this participant's data were removed or separately analyzed in the scaling.
  5. [III.C] The conversion to SDR is described only as using NBCU LUTs for non-VoD content and professional mastering for VoD content; please state whether the same LUT was applied to all source types and whether any tone mapping or gamut mapping was applied to the HDR versions for display on each TV, since this affects interpretation of the format comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the dataset construction and JOD scaling are self-contained; the prior-database anchors are calibration, not a load-bearing derivation.

full rationale

This is a dataset-construction paper, not a derivation. The 960 videos are generated from 54 source sequences with explicitly specified encodings, and the 22,000+ pairwise comparisons are collected from 145 participants on six televisions. The transformation from pairwise choices to JOD scores uses the external pwcmp maximum-likelihood tool [15]; no parameter is fitted to a target and then renamed as a prediction. The conclusions in Section V about when HDR is preferred are descriptive analyses of the collected JOD values, not outputs of a model whose inputs already contain those conclusions. The only self-reference is the inclusion of three anchor contents from the authors' earlier LIVE-HDRvsSDR database [9], described as 'calibrate and integrate data'; this is a transparency measure for scale alignment, and subjective judgments for those anchors are still collected in the present study. The cited prior datasets, including LIVE-HDR [8], are independent external resources; none is invoked as a uniqueness theorem or as a premise that forces the reported preferences. The skeptical concern about pooling across televisions with 252-2564 cd/m² peak brightness is an external-validity limitation, not a circularity, because the JOD scale is not defined in terms of the television mix in a way that makes the conclusions true by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The dataset claim rests on several domain assumptions about the subjective scaling, TV pooling, and SDR conversion; none of these are validated inside the paper, and they are load-bearing for the JOD scores.

assumptions (4)
  • domain assumption Pairwise comparison data scaled by pwcmp yields reliable interval-scale JOD scores.
    Section IV.D: JODs are obtained by maximum likelihood estimation from incomplete pairwise comparisons; the central scores depend on this assumption.
  • domain assumption Scores from six TVs with different peak brightness can be pooled into one scale.
    Section IV.A: TV diversity is intended to represent general HDR viewing, but no per-TV statistical analysis is provided; if preference is TV-dependent, pooled JODs may mislead.
  • domain assumption NBCU LUT conversion for non-VOD content and Amazon Studios grading for VOD produce comparable SDR versions.
    Section III.C: two different SDR creation paths are mixed without validation, so format comparisons may reflect conversion artifacts rather than intrinsic HDR/SDR quality.
  • domain assumption The ASAP active sampling strategy produces unbiased pairwise comparisons.
    Section IV.D: ASAP selects informative pairs and reduces redundancy; if the sampling strategy biases the comparisons, the scaled JODs inherit the bias.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation." pith.science (2026). https://pith.science/paper/LZHRM2JW

@misc{pith2026250521831,
  author       = {Pith},
  title        = {Pith review of: HDRSDR-VQA: A Subjective Video Quality Dataset for HDR and SDR Comparative Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LZHRM2JW}},
  note         = {Machine review of arXiv:2505.21831}
}
read the original abstract

We introduce HDRSDR-VQA, a large-scale video quality assessment dataset designed to facilitate comparative analysis between High Dynamic Range (HDR) and Standard Dynamic Range (SDR) content under realistic viewing conditions. The dataset comprises 960 videos generated from 54 diverse source sequences, each presented in both HDR and SDR formats across nine distortion levels. To obtain reliable perceptual quality scores, we conducted a comprehensive subjective study involving 145 participants and six consumer-grade HDR-capable televisions. A total of over 22,000 pairwise comparisons were collected and scaled into Just-Objectionable-Difference (JOD) scores. Unlike prior datasets that focus on a single dynamic range format or use limited evaluation protocols, HDRSDR-VQA enables direct content-level comparison between HDR and SDR versions, supporting detailed investigations into when and why one format is preferred over the other. The open-sourced part of the dataset is publicly available to support further research in video quality assessment, content-adaptive streaming, and perceptual model development.

Figures

Figures reproduced from arXiv: 2505.21831 by the authors.

Figure 1
Figure 1. Sample frames from the source sequences. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Spatial Information (SI) plotted against Temporal Information (TI) [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Min, max, and mean luminance metrics measured on all of the source [PITH_FULL_IMAGE:figures/full_fig_p002_3.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Examples Rate-Distortion Curve of four contents collected in our [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 4
Figure 4. Figure 4: Screenshot of the rating screen used to determine which video [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A 300+ device crowd-sourced VQA dataset plus Blade-Chest aggregation and a condition-adaptation MLP let standard metrics predict quality orderings under real mobile viewing conditions far better than unadapted baselines.

  2. Learning Quality from Complexity and Structure: A Feature-Fused XGBoost Model for Video Quality Assessment

    cs.MM 2025-06 reject novelty 3.0 of 10

    A concatenation of VCA complexity residuals and SSIM fed into an XGBoost regressor reaches PLCC 0.787 on the VQA Grand Challenge test set, though the method is not truly reduced-reference.

Reference graph

Works this paper leans on

15 extracted references · 13 canonical work pages · cited by 2 Pith papers

  1. [1]

    Study of subjective and objective quality assessment of video,

    K. Seshadrinathan, R. Soundararajan, A. C. Bovik, and L. K. Cormack, “Study of subjective and objective quality assessment of video,”IEEE transactions on Image Processing, vol. 19, no. 6, pp. 1427–1441, 2010

  2. [2]

    Video quality assessment on mobile devices: Subjective, behavioral and ob- jective studies,

    A. K. Moorthy, L. K. Choi, A. C. Bovik, and G. De Veciana, “Video quality assessment on mobile devices: Subjective, behavioral and ob- jective studies,”IEEE Journal of Selected Topics in Signal Processing, vol. 6, no. 6, pp. 652–671, 2012

  3. [3]

    Evaluating the Performance of Existing Full-Reference Quality Metrics on High Dynamic Range (HDR) Video Content

    M. Azimi, A. Banitalebi-Dehkordi, Y . Dong, M. T. Pourazad, and P. Nasiopoulos, “Evaluating the performance of existing full-reference quality metrics on high dynamic range (hdr) video content,”arXiv preprint arXiv:1803.04815, 2018

  4. [4]

    Hdr video quality assessment: Perceptual evaluation of compressed hdr video,

    X. Pan, J. Zhang, S. Wang, S. Wang, Y . Zhou, W. Ding, and Y . Yang, “Hdr video quality assessment: Perceptual evaluation of compressed hdr video,”Journal of Visual Communication and Image Representation, vol. 57, pp. 76–83, 2018

  5. [5]

    Verification test report for hdr/wcg video coding using hevc main 10 profile,

    V . Baroncini, K. Andersson, A. Ramasubramonian, and G. Sullivan, “Verification test report for hdr/wcg video coding using hevc main 10 profile,” inProc. JCTVC-X1018 24th JCT-VC Meeting, 2016, pp. 293– 303

  6. [6]

    Subjective and objective evaluation of hdr video compression,

    M. Rerabek, P. Hanhart, P. Korshunov, and T. Ebrahimi, “Subjective and objective evaluation of hdr video compression,” in9th International Workshop on Video Processing and Quality Metrics for Consumer Electronics (VPQM), 2015

  7. [7]

    Perceptual quality assessment of uhd-hdr-wcg videos,

    S. Athar, T. Costa, K. Zeng, and Z. Wang, “Perceptual quality assessment of uhd-hdr-wcg videos,” in2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019, pp. 1740–1744

  8. [8]

    A study of subjective and objec- tive quality assessment of hdr videos,

    Z. Shang, J. P. Ebenezer, A. K. Venkataramanan, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “A study of subjective and objec- tive quality assessment of hdr videos,”IEEE Transactions on Image Processing, vol. 33, pp. 42–57, 2023

Show all 15 references
  1. [9]

    Hdr or sdr? a subjective and objective study of scaled and compressed videos,

    J. P. Ebenezer, Z. Shang, Y . Chen, Y . Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “Hdr or sdr? a subjective and objective study of scaled and compressed videos,”IEEE Transactions on Image Processing, 2024

  2. [10]

    Avt-vqdb- uhd-2-hdr: An open 8k hdr source dataset for video quality research,

    D. Keller, T. Goebel, V . Siebenkees, J. Prenzel, and A. Raake, “Avt-vqdb- uhd-2-hdr: An open 8k hdr source dataset for video quality research,” in2024 16th International Conference on Quality of Multimedia Expe- rience (QoMEX), 2024, pp. 186–192

  3. [11]

    Ultra-high-definition television (rec. itu-r bt.2020): A generational leap in the evolution of television [standards in a nutshell],

    M. Sugawara, S.-Y . Choi, and D. Wood, “Ultra-high-definition television (rec. itu-r bt.2020): A generational leap in the evolution of television [standards in a nutshell],”IEEE Signal Processing Magazine, vol. 31, no. 3, pp. 170–174, 2014

  4. [12]

    Perceptual signal coding for more efficient usage of bit codes,

    S. Miller, M. Nezamabadi, and S. Daly, “Perceptual signal coding for more efficient usage of bit codes,”SMPTE Motion Imaging Journal, vol. 122, pp. 52–59, 05 2013

  5. [13]

    Methodology for the subjective assessment of the quality of television pictures,

    B. Series, “Methodology for the subjective assessment of the quality of television pictures,”Recommendation ITU-R BT, vol. 500, no. 13, 2012

  6. [14]

    Active sampling for pairwise comparisons via approximate message passing and information gain maximization,

    A. Mikhailiuk, C. Wilmot, M. Perez-Ortiz, D. Yue, and R. K. Mantiuk, “Active sampling for pairwise comparisons via approximate message passing and information gain maximization,” in2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021, pp. 2559– 2566

  7. [15]

    A practical guide and soft- ware for analysing pairwise comparison experiments,

    M. Perez-Ortiz and R. K. Mantiuk, “A practical guide and soft- ware for analysing pairwise comparison experiments,”arXiv preprint arXiv:1712.03686, 2017

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.