{"id":"d4212a40-ef30-48c6-9a33-68e0926cbbc8","arxiv_id":"2505.21831","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 960-video subjective quality dataset uses pairwise comparisons to assign JOD scores to HDR and SDR versions of the same content across six TVs.","lead":"This paper introduces a large video dataset with 960 HDR and SDR versions of the same content, rated by 145 people on six TVs through over 22,000 pairwise comparisons. It lets researchers directly compare when HDR video actually looks better than SDR video, which matters for streaming and display technology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pooled JOD scaling across six very different TVs (252-2564 cd/m² peak brightness) may hide display-dependent HDR/SDR preference reversals, so the 'general HDR TV' claim is not yet supported.","rationale":"The paper contributes a substantial resource: 960 videos from 54 sources, 145 participants, six consumer TVs, over 22,000 pairwise comparisons, and a standard ASAP/pwcmp pipeline. Counts are internally consistent: 51 non-anchor sources × 18 variants plus 3 anchors × 14 variants equals 960. The use of pairwise comparison and active sampling is appropriate for detecting fine-grained HDR/SDR differences. Credit is due for the multi-TV design, which is more realistic than a single display. However, the central claim that the results are 'representative of the general HDR TV' rests on pooling comparisons across displays whose peak brightness differs by an order of magnitude and whose color gamut coverage differs by more than a factor of two. The paper reports no per-TV JOD results or any statistical test for display-condition interactions, so the pooled scale could hide a preference reversal on modest TVs. This is the same weakness the reader identified, and it is load-bearing because the dataset's benchmark value depends on the JOD scores being a meaningful single scale. A secondary concern is that the content-dependent conclusions in Section V are illustrated with only four hand-picked videos rather than a systematic regression across all 54 sources; this strengthens the need for per-TV and per-content quantitative reporting, but it is secondary to the scale-validity question. The absence of a download link and confidence intervals also limits verification, consistent with the CONDITIONAL verdict. The proposed test, per-TV scaling and a regression on display metrics, would settle whether the pooling assumption holds; until then, CONDITIONAL is the right call.","tokens_in":8091,"tokens_out":4741,"duration_ms":47485,"concrete_test":"Fit the pairwise comparisons separately per TV (or with a Bradley-Terry/Thurstone model including TV and content as factors with interaction). For each of the 54 source contents, compute the per-TV HDR-SDR JOD gap and regress these gaps on TV peak brightness and Rec.2020 coverage. If the sign or magnitude of HDR preference shifts materially across TV brightness tiers (e.g., >1000 vs <500 cd/m²), the pooled JOD claims are display-mix dependent and the manuscript should report per-TV scales and qualify its generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the 22,000+ pairwise comparisons collected on six TVs with peak brightness ranging from 252 to 2564 cd/m² and Rec.2020 coverage from 23.8% to 54.8% (Table II) can be pooled into one JOD scale that speaks for a 'general HDR TV' (Section IV.A). The paper's own introduction says HDR quality depends 'most importantly, on the display capabilities of the end device.' Yet no per-TV JOD scales, no TV-as-covariate model, and no interaction analysis are reported. If HDR preference reverses on low-brightness or low-gamut TVs (e.g., CU8000, Vizio M6), then the aggregate HDR-SDR JOD gaps are a weighted average over an inhomogeneous display population, and the Section V conclusions about when HDR wins could change with a different TV mix. This is an external-validity concern, not an internal inconsistency: the pwcmp scaling is standard, but the inference target is 'general HDR TV,' which the data as analyzed do not establish.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HDRSDR-VQA, a subjective video quality dataset built for direct HDR-versus-SDR comparison. The dataset contains 960 videos from 54 source sequences, each source rendered in HDR10 and SDR at nine quality levels (one reference and eight distorted variants), and the authors report a subjective study with 145 participants on six consumer HDR-capable televisions, collecting over 22,000 pairwise comparisons that are scaled to JOD scores using the pwcmp algorithm. The paper presents four content-level case studies and concludes that HDR generally provides perceptual advantages for content with rich textures and high brightness or color gamut, but that these advantages diminish or reverse in low-dynamic-range or motion-intensive scenes. The authors position the dataset as a reusable benchmark for VQA, adaptive streaming, and perceptual model development.","tokens_in":8254,"tokens_out":3262,"duration_ms":36868,"significance":"If the dataset is sound, it fills a real gap: prior VQA datasets typically evaluate only one dynamic-range format, and the few HDR/SDR comparison studies use limited display conditions or absolute category rating rather than pairwise comparison. The scale of the study (960 videos, six TVs, 145 participants) is substantial, and the concrete protocol details, including the use of ASAP active sampling and pwcmp scaling, are credible. The count arithmetic is internally consistent: 51 non-anchor sources times 18 versions plus 3 anchors times 14 versions gives 960 videos, and 145 participants times 160 comparisons equals 23,200 comparisons, which is consistent with the claimed 'over 22,000'. The main scientific value would be in enabling display-aware study of when HDR is preferred, assuming the pooled JOD analysis is not hiding strong display dependence. The public release of scores and stimuli would also make this a practical resource for the community. The key weakness is external validity: the paper claims results representative of a 'general HDR TV' without reporting or modeling per-display variation.","major_comments":[{"comment":"The pooled JOD analysis combines comparisons from six televisions whose HDR peak brightness ranges from 252 to 2564 cd/m2 and whose Rec.2020 coverage ranges from 23.8% to 54.8%. The text in Section IV.A explicitly aims for results 'representative of a general HDR TV', but the paper reports no per-TV JOD scales, no display-as-covariate model, and no interaction analysis. If HDR versus SDR preference reverses on low-brightness or low-gamut displays such as the CU8000 or Vizio M6, the aggregate conclusions in Section VI about when HDR wins would be an artifact of the particular mix of TVs. The central claim of the paper therefore needs a per-TV analysis or a statistical model treating the display as a random effect, at minimum to show that the HDR-SDR JOD differences do not change sign or significance across the six displays.","section":"IV.A, Table II"},{"comment":"The concluding generalization that HDR advantages 'diminish or even reverse in low dynamic or motion-intensive scenes' is supported only by four selected content examples. No aggregate statistics are reported across the 54 contents, such as the distribution of HDR-minus-SDR JOD differences at each bitrate, the proportion of contents for which the difference is statistically distinguishable from zero, or correlations with SI, TI, colorfulness, and luminance metrics. For instance, the text describes the '28 Swan' HDR-SDR difference as negligible without reporting confidence intervals, so the reader cannot tell whether the claimed reversals are real effects or noise. The paper should provide confidence intervals for the JOD estimates and a content-level analysis linking content attributes to HDR-SDR preference.","section":"V, Fig. 5"},{"comment":"The three anchor contents are taken from the authors' own LIVE-HDRvsSDR database and are used to calibrate that database with the new dataset, but the paper does not report any result of this calibration, such as consistency checks between overlapping conditions, nor does it state how the anchor contents' 14-variation structure affects the pooled scaling. Since anchor contents have fewer versions than the other 51 sequences, the effective coverage is imbalanced, and the JOD scale may be less precisely determined for those contents. Please provide the anchor-calibration outcome and explain whether the pooled JOD scale is robust to excluding or reweighting the anchor contents.","section":"Abstract and III.A-Anchor Contents"}],"minor_comments":[{"comment":"There are typos and spacing issues in the text, including 'sourcd' instead of 'sourced' and inconsistent 'V oD' spacing; these should be corrected.","section":"III.A"},{"comment":"The text says 'eight distinct levels of distortion' and then Table I is described as detailing 'the specific categories of distortions', but Table I only lists resolutions and bitrates; the table caption and surrounding text should be made consistent.","section":"I and Table I"},{"comment":"The paper reports 145 participants but the listed gender counts (53 female, 91 male, one undisclosed) sum to 145, while the per-TV counts (19, 26, 21, 26, 26, 27) also sum to 145; this is internally consistent, but the 'over 22,000 comparisons' statement could be made exact (160 per participant times 145 participants gives 23,200).","section":"IV.B"},{"comment":"One color-deficient participant was retained in the study 'in line with our practice of accommodating diverse participants'; since colorfulness is discussed as a factor in HDR preference, please clarify whether this participant's data were removed or separately analyzed in the scaling.","section":"IV.B"},{"comment":"The conversion to SDR is described only as using NBCU LUTs for non-VoD content and professional mastering for VoD content; please state whether the same LUT was applied to all source types and whether any tone mapping or gamut mapping was applied to the HDR versions for display on each TV, since this affects interpretation of the format comparison.","section":"III.C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a dataset contribution and its internal consistency is good, but the external-validity concern about pooling across very different TVs is real and directly affects the headline conclusion. I would not reject, because the dataset itself appears usable and the missing analysis is additive, but the revision must address the per-display question explicitly. The use of three anchor contents from the authors' own earlier database is acceptable and disclosed, though a robustness check would strengthen confidence in the pooled scale."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a real, new resource—960 paired HDR/SDR videos with over 22k pairwise comparisons scaled to JODs, collected across six consumer TVs. The dataset itself is worth having, and the paper is honest about how it was built. But the headline claim about 'general HDR TV' outruns the analysis: all comparisons are pooled across displays that vary from 252 to 2564 cd/m² peak brightness, and there is no per-TV JOD or TV-as-covariate analysis. If HDR preference flips on low-brightness TVs, the aggregate JOD gaps could be an artifact of the display mix.\n\nWhat's new: no prior dataset offers large-scale pairwise HDR/SDR preference data across multiple TVs. The source content spans open 8K HDR, VoD, live sports, and anchors to the authors' earlier LIVE HDRvsSDR database, which is a legitimate calibration choice. The protocol is standard: pairwise comparison, ASAP active sampling, pwcmp scaling. The counting checks out—145 participants times 160 comparisons is about 23k, matching the reported 22k-plus. The distortion grid is concrete and realistic for streaming. The content analysis (SI/TI/CF, luminance) helps characterize the source space.\n\nSoft spots, in order of seriousness. First is the pooling issue above. The authors say they aim for results 'representative of the general HDR TV' but give no evidence that preferences are consistent across TVs. At minimum they should report per-TV JOD scales or a mixed-effects model with TV as a factor. Second, no confidence intervals on the JOD scores; pwcmp gives uncertainty, so it should be reported. Third, the content-preference conclusions (HDR wins on high-texture/high-brightness, loses on low-dynamic/motion scenes) rest on four selected examples, not on a systematic content-attribute regression. The dataset may support that analysis, but the paper doesn't do it. Fourth, minor: the abstract says 'nine distortion levels' while the intro says 'eight levels of distortion' plus a reference; Table I lists eight plus reference, so nine total versions. Also, no download link is given despite the 'publicly available' claim.\n\nWho this is for: VQA researchers and streaming engineers who want a benchmark for HDR/SDR preference. The dataset will likely be more influential than the specific analysis in the paper. It deserves a serious referee, but I would ask for per-TV results, confidence intervals, and a more systematic content-preference analysis—or at least an explicit limitation that the pooled results may not generalize across displays. Not a rejection.","headline":"A genuinely useful new HDR/SDR pairwise dataset, but the 'general HDR TV' claim needs per-TV analysis before the headline conclusions can be trusted.","tokens_in":8821,"tokens_out":2201,"would_cite":true,"duration_ms":21749,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a large-scale paired HDR/SDR video quality dataset and shows HDR's benefit depends on content brightness, texture, and motion.","keywords":["video quality assessment","HDR","SDR","subjective dataset","pairwise comparison","just-objectionable difference","adaptive streaming","quality of experience"],"falsifier":"Compute the HDR-minus-SDR JOD difference separately for each of the six televisions. If the crossover pattern disappears on the brightest TV, or if the dimmest TVs reverse the bright-content HDR advantage, then the pooled conclusion that HDR wins on bright scenes is an artifact of averaging different displays rather than a general property.","tokens_in":7872,"feed_emoji":"📺","tokens_out":6628,"duration_ms":66217,"temperature":0.7,"pith_summary":"This paper builds a benchmark for comparing video quality between High Dynamic Range (HDR10) and Standard Dynamic Range (SDR) versions of the same content. It collects 960 videos from 54 sources, each encoded at nine quality levels, and has 145 viewers make over 22,000 pairwise choices on six consumer televisions, converted into a single Just-Objectionable-Difference scale. The central finding is that HDR is not uniformly better: it wins on bright, textured, wide-gamut content, while its advantage shrinks and sometimes reverses on dark, low-color, or motion-heavy scenes and at low bitrates. If the dataset holds up, it gives streaming engineers and quality metrics a direct measurement of when a format switch actually pays off.","feed_headline":"22,000 viewer comparisons map when HDR beats SDR","feed_subtitle":"New dataset of 960 HDR/SDR videos finds HDR's edge depends on brightness, texture, and motion.","key_machinery":"The load-bearing object is the HDRSDR-VQA dataset itself: 54 source clips (with measured spatial, temporal, and colorfulness statistics plus luminance distributions) rendered in HDR10 and SDR, each at nine distortion levels, judged by pairwise comparison. The measurement engine is active sampling (ASAP) to pick informative pairs, followed by maximum-likelihood scaling (pwcmp) into Just-Objectionable-Differences, where one JOD means 75 percent of observers prefer one version; the JOD scale is what lets HDR and SDR be placed on one quality axis.","core_discovery":"On the paper's own terms, the discovery is a content-dependent preference law: HDR10 outperforms SDR when the source has highlight pixels beyond the SDR range, saturated colors outside the sRGB gamut, and fine texture that higher bit depth can preserve, but this advantage disappears or inverts in scenes with little brightness or color range and in high-motion content, where SDR's lower bit-depth demand is more robust to compression. The direction and size of the preference also depend on bitrate and resolution: HDR's gap widens at high bitrates and can turn negative at low bitrates.","pith_inferences":["If the per-TV data were released, one could test whether peak brightness moderates the preference: HDR's bright-scene advantage might shrink or reverse on the dimmest displays, which would qualify the pooled conclusion.","The SDR baseline depends on the chosen conversion and grading, so the comparison is one sample of SDR rather than SDR-in-general; a different SDR master could shift the bitrates where the crossover occurs.","The reported content statistics (SI, TI, CF, and luminance extremes) could train a simple predictor of format preference, generalizing the four-case analysis to unseen content.","Because active sampling prioritizes ranking information, the resulting scale may under-represent rare but strong disagreements; a re-analysis by viewer or device group could reveal whether the HDR preference on bright scenes is universal or driven by a subset."],"forward_implications":["Because all videos have paired HDR and SDR versions on the same JOD scale, objective VQA models can be tested against a direct format-preference signal rather than separate quality scores.","Content-adaptive streaming can use content statistics such as brightness, gamut, texture, and motion to decide when HDR is worth the bitrate, since the dataset maps where HDR's advantage appears.","Rate-distortion behavior differs by format: at low bitrates HDR can fall behind SDR on unfavorable content, so a fixed HDR-first ladder is not optimal.","The anchor contents link the new scale to an earlier HDR/SDR comparison database, making scores from both studies comparable.","The public subset of the dataset lets other labs reproduce the scaling and extend the analysis to new models."],"supporting_citations":[{"why":"Contributes the anchor contents and the earlier HDR/SDR comparison structure this dataset calibrates against.","marker":"[9]"},{"why":"Supplies 31 of the 54 source sequences and their original high-resolution HDR captures.","marker":"[10]"},{"why":"Provides the active-sampling algorithm that told the study which pairs to show next.","marker":"[14]"},{"why":"Provides the maximum-likelihood scaling used to turn pairwise choices into JOD scores.","marker":"[15]"},{"why":"Specifies the pairwise-comparison subjective-testing methodology the study follows.","marker":"[13]"},{"why":"Defines the PQ transfer function used to encode HDR luminance in the test content.","marker":"[12]"}],"fun_headline_variants":["HDR wins only on bright, colorful, low-motion scenes, dataset shows","960 HDR/SDR videos reveal when HDR beats SDR—and when it doesn't","HDR's edge is conditional: only bright, textured, low-motion content","In high motion, SDR beats HDR, per 22k pairwise comparisons","New VQA dataset: 960 videos, 22k comparisons, conditional HDR win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pooled JOD analysis assumes that viewer preferences from six TVs spanning 252 to 2,564 cd/m² peak brightness can be combined into one scale that represents the general HDR viewing experience.","fun_headline_variants_meta":{"raw":{"variants":["HDR wins only on bright, colorful, low-motion scenes, dataset shows","960 HDR/SDR videos reveal when HDR beats SDR—and when it doesn't","HDR's edge is conditional: only bright, textured, low-motion content","In high motion, SDR beats HDR, per 22k pairwise comparisons","New VQA dataset: 960 videos, 22k comparisons, conditional HDR win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000673,"raw_usage":{"total_tokens":3019,"prompt_tokens":854,"completion_tokens":2165,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":2054}},"tokens_in":470,"tokens_out":2165,"duration_ms":13590,"temperature":1.0,"reasoning_tokens":2054,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:22:26.300364+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the HDR-minus-SDR JOD difference separately for each of the six televisions. If the crossover pattern disappears on the brightest TV, or if the dimmest TVs reverse the bright-content HDR advantage, then the pooled conclusion that HDR wins on bright scenes is an artifact of averaging different displays rather than a general property.","supporting_citations":[{"cited_title":"Hdr or sdr? a subjective and objective study of scaled and compressed videos,","cited_arxiv_id":null,"evidence_quote":"Contributes the anchor contents and the earlier HDR/SDR comparison structure this dataset calibrates against."},{"cited_title":"Avt-vqdb- uhd-2-hdr: An open 8k hdr source dataset for video quality research,","cited_arxiv_id":null,"evidence_quote":"Supplies 31 of the 54 source sequences and their original high-resolution HDR captures."},{"cited_title":"Active sampling for pairwise comparisons via approximate message passing and information gain maximization,","cited_arxiv_id":null,"evidence_quote":"Provides the active-sampling algorithm that told the study which pairs to show next."},{"cited_title":"Perceptual signal coding for more efficient usage of bit codes,","cited_arxiv_id":null,"evidence_quote":"Defines the PQ transfer function used to encode HDR luminance in the test content."}],"review_version":1}