{"id":"4092e43f-c0bd-4fdf-927f-a119ef648da2","arxiv_id":"2412.21127","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SCOPE, a VR-annotated stereo preference dataset with 2,400 comparisons, and iSQoE, a model trained on it, outperform existing image-quality metrics on mono-to-stereo conversion ranking.","lead":"This paper introduces SCOPE, a dataset of 2,400 stereoscopic image pairs rated by 103 people in virtual reality, and iSQoE, a model trained to predict which stereo image a viewer prefers. It reports that VR-based preferences differ sharply from on-screen viewing, and that iSQoE matches human choices better than existing quality metrics when ranking mono-to-stereo conversion outputs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"External mono-to-stereo claim rests on 30 images and 10 participants without error bars; the apparent margin over baselines may not be statistically robust.","rationale":"The reader's weakest assumption contains two distinct sub-concerns: (1) construct validity of VR-based 2AFC preferences as SQoE, and (2) the representativeness of the 10 participants and 30 images in the external mono-to-stereo study. My stress-test focuses on the second, more concrete issue because it is directly tied to the central claim and is addressable with standard statistical resampling. The construct-validity question is largely a scope clarification: the paper explicitly defines SQoE in a VR context and demonstrates medium differences, so it is internally consistent even if one disagrees with the premise. However, the small external sample is a correctness risk: the headline percentage differences could be within sampling noise. The paper already received a CONDITIONAL verdict due to this and related issues, and my analysis reinforces rather than redirects that verdict. Hence the verdict remains UNCHANGED: the condition for acceptance should be strengthened to require explicit uncertainty quantification or a larger external study, as the reader already indicated.","tokens_in":18732,"tokens_out":7816,"duration_ms":83919,"concrete_test":"Bootstrap or leave-one-participant-out the Section 4.4 study: for 10,000 resamples, draw 30 images with replacement and, within each image, resample the 10 participant votes with replacement to form a bootstrap majority; recompute the 'Majority' agreement for iSQoE, StereoQA-Net, and MANIQA. If the 95% confidence intervals for iSQoE and StereoQA-Net overlap, or if iSQoE is not the top-ranked model in at least 90% of resamples, the claimed 56.7% vs 23.3% advantage is not statistically established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that iSQoE 'aligns better with human preferences than existing methods' for mono-to-stereo conversion is established exclusively by the experiment in Section 4.4 with 30 images from Spring and 10 participants on an Apple Vision Pro. The reported majority agreement of 56.7% for iSQoE corresponds to 17/30 images; StereoQA-Net's 23.3% is 7/30. The paper reports no confidence intervals, significance tests, or tie-handling rules for the 10-vote majority, and the human majority itself is a noisy estimate. If only 2-3 of the 30 images changed majority label under a different participant sample, the margin over StereoQA-Net could shrink to one image, and the ordering could plausibly flip under resampling. Because the abstract's headline assertion is precisely this alignment result, the absence of uncertainty quantification is the load-bearing weakness. A secondary concern is that the two baselines are used off-the-shelf in a task they were not designed for, which lowers the bar; but the primary issue is that even the claimed advantage is not shown to be statistically robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces SCOPE, a dataset of 2,400 stereo image pairs with two-alternative forced choice (2AFC) preference labels collected on an Apple Vision Pro headset, covering 19 distortion types from photometric, spatial, noise, compression, and novel-view synthesis categories. The authors propose iSQoE, a model that processes the left and right views with DINOv2 backbones with cross-attention fusion, trained with a hinge loss on the SCOPE labels. They report that iSQoE outperforms existing IQA and NR-SIQA baselines on the SCOPE test set, exhibits monotonic responses to unseen distortion severities, and, in an external evaluation using three mono-to-stereo conversion tools on 30 Spring images with 10 participants, achieves higher agreement with human majority votes than MANIQA and StereoQA-Net. The paper also includes a cross-medium analysis showing low correlation between VR and non-VR preference judgments.","tokens_in":18946,"tokens_out":6508,"duration_ms":58110,"significance":"The SCOPE dataset is a substantive contribution: it is the largest stereo preference dataset with VR-based annotations, and the paper's cross-headset consistency analysis (kappa = 0.48) is a useful first step toward understanding medium effects on stereo preference. The model architecture is sensible, and the in-distribution accuracy (84.8% on unanimous test cases) demonstrates that the dataset contains learnable signal. The external mono-to-stereo evaluation addresses a real use case. However, the headline claim of superior alignment with human preferences rests entirely on a small, statistically underpowered study, and the model's underperformance on established SIQA benchmarks suggests its scope is narrower than the abstract implies.","major_comments":[{"comment":"The external mono-to-stereo study uses 30 images and 10 participants, and the reported agreement scores are point estimates with no confidence intervals, significance tests, or tie-handling rules. The key result (iSQoE 56.7% vs StereoQA-Net 23.3% majority agreement) corresponds to 17/30 vs 7/30 images; a McNemar test or bootstrap over images and/or participants should be reported to show the margin is not plausibly due to sampling noise. The paper should also state how ties in the 10-vote human majority are treated when computing binary agreement. Because this experiment is the sole support for the abstract's claim that iSQoE aligns better with human preferences for mono-to-stereo conversion, this statistical robustness evidence is necessary.","section":"4.4 / Figure 7"},{"comment":"The construct validity of VR-based 2AFC preferences as a measure of SQoE is not established. The cross-medium analysis shows that VR and non-VR preferences correlate weakly (kappa ~0.14-0.22) and that two VR headsets agree moderately (kappa ~0.48), but this does not validate VR preferences as a ground truth; it only quantifies medium differences. Since SCOPE training labels and the external evaluation both use the same VR protocol, the risk of circularity in the central claim remains. The authors should provide evidence that VR preferences correlate with other SQoE-related constructs (e.g., reported comfort, perceived depth quality) or explicitly and consistently frame the model as predicting VR preference rather than general stereoscopic quality of experience.","section":"3.2 / 4.3 / Figure 6"},{"comment":"The paper's central claim is supported only by the small external study in Section 4.4, while Table 4 shows that iSQoE underperforms existing NR-SIQA methods on standard benchmarks (e.g., SROCC 0.774 vs 0.972 on LIVE 3D Phase I). The authors attribute this to annotation-medium differences, but this means the model's practical domain is narrow. To justify the unqualified claim in the abstract, additional external validation is needed—e.g., a larger or second out-of-distribution study, or at least a sensitivity analysis of the Section 4.4 results. Otherwise, the claim should be qualified to apply only to the specific conversion tools and test set used.","section":"4.4 vs. Table 4"}],"minor_comments":[{"comment":"In the sentence 'human preferences on Apple Vision Pro and Meta Quest Pro are have non-negligible correlation', 'are have' should be 'have'.","section":"4.3"},{"comment":"The caption uses 'Stereo-IQA' but the text and the rest of the paper refer to the baseline as 'StereoQA-Net'; please make the names consistent.","section":"Figure 7 caption"},{"comment":"In the phrase 'We use a with a margin of 0.05', there is a missing word; it should read 'We use a margin of 0.05' or 'We use a hinge loss with a margin of 0.05'.","section":"7"},{"comment":"In the sentence 'Various factors may effect the final quality', 'effect' should be 'affect'.","section":"Introduction"},{"comment":"The caption contains 'T est Accuracy' with a stray space; should be 'Test Accuracy'.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is solid and the paper is generally well organized. The main concern is that the headline claim is supported by a very small external study; I would expect the authors to either provide proper uncertainty quantification and ideally a larger validation, or to moderate the claim in the abstract. The paper's scope is appropriate for a computer vision venue, but the evidence basis for the central claim currently falls short of what I would consider conclusive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: SCOPE is a real contribution, and iSQoE is a reasonable first model, but the paper's central claim is much weaker than the abstract implies. The Section 4.4 experiment has 30 images, 10 participants, and no uncertainty quantification. That's a load-bearing soft spot, not a nit.\n\nWhat's actually new: SCOPE is the largest stereo preference dataset, 2400 samples with 2AFC annotations collected in VR, roughly twice the next largest. The VR medium comparison (kappa ~0.48 between headsets, ~0.15–0.22 between VR and screen) is useful evidence that annotation medium matters. The cross-attention extension of DINOv2 is a modest but sensible architectural tweak. The progressive degradation plots in Figure 5 are a good sanity check and show the model extrapolates monotonically over a wider range than trained.\n\nWhat's done well: the paper is honest. The limitations section acknowledges the narrow data sources and backbone constraints. Table 4 is reported even though it shows iSQoE underperforms all existing SIQA methods on LIVE and WIVC by a large margin; that transparency is rare and should be credited.\n\nWhere it's soft, in order:\n1. The mono-to-stereo claim (abstract and Section 4.4). 56.7% vs 23.3% on 30 images could shrink or flip under resampling. No confidence intervals, no tie-handling details, no significance test. This is the single biggest issue.\n2. No release link. The paper says SCOPE is public, but I see no URL for dataset or code. For a dataset paper, that's a major omission.\n3. Baselines are off-the-shelf, not adapted, which lowers the bar. StereoQA-Net wasn't designed for generative mono-to-stereo artifacts.\n4. Transfer to existing SIQA benchmarks is poor, attributed to annotation medium. That may be true, but it also means the construct is narrow; a larger VR-based external study would help.\n\nBottom line: the dataset and model deserve a serious referee. Don't desk-reject. But the external validation must be substantially expanded, or the claim softened, before the headline is trustworthy.","headline":"A genuinely useful VR stereo preference dataset and a reasonable model, but the headline mono-to-stereo claim rests on 30 images and 10 raters with no error bars; peer review should push for a bigger external study and released artifacts.","tokens_in":19517,"tokens_out":2546,"would_cite":true,"duration_ms":25337,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Stereoscopic quality of experience can be learned from VR preference votes, and the resulting model outperforms existing metrics.","keywords":["stereoscopic quality of experience","virtual reality","two-alternative forced choice","image quality assessment","perceptual metric","DINOv2","dataset","mono-to-stereo conversion"],"falsifier":"Run the external mono-to-stereo ranking study with a larger and more diverse participant pool (e.g., 50 people) and more source images (e.g., 100); if iSQoE's agreement with the majority human vote falls to the level of StereoQA-Net or below, the reported superiority would not survive. Another check: collect SCOPE-style preferences for the same images in a different VR headset; if the cross-headset Cohen's kappa drops to near zero, the claim that VR preferences are a stable, headset-independent ground truth is undermined.","tokens_in":18546,"feed_emoji":"🥽","tokens_out":15353,"duration_ms":125770,"temperature":0.7,"pith_summary":"This paper argues that the quality of a stereoscopic image should be measured by what humans prefer when they view it in a virtual-reality headset, not by how it scores on a 2D screen or by isolated low-level artifacts. To support this, the authors introduce SCOPE, a dataset of 2,400 distorted stereo-image pairs annotated by 103 participants in a two-alternative forced-choice VR study, and train iSQoE, a deep network that takes the left and right views and predicts a holistic quality score. On the SCOPE test set's unanimous-annotation subset, iSQoE reaches 84.8% agreement with human choices, beating the best existing no-reference stereo quality model (72.5%) and the top image-quality model (71.7%). They also show that in a small external study ranking three mono-to-stereo conversion tools on 30 images, iSQoE matches the majority human vote in 56.7% of cases, far above the same baselines. If correct, this provides a practical automated evaluator for the expanding field of stereo and VR content creation.","feed_headline":"Rank stereo images by what humans prefer in VR","feed_subtitle":"Trained on 2,400 VR preference votes, the new iSQoE model agrees with human judges far more often.","key_machinery":"The central machinery is the combination of a new preference dataset and a twin-branch (Siamese) network architecture. SCOPE provides 2,400 pairwise preference labels collected in VR, covering a wide range of distortion types and two creation pipelines (from monocular images via MotionCtrl or depth-based 2D lifting, and from multi-view scenes via 3D Gaussian splatting). The iSQoE model uses two DINOv2 backbones, one per view, with keys and values concatenated between the views' attention blocks at alternating layers (layers 2, 5, 8, 11); the pooled tokens are concatenated and passed through a small MLP with sigmoid output to produce a quality score, trained with hinge loss and low-rank (LoRA) adaptation of the backbone. The key design choice is early cross-view information fusion, which the ablations show to be consistently better than no fusion.","core_discovery":"On the paper's own terms, the central discovery is that stereoscopic quality of experience (SQoE) is a learnable preference signal that is best collected in VR, and that an SQoE model trained on such preferences generalizes beyond its training distortions. The SCOPE dataset contains 2,400 samples built from physically captured stereo images (Holopix50k) and multi-view scenes rendered with 3D Gaussian splatting, spanning 19 distortion types from classical noise and blur to diffusion-based editing and novel-view synthesis. Each sample is a pair of distorted versions of the same stereo image, labeled by five VR-headset participants who choose which version they prefer; about a third of the samples are unanimously agreed upon, a third have 4-1 splits, and a third are 3-2. The iSQoE model processes the two views with a shared DINOv2 backbone, fusing information across views by concatenating attention keys and values at alternating layers, and is trained with a hinge loss against the preference labels. The paper reports that on unanimous test samples iSQoE achieves 84.8% accuracy, and that in the external mono-to-stereo conversion study with three off-the-shelf tools it agrees with the majority human vote on 56.7% of cases, compared with 23.3% for StereoQA-Net and 6.7% for MANIQA.","pith_inferences":["A larger replication of the 30-image mono-to-stereo study would be needed to know whether the 56.7% agreement figure is stable or an overestimate from a small sample; the authors' own study used ten participants.","Because the model was trained on preferences collected with a single VR headset (Apple Vision Pro), its transfer to other displays with different optics and brightness is untested; collecting a small set of cross-headset preferences and fine-tuning could measure and close that gap.","The cross-view attention fusion likely encodes disparity and depth cues; analyzing which distortion types benefit most from fusion could identify what SQoE measures beyond 2D image quality.","Legacy stereo-quality datasets annotated on screens may measure a different construct; re-annotating subsets of those datasets in VR could show whether widely used benchmarks need to be re-collected."],"forward_implications":["Stereo content creators and VR platforms can use iSQoE to rank conversion methods and filter outputs without running additional user studies.","SQoE evaluation should move annotation protocols from screens and anaglyph presentations to VR headsets, since preferences collected on a 2D screen agree only weakly with those collected in VR.","A model trained on pairwise preference choices can extrapolate to distortion strengths and types not seen in training, such as downscaling, and behaves more monotonically than existing metrics.","The SCOPE dataset's size (2,400 samples, 19 distortion types) and its VR-based labels provide a benchmark that existing stereo quality datasets, which are smaller and annotated on screens, do not offer."],"supporting_citations":[{"why":"DINOv2 is the feature backbone for the model; its representations carry the perceptual signal.","marker":"[60]"},{"why":"DreamSim supplies the 2AFC training paradigm and the Siamese-with-LoRA design that iSQoE adapts.","marker":"[25]"},{"why":"LPIPS establishes perceptual metrics trained on human two-alternative choices, the conceptual basis for this work.","marker":"[101]"},{"why":"Holopix50k is the source of the physically captured stereo images that form the majority of SCOPE.","marker":"[29]"},{"why":"MotionCtrl generates stereo pairs from monocular video by imposing camera motion, providing synthetic novel-view distortions.","marker":"[90]"},{"why":"3D Gaussian splatting renders stereo pairs from multi-view scenes, generating the novel-view synthesis subset of SCOPE.","marker":"[36]"},{"why":"StereoQA-Net is the main existing no-reference stereo quality baseline, outperformed by iSQoE in both internal and external evaluations.","marker":"[104]"},{"why":"MANIQA is the strongest 2D image-quality baseline used for comparison in the study.","marker":"[95]"},{"why":"The Spring dataset provides the monoscopic source images for the external mono-to-stereo conversion evaluation.","marker":"[52]"}],"fun_headline_variants":["New VR dataset and model rank stereoscopic images by human taste","iSQoE: a stereoscopic quality metric learned from 2,400 VR votes","SCOPE dataset: 2,400 VR preference judgments for stereo images","Predicting stereoscopic quality of experience from human VR preferences"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument rests on the premise that forced-choice preferences collected from VR headset viewing are a valid and transferable measure of stereoscopic quality of experience, so that rankings that agree with those preferences are better rankings.","fun_headline_variants_meta":{"raw":{"variants":["New VR dataset and model rank stereoscopic images by human taste","iSQoE: a stereoscopic quality metric learned from 2,400 VR votes","SCOPE dataset: 2,400 VR preference judgments for stereo images","Predicting stereoscopic quality of experience from human VR preferences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000957,"raw_usage":{"total_tokens":4100,"prompt_tokens":985,"completion_tokens":3115,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":3036}},"tokens_in":601,"tokens_out":3115,"duration_ms":20738,"temperature":1.0,"reasoning_tokens":3036,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:01:31.000538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the external mono-to-stereo ranking study with a larger and more diverse participant pool (e.g., 50 people) and more source images (e.g., 100); if iSQoE's agreement with the majority human vote falls to the level of StereoQA-Net or below, the reported superiority would not survive. Another check: collect SCOPE-style preferences for the same images in a different VR headset; if the cross-headset Cohen's kappa drops to near zero, the claim that VR preferences are a stable, headset-independent ground truth is undermined.","supporting_citations":[{"cited_title":"Dinov2: Learning robust visual features without supervi- sion","cited_arxiv_id":null,"evidence_quote":"DINOv2 is the feature backbone for the model; its representations carry the perceptual signal."},{"cited_title":"The unreasonable effectiveness of deep features as a perceptual metric","cited_arxiv_id":null,"evidence_quote":"LPIPS establishes perceptual metrics trained on human two-alternative choices, the conceptual basis for this work."},{"cited_title":"Motionctrl: A unified and flexible motion controller for video generation","cited_arxiv_id":null,"evidence_quote":"MotionCtrl generates stereo pairs from monocular video by imposing camera motion, providing synthetic novel-view distortions."},{"cited_title":"Dual-stream inter- active networks for no-reference stereoscopic image quality assessment","cited_arxiv_id":null,"evidence_quote":"StereoQA-Net is the main existing no-reference stereo quality baseline, outperformed by iSQoE in both internal and external evaluations."},{"cited_title":"Maniqa: Multi-dimension attention network for no- reference image quality assessment","cited_arxiv_id":null,"evidence_quote":"MANIQA is the strongest 2D image-quality baseline used for comparison in the study."},{"cited_title":"Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo","cited_arxiv_id":null,"evidence_quote":"The Spring dataset provides the monoscopic source images for the external mono-to-stereo conversion evaluation."}],"review_version":1}