Pith. sign in

REVIEW 3 major objections 5 minor 120 references

What Makes for a Good Stereoscopic Image?

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Stereoscopic quality of experience can be learned from VR preference votes, and the resulting model outperforms existing metrics.

desk verdict A genuinely useful VR stereo preference dataset and a reasonable model, but the headline mono-to-stereo claim rests on 30 images and 10 raters with no error bars; peer review should push for a bigger external study and released artifacts. read the letter →

arxiv 2412.21127 v2 pith:CEY6P6KH submitted 2024-12-30 cs.CV

classification cs.CV
keywords stereoscopicqualityofexperiencevirtualrealitytwo-alternativeforcedchoiceimageassessmentperceptualmetricDINOv2datasetmono-to-stereoconversion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the quality of a stereoscopic image should be measured by what humans prefer when they view it in a virtual-reality headset, not by how it scores on a 2D screen or by isolated low-level artifacts. To support this, the authors introduce SCOPE, a dataset of 2,400 distorted stereo-image pairs annotated by 103 participants in a two-alternative forced-choice VR study, and train iSQoE, a deep network that takes the left and right views and predicts a holistic quality score. On the SCOPE test set's unanimous-annotation subset, iSQoE reaches 84.8% agreement with human choices, beating the best existing no-reference stereo quality model (72.5%) and the top image-quality model (71.7%). They also show that in a small external study ranking three mono-to-stereo conversion tools on 30 images, iSQoE matches the majority human vote in 56.7% of cases, far above the same baselines. If correct, this provides a practical automated evaluator for the expanding field of stereo and VR content creation.

What carries the argument

The central machinery is the combination of a new preference dataset and a twin-branch (Siamese) network architecture. SCOPE provides 2,400 pairwise preference labels collected in VR, covering a wide range of distortion types and two creation pipelines (from monocular images via MotionCtrl or depth-based 2D lifting, and from multi-view scenes via 3D Gaussian splatting). The iSQoE model uses two DINOv2 backbones, one per view, with keys and values concatenated between the views' attention blocks at alternating layers (layers 2, 5, 8, 11); the pooled tokens are concatenated and passed through a small MLP with sigmoid output to produce a quality score, trained with hinge loss and low-rank (LoRA) adaptation of the backbone. The key design choice is early cross-view information fusion, which the ablations show to be consistently better than no fusion.

What would settle it

Run the external mono-to-stereo ranking study with a larger and more diverse participant pool (e.g., 50 people) and more source images (e.g., 100); if iSQoE's agreement with the majority human vote falls to the level of StereoQA-Net or below, the reported superiority would not survive. Another check: collect SCOPE-style preferences for the same images in a different VR headset; if the cross-headset Cohen's kappa drops to near zero, the claim that VR preferences are a stable, headset-independent ground truth is undermined.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that stereoscopic quality of experience (SQoE) is a learnable preference signal that is best collected in VR, and that an SQoE model trained on such preferences generalizes beyond its training distortions. The SCOPE dataset contains 2,400 samples built from physically captured stereo images (Holopix50k) and multi-view scenes rendered with 3D Gaussian splatting, spanning 19 distortion types from classical noise and blur to diffusion-based editing and novel-view synthesis. Each sample is a pair of distorted versions of the same stereo image, labeled by five VR-headset participants who choose which version they prefer; about a third of the samples are unanimously agreed upon, a third have 4-1 splits, and a third are 3-2. The iSQoE model processes the two views with a shared DINOv2 backbone, fusing information across views by concatenating attention keys and values at alternating layers, and is trained with a hinge loss against the preference labels. The paper reports that on unanimous test samples iSQoE achieves 84.8% accuracy, and that in the external mono-to-stereo conversion study with three off-the-shelf tools it agrees with the majority human vote on 56.7% of cases, compared with 23.3% for StereoQA-Net and 6.7% for MANIQA.

Load-bearing premise

The whole argument rests on the premise that forced-choice preferences collected from VR headset viewing are a valid and transferable measure of stereoscopic quality of experience, so that rankings that agree with those preferences are better rankings.

Editorial extensions

If this is right

  • Stereo content creators and VR platforms can use iSQoE to rank conversion methods and filter outputs without running additional user studies.
  • SQoE evaluation should move annotation protocols from screens and anaglyph presentations to VR headsets, since preferences collected on a 2D screen agree only weakly with those collected in VR.
  • A model trained on pairwise preference choices can extrapolate to distortion strengths and types not seen in training, such as downscaling, and behaves more monotonically than existing metrics.
  • The SCOPE dataset's size (2,400 samples, 19 distortion types) and its VR-based labels provide a benchmark that existing stereo quality datasets, which are smaller and annotated on screens, do not offer.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A larger replication of the 30-image mono-to-stereo study would be needed to know whether the 56.7% agreement figure is stable or an overestimate from a small sample; the authors' own study used ten participants.
  • Because the model was trained on preferences collected with a single VR headset (Apple Vision Pro), its transfer to other displays with different optics and brightness is untested; collecting a small set of cross-headset preferences and fine-tuning could measure and close that gap.
  • The cross-view attention fusion likely encodes disparity and depth cues; analyzing which distortion types benefit most from fusion could identify what SQoE measures beyond 2D image quality.
  • Legacy stereo-quality datasets annotated on screens may measure a different construct; re-annotating subsets of those datasets in VR could show whether widely used benchmarks need to be re-collected.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper introduces SCOPE, a dataset of 2,400 stereo image pairs with two-alternative forced choice (2AFC) preference labels collected on an Apple Vision Pro headset, covering 19 distortion types from photometric, spatial, noise, compression, and novel-view synthesis categories. The authors propose iSQoE, a model that processes the left and right views with DINOv2 backbones with cross-attention fusion, trained with a hinge loss on the SCOPE labels. They report that iSQoE outperforms existing IQA and NR-SIQA baselines on the SCOPE test set, exhibits monotonic responses to unseen distortion severities, and, in an external evaluation using three mono-to-stereo conversion tools on 30 Spring images with 10 participants, achieves higher agreement with human majority votes than MANIQA and StereoQA-Net. The paper also includes a cross-medium analysis showing low correlation between VR and non-VR preference judgments.

Significance. The SCOPE dataset is a substantive contribution: it is the largest stereo preference dataset with VR-based annotations, and the paper's cross-headset consistency analysis (kappa = 0.48) is a useful first step toward understanding medium effects on stereo preference. The model architecture is sensible, and the in-distribution accuracy (84.8% on unanimous test cases) demonstrates that the dataset contains learnable signal. The external mono-to-stereo evaluation addresses a real use case. However, the headline claim of superior alignment with human preferences rests entirely on a small, statistically underpowered study, and the model's underperformance on established SIQA benchmarks suggests its scope is narrower than the abstract implies.

major comments (3)
  1. [4.4 / Figure 7] The external mono-to-stereo study uses 30 images and 10 participants, and the reported agreement scores are point estimates with no confidence intervals, significance tests, or tie-handling rules. The key result (iSQoE 56.7% vs StereoQA-Net 23.3% majority agreement) corresponds to 17/30 vs 7/30 images; a McNemar test or bootstrap over images and/or participants should be reported to show the margin is not plausibly due to sampling noise. The paper should also state how ties in the 10-vote human majority are treated when computing binary agreement. Because this experiment is the sole support for the abstract's claim that iSQoE aligns better with human preferences for mono-to-stereo conversion, this statistical robustness evidence is necessary.
  2. [3.2 / 4.3 / Figure 6] The construct validity of VR-based 2AFC preferences as a measure of SQoE is not established. The cross-medium analysis shows that VR and non-VR preferences correlate weakly (kappa ~0.14-0.22) and that two VR headsets agree moderately (kappa ~0.48), but this does not validate VR preferences as a ground truth; it only quantifies medium differences. Since SCOPE training labels and the external evaluation both use the same VR protocol, the risk of circularity in the central claim remains. The authors should provide evidence that VR preferences correlate with other SQoE-related constructs (e.g., reported comfort, perceived depth quality) or explicitly and consistently frame the model as predicting VR preference rather than general stereoscopic quality of experience.
  3. [4.4 vs. Table 4] The paper's central claim is supported only by the small external study in Section 4.4, while Table 4 shows that iSQoE underperforms existing NR-SIQA methods on standard benchmarks (e.g., SROCC 0.774 vs 0.972 on LIVE 3D Phase I). The authors attribute this to annotation-medium differences, but this means the model's practical domain is narrow. To justify the unqualified claim in the abstract, additional external validation is needed—e.g., a larger or second out-of-distribution study, or at least a sensitivity analysis of the Section 4.4 results. Otherwise, the claim should be qualified to apply only to the specific conversion tools and test set used.
minor comments (5)
  1. [4.3] In the sentence 'human preferences on Apple Vision Pro and Meta Quest Pro are have non-negligible correlation', 'are have' should be 'have'.
  2. [Figure 7 caption] The caption uses 'Stereo-IQA' but the text and the rest of the paper refer to the baseline as 'StereoQA-Net'; please make the names consistent.
  3. [7] In the phrase 'We use a with a margin of 0.05', there is a missing word; it should read 'We use a margin of 0.05' or 'We use a hinge loss with a margin of 0.05'.
  4. [Introduction] In the sentence 'Various factors may effect the final quality', 'effect' should be 'affect'.
  5. [Figure 4 caption] The caption contains 'T est Accuracy' with a stray space; should be 'Test Accuracy'.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: held-out SCOPE test and external Spring mono-to-stereo evaluation are independent of training fit.

full rationale

The derivation chain is a standard supervised-learning pipeline. iSQoE is trained on SCOPE 2AFC VR preference labels (Sections 3.2-3.3) and evaluated on held-out SCOPE splits (Section 4.1) plus an external mono-to-stereo study on 30 Spring images that the paper explicitly states the model never saw during training and that are out of SCOPE's distribution (Section 4.4). No parameter is fitted to the Spring human votes, so the headline comparison against MANIQA and StereoQA-Net is not forced by construction. The method-level citations to LPIPS and DreamSim share authors but are used as architectural inspiration, not as load-bearing theorems or uniqueness arguments; the paper additionally reports honest negative transfer to legacy SIQA datasets (SM Table 4), showing the model is not tuned to those benchmarks. The lack of confidence intervals for the 30-image/10-participant external study is a statistical robustness concern, not circularity. No equation in the paper reduces to its own input, and no fitted quantity is renamed as a prediction.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or conceptual entities. Its central claims rest on dataset design choices (VR annotation medium, five labels per sample, distortion pool) and model hyperparameters. The VR-medium assumption is the most consequential: all labels and all evaluations are defined relative to it.

free parameters (5)
  • LoRA rank = 8
    Chosen for DINOv2 finetuning in iSQoE (supplementary Section 7); affects model capacity and generalization.
  • LoRA alpha = 32
    Scaling factor for LoRA updates (supplementary Section 7).
  • LoRA dropout = 0.1
    Regularization during training (supplementary Section 7).
  • Hinge loss margin = 0.05
    Margin for the preference hinge loss (supplementary Section 7).
  • Learning rate = 3e-5
    Adam optimizer learning rate (supplementary Section 7).
assumptions (4)
  • domain assumption 2AFC preference collected on a VR headset is a valid operationalization of Stereo Quality of Experience.
    Sections 3.2 and 4.3 assume VR medium is the right ground truth; the kappa analysis supports cross-headset consistency but not construct validity.
  • domain assumption Unanimous (5-0) annotations are the most reliable and are used as the primary evaluation subset.
    Section 4.1 uses the 5-0 subset to rank methods, assuming these labels are effectively noise-free and cognitively impenetrable.
  • domain assumption Distortions applied to Holopix50k images and 13 NVS scenes cover the artifact space relevant to real stereo content.
    Section 3.1 and Table 1 list 19 distortions, but 41% of data comes from one low-disparity source and 400 NVS images come from 13 scenes; acknowledged in Section 5.
  • domain assumption DINOv2 features pretrained on 2D images transfer adequately to stereo quality when combined with cross-attention.
    Section 3.3 and Table 2 ablate backbones and find DINOv2 best; the choice is empirical, not derived.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What Makes for a Good Stereoscopic Image?." pith.science (2026). https://pith.science/paper/CEY6P6KH

@misc{pith2026241221127,
  author       = {Pith},
  title        = {Pith review of: What Makes for a Good Stereoscopic Image?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CEY6P6KH}},
  note         = {Machine review of arXiv:2412.21127}
}
read the original abstract

With rapid advancements in virtual reality (VR) headsets, effectively measuring stereoscopic quality of experience (SQoE) has become essential for delivering immersive and comfortable 3D experiences. However, most existing stereo metrics focus on isolated aspects of the viewing experience such as visual discomfort or image quality, and have traditionally faced data limitations. To address these gaps, we present SCOPE (Stereoscopic COntent Preference Evaluation), a new dataset comprised of real and synthetic stereoscopic images featuring a wide range of common perceptual distortions and artifacts. The dataset is labeled with preference annotations collected on a VR headset, with our findings indicating a notable degree of consistency in user preferences across different headsets. Additionally, we present iSQoE, a new model for stereo quality of experience assessment trained on our dataset. We show that iSQoE aligns better with human preferences than existing methods when comparing mono-to-stereo conversion methods.

Figures

Figures reproduced from arXiv: 2412.21127 by the authors.

Figure 1
Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Dataset examples. Each stereo image was subjected to two different distortions, applied consistently to either the left, right, or both images. Participants in a VR-based user study were then asked to choose their preferred version. On the right of each sample, we zoom to highlight the differences between the images. Some distortions are more easily visible in 2D (e.g. Gaussian White Noise, Rotation) while others ar… view at source ↗
Figure 3
Figure 3. Model architecture. The left and right images of a stereo pair are processed by a modified DINOv2 [60] network with cross￾attention between images. The resulting spatial tokens are pooled, concatenated, and passed through a small fully-connected network, outputting a single value indicating quality (lower is better). We train the model with a hinge loss and LoRA [28] for the DINOv2 network. data sample we ensure the… view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Test accuracy on SCOPE. We report the mean and stan￾dard deviation of the unanimous cases in the test set over several splits. Our model outperforms the other SIQA and IQA models. We partition the dataset into train, validation, and test subsets several times. When eva…
Figure 5
Figure 5. Figure 5: Progressive degradation evaluation. We report model scores on 200 stereo images for six different distortions. The gray regions indicate the distortion intensity used in SCOPE. Stereo images (a) – (d), presented as anaglyphs, exhibit progressive downscaling and are rep…
Figure 7
Figure 7. Figure 7: Alignment of stereo image preference between our participants and models. We calculate the agreement in two dif￾ferent ways. (i) Majority, where the model receives a binary score based on agreement with the majority human vote. (ii) Propor￾tional, in which the model sc…
Figure 8
Figure 8. Figure 8: Distortion examples. We show examples of image distortions in our dataset, exaggerated for illustration purposes. Q-Align BRISQUE CLIP-IQA MANIQA StereoQA-Net Ours Model 40 50 60 70 80 90 Accuracy (%) 49.7% 56.1% 52.8% 53.2% 53.9% 55.3% 58.0% 55.7% 61.0% 62.1% 62.4% 62…
Figure 9
Figure 9. Figure 9: Test set performance. We test the performance of existing IQA and NR-SIQA models as well as our proposed model on a held out test sets, and show the results on different splits [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Distortion strength comparison. Comparing maximum distortion strengths across existing datasets and our proposed dataset, demonstrating that the distortions applied in our dataset exhibit significantly lower intensity compared to existing SIQA datasets. LIVE Phase I L…
Figure 11
Figure 11. Figure 11: Example from our off-the-shelf, mono-to-stereo ex￾periment. Three versions of the same stereo image are generated using different off-the-shelf, mono-to-stereo conversion methods. (a) Depthify.ai (b) Immersity AI (c) Owl3D. The stereo images are presented as anaglyph …
Figure 12
Figure 12. Figure 12: User study setups: left/right image toggling (top) and anaglyph stereo (bottom) 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 10 1.00 0.23 1.00 0.19 0.63 1.00 0.08 0.12 0.24 1.00 0.48 0.52 0.48 0.12 1.00 0.41 -0.02 0.03 0.29 0.37 1.00 0.56 0.44 0.32 0.12 0.36 0.03 1.00 0.31 …
Figure 13
Figure 13. Figure 13: Inter-rater agreement for each viewing medium, measured using Cohen’s kappa coefficient. The heatmap displays the agreement scores between all pairs of the 10 participants, highlighting the correlation in subjective evaluations across different viewing conditions [PI…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

120 extracted references · 67 canonical work pages

  1. [1]

    No-reference stereoscopic image quality as- sessment

    Roushain Akhter, ZM Parvez Sazzad, Yuukou Horita, and Jacky Baltes. No-reference stereoscopic image quality as- sessment. In Stereoscopic Displays and Applications XXI , pages 271–282. SPIE, 2010. 2

  2. [2]

    Apple introduces spatial video capture on iphone 15 pro

    Apple. Apple introduces spatial video capture on iphone 15 pro. https://www.apple.com/il/newsroom/ 2023/12/apple-introduces-spatial-video- capture- on- iphone- 15- pro/, 2023. [Accessed 16-10-2024]. 2

  3. [3]

    Visual fatigue caused by stereoscopic images and the search for the requirement to prevent them: A review

    Takehiko Bando, Atsuhiko Iijima, and Sumio Yano. Visual fatigue caused by stereoscopic images and the search for the requirement to prevent them: A review. Displays, 33 (2):76–83, 2012. 2

  4. [4]

    Lumiere: A space- time diffusion model for video generation

    Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Yuanzhen Li, Tomer Michaeli, et al. Lumiere: A space- time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024. 2

  5. [5]

    Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P

    Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. ICCV, 2021. 2

  6. [6]

    Mip-nerf 360: Unbounded anti-aliased neural radiance fields

    Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In CVPR, 2022. 3

  7. [7]

    Barron, Ben Mildenhall, Dor Verbin, Pratul P

    Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid- based neural radiance fields. ICCV, 2023. 2

  8. [8]

    Signature verification using a” siamese” time delay neural network

    Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard S¨ackinger, and Roopak Shah. Signature verification using a” siamese” time delay neural network. NeurIPS, 6, 1993. 5

Show all 120 references
  1. [9]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators

  2. [10]

    D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black. A naturalistic open source movie for optical flow evaluation. In ECCV, 2012. 3

  3. [11]

    The cognitive impenetrability of cogni- tion

    Patrick Cavanagh. The cognitive impenetrability of cogni- tion. Behavioral and Brain Sciences, 22(3):370–371, 1999. 6

  4. [12]

    Visual discomfort prediction on stereoscopic 3d images without explicit disparities

    Jianyu Chen, Jun Zhou, Jun Sun, and Alan Conrad Bovik. Visual discomfort prediction on stereoscopic 3d images without explicit disparities. Signal Processing: Image Communication, 51:50–60, 2017. 2

  5. [13]

    Pea-pods: Perceptual eval- uation of algorithms for power optimization in xr displays

    Kenneth Chen, Thomas Wan, Nathan Matsuda, Ajit Ninan, Alexandre Chapiro, and Qi Sun. Pea-pods: Perceptual eval- uation of algorithms for power optimization in xr displays. ACM Transactions on Graphics (TOG), 43(4):1–17, 2024. 3

  6. [14]

    Study of 3d virtual reality picture qual- ity

    Meixu Chen, Yize Jin, Todd Goodall, Xiangxu Yu, and Alan Conrad Bovik. Study of 3d virtual reality picture qual- ity. IEEE Journal of Selected Topics in Signal Processing, 14(1):89–102, 2019. 2

  7. [15]

    No-reference quality assessment of natural stereopairs

    Ming-Jun Chen, Lawrence K Cormack, and Alan C Bovik. No-reference quality assessment of natural stereopairs. IEEE Transactions on Image Processing, 22(9):3379–3391,

  8. [16]

    Full-reference quality assessment of stereopairs accounting for rivalry

    Ming-Jun Chen, Che-Chun Su, Do-Kyoung Kwon, Lawrence K Cormack, and Alan C Bovik. Full-reference quality assessment of stereopairs accounting for rivalry. Signal Processing: Image Communication , 28(9):1143– 1155, 2013. 3, 4

  9. [17]

    Svg: 3d stereoscopic video generation via denoising frame matrix

    Peng Dai, Feitong Tan, Qiangeng Xu, David Futschik, Ruofei Du, Sean Fanello, Xiaojuan Qi, and Yinda Zhang. Svg: 3d stereoscopic video generation via denoising frame matrix. arXiv preprint arXiv:2407.00367, 2024. 2

  10. [18]

    Depthify.ai — convert 2d videos to 3d spatial videos

    Depthify.ai. Depthify.ai — convert 2d videos to 3d spatial videos. http://depthify.ai, 2024. 8

  11. [19]

    No-reference stereoscopic image quality assessment using convolutional neural network for adaptive feature extraction

    Yong Ding, Ruizhe Deng, Xin Xie, Xiaogang Xu, Yang Zhao, Xiaodong Chen, and Andrey S Krylov. No-reference stereoscopic image quality assessment using convolutional neural network for adaptive feature extraction. IEEE Ac- cess, 6:37595–37603, 2018. 3

  12. [20]

    Variation and extrema of human inter- pupillary distance

    Neil A Dodgson. Variation and extrema of human inter- pupillary distance. In Stereoscopic displays and virtual re- ality systems XI, pages 36–46. SPIE, 2004. 3

  13. [21]

    A projection method to generate anaglyph stereo images

    Eric Dubois. A projection method to generate anaglyph stereo images. pages 1661 – 1664 vol.3, 2001. 1

  14. [22]

    Daniel Duckworth, Peter Hedman, Christian Reiser, Pe- ter Zhizhin, Jean-Franc ¸ois Thibert, Mario Lu ˇci´c, Richard Szeliski, and Jonathan T. Barron. Smerf: Streamable mem- ory efficient radiance fields for real-time large-scene explo- ration, 2023. 2

  15. [23]

    Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds

    Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, et al. Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds. arXiv preprint arXiv:2403.20309, 2024. 2

  16. [24]

    Stereoscopic image quality assessment by deep convolu- tional neural network

    Yuming Fang, Jiebin Yan, Xuelin Liu, and Jiheng Wang. Stereoscopic image quality assessment by deep convolu- tional neural network. Journal of Visual Communication and Image Representation, 58:400–406, 2019. 2, 3

  17. [25]

    Dream- sim: Learning new dimensions of human visual similarity using synthetic data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similarity using synthetic data. NeurIPS, 2023. 5

  18. [26]

    Perceptual requirements for world- locked rendering in ar and vr

    Phillip Guan, Eric Penner, Joel Hegland, Benjamin Letham, and Douglas Lanman. Perceptual requirements for world- locked rendering in ar and vr. In SIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 3

  19. [27]

    Deep blending for free-viewpoint image-based rendering

    Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering. 37(6): 257:1–257:15, 2018. 3

  20. [28]

    Lora: Low-rank adaptation of large language mod- els

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language mod- els. ICLR, 2022. 5

  21. [29]

    Holopix50k: A large-scale in-the-wild stereo image dataset

    Yiwen Hua, Puneet Kohli, Pritish Uplavikar, Anand Ravi, Saravana Gunaseelan, Jason Orozco, and Edward Li. Holopix50k: A large-scale in-the-wild stereo image dataset. In CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Seattle, WA, 2020., 2020. 3, 8

  22. [30]

    Open- clip, 2021

    Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Han- naneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Open- clip, 2021. If you use this software, please cite it as below. 6

  23. [31]

    Immersity ai — convert image and video to 3d

    Immersity AI. Immersity ai — convert image and video to 3d. https://www.immersity.ai, 2024. 8

  24. [32]

    Virtual Reality [VR] Market Size, Growth, Share — Report, 2032 — fortunebusinessinsights.com

    Fortune Business Insights. Virtual Reality [VR] Market Size, Growth, Share — Report, 2032 — fortunebusinessinsights.com. https : / / www . fortunebusinessinsights . com / industry - reports / virtual - reality - market - 101378,

  25. [33]

    Zero-shot text-guided object gen- eration with dream fields

    Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object gen- eration with dream fields. In CVPR, pages 867–876, 2022. 2

  26. [34]

    Saliency-based deep convolu- tional neural network for no-reference image quality as- sessment

    Sen Jia and Yang Zhang. Saliency-based deep convolu- tional neural network for no-reference image quality as- sessment. Multimedia Tools and Applications , 77:14859– 14872, 2018. 2

  27. [35]

    Re- purposing diffusion-based image generators for monocular depth estimation

    Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Re- purposing diffusion-based image generators for monocular depth estimation. In CVPR, 2024. 3

  28. [36]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4):1–14, 2023. 2, 3

  29. [37]

    Visual fatigue pre- diction for stereoscopic image

    Donghyun Kim and Kwanghoon Sohn. Visual fatigue pre- diction for stereoscopic image. IEEE transactions on cir- cuits and systems for video technology , 21(2):231–236,

  30. [38]

    Binocular fusion net: Deep learning visual comfort assessment for stereoscopic 3d

    Hak Gu Kim, Hyunwook Jeong, Heoun taek Lim, and Yong Man Ro. Binocular fusion net: Deep learning visual comfort assessment for stereoscopic 3d. IEEE Transactions on Circuits and Systems for Video Technology, 29:956–967,

  31. [39]

    Tanks and temples: Benchmarking large-scale scene reconstruction

    Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics, 36 (4), 2017. 3

  32. [40]

    Visual comfort of binoc- ular and 3d displays

    Frank L Kooi and Alexander Toet. Visual comfort of binoc- ular and 3d displays. Displays, 25(2-3):99–108, 2004. 2

  33. [41]

    Optimizing depth perception in virtual and augmented real- ity through gaze-contingent stereo rendering

    Brooke Krajancich, Petr Kellnhofer, and Gordon Wetzstein. Optimizing depth perception in virtual and augmented real- ity through gaze-contingent stereo rendering. ACM Trans. Graph., 39(6), 2020. 3

  34. [42]

    Imagenet classification with deep convolutional neural net- works

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. NeurIPS, 2012. 6

  35. [43]

    Visual discomfort and visual fa- tigue of stereoscopic displays: A review

    Marc Lambooij, Wijnand IJsselsteijn, Marten Fortuin, In- grid Heynderickx, et al. Visual discomfort and visual fa- tigue of stereoscopic displays: A review. Journal of imag- ing science and technology, 53(3):30201–1, 2009. 2, 5

  36. [44]

    Gradient-based learning applied to document recognition

    Yann LeCun, L ´eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324,

  37. [45]

    No-reference stereoscopic image quality assessment based on visual attention and percep- tion

    Yafei Li, Feng Yang, Wenbo Wan, Jun Wang, Min Gao, Jia Zhang, and Jiande Sun. No-reference stereoscopic image quality assessment based on visual attention and percep- tion. IEEE Access, 7:46706–46716, 2019. 2, 3

  38. [46]

    Vr iqa net: Deep virtual reality image quality assessment using ad- versarial learning

    Heaun-Taek Lim, Hak Gu Kim, and Yang Man Ra. Vr iqa net: Deep virtual reality image quality assessment using ad- versarial learning. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6737–6741. IEEE, 2018. 2

  39. [47]

    Magic3d: High- resolution text-to-3d content creation

    Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High- resolution text-to-3d content creation. In CVPR, 2023. 2

  40. [48]

    Learning based no-reference metric for assessing quality of experience of stereoscopic images

    Tsung-Jung Liu, Kuan-Hsien Liu, and Kuan-Hung Shen. Learning based no-reference metric for assessing quality of experience of stereoscopic images. Journal of Visual Com- munication and Image Representation , 61:272–283, 2019. 2, 3, 4

  41. [49]

    No- reference stereoscopic image quality evaluator with seg- mented monocular features and perceptual binocular fea- tures

    Yun Liu, Chang Tang, Zhi Zheng, and Liyuan Lin. No- reference stereoscopic image quality evaluator with seg- mented monocular features and perceptual binocular fea- tures. Neurocomputing, 405:126–137, 2020. 2, 3

  42. [50]

    Spatialdreamer: Self-supervised stereo video synthesis from monocular in- put

    Zhen Lv, Yangqi Long, Congzhentao Huang, Cao Li, Chengfei Lv, Hao Ren, and Dian Zheng. Spatialdreamer: Self-supervised stereo video synthesis from monocular in- put. IEEE VR, 2024. 2

  43. [51]

    Realistic luminance in vr

    Nathan Matsuda, Alex Chapiro, Yang Zhao, Clinton Smith, Romain Bachy, and Douglas Lanman. Realistic luminance in vr. In SIGGRAPH Asia 2022 Conference Papers, pages 1–8, 2022. 3

  44. [52]

    Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo

    Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, and Andr ´es Bruhn. Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo. In CVPR, 2023. 3, 8, 4

  45. [53]

    Sdedit: Guided image synthesis and editing with stochastic differential equations

    Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jia- jun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. ICLR, 2022. 1, 3

  46. [54]

    Object scene flow for autonomous vehicles

    Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In CVPR, 2015. 3

  47. [55]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 2

  48. [56]

    No-reference image quality assessment in the spatial domain

    Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695–4708, 2012. 2, 6

  49. [57]

    Subjective evaluation of stereoscopic image quality

    Anush Krishna Moorthy, Che-Chun Su, Anish Mittal, and Alan Conrad Bovik. Subjective evaluation of stereoscopic image quality. Signal Processing: Image Communication , 28(8):870–883, 2013. 3, 6, 4

  50. [58]

    Beyond reality: Is the long-awaited vr revolution finally on the horizon?, 2022

    NRG. Beyond reality: Is the long-awaited vr revolution finally on the horizon?, 2022. 2

  51. [59]

    Deep visual discomfort predictor for stereo- scopic 3d images

    Heeseok Oh, Sewoong Ahn, Sanghoon Lee, and Alan Con- rad Bovik. Deep visual discomfort predictor for stereo- scopic 3d images. IEEE Transactions on Image Processing, 27(11):5420–5432, 2018. 2

  52. [60]

    Dinov2: Learning robust visual features without supervi- sion

    Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervi- sion. TMLR, 2024. 5, 6, 8

  53. [61]

    Owl3d — ai-powered 2d to 3d conversion soft- ware

    Owl3D. Owl3d — ai-powered 2d to 3d conversion soft- ware. https://www.owl3d.com, 2024. 8

  54. [62]

    3d visual discomfort predictor: Analysis of dis- parity and neural activity statistics

    Jincheol Park, Heeseok Oh, Sanghoon Lee, and Alan Con- rad Bovik. 3d visual discomfort predictor: Analysis of dis- parity and neural activity statistics. IEEE transactions on image processing, 24(3):1101–1114, 2014. 2, 3

  55. [63]

    Barron, and Ben Milden- hall

    Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. ICLR,

  56. [64]

    The dominant eye

    Clare Porac and Stanley Coren. The dominant eye. Psycho- logical bulletin, 83(5):880, 1976. 5

  57. [65]

    No-reference virtual reality im- age quality evaluator using global and local natural scene statistics

    Ajay Kumar Reddy Poreddy, Raja Bharath Chandra Ganeswaram, Balasubramanyam Appina, Priyanka Kokil, and Ram Bilas Pachori. No-reference virtual reality im- age quality evaluator using global and local natural scene statistics. IEEE Transactions on Instrumentation and Mea- surem...

  58. [66]

    A computa- tional model for perception of stereoscopic window viola- tions

    Steven Poulakos, Rafael Monroy, Tunc Aydin, Oliver Wang, Aljoscha Smolic, and Markus Gross. A computa- tional model for perception of stereoscopic window viola- tions. In 2015 Seventh International Workshop on Qual- ity of Multimedia Experience (QoMEX) , pages 1–6. IEEE,

  59. [67]

    Qual- ity of experience assessment for stereoscopic images

    Feng Qi, Tingting Jiang, Siwei Ma, and Debin Zhao. Qual- ity of experience assessment for stereoscopic images. In 2012 IEEE International Symposium on Circuits and Sys- tems (ISCAS), pages 1712–1715, 2012. 2

  60. [68]

    Learn- ing transferable visual models from natural language super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 5, 6

  61. [69]

    Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer

    Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(3), 2022. 3

  62. [70]

    RED HYDROGEN Holographic Smart- phone - first look — dxomark.com

    Lars Rehm. RED HYDROGEN Holographic Smart- phone - first look — dxomark.com. https://www. dxomark . com / red - hydrogen - holographic - smartphone-first-look/ , 2018. [Accessed 25-11- 2024]. 2

  63. [71]

    Capturing the Future: Exploring the Ever- Increasing Popularity of Dual Camera Smartphones — treendly.com

    Mike Rubini. Capturing the Future: Exploring the Ever- Increasing Popularity of Dual Camera Smartphones — treendly.com. https : / / treendly . com / blog / capturing - the - future - exploring - the - ever - increasing - popularity - of - dual - camera-smartphones, 2023. [Acce...

  64. [72]

    High-resolution stereo datasets with subpixel-accurate ground truth

    Daniel Scharstein, Heiko Hirschm ¨uller, York Kitajima, Greg Krathwohl, Nera Ne ˇsi´c, Xi Wang, and Porter West- ling. High-resolution stereo datasets with subpixel-accurate ground truth. In GCPR, 2014. 3

  65. [73]

    Dual Camera Smartphones Are Go- ing Mainstream: Explosive Growth in 2018 — counterpointresearch.com

    Neil Shah. Dual Camera Smartphones Are Go- ing Mainstream: Explosive Growth in 2018 — counterpointresearch.com. https : / / www . counterpointresearch.com/insights/dual- camera - smartphones - going - mainstream - explosive - growth - 2018/, 2018. [Accessed 16-10-2024]. 2

  66. [74]

    No-reference stereoscopic image quality as- sessment based on image distortion and stereo perceptual information

    Liquan Shen, Ruigang Fang, Yang Yao, Xianqiu Geng, and Dapeng Wu. No-reference stereoscopic image quality as- sessment based on image distortion and stereo perceptual information. IEEE Transactions on Emerging Topics in Computational Intelligence, 3(1):59–72, 2018. 2, 3

  67. [75]

    No-reference stereoscopic image qual- ity assessment based on global and local content character- istics

    Lili Shen, Xiongfei Chen, Zhaoqing Pan, Kefeng Fan, Fei Li, and Jianjun Lei. No-reference stereoscopic image qual- ity assessment based on global and local content character- istics. Neurocomputing, 424:132–142, 2021. 2, 3

  68. [76]

    Immersepro: End- to-end stereo video synthesis via implicit disparity learning

    Jian Shi, Zhenyu Li, and Peter Wonka. Immersepro: End- to-end stereo video synthesis via implicit disparity learning. arXiv preprint arXiv:2410.00262, 2024. 2

  69. [77]

    The zone of comfort: Predicting visual discomfort with stereo displays

    Takashi Shibata, Joohwan Kim, David M Hoffman, and Martin S Banks. The zone of comfort: Predicting visual discomfort with stereo displays. Journal of vision , 11(8): 11–11, 2011. 2

  70. [78]

    A no-reference stereoscopic image quality assessment network based on binocular interaction and fu- sion mechanisms

    Jianwei Si, Baoxiang Huang, Huan Yang, Weisi Lin, and Zhenkuan Pan. A no-reference stereoscopic image quality assessment network based on binocular interaction and fu- sion mechanisms. IEEE Transactions on Image Processing, 31:3066–3080, 2022. 2, 3

  71. [79]

    Cognitive penetrability of perception

    Dustin Stokes. Cognitive penetrability of perception. Phi- losophy Compass, 8(7):646–663, 2013. 6

  72. [80]

    Per- ceptual quality assessment of omnidirectional images as moving camera videos

    Xiangjie Sui, Kede Ma, Yiru Yao, and Yuming Fang. Per- ceptual quality assessment of omnidirectional images as moving camera videos. IEEE Transactions on Visualiza- tion and Computer Graphics, 28(8):3022–3034, 2021. 2

  73. [81]

    Resolution-robust large mask inpainting with fourier convolutions

    Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. In WACV, pages 2149–2159, 2022. 3

  74. [82]

    Examining user perception of the size of multiple objects in virtual reality.Applied Sciences, 10(11): 4049, 2020

    Bruce H Thomas. Examining user perception of the size of multiple objects in virtual reality.Applied Sciences, 10(11): 4049, 2020. 3

  75. [83]

    Visual fatigue caused by viewing stereoscopic motion images: Background, the- ories, and observations

    Kazuhiko Ukai and Peter A Howarth. Visual fatigue caused by viewing stereoscopic motion images: Background, the- ories, and observations. Displays, 29(2):106–116, 2008. 2

  76. [84]

    Quality predic- tion of asymmetrically distorted stereoscopic images from single views

    Jiheng Wang, Kai Zeng, and Zhou Wang. Quality predic- tion of asymmetrically distorted stereoscopic images from single views. In 2014 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE, 2014. 3, 4

  77. [85]

    Quality prediction of asymmetrically distorted stereoscopic 3d images

    Jiheng Wang, Abdul Rehman, Kai Zeng, Shiqi Wang, and Zhou Wang. Quality prediction of asymmetrically distorted stereoscopic 3d images. IEEE Transactions on Image Pro- cessing, 24(11):3400–3414, 2015. 3, 4

  78. [86]

    Ex- ploring clip for assessing the look and feel of images

    Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Ex- ploring clip for assessing the look and feel of images. In AAAI, 2023. 2, 6

  79. [87]

    Modelscope text-to-video technical report

    Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 2

  80. [88]

    Flickr1024: A large-scale dataset for stereo image super-resolution

    Yingqian Wang, Longguang Wang, Jungang Yang, Wei An, and Yulan Guo. Flickr1024: A large-scale dataset for stereo image super-resolution. In International Conference on Computer Vision Workshops, pages 3852–3857, 2019. 3

  81. [89]

    Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation

    Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongx- uan Li, Hang Su, and Jun Zhu. Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation. NeurIPS, 36, 2024. 2

  82. [90]

    Motionctrl: A unified and flexible motion controller for video generation

    Zhouxia Wang, Ziyang Yuan, Xintao Wang, Tianshui Chen, Menghan Xia, Ping Luo, and Yin Shan. Motionctrl: A unified and flexible motion controller for video generation

  83. [91]

    Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion, 2022

    Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Br´egier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and J ´erˆome Re- vaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion, 2022. 6

  84. [92]

    Croco v2: Improved cross-view completion pre- training for stereo matching and optical flow, 2023

    Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Br ´egier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, and J ´erˆome Revaud. Croco v2: Improved cross-view completion pre- training for stereo matching and optical flow, 2023. 6

  85. [93]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al. Q-align: Teaching lmms for visual scoring via discrete text-defined levels. arXiv preprint arXiv:2312.17090, 2023. 2, 6

  86. [94]

    Depth anything: Un- leashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Ji- ashi Feng, and Hengshuang Zhao. Depth anything: Un- leashing the power of large-scale unlabeled data. In CVPR,

  87. [95]

    Maniqa: Multi-dimension attention network for no- reference image quality assessment

    Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. Maniqa: Multi-dimension attention network for no- reference image quality assessment. In CVPR, 2022. 2, 6, 7, 8

  88. [96]

    Cogvideox: Text-to- video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to- video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2

  89. [97]

    Two factors in visual fatigue caused by stereoscopic hdtv im- ages

    Sumio Yano, Masaki Emoto, and Tetsuo Mitsuhashi. Two factors in visual fatigue caused by stereoscopic hdtv im- ages. Displays, 25(4):141–150, 2004. 2

  90. [98]

    Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models

    Taoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Qi Tian, and Xing- gang Wang. Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models. In CVPR, pages 6796–6807, 2024. 2

  91. [99]

    Mip-splatting: Alias-free 3d gaussian splatting

    Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3d gaussian splatting. In CVPR, 2024. 2

  92. [100]

    Towards top- down stereoscopic image quality assessment via stereo at- tention

    Huilin Zhang, Sumei Li, and Yongli Chang. Towards top- down stereoscopic image quality assessment via stereo at- tention. arXiv preprint arXiv:2308.04156, 2023. 2, 3

  93. [101]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 5

  94. [102]

    Learning structure of stereoscopic image for no- reference quality assessment with convolutional neural net- work

    Wei Zhang, Chenfei Qu, Lin Ma, Jingwei Guan, and Rui Huang. Learning structure of stereoscopic image for no- reference quality assessment with convolutional neural net- work. Pattern Recognition, 59:176–187, 2016. 2, 3

  95. [103]

    Stereocrafter: Diffusion-based generation of long and high-fidelity stereoscopic 3d from monocular videos

    Sijie Zhao, Wenbo Hu, Xiaodong Cun, Yong Zhang, Xi- aoyu Li, Zhe Kong, Xiangjun Gao, Muyao Niu, and Ying Shan. Stereocrafter: Diffusion-based generation of long and high-fidelity stereoscopic 3d from monocular videos. arXiv preprint arXiv:2409.07447, 2024. 2

  96. [104]

    Dual-stream inter- active networks for no-reference stereoscopic image quality assessment

    Wei Zhou, Zhibo Chen, and Weiping Li. Dual-stream inter- active networks for no-reference stereoscopic image quality assessment. IEEE Transactions on Image Processing , 28 (8):3946–3958, 2019. 2, 6, 7, 8, 3

  97. [105]

    Stereoscopic image discomfort prediction us- ing dual-stream multi-level interactive network

    Yang Zhou, Pingan Chen, Haibing Yin, Xiaofeng Huang, and Zhu Li. Stereoscopic image discomfort prediction us- ing dual-stream multi-level interactive network. Displays, 78:102444, 2023. 2

  98. [106]

    Esiqa: Perceptual quality assessment of vision-pro-based egocentric spatial images

    Xilei Zhu, Liu Yang, Huiyu Duan, Xiongkuo Min, Guang- tao Zhai, and Patrick Le Callet. Esiqa: Perceptual quality assessment of vision-pro-based egocentric spatial images. arXiv preprint arXiv:2407.21363, 2024. 6 What Makes for a Good Stereoscopic Image? Supplementary Material

  99. [108]

    Table 3 presents the total amount in which each distortion appears in the dataset

    Dataset Details Our dataset is comprised of2400 examples, each containing a pair of stereo images, resulting in a total of4800 stereo im- ages, all of which have undergone some form of distortion. Table 3 presents the total amount in which each distortion appears in the datase...

  100. [109]

    We maintain the original 1280 × 720 resolution, only applying a center crop to 1274 × 714 for compatibility with DINOv2-S’s patch size of 14

    Training Details Our model is trained on a single NVIDIA A100 using an Adam optimizer with a learning rate of3e−5 and batch size of 16. We maintain the original 1280 × 720 resolution, only applying a center crop to 1274 × 714 for compatibility with DINOv2-S’s patch size of 14....

  101. [110]

    The sizes of these splits are similar, comprising 32.9%, 34.1%, and 32.9% of the data respectively, confirming that our dataset contains a learnable signal

    Detailed Performance on the SCOPE dataset In Figure 9 we report test set accuracy across several differ- ent train, validation and test partitions, categorized by anno- tation consensus: unanimous ( 5 − 0 split), majority (4 − 1 split), and divided ( 3 − 2 split), with the lat...

  102. [111]

    SCOPE differs from them in several aspects:

    Performance on Existing SIQA Datasets Table 5 provides a comparison between our dataset and ex- isting stereo quality assessment datasets: LIVE 3D Phases I and II [16, 57], Waterloo IVC (WIVC) 3D Phases I and II [84, 85] and IEEE-SA [48]. SCOPE differs from them in several aspects:

  103. [112]

    Image Quantity: SCOPE is the largest of the datasets, with more than twice the amount of samples than IEEE- SA - the second largest dataset

  104. [113]

    In Section 4.3 and Figure 6 we demonstrate low correlation between preferences on VR devices and other stereo viewing methods

    Annotation medium: The annotations in all these datasets were collected using passive stereoscopic dis- plays or active shutter glasses, while ours were collected on a Vision Pro headset. In Section 4.3 and Figure 6 we demonstrate low correlation between preferences on VR devi...

  105. [114]

    Annotation Protocol: The other datasets collected Mean Opinion Score annotations, an absolute single- image protocol, while SCOPE collected 2AFC which are relative annotations

  106. [115]

    We evaluate our model on LIVE 3D Phase I and II [15, 16, 57] and Waterloo IVC (WIVC) 3D Phase I and II [84, 85]

    Distortion Strengths: The other datasets applied sig- nificantly stronger distortions than SCOPE, see Fig- ure 10. We evaluate our model on LIVE 3D Phase I and II [15, 16, 57] and Waterloo IVC (WIVC) 3D Phase I and II [84, 85]. For these evaluations, we use standard per- forma...

  107. [116]

    Fast-fading LIVE 360 360 8 DMOS Noise, Blur, Phase II Compression,

  108. [117]

    Stereoscopic Preference Datasets

    Fast-fading WIVC 330 330 6 MOS Noise, Blur Phase I [84] WIVC 460 460 10 MOS Noise, Blur, Phase II Compression, [85] IEEE-SA [48]800 800 160 MOS Horizontal disparity SCOPE 2400 4800 2400 2AFC 19types, (Ours) see Table 1 Table 5. Stereoscopic Preference Datasets. Prior datasets ...

  109. [118]

    Viewing stereoscopic images with the Apple Vision Pro was done through the native photos app in immersive mode

    Cross-Medium User Study Expanding on the user study outlined in Section 4.3, we de- tail the specific viewing setups for each device. Viewing stereoscopic images with the Apple Vision Pro was done through the native photos app in immersive mode. For the Meta Quest Pro, we empl...

  110. [119]

    Figure 11 shows an example from the user study

    Off-the-Shelf Mono-to-Stereo Evaluation We evaluated alignment of human opinion with the differ- ent SQoE candidates on the Spring [52] dataset. Figure 11 shows an example from the user study

  111. [120]

    Licenses The models and datasets we use are provided under the li- censes in Table 6. Dataset License Model License Tanks and Temples CC BY 4.0MotionCtrl Apache 2.0Deep Blending Apache 2.0 MiDaS MITMip-NeRF 360 Apache 2.0 Marigold Apache 2.0Holopix50k NC Depth Anything Apache ...

  112. [2024]

    [Accessed 16-10-2024]. 2

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.