Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision

T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This survey claims to be the first comprehensive map of deep learning for omnidirectional vision, organizing imaging principles, projection formats, datasets, representation learning, optimization strategies, and applications into one…

desk verdict A genuinely broad 360° vision survey with a useful taxonomy and benchmark tables, but the spherical projection equation in Sec. 2.2 is internally inconsistent and the 'comprehensive' claim lacks a documented selection process. read the letter →

arxiv 2502.10444 v1 pith:53KNJAMJ submitted 2025-02-11 cs.CV

classification cs.CV
keywords omnidirectionalvision360-degreeimagesdeeplearningsurveysphericalrepresentationequirectangularprojectionomni-directionalimagetaskspanoramicsceneunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish that deep learning for omnidirectional (360-degree) vision can be organized into a single coherent structure: how 360-degree images are captured and projected, how representations are learned from them, how models are optimized when annotations are scarce, and how the resulting methods sort into tasks from generation and super-resolution to depth, optical flow, and room layout estimation. The authors claim this is the first survey to cover all of these together, rather than a single sub-domain as earlier surveys did. A sympathetic reader's takeaway is a previously unavailable reference map: over 200 recent works placed in a taxonomy, benchmark tables for direct comparison, discussion of open problems, and a maintained code repository to make the map usable.

What carries the argument

The carrying object is the hierarchical taxonomy (Fig. 2 of the paper), which classifies methods along three axes: representation learning (Euclidean methods on ERP-style planes versus non-Euclidean methods on spherical meshes and graphs), optimization strategy (unsupervised and semi-supervised, transfer, multi-task, deep reinforcement learning), and task family (generation, super-resolution, quality assessment, detection, segmentation, saliency, depth, optical flow, room layout, SLAM). It rests on the imaging-projection framework of Sec. 2, where the sphere is the native domain and the equirectangular projection is the default planar format whose pole distortion is the recurring problem, and on the per-task comparison tables that rank representative methods. The taxonomy does the argumentative work: it turns many individual papers into a map and supports the paper's claims about which directions are settled and which are open.

What would settle it

A reader can test the map's completeness with a fixed protocol: enumerate every paper published in a defined set of major vision venues over the last five years whose title or abstract matches 360-degree, panoramic, or omnidirectional deep learning, then check whether the survey's taxonomy and benchmark tables systematically omit any of them. The claim of being the first comprehensive survey can be checked directly by searching for any earlier review that already covers representation learning, optimization, and applications for omnidirectional vision together.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that the development of deep learning for omnidirectional vision has matured enough to be reviewed end to end, and that it can be organized by three cross-cutting concerns: representation learning on spherical data (Sec. 3), optimization strategies beyond plain supervision (Sec. 4), and task families (Sec. 5, from visual enhancement through scene understanding to 3D geometry and motion estimation). The paper asserts that this is the first comprehensive review of that development, that its roughly 200 cited works are representative of top-tier output in the last five years, and that its hierarchical taxonomy plus quantitative tables give readers intra-task comparisons that were not previously collected in one place. If the paper is right, the field now has a unified entry point and a shared map of where it stands.

Load-bearing premise

The whole map stands on the assumption that the roughly 200 cited works were selected to be representative of top-tier research in the field, so that the taxonomy and the state-of-the-art tables are a fair picture rather than an author-selected subset.

Editorial extensions

If this is right

  • A newcomer to 360-degree vision can use the taxonomy to locate any method by its projection choice, learning strategy, and task, and the benchmark tables to compare reported results without reading dozens of papers.
  • The paper's structure makes the field's recurring trade-offs explicit, especially the tension between ERP's convenience and its pole distortion, and between planar projections' compatibility with pretrained models and their discontinuity, which future method design must navigate.
  • The identified open problems, such as data-efficient learning, panoramic optical aberration correction, multi-modal spherical understanding, and robustness to adversarial attacks, point to where the next wave of work is most likely to land.
  • The maintained open-source repository with code links gives the community a single access point to implementations, lowering the barrier to reproducing and building on existing work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My reading of the paper's own comparison tables is that the winning recipe in monocular depth estimation combines less-distorted projections (tangent rather than cubemap) with attention or transformer backbones, a template that plausibly generalizes to other spherical dense-prediction tasks.
  • The novelty claim of being the first comprehensive survey is directly testable: a systematic literature search for any earlier review that already unites representation learning, optimization, and applications for omnidirectional vision would settle it; the paper offers no search protocol to back the claim.
  • Because no inclusion or exclusion criteria are given for the roughly 200 works, the taxonomy's neutrality cannot be verified from the paper alone; annotating each entry with venue, year, and selection basis would make the map auditable.
  • The paper frames spherical transformers mostly as a future direction, yet its own tables show transformer-based methods leading in classification and depth estimation, suggesting this direction will grow faster than the survey's cautious framing implies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript surveys deep learning for omnidirectional vision, covering acquisition and projection formats (ERP, CP, TP, polyhedron, and others), datasets, representation learning (Euclidean and non-Euclidean), optimization strategies (unsupervised/semi-supervised, transfer, multi-task, and deep reinforcement learning), and a taxonomy of tasks ranging from visual enhancement and scene understanding to 3D geometry and motion estimation. It also discusses applications such as AR/VR, robot navigation, and autonomous driving, and concludes with challenges and future directions. The authors claim that this is the first comprehensive survey of deep learning for omnidirectional vision, summarize over 200 representative works, provide benchmark tables, and maintain an open-source repository with code links.

Significance. If the technical content is corrected, this survey could be a valuable unified entry point to the field. Its strengths include the hierarchical taxonomy in Fig. 2, the cross-task organization of methods, benchmark tables (Tables 2-8) that allow quick comparisons, and the openly maintained repository. The paper goes beyond listing papers; it draws useful cross-task insights, such as the trade-offs between distortion-aware convolutions and attention mechanisms, and identifies underexplored directions such as 3D robustness for panoramic segmentation and panoramic panoptic segmentation. However, the foundational projection equations in Sec. 2.2 contain a concrete mathematical inconsistency, and the claim of comprehensiveness is not backed by a documented selection methodology. Both issues must be resolved before the survey can be relied upon as an authoritative reference.

major comments (2)
  1. [§2.2, Eq. (1)] Eq. (1) is internally inconsistent with the spherical-coordinate convention stated immediately above it. The text defines p = [sinθ cosϕ, sinθ sinϕ, cosθ]^T, so θ is the polar angle and ϕ is the azimuth angle; the correct inverse is then θ = arccos(z/ρ), ϕ = arctan2(y,x), and the correct forward mapping is x = ρ sinθ cosϕ, y = ρ sinθ sinϕ, z = ρ cosθ. Instead, Eq. (1) gives θ = arctan(x/z), ϕ = arccos(y/ρ), and x = ρ sinθ sinϕ, y = ρ cosϕ, z = ρ cosθ sinϕ. These formulas do not round-trip: for θ=π/2 and ϕ=0 the stated definition gives p=(1,0,0), while the forward formula gives (0,ρ,0). Because the ERP mapping (u,v)→(θ,ϕ), the tangent projection in Eqs. (2)-(3), and many distortion-aware methods reviewed later depend on this convention, Eq. (1) must be corrected or the convention must be stated explicitly and used consistently.
  2. [§1, Contributions (I)-(II); §2.3] The central claim of being the "first comprehensive" survey and of covering "over 200 representative published top-tier works" is not supported by a systematic methodology. The manuscript does not report search databases, a time window, keywords, inclusion/exclusion criteria, screening steps, or a log of excluded works. Since the taxonomy and the state-of-the-art tables are presented as a map of the field, the authors should either document the selection process or explicitly qualify the selection as author-curated rather than comprehensive; otherwise a reader cannot assess possible selection bias.
minor comments (5)
  1. [Abstract / §1] The abstract says the paper covers "four main contents" and then lists five items, (i)-(v). This mismatch should be fixed, for example by replacing "four" with "five" or by merging two items.
  2. [Table 3] In the quantitative VQA comparison, the row for Assessor360 is cited as [72], but reference [72] is the AHGCN paper; the text identifies Assessor360 as [251]. This citation duplication should be corrected so that the table entries can be traced.
  3. [References [342] and [343]] References [342] and [343] are the same work ("Spherical view synthesis for self-supervised 360 depth estimation"), yet the text in §5.3.1 cites both as if they were distinct. The duplicate entry should be removed and the citations aligned.
  4. [Table 1] The data size for Deep360 is listed as "1,2000 (RGB)"; this appears to be a typo for 12,000. Please verify and correct the entry.
  5. [§5.2.3] The saliency method SalGAIL is attributed to Ma et al. [86], but reference [86] is the generative adversarial imitation learning paper; the actual SalGAIL work appears to be [265]/[266]. The citations should be re-checked and assigned to the correct publications.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's organization and comparisons do not reduce to their inputs, and the projection-equation inconsistency is a correctness issue, not a circularity issue.

full rationale

This is a survey paper, not a derivation chain, so there is no result whose conclusion is equivalent to its input. The central claims are organizational: providing a taxonomy, summarizing over 200 works, and listing benchmark tables. Those claims are supported by external literature and by comparative tables that include many non-author methods alongside the authors' own works (e.g., HRDFuse in Table 7 and GoodSAM in Table 4, both compared against independent baselines). The authors' self-citations appear in the benchmark tables and in the advertised open-source repository, but the survey's conclusions do not depend on accepting those self-cited papers as premises; the same taxonomy and comparisons would stand if the authors' own entries were removed. The absence of a formal search protocol is a completeness concern, not a circularity concern. The internally inconsistent spherical projection formulas in Eq. (1) are a definite correctness defect in the paper's foundational exposition, but a wrong mapping is not a case of deriving a conclusion from itself; it is a mathematical error that would undermine reproducibility, and it belongs in a correctness review rather than a circularity review. No step in the paper reduces by construction to its inputs, no fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work to force a choice. The verdict is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters, equations, or new entities are introduced; the survey is a synthesis of existing literature.

assumptions (4)
  • domain assumption The set of reviewed works is representative and comprehensive enough for the survey's claims.
    Used throughout the survey to support 'comprehensive' and 'state-of-the-art' claims; no systematic protocol is reported (Sec. 1, contribution II).
  • domain assumption Quantitative results in Tables 2-8 are correctly transcribed from the cited papers and are comparable.
    The tables compile single numbers from heterogeneous sources (e.g., Table 7 depth metrics); no code or raw outputs are provided, and citations appear to contain errors (Table 1).
  • domain assumption The taxonomy in Fig. 2 is a faithful organization of the field.
    The taxonomy is devised by the authors; there is no derivation or external validation.
  • ad hoc to paper The open-source repository exists and contains the claimed up-to-date taxonomy and code links.
    Contribution (V) relies on a URL without a commit hash or archived snapshot, so the artifact cannot be verified from the manuscript.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision." pith.science (2026). https://pith.science/paper/53KNJAMJ

@misc{pith2026250210444,
  author       = {Pith},
  title        = {Pith review of: A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/53KNJAMJ}},
  note         = {Machine review of arXiv:2502.10444}
}
read the original abstract

Omnidirectional image (ODI) data is captured with a field-of-view of 360x180, which is much wider than the pinhole cameras and captures richer surrounding environment details than the conventional perspective images. In recent years, the availability of customer-level 360 cameras has made omnidirectional vision more popular, and the advance of deep learning (DL) has significantly sparked its research and applications. This paper presents a systematic and comprehensive review and analysis of the recent progress of DL for omnidirectional vision. It delineates the distinct challenges and complexities encountered in applying DL to omnidirectional images as opposed to traditional perspective imagery. Our work covers four main contents: (i) A thorough introduction to the principles of omnidirectional imaging and commonly explored projections of ODI; (ii) A methodical review of varied representation learning approaches tailored for ODI; (iii) An in-depth investigation of optimization strategies specific to omnidirectional vision; (iv) A structural and hierarchical taxonomy of the DL methods for the representative omnidirectional vision tasks, from visual enhancement (e.g., image generation and super-resolution) to 3D geometry and motion estimation (e.g., depth and optical flow estimation), alongside the discussions on emergent research directions; (v) An overview of cutting-edge applications (e.g., autonomous driving and virtual reality), coupled with a critical discussion on prevailing challenges and open questions, to trigger more research in the community.

Figures

Figures reproduced from arXiv: 2502.10444 by the authors.

Figure 1
Figure 1. Overview of representation learning, optimization [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Hierarchical and structural taxonomy of omnidirectional vision with deep learning. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Examples of 360◦ cameras: (a) RICOH Theta Z1, and (b) GoPro Omni. reality (VR), autonomous driving, and robot naviga￾tion. As raw ODIs are captured on the sphere ( [PITH_FULL_IMAGE:figures/full_fig_p002_3.png] view at source ↗
Figures from the paper (18 more)
Figure 5
Figure 5. Figure 5: Illustration of the process for stitching a pair of [PITH_FULL_IMAGE:figures/full_fig_p003_5.png]
Figure 6
Figure 6. Figure 6: Spherical projection v.s. perspective projection. of the two fisheye cameras, and (2) limited information and severe distortions in the overlapping regions between the two cameras ( [PITH_FULL_IMAGE:figures/full_fig_p004_6.png]
Figure 7
Figure 7. Figure 7: Illustration of several representative projections: (a) ERP; (b) CP; (c) TP. [PITH_FULL_IMAGE:figures/full_fig_p005_7.png]
Figure 8
Figure 8. Figure 8: Top: Icosahedron projection; Bottom: Subdivi￾sion of an icosahedron’s face. The projection from Ps(θ, ϕ) to Pt(ut, vt) is defined as: ut = cos(ϕ) sin(θ − θc) cos c , vt = cos(ϕc) sin(ϕ) − sin(ϕc) cos(ϕ) cos(θ − θc) cos(c) , cos(c) = sin(ϕc) sin(ϕ) + cos(ϕc) cos(ϕ) cos(…
Figure 9
Figure 9. Figure 9: SPH-MNIST (top) & SPH-CIFAR10 (down). pitch attention to address the edge discontinuity and distortions of ERP images. 3.2 Non-Euclidean space As the raw spherical data possesses a non-Euclidean spherical data structure, some methods have explored the Non-Euclidean con…
Figure 10
Figure 10. Figure 10: Illustrations of representative representations on Euclidean space: (a) Spherical convolution [ [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: Illustrations of representative ODI representation learning methods on non-Euclidean space: (a) HexNet [ [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]
Figure 12
Figure 12. Figure 12: The input in unsupervised or semi-supervised [PITH_FULL_IMAGE:figures/full_fig_p009_12.png]
Figure 15
Figure 15. Figure 15: The common framework of DRL-based rate adaptation for 360◦ video streaming. resolution, limited network bandwidth is challenging for transmitting ODVs and proving high Quality of Experi￾ence (QoE) of users. Thus, it is to valuable to explore how to effectively allocat…
Figure 16
Figure 16. Figure 16: The illustrations of representative networks for ODI Generation: (a) GAN-based outpainting [ [PITH_FULL_IMAGE:figures/full_fig_p011_16.png]
Figure 17
Figure 17. Figure 17: Quantitative results on the SUN360 dataset. (a) [PITH_FULL_IMAGE:figures/full_fig_p013_17.png]
Figure 18
Figure 18. Figure 18: The illustrations of representative net [PITH_FULL_IMAGE:figures/full_fig_p013_18.png]
Figure 19
Figure 19. Figure 19: Typical pipelines of NR ODI-VQA: (a) CP patch-based approach; (b) TP patch-based approach; (c) Sequence-based approach. the spherical domain. It samples a limited number of points uniformly distributed on the sphere to calcu￾late the mean error. Additionally, L-PSNR […
Figure 20
Figure 20. Figure 20: (a) conventional planar bounding boxes (BBs) [PITH_FULL_IMAGE:figures/full_fig_p016_20.png]
Figure 21
Figure 21. Figure 21: Performance of SoTA panoramic segmentation [PITH_FULL_IMAGE:figures/full_fig_p017_21.png]
Figure 22
Figure 22. Figure 22: Typical pipelines of ODI saliency prediction: [PITH_FULL_IMAGE:figures/full_fig_p019_22.png]
Figure 23
Figure 23. Figure 23: Representative monocular depth estimation [PITH_FULL_IMAGE:figures/full_fig_p021_23.png]
Figure 24
Figure 24. Figure 24: Representative room layout estimation methods: [PITH_FULL_IMAGE:figures/full_fig_p023_24.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-branch optical flow network using a 90-degree rotated 'orthogonal' view reduces polar distortion errors and sets new state-of-the-art results on MPFDataset and FlowScape.

Reference graph

Works this paper leans on

299 extracted references · 78 canonical work pages · cited by 1 Pith paper

  1. [1]

    O’Connor

    Yasser Abdelaziz, Dahou Djilali, Tarun Krishna, Kevin McGuinness, and Noel E. O’Connor. Rethinking◦ 360 im- age visual attention modelling with unsupervised learning. ICCV, 2021

  2. [2]

    Hrdfuse: Monocular 360 depth estimation by collab- orativelylearningholistic-with-regionaldepthdistributions

    Hao Ai, Zidong Cao, Yan-Pei Cao, Ying Shan, and Lin Wang. Hrdfuse: Monocular 360 depth estimation by collab- orativelylearningholistic-with-regionaldepthdistributions. In CVPR, 2023

  3. [3]

    Lu, Chen Chen, Jiancang Ma, Pengyuan Zhou, Tae-Kyun Kim, Pan Hui, and Lin Wang

    Hao Ai, Zidong Cao, H. Lu, Chen Chen, Jiancang Ma, Pengyuan Zhou, Tae-Kyun Kim, Pan Hui, and Lin Wang. Dream360: Diverse and immersive outdoor virtual scene creation via transformer-based 360 image outpainting. TVCG, 2024

  4. [4]

    360-degree image completion by two- stage conditional gans

    Naofumi Akimoto, Seito Kasai, Masaki Hayashi, and Yoshimitsu Aoki. 360-degree image completion by two- stage conditional gans. InICIP, 2019

  5. [5]

    Diverse plausible 360-degree image outpainting for efficient 3dcg background creation.CVPR, 2022

    Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. Diverse plausible 360-degree image outpainting for efficient 3dcg background creation.CVPR, 2022

  6. [6]

    Gkitsas, Vladimiros Sterzentsenko, et al

    Georgios Albanis, Nikolaos Zioulis, Petros Drakoulis, V. Gkitsas, Vladimiros Sterzentsenko, et al. Pano3d: A holistic benchmark and a solid baseline for 360° depth estimation. CVPR Workshop, 2021

  7. [7]

    Cubes3D: Neural Network based Optical Flow in Omnidirectional Image Scenes

    André Apitzsch, Roman Seidel, and Gangolf Hirtz. Cubes3d: Neural network based optical flow in omnidi- rectional image scenes.arXiv preprint arXiv:1804.09004, 2018

  8. [8]

    Joint 2d-3d-semantic data for indoor scene understanding

    Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. ArXiv, 2017

Show all 299 references
  1. [9]

    Omniflownet: a perspective neural network adaptation for optical flow estimation in omnidirectional images

    Charles-Olivier Artizzu, Haozhou Zhang, Guillaume Allib- ert, and Cédric Demonceaux. Omniflownet: a perspective neural network adaptation for optical flow estimation in omnidirectional images. ICPR, 2021

  2. [10]

    Glpanodepth: Global-to-local panoramic depth estimation

    Jiayang Bai, Shuichang Lai, Haoyu Qin, Jie Guo, and Yanwen Guo. Glpanodepth: Global-to-local panoramic depth estimation. arXiv, 2022

  3. [11]

    Ma360: Multi-agent deep rein- forcement learning based live 360-degree video streaming on edge

    Yixuan Ban, Yuanxing Zhang, Haodan Zhang, Xinggong Zhang, and Zongming Guo. Ma360: Multi-agent deep rein- forcement learning based live 360-degree video streaming on edge. ICME, pages 1–6, 2020

  4. [12]

    Multidiffusion: Fusing diffusion paths for controlled image generation

    Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. Multidiffusion: Fusing diffusion paths for controlled image generation. In ICML, 2023

  5. [13]

    On the use of deep learning for computational imaging

    George Barbastathis, Aydogan Ozcan, and Guohai Situ. On the use of deep learning for computational imaging. Optica, 6(8):921–943, 2019

  6. [14]

    To- wards autonomous driving: a multi-modal 360◦ perception proposal

    Jorge Beltrán, Carlos Guindel, Irene Cortés, Alejandro Bar- rera, Armando Astudillo, Jesús Urdiales, Mario Álvarez, Farid Bekka, Vicente Milanés, and Fernando García. To- wards autonomous driving: a multi-modal 360◦ perception proposal. In ITSC, 2020

  7. [15]

    Scaled 360 layouts: Revisiting non-central panoramas

    Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, and Jose J Guerrero. Scaled 360 layouts: Revisiting non-central panoramas. In CVPR, 2021

  8. [16]

    Atlanta scaled layouts from non-central panoramas

    Bruno Berenguel-Baeta, Jesus Bermudez-Cameo, and Jose J Guerrero. Atlanta scaled layouts from non-central panoramas. PR, 2022

  9. [17]

    Learning omnidirectional flow in 360◦ video via siamese representation

    Keshav Bhandari, Bin Duan, Gaowen Liu, Hugo Latapie, Ziliang Zong, and Yan Yan. Learning omnidirectional flow in 360◦ video via siamese representation. InECCV, 2022

  10. [18]

    Revisiting optical flow estimation in 360 videos

    Keshav Bhandari, Ziliang Zong, and Yan Yan. Revisiting optical flow estimation in 360 videos. In2020 25th Inter- national Conference on Pattern Recognition (ICPR), pages 8196–8203. IEEE, 2021

  11. [19]

    nuscenes: A multi- modal dataset for autonomous driving

    Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multi- modal dataset for autonomous driving. InCVPR, 2020

  12. [20]

    Field- of-view iou for object detection in 360° images

    Miao Cao, Satoshi Ikehata, and Kiyoharu Aizawa. Field- of-view iou for object detection in 360° images. IEEE TIP, 2022

  13. [21]

    Ntire 2023 challenge on 360° omnidirectional image and video super-resolution: Datasets, methods and results.CVPR Workshop, 2023

    Ming Cao, Chong Mou, Fang Yu, et al. Ntire 2023 challenge on 360° omnidirectional image and video super-resolution: Datasets, methods and results.CVPR Workshop, 2023. A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision 29

  14. [22]

    Omnizoomer: Learning to move and zoom in on sphere at high-resolution

    Zidong Cao, Hao Ai, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Lin Wang. Omnizoomer: Learning to move and zoom in on sphere at high-resolution. InICCV, 2023

  15. [23]

    Engel, and Daniel Cremers

    David Caruso, Jakob J. Engel, and Daniel Cremers. Large- scale direct slam for omnidirectional cameras.IROS, 2015

  16. [24]

    Blind quality assessment of omnidirectional videos using spatio-temporal convolutional neural networks.Optik, 2021

    Xiongli Chai and Feng Shao. Blind quality assessment of omnidirectional videos using spatio-temporal convolutional neural networks.Optik, 2021

  17. [25]

    Matterport3d: Learning from rgb-d data in indoor environments.ArXiv, 2017

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments.ArXiv, 2017

  18. [26]

    Generating 360 outdoor panorama dataset with reliable sun position estimation.SIGGRAPH Asia 2018 Posters, 2018

    Shih-Hsiu Chang, Ching-Ya Chiu, Chia-Sheng Chang, et al. Generating 360 outdoor panorama dataset with reliable sun position estimation.SIGGRAPH Asia 2018 Posters, 2018

  19. [27]

    Depth estimation from indoor panoramas with neural scene repre- sentation

    Wenjie Chang, Yueyi Zhang, and Zhiwei Xiong. Depth estimation from indoor panoramas with neural scene repre- sentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 899–908, 2023

  20. [28]

    Borgwardt

    Dexiong Chen, Leslie O’Bray, and Karsten M. Borgwardt. Structure-aware transformer for graph representation learn- ing. In International Conference on Machine Learning, 2022

  21. [29]

    Intra- and inter- reasoning graph convolutional network for saliency predic- tion on 360° images

    Dongwen Chen, Chunmei Qing, Xuan Lin, Mengtao Ye, Xiangmin Xu, and Patrick Dickinson. Intra- and inter- reasoning graph convolutional network for saliency predic- tion on 360° images. IEEE TCVST, 2022

  22. [30]

    Salbinet360: Saliency prediction on 360◦ images with local-global bifurcated deep network.IEEE VR, 2020

    Dongwen Chen, Chunmei Qing, Xiangmin Xu, and Huan- sheng Zhu. Salbinet360: Saliency prediction on 360◦ images with local-global bifurcated deep network.IEEE VR, 2020

  23. [31]

    Multi-stage salient object detection in 360° omnidirectional image using complementary object-level semantic information

    Gang Chen, Feng Shao, Xiongli Chai, Qiuping Jiang, and Yo-Sung Ho. Multi-stage salient object detection in 360° omnidirectional image using complementary object-level semantic information. IEEE TETCI, 2024

  24. [32]

    360+ x: A panoptic multi- modal scene understanding dataset

    Hao Chen, Yuqi Hou, Chenyuan Qu, Irene Testini, Xiao- han Hong, and Jianbo Jiao. 360+ x: A panoptic multi- modal scene understanding dataset. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19373–19382, 2024

  25. [33]

    Dif- fusiondet: Diffusion model for object detection

    Shoufa Chen, Pei Sun, Yibing Song, and Ping Luo. Dif- fusiondet: Diffusion model for object detection. ICCV, 2022

  26. [34]

    Spherical structural similarity index for objective omnidirectional video quality assessment.ICME, 2018

    Sijia Chen, Yingxue Zhang, Yiming Li, Zhenzhong Chen, and Zhou Wang. Spherical structural similarity index for objective omnidirectional video quality assessment.ICME, 2018

  27. [35]

    Dynamic convolution: Attention over convolution kernels.CVPR, pages 11027– 11036, 2019

    Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic convolution: Attention over convolution kernels.CVPR, pages 11027– 11036, 2019

  28. [36]

    A unified and biologically-plausible re- lational graph representation of vision transformers.IEEE TNNLS, 2022

    Yuzhong Chen, Yu Du, Zhe Xiao, Lin Zhao, Lu Zhang, David Liu, Dajiang Zhu, Tuo Zhang, Xintao Hu, Tianming Liu, and Xi Jiang. A unified and biologically-plausible re- lational graph representation of vision transformers.IEEE TNNLS, 2022

  29. [37]

    Text2light: Zero-shot text-driven hdr panorama generation

    Zhaoxi Chen, Guangcong Wang, and Ziwei Liu. Text2light: Zero-shot text-driven hdr panorama generation. TOG, 2022

  30. [38]

    Recent advances in omnidirectional video coding for virtual reality: Projection and evaluation.Signal Process., 2018

    Zhenzhong Chen, Yiming Li, and Yingxue Zhang. Recent advances in omnidirectional video coding for virtual reality: Projection and evaluation.Signal Process., 2018

  31. [39]

    Unsupervised omnimvs: Efficient omnidirectional depth inference via establishing pseudo-stereo supervision

    Zisong Chen, Chunyu Lin, Nie Lang, Kang Liao, and Yao Zhao. Unsupervised omnimvs: Efficient omnidirectional depth inference via establishing pseudo-stereo supervision. IROS, 2023

  32. [40]

    Cube padding for weakly-supervised saliency prediction in 360 videos

    Hsien-Tzu Cheng, Chun-Hung Chao, Jin-Dong Dong, Hao- Kai Wen, Tyng-Luh Liu, and Min Sun. Cube padding for weakly-supervised saliency prediction in 360 videos. In CVPR, 2018

  33. [41]

    Sam- pling based spherical transformer for 360 degree image classification

    Sungmin Cho, Raehyuk Jung, and Junseok Kwon. Sam- pling based spherical transformer for 360 degree image classification. Expert Systems with Applications, 2024

  34. [42]

    360-indoor: towards learn- ing real-world objects in 360deg indoor equirectangular images

    Shih-Han Chou, Cheng Sun, Wen-Yen Chang, Wan-Ting Hsu, Min Sun, and Jianlong Fu. 360-indoor: towards learn- ing real-world objects in 360deg indoor equirectangular images. In WACV, 2020

  35. [43]

    Gauge equivariant convolutional networks and the icosahedral cnn

    Taco Cohen, Maurice Weiler, Berkay Kicanaoglu, and Max Welling. Gauge equivariant convolutional networks and the icosahedral cnn. InICML, 2019

  36. [44]

    Group equivariant convolu- tional networks

    Taco Cohen and Max Welling. Group equivariant convolu- tional networks. InInternational conference on machine learning. PMLR, 2016

  37. [45]

    Spherenet:Learningsphericalrepresentations for detection and classification in omnidirectional images

    Benjamin Coors, Alexandru Paul Condurache, and An- dreasGeiger. Spherenet:Learningsphericalrepresentations for detection and classification in omnidirectional images. In ECCV, 2018

  38. [46]

    Coughlan and Alan Loddon Yuille

    James M. Coughlan and Alan Loddon Yuille. The man- hattan world assumption: Regularities in scene statistics which enable bayesian inference. InNIPS, 2000

  39. [47]

    Zillow indoor dataset: Annotated floor plans with 360deg panoramas and 3d room layouts

    Steve Cruz, Will Hutchcroft, Yuguang Li, Naji Khosravan, Ivaylo Boyadzhiev, and Sing Bing Kang. Zillow indoor dataset: Annotated floor plans with 360deg panoramas and 3d room layouts. InCVPR, 2021

  40. [48]

    Thiago L. T. da Silveira, Paulo G. L. Pinto, Jeffri Murrugarra-Llerena, and Cl’audio Rosito Jung. 3d scene geometry estimation from 360◦ imagery: A survey.ACM CSUR, 2022

  41. [49]

    Dilated convolutional neural networks for panoramic image saliency prediction.ICASSP, 2020

    Feng Dai, Youqiang Zhang, Yike Ma, Hongliang Li, and Qiang Zhao. Dilated convolutional neural networks for panoramic image saliency prediction.ICASSP, 2020

  42. [50]

    Guided co-modulated gan for 360° field of view extrapolation.3DV, 2022

    MohammadRezaKarimiDastjerdi,YannickHold-Geoffroy, Jonathan Eisenmann, Siavash Khodadadeh, and Jean- François Lalonde. Guided co-modulated gan for 360° field of view extrapolation.3DV, 2022

  43. [51]

    Ev- erlight: Indoor-outdoor editable hdr lighting estimation

    MohammadRezaKarimiDastjerdi,YannickHold-Geoffroy, Jonathan Eisenmann, and Jean-François Lalonde. Ev- erlight: Indoor-outdoor editable hdr lighting estimation. ICCV, 2023

  44. [52]

    A viewport-driven multi-metric fusion approach for 360- degree video quality assessment.ICME, 2020

    Roberto Gerson de Albuquerque Azevedo, Neil Birkbeck, Ivan Janatra, Balu Adsumilli, and Pascal Frossard. A viewport-driven multi-metric fusion approach for 360- degree video quality assessment.ICME, 2020

  45. [53]

    Saliency prediction for omnidirectional images considering optimization on sphere domain.ICASSP, 2019

    Bhishma Dedhia, Jui-Chiu Chiang, and Yi-Fan Char. Saliency prediction for omnidirectional images considering optimization on sphere domain.ICASSP, 2019

  46. [54]

    Deepsphere: a graph-based spherical cnn

    Michaël Defferrard, Martino Milani, Frédérick Gusset, and Nathanaël Perraudin. Deepsphere: a graph-based spherical cnn. In ICLR, 2020

  47. [55]

    The mnist database of handwritten digit images for machine learning research

    Li Deng. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine, 29(6):141–142, 2012

  48. [56]

    Restricted deformable convolution- based road scene semantic segmentation using surround view cameras

    Liuyuan Deng, Ming Yang, Hao Li, Tianyi Li, Bing Hu, and Chunxiang Wang. Restricted deformable convolution- based road scene semantic segmentation using surround view cameras. IEEE TITS, 2020

  49. [57]

    Cnn based semantic segmentation for urban traffic scenes using fisheye camera.IV, 2017

    Liuyuan Deng, Ming Yang, Yeqiang Qian, Chunxiang Wang, and Bing Wang. Cnn based semantic segmentation for urban traffic scenes using fisheye camera.IV, 2017

  50. [58]

    Lau-net: Latitude adaptive upscaling network for omnidirectional imag super-resolution.CVPR, 2021

    Xin Deng, Hao Wang, Mai Xu, Yichen Guo, Yuhang Song, and Li Yang. Lau-net: Latitude adaptive upscaling network for omnidirectional imag super-resolution.CVPR, 2021. 30 Hao Ai 1 et al

  51. [59]

    Extending 2d saliency models for head movement prediction in 360-degree images using cnn- based fusion

    Ibrahim Djemai, Sid Ahmed Fezza, Wassim Hamidouche, and Olivier Déforges. Extending 2d saliency models for head movement prediction in 360-degree images using cnn- based fusion. ISCAS, 2020

  52. [60]

    Panocontext-former: Panoramic total scene understanding with a transformer.ArXiv, 2023

    Yuan Dong, Chuangjie Fang, Zilong Dong, Liefeng Bo, and Ping Tan. Panocontext-former: Panoramic total scene understanding with a transformer.ArXiv, 2023

  53. [61]

    An image is worth 16x16 words: Transformers for image recognition at scale.ICLR, 2020

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.ICLR, 2020

  54. [62]

    Perceptual quality assessment of omnidirectional images.ISCAS, 2018

    Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Yucheng Zhu, Yi Fang, and Xiaokang Yang. Perceptual quality assessment of omnidirectional images.ISCAS, 2018

  55. [63]

    Pano popups: Indoor 3d reconstruction with a plane-aware network

    Marc Eder, Pierre Moulon, and Li Guan. Pano popups: Indoor 3d reconstruction with a plane-aware network. In 3DV, 2019

  56. [64]

    Tangent images for mitigating spherical distortion

    Marc Eder, Mykhailo Shvets, John Lim, and Jan-Michael Frahm. Tangent images for mitigating spherical distortion. In CVPR, 2020

  57. [65]

    Engel, Thomas Schöps, and Daniel Cremers

    Jakob J. Engel, Thomas Schöps, and Daniel Cremers. Lsd- slam: Large-scale direct monocular slam. In European Conference on Computer Vision, 2014

  58. [66]

    Taming transformers for high-resolution image synthesis.CVPR, 2020

    Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis.CVPR, 2020

  59. [67]

    Deep depth estimation on 360 images with a double quaternion loss

    Brandon Yushan Feng, Wangjue Yao, Zheyuan Liu, and Amitabh Varshney. Deep depth estimation on 360 images with a double quaternion loss. In3DV, 2020

  60. [68]

    360 depth estimation in the wild-the depth360 dataset and the segfuse network

    Qi Feng, Hubert PH Shum, and Shigeo Morishima. 360 depth estimation in the wild-the depth360 dataset and the segfuse network. InVR, 2022

  61. [69]

    Fácil, Alejandro Pérez- Yus, Cédric Demonceaux, Javier Civera, and Josechu J

    Clara Fernandez-Labrador, José M. Fácil, Alejandro Pérez- Yus, Cédric Demonceaux, Javier Civera, and Josechu J. Guerrero. Corners for layout: End-to-end layout recovery from 360 images.RAL, 2020

  62. [70]

    Layouts from panoramic images with geometry and deep learning.RAL, 2018

    Clara Fernandez-Labrador, Alejandro Perez-Yus, Gon- zalo Lopez-Nicolas, and Jose J Guerrero. Layouts from panoramic images with geometry and deep learning.RAL, 2018

  63. [71]

    Generating a full spherical view by modeling the relation between two fisheye images.The Visual Computer, pages 1–26, 2024

    María Flores, David Valiente, Adrián Peidró, Oscar Reinoso, and Luis Payá. Generating a full spherical view by modeling the relation between two fisheye images.The Visual Computer, pages 1–26, 2024

  64. [72]

    Adaptive hypergraph convolutional network for no-reference 360-degree image quality assessment.ACM MM, 2021

    Jun Fu, Chengbin Hou, Wei Zhou, Jiahua Xu, and Zhibo Chen. Adaptive hypergraph convolutional network for no-reference 360-degree image quality assessment.ACM MM, 2021

  65. [73]

    Quality assessment for omnidirectional video: A spatio-temporal distortion modeling approach.TMM, 2022

    Pan Gao, Pengwei Zhang, and Aljosa Smolic. Quality assessment for omnidirectional video: A spatio-temporal distortion modeling approach.TMM, 2022

  66. [74]

    Review on panoramic imaging and its appli- cations in scene understanding

    Shaohua Gao, Kailun Yang, Hao Shi, Kaiwei Wang, and Jian Bai. Review on panoramic imaging and its appli- cations in scene understanding. IEEE Transactions on Instrumentation and Measurement, 71:1–34, 2022

  67. [75]

    Deep parametric indoor lighting estimation.ICCV, 2019

    Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné, and Jean-François Lalonde. Deep parametric indoor lighting estimation.ICCV, 2019

  68. [76]

    Learning to predict indoor illumination from a single image.TOG, 2017

    Marc-André Gardner, Kalyan Sunkavalli, Ersin Yumer, Xiaohui Shen, Emiliano Gambaretto, Christian Gagné, and Jean-François Lalonde. Learning to predict indoor illumination from a single image.TOG, 2017

  69. [77]

    Carr, and Jean-François Lalonde

    Mathieu Garon, Kalyan Sunkavalli, Sunil Hadap, Nathan A. Carr, and Jean-François Lalonde. Fast spatially- varying indoor lighting estimation.CVPR, 2019

  70. [78]

    Generative adversarial nets.NIPS, 2014

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets.NIPS, 2014

  71. [79]

    Gorski, Eric Hivon, A

    Krzysztof M. Gorski, Eric Hivon, A. J. Banday, Ben- jamin Dan Wandelt, Frode K. Hansen, Martin Reinecke, and M. Bartelman. Healpix: A framework for high- resolution discretization and fast analysis of data dis- tributed on the sphere.The Astrophysical Journal, 2005

  72. [80]

    Zero-shot learning for reflection removal of single 360-degree image

    Byeong-Ju Han and Jae-Young Sim. Zero-shot learning for reflection removal of single 360-degree image. InEuropean Conference on Computer Vision, pages 533–548. Springer, 2022

  73. [81]

    Piinet: A 360-degree panoramic image inpainting network using a cube map

    Seo Woo Han and Doug Young Suh. Piinet: A 360-degree panoramic image inpainting network using a cube map. arXiv, 2020

  74. [82]

    Enhancement of novel view synthesis using omnidirectional image completion

    Takayuki Hara and Tatsuya Harada. Enhancement of novel view synthesis using omnidirectional image completion. ArXiv, 2022

  75. [83]

    Spherical image generation from a single image by consid- ering scene symmetry

    Takayuki Hara, Yusuke Mukuta, and Tatsuya Harada. Spherical image generation from a single image by consid- ering scene symmetry. InAAAI, 2021

  76. [84]

    Zhang, Shaoqing Ren, and Jian Sun

    Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition.CVPR, 2016

  77. [85]

    Riecke, and Lillian Yang

    Yasamin Heshmat, Brennan Jones, Xiaoxuan Xiong, Car- man Neustaedter, Anthony Tang, Bernhard E. Riecke, and Lillian Yang. Geocaching with a beam: Shared outdoor activities through a telepresence robot with 360 degree viewing. CHI, 2018

  78. [86]

    Generative adversarial imitation learning

    Jonathan Ho and Stefano Ermon. Generative adversarial imitation learning. InNIPS, 2016

  79. [87]

    Deep sky modeling for single image outdoor lighting estimation

    Yannick Hold-Geoffroy, Akshaya Athawale, and Jean- François Lalonde. Deep sky modeling for single image outdoor lighting estimation. InCVPR, 2019

  80. [88]

    Deep outdoor illumination estimation.CVPR, 2017

    Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap, Emiliano Gambaretto, and Jean-François Lalonde. Deep outdoor illumination estimation.CVPR, 2017

  81. [89]

    Panoramic image reflection removal

    Yuchen Hong, Qian Zheng, Lingran Zhao, Xudong Jiang, Alex C Kot, and Boxin Shi. Panoramic image reflection removal. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7762– 7771, 2021

  82. [90]

    Par 2 net: End-to-end panoramic image reflection removal.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12192– 12205, 2023

    Yuchen Hong, Qian Zheng, Lingran Zhao, Xudong Jiang, Alex C Kot, and Boxin Shi. Par 2 net: End-to-end panoramic image reflection removal.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12192– 12205, 2023

  83. [91]

    Image quality metrics: Psnr vs

    Alain Horé and Djemel Ziou. Image quality metrics: Psnr vs. ssim. ICPR, 2010

  84. [92]

    C. Y. Hsu, Cheng Sun, and Hwann-Tzong Chen. Moving in a 360 world: Synthesizing panoramic parallaxes from a single panorama. ArXiv, 2021

  85. [93]

    LoRA: Low-rank adaptation of large language models

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022

  86. [94]

    Omnidirec- tional video quality assessment with causal intervention

    Zongyao Hu, Lixiong Liu, and Qingbing Sang. Omnidirec- tional video quality assessment with causal intervention. IEEE T-OB, 2024

  87. [95]

    360vo: Visual odometry using a single 360 camera.2022 International Conference on Robotics and Automation (ICRA), pages 5594–5600, 2022

    Hu Huang and Sai-Kit Yeung. 360vo: Visual odometry using a single 360 camera.2022 International Conference on Robotics and Automation (ICRA), pages 5594–5600, 2022

  88. [96]

    360loc: A dataset and benchmark for omnidirectional visual localization with cross-device queries

    Huajian Huang, Changkun Liu, Yipeng Zhu, Hui Cheng, Tristan Braud, and Sai-Kit Yeung. 360loc: A dataset and benchmark for omnidirectional visual localization with cross-device queries. In Proceedings of the IEEE/CVF A Survey of Representation Learning, Optimization Strategies,...

  89. [97]

    Surround-view fisheye optics in computer vision and simulation: Survey and challenges.IEEE Transactions on Intelligent Transportation Systems, 2024

    Daniel Jakab, Brian Michael Deegan, Sushil Sharma, Eoin Martino Grua, Jonathan Horgan, Enda Ward, Pepijn Van De Ven, Anthony Scanlan, and Ciarán Eis- ing. Surround-view fisheye optics in computer vision and simulation: Survey and challenges.IEEE Transactions on Intelligent Tra...

  90. [98]

    Panoramic panoptic segmentation: Towards complete sur- rounding understanding via unsupervised contrastive learn- ing

    Alexander Jaus, Kailun Yang, and Rainer Stiefelhagen. Panoramic panoptic segmentation: Towards complete sur- rounding understanding via unsupervised contrastive learn- ing. In 2021 IEEE Intelligent Vehicles Symposium (IV), pages 1421–1427. IEEE, 2021

  91. [99]

    Panoramic panoptic segmentation: Insights into surround- ing parsing for mobile agents via unsupervised contrastive learning

    Alexander Jaus, Kailun Yang, and Rainer Stiefelhagen. Panoramic panoptic segmentation: Insights into surround- ing parsing for mobile agents via unsupervised contrastive learning. IEEE Transactions on Intelligent Transportation Systems, 24(4):4438–4453, 2023

  92. [100]

    Active perception for outdoor localisation with an omnidirectional camera.2020 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 4567–4574, 2020

    Maleen Jayasuriya, Ravindra Ranasinghe, and Gamini Dissanayake. Active perception for outdoor localisation with an omnidirectional camera.2020 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 4567–4574, 2020

  93. [101]

    Panoramic slam from a multiple fisheye camera rig.ISPRS Journal of Photogrammetry and Remote Sensing, 159:169–183, 2020

    Shunping Ji, Zijie Qin, Jie Shan, and Meng Lu. Panoramic slam from a multiple fisheye camera rig.ISPRS Journal of Photogrammetry and Remote Sensing, 159:169–183, 2020

  94. [102]

    3d room layout recovery generalizing across manhattan and non-manhattan worlds

    Haijing Jia, Hong Yi, Hirochika Fujiki, Hengzhi Zhang, Wei Wang, and Makoto Odamaki. 3d room layout recovery generalizing across manhattan and non-manhattan worlds. CVPR Workshop, 2022

  95. [103]

    Spherical cnns on unstructured grids

    Chiyu Max Jiang, Jingwei Huang, Karthik Kashinath, Prabhat, Philip Marcus, and Matthias Niessner. Spherical cnns on unstructured grids. InICLR, 2019

  96. [104]

    Cubemap- based perception-driven blind quality assessment for 360- degree images

    Hao Jiang, Gang yi Jiang, Mei Yu, Yun Zhang, You Yang, Zongju Peng, Fen Chen, and Qingbo Zhang. Cubemap- based perception-driven blind quality assessment for 360- degree images. TIP, 2021

  97. [105]

    Unifuse: Unidirectional fusion for 360 panorama depth estimation

    Hualie Jiang, Zhe Sheng, Siyu Zhu, Zilong Dong, and Rui Huang. Unifuse: Unidirectional fusion for 360 panorama depth estimation. RAL, 2021

  98. [106]

    Minimalist and high-quality panoramic imaging with psf-aware transform- ers

    Qi Jiang, Shaohua Gao, Yao Gao, Kailun Yang, Zhonghua Yi, Hao Shi, Lei Sun, and Kaiwei Wang. Minimalist and high-quality panoramic imaging with psf-aware transform- ers. IEEE Transactions on Image Processing, 2024

  99. [107]

    Annular computational imaging: Capture clear panoramic images through simple lens.IEEE Trans- actions on Computational Imaging, 8:1250–1264, 2022

    QiJiang,HaoShi,LeiSun,ShaohuaGao,KailunYang,and Kaiwei Wang. Annular computational imaging: Capture clear panoramic images through simple lens.IEEE Trans- actions on Computational Imaging, 8:1250–1264, 2022

  100. [108]

    3d reconstruction of spherical images: A review of tech- niques, applications, and prospects.ArXiv, 2023

    San Jiang, Yaxin Li, Duojie Weng, Kan You, and Wu Chen. 3d reconstruction of spherical images: A review of tech- niques, applications, and prospects.ArXiv, 2023

  101. [109]

    Lgt-net: Indoor panoramic room layout estimation with geometry-aware transformer network

    Zhigang Jiang, Zhongzheng Xiang, Jinhua Xu, and Ming Zhao. Lgt-net: Indoor panoramic room layout estimation with geometry-aware transformer network. InCVPR, 2022

  102. [110]

    Reinforcement learning based rate adaptation for 360-degree video streaming.IEEE T-OB, 2020

    Zhiqian Jiang, Xu Zhang, Yiling Xu, Zhan Ma, Jun Sun, and Yunfei Zhang. Reinforcement learning based rate adaptation for 360-degree video streaming.IEEE T-OB, 2020

  103. [111]

    Geometric structure based and regularized depth estimation from 360 indoor imagery

    Lei Jin, Yanyu Xu, Jia Zheng, Junfei Zhang, Rui Tang, Shugong Xu, Jingyi Yu, and Shenghua Gao. Geometric structure based and regularized depth estimation from 360 indoor imagery. InCVPR, 2020

  104. [112]

    Rapt360: Reinforcement learning-based rate adaptation for 360-degree video streaming with adap- tive prediction and tiling.IEEE TCSVT, 2021

    Nuowen Kan, Junni Zou, Chenglin Li, Wenrui Dai, and Hongkai Xiong. Rapt360: Reinforcement learning-based rate adaptation for 360-degree video streaming with adap- tive prediction and tiling.IEEE TCSVT, 2021

  105. [113]

    Spherical fibonacci mapping

    Benjamin Keinert, Matthias Innmann, Michael Sänger, and Marc Stamminger. Spherical fibonacci mapping. TOG, 2015

  106. [114]

    3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (TOG), 42:1 – 14, 2023

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (TOG), 42:1 – 14, 2023

  107. [115]

    Geometry aware convolutional filters for omnidirectional images representa- tion

    Renata Khasanova and Pascal Frossard. Geometry aware convolutional filters for omnidirectional images representa- tion. In ICML, 2019

  108. [116]

    Deep virtual reality image quality assessment with human per- ception guider for omnidirectional image.IEEE TCSVT, 2020

    Hak Gu Kim, Heoun taek Lim, and Yong Man Ro. Deep virtual reality image quality assessment with human per- ception guider for omnidirectional image.IEEE TCSVT, 2020

  109. [117]

    Hansung Kim, Luca Hernaggi, Philip J. B. Jackson, and Adrian Hilton. Immersive spatial audio reproduction for VR/AR using room acoustic modelling from 360◦ images. In VR, 2019

  110. [118]

    3d scene reconstruction from multiple spherical stereo pairs.IJCV, 2013

    Hansung Kim and Adrian Hilton. 3d scene reconstruction from multiple spherical stereo pairs.IJCV, 2013

  111. [119]

    Junho Kim, Eungbean Lee, and Y. Kim. Calibrating panoramic depth estimation for practical localization and mapping. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 8796–8806, 2023

  112. [120]

    Panoptic segmentation from stitched panoramic view for automated driving

    Christian Kinzig, Henning Miller, Martin Lauer, and Christoph Stiller. Panoptic segmentation from stitched panoramic view for automated driving. In2024 IEEE In- telligent Vehicles Symposium (IV), pages 3342–3347. IEEE, 2024

  113. [121]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023

  114. [122]

    360 virtual reality: A swot analysis in comparison to virtual reality.Frontiers in Psychology, 2020

    Aden Kittel, Paul Larkin, Ian Cunningham, and Michael Spittle. 360 virtual reality: A swot analysis in comparison to virtual reality.Frontiers in Psychology, 2020

  115. [123]

    Krizhevsky and G

    A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images.Master’s thesis, Department of Computer Science, University of Toronto, 2009

  116. [124]

    Spherical light fields

    Bernd Krolla, Maximilian Diebold, Bastian Goldlücke, and Didier Stricker. Spherical light fields. InBMVC, 2014

  117. [125]

    Shreyas Kulkarni, Peng Yin, and Sebastian A. Scherer. 360fusionnerf: Panoramic neural radiance fields with joint guidance. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7202–7209, 2022

  118. [126]

    Real-time panoramic depth maps from omni- directional stereo images for 6 dof videos in virtual reality

    Po Kong Lai, Shuang Xie, Jochen Lang, and Robert La- ganière. Real-time panoramic depth maps from omni- directional stereo images for 6 dof videos in virtual reality. In 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pages 405–412. IEEE, 2019

  119. [127]

    Temporal ensembling for semi-supervised learning

    Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. InICLR, 2017

  120. [128]

    Semi-supervised 360° depth estimation from multiple fisheye cameras with pixel-level selective loss.ICASSP, 2022

    JaewooLee,DaeHyuckPark,DongwookLee,andDaehyun Ji. Semi-supervised 360° depth estimation from multiple fisheye cameras with pixel-level selective loss.ICASSP, 2022

  121. [129]

    Spherephd: Applying cnns on a spherical polyhedron representation of 360◦ images

    Yeonkun Lee, Jaeseok Jeong, Jong Seob Yun, Wonjune Cho, and Kuk jin Yoon. Spherephd: Applying cnns on a spherical polyhedron representation of 360◦ images. CVPR, 2019

  122. [130]

    Viewport proposal cnn for 360◦ video quality assessment

    Chen Li, Mai Xu, Lai Jiang, Shanyi Zhang, and Xiaom- ing Tao. Viewport proposal cnn for 360◦ video quality assessment. CVPR, 2019

  123. [131]

    Attentive deep stitching and quality assessment for 32 Hao Ai 1 et al

    Jia Li, Yifan Zhao, Weihua Ye, Kaiwen Yu, and Shiming Ge. Attentive deep stitching and quality assessment for 32 Hao Ai 1 et al. 360◦ omnidirectional images. IEEE Journal of Selected Topics in Signal Processing, 14(1):209–221, 2019

  124. [132]

    Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation

    Jialu Li and Mohit Bansal. Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation. Advances in Neural Information Processing Systems, 36:21878–21894, 2023

  125. [133]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730– 19742. PMLR, 2023

  126. [134]

    S2net: Accurate panorama depth estimation on spherical surface.IEEE RAL, 2023

    Meng Li, Senbo Wang, Weihao Yuan, Weichao Shen, Zhe Sheng, and Zilong Dong. S2net: Accurate panorama depth estimation on spherical surface.IEEE RAL, 2023

  127. [135]

    Mode: Multi-view omnidirectional depth estimation with 360◦ cameras

    Ming Li, Xueqian Jin, Xuejiao Hu, Jingzhao Dai, Sidan Du, and Yang Li. Mode: Multi-view omnidirectional depth estimation with 360◦ cameras. In ECCV, 2022

  128. [136]

    Spherical stereo for the construction of immersive vr environment

    Shigang Li and Kiyotaka Fukumori. Spherical stereo for the construction of immersive vr environment. InIEEE VR, 2005

  129. [137]

    Spherical stereo for the construction of immersive vr environment

    Shigang Li and Kiyotaka Fukumori. Spherical stereo for the construction of immersive vr environment. InIEEE Proceedings. VR 2005. Virtual Reality, 2005., pages 217–

  130. [138]

    Sgat4pass: Spherical geometry-aware transformer for panoramic semantic segmentation

    Xuewei Li, Tao Wu, Zhongang Qi, Gaoang Wang, Ying Shan, and Xi Li. Sgat4pass: Spherical geometry-aware transformer for panoramic semantic segmentation. InIJ- CAI, 2023

  131. [139]

    Deep 360◦ optical flow estimation based on multi- projection fusion

    Yiheng Li, Connelly Barnes, Kun Huang, and Fang-Lue Zhang. Deep 360◦ optical flow estimation based on multi- projection fusion. InEuropean Conference on Computer Vision, pages 336–352. Springer, 2022

  132. [140]

    Omnifusion: 360 monocular depth estimation via geometry-aware fusion.CVPR, 2022

    Yuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang, Ye Duan, and Liu Ren. Omnifusion: 360 monocular depth estimation via geometry-aware fusion.CVPR, 2022

  133. [141]

    Inverse rendering for complex indoor scenes: Shape, spatially- varying lighting and svbrdf from a single image.CVPR, 2020

    Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Inverse rendering for complex indoor scenes: Shape, spatially- varying lighting and svbrdf from a single image.CVPR, 2020

  134. [142]

    Sad360: Spheri- cal viewport-aware dynamic tiling for 360-degree video streaming

    Zhijun Li, Yumei Wang, and Yu Liu. Sad360: Spheri- cal viewport-aware dynamic tiling for 360-degree video streaming. VCIP, 2022

  135. [143]

    Cylin-painting: Seamless 360° panoramic image outpainting and beyond

    Kang Liao, Xiangyu Xu, Chunyu Lin, Wenqi Ren, Yun- chao Wei, and Yao Zhao. Cylin-painting: Seamless 360° panoramic image outpainting and beyond. IEEE TIP, 2022

  136. [144]

    Panoswin: A pano-style swin trans- former for panorama understanding

    Zhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao, and Guichun Zhou. Panoswin: A pano-style swin trans- former for panorama understanding. InCVPR, 2023

  137. [145]

    Pano-sfmlearner: Self-supervised multi-task learning of depth and semantics in panoramic videos.IEEE SPL, 2021

    Mengyi Liu, Shuhui Wang, Yulan Guo, Yuan He, and Hui Xue. Pano-sfmlearner: Self-supervised multi-task learning of depth and semantics in panoramic videos.IEEE SPL, 2021

  138. [146]

    Deep learning 3d shapes using alt-az anisotropic 2-sphere convolution

    Min Liu, Fupin Yao, Chiho Choi, Ayan Sinha, and Karthik Ramani. Deep learning 3d shapes using alt-az anisotropic 2-sphere convolution. InICLR, 2018

  139. [147]

    Cross-modal 360◦ depth completion and reconstruc- tion for large-scale indoor environment.IEEE TITS, 2022

    Ruyu Liu, Guodao Zhang, Jiangming Wang, and Shuwen Zhao. Cross-modal 360◦ depth completion and reconstruc- tion for large-scale indoor environment.IEEE TITS, 2022

  140. [148]

    Sph2pob: Boost- ing object detection on spherical images with planar ori- ented boxes methods

    Xinyuan Liu, Hang Xu, Bin Chen, Qiang Zhao, Yike Ma, Chenggang Clarence Yan, and Feng Dai. Sph2pob: Boost- ing object detection on spherical images with planar ori- ented boxes methods. InIJCAI, 2023

  141. [149]

    Perceptual quality assessment of omnidirectional images: A benchmark and computational model

    Xuelin Liu, Jiebin Yan, Liping Huang, Yuming Fang, Zheng Wan, and Yang Liu. Perceptual quality assessment of omnidirectional images: A benchmark and computational model. ACM TOMM, 2024

  142. [150]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021

  143. [151]

    Image stitching for dual fisheye cameras

    I-chan Lo, Kuang-tsu Shih, and Homer H Chen. Image stitching for dual fisheye cameras. In2018 25th IEEE In- ternational Conference on Image Processing (ICIP), pages 3164–3168. IEEE, 2018

  144. [152]

    Efficient and accurate stitching for 360° dual-fisheye images and videos

    I-Chan Lo, Kuang-Tsu Shih, and Homer H Chen. Efficient and accurate stitching for 360° dual-fisheye images and videos. IEEE Transactions on Image Processing, 31:251– 262, 2021

  145. [153]

    Photometric consistency for dual fisheye cameras

    I-Chan Lo, Kuang-Tsu Shih, Gwo-Hwa Ju, and Homer H Chen. Photometric consistency for dual fisheye cameras. In 2020 IEEE International Conference on Image Processing (ICIP), pages 261–265. IEEE, 2020

  146. [154]

    Autoregressive omni-aware outpainting for open-vocabulary 360-degree image generation

    Zhuqiang Lu, Kun Hu, Chaoyue Wang, Lei Bai, and Zhiy- ong Wang. Autoregressive omni-aware outpainting for open-vocabulary 360-degree image generation. AAAI, 2024

  147. [155]

    An iterative image reg- istration technique with an application to stereo vision

    Bruce D Lucas and Takeo Kanade. An iterative image reg- istration technique with an application to stereo vision. In IJCAI’81: 7th international joint conference on Artificial intelligence, volume 2, pages 674–679, 1981

  148. [156]

    Salgcn: Saliency prediction for 360-degree images based on spherical graph convolutional networks

    Haoran Lv, Qin Yang, Chenglin Li, Wenrui Dai, Junni Zou, and Hongkai Xiong. Salgcn: Saliency prediction for 360-degree images based on spherical graph convolutional networks. ACM MM, 2020

  149. [157]

    Densepass: Dense panoramicsemanticsegmentationviaunsuperviseddomain adaptation with attention-augmented context exchange

    Chaoxiang Ma, Jiaming Zhang, Kailun Yang, Alina Roitberg, and Rainer Stiefelhagen. Densepass: Dense panoramicsemanticsegmentationviaunsuperviseddomain adaptation with attention-augmented context exchange. ITSC, 2021

  150. [158]

    Viewport-aware deep reinforcement learning approach for 360 ◦ video caching

    Pantelis Maniotis and Nikolaos Thomos. Viewport-aware deep reinforcement learning approach for 360 ◦ video caching. TMM, 2020

  151. [159]

    Karvelis, Christoforos Kanellakis, Dariusz Kominiak, and George Nikolakopoulos

    Sina Sharif Mansouri, Petros S. Karvelis, Christoforos Kanellakis, Dariusz Kominiak, and George Nikolakopoulos. Vision-based mav navigation in underground mine using convolutional neural network.IECON, 2019

  152. [160]

    Usenko, J

    Hidenobu Matsuki, Lukas von Stumberg, Vladyslav C. Usenko, J. Stückler, and Daniel Cremers. Omnidirectional dso: Direct sparse odometry with fisheye cameras.IEEE Robotics and Automation Letters, 3:3693–3700, 2018

  153. [161]

    Recurrent neural networks

    Larry R Medsker and LC Jain. Recurrent neural networks. Design and Applications, 2001

  154. [162]

    Waymo open dataset: Panoramic video panoptic segmentation

    Jieru Mei, Alex Zihao Zhu, Xinchen Yan, Hang Yan, Siyuan Qiao, Liang-Chieh Chen, and Henrik Kretzschmar. Waymo open dataset: Panoramic video panoptic segmentation. In European Conference on Computer Vision, pages 53–72. Springer, 2022

  155. [163]

    Jeon, and Min H

    Andreas Meuleman, Hyeonjoong Jang, Daniel S. Jeon, and Min H. Kim. Real-time sphere sweeping stereo from multiview fisheye images.CVPR, 2021

  156. [164]

    Nerf: Representing scenes as neural radiance fields for view synthesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021

  157. [165]

    Conditional generative adversarial nets

    Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. ArXiv, abs/1411.1784, 2014

  158. [166]

    Salnet360: Saliency maps for omni- directional images with cnn.Signal Process

    Rafael Monroy, Sebastian Lutz, Tejo Chalasani, and Aljoscha Smolic. Salnet360: Saliency maps for omni- directional images with cnn.Signal Process. Image Com- mun., 2017

  159. [167]

    Anh Nguyen, Zhisheng Yan, and Klara Nahrstedt. Your attention is unique: Detecting 360-degree video saliency A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision 33 in head-mounted display for head movement prediction. ACM MM, 2018

  160. [168]

    Introduction to geometry.Physics Bulletin, 1962

    TH O’Beirne. Introduction to geometry.Physics Bulletin, 1962

  161. [169]

    Price, and Jason D

    Jeremy Ocampo, Matthew A. Price, and Jason D. McEwen. Scalable and equivariant spherical cnns by discrete- continuous convolutions. InICLR, 2023

  162. [170]

    Bips: Bi-modal indoor panorama synthesis via residual depth-aided adversarial learning

    Changgyoon Oh, Wonjune Cho, Daehee Park, Yujeong Chae, Lin Wang, and Kuk-Jin Yoon. Bips: Bi-modal indoor panorama synthesis via residual depth-aided adversarial learning. ArXiv, 2021

  163. [171]

    Visual attention-aware omnidirectional video streaming using op- timal tiles for virtual reality.IEEE JETCAS, 2019

    Cagri Ozcinar, Julián Cabrera, and Aljosa Smolic. Visual attention-aware omnidirectional video streaming using op- timal tiles for virtual reality.IEEE JETCAS, 2019

  164. [172]

    Super- resolution of omnidirectional images using adversarial learning

    Cagri Ozcinar, Aakanksha Rana, and Aljosa Smolic. Super- resolution of omnidirectional images using adversarial learning. In MMSP, 2019

  165. [173]

    Fully-automatic reflection removal for 360-degree images

    Jonghyuk Park, Hyeona Kim, Eunpil Park, and Jae-Young Sim. Fully-automatic reflection removal for 360-degree images. In Proceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision, pages 1609–1617, 2024

  166. [174]

    Adaptive streaming of 360-degree videos with rein- forcement learning

    Sohee Park, Minh Hoai, Arani Bhaacharya, and Samir R Das. Adaptive streaming of 360-degree videos with rein- forcement learning. WACV, 2021

  167. [175]

    Stereo panorama with a single camera.CVPR, 1999

    Shmuel Peleg and Moshe Ben-Ezra. Stereo panorama with a single camera.CVPR, 1999

  168. [176]

    High-resolution depth estimation for 360deg panoramas through perspective and panoramic depth images registration

    Chi-Han Peng and Jiayao Zhang. High-resolution depth estimation for 360deg panoramas through perspective and panoramic depth images registration. InProceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, pages 3116–3125, 2023

  169. [177]

    Gobbetti

    Giovanni Pintore, Marco Agus, and E. Gobbetti. At- lantanet: Inferring the 3d indoor layout from a single360◦ image beyond the manhattan world assumption. InECCV, 2020

  170. [178]

    Deep3dlayout: 3d reconstruction of an indoor layout from a spherical panoramic image.TOG, 2021

    Giovanni Pintore, Eva Almansa, Marco Agus, and Enrico Gobbetti. Deep3dlayout: 3d reconstruction of an indoor layout from a spherical panoramic image.TOG, 2021

  171. [179]

    Slicenet: deep dense depth estimation from a single in- door panorama using a slice-based representation.CVPR, 2021

    Giovanni Pintore, Eva Almansa, and Jens Schneider. Slicenet: deep dense depth estimation from a single in- door panorama using a slice-based representation.CVPR, 2021

  172. [180]

    Panoramic lens.Applied Optics, 33(31):7356– 7361, 1994

    Ian Powell. Panoramic lens.Applied Optics, 33(31):7356– 7361, 1994

  173. [181]

    Safe robot navigation using an omnidirectional camera

    Gyula Pudics, Miklos Zsolt Szabo-Resch, and Zoltán Vá- mossy. Safe robot navigation using an omnidirectional camera. CINTI, 2015

  174. [182]

    Survey on fish-eye cameras and their applications in intelligent vehicles

    Yeqiang Qian, Ming Yang, and John M Dolan. Survey on fish-eye cameras and their applications in intelligent vehicles. IEEE Transactions on Intelligent Transportation Systems, 23(12):22755–22771, 2022

  175. [183]

    Viewport-dependent saliency prediction in 360 ◦ video

    Minglang Qiao, Mai Xu, Zulin Wang, and Ali Borji. Viewport-dependent saliency prediction in 360 ◦ video. TMM, 2021

  176. [184]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. InICML, 2021

  177. [185]

    A dataset of head and eye movements for 360 degree images

    Yashas Rai, Jesús Gutiérrez, and Patrick Le Callet. A dataset of head and eye movements for 360 degree images. In ACM MMSys, 2017

  178. [186]

    Convolutional neural network-based robot navigation using uncalibrated spherical images†

    Lingyan Ran, Yanning Zhang, Qilin Zhang, and Tao Yang. Convolutional neural network-based robot navigation using uncalibrated spherical images†. Sensors, 2017

  179. [187]

    Vision transformers for dense prediction

    René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. InICCV, 2021

  180. [188]

    Girshick, and Jian Sun

    Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks.IEEE T-PAMI, 2015

  181. [189]

    360monodepth: High-resolution 360° monocular depth es- timation

    Manuel Rey-Area, Mingze Yuan, and Christian Richardt. 360monodepth: High-resolution 360° monocular depth es- timation. CVPR, 2022

  182. [190]

    Using local refine- ments on 360 stitching from dual-fisheye cameras

    Rafael Roberto, Daniel Perazzo, João Paulo Lima, Veronica Teichrieb, Jonysberg Peixoto Quintino, Fabio QB da Silva, Andre LM Santos, and Helder Pinho. Using local refine- ments on 360 stitching from dual-fisheye cameras. In VISIGRAPP (5: VISAPP), pages 17–26, 2020

  183. [191]

    Head-mounted display systems

    Jannick P Rolland and Hong Hua. Head-mounted display systems. Encyclopedia of optical engineering, 2005

  184. [192]

    Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer

    Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models.CVPR, 2021

  185. [193]

    Atlanta world: an expectation maximization framework for simultaneous low- level edge grouping and camera calibration in complex man-made environments

    Grant Schindler and Frank Dellaert. Atlanta world: an expectation maximization framework for simultaneous low- level edge grouping and camera calibration in complex man-made environments. InCVPR, 2004

  186. [194]

    Omni- flow: Human omnidirectional optical flow.CVPR Work- shops, 2021

    Roman Seidel, André Apitzsch, and Gangolf Hirtz. Omni- flow: Human omnidirectional optical flow.CVPR Work- shops, 2021

  187. [195]

    Rovo: Robust omnidirec- tional visual odometry for wide-baseline wide-fov camera systems

    Hochang Seok and Jongwoo Lim. Rovo: Robust omnidirec- tional visual odometry for wide-baseline wide-fov camera systems. 2019 International Conference on Robotics and Automation (ICRA), pages 6344–6350, 2019

  188. [196]

    Rovins: Robust omnidi- rectional visual inertial navigation system.IEEE Robotics and Automation Letters, 5:6225–6232, 2020

    Hochang Seok and Jongwoo Lim. Rovins: Robust omnidi- rectional visual inertial navigation system.IEEE Robotics and Automation Letters, 5:6225–6232, 2020

  189. [197]

    Equivari- ant networks for pixelized spheres

    Mehran Shakerinava and Siamak Ravanbakhsh. Equivari- ant networks for pixelized spheres. InICML, 2021

  190. [198]

    Viewport- oriented panoramic image inpainting

    Zhuoyi Shang, Yanwei Liu, Guoyi Li, Yunjian Zhang, Jingbo Miao, Jinxia Liu, and Liming Wang. Viewport- oriented panoramic image inpainting. InICIP, 2022

  191. [199]

    Training real-time panoramic object detectors with virtual dataset.ICASSP, 2021

    Qing-Yang Shen, Tian-Guo Huang, Pengxin Ding, and Jiantao He. Training real-time panoramic object detectors with virtual dataset.ICASSP, 2021

  192. [200]

    Pdo-es2cnns: Partial differential operator based equivariant spherical cnns

    Zhengyang Shen, Tiancheng Shen, Zhouchen Lin, and Jinwen Ma. Pdo-es2cnns: Partial differential operator based equivariant spherical cnns. InAAAI, 2021

  193. [201]

    Panoformer: Panorama transformer for indoor 360◦ depth estimation

    Zhijie Shen, Chunyu Lin, Kang Liao, Lang Nie, Zishuo Zheng, and Yao Zhao. Panoformer: Panorama transformer for indoor 360◦ depth estimation. InECCV, 2022

  194. [202]

    Disentangling orthogo- nal planes for indoor panoramic room layout estimation with cross-scale distortion awareness

    Zhijie Shen, Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Shuai Zheng, and Yao Zhao. Disentangling orthogo- nal planes for indoor panoramic room layout estimation with cross-scale distortion awareness. InCVPR, 2023

  195. [203]

    Panoflow: Learning 360° optical flow for surrounding tem- poral understanding

    Hao Shi, Yifan Zhou, Kailun Yang, Xiaoting Yin, Ze Wang, Yaozu Ye, Zhe Yin, Shi Meng, Peng Li, and Kaiwei Wang. Panoflow: Learning 360° optical flow for surrounding tem- poral understanding. IEEE Transactions on Intelligent Transportation Systems, 24(5):5570–5585, 2023

  196. [204]

    Convolutional lstm network: A machine learning approach for precipitation nowcasting

    Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. Convolutional lstm network: A machine learning approach for precipitation nowcasting. NIPS, 2015

  197. [205]

    Conditional 360-degree image synthesis for immersive indoor scene decoration

    Kashun Shum, Hong-Wing Pang, Binh-Son Hua, Duc Thanh Nguyen, and Sai-Kit Yeung. Conditional 360-degree image synthesis for immersive indoor scene decoration. ICCV, pages 4455–4465, 2023

  198. [206]

    Very deep convo- lutional networks for large-scale image recognition.ArXiv, 2014

    Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition.ArXiv, 2014. 34 Hao Ai 1 et al

  199. [207]

    An overview of multi-view fisheye for vision- first autonomous driving.Authorea Preprints, 2024

    Apoorv Singh. An overview of multi-view fisheye for vision- first autonomous driving.Authorea Preprints, 2024

  200. [208]

    Saliency in vr: How do people explore virtual environments? TVCG, 2018

    Vincent Sitzmann, Ana Serrano, Amy Pavel, Maneesh Agrawala, Diego Gutierrez, Belen Masia, and Gordon Wet- zstein. Saliency in vr: How do people explore virtual environments? TVCG, 2018

  201. [209]

    Learning structured output representation using deep conditional generative models

    Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. InNIPS, 2015

  202. [210]

    Hdr environment map estimation for real-time augmented reality.CVPR, 2021

    Gowri Somanath and Daniel Kurz. Hdr environment map estimation for real-time augmented reality.CVPR, 2021

  203. [211]

    Funkhouser

    Shuran Song and Thomas A. Funkhouser. Neural illumina- tion: Lighting prediction for indoor environments.CVPR, pages 6911–6919, 2019

  204. [212]

    The replica dataset: A digital replica of indoor spaces.ArXiv, 2019

    Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur- Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces.ArXiv, 2019

  205. [213]

    Gpr-net: Multi-view layout estimation via a geometry-aware panorama registration network

    Jheng-Wei Su, Chi-Han Peng, Peter Wonka, and Hung- Kuo Chu. Gpr-net: Multi-view layout estimation via a geometry-aware panorama registration network. InCVPR, 2023

  206. [214]

    Learning spherical convolution for fast features from 360◦ imagery

    Yu-Chuan Su and Kristen Grauman. Learning spherical convolution for fast features from 360◦ imagery. InNIPS, 2017

  207. [215]

    Kernel transformer networks for compact spherical convolution.CVPR, 2019

    Yu-Chuan Su and Kristen Grauman. Kernel transformer networks for compact spherical convolution.CVPR, 2019

  208. [216]

    Learning spherical convolution for 360 recognition.TPAMI, 2021

    Yu-Chuan Su and Kristen Grauman. Learning spherical convolution for 360 recognition.TPAMI, 2021

  209. [217]

    Perceptual quality assessment of 360° images based on generative scanpath representation

    Xiangjie Sui, Hanwei Zhu, Xuelin Liu, Yuming Fang, Shiqi Wang, and Zhou Wang. Perceptual quality assessment of 360° images based on generative scanpath representation. ArXiv, 2023

  210. [218]

    Horizonnet: Learning room layout with 1d repre- sentation and pano stretch data augmentation.CVPR, 2019

    Cheng Sun, Chi-Wei Hsiao, Min Sun, and Hwann-Tzong Chen. Horizonnet: Learning room layout with 1d repre- sentation and pano stretch data augmentation.CVPR, 2019

  211. [219]

    Hohonet: 360 indoor holistic understanding with latent horizontal features

    Cheng Sun, Min Sun, and Hwann-Tzong Chen. Hohonet: 360 indoor holistic understanding with latent horizontal features. In CVPR, 2021

  212. [220]

    Seg2reg: Differentiable 2d seg- mentation to 1d regression rendering for 360 room layout reconstruction

    Cheng Sun, Wei-En Tai, Yu-Lin Shih, Kuan-Wei Chen, Yong-Jing Syu, Kent Selwyn The, Yu-Chiang Frank Wang, and Hwann-Tzong Chen. Seg2reg: Differentiable 2d seg- mentation to 1d regression rendering for 360 room layout reconstruction. ArXiv, 2023

  213. [221]

    Cviqd: Subjective quality evaluation of compressed virtual reality images.ICIP, 2017

    Wei Sun, Ke Gu, Guangtao Zhai, Siwei Ma, Weisi Lin, and Patrick Le Callet. Cviqd: Subjective quality evaluation of compressed virtual reality images.ICIP, 2017

  214. [222]

    Mc360iqa: A multi-channel cnn for blind 360-degree image quality assessment.IEEE J-STSP, 2020

    Wei Sun, Xiongkuo Min, Guangtao Zhai, Ke Gu, Huiyu Duan, and Siwei Ma. Mc360iqa: A multi-channel cnn for blind 360-degree image quality assessment.IEEE J-STSP, 2020

  215. [223]

    Weighted-to-spherically- uniform quality evaluation for omnidirectional video.IEEE SPL, 2017

    Yule Sun, Ang Lu, and Lu Yu. Weighted-to-spherically- uniform quality evaluation for omnidirectional video.IEEE SPL, 2017

  216. [224]

    Saliency map esti- mation for omni-directional image considering prior distri- butions

    Tatsuya Suzuki and Takao Yamanaka. Saliency map esti- mation for omni-directional image considering prior distri- butions. SMC, 2018

  217. [225]

    Vr iqa net: Deep virtual reality image quality assessment using adversarial learning

    Heoun taek Lim, Hak Gu Kim, and Yong Man Ro. Vr iqa net: Deep virtual reality image quality assessment using adversarial learning. ICASSP, 2018

  218. [226]

    Distortion-aware convolutional filters for dense prediction in panoramic images

    Keisuke Tateno, Nassir Navab, and Federico Tombari. Distortion-aware convolutional filters for dense prediction in panoramic images. InECCV, 2018

  219. [227]

    Viewport-sphere-branch net- work for blind quality assessment of stitched 360° omnidi- rectional images

    Chongzhen Tian, Feng Shao, Xiongli Chai, Qiuping Jiang, Long Xu, and Yo-Sung Ho. Viewport-sphere-branch net- work for blind quality assessment of stitched 360° omnidi- rectional images. IEEE TCSVT, 2023

  220. [228]

    St360iq: No-reference omnidirectional image quality assessment with spherical vision transformers.ICASSP, 2023

    Nafiseh Jabbari Tofighi, Mohamed Hedi Elfkir, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, and Aykut Er- dem. St360iq: No-reference omnidirectional image quality assessment with spherical vision transformers.ICASSP, 2023

  221. [229]

    Object detection for panoramic images based on ms-rpn structure in traffic road scenes

    Guofeng Tong, Huairong Chen, Yong Li, Xiance Du, and Qingchun Zhang. Object detection for panoramic images based on ms-rpn structure in traffic road scenes. IET Comput. Vis., 2019

  222. [230]

    Saliency-driven omnidirectional imaging adaptive coding: Modeling and assessment

    Guilherme Luz Tortorella, João Ascenso, Catarina Brites, and Fernando Pereira. Saliency-driven omnidirectional imaging adaptive coding: Modeling and assessment. MMSP, 2017

  223. [231]

    Sslayout360: Semi-supervised indoor layout estimation from 360◦ panorama

    Phi Vu Tran. Sslayout360: Semi-supervised indoor layout estimation from 360◦ panorama. CVPR, 2021

  224. [232]

    Automatic 360 mono-stereo panorama gen- eration using a cost-effective multi-camera system.Sensors, 20(11):3097, 2020

    Hayat Ullah, Osama Zia, Jun Ho Kim, Kyungjin Han, and Jong Weon Lee. Automatic 360 mono-stereo panorama gen- eration using a cost-effective multi-camera system.Sensors, 20(11):3097, 2020

  225. [233]

    Testbed for subjective evaluation of omnidirectional visual content

    Evgeniy Upenik, Martin Rerábek, and Touradj Ebrahimi. Testbed for subjective evaluation of omnidirectional visual content. PCS, 2016

  226. [234]

    Subjective assessment of 360° image pro- jection formats

    Tran Thi Hai Uyen, Oh-Jin Kwon, Seungcheol Choi, and Ikram Hussain. Subjective assessment of 360° image pro- jection formats. IEEE Access, 2020

  227. [235]

    Lsd: A fast line segment detector with a false detection control.PAMI, 2008

    Rafael Grompone Von Gioi, Jeremie Jakubowicz, Jean- Michel Morel, and Gregory Randall. Lsd: A fast line segment detector with a false detection control.PAMI, 2008

  228. [236]

    Self-supervised learning of depth and camera motion from 360◦ videos

    Fu-En Wang, Hou-Ning Hu, Hsien-Tzu Cheng, Juan-Ting Lin, Shang-Ta Yang, Meng-Li Shih, Hung-Kuo Chu, and Min Sun. Self-supervised learning of depth and camera motion from 360◦ videos. In ACCV, 2018

  229. [237]

    Bifuse: Monocular 360 depth estima- tion via bi-projection fusion

    Fu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu, and Yi-Hsuan Tsai. Bifuse: Monocular 360 depth estima- tion via bi-projection fusion. InCVPR, 2020

  230. [238]

    Led2-net: Monocular 360 ◦ layout estimation via differentiable depth rendering.CVPR, 2021

    Fu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu, and Yi-Hsuan Tsai. Led2-net: Monocular 360 ◦ layout estimation via differentiable depth rendering.CVPR, 2021

  231. [239]

    Stylelight: Hdr panorama generation for lighting estimation and editing

    Guangcong Wang, Yinuo Yang, Chen Change Loy, and Ziwei Liu. Stylelight: Hdr panorama generation for lighting estimation and editing. InECCV, 2022

  232. [240]

    Customizing 360-degree panoramas through text-to-image diffusion models

    Hai Wang, Xiaoyu Xiang, Yuchen Fan, and Jing-Hao Xue. Customizing 360-degree panoramas through text-to-image diffusion models. InWACV, 2024

  233. [241]

    Psm- net: Position-aware stereo merging network for room layout estimation

    Haiyan Wang, Will Hutchcroft, Yuguang Li, Zhiqiang Wan, Ivaylo Boyadzhiev, Yingli Tian, and Sing Bing Kang. Psm- net: Position-aware stereo merging network for room layout estimation. In CVPR, 2022

  234. [242]

    360-degree panorama generation from few unreg- istered nfov images.ACM MM, 2023

    Jiong-Qi Wang, Ziyu Chen, Jun Ling, Rong Xie, and Li Song. 360-degree panorama generation from few unreg- istered nfov images.ACM MM, 2023

  235. [243]

    Object detection in curved space for 360-degree camera

    Kuan-Hsun Wang and Shang-Hong Lai. Object detection in curved space for 360-degree camera. InICASSP, 2019

  236. [244]

    360sd-net: 360 stereo depth estimation with learnable cost volume

    Ning-Hsu Wang, Bolivar Solarte, Yi-Hsuan Tsai, Wei-Chen Chiu, and Min Sun. 360sd-net: 360 stereo depth estimation with learnable cost volume. In2020 IEEE International Conference on Robotics and Automation (ICRA), pages 582–588. IEEE, 2020

  237. [245]

    Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

    Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. InICML, 2022

  238. [246]

    360dvd: Controllable panorama video generation with 360-degree video diffusion model.ArXiv, 2024

    Qian Wang, Weiqi Li, Chong Mou, Xinhua Cheng, and Jian Zhang. 360dvd: Controllable panorama video generation with 360-degree video diffusion model.ArXiv, 2024. A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision 35

  239. [247]

    Lf-vio: A visual-inertial-odometry framework for large field- of-view cameras with negative plane

    Ze Wang, Kailun Yang, Haowen Shi, and Kaiwei Wang. Lf-vio: A visual-inertial-odometry framework for large field- of-view cameras with negative plane. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4423–4430, 2022

  240. [248]

    Omnislam: Omnidirectional local- ization and dense mapping for wide-baseline multi-camera systems

    Changhee Won, Hochang Seok, Zhaopeng Cui, Marc Polle- feys, and Jongwoo Lim. Omnislam: Omnidirectional local- ization and dense mapping for wide-baseline multi-camera systems. 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 559–566, 2020

  241. [249]

    A spherical convolution approach for learning long term viewport prediction in 360 immersive video

    Chenglei Wu, Ruixiao Zhang, Zhi Wang, and Lifeng Sun. A spherical convolution approach for learning long term viewport prediction in 360 immersive video. In AAAI, 2020

  242. [250]

    Pan- odiffusion: 360-degree panorama outpainting via diffusion

    Tianhao Wu, Chuanxia Zheng, and Tat-Jen Cham. Pan- odiffusion: 360-degree panorama outpainting via diffusion. In ICLR, 2024

  243. [251]

    Assessor360: Multi-sequence network for blind omnidirectional image quality assessment

    Tianhe Wu, Shuwei Shi, Haoming Cai, Mingdeng Cao, Jing Xiao, Yinqiang Zheng, and Yujiu Yang. Assessor360: Multi-sequence network for blind omnidirectional image quality assessment. InNIPS, 2023

  244. [252]

    Zamir, Zhiyang He, Alexander Sax, Ji- tendra Malik, and Silvio Savarese

    Fei Xia, Amir R. Zamir, Zhiyang He, Alexander Sax, Ji- tendra Malik, and Silvio Savarese. Gibson Env: real-world perception for embodied agents. InCVPR, 2018

  245. [253]

    Ehinger, Aude Oliva, and Anto- nio Torralba

    Jianxiong Xiao, Krista A. Ehinger, Aude Oliva, and Anto- nio Torralba. Recognizing scene viewpoint using panoramic place representation. CVPR, 2012

  246. [254]

    Effective convolutional neural network layers in flow estimation for omnidirectional images

    Shuang Xie, Po Kong Lai, Robert Laganiere, and Jochen Lang. Effective convolutional neural network layers in flow estimation for omnidirectional images. In2019 Interna- tional Conference on 3D Vision (3DV), pages 671–680. IEEE, 2019

  247. [255]

    Omnidirectional dense slam for back-to-back fisheye cameras.2024 IEEE International Conference on Robotics and Automation (ICRA), pages 1653–1660, 2024

    Weijian Xie, Guanyi Chu, Quanhao Qian, Yihao Yu, Shangjin Zhai, Danpeng Chen, Nan Wang, Hujun Bao, and Guofeng Zhangv. Omnidirectional dense slam for back-to-back fisheye cameras.2024 IEEE International Conference on Robotics and Automation (ICRA), pages 1653–1660, 2024

  248. [256]

    Gaussian label distribu- tion learning for spherical image object detection.CVPR, pages 1033–1042, 2023

    Hang Xu, Xinyuan Liu, Qiang Zhao, Yike Ma, Cheng- gang Clarence Yan, and Feng Dai. Gaussian label distribu- tion learning for spherical image object detection.CVPR, pages 1033–1042, 2023

  249. [257]

    Pandora: A panoramic detection dataset for object with orientation

    Hang Xu, Qiang Zhao, Yike Ma, Xiao-Di Li, Peng Yuan, Bailan Feng, Chenggang Clarence Yan, and Feng Dai. Pandora: A panoramic detection dataset for object with orientation. In European Conference on Computer Vision, 2022

  250. [258]

    Blind omnidirec- tional image quality assessment with viewport oriented graph convolutional networks.IEEE TCSVT, 2020

    Jiahua Xu, Wei Zhou, and Zhibo Chen. Blind omnidirec- tional image quality assessment with viewport oriented graph convolutional networks.IEEE TCSVT, 2020

  251. [259]

    Viewport-based cnn: A multi-task approach for assessing 360◦ video quality.TPAMI, 2020

    Mai Xu, Lai Jiang, Chen Li, Zulin Wang, and Xiaom- ing Tao. Viewport-based cnn: A multi-task approach for assessing 360◦ video quality.TPAMI, 2020

  252. [260]

    Assessing visual quality of omnidirectional videos

    Mai Xu, Chen Li, Zhenzhong Chen, Zulin Wang, and Zhenyu Guan. Assessing visual quality of omnidirectional videos. IEEE TCSVT, 2019

  253. [261]

    State-of-the-art in 360◦ video/image processing: Percep- tion, assessment and compression.IEEE J-STSP, 2020

    Mai Xu, Chen Li, Shanyi Zhang, and Patrick Le Callet. State-of-the-art in 360◦ video/image processing: Percep- tion, assessment and compression.IEEE J-STSP, 2020

  254. [262]

    Predicting head move- ment in panoramic video: A deep reinforcement learning approach

    Mai Xu, Yuhang Song, Jianyi Wang, MingLang Qiao, Liangyu Huo, and Zulin Wang. Predicting head move- ment in panoramic video: A deep reinforcement learning approach. IEEE T-PAMI, 2018

  255. [263]

    Predicting head move- ment in panoramic video: A deep reinforcement learning approach

    Mai Xu, Yuhang Song, Jianyi Wang, Minglang Qiao, Liangyu Huo, and Zulin Wang. Predicting head move- ment in panoramic video: A deep reinforcement learning approach. TPAMI, 2019

  256. [264]

    Saliency prediction on omnidirectional image with generative adversarial imitation learning.IEEE TIP, 2019

    Mai Xu, Li Yang, Xiaoming Tao, Yiping Duan, and Zulin Wang. Saliency prediction on omnidirectional image with generative adversarial imitation learning.IEEE TIP, 2019

  257. [266]

    Saliency prediction on omnidirectional image with generative adversarial imitation learning.IEEE TIP, 2021

    Mai Xu, Li Yang, Xiaoming Tao, Yiping Duan, and Zulin Wang. Saliency prediction on omnidirectional image with generative adversarial imitation learning.IEEE TIP, 2021

  258. [267]

    Gaze prediction in dynamic 360◦ immersive videos

    Yanyu Xu, Yanbing Dong, Junru Wu, Zhengzhong Sun, Zhiru Shi, Jingyi Yu, and Shenghua Gao. Gaze prediction in dynamic 360◦ immersive videos. CVPR, 2018

  259. [268]

    Deep learning on image stitching with multi-viewpoint images: A survey.Neural Processing Letters, 55:3863–3898, 2023

    Ni Yan, Yupeng Mei, Ling Xu, Huihui Yu, Bo Sun, Zimao Wang, and Yingyi Chen. Deep learning on image stitching with multi-viewpoint images: A survey.Neural Processing Letters, 55:3863–3898, 2023

  260. [269]

    Distortion and uncertainty aware loss for panoramic depth completion

    Zhiqiang Yan, Xiang Li, Kun Wang, Shuo Chen, Jun Li, and Jian Yang. Distortion and uncertainty aware loss for panoramic depth completion. InInternational Conference on Machine Learning, pages 39099–39109. PMLR, 2023

  261. [270]

    Multi-modal masked pre-training for monocular panoramic depth completion

    Zhiqiang Yan, Xiang Li, Kun Wang, Zhenyu Zhang, Jun Li, and Jian Yang. Multi-modal masked pre-training for monocular panoramic depth completion. In European Conference on Computer Vision, pages 378–395. Springer, 2022

  262. [271]

    Can we pass beyond the field of view? panoramic annular semantic segmentation for real-world surrounding percep- tion

    Kailun Yang, Xinxin Hu, Luis M Bergasa, Eduardo Romera, Xiao Huang, Dongming Sun, and Kaiwei Wang. Can we pass beyond the field of view? panoramic annular semantic segmentation for real-world surrounding percep- tion. In IV, 2019

  263. [272]

    Pass: Panoramic annular semantic segmentation.IEEE TITS, 2020

    Kailun Yang, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, and Kaiwei Wang. Pass: Panoramic annular semantic segmentation.IEEE TITS, 2020

  264. [273]

    Ds-pass: Detail-sensitive panoramic annular semantic segmentation through swaft- net for surrounding sensing.IV, 2019

    Kailun Yang, Xinxin Hu, Hao Chen, Kaite Xiang, Kaiwei Wang, and Rainer Stiefelhagen. Ds-pass: Detail-sensitive panoramic annular semantic segmentation through swaft- net for surrounding sensing.IV, 2019

  265. [274]

    Omnisupervised omnidirectional semantic segmentation.IEEE TITS, 2020

    Kailun Yang, Xinxin Hu, Yicheng Fang, Kaiwei Wang, and Rainer Stiefelhagen. Omnisupervised omnidirectional semantic segmentation.IEEE TITS, 2020

  266. [275]

    Is context-aware cnn ready for the surroundings? panoramic semantic segmentation in the wild.IEEE TIP, 2021

    Kailun Yang, Xinxin Hu, and Rainer Stiefelhagen. Is context-aware cnn ready for the surroundings? panoramic semantic segmentation in the wild.IEEE TIP, 2021

  267. [276]

    Capturing omni-range context for omnidirectional segmentation

    Kailun Yang, Jiaming Zhang, Simon Reiß, Xinxin Hu, and Rainer Stiefelhagen. Capturing omni-range context for omnidirectional segmentation. CVPR, 2021

  268. [277]

    Spatialattention- based non-reference perceptual quality prediction network for omnidirectional images.ICME, 2021

    LiYang,MaiXu,XinDeng,andBoFeng. Spatialattention- based non-reference perceptual quality prediction network for omnidirectional images.ICME, 2021

  269. [278]

    Tvformer: Trajectory-guided visual quality assessment on 360° images with transformers.ACM MM, 2022

    Li Yang, Mai Xu, Tie Liu, Liangyu Huo, and Xinbo Gao. Tvformer: Trajectory-guided visual quality assessment on 360° images with transformers.ACM MM, 2022

  270. [279]

    Rotation equivariant graph convo- lutional network for spherical image classification.CVPR, 2020

    Qin Yang, Chenglin Li, Wenrui Dai, Junni Zou, Guo-Jun Qi, and Hongkai Xiong. Rotation equivariant graph convo- lutional network for spherical image classification.CVPR, 2020

  271. [280]

    Dula-net: A dual-projection network for estimating room layouts from a single rgb panorama

    Shang-Ta Yang, Fu-En Wang, Chi-Han Peng, Peter Wonka, Min Sun, and Hung kuo Chu. Dula-net: A dual-projection network for estimating room layouts from a single rgb panorama. CVPR, 2019

  272. [281]

    Pastiche master: Exemplar-based high-resolution portrait style transfer.CVPR, 2022

    Shuai Yang, Liming Jiang, Ziwei Liu, and Chen Change Loy. Pastiche master: Exemplar-based high-resolution portrait style transfer.CVPR, 2022

  273. [282]

    Object detection in equirectangular panorama

    Wenyan Yang, Yanlin Qian, Joni-Kristian Kämäräinen, Francesco Cricri, and Lixin Fan. Object detection in equirectangular panorama. InICPR, 2018. 36 Hao Ai 1 et al

  274. [283]

    Mcov-slam: A multicamera omnidirectional visual slam system.IEEE/ASME Trans- actions on Mechatronics, 29:3556–3567, 2024

    Yi Yang, Miaoxin Pan, Di Tang, Tao Wang, Yufeng Yue, Tong Liu, and Mengyin Fu. Mcov-slam: A multicamera omnidirectional visual slam system.IEEE/ASME Trans- actions on Mechatronics, 29:3556–3567, 2024

  275. [284]

    Salgfcn: Graph based fully convolutional network for panoramic saliency prediction.VCIP, 2021

    Yiwei Yang, Yucheng Zhu, Zhongpai Gao, and Guangtao Zhai. Salgfcn: Graph based fully convolutional network for panoramic saliency prediction.VCIP, 2021

  276. [285]

    A sur- vey on adaptive 360 video streaming: solutions, challenges and opportunities

    Abid Yaqoob, Ting Bi, and Gabriel-Miro Muntean. A sur- vey on adaptive 360 video streaming: solutions, challenges and opportunities. IEEE Commun. Surv. Tutor., 2020

  277. [286]

    Spheresr

    Youngho Yoon, Inchul Chung, Lin Wang, and Kuk-Jin Yoon. Spheresr. CVPR, 2022

  278. [287]

    Grid based spherical cnn for object detection from panoramic images.Sensors, 2019

    Dawen Yu and Shunping Ji. Grid based spherical cnn for object detection from panoramic images.Sensors, 2019

  279. [288]

    Osrt: Omnidirectional image super- resolution with distortion-aware transformer

    Fanghua Yu, Xintao Wang, Mingdeng Cao, Gen Li, Ying Shan, and Chao Dong. Osrt: Omnidirectional image super- resolution with distortion-aware transformer. InCVPR, 2023

  280. [289]

    Panelnet: Understanding 360 indoor environment via panel representation

    Haozheng Yu, Lu He, Bing Jian, Weiwei Feng, and Shan Liu. Panelnet: Understanding 360 indoor environment via panel representation. InCVPR, 2023

  281. [290]

    Ap- plications of deep learning for top-view omnidirectional imaging: A survey.CVPR Workshop, 2023

    Jingrui Yu, Ana Pérez Grassi, and Gangolf Hirtz. Ap- plications of deep learning for top-view omnidirectional imaging: A survey.CVPR Workshop, 2023

  282. [291]

    Panoramic image inpainting with gated convolution and contextual reconstruction loss.ArXiv, 2024

    Li Yu, Yanjun Gao, Farhad Pakdaman, and Moncef Gab- bouj. Panoramic image inpainting with gated convolution and contextual reconstruction loss.ArXiv, 2024

  283. [292]

    Yu, Haricharan Lakshman, and Bernd Girod

    Matt C. Yu, Haricharan Lakshman, and Bernd Girod. A framework to evaluate omnidirectional video coding schemes. ISMAR, 2015

  284. [293]

    360 optical flow using tangent images

    Mingze Yuan and Christian Richardt. 360 optical flow using tangent images. InBritish Machine Vision Confer- ence:(BMVC). Christian Richardt, 2021

  285. [294]

    Lf-vislam: A slam framework for large field-of-view cameras with negative imaging plane on mobile agents.IEEE Transactions on Automation Science and Engineering, 21:6321–6335, 2022

    Ze yuan Wang, Kailun Yang, Hao miao Shi, Peng Li, Fei Gao, Jian Bai, and Kaiwei Wang. Lf-vislam: A slam framework for large field-of-view cameras with negative imaging plane on mobile agents.IEEE Transactions on Automation Science and Engineering, 21:6321–6335, 2022

  286. [295]

    Im- proving 360 monocular depth estimation via non-local dense prediction transformer and joint supervised and self-supervised learning

    Il Dong Yun, Hyuk-Jae Lee, and Chae-Eun Rhee. Im- proving 360 monocular depth estimation via non-local dense prediction transformer and joint supervised and self-supervised learning. InAAAI, 2021

  287. [296]

    Egformer: Equirectangular geometry- biased transformer for 360 depth estimation

    Ilwi Yun, Chanyong Shin, Hyunku Lee, Hyuk-Jae Lee, and Chae Eun Rhee. Egformer: Equirectangular geometry- biased transformer for 360 depth estimation. InICCV, 2023

  288. [297]

    Quality metric for spherical panoramic video

    Vladyslav Zakharchenko, Kwang Pyo Choi, and Jeonghoon Park. Quality metric for spherical panoramic video. In Optical Engineering + Applications, 2016

  289. [298]

    Gmlight: Lighting estimation via geometric distribu- tion approximation

    Fangneng Zhan, Yingchen Yu, Rongliang Wu, Changgong Zhang, Shijian Lu, Ling Shao, Feiying Ma, and Xuansong Xie. Gmlight: Lighting estimation via geometric distribu- tion approximation. IEEE TIP, 2021

  290. [299]

    Em- light: Lighting estimation via spherical distribution ap- proximation

    Fangneng Zhan, Changgong Zhang, Yingchen Yu, Yuan Chang, Shijian Lu, Feiying Ma, and Xuansong Xie. Em- light: Lighting estimation via spherical distribution ap- proximation. ArXiv, 2020

  291. [300]

    Chao Zhang, Stephan Liwicki, Sen He, William H. B. Smith, and Roberto Cipolla. Hexnet: An orientation- aware deep learning framework for omni-directional input. TPAMI, 2023

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.