Pith. sign in

REVIEW 4 cited by

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03403 v2 pith:4EK4ZTV6 submitted 2023-06-06 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords sphericalpanoramicgeometry-awaresgat4passdatadisturbanceimagepass
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D properties of original $360^{\circ}$ data. Therefore, their performance will drop a lot when inputting panoramic images with the 3D disturbance. To be more robust to 3D disturbance, we propose our Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation (SGAT4PASS), considering 3D spherical geometry knowledge. Specifically, a spherical geometry-aware framework is proposed for PASS. It includes three modules, i.e., spherical geometry-aware image projection, spherical deformable patch embedding, and a panorama-aware loss, which takes input images with 3D disturbance into account, adds a spherical geometry-aware constraint on the existing deformable patch embedding, and indicates the pixel density of original $360^{\circ}$ data, respectively. Experimental results on Stanford2D3D Panoramic datasets show that SGAT4PASS significantly improves performance and robustness, with approximately a 2% increase in mIoU, and when small 3D disturbances occur in the data, the stability of our performance is improved by an order of magnitude. Our code and supplementary material are available at https://github.com/TencentARC/SGAT4PASS.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

    cs.CV 2026-03 conditional novelty 6.5 of 10

    A training-free pipeline jointly segments panoramic images and reconstructed point clouds with open-vocabulary language queries via tangential decomposition, instance proposals, CLIP alignment, and depth-based label t...

  2. Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Survey organizing panoramic scene analysis literature by architectural design and training paradigm, identifying the absence of methods achieving both strict spherical equivariance and full reuse of perspective-pretra...

  3. Deformable Mamba for Wide Field of View Segmentation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    On five wide-FoV segmentation benchmarks, a Mamba-plus-deformable-convolution decoder beats common segmentation heads with the same backbones while cutting decoder FLOPs by roughly 97% versus UperHead.

  4. 360-Degree Full-view Image Segmentation by Spherical Convolution compatible with Large-scale Planar Pre-trained Models

    cs.CV 2025-07 reject novelty 4.0 of 10

    A spherical input resampling trick transfers planar pretrained backbone weights to 360-degree segmentation, with a channel-attention branch adding moderate gains.

Pith tools