REVIEW 3 major objections 6 minor 1 cited by
PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read PanoSLAM is the first SLAM system to build panoptic 3D scene maps from raw RGB-D video without manual labels.
desk verdict A plausible label-free panoptic Gaussian SLAM with a real evaluation gap: the map is only scored in 2D on training views. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the Spatial-Temporal Lifting (STL) module, a voxel-based multi-view label refinement step: pixels from different frames whose unprojected 3D points fall in the same voxel are treated as correspondences, and their region predictions are replaced by the mean prediction inside that voxel. This averaging across views and time is the device that suppresses label noise from the 2D vision model and makes the pseudo-labels consistent enough to optimize the Gaussian map. The map itself is composed of 3D Gaussians with 13 parameters — RGB color, center, radius, opacity, plus a 3D semantic embedding, semantic radius, and semantic opacity — which allow differentiable splatting of depth, color, and panoptic predictions.
What would settle it
A direct test would be to sweep the voxel size $S_n$ on a single Replica scene and plot panoptic quality (PQ) and tracking ATE: if the method is right, there is a range of $S_n$ where both improve over the raw 2D pseudo-labels, and outside that range PQ should degrade sharply at object boundaries. A specific observation that would contradict the central claim is finding a scene where increasing $S_n$ never helps, i.e., where voxel averaging never beats the noisy 2D predictions.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that panoptic 3D scene reconstruction can be achieved online in a SLAM pipeline without any manual labels, by equipping each 3D Gaussian with a semantic embedding and using a Spatial-Temporal Lifting module to refine noisy 2D panoptic pseudo-labels across views. The lifting works by unprojecting pixels from multiple frames, assigning them to a shared voxel grid, and replacing each region prediction with the voxel-average prediction before re-rendering; this makes the 3D semantics coherent and improves tracking by reducing drift. With 13 parameters per Gaussian and a mask-classification panoptic formulation, the system renders depth, color, semantics, and instances from arbitrary viewpoints. The authors claim this is the first demonstration of panoptic 3D reconstruction of open-world environments directly from RGB-D video.
Load-bearing premise
The whole method rests on the assumption that pixels whose unprojected points land in the same voxel truly share one semantic label, so the voxel size must be small enough to respect object boundaries and large enough to collect multiple views — a value the paper never states.
Editorial extensions
If this is right
- Dense semantic SLAM no longer requires manually labeled data; a robot or AR device could map a novel indoor scene and immediately get object-level understanding.
- Because STL couples tracking and semantics, improved label consistency feeds back into lower camera drift, so semantic and geometric accuracy improve together.
- The same pipeline extends to any scene categories the 2D vision model can segment, so the panoptic map is open-world rather than limited to a fixed label set.
- Panoptic rendering from Gaussians gives a single representation for appearance, depth, and semantics, which could serve downstream tasks such as robotic manipulation or scene editing.
Reading between the lines
- The voxel-average operation is a form of multi-view agreement akin to test-time label smoothing; it should be most effective when many views cover each voxel, so performance likely varies with camera trajectory density — a testable hypothesis the paper does not report.
- If the assumption that voxels are smaller than object boundaries fails, instance boundaries will be blurred; measuring panoptic quality as a function of voxel size $S_n$ would provide a sensitivity curve the paper leaves implicit.
- The framework's reliance on a frozen 2D panoptic model means the ceiling of the 3D map is set by 2D prediction quality; fusing multi-view information back into the 2D model, as the authors suggest for future work, could push the system further.
- The shared voxel grid links this approach to classic volumetric fusion; one could unify STL with occupancy or truncated signed distance mapping to reduce memory overhead at large scales.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. PanoSLAM presents a dense RGB-D SLAM system built on 3D Gaussian Splatting that jointly estimates camera poses, geometry, and panoptic (semantic plus instance) information from unlabeled video. The method extends SplaTAM by adding per-Gaussian semantic embeddings and opacity, generates 2D panoptic pseudo-labels with the frozen SEEM model, and introduces a Spatial-Temporal Lifting (STL) module that averages region predictions across views within small 3D voxels before using them as supervision. Experiments on Replica and two ScanNet++ scenes report tracking ATE, depth L1, rendering metrics, and semantic/panoptic metrics, together with ablations on the STL module and running times.
Significance. If properly validated, the paper would make a useful contribution: it is a credible attempt at label-free panoptic mapping inside an online Gaussian SLAM system, with tracking accuracy comparable to SplaTAM and a clear system design. The release of code is a practical strength. However, the central claim of 'panoptic 3D scene reconstruction' is not yet supported by the evidence, because all segmentation metrics are computed by rendering the optimized Gaussians at the same training views used for optimization, which does not demonstrate a coherent 3D map. The STL module's evaluation also conflates label self-consistency with label accuracy. The novelty of the pipeline is real, but the evaluation needs to target the 3D representation directly or through held-out views.
major comments (3)
- [Section 4, Tables 1 and 3] The panoptic and semantic segmentation results are obtained by rendering the optimized Gaussians on the training frames and comparing those 2D renders to 2D ground truth. Because the Gaussians are optimized against exactly those frames—including the SEEM pseudo-labels refined by the STL module—the evaluation cannot rule out per-view memorization of 2D labels. To support the claim of 'panoptic 3D scene reconstruction', the authors should evaluate the 3D map itself, for example by rendering at held-out or novel viewpoints, by projecting the semantic/instance labels of the Gaussians into a 3D point cloud and comparing against 3D ground truth, or by measuring cross-view label consistency of the reconstructed map. As it stands, Tables 1 and 3 certify a 2D rendering capability, not a 3D panoptic reconstruction.
- [Section 3.2, Eq. (12), and Section 4.2, Table 7] The STL module produces refined pseudo-labels by averaging, within each voxel, the region predictions from the same SEEM model that later provides the supervision for the Gaussians (Eq. 12 and Eq. 13). The ablation in Table 7 shows a large improvement from adding STL (w/o STL PQ 7.3 vs. Ours PQ 19.9 on room0), but this reported gain reflects agreement with the smoothed target, not necessarily agreement with ground truth. The measured improvement is therefore partly circular; it demonstrates that the rendered outputs become more self-consistent with the spatially averaged pseudo-labels, but it does not establish that the refined labels or the resulting 3D map are more accurate. To substantiate the claimed benefit of STL, the authors should evaluate against ground truth at held-out views or in 3D, and ideally compare the smoothed pseudo-labels against ground truth as well.
- [Section 3.2, Eq. (12)] The voxel size Sn used for the spatial-temporal averaging in Eq. (12) is never specified in the methodology or the experiments. This parameter directly controls the trade-off between boundary preservation (small voxels) and multi-view denoising (large voxels), and it is essential for reproducing the reported results. The paper should state the value of Sn used for each dataset and provide an ablation over this parameter, at least on one Replica scene.
minor comments (6)
- [Table 1] The header 'mIoU(%' has an unbalanced parenthesis; it should read 'mIoU (%)'.
- [Section 4.2, Table 7] The ablation table uses 'STL-2' while the text defines 'STL-(n)' with n time-steps; please clarify what 'STL-2' denotes and why the full model uses n=4.
- [Table 6 caption] The caption contains a typo: 'Comparion' should be 'Comparison'.
- [Section 4, Implementation Details] The manuscript says 'When working with the Replica dataset, the image number set to 4.' Please specify the exact hyperparameter (the number T of frames used in the optimization window in Eq. 13) and give the corresponding value for the ScanNet++ experiments.
- [Section 4, Evaluation] The paper reports 'single-run results due to the high computational cost associated with training.' Without multiple seeds or error bars, the claimed improvements over baselines may not be statistically distinguishable, especially on the two ScanNet++ scenes; adding variance information or at least a discussion of stability would strengthen the comparison.
- [Section 3.2, Eq. (5)] In the text following Eq. (5), the explanation refers to 'areas where the map lacks density ( S < 0.5)', but the equation uses F(P) < 0.5 and \hat F(P) < 0.5; please make the notation consistent.
Circularity Check
Partial circularity: the panoptic metrics re-render the Gaussians at the same training views whose voxel-averaged SEEM labels (Eq. 12) supervise them (Eq. 13), so reported gains partly measure self-consistency; no load-bearing self-citation or definitional reduction found.
-
fitted input called prediction
[Sec. 3.2, Eq. 13; Sec. 4 Metrics and Tables 1/3/5 (training-view renders)]
""We optimize the Gaussian SLAM system by minimizing ... λ3CE (Ot(M), ˆOt(M)) + λ4DICE (Rt(P ), ˆRt(P ∗)) + λ5SigF (Rt(P ), ˆRt(P ∗))"; "For panoptic segmentation, we adopt the standard PQ (Panoptic Quality), SQ (Segmentation Quality), and RQ (Recognition Quality) metrics [23]"; "Table 5. Quantitative comparison of training view rendering performance on Replica dataset.""
Eq. 13 supervises the rendered panoptic predictions with pseudo-label targets R̂t and Ôt, which are the SEEM outputs refined by Eq. 12 (averaging the SEEM region predictions within each voxel). Gaussian embeddings are fit so that rendering at the chosen training keyframes reproduces these targets. Tables 1/3 then evaluate the claimed 3D panoptic reconstruction by rendering the same Gaussians at the same training viewpoints against ground truth; no novel-view or volumetric panoptic metric is reported. Up to optimization fidelity, the reported PQ/mIoU is therefore the GT accuracy of a smoothed version of the input 2D pseudo-labels reprojected onto the input viewpoints, so the evidence quantity is the fit quantity's close relative.
-
other
[Sec. 3.2, Eq. 12; Sec. 4.2, Table 7]
""we set all region predictions to be identical within the same local voxel gn, i.e., ˆR(P ∗) = 1/|gn| Σ ∗∈gn ˆR(P ∗)"; "Base represents the predictions from the 2D vision model alone" (Sec. 4.2, Table 7)."
Eq. 12 defines the 'refined' region labels as within-voxel averages of the raw SEEM predictions, and Eq. 13 uses those averages as the only panoptic supervision of the Gaussians. The ablation in Table 7 compares the full pipeline against 'base', the raw SEEM predictions themselves, so the measured STL gain partly reflects agreement between the optimized render and a smoothed version of its own supervision source, a self-consistency effect. This is not complete circularity: the improvement is also verified against external ground truth in Tables 1/3, and averaging can genuinely denoise. But the refinement operation never introduces information beyond the SEEM predictions it averages, so its reported benefit is partially self-referential.
full rationale
No load-bearing self-citation or imported uniqueness theorem is present: the construction rests on SplaTAM [18], SEEM [73], and MaskFormer [8], all external, and the many self-citations (e.g., [3,6]) serve as related-work background only. The paper's own Limitations section honestly concedes that the pseudo-labels "can be noisy, particularly in areas with fine and intricate details" (Sec. 5), and this caveat is weighed in the verdict as mitigating, not aggravating. Two partial circularities remain, both quotable. First, the claimed panoptic '3D reconstruction' is never measured in 3D or at novel views: Tables 1 and 3 score the Gaussians by rendering them at the same training viewpoints where Eq. 13 supervised the panoptic embeddings, with targets that are themselves voxel-averaged SEEM predictions (Eq. 12); the reported PQ/mIoU therefore partly measures how faithfully the fit reproduces a smoothed version of its own input labels. The external GT comparison keeps this from being a tautology and gives the central comparison real content. Second, Eq. 12's within-voxel averaging averages the very SEEM predictions that constitute the only panoptic supervision, so the STL ablation in Table 7 partly demonstrates self-consistency of a label source with its own average; the GT-anchored improvement (15.2 to 19.9 PQ on room0) nevertheless shows a genuine denoising effect. Reproducibility is also weakened because the voxel size Sn in Eq. 12 is never specified in Section 4, an omission flagged here though it is not a circularity. Overall the central claim has independent, externally benchmarked content, but the specifically '3D / panoptic' evidence is partially self-referential, hence a score of 4 rather than 0 or 6.
Assumptions & free parameters
free parameters (5)
- Loss weights lambda1..lambda5 =
1, 1, 1, 1, 20
- Number of frames T =
4 for Replica
- Densification threshold multiplier =
50
- Voxel size Sn =
not stated
- Keyframe stride u =
not stated
assumptions (5)
- standard math Differentiable alpha-compositing rendering of 3D Gaussians (SplaTAM) generalizes to semantic and instance channels.
- domain assumption SEEM 2D panoptic predictions are valid pseudo-labels for 3D scene understanding.
- ad hoc to paper Pixels whose unprojected 3D points fall in the same voxel must share a region label.
- domain assumption RGB-D depth is treated as ground truth for projection (Eq. 11) and depth loss.
- domain assumption Hungarian matching between rendered masks and pseudo masks yields a valid soft-assignment.
Cite this review
Pith. "Pith review of PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM." pith.science (2026). https://pith.science/paper/CKRW7SXN
@misc{pith2026250100352,
author = {Pith},
title = {Pith review of: PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM},
year = {2026},
howpublished = {\url{https://pith.science/paper/CKRW7SXN}},
note = {Machine review of arXiv:2501.00352}
}
read the original abstract
Understanding geometric, semantic, and instance information in 3D scenes from sequential video data is essential for applications in robotics and augmented reality. However, existing Simultaneous Localization and Mapping (SLAM) methods generally focus on either geometric or semantic reconstruction. In this paper, we introduce PanoSLAM, the first SLAM system to integrate geometric reconstruction, 3D semantic segmentation, and 3D instance segmentation within a unified framework. Our approach builds upon 3D Gaussian Splatting, modified with several critical components to enable efficient rendering of depth, color, semantic, and instance information from arbitrary viewpoints. To achieve panoptic 3D scene reconstruction from sequential RGB-D videos, we propose an online Spatial-Temporal Lifting (STL) module that transfers 2D panoptic predictions from vision models into 3D Gaussian representations. This STL module addresses the challenges of label noise and inconsistencies in 2D predictions by refining the pseudo labels across multi-view inputs, creating a coherent 3D representation that enhances segmentation accuracy. Our experiments show that PanoSLAM outperforms recent semantic SLAM methods in both mapping and tracking accuracy. For the first time, it achieves panoptic 3D reconstruction of open-world environments directly from the RGB-D video. (https://github.com/runnanchen/PanoSLAM)
Figures
Forward citations
Cited by 1 Pith paper
-
PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views
PanoImager is an SfM-free pipeline combining feed-forward priors, geometry-conditioned diffusion view completion, and depth-guided 3DGS optimization to reconstruct from sparse panoramic images.
Reference graph
Works this paper leans on
-
[1]
Zero-shot semantic segmentation.Advances in Neural Information Processing Systems, 32, 2019
Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick P´erez. Zero-shot semantic segmentation.Advances in Neural Information Processing Systems, 32, 2019. 2
2019
-
[2]
Zero-shot point cloud segmentation by transferring geometric primitives
Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li, Yuexin Ma, Ruigang Yang, and Wenping Wang. Zero-shot point cloud segmentation by transferring geometric primitives. arXiv preprint arXiv:2210.09923, 2022. 2
arXiv 2022
-
[3]
Clip2scene: Towards label-efficient 3d scene understanding by clip
Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li, Yuenan Hou, Yu Qiao, and Wen- ping Wang. Clip2scene: Towards label-efficient 3d scene understanding by clip. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 7020–7030, 2023. 1, 2
2023
-
[4]
Bridging language and geometric primitives for zero-shot point cloud segmen- tation
Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li, Yuexin Ma, Ruigang Yang, and Wenping Wang. Bridging language and geometric primitives for zero-shot point cloud segmen- tation. In Proceedings of the 31st ACM International Con- ference on Multimedia, pages 5380–5388, 2023. 2
work page 2023
-
[5]
Runnan Chen, Xinge Zhu, Nenglun Chen, Dawei Wang, Wei Li, Yuexin Ma, Ruigang Yang, Tongliang Liu, and Wenping Wang. Model2scene: Learning 3d scene representation via contrastive language-cad models pre-training.arXiv preprint arXiv:2309.16956, 2023. 2
arXiv 2023
-
[6]
Towards label-free scene understanding by vision foundation models
Runnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen, Xinge Zhu, Yuexin Ma, Tongliang Liu, and Wenping Wang. Towards label-free scene understanding by vision foundation models. Advances in Neural Information Pro- cessing Systems, 36, 2024. 1, 2
work page 2024
-
[7]
Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation
Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12475–12485, 2020. 2
2020
-
[8]
Per- pixel classification is not all you need for semantic segmen- tation
Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per- pixel classification is not all you need for semantic segmen- tation. Advances in neural information processing systems , 34:17864–17875, 2021. 4, 5
work page 2021
Show all 73 references
-
[9]
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022
2022
-
[10]
(af)2-s3net: Attentive feature fusion with adaptive feature selection for sparse semantic segmentation network
Ran Cheng, Ryan Razani, Ehsan Taghavi, Enxu Li, and Bingbing Liu. (af)2-s3net: Attentive feature fusion with adaptive feature selection for sparse semantic segmentation network. In IEEE Conference on Computer Vision and Pat- tern Recognition, pages 12547–12556, 2021
2021
-
[11]
Mmdetection3d: Openmm- lab next-generation platform for general 3d object detection,
MMDetection3D Contributors. Mmdetection3d: Openmm- lab next-generation platform for general 3d object detection,
-
[12]
Monoslam: Real-time single camera slam
Andrew J Davison, Ian D Reid, Nicholas D Molton, and Olivier Stasse. Monoslam: Real-time single camera slam. IEEE transactions on pattern analysis and machine intelli- gence, 29(6):1052–1067, 2007. 2
2007
-
[13]
Language-driven open- vocabulary 3d scene understanding
Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Language-driven open- vocabulary 3d scene understanding. arXiv preprint arXiv:2211.16312, 2022. 2
2022 arXiv
-
[14]
Neural implicit dense semantic slam
Yasaman Haghighi, Suryansh Kumar, Jean-Philippe Thiran, and Luc Van Gool. Neural implicit dense semantic slam. arXiv preprint arXiv:2304.14560, 2023. 1, 2, 6
2023 arXiv
-
[15]
Unified 3d and 4d panoptic segmentation via dynamic shifting network
Fangzhou Hong, Lingdong Kong, Hui Zhou, Xinge Zhu, Hongsheng Li, and Ziwei Liu. Unified 3d and 4d panoptic segmentation via dynamic shifting network. arXiv preprint arXiv:2203.07186, 2022. 2
2022 arXiv
-
[16]
Uncertainty-aware learning for zero-shot semantic segmentation
Ping Hu, Stan Sclaroff, and Kate Saenko. Uncertainty-aware learning for zero-shot semantic segmentation. Advances in Neural Information Processing Systems , 33:21713–21724,
-
[17]
Eslam: Efficient dense slam system based on hybrid representation of signed distance fields
Mohammad Mahdi Johari, Camilla Carta, and Francois Fleuret. Eslam: Efficient dense slam system based on hybrid representation of signed distance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17408–17419, 2023. 2, 6
2023
-
[18]
Splatam: Splat, track & map 3d gaus- sians for dense rgb-d slam.arXiv preprint arXiv:2312.02126,
Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallab- hula, Gengshan Yang, Sebastian Scherer, Deva Ramanan, and Jonathon Luiten. Splatam: Splat, track & map 3d gaus- sians for dense rgb-d slam.arXiv preprint arXiv:2312.02126,
-
[19]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4):1–14, 2023. 2
2023
-
[20]
Lerf: Language embedded radiance fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embedded radiance fields. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 19729–19739,
-
[21]
Approximate differ- entiable rendering with algebraic surfaces
Leonid Keselman and Martial Hebert. Approximate differ- entiable rendering with algebraic surfaces. InEuropean Con- ference on Computer Vision, pages 596–614. Springer, 2022. 2
2022
-
[22]
Flexible techniques for differentiable rendering with 3d gaussians.arXiv preprint arXiv:2308.14737, 2023
Leonid Keselman and Martial Hebert. Flexible techniques for differentiable rendering with 3d gaussians.arXiv preprint arXiv:2308.14737, 2023. 2
2023 arXiv
-
[23]
Panoptic segmentation
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Doll ´ar. Panoptic segmentation. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9404–9413, 2019. 6
2019
-
[24]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 2023. 1, 2
2023
-
[25]
Parallel tracking and map- ping for small ar workspaces
Georg Klein and David Murray. Parallel tracking and map- ping for small ar workspaces. In 2007 6th IEEE and ACM international symposium on mixed and augmented reality , pages 225–234. IEEE, 2007. 2
2007
-
[26]
Rethinking range view representation for lidar segmentation
Lingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma, Xinge Zhu, Yikang Li, Yuenan Hou, Yu Qiao, and Ziwei Liu. Rethinking range view representation for lidar segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 228–240, 2023. 2
2023
-
[27]
Robo3d: Towards robust and reliable 3d perception against corruptions
Lingdong Kong, Youquan Liu, Xin Li, Runnan Chen, Wen- wei Zhang, Jiawei Ren, Liang Pan, Kai Chen, and Ziwei Liu. Robo3d: Towards robust and reliable 3d perception against corruptions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19994–20006...
2023
-
[28]
Language-driven semantic seg- mentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. In International Conference on Learning Rep- resentations, 2022. 2
2022
-
[29]
Dns slam: Dense neural semantic-informed slam
Kunyi Li, Michael Niemeyer, Nassir Navab, and Federico Tombari. Dns slam: Dense neural semantic-informed slam. arXiv preprint arXiv:2312.00204, 2023. 1, 2, 6
2023 arXiv
-
[30]
Sgs-slam: Se- mantic gaussian splatting for neural dense slam
Mingrui Li, Shuhong Liu, and Heng Zhou. Sgs-slam: Se- mantic gaussian splatting for neural dense slam. arXiv preprint arXiv:2402.03246, 2024. 1, 2
2024 arXiv
-
[31]
Consistent structural relation learning for zero-shot segmentation
Peike Li, Yunchao Wei, and Yi Yang. Consistent structural relation learning for zero-shot segmentation. Advances in Neural Information Processing Systems , 33:10317–10327,
-
[32]
Urban4d: Semantic-guided 4d gaussian splatting for urban scene reconstruction
Ziwen Li, Jiaxin Huang, Runnan Chen, Yunlong Che, Yandong Guo, Tongliang Liu, Fakhri Karray, and Ming- ming Gong. Urban4d: Semantic-guided 4d gaussian splatting for urban scene reconstruction. arXiv preprint arXiv:2412.03473, 2024. 2
2024 arXiv
-
[33]
Uniseg: A unified multi-modal li- dar segmentation network and the openpcseg codebase
Youquan Liu, Runnan Chen, Xin Li, Lingdong Kong, Yuchen Yang, Zhaoyang Xia, Yeqi Bai, Xinge Zhu, Yuexin Ma, Yikang Li, et al. Uniseg: A unified multi-modal li- dar segmentation network and the openpcseg codebase. In Proceedings of the IEEE/CVF International Conference on Compu...
2023
-
[34]
Multi-space alignments towards universal lidar segmentation
Youquan Liu, Lingdong Kong, Xiaoyang Wu, Runnan Chen, Xin Li, Liang Pan, Ziwei Liu, and Yuexin Ma. Multi-space alignments towards universal lidar segmentation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14648–14661, 2024. 2
2024
-
[35]
See more and know more: Zero-shot point cloud segmentation via multi-modal visual data
Yuhang Lu, Qi Jiang, Runnan Chen, Yuenan Hou, Xinge Zhu, and Yuexin Ma. See more and know more: Zero-shot point cloud segmentation via multi-modal visual data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 21674–21684, 2023. 2
2023
-
[36]
Gaussian splatting slam
Hidenobu Matsuki, Riku Murai, Paul HJ Kelly, and An- drew J Davison. Gaussian splatting slam. arXiv preprint arXiv:2312.06741, 2023. 2
2023 arXiv
-
[37]
Generative zero-shot learning for semantic segmentation of 3d point clouds
Bj ¨orn Michele, Alexandre Boulch, Gilles Puy, Maxime Bucher, and Renaud Marlet. Generative zero-shot learning for semantic segmentation of 3d point clouds. In Interna- tional Conference on 3D Vision, pages 992–1002, 2021. 2
2021
-
[38]
Orb-slam: a versatile and accurate monocular slam system
Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE transactions on robotics , 31(5):1147–1163,
-
[39]
Openscene: 3d scene understanding with open vocabularies
Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. Openscene: 3d scene understanding with open vocabularies. arXiv preprint arXiv:2211.15654, 2022. 1
2022 arXiv
-
[40]
Openscene: 3d scene understanding with open vocabularies
Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 815–824, 2023. 2
2023
-
[41]
Learning to adapt sam for segmenting cross-domain point clouds
Xidong Peng, Runnan Chen, Feng Qiao, Lingdong Kong, Youquan Liu, Yujing Sun, Tai Wang, Xinge Zhu, and Yuexin Ma. Learning to adapt sam for segmenting cross-domain point clouds. In European Conference on Computer Vision, pages 54–71. Springer, 2025. 2
2025
-
[42]
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 652–660, 2017. 2
2017
-
[43]
Langsplat: 3d language gaussian splatting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. Langsplat: 3d language gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20051–20060, 2024. 1
2024
-
[44]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[45]
Novel class discovery for 3d point cloud semantic segmenta- tion
Luigi Riz, Cristiano Saltori, Elisa Ricci, and Fabio Poiesi. Novel class discovery for 3d point cloud semantic segmenta- tion. arXiv preprint arXiv:2303.11610, 2023. 2
2023 arXiv
-
[46]
Kimera: an open-source library for real-time metric- semantic localization and mapping
Antoni Rosinol, Marcus Abate, Yun Chang, and Luca Car- lone. Kimera: an open-source library for real-time metric- semantic localization and mapping. In IEEE International Conference on Robotics and Automation, pages 1689–1696. IEEE, 2020. 2
2020
-
[47]
Slam++: Si- multaneous localisation and mapping at the level of objects
Renato F Salas-Moreno, Richard A Newcombe, Hauke Strasdat, Paul HJ Kelly, and Andrew J Davison. Slam++: Si- multaneous localisation and mapping at the level of objects. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1352–1359, 2013. 2
2013
-
[48]
Point-slam: Dense neural point cloud-based slam
Erik Sandstr ¨om, Yue Li, Luc Van Gool, and Martin R Os- wald. Point-slam: Dense neural point cloud-based slam. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18433–18444, 2023. 2, 6
2023
-
[49]
Image-to-lidar self-supervised distillation for autonomous driving data
Corentin Sautier, Gilles Puy, Spyros Gidaris, Alexandre Boulch, Andrei Bursuc, and Renaud Marlet. Image-to-lidar self-supervised distillation for autonomous driving data. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 9891–9901, 2022. 2
2022
-
[50]
Language embedded 3d gaussians for open-vocabulary scene understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao- Hua Guan. Language embedded 3d gaussians for open-vocabulary scene understanding. arXiv preprint arXiv:2311.18482, 2023. 1
2023 arXiv
-
[51]
Panoptic lifting for 3d scene understanding with neural fields
Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bul ´o, Nor- man M ¨uller, Matthias Nießner, Angela Dai, and Peter Kontschieder. Panoptic lifting for 3d scene understanding with neural fields. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, p...
2023
-
[52]
The replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797,
1906 arXiv
-
[53]
Segmenter: Transformer for semantic segmenta- tion
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmenta- tion. In Proceedings of the IEEE/CVF international confer- ence on computer vision, pages 7262–7272, 2021. 2
2021
-
[54]
imap: Implicit mapping and positioning in real-time
Edgar Sucar, Shikun Liu, Joseph Ortiz, and Andrew J Davi- son. imap: Implicit mapping and positioning in real-time. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6229–6238, 2021. 2, 6
2021
-
[55]
An empirical study of training state-of-the-art lidar segmentation models
Jiahao Sun, Chunmei Qing, Xiang Xu, Lingdong Kong, Youquan Liu, Li Li, Chenming Zhu, Jingwei Zhang, Zeqi Xiao, Runnan Chen, et al. An empirical study of training state-of-the-art lidar segmentation models. arXiv preprint arXiv:2405.14870, 2024. 2
2024 arXiv
-
[56]
V oge: a differentiable volume renderer using gaussian ellipsoids for analysis-by-synthesis
Angtian Wang, Peng Wang, Jian Sun, Adam Kortylewski, and Alan Yuille. V oge: a differentiable volume renderer using gaussian ellipsoids for analysis-by-synthesis. arXiv preprint arXiv:2205.15401, 2022. 2
2022 arXiv
-
[57]
Co- slam: Joint coordinate and sparse parametric encodings for neural real-time slam
Hengyi Wang, Jingwen Wang, and Lourdes Agapito. Co- slam: Joint coordinate and sparse parametric encodings for neural real-time slam. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13293–13302, 2023. 2, 6
2023
-
[58]
Point transformer v2: Grouped vec- tor attention and partition-based pooling
Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao. Point transformer v2: Grouped vec- tor attention and partition-based pooling. arXiv preprint arXiv:2210.05666, 2022. 2
2022 arXiv
-
[59]
Rpvnet: A deep and efficient range-point- voxel fusion network for lidar point cloud segmentation
Jianyun Xu, Ruixiang Zhang, Jian Dou, Yushi Zhu, Jie Sun, and Shiliang Pu. Rpvnet: A deep and efficient range-point- voxel fusion network for lidar point cloud segmentation. In IEEE/CVF International Conference on Computer Vision , pages 16024–16033, 2021. 2
2021
-
[60]
A simple baseline for zero- shot semantic segmentation with pre-trained vision-language model
Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for zero- shot semantic segmentation with pre-trained vision-language model. arXiv preprint arXiv:2112.14757, 2021. 2
2021 arXiv
-
[61]
Human-centric scene understanding for 3d large-scale scenarios
Yiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen, Yue- nan Hou, Xinge Zhu, Xuming He, Jingyi Yu, and Yuexin Ma. Human-centric scene understanding for 3d large-scale scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20349–20359, 2023. 2
2023
-
[62]
Gs-slam: Dense visual slam with 3d gaussian splatting.arXiv preprint arXiv:2311.11700,
Chi Yan, Delin Qu, Dong Wang, Dan Xu, Zhigang Wang, Bin Zhao, and Xuelong Li. Gs-slam: Dense visual slam with 3d gaussian splatting.arXiv preprint arXiv:2311.11700,
-
[63]
2dpass: 2d pri- ors assisted semantic segmentation on lidar point clouds
Xu Yan, Jiantao Gao, Chaoda Zheng, Chaoda Zheng, Ruimao Zhang, Shenghui Cui, and Zhen Li. 2dpass: 2d pri- ors assisted semantic segmentation on lidar point clouds. In ECCV, 2022. 2
2022
-
[64]
Scannet++: A high-fidelity dataset of 3d in- door scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d in- door scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023. 2, 6
2023
-
[65]
Is-fusion: Instance-scene collaborative fusion for multimodal 3d ob- ject detection
Junbo Yin, Jianbing Shen, Runnan Chen, Wei Li, Ruigang Yang, Pascal Frossard, and Wenguan Wang. Is-fusion: Instance-scene collaborative fusion for multimodal 3d ob- ject detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1490...
2024
-
[66]
Gaussian-slam: Photo-realistic dense slam with gaus- sian splatting
Vladimir Yugay, Yue Li, Theo Gevers, and Martin R Os- wald. Gaussian-slam: Photo-realistic dense slam with gaus- sian splatting. arXiv preprint arXiv:2312.10070, 2023. 2
2023 arXiv
-
[67]
Prototypical matching and open set rejection for zero-shot semantic segmentation
Hui Zhang and Henghui Ding. Prototypical matching and open set rejection for zero-shot semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6974–6983, 2021. 2
2021
-
[68]
Sni-slam: Semantic neural implicit slam
Siting Zhu, Guangming Wang, Hermann Blum, Jiuming Liu, Liang Song, Marc Pollefeys, and Hesheng Wang. Sni-slam: Semantic neural implicit slam. arXiv preprint arXiv:2311.11016, 2023. 1, 2, 6
2023 arXiv
-
[69]
Semgauss-slam: Dense semantic gaussian splatting slam
Siting Zhu, Renjie Qin, Guangming Wang, Jiuming Liu, and Hesheng Wang. Semgauss-slam: Dense semantic gaussian splatting slam. arXiv preprint arXiv:2403.07494, 2024. 1, 2
2024 arXiv
-
[70]
Cylindrical and asymmetrical 3d convolution networks for lidar segmenta- tion
Xinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong, Yuexin Ma, Wei Li, Hongsheng Li, and Dahua Lin. Cylindrical and asymmetrical 3d convolution networks for lidar segmenta- tion. arXiv preprint arXiv:2011.10033, 2020. 2
2011 arXiv
-
[71]
Cylindrical and asymmetrical 3d convolution networks for lidar segmenta- tion
Xinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong, Yuexin Ma, Wei Li, Hongsheng Li, and Dahua Lin. Cylindrical and asymmetrical 3d convolution networks for lidar segmenta- tion. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 9939–9948, 2021. 2
2021
-
[72]
Nice-slam: Neural implicit scalable encoding for slam
Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hu- jun Bao, Zhaopeng Cui, Martin R Oswald, and Marc Polle- feys. Nice-slam: Neural implicit scalable encoding for slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12786–12796,...
2022
-
[73]
Segment everything everywhere all at once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. Advances in Neural Information Processing Systems, 36, 2024. 7
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.