REVIEW 4 major objections 5 minor 124 references
CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read A single Gaussian-rendering interface transfers 2D masks into persistent 3D labels without training a 3D segmentation network.
desk verdict Useful zero-training label-transfer interface, but the 'where' half rests on an underspecified visibility rule, and a plain projection baseline beats CDSeg on one benchmark. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the renderable Gaussian carrier: each primitive $g_i = (\mu_i, \Sigma_i, \alpha_i, \mathrm{sh}_i)$ can hold a label and be rendered. The identity that carries the argument is the binary pixel--primitive association $A_{i,v,p} = \mathbf{1}[i \in R_{v,p}] \, \mathbf{1}[\nu_{i,v,p} \geq \tau]$ with $\tau = 0.1$, recorded inside the splatting pass. It respects projected support, depth order, and accumulated transmittance, so it replaces any learned correspondence network. The association feeds a vote accumulator $S_i(l)$ over views and pixels and a $k$-nearest-neighbor majority filter that outputs the final label.
What would settle it
Render a synthetic scene with a thin object in front of a background plane, run CDSeg from one calibrated view with a mask of the thin object, and check whether any background primitive receives the object label; if it does, the transmittance threshold is mis-assigning surface ownership.
Extended reading notes
Core claim
CDSeg claims that rendering-based visibility, not learned correspondence, is sufficient to transfer labels from 2D masks to 3D. During each view's Gaussian splatting pass, it records for every pixel the primitives whose transmittance-based visibility score is at least 0.1; those primitives receive votes from that pixel's mask label. Votes are summed over views and labels, the strongest label wins, and a k-nearest-neighbor majority filter removes isolated disagreements. The same procedure operates on point-completed Gaussians (Mode I) and on optimized Gaussian scenes (Mode II), so the output is either labels in the original point order or labels retained on renderable primitives that can be mapped back to images.
Load-bearing premise
The method assumes that the renderer's transmittance threshold ($\nu \geq 0.1$) identifies the surface that truly owns each pixel's label, so at object boundaries, occlusions, and thin structures this threshold, not the mask source, determines where labels attach.
Editorial extensions
If this is right
- Any calibrated set of views plus a mask source can label an existing point cloud without training a 3D network, as long as the views cover the target surfaces.
- Labels fused on a Gaussian scene remain renderable, so a single 3D labeling can be queried from arbitrary cameras and produces consistent instance identities across viewpoints.
- Because masks are plug-in inputs, the same association and voting code is reused for promptable, automatic-instance, semantic, and LiDAR transfer tasks.
- Multi-view voting suppresses isolated mask errors and keeps accuracy near 93% mIoU down to about ten well-placed views, degrading only when coverage is lost.
- Runtime and memory grow approximately linearly with primitive and view counts, reaching a few seconds for seven million primitives, so the interface scales to large scenes.
Reading between the lines
- Because the association step is attribute-agnostic, the same carrier could plausibly fuse depth, occlusion order, affordances, or language embeddings rather than only discrete labels; the paper does not make this claim.
- A practical extension the paper leaves implicit is that replacing the 2D mask source with a temporally consistent video segmenter should improve boundary accuracy more than any change to the 3D voting rule.
- For dynamic scenes, the per-window static-carrier assumption suggests a natural extension: rebuild the point-completed carrier for each static window and track instance labels across windows, something the paper does not test.
- A stress test worth running is varying the visibility threshold $\tau$ at object boundaries on thin structures; if no threshold yields clean boundaries, the transmittance rule itself, rather than the masks, is the bottleneck.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CDSeg, a training-free label-transfer interface that uses Gaussian primitives as a 3D label carrier. An external 2D mask source supplies semantic or instance labels, and renderer-derived visibility during Gaussian splatting determines which primitives receive each pixel's label; multi-view voting and a local k-NN filter then fuse the labels onto points (Mode I) or native Gaussian scenes (Mode II). The method is evaluated on DesktopObjects-360, NeRDS-360, ScanNet-v2, and KITTI-360, reporting competitive mIoU against supervised 3D segmentation baselines and showing linear scaling to millions of primitives.
Significance. If the central visibility-association claim is validated, CDSeg would provide a simple, unified interface for reusing 2D masks across point clouds, Gaussian scenes, and arbitrary views without task-specific 3D training. The paper has several strengths: the method has no learned parameters and no optimization loop; it preserves exact point indices in Mode I; the association is performed inside the rendering pass; and the experiments cover four different task settings with ablations on view coverage and scaling. The code archive is mentioned as providing exact timing calls, which supports reproducibility. However, the core 'where' claim rests on an underspecified and insufficiently validated visibility rule, and the ScanNet evaluation appears to have a circularity risk because the provided 2D annotations may be derived from the same 3D ground truth used for scoring.
major comments (4)
- [§3.2, Eqs. (3)–(4)]
- [§4.5 and Supplement §E]
- [§3.3, Eq. (6)–(7) vs Supplement §C]
- [Table 7, §4.5]
minor comments (5)
- [§4.2, Table 4]
- [§3.1, Eq. (2)]
- [§4.5, KITTI-360]
- [§4.3–4.4, Tables 5–6]
- [Supplement §H] The text says 'Theresulstthereforedemonstrate3D-consistentinstanceidentities' — there is a typo ('Resulst').
Circularity Check
No derivational circularity: CDSeg's 3D labels are a fixed function of external masks plus renderer visibility, with no fitted parameter and with demonstrably lossy transfer (ScanNet 65.77% mIoU, not the 100% a tautological loop would force).
-
other
[Sec. 3.1 (Mode II); Sec. 4.1 (Experimental Protocol); Supp. B (Choice of Gaussian carrier)]
"We use the optimized scenes released with PointGauss Sun et al. (2025) for DesktopObjects-360."
Four of the five CDSeg authors (Sun, Chen, Li, Zelek) are also authors of PointGauss, which introduced the DesktopObjects-360 benchmark and released the optimized Gaussian scenes that CDSeg adopts as its Mode II carrier. The flagship 'native Gaussian scene' evaluation is therefore built on the authors' own same-group dataset, carrier, and prior work rather than on external infrastructure, so that one of four benchmark results provides no independent evidence for the central transfer claim.
full rationale
The derivation chain in Eqs. (2)-(7) is self-contained and not circular. Input labels M_v come from an external mask source (SAM2, YOLOv11, or dataset-provided 2D annotations); the association A_{i,v,p} in Eq. (4) is a fixed thresholding of renderer visibility built on the standard alpha-compositing terms of Eq. (3); voting in Eqs. (5)-(6) is discrete majority; filtering in Eq. (7) is a fixed k=3 KNN rule. No parameter is fitted to any evaluation label, and the output is not equal to the input by construction: on ScanNet, where the input 2D masks and the 3D ground truth come from the same dataset annotations, the pipeline recovers only 65.77% mIoU rather than the 100% a tautological loop would force, so the transfer demonstrably loses information through occlusion and the coarse mask boundaries the paper itself admits in Supp. E ('Boundaries in the image annotations and reconstructed point geometry are not perfectly aligned, which creates contradictory labels near object contours'). The ScanNet protocol has an evaluation-loop flavor - the provided 2D annotations serve as controlled masks while the same dataset's 3D labels serve as ground truth - but the paper states this explicitly as an isolation of the correspondence stage, which is a legitimate if favorable protocol rather than a derivational circularity. Eq. (4)'s visibility score nu is left undefined and tau=0.1 is never ablated; this is an under-specification and correctness risk in the 'where' half of the claim, not a circular step. The one genuine coupling is same-group: DesktopObjects-360 and its optimized scenes come from PointGauss with heavy author overlap, which reduces the external-evidence weight of one flagship Mode II result but is not load-bearing because the central mechanism is also demonstrated on external benchmarks. Score 2 matches the rubric's 'one minor self-citation that is not load-bearing.'
Assumptions & free parameters
free parameters (3)
- visibility threshold tau =
0.1
- neighborhood filter size k =
3
- renderer stopping threshold tau_stop =
0.1
assumptions (5)
- domain assumption 3DGS alpha-compositing and accumulated transmittance model the physical visibility of surfaces.
- domain assumption The Gaussian carrier corresponds to the geometry being labeled.
- domain assumption ScanNet's provided 2D semantic annotations are independent valid masks for the same geometry as the withheld 3D labels.
- domain assumption SAM2 and YOLO masks provide correct cross-view instance identity after pseudo-video tracking.
- domain assumption The selected KITTI-360 static windows and stationary vehicles are representative of LiDAR label transfer.
Cite this review
Pith. "Pith review of CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer." pith.science (2026). https://pith.science/paper/VWJKNXK7
@misc{pith2026260805482,
author = {Pith},
title = {Pith review of: CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer},
year = {2026},
howpublished = {\url{https://pith.science/paper/VWJKNXK7}},
note = {Machine review of arXiv:2608.05482}
}
read the original abstract
Modern image models provide strong cues about \emph{what} should be segmented in each view, but their masks do not by themselves determine \emph{where} those labels should persist in 3D. We present Cross-Domain Segmentation via Gaussian Splatting (CDSeg), a label-transfer interface that requires no task-specific 3D segmentation training and uses Gaussian primitives as a renderable label carrier. An external mask source supplies the labels, while renderer-derived visibility determines which 3D primitives receive them. The carrier is instantiated either by completing each input point into one Gaussian, preserving its index, or by reusing the native primitives of an optimized Gaussian scene. CDSeg records pixel--primitive associations during rendering and fuses multi-view masks through voting and a local filter. The resulting labels can be returned to the original points, retained on the native Gaussian scene, or rendered into other views. CDSeg covers promptable, automatic instance, semantic, and LiDAR settings and processes scenes with millions of primitives in seconds. It obtains 92.35\% mIoU on DesktopObjects-360, 95.89\% on NeRDS-360, and 65.77\% on the full ScanNet-v2 validation split using the provided 2D semantic annotations. CDSeg thereby provides one interface for reusing 2D masks across point clouds, Gaussian scenes, and image views without a task-specific 3D segmentation network.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Pointdc: Unsupervised semantic segmentation of 3D point clouds via cross-modal distillation and super-voxel clustering
Zisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou, Haihong Xiao, Baigui Sun, Xuansong Xie, and Wenxiong Kang. Pointdc: Unsupervised semantic segmentation of 3D point clouds via cross-modal distillation and super-voxel clustering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 14290--14299, October 2023
2023
-
[2]
Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nie ner
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nie ner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2432--2443, 2017. doi:10.1109/CVPR.2017.261
-
[3]
Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans
Angela Dai, Daniel Ritchie, Martin Bokeloh, Scott Reed, Jürgen Sturm, and Matthias Nießner. Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578--4587, 2018. doi:10.1109/CVPR.2018.00481
arXiv 2018
-
[4]
Deep learning for 3d point clouds: A survey
Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu, Li Liu, and Mohammed Bennamoun. Deep learning for 3d point clouds: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43 0 (12): 0 4338--4364, 2021. doi:10.1109/TPAMI.2020.3005434
arXiv 2021
-
[5]
Randla-net: Efficient semantic segmentation of large-scale point clouds
Qingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang, Niki Trigoni, and Andrew Markham. Randla-net: Efficient semantic segmentation of large-scale point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11108--11117, 2020
2020
-
[6]
2D gaussian splatting for geometrically accurate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D gaussian splatting for geometrically accurate radiance fields. In ACM SIGGRAPH 2024 Conference Papers. Association for Computing Machinery, 2024. doi:10.1145/3641519.3657428
arXiv 2024
-
[7]
Neo 360: Neural fields for sparse view synthesis of outdoor scenes
Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Guizilini, Thomas Kollar, Adrien Gaidon, Zsolt Kira, and Rares Ambrus. Neo 360: Neural fields for sparse view synthesis of outdoor scenes. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9187--9198, 2023. doi:10.1109/ICCV51070.2023.00843
arXiv 2023
-
[8]
Ultralytics yolo11, 2024
Glenn Jocher and Jing Qiu. Ultralytics yolo11, 2024. URL https://github.com/ultralytics/ultralytics
2024
Show all 124 references
-
[10]
Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision ...
2023
-
[11]
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Yiyi Liao, Jun Xie, and Andreas Geiger. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (3): 0 3292--3310, 2023. doi:10.1109/TPAMI.2022.3179507
2023
-
[12]
Meta architecture for point cloud analysis
Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao, Shanshan Wang, Yan Wang, Yonghong Tian, and Rongrong Ji. Meta architecture for point cloud analysis. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17682--17691, 2023. doi:10.1109/CVPR52729.2023.01696
2023
-
[13]
One thing one click: A self-training approach for weakly supervised 3d semantic segmentation
Zhengzhe Liu, Xiaojuan Qi, and Chi-Wing Fu. One thing one click: A self-training approach for weakly supervised 3d semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1726--1736, 2021
2021
-
[14]
LUDVIG : Learning-free uplifting of 2D visual features to gaussian splatting scenes
Juliette Marrie, Romain Menegaux, Michael Arbel, Diane Larlus, and Julien Mairal. LUDVIG : Learning-free uplifting of 2D visual features to gaussian splatting scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7440--7450, 2025
2025
-
[15]
DINOv2 : Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, ...
2024
-
[17]
Qi, Li Yi, Hao Su, and Leonidas J
Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. PointNet++ : Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, volume 30, pages 5105--5114, 2017
2017
-
[18]
Pointnext: Revisiting pointnet++ with improved training and scaling strategies
Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in N...
2022
-
[19]
Langsplat: 3d language gaussian splatting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. Langsplat: 3d language gaussian splatting. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20051--20060, 2024
2024
-
[20]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and ...
2021
-
[21]
SAM 2 : Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenh...
2025
-
[22]
FlashSplat : 2D to 3D gaussian splatting segmentation solved optimally
Qiuhong Shen, Xingyi Yang, and Xinchao Wang. FlashSplat : 2D to 3D gaussian splatting segmentation solved optimally. In European Conference on Computer Vision, pages 456--472. Springer, 2024. doi:10.1007/978-3-031-72670-5_26
2024 doi
-
[23]
Language embedded 3d gaussians for open-vocabulary scene understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao-Hua Guan. Language embedded 3d gaussians for open-vocabulary scene understanding. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5333--5343, 2024
2024
-
[24]
PointGS : Semantic-consistent unsupervised 3D point cloud segmentation with 3D gaussian splatting
Yixiao Song, Qingyong Li, Wen Wang, and Zhicheng Yan. PointGS : Semantic-consistent unsupervised 3D point cloud segmentation with 3D gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 33343--33352, 2026
2026
-
[25]
Zelek, and Jonathan Li
Wentao Sun, Hanqing Xu, Quanyun Wu, Dedong Zhang, Yiping Chen, Lingfei Ma, John S. Zelek, and Jonathan Li. PointGauss : Point cloud-guided multi-object segmentation for gaussian splatting, 2025. URL https://arxiv.org/abs/2508.00259
2025 arXiv
-
[26]
Contrastive boundary learning for point cloud segmentation
Liyao Tang, Yibing Zhan, Zhe Chen, Baosheng Yu, and Dacheng Tao. Contrastive boundary learning for point cloud segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8489--8499, 2022
2022
-
[27]
Point transformer v2: Grouped vector attention and partition-based pooling
Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao. Point transformer v2: Grouped vector attention and partition-based pooling. In Advances in Neural Information Processing Systems, volume 35, pages 33330--33342, 2022
2022
-
[28]
Point transformer v3: Simpler faster stronger
Xiaoyang Wu, Li Jiang, Peng-Shuai Wang, Zhijian Liu, Xihui Liu, Yu Qiao, Wanli Ouyang, Tong He, and Hengshuang Zhao. Point transformer v3: Simpler faster stronger. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4840--4851, 2024
2024
-
[29]
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany. Pointcontrast: Unsupervised pre-training for 3d point cloud understanding. In European conference on computer vision, pages 574--591. Springer, 2020
2020
-
[30]
3d weakly supervised semantic segmentation with 2d vision-language guidance
Xiaoxu Xu, Yitian Yuan, Jinlong Li, Qiudan Zhang, Zequn Jie, Lin Ma, Hao Tang, Nicu Sebe, and Xu Wang. 3d weakly supervised semantic segmentation with 2d vision-language guidance. In European Conference on Computer Vision, pages 87--104. Springer, 2024
2024
-
[31]
Gaussian G rouping: Segment and edit anything in 3D scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu, and Lei Ke. Gaussian G rouping: Segment and edit anything in 3D scenes. In Proceedings of the European Conference on Computer Vision (ECCV), pages 162--179, 2024
2024
-
[32]
Feature 3DGS : Supercharging 3D gaussian splatting to enable distilled feature fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3DGS : Supercharging 3D gaussian splatting to enable distilled feature fields. In Proceedings of the IEEE/CVF Conference on Compute...
2024
-
[33]
FirstName LastName , title =
-
[34]
FirstName Alpher , title =
-
[35]
Journal of Foo , volume = 13, number = 1, pages =
FirstName Alpher and FirstName Fotheringham-Smythe , title =. Journal of Foo , volume = 13, number = 1, pages =
-
[36]
Journal of Foo , volume = 14, number = 1, pages =
FirstName Alpher and FirstName Fotheringham-Smythe and FirstName Gamow , title =. Journal of Foo , volume = 14, number = 1, pages =
-
[37]
FirstName Alpher and FirstName Gamow , title =
-
[38]
KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D , year=
Liao, Yiyi and Xie, Jun and Geiger, Andreas , journal=. KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D , year=
-
[39]
OneFormer: One Transformer to Rule Universal Image Segmentation , year=
Jain, Jitesh and Li, Jiachen and Chiu, MangTik and Hassani, Ali and Orlov, Nikita and Shi, Humphrey , booktitle=. OneFormer: One Transformer to Rule Universal Image Segmentation , year=
-
[40]
2024 , url =
Glenn Jocher and Jing Qiu , title =. 2024 , url =
2024
-
[41]
arXiv preprint arXiv:2402.13616 , year=
YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information , author=. arXiv preprint arXiv:2402.13616 , year=
-
[42]
arXiv preprint arXiv:2207.02696 , year=
YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors , author=. arXiv preprint arXiv:2207.02696 , year=
-
[43]
2023 , eprint=
YOLOv6 v3.0: A Full-Scale Reloading , author=. 2023 , eprint=
2023
-
[44]
Deep Learning for 3D Point Clouds: A Survey , year=
Guo, Yulan and Wang, Hanyun and Hu, Qingyong and Liu, Hao and Liu, Li and Bennamoun, Mohammed , journal=. Deep Learning for 3D Point Clouds: A Survey , year=
-
[45]
Foundation Models Defining a New Era in Vision: A Survey and Outlook , year=
Awais, Muhammad and Naseer, Muzammal and Khan, Salman and Anwer, Rao Muhammad and Cholakkal, Hisham and Shah, Mubarak and Yang, Ming-Hsuan and Khan, Fahad Shahbaz , journal=. Foundation Models Defining a New Era in Vision: A Survey and Outlook , year=
-
[46]
2024 , eprint=
Image Segmentation in Foundation Model Era: A Survey , author=. 2024 , eprint=
2024
-
[47]
and Saux, B
Boulch, A. and Saux, B. Le and Audebert, N. , title =. 2017 , publisher =. doi:10.2312/3dor.20171047 , booktitle =
2017 doi
-
[48]
Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks
Audebert, Nicolas and Le Saux, Bertrand and Lef \`e vre, S \'e bastien. Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks. Computer Vision -- ACCV 2016. 2017
2016
-
[49]
Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Tangent convolutions for dense prediction in 3d , author=. Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
-
[50]
2018 IEEE international conference on robotics and automation (ICRA) , pages=
Squeezeseg: Convolutional neural nets with recurrent crf for real-time road-object segmentation from 3d lidar point cloud , author=. 2018 IEEE international conference on robotics and automation (ICRA) , pages=. 2018 , organization=
2018
-
[51]
2019 , volume=
Milioto, Andres and Vizzo, Ignacio and Behley, Jens and Stachniss, Cyrill , booktitle=. 2019 , volume=
2019
-
[52]
Learning Multi-View Aggregation in the Wild for Large-Scale
Robert, Damien and Vallet, Bruno and Landrieu, Loic , booktitle=. Learning Multi-View Aggregation in the Wild for Large-Scale
-
[53]
Exploiting Multi-Layer Grid Maps for Surround-View Semantic Segmentation of Sparse
Bieder, Frank and Wirges, Sascha and Janosovits, Johannes and Richter, Sven and Wang, Zheyuan and Stiller, Christoph , booktitle=. Exploiting Multi-Layer Grid Maps for Surround-View Semantic Segmentation of Sparse. 2020 , organization=
2020
-
[54]
ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans , year=
Dai, Angela and Ritchie, Daniel and Bokeloh, Martin and Reed, Scott and Sturm, Jürgen and Nießner, Matthias , booktitle=. ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans , year=
-
[55]
PointGrid: A Deep Network for 3D Shape Understanding , year=
Le, Truc and Duan, Ye , booktitle=. PointGrid: A Deep Network for 3D Shape Understanding , year=
-
[56]
PCSCNet: Fast 3D semantic segmentation of LiDAR point cloud for autonomous car using point convolution and sparse convolution network , journal =
Jaehyun Park and Chansoo Kim and Soyeong Kim and Kichun Jo , keywords =. PCSCNet: Fast 3D semantic segmentation of LiDAR point cloud for autonomous car using point convolution and sparse convolution network , journal =. 2023 , issn =. doi:10.1016/j.eswa.2022.118815 , url =
2023
-
[57]
SpSequenceNet: Semantic Segmentation Network on 4D Point Clouds , year=
Shi, Hanyu and Lin, Guosheng and Wang, Hao and Hung, Tzu-Yi and Wang, Zhenhua , booktitle=. SpSequenceNet: Semantic Segmentation Network on 4D Point Clouds , year=
-
[58]
SIEV-Net: A Structure-Information Enhanced Voxel Network for 3D Object Detection From LiDAR Point Clouds , year=
Yu, Chuanbo and Lei, Jianjun and Peng, Bo and Shen, Haifeng and Huang, Qingming , journal=. SIEV-Net: A Structure-Information Enhanced Voxel Network for 3D Object Detection From LiDAR Point Clouds , year=
-
[59]
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , month =
Chen, Zisheng and Xu, Hongbin and Chen, Weitao and Zhou, Zhipeng and Xiao, Haihong and Sun, Baigui and Xie, Xuansong and Kang, Wenxiong , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , month =. 2023 , pages =
2023
-
[60]
and Yi, Li and Su, Hao and Guibas, Leonidas J
Qi, Charles R. and Yi, Li and Su, Hao and Guibas, Leonidas J. , title =. Advances in Neural Information Processing Systems , volume =
-
[61]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Point Transformer V3: Simpler Faster Stronger , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[62]
IEEE Transactions on Geoscience and Remote Sensing , volume=
A click-based interactive segmentation network for point clouds , author=. IEEE Transactions on Geoscience and Remote Sensing , volume=
-
[63]
Advances in Neural Information Processing Systems , volume =
Wu, Xiaoyang and Lao, Yixing and Jiang, Li and Liu, Xihui and Zhao, Hengshuang , title =. Advances in Neural Information Processing Systems , volume =
-
[64]
3DGTN: 3-D Dual-Attention GLocal Transformer Network for Point Cloud Classification and Segmentation , year=
Lu, Dening and Gao, Kyle and Xie, Qian and Xu, Linlin and Li, Jonathan , journal=. 3DGTN: 3-D Dual-Attention GLocal Transformer Network for Point Cloud Classification and Segmentation , year=
-
[65]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages=
ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages=. 2017 , doi=
2017
-
[66]
The Thirteenth International Conference on Learning Representations , year =
Nikhila Ravi and Valentin Gabeur and Yuan-Ting Hu and Ronghang Hu and Chaitanya Ryali and Tengyu Ma and Haitham Khedr and Roman Rädle and Chloe Rolland and Laura Gustafson and Eric Mintun and Junting Pan and Kalyan Vasudev Alwala and Nicolas Carion and Chao-Yuan Wu and Ross Gi...
-
[67]
and Lo, Wan-Yen and Dollar, Piotr and Girshick, Ross , title =
Kirillov, Alexander and Mintun, Eric and Ravi, Nikhila and Mao, Hanzi and Rolland, Chloe and Gustafson, Laura and Xiao, Tete and Whitehead, Spencer and Berg, Alexander C. and Lo, Wan-Yen and Dollar, Piotr and Girshick, Ross , title =. Proceedings of the IEEE/CVF International ...
2023
-
[68]
Change Detection between Optical Remote Sensing Imagery and Map Data Via Segment Anything Model (SAM) , year=
Chen, Hongruixuan and Song, Jian and Yokoya, Naoto , booktitle=. Change Detection between Optical Remote Sensing Imagery and Map Data Via Segment Anything Model (SAM) , year=
-
[69]
and Hiremath, Manjunatha , booktitle=
Patil, Ashwini P. and Hiremath, Manjunatha , booktitle=. Segment Anything Model (SAM) to Segment lymphocyte from Blood Smear Images , year=
-
[70]
SAM-I2I: Unleash the Power of Segment Anything Model for Medical Image Translation , year=
Huo, Jiayu and Ourselin, Sébastien and Sparks, Rachel , booktitle=. SAM-I2I: Unleash the Power of Segment Anything Model for Medical Image Translation , year=
-
[71]
Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection , year=
Hu, Shenglong and Han, Zhidong and Dong, Gang and Liang, Lingyan and Wen, Dongchao and Zhang, Kaihua , booktitle=. Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change Detection , year=
-
[72]
2024 , issn=
Maxime Oquab and Timothée Darcet and Théo Moutakanni and Huy Vo and Marc Szafraniec and Vasil Khalidov and Pierre Fernandez and Daniel Haziza and Francisco Massa and Alaaeldin El-Nouby and Mido Assran and Nicolas Ballas and Wojciech Galuba and Russell Howes and Po-Yao Huang an...
2024
-
[73]
BrainCrossFed CNN Model for Alzheimer Classification using MRI data and Comparison and Benchmarking proposed model with DINOv2 and ExplainableAI using GradCAM , year=
Rashmi, Uppin and Beena, B M and Ambesange, Sateesh , booktitle=. BrainCrossFed CNN Model for Alzheimer Classification using MRI data and Comparison and Benchmarking proposed model with DINOv2 and ExplainableAI using GradCAM , year=
-
[74]
DINOv2 Based Self Supervised Learning for Few Shot Medical Image Segmentation , year=
Ayzenberg, Lev and Giryes, Raja and Greenspan, Hayit , booktitle=. DINOv2 Based Self Supervised Learning for Few Shot Medical Image Segmentation , year=
-
[75]
DINOv2-Based UAV Visual Self-Localization in Low-Altitude Urban Environments , year=
Yang, Jiaqiang and Qin, Danyang and Tang, Huapeng and Tao, Sili and Bie, Haoze and Ma, Lin , journal=. DINOv2-Based UAV Visual Self-Localization in Low-Altitude Urban Environments , year=
-
[76]
SegGPT: Towards Segmenting Everything In Context , year=
Wang, Xinlong and Zhang, Xiaosong and Cao, Yue and Wang, Wen and Shen, Chunhua and Huang, Tiejun , booktitle=. SegGPT: Towards Segmenting Everything In Context , year=
-
[77]
Proceedings of the 38th International Conference on Machine Learning , pages=
Learning Transferable Visual Models From Natural Language Supervision , author=. Proceedings of the 38th International Conference on Machine Learning , pages=. 2021 , editor=
2021
-
[78]
ACM Transactions on Graphics , month = dec, articleno =
Yang, Chen and Li, Sikuang and Fang, Jiemin and Liang, Ruofan and Xie, Lingxi and Zhang, Xiaopeng and Shen, Wei and Tian, Qi , title =. ACM Transactions on Graphics , month = dec, articleno =. 2024 , volume =
2024
-
[79]
MVSplat : Efficient 3D Gaussian Splatting from Sparse Multi-view Images
Chen, Yuedong and Xu, Haofei and Zheng, Chuanxia and Zhuang, Bohan and Pollefeys, Marc and Geiger, Andreas and Cham, Tat-Jen and Cai, Jianfei. MVSplat : Efficient 3D Gaussian Splatting from Sparse Multi-view Images. Proceedings of the European Conference on Computer Vision (EC...
2025
-
[80]
DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization , year=
Li, Jiahe and Zhang, Jiawei and Bai, Xiao and Zheng, Jin and Ning, Xin and Zhou, Jun and Gu, Lin , booktitle=. DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization , year=
-
[81]
FSGS: Real-Time Few-Shot View Synthesis Using Gaussian Splatting
Zhu, Zehao and Fan, Zhiwen and Jiang, Yifan and Wang, Zhangyang. FSGS: Real-Time Few-Shot View Synthesis Using Gaussian Splatting. Computer Vision -- ECCV 2024. 2025
2024
-
[82]
PixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction , year=
Charatan, David and Li, Sizhe Lester and Tagliasacchi, Andrea and Sitzmann, Vincent , booktitle=. PixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction , year=
-
[83]
Lee, Joo Chan and Rho, Daniel and Sun, Xiangyu and Ko, Jong Hwan and Park, Eunbyung , booktitle=. Compact. 2024 , volume=
2024
-
[84]
Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis , year=
Niedermayr, Simon and Stumpfegger, Josef and Westermann, Rüdiger , booktitle=. Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis , year=
-
[85]
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Zhang, Xinjie and Ge, Xingtong and Xu, Tongda and He, Dailan and Wang, Yan and Qin, Hongwei and Lu, Guo and Geng, Jing and Zhang, Jun. GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting. Computer Vision -- ECCV 2024. 2025
2024
-
[86]
Mip-Splatting: Alias-Free 3D Gaussian Splatting , year=
Yu, Zehao and Chen, Anpei and Huang, Binbin and Sattler, Torsten and Geiger, Andreas , booktitle=. Mip-Splatting: Alias-Free 3D Gaussian Splatting , year=
-
[87]
Multi-Scale 3D Gaussian Splatting for Anti-Aliased Rendering , year=
Yan, Zhiwen and Low, Weng Fei and Chen, Yu and Lee, Gim Hee , booktitle=. Multi-Scale 3D Gaussian Splatting for Anti-Aliased Rendering , year=
-
[88]
GaussianShader: 3D Gaussian Splatting with Shading Functions for Reflective Surfaces , year=
Jiang, Yingwenqi and Tu, Jiadong and Liu, Yuan and Gao, Xifeng and Long, Xiaoxiao and Wang, Wenping and Ma, Yuexin , booktitle=. GaussianShader: 3D Gaussian Splatting with Shading Functions for Reflective Surfaces , year=
-
[89]
Gaussian Shadow Casting for Neural Characters , year=
Bolanos, Luis and Su, Shih-Yang and Rhodin, Helge , booktitle=. Gaussian Shadow Casting for Neural Characters , year=
-
[90]
2024 , booktitle =
Huang, Binbin and Yu, Zehao and Chen, Anpei and Geiger, Andreas and Gao, Shenghua , title =. 2024 , booktitle =
2024
-
[91]
2024 , volume=
Guédon, Antoine and Lepetit, Vincent , booktitle=. 2024 , volume=
2024
-
[92]
, booktitle=
Fu, Yang and Wang, Xiaolong and Liu, Sifei and Kulkarni, Amey and Kautz, Jan and Efros, Alexei A. , booktitle=. 2024 , volume=
2024
-
[93]
FreGS: 3D Gaussian Splatting with Progressive Frequency Regularization , year=
Zhang, Jiahui and Zhan, Fangneng and Xu, Muyu and Lu, Shijian and Xing, Eric , booktitle=. FreGS: 3D Gaussian Splatting with Progressive Frequency Regularization , year=
-
[94]
Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering , year=
Lu, Tao and Yu, Mulin and Xu, Linning and Xiangli, Yuanbo and Wang, Limin and Lin, Dahua and Dai, Bo , booktitle=. Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering , year=
-
[95]
Zhou, Shijie and Chang, Haoran and Jiang, Sicheng and Fan, Zhiwen and Zhu, Zehao and Xu, Dejia and Chari, Pradyumna and You, Suya and Wang, Zhangyang and Kadambi, Achuta , booktitle=. Feature. 2024 , volume=
2024
-
[96]
Gaussian G rouping: Segment and Edit Anything in 3D Scenes
Ye, Mingqiao and Danelljan, Martin and Yu, Fisher and Ke, Lei. Gaussian G rouping: Segment and Edit Anything in 3D Scenes. Proceedings of the European Conference on Computer Vision (ECCV). 2024
2024
-
[97]
Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding , year=
Shi, Jin-Chuan and Wang, Miao and Duan, Hao-Bin and Guan, Shao-Hua , booktitle=. Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding , year=
-
[98]
LangSplat: 3D Language Gaussian Splatting , year=
Qin, Minghan and Li, Wanhua and Zhou, Jiawei and Wang, Haoqian and Pfister, Hanspeter , booktitle=. LangSplat: 3D Language Gaussian Splatting , year=
-
[99]
Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction , year=
Yang, Ziyi and Gao, Xinyu and Zhou, Wen and Jiao, Shaohui and Zhang, Yuqing and Jin, Xiaogang , booktitle=. Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction , year=
-
[100]
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering , year=
Wu, Guanjun and Yi, Taoran and Fang, Jiemin and Xie, Lingxi and Zhang, Xiaopeng and Wei, Wei and Liu, Wenyu and Tian, Qi and Wang, Xinggang , booktitle=. 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering , year=
-
[101]
Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic Gaussians , year=
Xu, Yuelang and Chen, Bengwang and Li, Zhe and Zhang, Hongwen and Wang, Lizhen and Zheng, Zerong and Liu, Yebin , booktitle=. Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic Gaussians , year=
-
[102]
ACM Trans
Moenne-Loccoz, Nicolas and Mirzaei, Ashkan and Perel, Or and de Lutio, Riccardo and Martinez Esturo, Janick and State, Gavriel and Fidler, Sanja and Sharp, Nicholas and Gojcic, Zan , title =. ACM Trans. Graph. , month = nov, articleno =. 2024 , issue_date =
2024
-
[103]
ACM Trans
Condor, Jorge and Speierer, Sebastien and Bode, Lukas and Bozic, Aljaz and Green, Simon and Didyk, Piotr and Jarabo, Adrian , title =. ACM Trans. Graph. , month = feb, articleno =. 2025 , issue_date =
2025
-
[104]
2025 , journal =
A Survey on 3D Gaussian Splatting , author=. 2025 , journal =
2025
-
[105]
Deep Learning for Image and Point Cloud Fusion in Autonomous Driving: A Review , year=
Cui, Yaodong and Chen, Ren and Chu, Wenbo and Chen, Long and Tian, Daxin and Li, Ying and Cao, Dongpu , journal=. Deep Learning for Image and Point Cloud Fusion in Autonomous Driving: A Review , year=
-
[106]
and Keutzer, Kurt and Sangiovanni-Vincentelli, Alberto L
Yue, Xiangyu and Wu, Bichen and Seshia, Sanjit A. and Keutzer, Kurt and Sangiovanni-Vincentelli, Alberto L. , title =. 2018 , isbn =. doi:10.1145/3206025.3206080 , booktitle =
2018
-
[107]
Remote Sensing , VOLUME =
Yang, Su and Hou, Miaole and Li, Songnian , TITLE =. Remote Sensing , VOLUME =. 2023 , NUMBER =
2023
-
[108]
A Novel Radar Point Cloud Generation Method for Robot Environment Perception , year=
Cheng, Yuwei and Su, Jingran and Jiang, Mengxin and Liu, Yimin , journal=. A Novel Radar Point Cloud Generation Method for Robot Environment Perception , year=
-
[109]
, TITLE =
Alaba, Simegnew Yihunie and Ball, John E. , TITLE =. Sensors , VOLUME =. 2022 , NUMBER =
2022
-
[110]
A Comparative Survey of LiDAR-SLAM and LiDAR based Sensor Technologies , year=
Khan, Misha Urooj and Zaidi, Syed Azhar Ali and Ishtiaq, Arslan and Bukhari, Syeda Ume Rubab and Samer, Sana and Farman, Ayesha , booktitle=. A Comparative Survey of LiDAR-SLAM and LiDAR based Sensor Technologies , year=
-
[111]
and Tang, Lisa and Li, Jonathan , TITLE =
Wang, Junbo and Wang, Lanying and Feng, Shufang and Peng, Benrong and Huang, Lingfeng and Fatholahi, Sarah N. and Tang, Lisa and Li, Jonathan , TITLE =. Remote Sensing , VOLUME =. 2023 , NUMBER =
2023
-
[112]
3D Gaussian Splatting for Real-Time Radiance Field Rendering , journal =
Kerbl, Bernhard and Kopanas, Georgios and Leimk. 3D Gaussian Splatting for Real-Time Radiance Field Rendering , journal =. 2023 , articleno =. doi:10.1145/3592433 , url =
2023 doi
-
[113]
PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies , url =
Qian, Guocheng and Li, Yuchen and Peng, Houwen and Mai, Jinjie and Hammoud, Hasan and Elhoseiny, Mohamed and Ghanem, Bernard , booktitle =. PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies , url =
-
[114]
Meta Architecture for Point Cloud Analysis , year=
Lin, Haojia and Zheng, Xiawu and Li, Lijiang and Chao, Fei and Wang, Shanshan and Wang, Yan and Tian, Yonghong and Ji, Rongrong , booktitle=. Meta Architecture for Point Cloud Analysis , year=
-
[115]
NeO 360: Neural Fields for Sparse View Synthesis of Outdoor Scenes , year=
Irshad, Muhammad Zubair and Zakharov, Sergey and Liu, Katherine and Guizilini, Vitor and Kollar, Thomas and Gaidon, Adrien and Kira, Zsolt and Ambrus, Rares , booktitle=. NeO 360: Neural Fields for Sparse View Synthesis of Outdoor Scenes , year=
- [116]
- [117]
-
[118]
European Conference on Computer Vision , pages =
Shen, Qiuhong and Yang, Xingyi and Wang, Xinchao , title =. European Conference on Computer Vision , pages =. 2024 , publisher =
2024
-
[119]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Marrie, Juliette and Menegaux, Romain and Arbel, Michael and Larlus, Diane and Mairal, Julien , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[120]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Song, Yixiao and Li, Qingyong and Wang, Wen and Yan, Zhicheng , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[121]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Park, Chunghyun and Jeong, Yoonwoo and Cho, Minsu and Park, Jaesik , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[122]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Randla-net: Efficient semantic segmentation of large-scale point clouds , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[123]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
One thing one click: A self-training approach for weakly supervised 3d semantic segmentation , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[124]
European conference on computer vision , pages=
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding , author=. European conference on computer vision , pages=. 2020 , organization=
2020
-
[125]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Contrastive boundary learning for point cloud segmentation , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[126]
European Conference on Computer Vision , pages=
3d weakly supervised semantic segmentation with 2d vision-language guidance , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.