REVIEW 3 major objections 3 minor 89 references
Ultra Ethernet's Design Principles and Architectural Innovations
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Ultra Ethernet 1.0 bets on a fully hardware-accelerated transport
desk verdict The submission is broken: the full text is an unrelated point-cloud paper, so the Ultra Ethernet overview cannot be evaluated; desk-reject unless the correct manuscript is provided. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Ultra Ethernet Transport (UET), the new transport-layer protocol at the core of Ultra Ethernet 1.0. The paper presents UET as a potentially fully hardware-accelerated protocol that handles reliability, speed, and efficiency at extreme scale, and it ties the feasibility of this design to two enabling forces: the broad Ethernet ecosystem and the roughly 1,000x improvement in computational efficiency per moved bit since InfiniBand's standardization. UET carries the argument by being the mechanism through which the spec's high-level goals become concrete network behavior.
What would settle it
Observe a production-scale UET implementation (for example, a cluster with tens of thousands of endpoints) under realistic AI training traffic and check whether any packet loss, congestion response, or reliability function falls back to software or drops below line rate. The central claim would be falsified if reliable full-hardware operation cannot be maintained at extreme scale, or if the measured per-bit compute cost is not close to the paper's 1,000x-efficiency premise.
Extended reading notes
Core claim
The paper argues that the recently released Ultra Ethernet 1.0 specification is a transformative high-performance Ethernet standard for AI and HPC systems, with contributions across the whole stack. Its central discovery is the Ultra Ethernet Transport (UET), a protocol engineered for reliable, fast, and efficient communication in extreme-scale systems and described as potentially fully hardware-accelerated. The intended consequence is that high-performance networking can be standardized on Ethernet rather than InfiniBand, leveraging the ecosystem scale of Ethernet and the dramatic growth in compute efficiency per bit moved over the past two decades.
Load-bearing premise
The load-bearing premise is that UET can be fully hardware-accelerated while still delivering reliable, fast, and efficient communication at extreme scale; if hardware offload forces software fallbacks or weakens reliability, the paper's stated advantage over InfiniBand collapses.
Editorial extensions
If this is right
- Extreme-scale AI and HPC clusters could run on standard Ethernet hardware carrying a reliable, loss-managed transport rather than requiring InfiniBand.
- If UET is fully hardware-accelerated as claimed, end hosts would offload transport processing and free CPU cycles for application work.
- The Ultra Ethernet 1.0 standard gives the industry a fresh, open, high-performance networking baseline to build on after more than two decades without a major standardization effort.
- Adoption of UET would let AI/HPC networking ride the cost and innovation curve of the much larger Ethernet ecosystem.
Reading between the lines
- A testable extension would be an apples-to-apples benchmark of UET against InfiniBand on a large GPU cluster, measuring achieved bandwidth, tail latency, and CPU cost under load; the paper itself does not report such measurements.
- If full hardware acceleration holds, UET could push high-performance transport features into commodity switches and NICs, potentially lowering the entry cost for smaller research clusters that currently cannot justify InfiniBand.
- The 1,000x compute-efficiency argument implies that the main risk to UET is not raw processing power but the complexity of reliability and congestion logic; that is where a software fallback would likely appear if the central claim fails.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, arXiv:2508.08906, is presented with an abstract claiming that the Ultra Ethernet (UE) 1.0 specification is a transformative high-performance Ethernet standard for AI and HPC systems, and that its standout contribution is the novel Ultra Ethernet Transport (UET), a potentially fully hardware-accelerated protocol engineered for reliable, fast, and efficient communication at extreme scale, positioned as an alternative to InfiniBand. However, the full text provided with the submission is arXiv:2508.08910v1, an unrelated computer-vision paper on masked clustering prediction for unsupervised point cloud pre-training. That full text contains no mention of Ultra Ethernet, UET, InfiniBand, PFC, DCQCN, transport offload, or any networking concept. As a result, the only content attributable to the claimed topic is the abstract, which offers assertions without any design description, protocol mechanisms, equations, benchmarks, or implementation details.
Significance. If the claimed UE overview were actually present, this paper would be a potentially significant community resource: a high-level overview of a new industry standard written by the specification's authors, including the falsifiable claim that UET can achieve full hardware acceleration while remaining reliable at scale. The paper should be credited for its explicit authorial connection to the specification and for raising a concrete architectural question with practical consequences for AI/HPC networking. However, none of this is verifiable in the submitted artifact: there are no derivable mechanisms, no machine-checked proofs, no reproducible code, no benchmarks, and no technical description of UET whatsoever. The significance of the claims therefore cannot be assessed because the manuscript as written does not contain the claimed content.
major comments (3)
- [Full Text (all sections)] The full text of this submission is arXiv:2508.08910v1, 'Masked Clustering Prediction for Unsupervised Point Cloud Pre-training,' which never mentions Ultra Ethernet, UET, InfiniBand, PFC, DCQCN, transport offload, or HPC networking; the abstract's central claim about UET is therefore entirely unsupported by the manuscript as submitted.
- [Abstract] The statement that UET is 'a potentially fully hardware-accelerated protocol engineered for reliable, fast, and efficient communication in extreme-scale systems' is an assertion with no accompanying description of UET's reliability mechanisms, offload architecture, congestion control, or performance results; because the word 'potentially' concedes the hardware-acceleration premise is not established, the central load-bearing claim is unverified.
- [Entire manuscript] The manuscript contains no equations, no protocol mechanisms, no state machines, no comparison methodology, no measurements, and no limitations section related to UE 1.0 or UET; the only limitation and broader-impact text present belongs to the unrelated point-cloud paper, so the paper cannot function as a technical overview of the UE 1.0 specification.
minor comments (3)
- [Abstract] The abstract's claim that UE 1.0 is 'transformative' is stated without comparative data or a definition of the comparison baseline; if the paper is revised with the correct full text, this term should be substantiated or qualified.
- [Abstract] The abstract compares UE to InfiniBand but does not cite any InfiniBand specification or prior comparative study; a revised version should include such references.
- [Manuscript structure] The submission contains no section headings or reference list for the claimed UE overview, making it impossible to navigate the purported material; the root cause appears to be the text mismatch between the abstract and the provided full text.
Circularity Check
No circular step is identifiable: the submitted full text is an unrelated point-cloud paper, so the UET overview's assertions are unevidenced rather than derivational-circular.
full rationale
The abstract claims UET is "a potentially fully hardware-accelerated protocol engineered for reliable, fast, and efficient communication in extreme-scale systems," but the full text provided is arXiv:2508.08910v1, "Masked Clustering Prediction for Unsupervised Point Cloud Pre-training," which contains no description of Ultra Ethernet, UET, InfiniBand, or transport mechanisms. There is therefore no derivation chain, no fitted parameter renamed as a prediction, and no equation whose output equals its input. The hard rules require quoting a specific reduction (Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction) to assign a nonzero circularity score; no such reduction can be exhibited from the supplied artifact. Asserting a protocol capability without providing the protocol's technical content is an evidentiary gap and a correctness/verifiability concern, not circularity in the sense of this axis. Consequently the appropriate score is 0, with the caveat that the paper's central claim cannot be evaluated from the submitted text at all.
Assumptions & free parameters
assumptions (3)
- domain assumption UET can be fully hardware-accelerated while providing reliable, fast, and efficient communication.
- domain assumption The Ethernet ecosystem and roughly 1000x computational-efficiency gains make a high-performance Ethernet-based alternative to InfiniBand viable.
- ad hoc to paper UE 1.0 represents a 'transformative' improvement over prior Ethernet and InfiniBand.
invented entities (1)
-
Ultra Ethernet Transport (UET)
independent evidence
Cite this review
Pith. "Pith review of Ultra Ethernet's Design Principles and Architectural Innovations." pith.science (2026). https://pith.science/paper/3YE6S4SI
@misc{pith2026250808906,
author = {Pith},
title = {Pith review of: Ultra Ethernet's Design Principles and Architectural Innovations},
year = {2026},
howpublished = {\url{https://pith.science/paper/3YE6S4SI}},
note = {Machine review of arXiv:2508.08906}
}
read the original abstract
The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance Computing (HPC) systems. This paper, written by the specification's authors, provides a high-level overview of UE's design, offering crucial motivations and scientific context to understand its innovations. While UE introduces advancements across the entire Ethernet stack, its standout contribution is the novel Ultra Ethernet Transport (UET), a potentially fully hardware-accelerated protocol engineered for reliable, fast, and efficient communication in extreme-scale systems. Unlike InfiniBand, the last major standardization effort in high-performance networking over two decades ago, UE leverages the expansive Ethernet ecosystem and the 1,000x gains in computational efficiency per moved bit to deliver a new era of high-performance networking.
Reference graph
Works this paper leans on
-
[1]
Learning represen- tations and generative models for 3d point clouds
Panos Achlioptas, Olga Diamanti, et al. Learning represen- tations and generative models for 3d point clouds. In ICML, pages 40–49, 2018. 2
2018
-
[2]
Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding
Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri, Kanchana Thilakarathna, and Ranga Ro- drigo. Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding. In CVPR, pages 9902–9912, 2022. 5, 6, 3
2022
-
[3]
Zamir, Helen Jiang, Ioan- nis Brilakis, Martin Fischer, and Silvio Savarese
Iro Armeni, Ozan Sener, Amir R. Zamir, Helen Jiang, Ioan- nis Brilakis, Martin Fischer, and Silvio Savarese. 3d seman- tic parsing of large-scale indoor spaces. In CVPR, 2016. 2, 6
2016
-
[4]
Deep clustering for unsupervised learning 5 GTPoint-MAEMaskClu (Ours) Figure B
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning 5 GTPoint-MAEMaskClu (Ours) Figure B. Part segmentation results on the ShapeNetPart dataset [6]. The visualization compares the predictions from MaskClu and PointMAE [7] against the ground truth annotations ( GT). The examples illustrate the supe...
2018
-
[5]
SL3D: Self-supervised-Self-labeled 3D Recognition
Fernando Julio Cendra, Lan Ma, Jiajun Shen, and Xiaojuan Qi. Sl3d: Self-supervised-self-labeled 3d recognition. arXiv preprint arXiv:2210.16810, 2022. 2
work page Pith review arXiv 2022
-
[6]
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 2, 5, 7, 3, 4, 6
arXiv 2015
-
[7]
Pimae: Point cloud and image interactive masked autoencoders for 3d object detection
Anthony Chen, Kevin Zhang, Renrui Zhang, Zihan Wang, Yuheng Lu, Yandong Guo, and Shanghang Zhang. Pimae: Point cloud and image interactive masked autoencoders for 3d object detection. In CVPR, pages 5291–5301, 2023. 3, 6
2023
-
[8]
Pointgpt: Auto-regressively generative pre- training from point clouds
Guangyan Chen, Meiling Wang, Yi Yang, Kai Yu, Li Yuan, and Yufeng Yue. Pointgpt: Auto-regressively generative pre- training from point clouds. NeurIPS, 36, 2024. 3, 6, 7, 8
2024
Show all 89 references
-
[9]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, et al. A simple framework for contrastive learning of visual representations. ICML, 2020. 2
2020
-
[10]
Exploring simple siamese rep- resentation learning
Xinlei Chen and Kaiming He. Exploring simple siamese rep- resentation learning. In CVPR, 2021. 3, 5
2021
-
[11]
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. NeurIPS, 26, 2013. 5
2013
-
[12]
Scannet: 6 Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: 6 Richly-annotated 3d reconstructions of indoor scenes. In CVPR, pages 5828–5839, 2017. 2, 6, 4
2017
-
[13]
Autoen- coders as cross-modal teachers: Can pretrained 2d image transformers help 3d representation learning? arXiv preprint arXiv:2212.08320, 2022
Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang, Jian- jian Sun, Zheng Ge, Li Yi, and Kaisheng Ma. Autoen- coders as cross-modal teachers: Can pretrained 2d image transformers help 3d representation learning? arXiv preprint arXiv:2212.08320, 2022. 6, 7, 8, 3
2022 arXiv
-
[14]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. InInternational ...
2020
-
[15]
Self-supervised learning on 3d point clouds by learning dis- crete generative models
Benjamin Eckart, Wentao Yuan, Chao Liu, and Jan Kautz. Self-supervised learning on 3d point clouds by learning dis- crete generative models. In CVPR, pages 8248–8257, 2021. 1
2021
-
[16]
Point transformer
Nico Engel, Vasileios Belagiannis, and Klaus Dietmayer. Point transformer. IEEE access , 9:134826–134840, 2021. 5, 2
2021
-
[17]
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Moham- mad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. In N...
2020
-
[18]
Efficiently mod- eling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher R´e. Efficiently mod- eling long sequences with structured state spaces. In ICLR,
-
[19]
Adamw and super- convergence is now the fastest way to train neural nets
Sylvain Gugger and Jeremy Howard. Adamw and super- convergence is now the fastest way to train neural nets. last accessed, 19, 2018. 2
2018
-
[20]
Pct: Point cloud transformer
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu. Pct: Point cloud transformer. Computational Visual Media, 7:187–199, 2021. 2
2021
-
[21]
Joint-mae: 2d-3d joint masked autoen- coders for 3d point cloud pre-training
Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzhi Li, and Pheng-Ann Heng. Joint-mae: 2d-3d joint masked autoen- coders for 3d point cloud pre-training. In Proceedings of the Thirty-Second International Joint Conference on Artifi- cial Intelligence, pages 791–799, 2023. 3, 8
2023
-
[22]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, pages 16000–16009, 2022. 3
2022
-
[23]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. NeurIPS, 33:6840–6851, 2020. 3
2020
-
[24]
Ponder: Point cloud pre-training via neural rendering
Di Huang, Sida Peng, Tong He, Honghui Yang, Xiaowei Zhou, and Wanli Ouyang. Ponder: Point cloud pre-training via neural rendering. In CVPR, pages 16089–16098, 2023. 3
2023
-
[25]
Spatio-temporal self-supervised representation learning for 3d point clouds
Siyuan Huang, Yichen Xie, Song-Chun Zhu, and Yixin Zhu. Spatio-temporal self-supervised representation learning for 3d point clouds. In CVPR, pages 6535–6545, 2021. 2
2021
-
[26]
Free-form language-based robotic reasoning and grasping
Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bor- tolon, Sergio Povoli, Guofeng Mei, Yiming Wang, and Fabio Poiesi. Free-form language-based robotic reasoning and grasping. arXiv preprint arXiv:2503.13082, 2025. 1
2025 arXiv
-
[27]
Self-supervised feature learning by cross-modality and cross-view corre- spondences
Longlong Jing, Ling Zhang, and Yingli Tian. Self-supervised feature learning by cross-modality and cross-view corre- spondences. In CVPR, pages 1581–1591, 2021. 2
2021
-
[28]
Point cloud gan
Chun-Liang Li, Manzil Zaheer, Yang Zhang, Barnabas Poc- zos, and Ruslan Salakhutdinov. Point cloud gan. CoRR, abs/1810.05795, 2018. 3
2018 arXiv
-
[29]
Pointcnn: Convolution on x- transformed points
Yangyan Li, Rui Bu, et al. Pointcnn: Convolution on x- transformed points. In NeurIPS, pages 820–830, 2018. 2
2018
-
[30]
Scenesplat: Gaussian splatting-based scene understanding with vision-language pretraining
Yue Li, Qi Ma, Runyi Yang, Huapeng Li, Mengjiao Ma, Bin Ren, Nikola Popovic, Nicu Sebe, Ender Konukoglu, Theo Gevers, et al. Scenesplat: Gaussian splatting-based scene understanding with vision-language pretraining. In ICCV,
-
[31]
Pvafn: Point-voxel at- tention fusion network with multi-pooling enhancing for 3d object detection
Yidi Li, Jiahao Wen, Rui Gong, Bin Ren, Wenhao Li, Chen Cheng, Hong Liu, and Nicu Sebe. Pvafn: Point-voxel at- tention fusion network with multi-pooling enhancing for 3d object detection. Expert Systems with Applications , 281: 127608, 2025. 1
2025
-
[32]
Pointmamba: A simple state space model for point cloud analysis
Dingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu, Zhikang Zou, Xiaoqing Ye, Xiao Tan, and Xiang Bai. Pointmamba: A simple state space model for point cloud analysis. In NeurIPS, 2024. 1, 3
2024
-
[33]
Spatio-temporal graph diffusion for text-driven human motion generation
Chang Liu, Mengyi Zhao, Bin Ren, Mengyuan Liu, Nicu Sebe, et al. Spatio-temporal graph diffusion for text-driven human motion generation. In BMVC, pages 722–729, 2023. 3
2023
-
[34]
Point discriminative learning for data- efficient 3d point cloud analysis
Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Chaitanya K Joshi, and Jie Lin. Point discriminative learning for data- efficient 3d point cloud analysis. In3DV, pages 42–51. IEEE,
-
[35]
Masked discrimina- tion for self-supervised learning on point clouds
Haotian Liu, Mu Cai, and Yong Jae Lee. Masked discrimina- tion for self-supervised learning on point clouds. In ECCV, pages 657–675. Springer, 2022. 1, 2, 3, 5, 6, 8
2022
-
[36]
L2g auto-encoder: Under- standing point clouds by local-to-global reconstruction with hierarchical self-attention
Xinhai Liu, Zhizhong Han, et al. L2g auto-encoder: Under- standing point clouds by local-to-global reconstruction with hierarchical self-attention. In ACM MM , pages 989–997,
-
[37]
Densepoint: Learning densely contextual representation for efficient point cloud process- ing
Yongcheng Liu, Bin Fan, Gaofeng Meng, Jiwen Lu, Shiming Xiang, and Chunhong Pan. Densepoint: Learning densely contextual representation for efficient point cloud process- ing. In CVPR, pages 5239–5248, 2019. 2
2019
-
[38]
Relation-shape convolutional neural network for point cloud analysis
Yongcheng Liu, Bin Fan, et al. Relation-shape convolutional neural network for point cloud analysis. In CVPR, pages 8895–8904, 2019. 6
2019
-
[39]
Pointclustering: Unsupervised point cloud pre-training using transformation invariance in clustering
Fuchen Long, Ting Yao, Zhaofan Qiu, Lusong Li, and Tao Mei. Pointclustering: Unsupervised point cloud pre-training using transformation invariance in clustering. In CVPR, pages 21824–21834, 2023. 2, 6, 8
2023
-
[40]
Scenesplat++: A large dataset and compre- hensive benchmark for language gaussian splatting
Mengjiao Ma, Qi Ma, Yue Li, Jiahuan Cheng, Runyi Yang, Bin Ren, Nikola Popovic, Mingqiang Wei, Nicu Sebe, Luc Van Gool, et al. Scenesplat++: A large dataset and compre- hensive benchmark for language gaussian splatting. arXiv preprint arXiv:2506.08710, 2025. 1
2025
-
[41]
Shapes- plat: A large-scale dataset of gaussian splats and their self- supervised pretraining
Qi Ma, Yue Li, Bin Ren, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, and Danda Pani Paudel. Shapes- plat: A large-scale dataset of gaussian splats and their self- supervised pretraining. arXiv preprint arXiv:2408.10906 ,
-
[42]
Data augmentation- free unsupervised learning for 3d point cloud understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang, Elisa Ricci, Nicu Sebe, and Qiang Wu. Data augmentation- free unsupervised learning for 3d point cloud understanding. BMVC, 2022. 1, 2, 6
2022
-
[43]
Unsupervised point cloud representation learning by clustering and neural ren- dering
Guofeng Mei, Cristiano Saltori, Elisa Ricci, Nicu Sebe, Qiang Wu, Jian Zhang, and Fabio Poiesi. Unsupervised point cloud representation learning by clustering and neural ren- dering. International Journal of Computer Vision, pages 1– 19, 2024. 1, 2, 6
2024
-
[44]
Self-supervised and gen- eralizable tokenization for clip-based 3d understanding
Guofeng Mei, Bin Ren, Juan Liu, Luigi Riz, Xiaoshui Huang, Xu Zheng, Yongshun Gong, Ming-Hsuan Yang, Nicu Sebe, and Fabio Poiesi. Self-supervised and gen- eralizable tokenization for clip-based 3d understanding. arXiv:2505.18819, 2025. 7
2025 arXiv
-
[45]
Scene- graphloc: Cross-modal coarse visual localization on 3d scene graphs
Yang Miao, Francis Engelmann, Olga Vysotska, Federico Tombari, Marc Pollefeys, and D ´aniel B ´ela Bar ´ath. Scene- graphloc: Cross-modal coarse visual localization on 3d scene graphs. In European Conference on Computer Vision, pages 127–150. Springer, 2024. 1
2024
-
[46]
An end- to-end transformer model for 3d object detection
Ishan Misra, Rohit Girdhar, and Armand Joulin. An end- to-end transformer model for 3d object detection. In CVPR, pages 2906–2917, 2021. 6
2021
-
[47]
Masked autoencoders for point cloud self-supervised learning
Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part II, pages 604–621. Springer,
2022
-
[48]
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, et al. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, pages 652–660, 2017. 2
2017
-
[49]
Pointnet++: Deep hier- archical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, et al. Pointnet++: Deep hier- archical feature learning on point sets in a metric space. In NeurIPS, pages 5099–5108, 2017. 5, 2
2017
-
[50]
Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining
Zekun Qi, Runpei Dong, Guofan Fan, Zheng Ge, Xiangyu Zhang, Kaisheng Ma, and Li Yi. Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining. In ICML, pages 28223–28243. PMLR, 2023. 2, 6
2023
-
[51]
Dense-resolution network for point cloud classification and segmentation
Shi Qiu, Saeed Anwar, and Nick Barnes. Dense-resolution network for point cloud classification and segmentation. In WACV, pages 3813–3822, 2021. 2
2021
-
[52]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, pages 8748–8763. PMLR, 2021. 2
2021
-
[53]
Global-local bidi- rectional reasoning for unsupervised representation learning of 3d point clouds
Yongming Rao, Jiwen Lu, and Jie Zhou. Global-local bidi- rectional reasoning for unsupervised representation learning of 3d point clouds. In CVPR, 2020. 1, 2
2020
-
[54]
Cascaded cross mlp- mixer gans for cross-view image translation
Bin Ren, Hao Tang, Nicu Sebe, et al. Cascaded cross mlp- mixer gans for cross-view image translation. In British Ma- chine Vision Conference (BMVC’21) , pages 1–14. British Machine Vision Association, BMV A, 2021. 1
2021
-
[55]
Masked jigsaw puzzle: A versatile po- sition embedding for vision transformers
Bin Ren, Yahui Liu, Yue Song, Wei Bi, Rita Cucchiara, Nicu Sebe, and Wei Wang. Masked jigsaw puzzle: A versatile po- sition embedding for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20382–20391, 2023. 1
2023
-
[56]
Bringing masked autoencoders explicit con- trastive properties for point cloud self-supervised learning
Bin Ren, Guofeng Mei, Danda Pani Paudel, Weijie Wang, Yawei Li, Mengyuan Liu, Rita Cucchiara, Luc Van Gool, and Nicu Sebe. Bringing masked autoencoders explicit con- trastive properties for point cloud self-supervised learning. In ACCV, 2024. 1, 3, 5, 6, 8, 2
2024
-
[57]
Info3d: Representation learning on 3d ob- jects using mutual information maximization and contrastive learning
Aditya Sanghi. Info3d: Representation learning on 3d ob- jects using mutual information maximization and contrastive learning. In European Conference on Computer Vision , pages 626–642. Springer, 2020. 1, 2, 3
2020
-
[58]
Rl-gan-net: A reinforcement learning agent controlled gan network for real-time point cloud shape completion
Muhammad Sarmad, Hyunjoo Jenny Lee, et al. Rl-gan-net: A reinforcement learning agent controlled gan network for real-time point cloud shape completion. In CVPR, pages 5898–5907, 2019. 1
2019
-
[59]
Self-supervised deep learning on point clouds by reconstructing space
Jonathan Sauder and Bjarne Sievers. Self-supervised deep learning on point clouds by reconstructing space. In NeurIPS, pages 12942–12952, 2019. 5
2019
-
[60]
Earth- mind: Towards multi-granular and multi-sensor earth ob- servation with large multimodal models
Yan Shu, Bin Ren, Zhitong Xiong, Danda Pani Paudel, Luc Van Gool, Begum Demir, Nicu Sebe, and Paolo Rota. Earth- mind: Towards multi-granular and multi-sensor earth ob- servation with large multimodal models. arXiv preprint arXiv:2506.01667, 2025. 2
2025
-
[61]
Pointgrow: Autoregressively learned point cloud generation with self- attention
Yongbin Sun, Yue Wang, Ziwei Liu, et al. Pointgrow: Autoregressively learned point cloud generation with self- attention. In WACV, pages 61–70, 2020. 1
2020
-
[62]
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. InCVPR, pages 1588–1597, 2019. 2
2019
-
[63]
Revisit- ing point cloud classification: A new benchmark dataset and classification model on real-world data
Mikaela Angelina Uy, Quang-Hieu Pham, et al. Revisit- ing point cloud classification: A new benchmark dataset and classification model on real-world data. In ICCV, pages 1588–1597, 2019. 6
2019
-
[64]
Attention is all you need
A Vaswani. Attention is all you need. NeurIPS, 2017. 1
2017
-
[65]
Unsupervised point cloud pre-training via occlusion completion
Hanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby, and Matt J Kusner. Unsupervised point cloud pre-training via occlusion completion. In CVPR, pages 9782–9792, 2021. 5, 6, 8, 2, 3
2021
-
[66]
Zero-shot point cloud registration
Weijie Wang, Guofeng Mei, Bin Ren, Xiaoshui Huang, Fabio Poiesi, Luc Van Gool, Nicu Sebe, and Bruno Lepri. Zero-shot point cloud registration. arXiv preprint arXiv:2312.03032, 2023. 2
2023 arXiv
-
[67]
Dynamic graph cnn for learning on point clouds
Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. ACM TOG, 38(5): 1–12, 2019. 5, 2, 3
2019
-
[68]
Rolo-slam: rotation-optimized lidar-only slam in uneven ter- rain with ground vehicle
Yinchuan Wang, Bin Ren, Xiang Zhang, Pengyu Wang, Chaoqun Wang, Rui Song, Yibin Li, and Max Q-H Meng. Rolo-slam: rotation-optimized lidar-only slam in uneven ter- rain with ground vehicle. Journal of Field Robotics, 42(3): 880–902, 2025. 1
2025
-
[69]
Take-a-photo: 3d-to-2d generative pre-training of point cloud models
Ziyi Wang, Xumin Yu, Yongming Rao, Jie Zhou, and Jiwen Lu. Take-a-photo: 3d-to-2d generative pre-training of point cloud models. In CVPR, pages 5640–5650, 2023. 3 8
2023
-
[70]
Diffusion models as masked autoencoders
Chen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li, Haoqi Fan, Hu Xu, Huiyu Wang, Cihang Xie, Alan Yuille, and Christoph Feichtenhofer. Diffusion models as masked autoencoders. In CVPR, pages 16284–16294, 2023. 3
2023
-
[71]
Pointconv: Deep convolutional networks on 3d point clouds
Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In CVPR, pages 9621–9630, 2019. 2
2019
-
[72]
3d shapenets: A deep repre- sentation for volumetric shapes
Zhirong Wu, Shuran Song, et al. 3d shapenets: A deep repre- sentation for volumetric shapes. InCVPR, pages 1912–1920,
1912
-
[73]
Pointcontrast: Unsupervised pre- training for 3d point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany. Pointcontrast: Unsupervised pre- training for 3d point cloud understanding. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 , pages...
2020
-
[74]
Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Mart ´ın-Mart´ın, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese. Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding. In CVPR, pages 1179–1189, 2023. 2
2023
-
[75]
Gd-mae: gener- ative decoder for mae pre-training on lidar point clouds
Honghui Yang, Tong He, Jiaheng Liu, Hua Chen, Boxi Wu, Binbin Lin, Xiaofei He, and Wanli Ouyang. Gd-mae: gener- ative decoder for mae pre-training on lidar point clouds. In CVPR, pages 9403–9414, 2023. 3
2023
-
[76]
Fold- ingnet: Point cloud auto-encoder via deep grid deformation
Yaoqing Yang, Chen Feng, Yiru Shen, and Dong Tian. Fold- ingnet: Point cloud auto-encoder via deep grid deformation. In CVPR, pages 206–215, 2018. 1, 2
2018
-
[77]
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In CVPR, pages 19313–19322, 2022. 3, 6, 8, 1, 2
2022
-
[78]
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In CVPR, pages 19313–19322, 2022. 5, 6
2022
-
[79]
Towards compact 3d representations via point feature enhancement masked au- toencoders
Yaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li, Tao Dai, Bin Chen, Zhi Wang, and Shu-Tao Xia. Towards compact 3d representations via point feature enhancement masked au- toencoders. In AAAI, pages 6962–6970, 2024. 6
2024
-
[80]
Online deep clustering for unsupervised representation learning
Xiaohang Zhan, Jiahao Xie, Ziwei Liu, Yew-Soon Ong, and Chen Change Loy. Online deep clustering for unsupervised representation learning. In CVPR, pages 6688–6697, 2020. 1
2020
-
[81]
Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training
Renrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang, Bin Zhao, Dong Wang, Yu Qiao, and Hongsheng Li. Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training. arXiv preprint arXiv:2205.14401, 2022. 3, 5, 6, 8
2022 arXiv
-
[82]
Pointclip: Point cloud understanding by clip
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xu- peng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In CVPR, pages 8552–8562, 2022. 2
2022
-
[83]
Pcp- mae: Learning to predict centers for point masked autoen- coders
Xiangdong Zhang, Shaofeng Zhang, and Junchi Yan. Pcp- mae: Learning to predict centers for point masked autoen- coders. arXiv preprint arXiv:2408.08753, 2024. 6
2024 arXiv
-
[84]
Masked surfel prediction for self-supervised point cloud learning
Yabin Zhang, Jiehong Lin, Chenhang He, Yongwei Chen, Kui Jia, and Lei Zhang. Masked surfel prediction for self-supervised point cloud learning. arXiv preprint arXiv:2207.03111, 2022. 6
2022 arXiv
-
[85]
Point-dae: Denoising autoencoders for self-supervised point cloud learning
Yabin Zhang, Jiehong Lin, Ruihuang Li, Kui Jia, and Lei Zhang. Point-dae: Denoising autoencoders for self-supervised point cloud learning. arXiv preprint arXiv:2211.06841, 2022. 1, 5, 6
2022 arXiv
-
[86]
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In CVPR, pages 16259– 16268, 2021. 2
2021
-
[87]
Denoising diffusion probabilistic models for action-conditioned 3d motion generation
Mengyi Zhao, Mengyuan Liu, Bin Ren, Shuling Dai, and Nicu Sebe. Denoising diffusion probabilistic models for action-conditioned 3d motion generation. In ICASSP, pages 4225–4229. IEEE, 2024. 3
2024
-
[88]
Deep image clustering based on curriculum learning and density information
Haiyang Zheng, Ruilin Zhang, and Hongpeng Wang. Deep image clustering based on curriculum learning and density information. In Proceedings of the 2024 International Con- ference on Multimedia Retrieval, pages 330–338, 2024. 1
2024
-
[89]
Point cloud pre-training with diffusion models
Xiao Zheng, Xiaoshui Huang, Guofeng Mei, Yuenan Hou, Zhaoyang Lyu, Bo Dai, Wanli Ouyang, and Yongshun Gong. Point cloud pre-training with diffusion models. In CVPR, pages 22935–22945, 2024. 1, 3 9
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.