REVIEW 1 minor 181 references
A Cookbook of 3D Vision: Data, Learning Paradigms, and Application
T0 review · 0 major / 1 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read A data-centric taxonomy connects 3D geometric representations, datasets, and learning paradigms into one map.
desk verdict This is a survey paper that builds a taxonomy of 3D vision around representations like point clouds and Gaussians plus supervision types, but adds no new results or derivations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The data-centric taxonomy that integrates principal structural representations such as point clouds, meshes, voxels and 3D Gaussians with dataset design and supervision regimes to map connections to downstream tasks.
What would settle it
A survey or experiment finding no consistent patterns linking specific representation choices to measurable differences in task efficiency or fidelity would show that the taxonomy does not clarify the claimed relationships.
Extended reading notes
Core claim
We provide a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video mode
Load-bearing premise
The relationships among representations, learning paradigms, and downstream tasks can be clarified by analyzing principal structural representations and supervision regimes as described.
Editorial extensions
If this is right
- The taxonomy reveals how choices among point clouds, meshes, voxels, and 3D Gaussians interact with 2D-supervised learning to affect reconstruction quality.
- Dataset design and benchmark construction directly shape advances in implicit neural representations and 4D world modeling.
- Clearer links between supervision regimes and applications support trends that balance efficiency against fidelity in generation and video modeling.
- The map points to multimodal geometric grounding as a direction that ties representations to new task types.
Reading between the lines
- The taxonomy could be used to identify gaps where certain representation-supervision pairs lack dedicated benchmarks.
- Developers might apply the map to choose representations that match hardware or latency constraints in new applications.
- The same data-centric approach could be tested on related domains such as dynamic scene understanding to check if similar connections appear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents a data-centric taxonomy of 3D vision that organizes geometric representations (point clouds, meshes, voxels, 3D Gaussians) and their acquisition pipelines, then connects these to dataset design, benchmark construction, and supervision regimes (2D-supervised 3D learning, implicit neural representations, 4D world modeling), ultimately mapping the elements to applications in reconstruction, generation, and video modeling.
Significance. If the taxonomy accurately and comprehensively links representations, supervision regimes, and tasks, the work would supply a useful integrative map for a fragmented field, clarifying relationships and trends toward efficiency-fidelity trade-offs and multimodal grounding without advancing new theorems or empirical results.
minor comments (1)
- [Abstract] The abstract is dense with terminology; a short overview paragraph or figure in the introduction that visually summarizes the taxonomy axes would improve accessibility for readers new to the subfield.
Simulated Author's Rebuttal
We thank the referee for their positive review and recommendation to accept the manuscript. We appreciate the recognition that the data-centric taxonomy can serve as an integrative map linking representations, supervision regimes, and applications in 3D vision.
Circularity Check
No significant circularity; purely descriptive survey
full rationale
The paper constructs a data-centric taxonomy of 3D vision by surveying existing representations (point clouds, meshes, voxels, 3D Gaussians), acquisition methods, datasets, supervision regimes (2D-supervised, implicit, 4D), and applications. No equations, derivations, predictions, fitted parameters, or uniqueness theorems are present. The central contribution is an organizational map of the literature rather than any claim that reduces to its own inputs by construction. Self-citations, if present, are not load-bearing for any technical result because no technical results are derived. This matches the default expectation for a review paper with no derivation chain.
Assumptions & free parameters
Cite this review
Pith. "Pith review of A Cookbook of 3D Vision: Data, Learning Paradigms, and Application." pith.science (2026). https://pith.science/paper/LWRYQQNG
@misc{pith2026260604291,
author = {Pith},
title = {Pith review of: A Cookbook of 3D Vision: Data, Learning Paradigms, and Application},
year = {2026},
howpublished = {\url{https://pith.science/paper/LWRYQQNG}},
note = {Machine review of arXiv:2606.04291}
}
read the original abstract
3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains fragmented across representations and benchmarks, making it difficult to develop unified perspectives on efficiency, fidelity, and scalability. This work provides a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video modeling, offering a consolidated view of emerging trends toward balancing efficiency and fidelity and toward multimodal geometric grounding.
Reference graph
Works this paper leans on
-
[1]
Pointnetlk: Robust & efficient point cloud registration using pointnet
Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan, and Simon Lucey. Pointnetlk: Robust & efficient point cloud registration using pointnet. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[2]
Navigation world models, 2025
Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models, 2025
2025
-
[3]
A dataset for semantic scene understanding of lidar sequences
J Behley, M Garbade, A Milioto, J Quenzel, S Behnke, C Stachniss, J Gall, and Semantickitti. A dataset for semantic scene understanding of lidar sequences. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9297–9307
-
[4]
Virtual kitti 2, 2020
Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2, 2020
2020
-
[5]
Monoscene: Monocular 3d semantic scene completion, 2022
Anh-Quan Cao and Raoul de Charette. Monoscene: Monocular 3d semantic scene completion, 2022
2022
-
[6]
Matterport3d: Learning from rgb-d data in indoor environments, 2017
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments, 2017
2017
-
[7]
Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository, 2015
2015
-
[8]
Tensorf: Tensorial radiance fields
Anpei Chen, Zexiang Xu, Matthew Tancik, Jingyi Xu, Xiuming Zhang, Hiroharu Kato, and Jingyi Yu. Tensorf: Tensorial radiance fields. InECCV, pages 333–350, 2022
2022
Show all 181 references
-
[9]
A survey on 3d gaussian splatting, 2025
Guikun Chen and Wenguan Wang. A survey on 3d gaussian splatting, 2025
2025
-
[10]
Sam 3d: 3dfy anything in images.arXiv preprint arXiv:2511.16624, 2025
Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, et al. Sam 3d: 3dfy anything in images.arXiv preprint arXiv:2511.16624, 2025
2025 arXiv
-
[11]
Parametric 20000
Xi Cheng. Parametric 20000. Mendeley Data, V1, 2024
2024
-
[12]
Robust reconstruction of indoor scenes
Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun. Robust reconstruction of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5556–5565, 2015
2015
-
[13]
A large dataset of object scans
Sungjoon Choi, Qian-Yi Zhou, Stephen Miller, and Vladlen Koltun. A large dataset of object scans. arXiv:1602.02481, 2016
2016 arXiv
-
[14]
4d spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. InCVPR, pages 3075–3084, 2019
2019
-
[15]
Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese
Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction, 2016
2016
-
[16]
3d u-net: Learning dense volumetric segmentation from sparse annotation
Ozgun Cicek, Ahmed Abdulkadir, Soeren S Lienkamp, Thomas Brox, and Olaf Ronneberger. 3d u-net: Learning dense volumetric segmentation from sparse annotation. InMICCAI, pages 424–432, 2016
2016
-
[17]
Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022
Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik. Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022
2022
-
[18]
M. G. Cox. The numerical evaluation of b-splines.IMA Journal of Applied Mathematics, 10(2):134–149, 1972
1972
-
[19]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2432–2443, 2017
2017
-
[20]
Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017
Angela Dai, Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Christian Theobalt. Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017
2017
-
[21]
Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans
Angela Dai, Maximilian Dahnert, and Matthias Nießner. Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans. InCVPR, pages 4578–4587, 2018
2018
-
[22]
Brepformer: Transformer-based b-rep geometric feature recognition
Yongkang Dai, Xiaoshui Huang, Yunpeng Bai, Hao Guo, Hongping Gan, Ling Yang, and Yilei Shi. Brepformer: Transformer-based b-rep geometric feature recognition. InProceedings of the 2025 International Conference on Multimedia Retrieval, page 155–163, New York, NY, USA, 2025. Ass...
2025
-
[23]
Springer, New York, revised 2001 edition, 1978
Carl de Boor.A Practical Guide to Splines. Springer, New York, revised 2001 edition, 1978
2001
-
[24]
Objaverse-xl: A universe of 10m+ 3d objects.arXiv preprint arXiv:2307.05663, 2023
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objavers...
2023 arXiv
-
[25]
Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026
Hongyang Du, Junjie Ye, Xiaoyan Cong, Runhao Li, Jingcheng Ni, Aman Agarwal, Zeqi Zhou, Zekun Li, Randall Balestriero, and Yue Wang. Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026
2026
-
[26]
The mapillary traffic sign dataset for detection and classification on a global scale, 2020
Christian Ertler, Jerneja Mislej, Tobias Ollmann, Lorenzo Porzi, Gerhard Neuhold, and Yubin Kuang. The mapillary traffic sign dataset for detection and classification on a global scale, 2020
2020
-
[27]
A point set generation network for 3d object reconstruction from a single image, 2016
Haoqiang Fan, Hao Su, and Leonidas Guibas. A point set generation network for 3d object reconstruction from a single image, 2016
2016
-
[28]
A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025
Rubin Fan, Fazhi He, Yuxin Liu, and Jing Lin. A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025
2025
-
[29]
Morgan Kaufmann, San Diego, 5 edition, 2002
Gerald Farin.Curves and Surfaces for CAGD: A Practical Guide. Morgan Kaufmann, San Diego, 5 edition, 2002
2002
-
[30]
Alex Fisher, Ricardo Cannizzaro, Madeleine Cochrane, Chatura Nagahawatte, and Jennifer L. Palmer. Colmap: A memory-efficient occupancy grid mapping framework.Robotics and Autonomous Systems, 142:103755, 2021
2021
-
[31]
3d-future: 3d furniture shape with texture, 2020
Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture, 2020
2020
-
[32]
3d-front: 3d furnished rooms with layouts and semantics, 2021
Huan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang, Jiaming Wang Cao Li, Zengqi Xun, Chengyue Sun, Rongfei Jia, Binqiang Zhao, and Hao Zhang. 3d-front: 3d furnished rooms with layouts and semantics, 2021
2021
-
[33]
Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024
Rao Fu, Zehao Wen, Zichen Liu, and Srinath Sridhar. Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024
2024
-
[34]
Gigahands: A massive annotated dataset of bimanual hand activities
Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Fund, Daniel Ritchie, and Srinath Sridhar. Gigahands: A massive annotated dataset of bimanual hand activities. 2025
2025
-
[35]
Efros, and Xiaolong Wang
Yang Fu, Sifei Liu, Amey Kulkarni, Jan Kautz, Alexei A. Efros, and Xiaolong Wang. Colmap-free 3d gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20796–20805, 2024
2024
-
[36]
Accurate, dense, and robust multi-view stereopsis.IEEE Trans
Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multi-view stereopsis.IEEE Trans. on Pattern Analysis and Machine Intelligence, 32(8):1362–1376, 2010
2010
-
[37]
Virtual worlds as proxy for multi-object tracking analysis, 2016
Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual worlds as proxy for multi-object tracking analysis, 2016
2016
-
[38]
Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025
Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025
2025
-
[39]
Submanifold sparse convolutional networks
Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. Submanifold sparse convolutional networks. InCVPR, pages 9224–9232, 2018
2018
-
[40]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Eugene Patterson, Tsung-Yi Fu, Gijsbert Halbertsma, Lijun Zhao, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2024
-
[41]
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam Laradji, Hsueh-Ti (Derek) Liu, Henning Meyer, Yishu Miao, Derek Nowrouzeza...
2022
-
[42]
Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025
Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu. Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025. 12
2025
-
[43]
Roca: Robust cad model retrieval and alignment from a single image
Can Gumeli, Angela Dai, and Matthias Niebner. Roca: Robust cad model retrieval and alignment from a single image. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 4012–4021. IEEE, 2022
2022
-
[44]
Martin, and Shi-Min Hu
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, and Shi-Min Hu. Pct: Point cloud transformer.Computational Visual Media, 7(2):187–199, 2021
2021
-
[45]
Learning rich features from rgb-d images for object detection and segmentation
Saurabh Gupta, Ross Girshick, Pablo Arbeláez, and Jitendra Malik. Learning rich features from rgb-d images for object detection and segmentation. InEuropean Conference on Computer Vision (ECCV), pages 345–360, 2014
2014
-
[46]
Savinov, L
Timo Hackel, N. Savinov, L. Ladicky, Jan D. Wegner, K. Schindler, and M. Pollefeys. SEMANTIC3D.NET: A new large-scale point cloud classification benchmark. InISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, pages 91–98, 2017
2017
-
[47]
Dual transformer for point cloud analysis
Xian-Feng Han, Yi-Fei Jin, Hui-Xian Cheng, and Guo-Qiang Xiao. Dual transformer for point cloud analysis. IEEE Transactions on Multimedia, 25:5638–5648, 2023
2023
-
[48]
Meshcnn: A network with an edge
Rana Hanocka, Amir Hertz, Noa Fish, Raja Giryes, Shachar Fleishman, and Daniel Cohen-Or. Meshcnn: A network with an edge. InACM Transactions on Graphics (TOG), pages 1–12, 2019
2019
-
[49]
Deep learning based 3d segmentation: A survey, 2024
Yong He, Hongshan Yu, Xiaoyan Liu, Zhengeng Yang, Wei Sun, Saeed Anwar, and Ajmal Mian. Deep learning based 3d segmentation: A survey, 2024
2024
-
[50]
Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012
Peter Henry, Michael Krainin, Evan Herbst, Xiaofeng Ren, and Dieter Fox. Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012
2012
-
[51]
Hoffmann.Geometric and Solid Modeling: An Introduction
Christoph M. Hoffmann.Geometric and Solid Modeling: An Introduction. Morgan Kaufmann, San Mateo, CA, 1989
1989
-
[52]
Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025
Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jisang Han, Jiaolong Yang, Chong Luo, and Seungryong Kim. Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025
2025
-
[53]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022
2022 arXiv
-
[54]
Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023
2023 arXiv
-
[55]
3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023
2023
-
[56]
Scenenn: A scene meshes dataset with annotations
Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung. Scenenn: A scene meshes dataset with annotations. InInternational Conference on 3D Vision (3DV), 2016
2016
-
[57]
An embodied generalist agent in 3d world
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. InProceedings of the 41st International Conference on Machine Learning. JMLR.org, 2024
2024
-
[58]
Deepmvs: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. Deepmvs: Learning multi-view stereopsis. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[59]
Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025
Suning Huang, Qianzhong Chen, Xiaohan Zhang, Jiankai Sun, and Mac Schwager. Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025
2025
-
[60]
Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026
Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, and Li Fei-Fei. Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026
2026
-
[61]
Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera
Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Daniel Freeman, Andrew Davison, et al. Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera. InProceedings of the 24th Ann...
2011
-
[62]
Rayzer: A self-supervised large view synthesis model, 2025
Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, and Georgios Pavlakos. Rayzer: A self-supervised large view synthesis model, 2025
2025
-
[63]
Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data
Hanwen Jiang, Zexiang Xu, Desai Xie, Ziwen Chen, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang, Sai Bi, Xin Sun, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, and Hao Tan. Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data. InProceedings of the IEEE/CVF Conf...
2025
-
[64]
Rellis-3d dataset: Data, benchmarks and analysis, 2020
Peng Jiang, Philip Osteen, Maggie Wigness, and Srikanth Saripalli. Rellis-3d dataset: Data, benchmarks and analysis, 2020
2020
-
[65]
Tensoir: Tensorial inverse rendering, 2024
Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Songfang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. Tensoir: Tensorial inverse rendering, 2024
2024
-
[66]
Neural 3d mesh renderer, 2017
Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neural 3d mesh renderer, 2017
2017
-
[67]
Differentiable rendering: A survey, 2020
Hiroharu Kato, Deniz Beker, Mihai Morariu, Takahiro Ando, Toru Matsuoka, Wadim Kehl, and Adrien Gaidon. Differentiable rendering: A survey, 2020
2020
-
[68]
Screened poisson surface reconstruction.ACM Trans
Michael Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction.ACM Trans. Graph., 32(3), 2013
2013
-
[69]
Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006
Michael Kazhdan, Michael Bolitho, and Hugues Hoppe. Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006
2006
-
[70]
Posenet: A convolutional network for real-time 6-dof camera relocalization
Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2938–2946, 2015
2015
-
[71]
3d gaussian splatting for real-time radiance field rendering, 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering, 2023
2023
-
[72]
Egohumans: An egocentric 3d multi-human benchmark, 2023
Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. Egohumans: An egocentric 3d multi-human benchmark, 2023
2023
-
[73]
Parallel tracking and mapping for small ar workspaces
Georg Klein and David Murray. Parallel tracking and mapping for small ar workspaces. In2007 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, pages 225–234, 2007
2007
-
[74]
Abc: A big cad model dataset for geometric deep learning
Sebastian Koch, Albert Matveev, Zhongshi Jiang, Francis Williams, Alexey Artemov, Evgeny Burnaev, Marc Alexa, Denis Zorin, and Daniele Panozzo. Abc: A big cad model dataset for geometric deep learning. In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...
2019
-
[75]
Oneformer3d: One transformer for unified point cloud segmentation
Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, and Danila Rukhovich. Oneformer3d: One transformer for unified point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20943–20953, 2024
2024
-
[76]
Epipolar geometry improves video generation models, 2025
Orest Kupyn, Fabian Manhardt, Federico Tombari, and Christian Rupprecht. Epipolar geometry improves video generation models, 2025
2025
-
[77]
3d vision with transformers: A survey, 2022
Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, and Ming-Hsuan Yang. 3d vision with transformers: A survey, 2022
2022
-
[78]
Stratified transformer for 3d point cloud segmentation
Xin Lai, Jianhui Liu, Li Jiang, Liwei Wang, Hengshuang Zhao, Shu Liu, Xiaojuan Qi, and Jiaya Jia. Stratified transformer for 3d point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8500–8509, 2022
2022
-
[79]
Advances in 3d generation: A survey, 2024
Xiaoyu Li, Qi Zhang, Di Kang, Weihao Cheng, Yiming Gao, Jingbo Zhang, Zhihao Liang, Jing Liao, Yan-Pei Cao, and Ying Shan. Advances in 3d generation: A survey, 2024
2024
-
[80]
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[81]
Pointmamba: A simple state space model for point cloud analysis, 2024
Dingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu, Zhikang Zou, Xiaoqing Ye, Xiao Tan, and Xiang Bai. Pointmamba: A simple state space model for point cloud analysis, 2024
2024
-
[82]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[83]
Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang
Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025
2025 arXiv
-
[84]
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision, 2023
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, 14 Tianyi Zhang, Bedrich Benes, and Aniket Bera. Dl3dv-10k: A large-scale ...
2023
-
[85]
Point mamba: A novel point cloud backbone based on state space model with octree-based ordering strategy, 2024
Jiuming Liu, Ruiji Yu, Yian Wang, Yu Zheng, Tianchen Deng, Weicai Ye, and Hesheng Wang. Point mamba: A novel point cloud backbone based on state space model with octree-based ordering strategy, 2024
2024
-
[86]
Neural sparse voxel fields
Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Lin Bao, Jan Kautz, and Christian Theobalt. Neural sparse voxel fields. InNeurIPS, pages 15651–15663, 2020
2020
-
[87]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. InProceedings of the IEEE/CVF international conference on computer vision, pages 9298–9309, 2023
2023
-
[88]
Soft rasterizer: A differentiable renderer for image-based 3d reasoning, 2019
Shichen Liu, Tianye Li, Weikai Chen, and Hao Li. Soft rasterizer: A differentiable renderer for image-based 3d reasoning, 2019
2019
-
[89]
Loper and Michael J
Matthew M. Loper and Michael J. Black. Opendr: An approximate differentiable renderer. InComputer Vision – ECCV 2014, pages 154–169, Cham, 2014. Springer International Publishing
2014
-
[90]
Comport, Kefan Chen, and Srinath Sridhar
Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Mavidipalli, Dylan Hu, Andrew I. Comport, Kefan Chen, and Srinath Sridhar. Diva-360: The dynamic visual dataset for immersive neural fields. In2024 IEEE/CVF Conference on Compu...
2024
-
[91]
3dctn: 3d convolution-transformer network for point cloud classification.IEEE Transactions on Intelligent Transportation Systems, 23(12):24854–24865, 2022
Dening Lu, Qian Xie, Kyle Gao, Linlin Xu, and Jonathan Li. 3dctn: 3d convolution-transformer network for point cloud classification.IEEE Transactions on Intelligent Transportation Systems, 23(12):24854–24865, 2022
2022
-
[92]
Transformers in 3d point clouds: A survey, 2022
Dening Lu, Qian Xie, Mingqiang Wei, Kyle Gao, Linlin Xu, and Jonathan Li. Transformers in 3d point clouds: A survey, 2022
2022
-
[93]
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis, 2023
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis, 2023
2023
-
[94]
Ojea, and Ken Goldberg
Jeffrey Mahler, Jacky Liang, Siddhartha Niyaz, Michael Laskey, Richard Doan, Xue Bin Liu, Jose A. Ojea, and Ken Goldberg. Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics. InRobotics: Science and Systems (RSS), 2017
2017
-
[95]
Computer Science Press, Rockville, MD, 1988
Martti Mäntylä.An Introduction to Solid Modeling. Computer Science Press, Rockville, MD, 1988
1988
-
[96]
Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, and Andrew J. Davison. Gaussian splatting slam, 2024
2024
-
[97]
Voxnet: A 3d convolutional neural network for real-time object recognition
Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. InIROS, pages 922–928, 2015
2015
-
[98]
Occupancy networks: Learning 3d reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4460–4470, 2019
2019
-
[99]
Latent-nerf for shape-guided generation of 3d shapes and textures
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12663–12673, 2023
2023
-
[100]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: representing scenes as neural radiance fields for view synthesis.Commun. ACM, 65(1):99–106, 2021
2021
-
[101]
Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras.IEEE Transactions on Robotics, 33(5):1255–1262, 2017
Raul Mur-Artal and Juan D Tardos. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras.IEEE Transactions on Robotics, 33(5):1255–1262, 2017
2017
-
[102]
Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time
Richard A Newcombe, Dieter Fox, and Steven M Seitz. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 343–352, 2015
2015
-
[103]
Aria digital twin: A new benchmark dataset for egocentric 3d machine perception, 2023
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Carl Yuheng Ren. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception, 2023
2023
-
[104]
Patrikalakis and Takashi Maekawa.Shape Interrogation for Computer Aided Design and Manufac- turing
Nicholas M. Patrikalakis and Takashi Maekawa.Shape Interrogation for Computer Aided Design and Manufac- turing. Springer, Berlin, 2002. 15
2002
-
[105]
Springer, Berlin, 2 edition, 1997
Les Piegl and Wayne Tiller.The NURBS Book. Springer, Berlin, 2 edition, 1997
1997
-
[106]
Barron, and Ben Mildenhall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv, 2022
2022
-
[107]
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 652–660, 2017
2017
-
[108]
Qi, Hao Su, Kaichun Mo, and Leonidas J
Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation, 2017
2017
-
[109]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. InAdvances in Neural Information Processing Systems (NeurIPS), pages 5099–5108, 2017
2017
-
[110]
Frustum pointnets for 3d object detection from rgb-d data
Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 918–927, 2018
2018
-
[111]
3d object detection for autonomous driving: A survey.Pattern Recognition, 130:108796, 2022
Rui Qian, Xin Lai, and Xirong Li. 3d object detection for autonomous driving: A survey.Pattern Recognition, 130:108796, 2022
2022
-
[112]
Worldsimbench: Towards video generation models as world simulators, 2024
Yiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang, Enshen Zhou, Lijun Li, Zhenfei Yin, Xihui Liu, Lu Sheng, Jing Shao, Lei Bai, Wanli Ouyang, and Ruimao Zhang. Worldsimbench: Towards video generation models as world simulators, 2024
2024
-
[113]
Geometric transformer for fast and robust point cloud registration
Zheng Qin, Hao Yu, Changjian Wang, Yulan Guo, Yuxing Peng, and Kai Xu. Geometric transformer for fast and robust point cloud registration. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11143–11152, 2022
2022
-
[114]
Pu-transformer: Point cloud upsampling transformer
Shi Qiu, Saeed Anwar, and Nick Barnes. Pu-transformer: Point cloud upsampling transformer. InProceedings of the Asian Conference on Computer Vision (ACCV), pages 2475–2493, 2022
2022
-
[115]
Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X
Santhosh K. Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X. Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d enviro...
2021
-
[116]
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. InInternational Conference on Computer Vision, 2021
2021
-
[117]
Octnet: Learning deep 3d representations at high resolutions
Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. Octnet: Learning deep 3d representations at high resolutions. InCVPR, pages 3577–3586, 2017
2017
-
[118]
Susskind
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding, 2021
2021
-
[119]
Cofusion: Real-time segmentation, tracking and fusion of multiple objects
Nicholas Runz and Lourdes Agapito. Cofusion: Real-time segmentation, tracking and fusion of multiple objects. IEEE Transactions on Visualization and Computer Graphics, 24(11):2957–2968, 2018
2018
-
[120]
Pcrnet: Point cloud registration network using pointnet encoding, 2019
Vinit Sarode, Xueqian Li, Hunter Goforth, Yasuhiro Aoki, Rangaprasad Arun Srivatsan, Simon Lucey, and Howie Choset. Pcrnet: Point cloud registration network using pointnet encoding, 2019
2019
-
[121]
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-motion revisited. InConference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[122]
Pixelwise view selection for unstructured multi-view stereo
Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. InEuropean Conference on Computer Vision (ECCV), 2016
2016
-
[123]
Ari Seff, Yaniv Ovadia, Wenda Zhou, and Ryan P. Adams. Sketchgraphs: A large-scale dataset for modeling relational geometry in computer-aided design, 2020
2020
-
[124]
Mvdream: Multi-view diffusion for 3d generation, 2024
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation, 2024
2024
-
[125]
Real-time human pose recognition in parts from a single depth image
Jamie Shotton, Andrew Fitzgibbon, Mat Cook, Toby Sharp, Mark Finocchio, Richard Moore, Alex Kipman, and Andrew Blake. Real-time human pose recognition in parts from a single depth image. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pa...
2011
-
[126]
Indoor segmentation and support inference from rgb-d images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgb-d images. InEuropean Conference on Computer Vision (ECCV), pages 746–760. Springer, 2012
2012
-
[127]
Deep sliding shapes for amodal 3d object detection in rgb-d images
Shuran Song and Jianxiong Xiao. Deep sliding shapes for amodal 3d object detection in rgb-d images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 808–816, 2016
2016
-
[128]
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 567–576, 2015
2015
-
[129]
SpatialVerse Research Team
Manycore Tech Inc. SpatialVerse Research Team. Interiorgs: A 3d gaussian splatting dataset of semantically labeled indoor scenes.https://huggingface.co/datasets/spatialverse/InteriorGS, 2025
2025
-
[130]
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gi...
1906 arXiv
-
[131]
Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022
Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022
2022
-
[132]
Habitat 2.0: Training home assistants to rearrange their habitat, 2022
Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Ji...
2022
-
[133]
Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior.arXiv preprint arXiv:2303.14184, 2023
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior.arXiv preprint arXiv:2303.14184, 2023
2023
-
[134]
Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025
Tencent Hunyuan3D Team. Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025
2025
-
[135]
Deepv2d: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng. Deepv2d: Video to depth with differentiable structure from motion. InInternational Conference on Learning Representations (ICLR), 2020
2020
-
[136]
Consistent view synthesis with pose-guided diffusion models, 2023
Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models, 2023
2023
-
[137]
Zippered polygon meshes from range images
Greg Turk and Marc Levoy. Zippered polygon meshes from range images. InProceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques, page 311–318, New York, NY, USA, 1994. Association for Computing Machinery
1994
-
[138]
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data, 2019
Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data, 2019
2019
-
[139]
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. InProceedings of the European Conference on Computer Vision (ECCV), pages 601–617, 2018
2018
-
[140]
Wan: Open and advanced large-scale video generative models
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pande...
2025 arXiv
-
[141]
Vggt: Visual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Visual geometry grounded transformer. 2025
2025
-
[142]
O-cnn: Octree-based convolutional neural networks for 3d shape analysis.ACM TOG, 36(4):1–11, 2017
Peng-Shuai Wang, Yang Liu, Yueshan Guo, Chun-Yu Sun, and Xiao Tong. O-cnn: Octree-based convolutional neural networks for 3d shape analysis.ACM TOG, 36(4):1–11, 2017. 17
2017
-
[143]
Moge-2: Accurate monocular geometry with metric scale and sharp details, 2025
Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. Moge-2: Accurate monocular geometry with metric scale and sharp details, 2025
2025
-
[144]
DUSt3R: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jérôme Revaud. DUSt3R: Geometric 3d vision made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20697–20709, 2024
2024
-
[145]
Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, and Tong He.π3: Scalable permutation-equivariant visual geometry learning, 2025
2025
-
[146]
Pointramba: A hybrid transformer-mamba framework for point cloud analysis, 2024
Zicheng Wang, Zhenghao Chen, Yiming Wu, Zhen Zhao, Luping Zhou, and Dong Xu. Pointramba: A hybrid transformer-mamba framework for point cloud analysis, 2024
2024
-
[147]
Truncated signed distance function: Experiments on voxel size
Diana Werner, Ayoub Al-Hamadi, and Philipp Werner. Truncated signed distance function: Experiments on voxel size. InImage Analysis and Recognition, pages 357–364, Cham, 2014. Springer International Publishing
2014
-
[148]
Karl D. D. Willis, Yewen Pu, Jieliang Luo, Hang Chu, Tao Du, Joseph G. Lambourne, Armando Solar-Lezama, and Wojciech Matusik. Fusion 360 gallery: A dataset and environment for programmatic cad construction from human design sequences.ACM Transactions on Graphics (TOG), 40(4), 2021
2021
-
[149]
4d gaussian splatting for real-time dynamic scene rendering, 2024
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering, 2024
2024
-
[150]
Geometry forcing: Marrying video diffusion and 3d representation for consistent world modeling, 2025
Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, and Jiang Bian. Geometry forcing: Marrying video diffusion and 3d representation for consistent world modeling, 2025
2025
-
[151]
Deepcad: A deep generative network for computer-aided design models
Rundi Wu, Chang Xiao, and Changxi Zheng. Deepcad: A deep generative network for computer-aided design models. InICCV, pages 6762–6772, 2021
2021
-
[152]
Barron, and Aleksander Holynski
Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, and Aleksander Holynski. Cat4d: Create anything in 4d with multi-view video diffusion models, 2024
2024
-
[153]
Point transformer v2: Grouped vector attention and partition-based pooling, 2022
Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao. Point transformer v2: Grouped vector attention and partition-based pooling, 2022
2022
-
[154]
Point transformer v3: Simpler faster stronger
Xiaoyang Wu, Li Jiang, Peng-Shuai Wang, Zhijian Liu, Xihui Liu, Yu Qiao, Wanli Ouyang, Tong He, and Hengshuang Zhao. Point transformer v3: Simpler faster stronger. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4840–4851, 2024
2024
-
[155]
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. InCVPR, pages 1912–1920, 2015
1912
-
[156]
Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024
2024
-
[157]
Structured 3d latents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506, 2024
Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. Structured 3d latents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506, 2024
2024 arXiv
-
[158]
Sun3d: A database of big spaces reconstructed using sfm and object labels
Jianxiong Xiao, Andrew Owens, and Antonio Torralba. Sun3d: A database of big spaces reconstructed using sfm and object labels. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 1625–1632, 2013
2013
-
[159]
Instantmesh: Effi- cient 3d mesh generation from a single image with sparse-view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Effi- cient 3d mesh generation from a single image with sparse-view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024
2024 arXiv
-
[160]
Pointllm: Empowering large language models to understand point clouds.arXiv preprint arXiv:2308.16966, 2023
Runsen Xu, Xiaojuan Wang, Tai Wang, Kai Chen, Jiangmiao Pang, and Dahua Lin. Pointllm: Empowering large language models to understand point clouds.arXiv preprint arXiv:2308.16966, 2023
2023
-
[161]
Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction
Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[162]
Teaser: Fast and certifiable point cloud registration.IEEE Transactions on Robotics, 37(2):314–333, 2021
Heng Yang, Jingnan Shi, and Luca Carlone. Teaser: Fast and certifiable point cloud registration.IEEE Transactions on Robotics, 37(2):314–333, 2021. 18
2021
-
[163]
Sam 3d body: Robust full-body human mesh recovery.arXiv preprint, 2025
Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Dollar, and Kris Kitani. Sam 3d body: Robust full-body human mesh recovery.arXiv preprint, 2025
2025
-
[164]
Blendedmvs: A large-scale dataset for generalized multi-view stereo networks, 2020
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks, 2020
2020
-
[165]
No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images, 2024
Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images, 2024
2024
-
[166]
Scannet++: A high-fidelity dataset of 3d indoor scenes, 2023
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d indoor scenes, 2023
2023
-
[167]
Plenoctrees for real-time rendering of neural radiance fields, 2021
Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields, 2021
2021
-
[168]
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19291–19300, 2022
2022
-
[169]
Strobenet: Category-level multiview reconstruction of articulated objects, 2021
Ge Zhang, Or Litany, Srinath Sridhar, and Leonidas Guibas. Strobenet: Category-level multiview reconstruction of articulated objects, 2021
2021
-
[170]
Shuming Zhang, Zhidong Guan, Hao Jiang, Tao Ning, Xiaodong Wang, and Pingan Tan. Brep2seq: a dataset and hierarchical deep learning network for reconstruction and generation of computer-aided design models.Journal of Computational Design and Engineering, 11(1):110–134, 2024
2024
-
[171]
3d-scenedreamer: Text-driven 3d-consistent scene generation
Songchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma, Wei Hua, Hujun Bao, Weiwei Xu, and Changqing Zou. 3d-scenedreamer: Text-driven 3d-consistent scene generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10170–10180, 2024
2024
-
[172]
Point cloud mamba: Point cloud learning via state space model, 2024
Tao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang, Qianyu Zhou, Shunping Ji, Shuicheng Yan, and Xiangtai Li. Point cloud mamba: Point cloud learning via state space model, 2024
2024
-
[173]
Microsoft kinect sensor and its effect.IEEE Multimedia, 19(2):4–10, 2012
Zhengyou Zhang. Microsoft kinect sensor and its effect.IEEE Multimedia, 19(2):4–10, 2012
2012
-
[174]
Point transformer, 2021
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun. Point transformer, 2021
2021
-
[175]
Torr, and Vladlen Koltun
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H.S. Torr, and Vladlen Koltun. Point transformer. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 16259–16268, 2021
2021
-
[176]
Structured3d: A large photo-realistic dataset for structured 3d modeling, 2020
Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling, 2020
2020
-
[177]
Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J
Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J. Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking, 2023
2023
-
[178]
A comprehensive review of vision-based 3d reconstruction methods.Sensors, 24(7), 2024
Linglong Zhou, Guoxin Wu, Yunbo Zuo, Xuanyu Chen, and Hongle Hu. A comprehensive review of vision-based 3d reconstruction methods.Sensors, 24(7), 2024
2024
-
[179]
Thingi10k: A dataset of 10,000 3d-printing models.arXiv preprint arXiv:1605.04797, 2016
Qingnan Zhou and Alec Jacobson. Thingi10k: A dataset of 10,000 3d-printing models.arXiv preprint arXiv:1605.04797, 2016
2016 arXiv
-
[180]
Stereo magnification: Learning view synthesis using multiplane images, 2018
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images, 2018
2018
-
[181]
H3wb: Human3.6m 3d wholebody dataset and benchmark
Yue Zhu, Nermin Samet, and David Picard. H3wb: Human3.6m 3d wholebody dataset and benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 20166–20177, 2023. 19
2023
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.