Pith. sign in

REVIEW 1 minor 181 references

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

T0 review · 0 major / 1 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read A data-centric taxonomy connects 3D geometric representations, datasets, and learning paradigms into one map.

desk verdict This is a survey paper that builds a taxonomy of 3D vision around representations like point clouds and Gaussians plus supervision types, but adds no new results or derivations. read the letter →

arxiv 2606.04291 v1 pith:LWRYQQNG submitted 2026-06-02 cs.CV

classification cs.CV
keywords 3Dvisiondata-centrictaxonomygeometricrepresentationspointcloudsmeshesvoxelsGaussiansimplicitneural
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper establishes a unified conceptual map for 3D vision by organizing it around geometric data representations and how they pair with different learning approaches. A reader would care because the field is currently split across many benchmarks and methods, making it hard to see how to build scalable systems. The taxonomy starts with core representations including point clouds, meshes, voxels, and 3D Gaussians and their data collection methods. It then connects these to dataset designs, supervision types like 2D-supervised learning and implicit neural representations, and applications in reconstruction, generation, and 4D modeling. The result is a clearer picture of trends that aim to improve both efficiency and accuracy in 3D tasks.

What carries the argument

The data-centric taxonomy that integrates principal structural representations such as point clouds, meshes, voxels and 3D Gaussians with dataset design and supervision regimes to map connections to downstream tasks.

What would settle it

A survey or experiment finding no consistent patterns linking specific representation choices to measurable differences in task efficiency or fidelity would show that the taxonomy does not clarify the claimed relationships.

Watch

Extended reading notes

Core claim

We provide a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video mode

Load-bearing premise

The relationships among representations, learning paradigms, and downstream tasks can be clarified by analyzing principal structural representations and supervision regimes as described.

Editorial extensions

If this is right

  • The taxonomy reveals how choices among point clouds, meshes, voxels, and 3D Gaussians interact with 2D-supervised learning to affect reconstruction quality.
  • Dataset design and benchmark construction directly shape advances in implicit neural representations and 4D world modeling.
  • Clearer links between supervision regimes and applications support trends that balance efficiency against fidelity in generation and video modeling.
  • The map points to multimodal geometric grounding as a direction that ties representations to new task types.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could be used to identify gaps where certain representation-supervision pairs lack dedicated benchmarks.
  • Developers might apply the map to choose representations that match hardware or latency constraints in new applications.
  • The same data-centric approach could be tested on related domains such as dynamic scene understanding to check if similar connections appear.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 1 minor

Summary. The manuscript presents a data-centric taxonomy of 3D vision that organizes geometric representations (point clouds, meshes, voxels, 3D Gaussians) and their acquisition pipelines, then connects these to dataset design, benchmark construction, and supervision regimes (2D-supervised 3D learning, implicit neural representations, 4D world modeling), ultimately mapping the elements to applications in reconstruction, generation, and video modeling.

Significance. If the taxonomy accurately and comprehensively links representations, supervision regimes, and tasks, the work would supply a useful integrative map for a fragmented field, clarifying relationships and trends toward efficiency-fidelity trade-offs and multimodal grounding without advancing new theorems or empirical results.

minor comments (1)
  1. [Abstract] The abstract is dense with terminology; a short overview paragraph or figure in the introduction that visually summarizes the taxonomy axes would improve accessibility for readers new to the subfield.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive review and recommendation to accept the manuscript. We appreciate the recognition that the data-centric taxonomy can serve as an integrative map linking representations, supervision regimes, and applications in 3D vision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; purely descriptive survey

full rationale

The paper constructs a data-centric taxonomy of 3D vision by surveying existing representations (point clouds, meshes, voxels, 3D Gaussians), acquisition methods, datasets, supervision regimes (2D-supervised, implicit, 4D), and applications. No equations, derivations, predictions, fitted parameters, or uniqueness theorems are present. The central contribution is an organizational map of the literature rather than any claim that reduces to its own inputs by construction. Self-citations, if present, are not load-bearing for any technical result because no technical results are derived. This matches the default expectation for a review paper with no derivation chain.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is a survey paper. No free parameters, axioms, or invented entities are introduced because the work does not advance new technical claims.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Cookbook of 3D Vision: Data, Learning Paradigms, and Application." pith.science (2026). https://pith.science/paper/LWRYQQNG

@misc{pith2026260604291,
  author       = {Pith},
  title        = {Pith review of: A Cookbook of 3D Vision: Data, Learning Paradigms, and Application},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LWRYQQNG}},
  note         = {Machine review of arXiv:2606.04291}
}
read the original abstract

3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains fragmented across representations and benchmarks, making it difficult to develop unified perspectives on efficiency, fidelity, and scalability. This work provides a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video modeling, offering a consolidated view of emerging trends toward balancing efficiency and fidelity and toward multimodal geometric grounding.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

181 extracted references · 13 canonical work pages

  1. [1]

    Pointnetlk: Robust & efficient point cloud registration using pointnet

    Yasuhiro Aoki, Hunter Goforth, Rangaprasad Arun Srivatsan, and Simon Lucey. Pointnetlk: Robust & efficient point cloud registration using pointnet. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  2. [2]

    Navigation world models, 2025

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models, 2025

  3. [3]

    A dataset for semantic scene understanding of lidar sequences

    J Behley, M Garbade, A Milioto, J Quenzel, S Behnke, C Stachniss, J Gall, and Semantickitti. A dataset for semantic scene understanding of lidar sequences. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9297–9307

  4. [4]

    Virtual kitti 2, 2020

    Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2, 2020

  5. [5]

    Monoscene: Monocular 3d semantic scene completion, 2022

    Anh-Quan Cao and Raoul de Charette. Monoscene: Monocular 3d semantic scene completion, 2022

  6. [6]

    Matterport3d: Learning from rgb-d data in indoor environments, 2017

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments, 2017

  7. [7]

    Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu

    Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository, 2015

  8. [8]

    Tensorf: Tensorial radiance fields

    Anpei Chen, Zexiang Xu, Matthew Tancik, Jingyi Xu, Xiuming Zhang, Hiroharu Kato, and Jingyi Yu. Tensorf: Tensorial radiance fields. InECCV, pages 333–350, 2022

Show all 181 references
  1. [9]

    A survey on 3d gaussian splatting, 2025

    Guikun Chen and Wenguan Wang. A survey on 3d gaussian splatting, 2025

  2. [10]

    Sam 3d: 3dfy anything in images.arXiv preprint arXiv:2511.16624, 2025

    Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, et al. Sam 3d: 3dfy anything in images.arXiv preprint arXiv:2511.16624, 2025

  3. [11]

    Parametric 20000

    Xi Cheng. Parametric 20000. Mendeley Data, V1, 2024

  4. [12]

    Robust reconstruction of indoor scenes

    Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun. Robust reconstruction of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5556–5565, 2015

  5. [13]

    A large dataset of object scans

    Sungjoon Choi, Qian-Yi Zhou, Stephen Miller, and Vladlen Koltun. A large dataset of object scans. arXiv:1602.02481, 2016

  6. [14]

    4d spatio-temporal convnets: Minkowski convolutional neural networks

    Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. InCVPR, pages 3075–3084, 2019

  7. [15]

    Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese

    Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction, 2016

  8. [16]

    3d u-net: Learning dense volumetric segmentation from sparse annotation

    Ozgun Cicek, Ahmed Abdulkadir, Soeren S Lienkamp, Thomas Brox, and Olaf Ronneberger. 3d u-net: Learning dense volumetric segmentation from sparse annotation. InMICCAI, pages 424–432, 2016

  9. [17]

    Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022

    Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik. Abo: Dataset and benchmarks for real-world 3d object understanding.CVPR, 2022

  10. [18]

    M. G. Cox. The numerical evaluation of b-splines.IMA Journal of Applied Mathematics, 10(2):134–149, 1972

  11. [19]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2432–2443, 2017

  12. [20]

    Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017

    Angela Dai, Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Christian Theobalt. Bundlefusion: Real- time globally consistent 3d reconstruction using on-the-fly surface reintegration.ACM Transactions on Graphics (TOG), 36(4):1–18, 2017

  13. [21]

    Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans

    Angela Dai, Maximilian Dahnert, and Matthias Nießner. Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans. InCVPR, pages 4578–4587, 2018

  14. [22]

    Brepformer: Transformer-based b-rep geometric feature recognition

    Yongkang Dai, Xiaoshui Huang, Yunpeng Bai, Hao Guo, Hongping Gan, Ling Yang, and Yilei Shi. Brepformer: Transformer-based b-rep geometric feature recognition. InProceedings of the 2025 International Conference on Multimedia Retrieval, page 155–163, New York, NY, USA, 2025. Ass...

  15. [23]

    Springer, New York, revised 2001 edition, 1978

    Carl de Boor.A Practical Guide to Splines. Springer, New York, revised 2001 edition, 1978

  16. [24]

    Objaverse-xl: A universe of 10m+ 3d objects.arXiv preprint arXiv:2307.05663, 2023

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objavers...

  17. [25]

    Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026

    Hongyang Du, Junjie Ye, Xiaoyan Cong, Runhao Li, Jingcheng Ni, Aman Agarwal, Zeqi Zhou, Zekun Li, Randall Balestriero, and Yue Wang. Videogpa: Distilling geometry priors for 3d-consistent video generation, 2026

  18. [26]

    The mapillary traffic sign dataset for detection and classification on a global scale, 2020

    Christian Ertler, Jerneja Mislej, Tobias Ollmann, Lorenzo Porzi, Gerhard Neuhold, and Yubin Kuang. The mapillary traffic sign dataset for detection and classification on a global scale, 2020

  19. [27]

    A point set generation network for 3d object reconstruction from a single image, 2016

    Haoqiang Fan, Hao Su, and Leonidas Guibas. A point set generation network for 3d object reconstruction from a single image, 2016

  20. [28]

    A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025

    Rubin Fan, Fazhi He, Yuxin Liu, and Jing Lin. A history-based parametric cad sketch dataset with advanced engineering commands.Computer-Aided Design, 182:103848, 2025

  21. [29]

    Morgan Kaufmann, San Diego, 5 edition, 2002

    Gerald Farin.Curves and Surfaces for CAGD: A Practical Guide. Morgan Kaufmann, San Diego, 5 edition, 2002

  22. [30]

    Alex Fisher, Ricardo Cannizzaro, Madeleine Cochrane, Chatura Nagahawatte, and Jennifer L. Palmer. Colmap: A memory-efficient occupancy grid mapping framework.Robotics and Autonomous Systems, 142:103755, 2021

  23. [31]

    3d-future: 3d furniture shape with texture, 2020

    Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture, 2020

  24. [32]

    3d-front: 3d furnished rooms with layouts and semantics, 2021

    Huan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang, Jiaming Wang Cao Li, Zengqi Xun, Chengyue Sun, Rongfei Jia, Binqiang Zhao, and Hao Zhang. 3d-front: 3d furnished rooms with layouts and semantics, 2021

  25. [33]

    Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024

    Rao Fu, Zehao Wen, Zichen Liu, and Srinath Sridhar. Anyhome: Open-vocabulary generation of structured and textured 3d homes, 2024

  26. [34]

    Gigahands: A massive annotated dataset of bimanual hand activities

    Rao Fu, Dingxi Zhang, Alex Jiang, Wanjia Fu, Austin Fund, Daniel Ritchie, and Srinath Sridhar. Gigahands: A massive annotated dataset of bimanual hand activities. 2025

  27. [35]

    Efros, and Xiaolong Wang

    Yang Fu, Sifei Liu, Amey Kulkarni, Jan Kautz, Alexei A. Efros, and Xiaolong Wang. Colmap-free 3d gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20796–20805, 2024

  28. [36]

    Accurate, dense, and robust multi-view stereopsis.IEEE Trans

    Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multi-view stereopsis.IEEE Trans. on Pattern Analysis and Machine Intelligence, 32(8):1362–1376, 2010

  29. [37]

    Virtual worlds as proxy for multi-object tracking analysis, 2016

    Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual worlds as proxy for multi-object tracking analysis, 2016

  30. [38]

    Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025

    Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. Nerf: Neural radiance field in 3d vision: A comprehensive review (updated post-gaussian splatting), 2025

  31. [39]

    Submanifold sparse convolutional networks

    Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. Submanifold sparse convolutional networks. InCVPR, pages 9224–9232, 2018

  32. [40]

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, Eugene Patterson, Tsung-Yi Fu, Gijsbert Halbertsma, Lijun Zhao, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  33. [41]

    Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam Laradji, Hsueh-Ti (Derek) Liu, Henning Meyer, Yishu Miao, Derek Nowrouzeza...

  34. [42]

    Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025

    Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu. Diffusion as shader: 3d-aware video diffusion for versatile video generation control, 2025. 12

  35. [43]

    Roca: Robust cad model retrieval and alignment from a single image

    Can Gumeli, Angela Dai, and Matthias Niebner. Roca: Robust cad model retrieval and alignment from a single image. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 4012–4021. IEEE, 2022

  36. [44]

    Martin, and Shi-Min Hu

    Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, and Shi-Min Hu. Pct: Point cloud transformer.Computational Visual Media, 7(2):187–199, 2021

  37. [45]

    Learning rich features from rgb-d images for object detection and segmentation

    Saurabh Gupta, Ross Girshick, Pablo Arbeláez, and Jitendra Malik. Learning rich features from rgb-d images for object detection and segmentation. InEuropean Conference on Computer Vision (ECCV), pages 345–360, 2014

  38. [46]

    Savinov, L

    Timo Hackel, N. Savinov, L. Ladicky, Jan D. Wegner, K. Schindler, and M. Pollefeys. SEMANTIC3D.NET: A new large-scale point cloud classification benchmark. InISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, pages 91–98, 2017

  39. [47]

    Dual transformer for point cloud analysis

    Xian-Feng Han, Yi-Fei Jin, Hui-Xian Cheng, and Guo-Qiang Xiao. Dual transformer for point cloud analysis. IEEE Transactions on Multimedia, 25:5638–5648, 2023

  40. [48]

    Meshcnn: A network with an edge

    Rana Hanocka, Amir Hertz, Noa Fish, Raja Giryes, Shachar Fleishman, and Daniel Cohen-Or. Meshcnn: A network with an edge. InACM Transactions on Graphics (TOG), pages 1–12, 2019

  41. [49]

    Deep learning based 3d segmentation: A survey, 2024

    Yong He, Hongshan Yu, Xiaoyan Liu, Zhengeng Yang, Wei Sun, Saeed Anwar, and Ajmal Mian. Deep learning based 3d segmentation: A survey, 2024

  42. [50]

    Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012

    Peter Henry, Michael Krainin, Evan Herbst, Xiaofeng Ren, and Dieter Fox. Rgb-d mapping: Using depth cameras for dense 3d modeling of indoor environments.The International Journal of Robotics Research, 31(5):647–663, 2012

  43. [51]

    Hoffmann.Geometric and Solid Modeling: An Introduction

    Christoph M. Hoffmann.Geometric and Solid Modeling: An Introduction. Morgan Kaufmann, San Mateo, CA, 1989

  44. [52]

    Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025

    Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jisang Han, Jiaolong Yang, Chong Luo, and Seungryong Kim. Pf3plat: Pose-free feed-forward 3d gaussian splatting, 2025

  45. [53]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022

  46. [54]

    Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d.arXiv preprint arXiv:2311.04400, 2023

  47. [55]

    3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023

    Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36, 2023

  48. [56]

    Scenenn: A scene meshes dataset with annotations

    Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung. Scenenn: A scene meshes dataset with annotations. InInternational Conference on 3D Vision (3DV), 2016

  49. [57]

    An embodied generalist agent in 3d world

    Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. InProceedings of the 41st International Conference on Machine Learning. JMLR.org, 2024

  50. [58]

    Deepmvs: Learning multi-view stereopsis

    Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. Deepmvs: Learning multi-view stereopsis. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  51. [59]

    Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025

    Suning Huang, Qianzhong Chen, Xiaohan Zhang, Jiankai Sun, and Mac Schwager. Particleformer: A 3d point cloud world model for multi-object, multi-material robotic manipulation, 2025

  52. [60]

    Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026

    Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, and Li Fei-Fei. Pointworld: Scaling 3d world models for in-the-wild robotic manipulation, 2026

  53. [61]

    Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera

    Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Daniel Freeman, Andrew Davison, et al. Kinectfusion: Real-time 3d reconstruction and interaction using a moving depth camera. InProceedings of the 24th Ann...

  54. [62]

    Rayzer: A self-supervised large view synthesis model, 2025

    Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, and Georgios Pavlakos. Rayzer: A self-supervised large view synthesis model, 2025

  55. [63]

    Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data

    Hanwen Jiang, Zexiang Xu, Desai Xie, Ziwen Chen, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang, Sai Bi, Xin Sun, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, and Hao Tan. Megasynth: Scaling up 3d scene 13 reconstruction with synthesized data. InProceedings of the IEEE/CVF Conf...

  56. [64]

    Rellis-3d dataset: Data, benchmarks and analysis, 2020

    Peng Jiang, Philip Osteen, Maggie Wigness, and Srikanth Saripalli. Rellis-3d dataset: Data, benchmarks and analysis, 2020

  57. [65]

    Tensoir: Tensorial inverse rendering, 2024

    Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Songfang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. Tensoir: Tensorial inverse rendering, 2024

  58. [66]

    Neural 3d mesh renderer, 2017

    Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neural 3d mesh renderer, 2017

  59. [67]

    Differentiable rendering: A survey, 2020

    Hiroharu Kato, Deniz Beker, Mihai Morariu, Takahiro Ando, Toru Matsuoka, Wadim Kehl, and Adrien Gaidon. Differentiable rendering: A survey, 2020

  60. [68]

    Screened poisson surface reconstruction.ACM Trans

    Michael Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction.ACM Trans. Graph., 32(3), 2013

  61. [69]

    Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006

    Michael Kazhdan, Michael Bolitho, and Hugues Hoppe. Poisson surface reconstruction.Proceedings of the Fourth Eurographics Symposium on Geometry Processing, 7:61–70, 2006

  62. [70]

    Posenet: A convolutional network for real-time 6-dof camera relocalization

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2938–2946, 2015

  63. [71]

    3d gaussian splatting for real-time radiance field rendering, 2023

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering, 2023

  64. [72]

    Egohumans: An egocentric 3d multi-human benchmark, 2023

    Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. Egohumans: An egocentric 3d multi-human benchmark, 2023

  65. [73]

    Parallel tracking and mapping for small ar workspaces

    Georg Klein and David Murray. Parallel tracking and mapping for small ar workspaces. In2007 6th IEEE and ACM International Symposium on Mixed and Augmented Reality, pages 225–234, 2007

  66. [74]

    Abc: A big cad model dataset for geometric deep learning

    Sebastian Koch, Albert Matveev, Zhongshi Jiang, Francis Williams, Alexey Artemov, Evgeny Burnaev, Marc Alexa, Denis Zorin, and Daniele Panozzo. Abc: A big cad model dataset for geometric deep learning. In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...

  67. [75]

    Oneformer3d: One transformer for unified point cloud segmentation

    Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, and Danila Rukhovich. Oneformer3d: One transformer for unified point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20943–20953, 2024

  68. [76]

    Epipolar geometry improves video generation models, 2025

    Orest Kupyn, Fabian Manhardt, Federico Tombari, and Christian Rupprecht. Epipolar geometry improves video generation models, 2025

  69. [77]

    3d vision with transformers: A survey, 2022

    Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, and Ming-Hsuan Yang. 3d vision with transformers: A survey, 2022

  70. [78]

    Stratified transformer for 3d point cloud segmentation

    Xin Lai, Jianhui Liu, Li Jiang, Liwei Wang, Hengshuang Zhao, Shu Liu, Xiaojuan Qi, and Jiaya Jia. Stratified transformer for 3d point cloud segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8500–8509, 2022

  71. [79]

    Advances in 3d generation: A survey, 2024

    Xiaoyu Li, Qi Zhang, Di Kang, Weihao Cheng, Yiming Gao, Jingbo Zhang, Zhihao Liang, Jing Liao, Yan-Pei Cao, and Ying Shan. Advances in 3d generation: A survey, 2024

  72. [80]

    Megadepth: Learning single-view depth prediction from internet photos

    Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018

  73. [81]

    Pointmamba: A simple state space model for point cloud analysis, 2024

    Dingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu, Zhikang Zou, Xiaoqing Ye, Xiao Tan, and Xiang Bai. Pointmamba: A simple state space model for point cloud analysis, 2024

  74. [82]

    Magic3d: High-resolution text-to-3d content creation

    Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  75. [83]

    Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang

    Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025

  76. [84]

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision, 2023

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, 14 Tianyi Zhang, Bedrich Benes, and Aniket Bera. Dl3dv-10k: A large-scale ...

  77. [85]

    Point mamba: A novel point cloud backbone based on state space model with octree-based ordering strategy, 2024

    Jiuming Liu, Ruiji Yu, Yian Wang, Yu Zheng, Tianchen Deng, Weicai Ye, and Hesheng Wang. Point mamba: A novel point cloud backbone based on state space model with octree-based ordering strategy, 2024

  78. [86]

    Neural sparse voxel fields

    Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Lin Bao, Jan Kautz, and Christian Theobalt. Neural sparse voxel fields. InNeurIPS, pages 15651–15663, 2020

  79. [87]

    Zero-1-to-3: Zero-shot one image to 3d object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. InProceedings of the IEEE/CVF international conference on computer vision, pages 9298–9309, 2023

  80. [88]

    Soft rasterizer: A differentiable renderer for image-based 3d reasoning, 2019

    Shichen Liu, Tianye Li, Weikai Chen, and Hao Li. Soft rasterizer: A differentiable renderer for image-based 3d reasoning, 2019

  81. [89]

    Loper and Michael J

    Matthew M. Loper and Michael J. Black. Opendr: An approximate differentiable renderer. InComputer Vision – ECCV 2014, pages 154–169, Cham, 2014. Springer International Publishing

  82. [90]

    Comport, Kefan Chen, and Srinath Sridhar

    Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Mavidipalli, Dylan Hu, Andrew I. Comport, Kefan Chen, and Srinath Sridhar. Diva-360: The dynamic visual dataset for immersive neural fields. In2024 IEEE/CVF Conference on Compu...

  83. [91]

    3dctn: 3d convolution-transformer network for point cloud classification.IEEE Transactions on Intelligent Transportation Systems, 23(12):24854–24865, 2022

    Dening Lu, Qian Xie, Kyle Gao, Linlin Xu, and Jonathan Li. 3dctn: 3d convolution-transformer network for point cloud classification.IEEE Transactions on Intelligent Transportation Systems, 23(12):24854–24865, 2022

  84. [92]

    Transformers in 3d point clouds: A survey, 2022

    Dening Lu, Qian Xie, Mingqiang Wei, Kyle Gao, Linlin Xu, and Jonathan Li. Transformers in 3d point clouds: A survey, 2022

  85. [93]

    Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis, 2023

    Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis, 2023

  86. [94]

    Ojea, and Ken Goldberg

    Jeffrey Mahler, Jacky Liang, Siddhartha Niyaz, Michael Laskey, Richard Doan, Xue Bin Liu, Jose A. Ojea, and Ken Goldberg. Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics. InRobotics: Science and Systems (RSS), 2017

  87. [95]

    Computer Science Press, Rockville, MD, 1988

    Martti Mäntylä.An Introduction to Solid Modeling. Computer Science Press, Rockville, MD, 1988

  88. [96]

    Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, and Andrew J. Davison. Gaussian splatting slam, 2024

  89. [97]

    Voxnet: A 3d convolutional neural network for real-time object recognition

    Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. InIROS, pages 922–928, 2015

  90. [98]

    Occupancy networks: Learning 3d reconstruction in function space

    Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4460–4470, 2019

  91. [99]

    Latent-nerf for shape-guided generation of 3d shapes and textures

    Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12663–12673, 2023

  92. [100]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: representing scenes as neural radiance fields for view synthesis.Commun. ACM, 65(1):99–106, 2021

  93. [101]

    Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras.IEEE Transactions on Robotics, 33(5):1255–1262, 2017

    Raul Mur-Artal and Juan D Tardos. Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras.IEEE Transactions on Robotics, 33(5):1255–1262, 2017

  94. [102]

    Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time

    Richard A Newcombe, Dieter Fox, and Steven M Seitz. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 343–352, 2015

  95. [103]

    Aria digital twin: A new benchmark dataset for egocentric 3d machine perception, 2023

    Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Carl Yuheng Ren. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception, 2023

  96. [104]

    Patrikalakis and Takashi Maekawa.Shape Interrogation for Computer Aided Design and Manufac- turing

    Nicholas M. Patrikalakis and Takashi Maekawa.Shape Interrogation for Computer Aided Design and Manufac- turing. Springer, Berlin, 2002. 15

  97. [105]

    Springer, Berlin, 2 edition, 1997

    Les Piegl and Wayne Tiller.The NURBS Book. Springer, Berlin, 2 edition, 1997

  98. [106]

    Barron, and Ben Mildenhall

    Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv, 2022

  99. [107]

    Pointnet: Deep learning on point sets for 3d classification and segmentation

    Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 652–660, 2017

  100. [108]

    Qi, Hao Su, Kaichun Mo, and Leonidas J

    Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation, 2017

  101. [109]

    Pointnet++: Deep hierarchical feature learning on point sets in a metric space

    Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. InAdvances in Neural Information Processing Systems (NeurIPS), pages 5099–5108, 2017

  102. [110]

    Frustum pointnets for 3d object detection from rgb-d data

    Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 918–927, 2018

  103. [111]

    3d object detection for autonomous driving: A survey.Pattern Recognition, 130:108796, 2022

    Rui Qian, Xin Lai, and Xirong Li. 3d object detection for autonomous driving: A survey.Pattern Recognition, 130:108796, 2022

  104. [112]

    Worldsimbench: Towards video generation models as world simulators, 2024

    Yiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang, Enshen Zhou, Lijun Li, Zhenfei Yin, Xihui Liu, Lu Sheng, Jing Shao, Lei Bai, Wanli Ouyang, and Ruimao Zhang. Worldsimbench: Towards video generation models as world simulators, 2024

  105. [113]

    Geometric transformer for fast and robust point cloud registration

    Zheng Qin, Hao Yu, Changjian Wang, Yulan Guo, Yuxing Peng, and Kai Xu. Geometric transformer for fast and robust point cloud registration. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11143–11152, 2022

  106. [114]

    Pu-transformer: Point cloud upsampling transformer

    Shi Qiu, Saeed Anwar, and Nick Barnes. Pu-transformer: Point cloud upsampling transformer. InProceedings of the Asian Conference on Computer Vision (ACCV), pages 2475–2493, 2022

  107. [115]

    Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X

    Santhosh K. Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X. Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d enviro...

  108. [116]

    Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

    Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. InInternational Conference on Computer Vision, 2021

  109. [117]

    Octnet: Learning deep 3d representations at high resolutions

    Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. Octnet: Learning deep 3d representations at high resolutions. InCVPR, pages 3577–3586, 2017

  110. [118]

    Susskind

    Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding, 2021

  111. [119]

    Cofusion: Real-time segmentation, tracking and fusion of multiple objects

    Nicholas Runz and Lourdes Agapito. Cofusion: Real-time segmentation, tracking and fusion of multiple objects. IEEE Transactions on Visualization and Computer Graphics, 24(11):2957–2968, 2018

  112. [120]

    Pcrnet: Point cloud registration network using pointnet encoding, 2019

    Vinit Sarode, Xueqian Li, Hunter Goforth, Yasuhiro Aoki, Rangaprasad Arun Srivatsan, Simon Lucey, and Howie Choset. Pcrnet: Point cloud registration network using pointnet encoding, 2019

  113. [121]

    Structure-from-motion revisited

    Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-motion revisited. InConference on Computer Vision and Pattern Recognition (CVPR), 2016

  114. [122]

    Pixelwise view selection for unstructured multi-view stereo

    Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. InEuropean Conference on Computer Vision (ECCV), 2016

  115. [123]

    Ari Seff, Yaniv Ovadia, Wenda Zhou, and Ryan P. Adams. Sketchgraphs: A large-scale dataset for modeling relational geometry in computer-aided design, 2020

  116. [124]

    Mvdream: Multi-view diffusion for 3d generation, 2024

    Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation, 2024

  117. [125]

    Real-time human pose recognition in parts from a single depth image

    Jamie Shotton, Andrew Fitzgibbon, Mat Cook, Toby Sharp, Mark Finocchio, Richard Moore, Alex Kipman, and Andrew Blake. Real-time human pose recognition in parts from a single depth image. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pa...

  118. [126]

    Indoor segmentation and support inference from rgb-d images

    Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgb-d images. InEuropean Conference on Computer Vision (ECCV), pages 746–760. Springer, 2012

  119. [127]

    Deep sliding shapes for amodal 3d object detection in rgb-d images

    Shuran Song and Jianxiong Xiao. Deep sliding shapes for amodal 3d object detection in rgb-d images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 808–816, 2016

  120. [128]

    Sun rgb-d: A rgb-d scene understanding benchmark suite

    Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 567–576, 2015

  121. [129]

    SpatialVerse Research Team

    Manycore Tech Inc. SpatialVerse Research Team. Interiorgs: A 3d gaussian splatting dataset of semantically labeled indoor scenes.https://huggingface.co/datasets/spatialverse/InteriorGS, 2025

  122. [130]

    Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gi...

  123. [131]

    Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022

    Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022

  124. [132]

    Habitat 2.0: Training home assistants to rearrange their habitat, 2022

    Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Ji...

  125. [133]

    Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior.arXiv preprint arXiv:2303.14184, 2023

    Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior.arXiv preprint arXiv:2303.14184, 2023

  126. [134]

    Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025

    Tencent Hunyuan3D Team. Hunyuan3d 2.5: Towards high-fidelity 3d assets generation with ultimate details, 2025

  127. [135]

    Deepv2d: Video to depth with differentiable structure from motion

    Zachary Teed and Jia Deng. Deepv2d: Video to depth with differentiable structure from motion. InInternational Conference on Learning Representations (ICLR), 2020

  128. [136]

    Consistent view synthesis with pose-guided diffusion models, 2023

    Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models, 2023

  129. [137]

    Zippered polygon meshes from range images

    Greg Turk and Marc Levoy. Zippered polygon meshes from range images. InProceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques, page 311–318, New York, NY, USA, 1994. Association for Computing Machinery

  130. [138]

    Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data, 2019

    Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data, 2019

  131. [139]

    Recovering accurate 3d human pose in the wild using imus and a moving camera

    Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. InProceedings of the European Conference on Computer Vision (ECCV), pages 601–617, 2018

  132. [140]

    Wan: Open and advanced large-scale video generative models

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pande...

  133. [141]

    Vggt: Visual geometry grounded transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Visual geometry grounded transformer. 2025

  134. [142]

    O-cnn: Octree-based convolutional neural networks for 3d shape analysis.ACM TOG, 36(4):1–11, 2017

    Peng-Shuai Wang, Yang Liu, Yueshan Guo, Chun-Yu Sun, and Xiao Tong. O-cnn: Octree-based convolutional neural networks for 3d shape analysis.ACM TOG, 36(4):1–11, 2017. 17

  135. [143]

    Moge-2: Accurate monocular geometry with metric scale and sharp details, 2025

    Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. Moge-2: Accurate monocular geometry with metric scale and sharp details, 2025

  136. [144]

    DUSt3R: Geometric 3d vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jérôme Revaud. DUSt3R: Geometric 3d vision made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20697–20709, 2024

  137. [145]

    Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, and Tong He.π3: Scalable permutation-equivariant visual geometry learning, 2025

  138. [146]

    Pointramba: A hybrid transformer-mamba framework for point cloud analysis, 2024

    Zicheng Wang, Zhenghao Chen, Yiming Wu, Zhen Zhao, Luping Zhou, and Dong Xu. Pointramba: A hybrid transformer-mamba framework for point cloud analysis, 2024

  139. [147]

    Truncated signed distance function: Experiments on voxel size

    Diana Werner, Ayoub Al-Hamadi, and Philipp Werner. Truncated signed distance function: Experiments on voxel size. InImage Analysis and Recognition, pages 357–364, Cham, 2014. Springer International Publishing

  140. [148]

    Karl D. D. Willis, Yewen Pu, Jieliang Luo, Hang Chu, Tao Du, Joseph G. Lambourne, Armando Solar-Lezama, and Wojciech Matusik. Fusion 360 gallery: A dataset and environment for programmatic cad construction from human design sequences.ACM Transactions on Graphics (TOG), 40(4), 2021

  141. [149]

    4d gaussian splatting for real-time dynamic scene rendering, 2024

    Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering, 2024

  142. [150]

    Geometry forcing: Marrying video diffusion and 3d representation for consistent world modeling, 2025

    Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, and Jiang Bian. Geometry forcing: Marrying video diffusion and 3d representation for consistent world modeling, 2025

  143. [151]

    Deepcad: A deep generative network for computer-aided design models

    Rundi Wu, Chang Xiao, and Changxi Zheng. Deepcad: A deep generative network for computer-aided design models. InICCV, pages 6762–6772, 2021

  144. [152]

    Barron, and Aleksander Holynski

    Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, and Aleksander Holynski. Cat4d: Create anything in 4d with multi-view video diffusion models, 2024

  145. [153]

    Point transformer v2: Grouped vector attention and partition-based pooling, 2022

    Xiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu, and Hengshuang Zhao. Point transformer v2: Grouped vector attention and partition-based pooling, 2022

  146. [154]

    Point transformer v3: Simpler faster stronger

    Xiaoyang Wu, Li Jiang, Peng-Shuai Wang, Zhijian Liu, Xihui Liu, Yu Qiao, Wanli Ouyang, Tong He, and Hengshuang Zhao. Point transformer v3: Simpler faster stronger. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4840–4851, 2024

  147. [155]

    3d shapenets: A deep representation for volumetric shapes

    Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. InCVPR, pages 1912–1920, 2015

  148. [156]

    Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024

    Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024

  149. [157]

    Structured 3d latents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506, 2024

    Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. Structured 3d latents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506, 2024

  150. [158]

    Sun3d: A database of big spaces reconstructed using sfm and object labels

    Jianxiong Xiao, Andrew Owens, and Antonio Torralba. Sun3d: A database of big spaces reconstructed using sfm and object labels. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 1625–1632, 2013

  151. [159]

    Instantmesh: Effi- cient 3d mesh generation from a single image with sparse-view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Effi- cient 3d mesh generation from a single image with sparse-view large reconstruction models.arXiv preprint arXiv:2404.07191, 2024

  152. [160]

    Pointllm: Empowering large language models to understand point clouds.arXiv preprint arXiv:2308.16966, 2023

    Runsen Xu, Xiaojuan Wang, Tai Wang, Kai Chen, Jiangmiao Pang, and Dahua Lin. Pointllm: Empowering large language models to understand point clouds.arXiv preprint arXiv:2308.16966, 2023

  153. [161]

    Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction

    Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020

  154. [162]

    Teaser: Fast and certifiable point cloud registration.IEEE Transactions on Robotics, 37(2):314–333, 2021

    Heng Yang, Jingnan Shi, and Luca Carlone. Teaser: Fast and certifiable point cloud registration.IEEE Transactions on Robotics, 37(2):314–333, 2021. 18

  155. [163]

    Sam 3d body: Robust full-body human mesh recovery.arXiv preprint, 2025

    Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Dollar, and Kris Kitani. Sam 3d body: Robust full-body human mesh recovery.arXiv preprint, 2025

  156. [164]

    Blendedmvs: A large-scale dataset for generalized multi-view stereo networks, 2020

    Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks, 2020

  157. [165]

    No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images, 2024

    Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images, 2024

  158. [166]

    Scannet++: A high-fidelity dataset of 3d indoor scenes, 2023

    Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d indoor scenes, 2023

  159. [167]

    Plenoctrees for real-time rendering of neural radiance fields, 2021

    Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields, 2021

  160. [168]

    Point-bert: Pre-training 3d point cloud transformers with masked point modeling

    Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19291–19300, 2022

  161. [169]

    Strobenet: Category-level multiview reconstruction of articulated objects, 2021

    Ge Zhang, Or Litany, Srinath Sridhar, and Leonidas Guibas. Strobenet: Category-level multiview reconstruction of articulated objects, 2021

  162. [170]

    Shuming Zhang, Zhidong Guan, Hao Jiang, Tao Ning, Xiaodong Wang, and Pingan Tan. Brep2seq: a dataset and hierarchical deep learning network for reconstruction and generation of computer-aided design models.Journal of Computational Design and Engineering, 11(1):110–134, 2024

  163. [171]

    3d-scenedreamer: Text-driven 3d-consistent scene generation

    Songchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma, Wei Hua, Hujun Bao, Weiwei Xu, and Changqing Zou. 3d-scenedreamer: Text-driven 3d-consistent scene generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10170–10180, 2024

  164. [172]

    Point cloud mamba: Point cloud learning via state space model, 2024

    Tao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang, Qianyu Zhou, Shunping Ji, Shuicheng Yan, and Xiangtai Li. Point cloud mamba: Point cloud learning via state space model, 2024

  165. [173]

    Microsoft kinect sensor and its effect.IEEE Multimedia, 19(2):4–10, 2012

    Zhengyou Zhang. Microsoft kinect sensor and its effect.IEEE Multimedia, 19(2):4–10, 2012

  166. [174]

    Point transformer, 2021

    Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun. Point transformer, 2021

  167. [175]

    Torr, and Vladlen Koltun

    Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H.S. Torr, and Vladlen Koltun. Point transformer. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 16259–16268, 2021

  168. [176]

    Structured3d: A large photo-realistic dataset for structured 3d modeling, 2020

    Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling, 2020

  169. [177]

    Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J

    Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J. Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking, 2023

  170. [178]

    A comprehensive review of vision-based 3d reconstruction methods.Sensors, 24(7), 2024

    Linglong Zhou, Guoxin Wu, Yunbo Zuo, Xuanyu Chen, and Hongle Hu. A comprehensive review of vision-based 3d reconstruction methods.Sensors, 24(7), 2024

  171. [179]

    Thingi10k: A dataset of 10,000 3d-printing models.arXiv preprint arXiv:1605.04797, 2016

    Qingnan Zhou and Alec Jacobson. Thingi10k: A dataset of 10,000 3d-printing models.arXiv preprint arXiv:1605.04797, 2016

  172. [180]

    Stereo magnification: Learning view synthesis using multiplane images, 2018

    Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images, 2018

  173. [181]

    H3wb: Human3.6m 3d wholebody dataset and benchmark

    Yue Zhu, Nermin Samet, and David Picard. H3wb: Human3.6m 3d wholebody dataset and benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 20166–20177, 2023. 19

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.