REVIEW 3 major objections 4 minor 205 references
From 2D to 3D Cognition: A Brief Survey of General World Models
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that world models are evolving from 2D video prediction toward systems that perceive, reason, and interact in 3D, and it supplies a framework that organizes the field around that transition.
desk verdict A useful organizing survey of 3D world models, but the 'physical' qualifier in its capability triad is applied unevenly and needs tightening before it can serve as a reliable taxonomy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework rests on two pillars and a triad of capabilities. The first pillar is explicit 3D representation—volumetric or surface-based formats such as neural radiance fields (NeRF, a function mapping position and viewing direction to color and density) and 3D Gaussian splatting (a radiance field of elliptical kernels rendered in real time)—that give models actual geometry rather than pixel correlations. The second pillar is world knowledge, supplied either by physics simulation or by commonsense and semantic priors drawn from large language and vision-language models. These pillars support the triad: 3D physical scene generation (static and dynamic, with physical constraints), 3D spatial reasoning (static understanding and dynamic forecasting, often by aligning point clouds or radiance fields with language models), and 3D spatial interaction (egocentric embodied action and exocentric scene editing). The framework's work is to classify and connect nearly all recent methods in the surveyed space under one structure.
What would settle it
A model that achieves physically consistent, interactive 3D behavior while using only 2D pixel-level representations and no explicit 3D structure or world-knowledge priors would undercut the claim that those pillars are necessary for 3D cognition; one could search current video-generation and reinforcement-learning benchmarks for such a counterexample.
Extended reading notes
Core claim
The central claim is that world models are undergoing a paradigm shift from 2D simulations to 3D cognitive systems that can perceive, reason, and interact with complex 3D environments. The paper establishes this by organizing recent work into a conceptual framework: advances in 3D representations (point clouds, meshes, occupancy grids, signed distance functions, neural radiance fields, and 3D Gaussian splatting) and the incorporation of world knowledge (physics simulation and priors extracted from large pretrained models) together support three cognitive capabilities—3D physical scene generation, 3D spatial reasoning, and 3D spatial interaction. The authors present this capability triad as the organizing principle for the field, tracing how current methods realize each capability and how they are deployed in embodied AI, autonomous driving, digital twin cities, and gaming/VR, then enumerate open challenges in data, modeling, and deployment.
Load-bearing premise
The framework assumes that every important method in the field can be cleanly sorted into the three capability buckets—generation, reasoning, interaction—and that this triad is the right decomposition of 3D cognition; if many methods straddle categories or a better decomposition exists, the survey's map becomes a subjective grouping rather than a natural taxonomy.
Editorial extensions
If this is right
- Future world models will likely be built directly on explicit 3D representations rather than on pixel-level video prediction, because geometry is what makes physical consistency and interaction possible.
- World knowledge in the form of physics simulation and pretrained-model priors will become a standard component, not an optional extra, in 3D scene generation, reasoning, and editing systems.
- The three capabilities—generation, reasoning, interaction—offer a shared vocabulary and evaluation lens across embodied AI, autonomous driving, digital twins, and gaming/VR, so progress in one domain can be compared with progress in another.
- The challenges the survey lists (multimodal data alignment, scalability of 3D representations, real-time instruction-to-action pipelines, deployment latency) define the concrete bottlenecks that must be solved for 3D cognitive world models to reach real-world use.
- If the framework is adopted, new systems will increasingly be described by which of the three capabilities they deliver and which pillars they lean on, making the field more amenable to systematic benchmarking.
Reading between the lines
- The perceive–think–act structure of the proposed triad mirrors a long-standing view of intelligent systems, so the framework may be better read as a way of organizing existing effort rather than a prediction of what new types of models will emerge.
- A testable consequence of the framework is that models lacking an explicit 3D representation will consistently underperform on physically interactive benchmarks compared with models that have one; someone could design a controlled comparison to check whether the pillar is truly necessary or merely convenient.
- The survey's binary division of world knowledge into physics simulation and learned priors leaves out other possible sources, such as structured databases or programmatic simulators; a richer ontology might be needed as the field grows.
- The four application domains all lean on the same triad, which suggests that progress in, say, autonomous driving occupancy forecasting could transfer to embodied manipulation—an implication worth testing directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a conceptual framework for organizing recent work on world models as they move from 2D video generation and prediction toward 3D cognitive systems. The framework identifies two foundational pillars—3D representations and world knowledge—and three cognitive capabilities: 3D physical scene generation, 3D spatial reasoning, and 3D spatial interaction. The paper reviews representative methods for each capability, with tables of static and dynamic scene generation, spatial reasoning methods, and exocentric scene manipulation, and then surveys applications in embodied AI, autonomous driving, digital twins, and gaming/VR. It concludes with challenges in data, modeling, and deployment.
Significance. The paper's main contribution is organizational: it provides a readable map of a fast-moving and fragmented literature, with useful tables that group methods by representation, priors, and supported tasks. If the proposed triad of generation, reasoning, and interaction is accepted as the right decomposition of 3D cognition, the survey gives practitioners and newcomers a convenient entry point and highlights the growing role of physical priors and foundation-model knowledge in 3D scene modeling. The strengths are the breadth of covered methods, the explicit comparison with prior surveys, and the application-oriented discussion. The paper does not derive new results or quantitative comparisons, so its value rests on the accuracy and consistency of its taxonomy; the major issues below concern places where that taxonomy overstates the physical grounding of the surveyed methods.
major comments (3)
- [Section 1.2 and Section 5.2 / Table 5] The capability "3D spatial interaction" is defined in Section 1.2 as "goal-directed, physically consistent interaction," but the majority of Table 5 entries (Instruct-NeRF2NeRF, SIn-NeRF2NeRF, CLIP-NeRF, GaussianEditor, Point'n Move, Instruct 4D-to-4D, 4D-Editor) are appearance- or geometry-editing methods driven by 2D diffusion or CLIP/DINO priors, with no physics simulation or physical consistency check. This conflates user-driven editing with physical interaction and weakens the internal consistency of the triad. Please either rename the capability to something like "3D scene manipulation" or explicitly distinguish interactive editing from physically grounded interaction and mark Table 5 entries accordingly.
- [Section 3.2.1 / Table 3] The category "Physics-regularized generation" includes GausSim and DeformGS, which are better described as physics-regularized reconstruction or simulation of observed dynamic scenes from video, not as generative models that synthesize novel physical scenes. The section title "Dynamic Scene Generation" and the broader capability "3D physical scene generation" therefore overstate what these methods do. Please clarify the reconstruction-vs-generation distinction in Section 3.2.1 and adjust the table captions or category names to avoid implying that all listed methods are generative.
- [Section 2.3.1] The sentence "the non-multimodal version of GPT-4V[2] exhibit emergent capabilities in interpreting the 3D environments through codes[13]" is internally contradictory because GPT-4V is by definition the multimodal (vision) version; the paper likely means the text-only GPT-4 model. The reference to [2] (the GPT-4 technical report) and [13] ("Sparks of AGI") should be checked against the intended claim, since this sentence is used as evidence for spatial commonsense in pretrained models.
minor comments (4)
- [Table 1] The coverage symbols in Table 1 appear corrupted in the manuscript: the legend reads "H #" and "#" for comprehensive/partial/no coverage, and the row for "Ours" shows blank symbols. Please re-render the table with standard symbols (e.g., filled, half-filled, empty circles) so that the comparison is legible.
- [Section 4.1.2] The names "LeRF" and "ReasonGronder" should be "LERF" and "ReasonGrounder" to match the cited works and the reference list.
- [Section 5.1.1] There is a typo in the sentence about ReasonGrounder: "his hierarchical supervision" should read "This hierarchical supervision."
- [Section 3.2.2] The phrase "A growling line of methods" should be "A growing line of methods."
Circularity Check
No circularity: the paper is a survey that organizes externally published methods; it makes no prediction derived from fitted parameters or from a self-citation chain.
full rationale
This paper is a literature survey, not a derivation. The central claim that world models are evolving from 2D simulation toward 3D cognitive systems with generation, reasoning, and interaction capabilities is supported by citing independent, externally published methods and by organizing them under an explicitly stated conceptual framework. No quantity is fitted and then reported as a prediction; no result is derived from equations that are themselves the target claim; and the capability triad is presented as an organizing principle, not as a theorem forced by prior work. I found no instance where a cited uniqueness theorem or author-specific prior result carries load-bearing weight: the authors do not cite their own previous work at all, and the external works cited, such as Ha and Schmidhuber's world models, JEPA, NeRF, and 3DGS, are independent evidence. The weakest scientific point is taxonomic consistency, namely whether all methods in Tables 2 through 5 genuinely instantiate physical generation and physical interaction, but that is a categorization concern, not circularity. Accordingly the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption 3D representations such as NeRF and 3D Gaussian Splatting are foundational for 3D world modeling.
- domain assumption World knowledge from physics simulation and pretrained models is a necessary complement to 3D representations.
- ad hoc to paper The triad of generation, reasoning, and interaction is the correct decomposition of 3D cognition.
Cite this review
Pith. "Pith review of From 2D to 3D Cognition: A Brief Survey of General World Models." pith.science (2026). https://pith.science/paper/ZNDZNWT6
@misc{pith2026250620134,
author = {Pith},
title = {Pith review of: From 2D to 3D Cognition: A Brief Survey of General World Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZNDZNWT6}},
note = {Machine review of arXiv:2506.20134}
}
read the original abstract
World models have garnered increasing attention in the development of artificial general intelligence (AGI), serving as computational frameworks for learning representations of the external world and forecasting future states. While early efforts focused on 2D visual perception and simulation, recent 3D-aware generative world models have demonstrated the ability to synthesize geometrically consistent, interactive 3D environments, marking a shift toward 3D spatial cognition. Despite rapid progress, the field lacks systematic analysis to categorize emerging techniques and clarify their roles in advancing 3D cognitive world models. This survey addresses this need by introducing a conceptual framework, providing a structured and forward-looking review of world models transitioning from 2D perception to 3D cognition. Within this framework, we highlight two key technological drivers, particularly advances in 3D representations and the incorporation of world knowledge, as fundamental pillars. Building on these, we dissect three core cognitive capabilities that underpin 3D world modeling: 3D physical scene generation, 3D spatial reasoning, and 3D spatial interaction. We further examine the deployment of these capabilities in real-world applications, including embodied AI, autonomous driving, digital twin, and gaming/VR. Finally, we identify challenges across data, modeling, and deployment, and outline future directions for advancing more robust and generalizable 3D world models.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, et al. 2024. Gpt-4 Technical Report. arXiv:2303.08774 [cs.CL]
arXiv 2024
-
[13]
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, et al. 2023. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv:2303.12712 [cs.CL]
arXiv 2023
-
[1]
Jad Abou-Chakra, Feras Dayoub, and Niko Sünderhauf. 2024. ParticleNeRF: A Particle-Based Encoding for Online Neural Radiance Fields. InProceedings of the 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 5975–5984
2024
-
[3]
Andrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti, and Luigi Di Stefano. 2024. LLaNA: Large Language and NeRF Assistant.Advances in Neural Information Processing Systems37 (2024), 1162–1195
2024
-
[4]
Haotian Bai, Yuanhuiyi Lyu, Lutao Jiang, Sijia Li, Haonan Lu, Xiaodong Lin, and Lin Wang. 2023. CompoNeRF: Text-Guided Multi-Object Compositional NeRF with Editable 3D Scene Layout. arXiv:2303.13843 [cs.CV] 26 Xie et al
arXiv 2023
-
[5]
Francesco Ballerini, Pierluigi Zama Ramirez, Roberto Mirabella, Samuele Salti, and Luigi Di Stefano. 2024. Connecting NeRFs Images and Text. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 866–876
2024
-
[6]
Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, et al. 2023. One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale. InProceedings of the 40th International Conference on Machine Learning (ICML). 1692–1717
2023
-
[7]
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mido Assran, and Nicolas Ballas. 2024. Revisiting Feature Prediction for Learning Visual Representations from Video.Transactions on Machine Learning Research(2024)
2024
Show all 205 references
-
[8]
Deniz A Bezgin, Aaron B Buhendwa, and Nikolaus A Adams. 2023. JAX-Fluids: A Fully-Differentiable High-Order Computational Fluid Dynamics Solver for Compressible Two-Phase Flows.Computer Physics Communications282 (2023), 108527
2023
-
[9]
Hengwei Bian, Lingdong Kong, Haozhe Xie, Liang Pan, Yu Qiao, and Ziwei Liu. 2025. DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes. InProceedings of the 13th International Conference on Learning Representations (ICLR)
2025
-
[10]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. PIQA: Reasoning about Physical Commonsense in Natural Language. InProceedings of the AAAI conference on artificial intelligence, Vol. 34. 7432–7439
2020
-
[11]
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis
-
[12]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-Shot Learners.Advances in Neural Information Processing Systems33 (2020), 1877–1901
2020
-
[14]
Junyi Cao, Shanyan Guan, Yanhao Ge, Wei Li, Xiaokang Yang, and Chao Ma. 2024. Neuma: Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics.Advances in Neural Information Processing Systems37 (2024), 65643–65669
2024
-
[15]
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin
-
[16]
Amirhosein Chahe and Lifeng Zhou. 2025. Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussians. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 1051–1060
2025
-
[17]
Chang Chen, Yi-Fu Wu, Jaesik Yoon, and Sungjin Ahn. 2024. TransDreamer: Reinforcement Learning with Transformer World Models. arXiv:2202.09481 [cs.LG]
2024 arXiv
-
[18]
Hong Chen, Xin Wang, Yuwei Zhou, Bin Huang, Yipeng Zhang, Wei Feng, et al. 2024. Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond. arXiv:2409.14993 [cs.AI]
2024
-
[19]
Hanlin Chen, Fangyin Wei, and Gim Hee Lee. 2024. ChatSplat: 3D Conversational Gaussian Splatting. arXiv:2412.00734 [cs.CV]
2024 arXiv
-
[20]
Sijin Chen, Xin Chen, Chi Zhang, Mingsheng Li, Gang Yu, Hao Fei, et al. 2024. Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning. InProceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 26428–26438
2024
-
[21]
Shizhe Chen, Ricardo Garcia, Cordelia Schmid, and Ivan Laptev. 2023. PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation. InProceedings of the 7th Annual Conference on Robot Learning (CoRL)
2023
-
[22]
Yilun Chen, Shuai Yang, Haifeng Huang, Tai Wang, Runsen Xu, Ruiyuan Lyu, Dahua Lin, and Jiangmiao Pang. 2024. Grounded 3D-LLM with Referent Tokens. arXiv:2405.10370 [cs.CV]
2024 arXiv
-
[23]
Joseph Cho, Fachrina Dewi Puspitasari, Sheng Zheng, Jingyao Zheng, Lik-Hang Lee, Tae-Ho Kim, et al. 2024. Sora as an AGI World Model? A Complete Survey on Text-to-Video Generation. arXiv:2403.05131 [cs.AI]
2024
-
[24]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, et al. 2023. Palm: Scaling language modeling with pathways.Journal of Machine Learning Research24, 240 (2023), 1–113
2023
-
[25]
Dana Cohen-Bar, Elad Richardson, Gal Metzer, Raja Giryes, and Daniel Cohen-Or. 2023. Set-the-Scene: Global-Local Training for Generating Controllable NeRF Scenes. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 2920–2929
2023
-
[26]
Erwin Coumans. 2015. Bullet Physics Simulation.ACM SIGGRAPH 2015 Courses(2015), 1
2015
-
[27]
Filipe de Avila Belbute-Peres, Kevin Smith, Kelsey Allen, Josh Tenenbaum, and J Zico Kolter. 2018. End-to-End Differentiable Physics for Learning and Control.Advances in Neural Information Processing Systems31 (2018). From 2D to 3D Cognition: A Brief Survey of General World Models 27
2018
-
[28]
Jie Deng, Wenhao Chai, Junsheng Huang, Zhonghan Zhao, Qixuan Huang, Mingyan Gao, Jianshu Guo, et al. 2024. CityCraft: A Real Crafter for 3D City Generation. arXiv:2406.04983 [cs.CV]
2024 arXiv
-
[29]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019
-
[30]
Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, et al. 2024. Understanding World or Predicting Future? A Comprehensive Survey of World Models. arXiv:2411.14499 [cs.CL]
2024
-
[31]
Xinlong Dong, Peicheng Shi, Heng Qi, Aixi Yang, and Taonian Liang. 2024. TS-BEV: BEV Object Detection Algorithm Based on Temporal-Spatial Feature Fusion.Displays84 (2024), 102814
2024
-
[32]
Bardienus P Duisterhof, Zhao Mandi, Yunchao Yao, Jia-Wei Liu, Jenny Seidenschwarz, Mike Zheng Shou, et al. 2024. DeformGS: Scene Flow in Highly Deformable Scenes for Deformable Object Manipulation.W AFR(2024)
2024
-
[33]
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. 2023. Structure and Content-Guided Video Synthesis with Diffusion Models. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 7346–7356
2023
-
[34]
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, et al. 2022. Fast Dynamic Radiance Fields with Time-Aware Neural Voxels. InSIGGRAPH Asia 2022 Conference Papers
2022
-
[35]
Yutao Feng, Xiang Feng, Yintong Shang, Ying Jiang, Chang Yu, Zeshun Zong, Tianjia Shao, Hongzhi Wu, Kun Zhou, Chenfanfu Jiang, et al. 2025. Gaussian Splashing: Unified Particles for Versatile Motion Synthesis and Rendering. In Proceedings of the 2025 IEEE/CVF Conference on Com...
2025
-
[36]
Yutao Feng, Yintong Shang, Xuan Li, Tianjia Shao, Chenfanfu Jiang, and Yin Yang. 2024. Pie-NeRF: Physics-Based Interactive Elastodynamics with NeRF. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4450–4461
2024
-
[37]
J. W. Forrester. 1971. Counterintuitive Behavior of Social Systems.Theory and Decision2, 2 (1971), 109–140
1971
-
[38]
Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, et al. 2021. ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation. InProceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS)
2021
-
[39]
Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. 2021. Dynamic View Synthesis from Dynamic Monocular Video. InProceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). 5712–5721
2021
-
[40]
Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes, Laurent Najman, Michael Rabbat, Emmanuel Dupoux, and Yann LeCun. 2025. Intuitive physics understanding emerges from self-supervised pretraining on natural videos. arXiv:2502.11831 [cs.CV]
2025 arXiv
-
[41]
Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li, Guoming Wang, Siliang Tang, and Yueting Zhuang. 2024. WorldGPT: Empowering LLM as Multimodal World Model. InProceedings of the 32nd ACM International Conference on Multimedia. 7346–7355
2024
-
[42]
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, et al. 2024. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv:2311.10709 [cs.CV]
2024 arXiv
-
[43]
Minghao Guo, Bohan Wang, Pingchuan Ma, Tianyuan Zhang, Crystal Owens, Chuang Gan, et al. 2024. Physically Compatible 3D Object Modeling from A Single Image.Advances in Neural Information Processing Systems37 (2024), 119260–119282
2024
-
[44]
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Fei-Fei Li, et al. 2024. Photorealistic Video Generation with Diffusion Models. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 393–411
2024
-
[45]
Tarun Gupta, Wenbo Gong, Chao Ma, Nick Pawlowski, Agrin Hilmkil, Meyer Scetbon, Marc Rigter, Ade Famoti, Ashley Juan Llorens, Jianfeng Gao, Stefan Bauer, Danica Kragic, Bernhard Schölkopf, and Cheng Zhang. 2024. The Essential Role of Causality in Foundation World Models for Em...
2024 arXiv
-
[46]
David Ha and Jürgen Schmidhuber. 2018. Recurrent World Models Facilitate Policy Evolution.Advances in Neural Information Processing Systems31 (2018), 1
2018
-
[47]
David Ha and Jürgen Schmidhuber. 2018. World Models. (2018)
2018
-
[48]
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and et al. 2020. Dream to Control: Learning Behaviors by Latent Imagination. InProceedings of the 37th International Conference on Machine Learning (ICML)
2020
-
[49]
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. 2021. Mastering Atari with Discrete World Models. InProceedings of the 38th International Conference on Machine Learning (ICML)
2021
-
[50]
Ayaan Haque, Matthew Tancik, Alexei A Efros, Aleksander Holynski, and Angjoo Kanazawa. 2023. Instruct- NeRF2NeRF: Editing 3D Scenes with Instructions. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 19740–19750
2023
-
[51]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022. Video Diffusion Models. 8633–8646 pages. 28 Xie et al
2022
-
[52]
Philipp Holl, Nils Thuerey, and Vladlen Koltun. 2020. Learning to Control PDEs with Differentiable Physics. In Proceedings of the 8th International Conference on Learning Representations (ICLR)
2020
-
[53]
Jiseung Hong, Changmin Lee, and Gyusang Yu. 2024. SIn-NeRF2NeRF: Editing 3D Scenes with Instructions through Segmentation and Inpainting. arXiv:2408.13285 [cs.CV]
2024 arXiv
-
[54]
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 2023. 3D-LLM: Injecting the 3D World into Large Language Models.Advances in Neural Information Processing Systems36 (2023), 20482–20494
2023
-
[55]
Howell, Simon Le Cleac’h, Jan Brüdigam, J
Taylor A. Howell, Simon Le Cleac’h, Jan Brüdigam, J. Zico Kolter, Mac Schwager, and Zachary Manchester. 2023. Dojo: A Differentiable Physics Engine for Robotics. arXiv:2203.00806 [cs.RO]
2023 arXiv
-
[56]
Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Fredo Durand. 2020. DiffTaichi: Differentiable Programming for Physical Simulation. InProceedings of the 8th International Conference on Learning Representations (ICLR)
2020
-
[57]
Yuanming Hu, Tzu-Mao Li, Luke Anderson, Jonathan Ragan-Kelley, and Frédo Durand. 2019. Taichi: A Language for High-Performance Computation on Spatially Sparse Data Structures.ACM Transactions on Graphics (TOG)38, 6 (2019), 1–16
2019
-
[58]
Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, et al. 2025. Simulating the Real World: A Unified Survey of Multimodal Generative Models. arXiv:2503.04641 [cs.CV]
2025
-
[59]
Haifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang, Runsen Xu, Tai Wang, Luping Liu, Xize Cheng, Yang Zhao, Jiangmiao Pang, and Zhou Zhao. 2024. Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers. InProceedings of the 38th Annual Conference on Ne...
2024
-
[60]
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. 2024. An Embodied Generalist Agent in 3D World. InProceedings of the 41st International Conference on Machine Learning (ICML)
2024
-
[61]
Jiajun Huang, Hongchuan Yu, Jianjun Zhang, and Hammadi Nait-Charif. 2024. Point’n Move: Interactive Scene Object Manipulation on Gaussian Splatting Radiance Fields.IET Image Processing18, 12 (2024), 3507–3517
2024
-
[62]
Kuan-Chih Huang, Xiangtai Li, Lu Qi, Shuicheng YAN, and Ming-Hsuan Yang. 2025. Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model. InProceedings of the 2025 International Conference on 3D Vision (3DV)
2025
-
[63]
Yi-Hua Huang, Yang-Tian Sun, Ziyi Yang, Xiaoyang Lyu, Yan-Pei Cao, and Xiaojuan Qi. 2024. SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4220–4230
2024
-
[64]
Tenenbaum, and Chuang Gan
Zhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou, Hao Su, Joshua B. Tenenbaum, and Chuang Gan. 2021. Plas- ticineLab: A Soft-Body Manipulation Benchmark with Differentiable Physics. arXiv:2104.03311 [cs.LG]
2021 arXiv
-
[65]
Zanming Huang, Jimuyang Zhang, and Eshed Ohn-Bar. 2024. Neural Volumetric World Models for Autonomous Driving. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 195–213
2024
-
[66]
Dadong Jiang, Zhihui Ke, Xiaobo Zhou, and Xidong Shi. 2025. 4D-Editor: Interactive Object-Level Editing in Dynamic Neural Radiance Fields via Semantic Distillation. InProceedings of the 2025 IEEE/CVF International Conference on 3D Vision (3DV)
2025
-
[67]
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, et al . 2024. Scaling Up Dynamic Human-Scene Interaction Modeling. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1737–1747
2024
-
[68]
Ying Jiang, Chang Yu, Tianyi Xie, Xuan Li, Yutao Feng, Huamin Wang, Minchen Li, et al. 2024. VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality. InProceedings of the ACM SIGGRAPH Conference on Computer Graphics and Interactive Techniques ...
2024
-
[69]
P. N. Johnson-Laird. 1980. Mental Models in Cognitive Science.Cognitive Science4, 1 (1980), 71–115
1980
-
[70]
James Kennedy and Russell Eberhart. 1995. Particle Swarm Optimization. InProceedings of the International Conference on Neural Networks (ICNN’95), Vol. 4. IEEE, 1942–1948
1995
-
[71]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 2023. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Transactions on Graphics42, 4 (July 2023)
2023
-
[72]
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. 2023. LERF: Language Embedded Radiance Fields. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (CVPR). 19729–19739
2023
-
[73]
Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. 2023. Segment Anything. InProceedings of the 2023 IEEE/CVF International Conference on Comput...
2023
-
[74]
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, and othersi. 2022. AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv:1712.05474 [cs.CV] From 2D to 3D Cognition: A Brief Survey of General World Models 29
2022 arXiv
-
[75]
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, et al. 2024. VideoPoet: A Large Language Model for Zero-Shot Video Generation. InProceedings of the 41st International Conference on Machine Learning (ICML), Vol. 235. 25105–25124
2024
-
[76]
Chen Kong, Chen-Hsuan Lin, and Simon Lucey. 2017. Using Locally Corresponding CAD Models for Dense 3D Reconstructions from A Single Image. InProceedings of the 2017 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4857–4865
2017
-
[77]
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. 2024. LISA: Reasoning Segmen- tation via Large Language Model. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[78]
Simon Le Cleac’h, Hong-Xing Yu, Michelle Guo, Taylor Howell, Ruohan Gao, Jiajun Wu, Zachary Manchester, and Mac Schwager. 2023. Differentiable Physics Simulation of Dynamics-Augmented Neural Objects.IEEE Robotics and Automation Letters8, 5 (2023), 2780–2787
2023
-
[79]
2022.A Path Towards Autonomous Machine Intelligence Version 0.9.2
Yann LeCun. 2022.A Path Towards Autonomous Machine Intelligence Version 0.9.2. Tech. Rep. 62(1). OpenReview. 1–62 pages. Version 0.9.2, published June 27, 2022
2022
-
[80]
Hongjie Li, Hong-Xing Yu, Jiaman Li, and Jiajun Wu. 2025. ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation. arXiv:2412.18600 [cs.CV]
2025 arXiv
-
[81]
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. 2024. Controllable Human-Object Interaction Synthesis. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 54–72
2024
-
[82]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. InProceedings of the 40th International Conference on Machine Learning (ICML)
2023
-
[83]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation. InProceedings of the 39th International Conference on Machine Learning (ICML)
2022
-
[84]
Xuan Li, Yi-Ling Qiao, Peter Yichen Chen, Krishna Murthy Jatavallabhula, Ming Lin, Chenfanfu Jiang, and Chuang Gan. 2023. PAC-NeRF: Physics Augmented Continuum Neural Radiance Fields for Geometry-Agnostic System Identification. InProceedings of the 11th International Conferenc...
2023
-
[85]
Yixun Liang, Xin Yang, Jiantao Lin, Haodong Li, Xiaogang Xu, and Yingcong Chen. 2024. LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6517–6526
2024
-
[86]
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, et al. 2023. Magic3D: High- Resolution Text-to-3D Content Creation. InProceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[87]
Jiajing Lin, Zhenzhong Wang, Yongjie Hou, Yuzhou Tang, and Min Jiang. 2024. Phy124: Fast Physics-Driven 4D Content Generation from a Single Image. arXiv:2409.07179 [cs.CV]
2024 arXiv
-
[88]
Minghui Lin, Xiang Wang, Yishan Wang, Shu Wang, Fengqi Dai, Pengxiang Ding, et al. 2025. Exploring the Evolution of Physics Cognition in Video Generation: A Survey. arXiv:2503.21765 [cs.CV]
2025 arXiv
-
[89]
Tenenbaum, and Chuang Gan
Xingyu Lin, Zhiao Huang, Yunzhu Li, David Held, Joshua B. Tenenbaum, and Chuang Gan. 2022. DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools. InProceedings of the 10th International Conference on Learning Representations (ICLR)
2022
-
[90]
Yuchen Lin, Chenguo Lin, Jianjin Xu, and Yadong MU. 2025. OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation. InProceedings of the 13th International Conference on Learning Representations (ICLR)
2025
-
[91]
Yiqi Lin, Hao Wu, Ruichen Wang, Haonan Lu, Xiaodong Lin, Hui Xiong, and Lin Wang. 2023. Towards Language- guided Interactive 3D Generation: LLMs as Layout Interpreter with Generative Feedback. arXiv:2305.15808 [cs.CV]
2023 arXiv
-
[92]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning.Advances in Neural Information Processing Systems36 (2023), 34892–34916
2023
-
[93]
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023. Zero-1-to-3: Zero-Shot One Image to 3D Object. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 9298–9309
2023
-
[94]
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. 2024. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. InProceedings of the 2024 European Confer...
2024
-
[95]
Yang Liu, Chuanchen Luo, Lue Fan, Naiyan Wang, Junran Peng, and Zhaoxiang Zhang. 2024. CityGaussian: Real-time High-Quality Large-Scale Scene Rendering with Gaussians. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 265–282. 30 Xie et al
2024
-
[96]
Yili Liu, Linzhan Mou, Xuan Yu, Chenrui Han, Sitong Mao, Rong Xiong, and Yue Wang. 2024. Let Occ Flow: Self-Supervised 3D Occupancy Flow Prediction. InProceedings of the 8th Annual Conference on Robot Learning (CoRL)
2024
-
[97]
Zhenyang Liu, Yikai Wang, Sixiao Zheng, Tongying Pan, Longfei Liang, Yanwei Fu, and Xiangyang Xue. 2025. ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning. InProceedings of the 2025 IEEE/CVF Conference on Computer ...
2025
-
[98]
TaiMing Lu, Tianmin Shu, Alan Yuille, Daniel Khashabi, and Jieneng Chen. 2025. GenEx: Generating an Explorable World. InProceedings of the 13th International Conference on Learning Representations (ICLR)
2025
-
[99]
Xiaojian Ma, Silong Yong, Zilong Zheng, Qing Li, Yitao Liang, Song-Chun Zhu, and Siyuan Huang. 2023. SQA3D: Situ- ated Question Answering in 3D Scenes. InProceedings of the 11th International Conference on Learning Representations (ICLR)
2023
-
[100]
Miles Macklin. 2022. Warp: A High-Performance Python Framework for GPU Simulation and Graphics. InNVIDIA GPU Technology Conference (GTC). https://github.com/nvidia/warp
2022
-
[101]
Miles Macklin, Matthias Müller, and Nuttapong Chentanez. 2016. XPBD: Position-Based Simulation of Compliant Constrained Dynamics. InProceedings of the 9th International Conference on Motion in Games. 49–54
2016
-
[102]
Xinji Mai, Zeng Tao, Junxiong Lin, Haoran Wang, Yang Chang, Yanlan Kang, et al. 2024. From Efficient Multimodal Models to World Models: A Survey. arXiv:2407.00118 [cs.LG]
2024 arXiv
-
[103]
Haotian Mao, Zhuoxiong Xu, Siyue Wei, Yule Quan, Nianchen Deng, and Xubo Yang. 2025. LIVE-GS: LLM Powers Interactive VR by Enhancing Gaussian Splatting. InProceedings of the 2025 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). 1234–1235
2025
-
[104]
Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng, Rui Tang, Hao Zhu, Ping Tan, and Zihan Zhou. 2025. SpatialLM: Training Large Language Models for Structured Indoor Modeling. arXiv:2506.07491 [cs.CV]
2025
-
[105]
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. 2019. Occupancy Networks: Learning 3D Reconstruction in Function Space. InProceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4460–4470
2019
-
[106]
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. 2023. Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures. InProceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12663–12673
2023
-
[107]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. NeRF: Representing Scenes As Neural Radiance Fields for View Synthesis.Commun. ACM65, 1 (2021), 99–106
2021
-
[108]
Linzhan Mou, Jun-Kun Chen, and Yu-Xiong Wang. 2024. Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20176–20185
2024
-
[109]
Matthias Müller, Bruno Heidelberger, Marcus Hennix, and John Ratcliff. 2007. Position Based Dynamics.Journal of Visual Communication and Image Representation18, 2 (2007), 109–118
2007
-
[110]
Toan Nguyen, Minh Nhat Vu, Baoru Huang, Tuan Van Vo, Vy Truong, Ngan Le, et al. 2024. Language-Conditioned Affordance-Pose Detection in 3D Point Clouds. InProceedings of 2024 IEEE International Conference on Robotics and Automation (ICRA). 3071–3078
2024
-
[111]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2022. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. InProceedings of the 39th International Conference on Mac...
2022
-
[112]
Başak Melis Öcal, Maxim Tatarchenko, Sezer Karaoğlu, and Theo Gevers. 2024. SceneTeller: Language-to-3D Scene Generation. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 362–378
2024
-
[113]
Odyssey. 2024. World Models for Film, Gaming, and Beyond. Online. https://odyssey.world/introducing-explorer Accessed: 2025-04-28
2024
-
[114]
OpenAI. 2024. Sora: Creating Video from Text. Online. https://openai.com/research/sora Accessed: 2025-04-28
2024
-
[115]
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. 2019. DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation. InProceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 165–174
2019
-
[116]
Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. 2021. Nerfies: Deformable Neural Radiance Fields. InProceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV). 5865–5874
2021
-
[117]
Jack Parker-Holder and et al. 2024. Genie 2: A Large -Scale Foundation World Model. Online. https://deepmind.com/ research/publications/2024/genie-2-a-large-scale-foundation-world-model Accessed: 2025-04-28
2024
-
[118]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. 2019. Expressive Body Capture: 3D Hands, Face, and Body from A Single Image. InProceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern ...
2019
-
[119]
Ryan Po and Gordon Wetzstein. 2024. Compositional 3D Scene Generation Using Locally Conditioned Diffusion. In Proceedings of the 2024 International Conference on 3D Vision (3DV). IEEE, 651–663
2024
-
[120]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. InProceedings of the 12th International Conference on Learning Representatio...
2024
-
[121]
Barron, and Ben Mildenhall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2023. DreamFusion: Text-to-3D using 2D Diffusion. In Proceedings of the 11th International Conference on Learning Representations (ICLR)
2023
-
[122]
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. 2021. D-NeRF: Neural Radiance Fields for Dynamic Scenes. InProceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10318–10327
2021
-
[123]
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. 2017. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. InProceedings of the 2017 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 652–660
2017
-
[124]
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017. PointNet++: Deep Hierarchical Feature Learning on Point Sets in A Metric Space.Advances in Neural Information Processing Systems30 (2017)
2017
-
[125]
Zhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu, Tong Wu, Jiaqi Wang, Dahua Lin, and Hengshuang Zhao. 2024. Gpt4point: A Unified Framework for Point-Language Understanding and Generation. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2024
-
[126]
Yi-Ling Qiao, Alexander Gao, and Ming Lin. 2022. NeuPhysics: Editable Neural Geometry and Physics from Monocular Videos.Advances in Neural Information Processing Systems35 (2022), 12841–12854
2022
-
[127]
Yi-Ling Qiao, Junbang Liang, Vladlen Koltun, and Ming C Lin. 2020. Scalable Differentiable Physics for Learning and Control. InProceedings of the 37th International Conference on Machine Learning (ICML)
2020
-
[128]
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. 2024. LangSplat: 3D Language Gaussian Splatting. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20051–20060
2024
-
[129]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, et al. 2021. Learning Transferable Visual Models from Natural Language Supervision. InProceedings of the 38th International Conference on Machine Learning (ICML). 8748–8763
2021
-
[130]
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving Language Understanding by Generative Pre-Training. (2018)
2018
-
[131]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language Models are Unsupervised Multitask Learners.OpenAI blog1, 8 (2019), 9
2019
-
[132]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever
-
[133]
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang. 2024. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. ...
2024 arXiv
-
[134]
Jan Robine, Marc Höftmann, Tobias Uelwer, and Stefan Harmeling. 2023. Transformer-based World Models are Happy with 100k Interactions. InProceedings of the 11th International Conference on Learning Representations (ICLR)
2023
-
[135]
InProceedings of the 38th International Conference on Machine Learning (ICML)
Zero-Shot Text-to-Image Generation. InProceedings of the 38th International Conference on Machine Learning (ICML). 8821–8831
-
[136]
2016.Artificial Intelligence: A Modern Approach
Stuart J Russell and Peter Norvig. 2016.Artificial Intelligence: A Modern Approach. pearson
2016
-
[137]
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, et al. 2019. Habitat: A platform for embodied ai research. InProceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV). 9339–9347
2019
-
[138]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. InProceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[139]
Yidi Shao, Mu Huang, Chen Change Loy, and Bo Dai. 2025. GausSim: Foreseeing Reality by Gaussian Simulator for Elastic Objects. arXiv:2412.17804 [cs.CV]
2025 arXiv
-
[140]
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao-Hua Guan. 2024. Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5333–5343
2024
-
[141]
Yu Shang, Yuming Lin, Yu Zheng, Hangyu Fan, Jingtao Ding, Jie Feng, Jiansheng Chen, Li Tian, and Yong Li. 2024. UrbanWorld: An Urban World Model for 3D City Generation. arXiv:2407.11965 [cs.CV]
2024 arXiv
-
[142]
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2023. Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation. InProceedings of the 7th Annual Conference on Robot Learning (CoRL). PMLR, 785–799
2023
-
[143]
Deborah Sulsky, Zhen Chen, and Howard L Schreyer. 1994. A Particle Method for History-Dependent Materials. Computer Methods in Applied Mechanics and Engineering118, 1-2 (1994), 179–196
1994
-
[144]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). ...
2020
-
[145]
Aether Team, Haoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang, Yang Zhou, Zizun Li, et al. 2025. Aether: Geometric-Aware Unified World Modeling. arXiv:2503.18945 [cs.CV]
2025 arXiv
-
[146]
Anh Thai, Songyou Peng, Kyle Genova, Leonidas Guibas, and Thomas Funkhouser. 2025. SplatTalk: 3D VQA with Gaussian Splatting. arXiv:2503.06271 [cs.CV]
2025 arXiv
-
[147]
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2024. DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation. InProceedings of the 12th International Conference on Learning Representations (ICLR)
2024
-
[148]
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012. MuJoCo: A physics engine for model-based control. InProceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 5026–5033
2012
-
[149]
Wenwen Tong, Chonghao Sima, Tai Wang, Li Chen, Silei Wu, Hanming Deng, et al. 2023. Scene As Occupancy. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 8406–8415
2023
-
[150]
Xiaoyu Tian, Tao Jiang, Longfei Yun, Yucheng Mao, Huitong Yang, Yue Wang, Yilun Wang, and Hang Zhao. 2023. Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving.Advances in Neural Information Processing Systems36 (2023), 64318–64330
2023
-
[151]
Simeng Tu, Xiaowei Zhou, Dehong Liang, Xiaojie Jiang, Yifan Zhang, Xiaojun Li, and Xiaodong Bai. 2025. The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey. arXiv:2502.10498 [cs.CV]
2025
-
[152]
Nico Uhlemann, Felix Fent, and Markus Lienkamp. 2023. Evaluating Pedestrian Trajectory Prediction Methods with Respect to Autonomous Driving.Computing Research Repository (CoRR)(2023)
2023
-
[153]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288 [cs.CL]
2023 arXiv
-
[154]
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. 2022. CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields. InProceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3835–3844
2022
-
[155]
Junjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, and Qi Tian. 2024. GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20902–20911
2024
-
[156]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019. Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748 [cs.LG]
2019 arXiv
-
[157]
Tenenbaum, et al
Tsun-Hsuan Wang, Pingchuan Ma, Andrew Everett Spielberg, Zhou Xian, Hao Zhang, Joshua B. Tenenbaum, et al
-
[158]
Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. 2022. HUMANISE: Language- conditioned Human Motion Generation in 3D Scenes.Advances in Neural Information Processing Systems35 (2022), 14959–14971
2022
-
[159]
Lening Wang, Wenzhao Zheng, Yilong Ren, Han Jiang, Zhiyong Cui, Haiyang Yu, and Jiwen Lu. 2024. OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving. arXiv:2405.20337 [cs.CV]
2024 arXiv
-
[160]
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. 2023. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation.Advances in Neural Information Processing Systems36 (2023), 8406–8441
2023
-
[161]
InProceedings of the 11th International Conference on Learning Representations (ICLR)
SoftZoo: A Soft Robot Co-design Benchmark For Locomotion In Diverse Environments. InProceedings of the 11th International Conference on Learning Representations (ICLR)
-
[162]
Julong Wei, Shanshuai Yuan, Pengfei Li, Qingda Hu, Zhongxue Gan, and Wenchao Ding. 2024. OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving. arXiv:2409.03272 [cs.CV]
2024 arXiv
-
[163]
Zehan Wang, Haifeng Huang, Yang Zhao, Ziang Zhang, and Zhou Zhao. 2023. Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes. arXiv:2308.08769 [cs.CV]
2023 arXiv
-
[164]
Keenon Werling, Dalton Omens, Jeongseok Lee, Ioannis Exarchos, and C Karen Liu. 2021. Fast and Feature-Complete Differentiable Physics for Articulated Rigid Bodies with Contact. InProceedings of the 17th Robotics: Science and Systems
2021
-
[165]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, et al. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.Advances in neural information processing systems35 (2022), 24824–24837
2022
-
[166]
Wayne Wu, Honglin He, Jack He, Yiran Wang, Chenda Duan, Zhizheng Liu, Quanyi Li, and Bolei Zhou. 2025. MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility. InProceedings of the 13th International Conference on Learning Representations (ICLR)
2025
-
[167]
Yi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023. SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Driving. InProceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 21729–21740
2023
-
[168]
2024.Genesis: A Universal and Generative Physics Engine for Robotics and Beyond
Zhou Xian, Qiao Yiling, Xu Zhenjia, Wang Tsun-Hsuan, Chen Zhehuan, Zheng Juntian, et al . 2024.Genesis: A Universal and Generative Physics Engine for Robotics and Beyond. https://github.com/Genesis-Embodied-AI/Genesis
2024
-
[169]
Jiajun Wu, Chengkai Zhang, Xiuming Zhang, Zhoutong Zhang, William T Freeman, and Joshua B Tenenbaum. 2018. Learning Shape Priors for Single-View 3D Completion and Reconstruction. InProceedings of the 2018 European From 2D to 3D Cognition: A Brief Survey of General World Models...
2018
-
[170]
Haozhe Xie, Zhaoxi Chen, Fangzhou Hong, and Ziwei Liu. 2024. CityDreamer: Compositional Generative Model of Unbounded 3D Cities. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 9666–9675
2024
-
[171]
Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, and Shenlong Wang. 2024. Video2Game: Real-Time Interactive Realistic and Browser-Compatible Environment from a Single Video. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4578–4588
2024
-
[172]
Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. 2024. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4389–4398
2024
-
[173]
Zhou Xian, Bo Zhu, Zhenjia Xu, Hsiao-Yu Tung, Antonio Torralba, Katerina Fragkiadaki, and Chuang Gan. 2023. FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation. InProceedings of the 11th International Conference on Learning Representations (ICLR)
2023
-
[174]
Jie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos, Wojciech Matusik, Animesh Garg, and Miles Macklin
-
[175]
Haozhe Xie, Zhaoxi Chen, Fangzhou Hong, and Ziwei Liu. 2025. CityDreamer4D: Compositional Generative Model of Unbounded 4D Cities. arXiv:2501.08983 [cs.CV]
2025 arXiv
-
[176]
Runsen Xu, Xiaolong Wang, Tai Wang, Yilun Chen, Jiangmiao Pang, and Dahua Lin. 2024. PointLLM: Empowering Large Language Models to Understand Point Clouds. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 131–147
2024
-
[177]
Huaiyuan Xu, Junliang Chen, Shiyu Meng, Yi Wang, and Lap-Pui Chau. 2025. A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective.Information Fusion114 (2025), 102671
2025
-
[178]
Zetong Yang, Li Chen, Yanan Sun, and Hongyang Li. 2024. Visual Point Cloud Forecasting Enables Scalable Au- tonomous Driving. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14673–14684
2024
-
[179]
Hong-Xing Yu, Haoyi Duan, Charles Herrmann, William T Freeman, and Jiajun Wu. 2025. WonderWorld: Interactive 3D Scene Generation from a Single Image. InProceedings of the 2025 IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR). 5916–5926
2025
-
[180]
Qingshan Xu, Jiao Liu, Melvin Wong, Caishun Chen, and Yew-Soon Ong. 2024. Precise-Physics Driven Text-to-3D Generation. arXiv:2403.12438 [cs.CV]
2024 arXiv
-
[181]
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, et al. 2024. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. InProceedings of the 12th International Conference on Learning Representations (ICLR)
2024
-
[182]
Yandan Yang, Baoxiong Jia, Peiyuan Zhi, and Siyuan Huang. 2024. PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16262–16272
2024
-
[183]
Chubin Zhang, Juncheng Yan, Yi Wei, Jiaxin Li, Li Liu, Yansong Tang, Yueqi Duan, and Jiwen Lu. 2024. OccNeRF: Self-Supervised Multi-Camera Occupancy Prediction with Neural Radiance Fields. arXiv:2312.09243 [cs.CV]
2024 arXiv
-
[184]
Haiming Zhang, Ying Xue, Xu Yan, Jiacheng Zhang, Weichao Qiu, Dongfeng Bai, Bingbing Liu, Shuguang Cui, and Zhen Li. 2024. An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training. arXiv:2412.13772 [cs.CV]
2024 arXiv
-
[185]
Hong-Xing Yu, Haoyi Duan, Junhwa Hur, Kyle Sargent, Michael Rubinstein, William T Freeman, et al. 2024. Wonder- Journey: Going from Anywhere to Everywhere. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6658–6667
2024
-
[186]
Qihang Zhang, Chaoyang Wang, Aliaksandr Siarohin, Peiye Zhuang, Yinghao Xu, Ceyuan Yang, Dahua Lin, Bolei Zhou, Sergey Tulyakov, and Hsin-Ying Lee. 2024. Towards Text-Guided 3D Scene Composition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2024
-
[187]
Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. 2024. Make Pixels Dance: High-Dynamic Video Generation. InProceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8850–8860
2024
-
[188]
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. 2021. Point Transformer. InProceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 16259–16268
2021
-
[189]
Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, and Hua Zou. 2025. PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting. arXiv:2411.12789 [cs.CV]
2025 arXiv
-
[190]
Lunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas, Rui Hu, and Raquel Urtasun. 2024. Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion. InProceedings of the 12th International Conference on Learning Representations (ICLR). 34 Xie et al
2024
-
[191]
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 2024. 3D-VLA: A 3D Vision-Language-Action Generative World Model. InProceedings of the 41st International Conference on Machine Learning (ICML)
2024
-
[192]
Shougao Zhang, Mengqi Zhou, Yuxi Wang, Chuanchen Luo, Rongyu Wang, Yiwei Li, Zhaoxiang Zhang, and Junran Peng. 2024. CityX: Controllable Procedural Content Generation for Unbounded 3D Cities.Computing Research Repository (CoRR)(2024)
2024
-
[193]
Licheng Zhong, Hong-Xing Yu, Jiajun Wu, and Yunzhu Li. 2024. Reconstruction and Simulation of Elastic Objects with Spring-Mass 3D Gaussians. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 407–423
2024
-
[194]
Mengqi Zhou, Yuxi Wang, Jun Hou, Shougao Zhang, Yiwei Li, Chuanchen Luo, Junran Peng, and Zhaoxiang Zhang
-
[195]
Junhui Zhao, Jingyue Shi, and Li Zhuo. 2024. BEV Perception for Autonomous Driving: State of The Art and Future Perspectives.Expert Systems with Applications258 (2024), 125103
2024
-
[196]
Xiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei, Nan Dong, and Ming-Hsuan Yang. 2025. OccGS: Zero-shot 3D Occupancy Reconstruction with Semantic and Geometric-Aware Gaussian Splatting. arXiv:2502.04981 [cs.CV]
2025 arXiv
-
[197]
Wenzhao Zheng, Weiliang Chen, Yuanhui Huang, Borui Zhang, Yueqi Duan, and Jiwen Lu. 2024. OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving. InProceedings of the 2024 European Conference on Computer Vision (ECCV). 55–72
2024
-
[198]
2005.The Finite Element Method for Solid and Structural Mechanics
Olgierd Cecil Zienkiewicz and Robert Leroy Taylor. 2005.The Finite Element Method for Solid and Structural Mechanics. Elsevier
2005
-
[201]
Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhiwei Lin, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang
-
[204]
Yang Zhou, Zongjin He, Qixuan Li, and Chao Wang. 2025. LayoutDreamer: Physics-Guided Layout for Text-to-3D Compositional Scene Generation. arXiv:2502.01949 [cs.CV]
2025
-
[2021]
InProceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV)
Emerging Properties in Self-Supervised Vision Transformers. InProceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[2022]
InProceedings of the 10th International Conference on Learning Representations (ICLR)
Accelerated Policy Learning with Parallel Differentiable Simulation. InProceedings of the 10th International Conference on Learning Representations (ICLR)
-
[2023]
InProceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. InProceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 22563–22575
2023
-
[2024]
In Proceedings of the 41st International Conference on Machine Learning (ICML)
Gala3D: Towards Text-to-3D Complex Scene Generation via Layout-Guided Generative Gaussian Splatting. In Proceedings of the 41st International Conference on Machine Learning (ICML)
-
[2025]
InProceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vol
SceneX: Procedural Controllable Large -Scale Scene Generation. InProceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vol. 39. 10806–10814
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.