REVIEW 3 major objections 3 minor 12 cited by
This survey proposes that large-model-empowered embodied AI is best understood through two decision-making paradigms—hierarchical and end-to-end vision-language-action models—with imitation learning and reinforcement learning as the learnin
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A useful survey skeleton with a healthy world-models section, but the citation chain is broken in several places, which guts the reference-traceability function in its current form. the 3 major comments →
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that recent large-model-empowered embodied systems form a structured landscape rather than a scattered collection. The landscape has two decision-making paradigms: hierarchical, where large language models handle planning and separate modules handle execution and feedback, and end-to-end, where vision-language-action models map observations and instructions directly to action tokens. It has two learning engines, imitation learning and reinforcement learning, each enhanced by large models through diffusion or Transformer policies and through automated reward design. The survey further claims, for the first time in this line of surveys, that world models are a fourth pilla
What carries the argument
The organizing device is a taxonomy with three axes: decision-making paradigm (hierarchical versus end-to-end), embodied learning method (imitation versus reinforcement), and world-model design (latent-state, Transformer-based, diffusion-based, and joint-embedding predictive architectures). The named central object is the vision-language-action model, defined as a single network that maps multimodal observations and language instructions directly to action tokens; it anchors the end-to-end paradigm. The survey pairs this grid with a dual analytical approach that compares approaches horizontally across paradigms and traces core models vertically through their development, which is what carrie
Load-bearing premise
The survey is useful as a map only if the reader can trust that each bracket number points to the paper actually supporting that sentence; a few markers do not line up (TinyVLA is cited through the RT-2 reference, DiffusionVLA through the Sora reference, and DEPS is assigned two different numbers), so a citation audit is the load-bearing check.
What would settle it
Open the survey at the load-bearing citations and verify them against the primary literature: Section 4.2.3's claim about TinyVLA should lead to the TinyVLA paper, not RT-2 [234], and Section 7.3's claim about DiffusionVLA should lead to DiffusionVLA, not Sora [22]. If those markers mispoint, the map cannot be used to reach the cited methods.
If this is right
- A researcher encountering a new embodied system can classify it by paradigm and learning method and immediately see which prior work it builds on and which comparison table applies.
- The hierarchical-versus-end-to-end comparison gives practitioners explicit selection criteria: interpretability and reliability versus generalization and speed, at different computational costs.
- Treating world models as a first-class component turns learned simulators into a standard tool for both decision validation and data generation in embodied learning.
- The four open challenges identified by the survey, embodied data scarcity, continual learning, deployment efficiency, and the sim-to-real gap, become a shared priority list with world models and large-model reward design as candidate remedies.
Where Pith is reading between the lines
- If the taxonomy holds, a useful next step would be to mine the literature for cells that are still empty, such as world-model-augmented reinforcement learning on physical robots, since the survey's own tables suggest those gaps.
- The survey's comparison of hierarchical and end-to-end paradigms implies a hybrid space, LLM-level planning with VLA-level control, that it describes only briefly; a systematic hybrid benchmark would be a natural extension.
- The world-model sections suggest a testable prediction the authors do not demonstrate: agents that imagine rollouts in a learned world model before acting should outperform purely reactive VLA policies on long-horizon tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a survey of large-model-empowered embodied AI focused on autonomous decision-making and embodied learning. It proposes a taxonomy that separates hierarchical and end-to-end decision-making, organizes embodied learning into imitation learning, reinforcement learning, transfer learning, and meta-learning, and includes world models as a first-time contribution to embodied-AI surveys. The paper's stated contributions are organizational and pedagogical: no new algorithms or derived quantities are introduced.
Significance. If the content were accurately anchored to the literature, this survey would be a useful map of a fast-moving field. The taxonomy is coherent, the dual horizontal/vertical perspective is reasonable, and the comparison tables cover a broad set of recent VLA and world-model works. The claim of first integration of world models into an embodied-AI survey is notable. However, the central deliverable of a survey is traceability: a reader must be able to move from a claim to the cited reference number to the correct bibliography entry. That chain breaks repeatedly in the current manuscript, so the survey cannot currently serve its primary function. The errors are systematic rather than isolated and affect load-bearing comparative claims.
major comments (3)
- [§3.2.3 and §3.3.1] The same system DEPS is cited as [192] in §3.2.3 and as [153] in §3.3.1, while [153] is the Re-Prompting paper (Raman et al.) and [192] is the DEPS paper (Wang et al.). This is an internal inconsistency in the claim-to-reference chain and directly undermines the survey's stated purpose of pointing readers to the primary literature.
- [§4.2.3, §7.3, and Table 2] TinyVLA is repeatedly cited as [234], but [234] is the RT-2 paper; TinyVLA is [198]. The inference-speed and data-efficiency claims attributed to TinyVLA in §4.2.3 and §7.3 are therefore attached to the wrong source. Similarly, §7.3 cites DiffusionVLA as [22], which is the Sora paper, while Diffusion-VLA is [196]. These are not minor number typos; they misattribute the very technical results the survey uses for comparative claims.
- [§7.1 and §5.2.2] In §7.1 RT-1 is cited as [28], but [28] is the GLAM paper (Carta et al.); RT-1 is [20]. In §5.2.2 HiveFormer is cited as [224], but [224] is ALOHA. The simultaneous use of [224] for both ALOHA and HiveFormer in the same paragraph shows that the citation-number mapping is unreliable and requires a full audit, not just patch-up of individual entries.
minor comments (3)
- [§5.1.4, Eq. (5)] The meta-learning objective is garbled: the notation "T⟩" appears in the summation, and the placement of the inner update relative to the outer minimization is unclear. Please reformat the equation and define all symbols consistently with the text.
- [References] Several bibliography entries are duplicated under different numbers: OpenVLA appears as both [93] and [94], SAM appears as both [96] and [97], Embodied-GPT appears as both [131] and [132], and Dreamer appears as both [69] and [70]. This makes the reference list hard to use and is consistent with an incomplete citation-management pass.
- [§7.3] Typo: "knowledge distillation ad quantization" should read "knowledge distillation and quantization".
Circularity Check
No circularity: the survey is a bibliographic/organizational review, not a derivational claim; citation-number errors are accuracy issues, not circularity.
full rationale
This is a survey paper. Its central claims are taxonomic and organizational (e.g., introducing a hierarchy of decision-making paradigms, categorizing VLA enhancements, and integrating world models into the survey). It contains no fitted parameters, no derived quantitative predictions, and no mathematical derivation whose conclusion is assumed in its premises. The abstract claim of being 'the first time' to integrate world models into an embodied-AI survey is a bibliographic/novelty claim, not a result derived from inputs. The paper does not invoke any load-bearing self-citation, uniqueness theorem, or ansatz imported from the authors' prior work; no author of this survey appears in the reference list as a source of a foundational premise. The weaknesses identified in the skeptical review—inconsistent reference numbering (DEPS cited as both [153] and [192], TinyVLA cited as [234] = RT-2, Diffusion-VLA cited as [22] = Sora, etc.)—are genuine quality defects that undermine the survey's reference-traceability function, but they are not circularity. A wrong citation does not make a claim equivalent to its own input; it makes the claim harder to audit. Under the rule that circularity requires a specific reduction by definition, fitted input, or self-citation chain, none is present here. The honest finding is therefore 0.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption The proposed taxonomy (hierarchical vs end-to-end decision-making, IL vs RL, plus world models) is exhaustive enough to organize the field
- domain assumption The cited references accurately support the claims attached to them
Cite this review
Pith. "Pith review of Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning." pith.science (2026). https://pith.science/paper/AFYJJ57B
@misc{pith2026250810399,
author = {Pith},
title = {Pith review of: Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/AFYJJ57B}},
note = {Machine review of arXiv:2508.10399}
}
read the original abstract
Embodied AI aims to develop intelligent systems with physical forms capable of perceiving, decision-making, acting, and learning in real-world environments, providing a promising way to Artificial General Intelligence (AGI). Despite decades of explorations, it remains challenging for embodied agents to achieve human-level intelligence for general-purpose tasks in open dynamic environments. Recent breakthroughs in large models have revolutionized embodied AI by enhancing perception, interaction, planning and learning. In this article, we provide a comprehensive survey on large model empowered embodied AI, focusing on autonomous decision-making and embodied learning. We investigate both hierarchical and end-to-end decision-making paradigms, detailing how large models enhance high-level planning, low-level execution, and feedback for hierarchical decision-making, and how large models enhance Vision-Language-Action (VLA) models for end-to-end decision making. For embodied learning, we introduce mainstream learning methodologies, elaborating on how large models enhance imitation learning and reinforcement learning in-depth. For the first time, we integrate world models into the survey of embodied AI, presenting their design methods and critical roles in enhancing decision-making and learning. Though solid advances have been achieved, challenges still exist, which are discussed at the end of this survey, potentially as the further research directions.
Figures
Forward citations
Cited by 12 Pith papers
-
HumanCLAW: Can Vision-Language Models Act Through a Body?
A new full-body benchmark shows that current VLMs can recognize targets but cannot reliably tell where their own body is, whether it arrived, or whether it collided; the best solves only 16.8% of episodes.
-
World Models as Group Actions
Formalizes video world models as group actions on states and uses latent regularization with synthesized supervision to enforce consistency, introducing GAC and GAR metrics that improve structural correctness in SOTA models.
-
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
GridS is a plug-and-play differentiable module for geometry-aware visual token resampling in VLA models that achieves under 10% token retention and 76% FLOPs reduction with no success-rate loss.
-
HumanCLAW: Can Vision-Language Models Act Through a Body?
Off-the-shelf VLMs fail closed-loop whole-body find-navigate-interact tasks (best 16.8%) because they lack embodied self-awareness, not target recognition.
-
Analytic Concept-Centric Memory for Agentic Embodied Manipulation
Proposes a structured concept-centric memory system for embodied agents that connects object, scene, transition, and skill memories to support coarse-to-fine retrieval and improve task performance over baselines.
-
Joint Learning of Experiential Rules and Policies for Large Language Model Agents
JERP jointly updates experiential rules and policies for LLM agents from shared trajectories, keeping rules aligned with the policy and yielding gains on AlfWorld and WebShop.
-
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
GridS reduces visual tokens in VLA models to under 10% of the original count via task-aware differentiable resampling, delivering 76% lower FLOPs with no drop in task success rate on benchmarks and real robots.
-
Social Structure Matters in 3D Human-Human Interaction Generation
Introduces a Solo-to-Social planner-executor framework where LLMs decompose HHI into phases and roles, then a LoRA-adapted solo motion model grounds them into partner-aware 3D motion.
-
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning
A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.
-
Position: Embodied AI Requires a Privacy-Utility Trade-off
Embodied AI requires treating privacy as a lifecycle architectural constraint rather than a stage-local feature, addressed via the proposed SPINE framework with a multi-criterion privacy classification matrix.
-
Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
A survey of UAV vision-and-language navigation that establishes a methodological taxonomy, reviews resources and challenges, and proposes a forward-looking research roadmap.
-
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
A survey of physical AI that distinguishes theoretical physics reasoning from applied understanding and synthesizes advances in symbolic reasoning, embodied systems, and generative models to advocate for physics-groun...
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
Pith/arXiv arXiv 2023
-
[2]
Abdul Afram and Farrokh Janabi-Sharifi. 2014. Theory and applications of HVAC control systems–A review of model predictive control (MPC). Building and environment 72 (2014), 343–355
2014
-
[3]
Ali Agha, Kyohei Otsu, Benjamin Morrell, David D Fan, Rohan Thakker, Angel Santamaria-Navarro, Sung-Kyun Kim, Amanda Bouman, Xianmei Lei, Jeffrey Edlund, et al. 2021. Nebula: Quest for robotic autonomy in challenging environments; team costar at the darpa subterranean challenge. arXiv preprint arXiv:2103.11470 (2021)
Pith/arXiv arXiv 2021
-
[4]
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 (2022)
Pith/arXiv arXiv 2022
-
[5]
Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, et al. 2024. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv preprint arXiv:2401.12963 (2024)
arXiv 2024
-
[6]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736
2022
-
[7]
Shan An, Ziyu Meng, Chao Tang, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu, Ran Song, Wei Zhang, et al. 2025. Dexterous manipulation through imitation learning: A survey. arXiv preprint arXiv:2504.03515 (2025)
arXiv 2025
-
[8]
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403 (2023). Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning 39
Pith/arXiv arXiv 2023
-
[9]
Daman Arora and Subbarao Kambhampati. 2023. Learning and leveraging verifiers to improve planning capabilities of pre-trained language models. arXiv preprint arXiv:2305.17077 (2023)
Pith/arXiv arXiv 2023
-
[10]
Benedikt Bagus and Alexander Gepperth. 2021. An investigation of replay-based approaches for continual learning. In 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–9
2021
-
[11]
Andrew G Barto. 2021. Reinforcement learning: An introduction. by richard’s sutton. SIAM Rev 6, 2 (2021), 423
2021
-
[12]
Suneel Belkhale, Yuchen Cui, and Dorsa Sadigh. 2023. Hydra: Hybrid robot actions for imitation learning. InConference on Robot Learning. PMLR, 2113–2133
2023
-
[13]
Guillaume Bellegarda, Yiyu Chen, Zhuochen Liu, and Quan Nguyen. 2022. Robust high-speed running for quadruped robots via deep reinforcement learning. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 10364–10370
2022
-
[14]
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley. 2023. Unlimiformer: Long-range transformers with unlimited length input. Advances in Neural Information Processing Systems 36 (2023), 35522–35543
2023
-
[15]
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the AAAI conference on artificial intelligence , Vol. 38. 17682–17690
2024
-
[16]
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. 2023. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2, 3 (2023), 8
2023
-
[17]
Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. 2024. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 4788–4795
2024
-
[18]
2024.𝜋0: A Vision-Language-Action Flow Model for General Robot Control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. 2024.𝜋0: A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164 (2024)
Pith/arXiv arXiv 2024
-
[19]
Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Devin, Alex X Lee, Maria Bauzá, Todor Davchev, Yuxi- ang Zhou, Agrim Gupta, Akhil Raju, et al. 2023. Robocat: A self-improving generalist agent for robotic manipulation. arXiv preprint arXiv:2306.11706 (2023)
Pith/arXiv arXiv 2023
-
[20]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakr- ishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. 2022. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817 (2022)
Pith/arXiv arXiv 2022
-
[21]
Rodney Brooks. 2003. A robust layered control system for a mobile robot. IEEE journal on robotics and automation 2, 1 (2003), 14–23
2003
-
[22]
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. 2024. Video generation models as world simulators. OpenAI Blog 1, 8 (2024), 1
2024
-
[23]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901
2020
-
[24]
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. 2024. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning
2024
-
[25]
Paweł Budzianowski, Wesley Maa, Matthew Freed, Jingxiang Mo, Winston Hsiao, Aaron Xie, Tomasz Młodu- chowski, Viraj Tipnis, and Benjamin Bolte. 2025. EdgeVLA: Efficient Vision-Language-Action Models. arXiv preprint arXiv:2507.14049 (2025)
Pith/arXiv arXiv 2025
-
[26]
Yuji Cao, Huan Zhao, Yuheng Cheng, Ting Shu, Yue Chen, Guolong Liu, Gaoqi Liang, Junhua Zhao, Jinyue Yan, and Yun Li. 2024. Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods. IEEE Transactions on Neural Networks and Learning Systems (2024)
2024
-
[27]
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision. 9650–9660
2021
-
[28]
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer. 2023. Grounding large language models in interactive environments with online reinforcement learning. In International Conference on Machine Learning . PMLR, 3676–3713
2023
-
[29]
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology 15, 3 (2024), 1–45
2024
-
[30]
Yevgen Chebotar, Quan Vuong, Karol Hausman, Fei Xia, Yao Lu, Alex Irpan, Aviral Kumar, Tianhe Yu, Alexander Herzog, Karl Pertsch, et al. 2023. Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions. 40 Liang et al. In Conference on Robot Learning . PMLR, 3909–3928
2023
-
[31]
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems 34 (2021), 15084–15097
2021
-
[32]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)
Pith/arXiv arXiv 2021
-
[33]
Po-Lin Chen and Cheng-Shang Chang. 2023. Interact: Exploring the potentials of chatgpt as a cooperative agent. arXiv preprint arXiv:2308.01552 (2023)
Pith/arXiv arXiv 2023
-
[34]
Shang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai, and Shao-Hua Sun. 2023. Diffusion model- augmented behavioral cloning. arXiv preprint arXiv:2302.13335 (2023)
Pith/arXiv arXiv 2023
-
[35]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning . PmLR, 1597–1607
2020
-
[36]
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song
-
[37]
Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu, Mithun George Jacob, Tingnan Zhang, Tsang-Wei Edward Lee, Wenhao Yu, Connor Schenck, David Rendleman, Dhruv Shah, et al. 2024. Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs. arXiv preprint arXiv:2407.07775 (2024)
Pith/arXiv arXiv 2024
-
[38]
WL Chiang et al. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 With 90% ChatGPT Quality. Accessed: Apr. 14, 2023
2023
-
[39]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1–113
2023
-
[40]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. 2023. Diffusion models in vision: A survey. IEEE transactions on pattern analysis and machine intelligence 45, 9 (2023), 10850–10869
2023
-
[41]
Sahith Dambekodi, Spencer Frazier, Prithviraj Ammanabrolu, and Mark O Riedl. 2020. Playing text-based games with common sense. arXiv preprint arXiv:2012.02757 (2020)
Pith/arXiv arXiv 2020
-
[42]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) . 4171–4186
2019
-
[43]
Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al. 2024. Understanding world or predicting future? a comprehensive survey of world models. Comput. Surveys (2024)
2024
-
[44]
Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. 2024. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. arXiv preprint arXiv:2408.11812 (2024)
Pith/arXiv arXiv 2024
-
[45]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)
Pith/arXiv arXiv 2020
-
[46]
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, et al. 2023. Palm-e: An embodied multimodal language model. (2023)
2023
-
[47]
Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel
-
[48]
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. 2022. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6, 2 (2022), 230–244
2022
-
[49]
Advances in neural information processing systems 36 (2023), 9156–9172
Learning universal policies via text-guided video generation. Advances in neural information processing systems 36 (2023), 9156–9172
2023
-
[50]
Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12873–12883
2021
-
[51]
Jonas Eschmann. 2021. Reward function design in reinforcement learning.Reinforcement learning algorithms: Analysis and Applications (2021), 25–33
2021
-
[52]
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning . PMLR, 1126–1135
2017
-
[53]
Rasool Fakoor, Pratik Chaudhari, Stefano Soatto, and Alexander J Smola. 2019. Meta-q-learning. arXiv preprint arXiv:1910.00125 (2019)
Pith/arXiv arXiv 2019
-
[54]
Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its nature, scope, limits, and consequences. Minds and machines 30, 4 (2020), 681–694
2020
-
[55]
Pete Florence, Corey Lynch, Andy Zeng, Oscar A Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson. 2022. Implicit behavioral cloning. In Conference on robot learning . PMLR, 158–168. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning 41
2022
-
[56]
Gene F Franklin, J David Powell, Abbas Emami-Naeini, and J David Powell. 2002. Feedback control of dynamic systems . Vol. 4. Prentice hall Upper Saddle River
2002
-
[57]
Maria Fox and Derek Long. 2003. PDDL2. 1: An extension to PDDL for expressing temporal planning domains. Journal of artificial intelligence research 20 (2003), 61–124
2003
-
[58]
Zipeng Fu, Tony Z Zhao, and Chelsea Finn. 2024. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117 (2024)
Pith/arXiv arXiv 2024
-
[59]
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219 (2020)
Pith/arXiv arXiv 2020
-
[60]
Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox. 2023. Rvt: Robotic view transformer for 3d object manipulation. In Conference on Robot Learning . PMLR, 694–710
2023
-
[61]
Michael Gelfond and Yulia Kahl. 2014. Knowledge representation, reasoning, and the design of intelligent agents: The answer-set programming approach. Cambridge University Press
2014
-
[62]
Jiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu, Montserrat Gonzalez Arenas, Kanishka Rao, Wenhao Yu, Chuyuan Fu, Keerthana Gopalakrishnan, Zhuo Xu, et al. 2023. Rt-trajectory: Robotic task generalization via hindsight trajectory sketches. arXiv preprint arXiv:2311.01977 (2023)
Pith/arXiv arXiv 2023
-
[63]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 18995–19012
2022
-
[64]
Lin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. 2023. Leveraging pre-trained large language models to construct and utilize world models for model-based task planning.Advances in Neural Information Processing Systems 36 (2023), 79081–79094
2023
-
[65]
Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, and Yang Gao. 2023. Seer: Language instructed video prediction with latent diffusion models. arXiv preprint arXiv:2303.14897 (2023)
Pith/arXiv arXiv 2023
-
[66]
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine. 2018. Meta-reinforcement learning of structured exploration strategies. Advances in neural information processing systems 31 (2018)
2018
-
[67]
Yanjiang Guo, Yen-Jen Wang, Lihan Zha, and Jianyu Chen. 2024. Doremi: Grounding language model by detecting and recovering from plan-execution misalignment. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 12124–12131
2024
-
[68]
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning . Pmlr, 1861–1870
2018
-
[69]
David Ha and Jürgen Schmidhuber. 2018. Recurrent world models facilitate policy evolution. Advances in neural information processing systems 31 (2018)
2018
-
[70]
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2019. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603 (2019)
Pith/arXiv arXiv 2019
-
[72]
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. 2020. Mastering atari with discrete world models. arXiv preprint arXiv:2010.02193 (2020)
Pith/arXiv arXiv 2020
-
[73]
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. 2019. Learning latent dynamics for planning from pixels. InInternational conference on machine learning. PMLR, 2555–2565
2019
-
[74]
Asher J Hancock, Allen Z Ren, and Anirudha Majumdar. 2024. Run-time observation interventions make vision- language-action models more visually robust. arXiv preprint arXiv:2410.01971 (2024)
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[75]
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104 (2023)
Pith/arXiv arXiv 2023
-
[76]
Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang, Bin Zhao, and Xuelong Li. 2024. Large-scale actionless video pre-training via discrete diffusion for efficient policy learning. CoRR (2024)
2024
-
[77]
Patrik Haslum, Nir Lipovetzky, Daniele Magazzeni, Christian Muise, Ronald Brachman, Francesca Rossi, and Peter Stone. 2019. An introduction to the planning domain definition language . Vol. 13. Springer
2019
-
[78]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16000–16009. 42 Liang et al
2022
-
[79]
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li. 2023. Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning. Advances in neural information processing systems 36 (2023), 64896–64917
2023
-
[80]
Kazuki Hori, Kanata Suzuki, and Tetsuya Ogata. 2024. Interactively robot action planning with uncertainty analysis and active questioning by large language model. In 2024 IEEE/SICE International Symposium on System Integration (SII). IEEE, 85–91
2024
-
[81]
Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29 (2016)
2016
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.