Pith. sign in

REVIEW 3 major objections 3 minor 12 cited by

This survey proposes that large-model-empowered embodied AI is best understood through two decision-making paradigms—hierarchical and end-to-end vision-language-action models—with imitation learning and reinforcement learning as the learnin

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection A useful survey skeleton with a healthy world-models section, but the citation chain is broken in several places, which guts the reference-traceability function in its current form. the 3 major comments →

arxiv 2508.10399 v1 pith:AFYJJ57B submitted 2025-08-14 cs.RO

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

classification cs.RO
keywords embodied AIlarge modelshierarchical decision-makingend-to-end decision-makingvision-language-action modelsimitation learningreinforcement learningworld models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The survey sets out to organize the fast-moving field of large-model-empowered embodied AI into one coherent map. It claims that embodied decision-making splits into hierarchical and end-to-end paradigms, that embodied learning is driven mainly by imitation and reinforcement learning, and that world models belong inside this picture as a first-class component that strengthens both decision-making and learning. If the map is right, a researcher can place any recent method by asking how it handles planning, execution, feedback, and learning, and can see where world models fit. The claim matters because existing reviews cover pieces, such as language models, vision-language models, planning, or robotics, but not the whole loop from instruction to physical action to experience-based improvement.

Core claim

The central claim is that recent large-model-empowered embodied systems form a structured landscape rather than a scattered collection. The landscape has two decision-making paradigms: hierarchical, where large language models handle planning and separate modules handle execution and feedback, and end-to-end, where vision-language-action models map observations and instructions directly to action tokens. It has two learning engines, imitation learning and reinforcement learning, each enhanced by large models through diffusion or Transformer policies and through automated reward design. The survey further claims, for the first time in this line of surveys, that world models are a fourth pilla

What carries the argument

The organizing device is a taxonomy with three axes: decision-making paradigm (hierarchical versus end-to-end), embodied learning method (imitation versus reinforcement), and world-model design (latent-state, Transformer-based, diffusion-based, and joint-embedding predictive architectures). The named central object is the vision-language-action model, defined as a single network that maps multimodal observations and language instructions directly to action tokens; it anchors the end-to-end paradigm. The survey pairs this grid with a dual analytical approach that compares approaches horizontally across paradigms and traces core models vertically through their development, which is what carrie

Load-bearing premise

The survey is useful as a map only if the reader can trust that each bracket number points to the paper actually supporting that sentence; a few markers do not line up (TinyVLA is cited through the RT-2 reference, DiffusionVLA through the Sora reference, and DEPS is assigned two different numbers), so a citation audit is the load-bearing check.

What would settle it

Open the survey at the load-bearing citations and verify them against the primary literature: Section 4.2.3's claim about TinyVLA should lead to the TinyVLA paper, not RT-2 [234], and Section 7.3's claim about DiffusionVLA should lead to DiffusionVLA, not Sora [22]. If those markers mispoint, the map cannot be used to reach the cited methods.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A researcher encountering a new embodied system can classify it by paradigm and learning method and immediately see which prior work it builds on and which comparison table applies.
  • The hierarchical-versus-end-to-end comparison gives practitioners explicit selection criteria: interpretability and reliability versus generalization and speed, at different computational costs.
  • Treating world models as a first-class component turns learned simulators into a standard tool for both decision validation and data generation in embodied learning.
  • The four open challenges identified by the survey, embodied data scarcity, continual learning, deployment efficiency, and the sim-to-real gap, become a shared priority list with world models and large-model reward design as candidate remedies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy holds, a useful next step would be to mine the literature for cells that are still empty, such as world-model-augmented reinforcement learning on physical robots, since the survey's own tables suggest those gaps.
  • The survey's comparison of hierarchical and end-to-end paradigms implies a hybrid space, LLM-level planning with VLA-level control, that it describes only briefly; a systematic hybrid benchmark would be a natural extension.
  • The world-model sections suggest a testable prediction the authors do not demonstrate: agents that imagine rollouts in a learned world model before acting should outperform purely reactive VLA policies on long-horizon tasks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript is a survey of large-model-empowered embodied AI focused on autonomous decision-making and embodied learning. It proposes a taxonomy that separates hierarchical and end-to-end decision-making, organizes embodied learning into imitation learning, reinforcement learning, transfer learning, and meta-learning, and includes world models as a first-time contribution to embodied-AI surveys. The paper's stated contributions are organizational and pedagogical: no new algorithms or derived quantities are introduced.

Significance. If the content were accurately anchored to the literature, this survey would be a useful map of a fast-moving field. The taxonomy is coherent, the dual horizontal/vertical perspective is reasonable, and the comparison tables cover a broad set of recent VLA and world-model works. The claim of first integration of world models into an embodied-AI survey is notable. However, the central deliverable of a survey is traceability: a reader must be able to move from a claim to the cited reference number to the correct bibliography entry. That chain breaks repeatedly in the current manuscript, so the survey cannot currently serve its primary function. The errors are systematic rather than isolated and affect load-bearing comparative claims.

major comments (3)
  1. [§3.2.3 and §3.3.1] The same system DEPS is cited as [192] in §3.2.3 and as [153] in §3.3.1, while [153] is the Re-Prompting paper (Raman et al.) and [192] is the DEPS paper (Wang et al.). This is an internal inconsistency in the claim-to-reference chain and directly undermines the survey's stated purpose of pointing readers to the primary literature.
  2. [§4.2.3, §7.3, and Table 2] TinyVLA is repeatedly cited as [234], but [234] is the RT-2 paper; TinyVLA is [198]. The inference-speed and data-efficiency claims attributed to TinyVLA in §4.2.3 and §7.3 are therefore attached to the wrong source. Similarly, §7.3 cites DiffusionVLA as [22], which is the Sora paper, while Diffusion-VLA is [196]. These are not minor number typos; they misattribute the very technical results the survey uses for comparative claims.
  3. [§7.1 and §5.2.2] In §7.1 RT-1 is cited as [28], but [28] is the GLAM paper (Carta et al.); RT-1 is [20]. In §5.2.2 HiveFormer is cited as [224], but [224] is ALOHA. The simultaneous use of [224] for both ALOHA and HiveFormer in the same paragraph shows that the citation-number mapping is unreliable and requires a full audit, not just patch-up of individual entries.
minor comments (3)
  1. [§5.1.4, Eq. (5)] The meta-learning objective is garbled: the notation "T⟩" appears in the summation, and the placement of the inner update relative to the outer minimization is unclear. Please reformat the equation and define all symbols consistently with the text.
  2. [References] Several bibliography entries are duplicated under different numbers: OpenVLA appears as both [93] and [94], SAM appears as both [96] and [97], Embodied-GPT appears as both [131] and [132], and Dreamer appears as both [69] and [70]. This makes the reference list hard to use and is consistent with an incomplete citation-management pass.
  3. [§7.3] Typo: "knowledge distillation ad quantization" should read "knowledge distillation and quantization".

Circularity Check

0 steps flagged

No circularity: the survey is a bibliographic/organizational review, not a derivational claim; citation-number errors are accuracy issues, not circularity.

full rationale

This is a survey paper. Its central claims are taxonomic and organizational (e.g., introducing a hierarchy of decision-making paradigms, categorizing VLA enhancements, and integrating world models into the survey). It contains no fitted parameters, no derived quantitative predictions, and no mathematical derivation whose conclusion is assumed in its premises. The abstract claim of being 'the first time' to integrate world models into an embodied-AI survey is a bibliographic/novelty claim, not a result derived from inputs. The paper does not invoke any load-bearing self-citation, uniqueness theorem, or ansatz imported from the authors' prior work; no author of this survey appears in the reference list as a source of a foundational premise. The weaknesses identified in the skeptical review—inconsistent reference numbering (DEPS cited as both [153] and [192], TinyVLA cited as [234] = RT-2, Diffusion-VLA cited as [22] = Sora, etc.)—are genuine quality defects that undermine the survey's reference-traceability function, but they are not circularity. A wrong citation does not make a claim equivalent to its own input; it makes the claim harder to audit. Under the rule that circularity requires a specific reduction by definition, fitted input, or self-citation chain, none is present here. The honest finding is therefore 0.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

This survey introduces no free parameters, no fitted values, and no invented entities. Its claims rest on the accuracy of its citations and the completeness of its taxonomy, both of which are problematic given the citation errors and the absence of a stated literature search protocol.

axioms (2)
  • domain assumption The proposed taxonomy (hierarchical vs end-to-end decision-making, IL vs RL, plus world models) is exhaustive enough to organize the field
    The survey's central claim of comprehensiveness rests on this categorization, but no search protocol or inclusion criteria are provided to justify exhaustiveness. The claim appears in the Abstract and Section 1.
  • domain assumption The cited references accurately support the claims attached to them
    The survey is a secondary source, so every description depends on this assumption. It is violated by mis-citations such as DEPS[153] vs [192] and TinyVLA[234] vs [198], meaning the assumption does not hold.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning." pith.science (2026). https://pith.science/paper/AFYJJ57B

@misc{pith2026250810399,
  author       = {Pith},
  title        = {Pith review of: Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AFYJJ57B}},
  note         = {Machine review of arXiv:2508.10399}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Embodied AI aims to develop intelligent systems with physical forms capable of perceiving, decision-making, acting, and learning in real-world environments, providing a promising way to Artificial General Intelligence (AGI). Despite decades of explorations, it remains challenging for embodied agents to achieve human-level intelligence for general-purpose tasks in open dynamic environments. Recent breakthroughs in large models have revolutionized embodied AI by enhancing perception, interaction, planning and learning. In this article, we provide a comprehensive survey on large model empowered embodied AI, focusing on autonomous decision-making and embodied learning. We investigate both hierarchical and end-to-end decision-making paradigms, detailing how large models enhance high-level planning, low-level execution, and feedback for hierarchical decision-making, and how large models enhance Vision-Language-Action (VLA) models for end-to-end decision making. For embodied learning, we introduce mainstream learning methodologies, elaborating on how large models enhance imitation learning and reinforcement learning in-depth. For the first time, we integrate world models into the survey of embodied AI, presenting their design methods and critical roles in enhancing decision-making and learning. Though solid advances have been achieved, challenges still exist, which are discussed at the end of this survey, potentially as the further research directions.

Figures

Figures reproduced from arXiv: 2508.10399 by Bing Zhang, Ping Kuang, Rui Zhou, Songlin Li, Wenlong Liang, Yang Ma, Yijia Liao.

Figure 1
Figure 1. Figure 1: Organization of this survey. intelligent vehicles[160], execute actions and receive feedback, serving as the interface between the physical and the digital worlds. Intelligent agents form the cognitive core, enabling autonomous decision-making and learning. To perform embodied tasks, embodied AI systems interpret human intentions from language instructions, actively explore their surroundings, perceive mul… view at source ↗
Figure 2
Figure 2. Figure 2: Embodied AI: from prospect of capabilities required during the whole process. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Timeline of major large models. Subsequently, OpenAI released GPT[149], a generative model based on Transformer architecture, which used autoregressive training on large-scale unsupervised corpora to produce coherent text, marking a breakthrough in generative models. GPT-2[150] further scaled up model size and training data, enhancing text coherence and naturalness. In 2020, GPT-3[54] set a milestone with … view at source ↗
Figure 4
Figure 4. Figure 4: General capability enhancement of large models. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Hierarchical decision-making paradigm, consisting of perception and interaction, high-level planning, [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: High-level planning empowered by large models. [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Low-level execution [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Feedback and enhancement. 3.3.1 Self-Reflection of Large Models. The large models can act as task planners, evaluators and optimizers, thus iteratively refine decision-making processes without external interventions. The agents get feedback of actions, autonomously detect and analyze failed executions, and continuously learn from previous tasks. By this self-reflection and optimization mechanism, the large… view at source ↗
Figure 9
Figure 9. Figure 9: End-to-end decision-making by VLA [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Vision-Language-Action Models. (1) Tokenization and Representation. VLA models use four token types: vision, language, state, and action, to encode multimodal inputs for context-aware action generation. Vision tokens and language tokens encode environmental scenes and instructions into embeddings, forming the foundation of task and context. State tokens capture the agent’s physical configuration, includin… view at source ↗
Figure 11
Figure 11. Figure 11: 4.2.1 Perception Capability Enhancement. To improve the perception capability, BYO-VLA[74] optimizes the tokenization and representation component by implementing a runtime observation intervention mechanism, which leverages automated image preprocessing to filter out visual noise originating from occluding objects and cluttered backgrounds. TraceVLA[229] focuses on the multimodal information fusion compo… view at source ↗
Figure 11
Figure 11. Figure 11: Enhancements on Vision-Language-Action Models. [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Embodied learning: process and methodologies. [PITH_FULL_IMAGE:figures/full_fig_p024_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Imitation learning empowered by diffusion models or Transformers. [PITH_FULL_IMAGE:figures/full_fig_p028_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Reinforcement learning empowered by large models. [PITH_FULL_IMAGE:figures/full_fig_p030_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Policy network construction empowered by large models. [PITH_FULL_IMAGE:figures/full_fig_p031_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: World models and applications in decision-making and embodied learning. [PITH_FULL_IMAGE:figures/full_fig_p032_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HumanCLAW: Can Vision-Language Models Act Through a Body?

    cs.CV 2026-07 conditional novelty 7.0

    A new full-body benchmark shows that current VLMs can recognize targets but cannot reliably tell where their own body is, whether it arrived, or whether it collided; the best solves only 16.8% of episodes.

  2. World Models as Group Actions

    cs.CV 2026-05 unverdicted novelty 7.0

    Formalizes video world models as group actions on states and uses latent regularization with synthesized supervision to enforce consistency, introducing GAC and GAR metrics that improve structural correctness in SOTA models.

  3. See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

    cs.RO 2026-05 conditional novelty 7.0

    GridS is a plug-and-play differentiable module for geometry-aware visual token resampling in VLA models that achieves under 10% token retention and 76% FLOPs reduction with no success-rate loss.

  4. HumanCLAW: Can Vision-Language Models Act Through a Body?

    cs.CV 2026-07 conditional novelty 6.5

    Off-the-shelf VLMs fail closed-loop whole-body find-navigate-interact tasks (best 16.8%) because they lack embodied self-awareness, not target recognition.

  5. Analytic Concept-Centric Memory for Agentic Embodied Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    Proposes a structured concept-centric memory system for embodied agents that connects object, scene, transition, and skill memories to support coarse-to-fine retrieval and improve task performance over baselines.

  6. Joint Learning of Experiential Rules and Policies for Large Language Model Agents

    cs.AI 2026-06 unverdicted novelty 6.0

    JERP jointly updates experiential rules and policies for LLM agents from shared trajectories, keeping rules aligned with the policy and yielding gains on AlfWorld and WebShop.

  7. See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

    cs.RO 2026-05 unverdicted novelty 6.0

    GridS reduces visual tokens in VLA models to under 10% of the original count via task-aware differentiable resampling, delivering 76% lower FLOPs with no drop in task success rate on benchmarks and real robots.

  8. Social Structure Matters in 3D Human-Human Interaction Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    Introduces a Solo-to-Social planner-executor framework where LLMs decompose HHI into phases and roles, then a LoRA-adapted solo motion model grounds them into partner-aware 3D motion.

  9. RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

    cs.AI 2025-10 conditional novelty 5.0

    A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.

  10. Position: Embodied AI Requires a Privacy-Utility Trade-off

    cs.AI 2026-05 unverdicted novelty 4.0

    Embodied AI requires treating privacy as a lifecycle architectural constraint rather than a stage-local feature, addressed via the proposed SPINE framework with a multi-criterion privacy classification matrix.

  11. Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

    cs.RO 2026-04 unverdicted novelty 4.0

    A survey of UAV vision-and-language navigation that establishes a methodological taxonomy, reviews resources and challenges, and proposes a forward-looking research roadmap.

  12. Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI

    cs.AI 2025-10 unverdicted novelty 4.0

    A survey of physical AI that distinguishes theoretical physics reasoning from applied understanding and synthesizes advances in symbolic reasoning, embodied systems, and generative models to advocate for physics-groun...

Reference graph

Works this paper leans on

233 extracted references · 2 canonical work pages · cited by 10 Pith papers · 2 internal anchors

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Abdul Afram and Farrokh Janabi-Sharifi. 2014. Theory and applications of HVAC control systems–A review of model predictive control (MPC). Building and environment 72 (2014), 343–355

  3. [3]

    Ali Agha, Kyohei Otsu, Benjamin Morrell, David D Fan, Rohan Thakker, Angel Santamaria-Navarro, Sung-Kyun Kim, Amanda Bouman, Xianmei Lei, Jeffrey Edlund, et al. 2021. Nebula: Quest for robotic autonomy in challenging environments; team costar at the darpa subterranean challenge. arXiv preprint arXiv:2103.11470 (2021)

  4. [4]

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 (2022)

  5. [5]

    Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, et al. 2024. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv preprint arXiv:2401.12963 (2024)

  6. [6]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736

  7. [7]

    Shan An, Ziyu Meng, Chao Tang, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu, Ran Song, Wei Zhang, et al. 2025. Dexterous manipulation through imitation learning: A survey. arXiv preprint arXiv:2504.03515 (2025)

  8. [8]

    Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403 (2023). Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning 39

  9. [9]

    Daman Arora and Subbarao Kambhampati. 2023. Learning and leveraging verifiers to improve planning capabilities of pre-trained language models. arXiv preprint arXiv:2305.17077 (2023)

  10. [10]

    Benedikt Bagus and Alexander Gepperth. 2021. An investigation of replay-based approaches for continual learning. In 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–9

  11. [11]

    Andrew G Barto. 2021. Reinforcement learning: An introduction. by richard’s sutton. SIAM Rev 6, 2 (2021), 423

  12. [12]

    Suneel Belkhale, Yuchen Cui, and Dorsa Sadigh. 2023. Hydra: Hybrid robot actions for imitation learning. InConference on Robot Learning. PMLR, 2113–2133

  13. [13]

    Guillaume Bellegarda, Yiyu Chen, Zhuochen Liu, and Quan Nguyen. 2022. Robust high-speed running for quadruped robots via deep reinforcement learning. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 10364–10370

  14. [14]

    Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley. 2023. Unlimiformer: Long-range transformers with unlimited length input. Advances in Neural Information Processing Systems 36 (2023), 35522–35543

  15. [15]

    Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the AAAI conference on artificial intelligence , Vol. 38. 17682–17690

  16. [16]

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. 2023. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2, 3 (2023), 8

  17. [17]

    Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. 2024. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 4788–4795

  18. [18]

    2024.𝜋0: A Vision-Language-Action Flow Model for General Robot Control

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. 2024.𝜋0: A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164 (2024)

  19. [19]

    Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Devin, Alex X Lee, Maria Bauzá, Todor Davchev, Yuxi- ang Zhou, Agrim Gupta, Akhil Raju, et al. 2023. Robocat: A self-improving generalist agent for robotic manipulation. arXiv preprint arXiv:2306.11706 (2023)

  20. [20]

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakr- ishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. 2022. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817 (2022)

  21. [21]

    Rodney Brooks. 2003. A robust layered control system for a mobile robot. IEEE journal on robotics and automation 2, 1 (2003), 14–23

  22. [22]

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. 2024. Video generation models as world simulators. OpenAI Blog 1, 8 (2024), 1

  23. [23]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  24. [24]

    Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. 2024. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning

  25. [25]

    Paweł Budzianowski, Wesley Maa, Matthew Freed, Jingxiang Mo, Winston Hsiao, Aaron Xie, Tomasz Młodu- chowski, Viraj Tipnis, and Benjamin Bolte. 2025. EdgeVLA: Efficient Vision-Language-Action Models. arXiv preprint arXiv:2507.14049 (2025)

  26. [26]

    Yuji Cao, Huan Zhao, Yuheng Cheng, Ting Shu, Yue Chen, Guolong Liu, Gaoqi Liang, Junhua Zhao, Jinyue Yan, and Yun Li. 2024. Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods. IEEE Transactions on Neural Networks and Learning Systems (2024)

  27. [27]

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision. 9650–9660

  28. [28]

    Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer. 2023. Grounding large language models in interactive environments with online reinforcement learning. In International Conference on Machine Learning . PMLR, 3676–3713

  29. [29]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology 15, 3 (2024), 1–45

  30. [30]

    Yevgen Chebotar, Quan Vuong, Karol Hausman, Fei Xia, Yao Lu, Alex Irpan, Aviral Kumar, Tianhe Yu, Alexander Herzog, Karl Pertsch, et al. 2023. Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions. 40 Liang et al. In Conference on Robot Learning . PMLR, 3909–3928

  31. [31]

    Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems 34 (2021), 15084–15097

  32. [32]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  33. [33]

    Po-Lin Chen and Cheng-Shang Chang. 2023. Interact: Exploring the potentials of chatgpt as a cooperative agent. arXiv preprint arXiv:2308.01552 (2023)

  34. [34]

    Shang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai, and Shao-Hua Sun. 2023. Diffusion model- augmented behavioral cloning. arXiv preprint arXiv:2302.13335 (2023)

  35. [35]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning . PmLR, 1597–1607

  36. [36]

    Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song

  37. [37]

    Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu, Mithun George Jacob, Tingnan Zhang, Tsang-Wei Edward Lee, Wenhao Yu, Connor Schenck, David Rendleman, Dhruv Shah, et al. 2024. Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs. arXiv preprint arXiv:2407.07775 (2024)

  38. [38]

    WL Chiang et al. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 With 90% ChatGPT Quality. Accessed: Apr. 14, 2023

  39. [39]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1–113

  40. [40]

    Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. 2023. Diffusion models in vision: A survey. IEEE transactions on pattern analysis and machine intelligence 45, 9 (2023), 10850–10869

  41. [41]

    Sahith Dambekodi, Spencer Frazier, Prithviraj Ammanabrolu, and Mark O Riedl. 2020. Playing text-based games with common sense. arXiv preprint arXiv:2012.02757 (2020)

  42. [42]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) . 4171–4186

  43. [43]

    Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al. 2024. Understanding world or predicting future? a comprehensive survey of world models. Comput. Surveys (2024)

  44. [44]

    Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. 2024. Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation. arXiv preprint arXiv:2408.11812 (2024)

  45. [45]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)

  46. [46]

    Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, et al. 2023. Palm-e: An embodied multimodal language model. (2023)

  47. [47]

    Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel

  48. [48]

    Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. 2022. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6, 2 (2022), 230–244

  49. [49]

    Advances in neural information processing systems 36 (2023), 9156–9172

    Learning universal policies via text-guided video generation. Advances in neural information processing systems 36 (2023), 9156–9172

  50. [50]

    Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12873–12883

  51. [51]

    Jonas Eschmann. 2021. Reward function design in reinforcement learning.Reinforcement learning algorithms: Analysis and Applications (2021), 25–33

  52. [52]

    Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning . PMLR, 1126–1135

  53. [53]

    Rasool Fakoor, Pratik Chaudhari, Stefano Soatto, and Alexander J Smola. 2019. Meta-q-learning. arXiv preprint arXiv:1910.00125 (2019)

  54. [54]

    Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its nature, scope, limits, and consequences. Minds and machines 30, 4 (2020), 681–694

  55. [55]

    Pete Florence, Corey Lynch, Andy Zeng, Oscar A Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson. 2022. Implicit behavioral cloning. In Conference on robot learning . PMLR, 158–168. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning 41

  56. [56]

    Gene F Franklin, J David Powell, Abbas Emami-Naeini, and J David Powell. 2002. Feedback control of dynamic systems . Vol. 4. Prentice hall Upper Saddle River

  57. [57]

    Maria Fox and Derek Long. 2003. PDDL2. 1: An extension to PDDL for expressing temporal planning domains. Journal of artificial intelligence research 20 (2003), 61–124

  58. [58]

    Zipeng Fu, Tony Z Zhao, and Chelsea Finn. 2024. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation. arXiv preprint arXiv:2401.02117 (2024)

  59. [59]

    Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219 (2020)

  60. [60]

    Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox. 2023. Rvt: Robotic view transformer for 3d object manipulation. In Conference on Robot Learning . PMLR, 694–710

  61. [61]

    Michael Gelfond and Yulia Kahl. 2014. Knowledge representation, reasoning, and the design of intelligent agents: The answer-set programming approach. Cambridge University Press

  62. [62]

    Jiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu, Montserrat Gonzalez Arenas, Kanishka Rao, Wenhao Yu, Chuyuan Fu, Keerthana Gopalakrishnan, Zhuo Xu, et al. 2023. Rt-trajectory: Robotic task generalization via hindsight trajectory sketches. arXiv preprint arXiv:2311.01977 (2023)

  63. [63]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 18995–19012

  64. [64]

    Lin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. 2023. Leveraging pre-trained large language models to construct and utilize world models for model-based task planning.Advances in Neural Information Processing Systems 36 (2023), 79081–79094

  65. [65]

    Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, and Yang Gao. 2023. Seer: Language instructed video prediction with latent diffusion models. arXiv preprint arXiv:2303.14897 (2023)

  66. [66]

    Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine. 2018. Meta-reinforcement learning of structured exploration strategies. Advances in neural information processing systems 31 (2018)

  67. [67]

    Yanjiang Guo, Yen-Jen Wang, Lihan Zha, and Jianyu Chen. 2024. Doremi: Grounding language model by detecting and recovering from plan-execution misalignment. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 12124–12131

  68. [68]

    Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning . Pmlr, 1861–1870

  69. [69]

    David Ha and Jürgen Schmidhuber. 2018. Recurrent world models facilitate policy evolution. Advances in neural information processing systems 31 (2018)

  70. [70]

    Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2019. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603 (2019)

  71. [72]

    Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. 2020. Mastering atari with discrete world models. arXiv preprint arXiv:2010.02193 (2020)

  72. [73]

    Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. 2019. Learning latent dynamics for planning from pixels. InInternational conference on machine learning. PMLR, 2555–2565

  73. [74]

    Asher J Hancock, Allen Z Ren, and Anirudha Majumdar. 2024. Run-time observation interventions make vision- language-action models more visually robust. arXiv preprint arXiv:2410.01971 (2024)

  74. [75]

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104 (2023)

  75. [76]

    Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang, Bin Zhao, and Xuelong Li. 2024. Large-scale actionless video pre-training via discrete diffusion for efficient policy learning. CoRR (2024)

  76. [77]

    Patrik Haslum, Nir Lipovetzky, Daniele Magazzeni, Christian Muise, Ronald Brachman, Francesca Rossi, and Peter Stone. 2019. An introduction to the planning domain definition language . Vol. 13. Springer

  77. [78]

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 16000–16009. 42 Liang et al

  78. [79]

    Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li. 2023. Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning. Advances in neural information processing systems 36 (2023), 64896–64917

  79. [80]

    Kazuki Hori, Kanata Suzuki, and Tetsuya Ogata. 2024. Interactively robot action planning with uncertainty analysis and active questioning by large language model. In 2024 IEEE/SICE International Symposium on System Integration (SII). IEEE, 85–91

  80. [81]

    Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29 (2016)

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.