Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Toward Embodied AGI: A Review of Embodied AI and the Road Ahead

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper proposes a five-level roadmap for embodied AGI and argues that current embodied AI systems sit between levels 1 and 2, lacking the omnimodal, humanoid, real-time, and generalization abilities that level 3 requires.

desk verdict A useful, honestly-hedged roadmap that recombines existing level scales for embodied AI; the L1–L2 placement is defensible but rests on an asserted omnimodality requirement. read the letter →

arxiv 2505.14235 v1 pith:WOHLSB4H submitted 2025-05-20 cs.AI

classification cs.AI
keywords EmbodiedAGIfive-leveltaxonomyroboticbrainvision-language-actionmodelomnimodalprocessinghumanoidcognitionlifelonglearninggeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Embodied artificial general intelligence, the paper argues, should be measured on a five-level ladder (L1-L5) akin to autonomous driving, from single-task robots to all-purpose humanoid agents. The authors review current embodied AI research and conclude that today's systems, including LLM-, VLM-, and VLA-based robots, fall between levels 1 and 2: they reliably complete single or composed tasks but cannot handle genuinely different task categories in real time. The paper's contribution is a systematic taxonomy built on four capability dimensions (omnimodal processing, humanoid cognition, real-time responsiveness, and generalization), plus a conceptual framework for an L3+ robotic brain that would stream all modalities at once and learn continuously. A sympathetic reader cares because the ladder gives the field a shared vocabulary for progress and a concrete diagnosis of why current models plateau below general-purpose competence.

What carries the argument

The load-bearing machinery is the five-level taxonomy (L1-L5) together with the four capability dimensions that define each level: modalities, humanoid cognitive abilities, real-time responsiveness, and generalization. The roadmap provides the measurement apparatus, since each level is a conjunction of thresholds on those four dimensions, and it is what lets the authors convert a scattered literature into a position claim that current systems sit at L1-L2. A second piece of machinery is the conceptual L3+ robotic brain, formalized as a model that at each timestep conditions its outputs (thoughts, speech, actions, mobility) on all prior multimodal inputs, which is the architectural expression of the real-time duplex requirement.

What would settle it

Take a modern VLA or omnimodal robot and test it on several unrelated task categories (for example, cooking, conversation, and navigation) in a novel environment with mid-task instruction changes and multimodal inputs including audio; if any single system completes these reliably in real time, the paper's placement of all current Embodied AI between L1 and L2 is wrong, since the paper predicts no such system exists today.

Watch

Extended reading notes

Core claim

The central discovery is a status report and a target: current Embodied AI is positioned between L1 and L2 on a proposed five-level roadmap, because no existing system meets the L3 threshold of conditional general-purpose task completion, which means handling substantially different task categories with full-spectrum multimodal perception, real-time duplex interaction, and generalization across tasks rather than just across environments. The paper argues that existing architectures, including LLMs, VLMs, VLA models, and recent omnimodal models, fall short of L3+ multimodal processing and precise real-time action execution, and that current learning paradigms such as supervised and reinforcement learning are insufficient for human-like behavior. It then derives the requirements for L3+ from four dimensions, namely omnimodal capabilities, humanoid cognitive abilities (self-awareness, social connection understanding, procedural memory, and memory reconsolidation), real-time responsiveness, and generalization, and sketches a conceptual L3+ robotic brain with a streaming joint multimodal architecture and a training paradigm combining multimodal-from-scratch pre-training, lifelong learning, and physical-oriented training.

Load-bearing premise

The framework assumes that the four dimensions, and especially the four humanoid cognitive traits (self-awareness, social connection understanding, procedural memory, and memory reconsolidation), are genuinely necessary for general-purpose embodied intelligence; if a system could achieve broad real-world competence without those traits, the L3+ requirements and the roadmap would be built on the wrong foundation.

Editorial extensions

If this is right

  • Any embodied system that claims general-purpose competence can be located on the L1-L5 ladder by scoring the four dimensions, giving the field a common benchmark language.
  • Because current systems sit at L1-L2, no existing robot should be treated as broadly general-purpose; L3+ requires crossing a qualitative threshold in task diversity, not just scaling data.
  • Future L3+ architectures will need to process streaming vision, audio, touch, and language jointly while outputting actions, speech, and internal reasoning in parallel, rather than in turn-based or semi-duplex pipelines.
  • Training must move beyond pretrain-finetune-deploy toward lifelong learning with procedural memory and memory reconsolidation, plus physical-world prediction objectives, to support inter-task generalization.
  • If the roadmap is right, progress toward Embodied AGI should be evaluated on open-ended task diversity and real-time human interaction, not only on manipulation accuracy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the taxonomy implies a concrete benchmarking program that scores each dimension separately on standardized suites, which the paper does not itself build; such a suite would make the L1-L2 claim testable.
  • Editorial inference: if the four humanoid cognitive traits are truly required for L4/L5, then social and self-referential long-horizon tasks (for example, maintaining a consistent identity and relationships over months) belong in any serious Embodied AGI benchmark, which current manipulation-centric benchmarks under-represent.
  • Editorial inference: the paper's own framework suggests a testable extension: a system with the proposed streaming architecture but without explicit self-awareness modules could reveal whether self-awareness is an emergent necessity or an optional add-on.
  • Editorial inference: the claim that existing model families fall short could be probed by attempting to fine-tune a current VLA on a radically new task category in real time; if that succeeds, the L1-L2 placement would need revision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper defines Embodied AGI and proposes a five-level roadmap (L1–L5) adapted from autonomous driving levels, with four capability dimensions (modalities, humanoid cognition, real-time responsiveness, generalization). It reviews recent embodied AI systems (GraspVLA, Helix, OpenVLA, π0.5, GO-1, etc.) and concludes that current capabilities sit between L1 and L2. It then identifies challenges in each dimension and outlines a conceptual L3+ robotic brain architecture with multimodal streaming input/output and a multi-stage training paradigm including lifelong learning and physical-oriented training.

Significance. If the taxonomy is adopted, the paper could provide a useful shared vocabulary for benchmarking embodied AI. Its strengths are the concreteness of the surveyed systems, the explicit level definitions, and the clearly labeled conceptual framework. The claim that current systems are at L1–L2 is a falsifiable assessment in principle, and the placement of named systems is plausible under the paper's own definitions. However, because the level boundaries are partly stipulated rather than derived, the empirical force of the headline claim is currently weaker than the prose suggests.

major comments (3)
  1. [Sections 2–3; Table 1] The central claim that current systems are between L1 and L2 is not independent of a threshold that the paper does not defend. Table 1 assigns L3 the requirement of 'Full' modalities and L2 only 'Partial', yet the functional definition of L3 in Section 2 is 'conditional generalization across tasks, environments, and human instructions' plus substantial real-time responsiveness, and the L3 paragraph says only that comprehensive sensory input is 'required' while listing vision, audition, and 'optionally touch and proprioception.' The systems discussed as near-L2, notably Helix and π0.5, are judged below L3 in part because they are bimodal or trimodal, not because they fail the functional criteria. A vision-language-action system with broad cross-category generalization and real-time closed-loop control would, by the Section 2 text, qualify for L3 yet be excluded by the Table 1 modalities row. The authors should either provide a task-level justification for why omnimodality is necessary for L3's functional definition, or explicitly present the L1–L2 placement as a consequence of adopting their proposed normative taxonomy rather than as an empirical measurement.
  2. [Section 4; Figure 2; Section 1] The four capability dimensions, and especially the four subcomponents of humanoid cognition (self-awareness, social connection understanding, procedural memory, memory reconsolidation), are introduced as the core constituents of Embodied AGI and then used to derive L3+ requirements. This is a stipulation, not a derivation: Section 4 says these constituents are 'derived from their definitions,' but the definitions in Table 1 already include the dimensions, so the requirements follow by construction. If one does not accept that humanoid cognition requires these four traits, the L3+ framework loses its stated foundation. The paper should either cite direct empirical evidence for the necessity of these traits for general-purpose embodied competence, or explicitly frame them as design assumptions to be validated, while softening the 'necessary' and 'requires' language in Sections 1 and 4.
  3. [Section 4 vs Table 1] There is an internal inconsistency about the role of humanoid cognition across levels. Section 4 states that 'human-like cognitive behaviors are essential across all levels (L1–L5),' while Table 1 assigns Humanoid = 'No' for L1–L3 and 'Partial' for L4. The authors should reconcile the table with this statement, for example by distinguishing between the presence of humanoid cognitive mechanisms and the degree to which they are externally exhibited, or by revising one of the two claims.
minor comments (4)
  1. [Section 1] The phrase 'for nuance social comprehension' should read 'for nuanced social comprehension.'
  2. [Section 5.1, Eq. (1)] The notation x_{0∼t}^{b1}, ..., x_{0∼t}^{bm} is undefined; please define the superscript convention and the meaning of the subscript range.
  3. [References] The in-text citation 'dri 2025' is incomplete; the full reference is to the Bosch page on automated driving, which should be cited consistently.
  4. [Figure 3] Figure 3 is dense and the idle/pointing annotations are hard to read; consider enlarging or simplifying the figure to make the modality streams legible.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the L1-L2 status claim is a classification against the paper's own stipulated taxonomy, not a fitted or self-referential prediction.

full rationale

This paper is a review and position piece: it proposes a five-level taxonomy (Table 1) and then classifies current embodied AI systems relative to that taxonomy. The central claim that current models sit between L1 and L2 is an application of the paper's own definitions, not a derivation from an external first principle, so no 'prediction' reduces to a fitted input by construction. The L3 requirement of 'Full' modalities is a stipulated threshold, and the assessment that bimodal or trimodal systems fall short is a direct application of that threshold; whether the threshold is well-motivated is a validity concern, not circularity. Section 4 states that its constituents are 'derived from their definitions,' which is definitional unpacking rather than a surprising empirical result. The only self-citation (Fan et al. 2025, an overlapping-author paper on lifelong learning) is used to support the general desideratum that social connection understanding should be lifelong, dynamic, and stateful, and it is co-cited with Zheng et al. 2025; removing it would not alter the L1-L2 assessment or the proposed L3+ framework. There are no equations fitted to data, no uniqueness theorem imported from the authors' prior work, and no renamed known result masquerading as a fresh derivation. Accordingly, no specific circular step can be quoted, and the paper is best described as non-circular with a low score reflecting only the presence of a non-load-bearing self-citation.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The central claims rest on definitions and domain assumptions, not on measured quantities; no parameters are fitted and no new physical entities are introduced.

assumptions (4)
  • domain assumption Embodied AGI should demonstrate human-like interaction and behavior.
    Definition 1 in Section 1 frames the entire roadmap on human-like proficiency and human-like settings without justifying why human-like rather than non-human-like general competence is the target.
  • domain assumption The four dimensions, modalities, humanoid cognitive abilities, real-time responsiveness, and generalization, are the core constituents of Embodied AGI capability.
    Section 1 and Figure 2 assert these dimensions without derivation; they determine the level definitions and the entire gap analysis.
  • domain assumption Humanoid cognitive behaviors require lifelong experiential learning with continuous internal state updates.
    Section 4 equates self-awareness and memory reconsolidation with lifelong learning in parameters, but no evidence is provided that such mechanisms are necessary for L3+ capability.
  • ad hoc to paper Lessons from autonomous driving levels transfer to embodied AGI.
    Section 2 borrows the five-level structure from autonomous driving, but no argument establishes that robot general intelligence scales monotonically in the same way.
invented entities (1)
  • L3+ robotic brain (conceptual architecture)
    purpose: Illustrative architecture combining omnimodal streaming input and output (Eq. 1) with lifelong and physical-oriented training.
    Explicitly described as illustrative and replaceable in Section 5; no implementation or benchmark is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Embodied AGI: A Review of Embodied AI and the Road Ahead." pith.science (2026). https://pith.science/paper/WOHLSB4H

@misc{pith2026250514235,
  author       = {Pith},
  title        = {Pith review of: Toward Embodied AGI: A Review of Embodied AI and the Road Ahead},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WOHLSB4H}},
  note         = {Machine review of arXiv:2505.14235}
}
read the original abstract

Artificial General Intelligence (AGI) is often envisioned as inherently embodied. With recent advances in robotics and foundational AI models, we stand at the threshold of a new era-one marked by increasingly generalized embodied AI systems. This paper contributes to the discourse by introducing a systematic taxonomy of Embodied AGI spanning five levels (L1-L5). We review existing research and challenges at the foundational stages (L1-L2) and outline the key components required to achieve higher-level capabilities (L3-L5). Building on these insights and existing technologies, we propose a conceptual framework for an L3+ robotic brain, offering both a technical outlook and a foundation for future exploration.

Figures

Figures reproduced from arXiv: 2505.14235 by the authors.

Figure 1
Figure 1. Roadmap of the five levels of Embodied AGI, inspired by the established levels of autonomous driving. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of four basic constituents of Embodied AGI. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Omnimodal Capabilities and Illustrative Model Structure. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Conceptual Framework: Illustrative Training Paradigms. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

    cs.AI 2025-06 conditional novelty 5.0 of 10

    RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...

Reference graph

Works this paper leans on

62 extracted references · 19 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    The Five Steps of Automated Driving

    2025. The Five Steps of Automated Driving. https://www.bosch-mobility.com/en/mobility-topics/the-five-steps-of-automated-driving/

  4. [4]

    Aghajanyan, A.; Yu, L.; Conneau, A.; Hsu, W.-N.; Hambardzumyan, K.; Zhang, S.; Roller, S.; Goyal, N.; Levy, O.; and Zettlemoyer, L. 2023. Scaling laws for generative mixed-modal language models. In International Conference on Machine Learning, 265--279. PMLR

  5. [5]

    AgiBot - World - Contributors; Bu, Q.; Cai, J.; Chen, L.; Cui, X.; Ding, Y.; Feng, S.; Gao, S.; He, X.; Huang, X.; Jiang, S.; Jiang, Y.; Jing, C.; Li, H.; Li, J.; Liu, C.; Liu, Y.; Lu, Y.; Luo, J.; Luo, P.; Mu, Y.; Niu, Y.; Pan, Y.; Pang, J.; Qiao, Y.; Ren, G.; Ruan, C.; Shan, J.; Shen, Y.; Shi, C.; Shi, M.; Shi, M.; Sima, C.; Song, J.; Wang, H.; Wang, W....

  6. [6]

    B.; Bout, B.; Chaplot, D

    Agrawal, P.; Antoniak, S.; Hanna, E. B.; Bout, B.; Chaplot, D. S.; Chudnovsky, J.; Costa, D.; Monicault, B. D.; Garg, S.; Gervet, T.; Ghosh, S.; H \' e liou, A.; Jacob, P.; Jiang, A. Q.; Khandelwal, K.; Lacroix, T.; Lample, G.; de Las Casas, D.; Lavril, T.; Scao, T. L.; Lo, A.; Marshall, W.; Martin, L.; Mensch, A.; Muddireddy, P.; Nemychnikova, V.; Pellat...

  7. [7]

    M.; and LeDoux, J

    Alberini, C. M.; and LeDoux, J. E. 2013. Memory reconsolidation. Current Biology, 23(17): R746--R750

  8. [8]

    Bar, A.; Zhou, G.; Tran, D.; Darrell, T.; and LeCun, Y. 2024. Navigation world models. arXiv preprint arXiv:2412.03572

Show all 62 references
  1. [9]

    Bayer, M.; and Reuter, C. 2024. Activellm: Large language model-based active learning for textual few-shot scenarios. arXiv preprint arXiv:2405.10808

  2. [10]

    G.; Gopalakrishnan, K.; Han, K.; Hausman, K.; Herzog, A.; Hsu, J.; Ichter, B.; Irpan, A.; Joshi, N

    Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; Florence, P.; Fu, C.; Arenas, M. G.; Gopalakrishnan, K.; Han, K.; Hausman, K.; Herzog, A.; Hsu, J.; Ichter, B.; Irpan, A.; Joshi, N. J.; Julian, R.; Kalashn...

  3. [11]

    D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al

    Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877--1901

  4. [12]

    T.; Li, Y.; Lundberg, S

    Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S. M.; Nori, H.; Palangi, H.; Ribeiro, M. T.; and Zhang, Y. 2023. Sparks of Artificial General Intelligence: Early experiments with GPT-4 . CoRR, abs/2303.12712

  5. [13]

    W.; Allen, J

    Cavaco, S.; Anderson, S. W.; Allen, J. S.; Castro-Caldas, A.; and Damasio, H. 2004. The scope of preserved procedural memory in amnesia. Brain, 127(8): 1853--1867

  6. [14]

    Cheang, C.; Chen, G.; Jing, Y.; Kong, T.; Li, H.; Li, Y.; Liu, Y.; Wu, H.; Xu, J.; Yang, Y.; Zhang, H.; and Zhu, M. 2024. GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation. CoRR, abs/2410.06158

  7. [15]

    Chen, H.; Chen, H.; Yan, M.; Xu, W.; Gao, X.; Shen, W.; Quan, X.; Li, C.; Zhang, J.; Huang, F.; et al. 2024. Socialbench: Sociality evaluation of role-playing conversational agents. arXiv preprint arXiv:2403.13679

  8. [16]

    C.; and Croft, W

    Clausner, T. C.; and Croft, W. 1999. Domains and image schemas. Cognitive linguistics, 10: 1--32

  9. [17]

    Dahl, E. 2024. Incarnation, pain, theology: a phenomenology of the body. Northwestern University Press

  10. [18]

    Dao, T. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691

  11. [19]

    DeepSeek - AI; Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; Zhang, X.; Yu, X.; Wu, Y.; Wu, Z. F.; Gou, Z.; Shao, Z.; Li, Z.; Gao, Z.; Liu, A.; Xue, B.; Wang, B.; Wu, B.; Feng, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.;...

  12. [20]

    D \' e fossez, A.; Mazar \' e , L.; Orsini, M.; Royer, A.; P \' e rez, P.; J \' e gou, H.; Grave, E.; and Zeghidour, N. 2024. Moshi: a speech-text foundation model for real-time dialogue. CoRR, abs/2410.00037

  13. [21]

    Deng, S.; Yan, M.; Wei, S.; Ma, H.; Yang, Y.; Chen, J.; Zhang, Z.; Yang, T.; Zhang, X.; Cui, H.; Zhang, Z.; and Wang, H. 2025. GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data. CoRR, 2505.03233

  14. [22]

    Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for ...

  15. [23]

    Fan, S.; Huang, X.; Yao, Y.; Fang, X.; Liu, K.; Han, P.; Shang, S.; Sun, A.; and Wang, Y. 2025. If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs. arXiv preprint arXiv:2503.23514

  16. [24]

    Feng, T.; Jin, C.; Liu, J.; Zhu, K.; Tu, H.; Cheng, Z.; Lin, G.; and You, J. 2024. How Far Are We From AGI: Are LLMs All We Need? arXiv preprint arXiv:2405.10313

  17. [25]

    Gallagher, S. 2000. Philosophical conceptions of the self: implications for cognitive science. Trends in cognitive sciences, 4(1): 14--21

  18. [26]

    Garrido, Q.; Assran, M.; Ballas, N.; Bardes, A.; Najman, L.; and LeCun, Y. 2024. Learning and Leveraging World Models in Visual Representation Learning. CoRR, abs/2403.00504

  19. [27]

    Gemini. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530

  20. [28]

    K.; et al

    Gu, Z.; Li, J.; Shen, W.; Yu, W.; Xie, Z.; McCrory, S.; Cheng, X.; Shamsah, A.; Griffin, R.; Liu, C. K.; et al. 2025. Humanoid locomotion and manipulation: Current progress and challenges in control, planning, and learning. arXiv preprint arXiv:2501.02116

  21. [29]

    Hu, Y.; Guo, Y.; Wang, P.; Chen, X.; Wang, Y.-J.; Zhang, J.; Sreenath, K.; Lu, C.; and Chen, J. 2024. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. arXiv preprint arXiv:2412.14803

  22. [30]

    Hu, Y.; Lin, F.; Zhang, T.; Yi, L.; and Gao, Y. 2023. Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning. CoRR, abs/2311.17842

  23. [31]

    Huang, W.; Wang, C.; Zhang, R.; Li, Y.; Wu, J.; and Fei - Fei, L. 2023. VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models. In Tan, J.; Toussaint, M.; and Darvish, K., eds., Conference on Robot Learning, CoRL 2023, 6-9 November 2023, Atlanta, GA, ...

  24. [32]

    Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El - Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; Iftimie, A.; Karpenko, A.; Passos, A. T.; Neitz, A.; Prokofiev, A.; Wei, A.; Tam, A.; Bennett, A.; Kumar, A.; Saraiva, A.; Vallone, A.; Duberstein, A.; Kon...

  25. [33]

    S.; Venkatesh, V

    Kannan, S. S.; Venkatesh, V. L. N.; and Min, B. 2024. SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2024, Abu Dhabi, United Arab Emirates, October 14-18, 2024 , 12140--...

  26. [34]

    J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E

    Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E. P.; Sanketi, P. R.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2024. OpenVLA: An Open-Source Vision-Language-Action Mo...

  27. [35]

    A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al

    Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A. A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13): 3521--3526

  28. [36]

    Liang, Y.; Song, Z.; Wang, H.; and Zhang, J. 2024. Learning to trust your feelings: Leveraging self-awareness in llms for hallucination mitigation. arXiv preprint arXiv:2401.15449

  29. [37]

    Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual Instruction Tuning. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, N...

  30. [38]

    D.; Zhang, H.; Zhang, T.; Sun, W.; Li, Y.; Vasilakos, A

    Liu, J.; Shi, X.; Nguyen, T. D.; Zhang, H.; Zhang, T.; Sun, W.; Li, Y.; Vasilakos, A. V.; Iacca, G.; Khan, A. A.; et al. 2025. Neural Brain: A Neuroscience-inspired Framework for Embodied Agents. arXiv preprint arXiv:2505.07634

  31. [39]

    Liu, Y.; Chen, W.; Bai, Y.; Liang, X.; Li, G.; Gao, W.; and Lin, L. 2024. Aligning cyber space with physical world: A comprehensive survey on embodied ai. arXiv preprint arXiv:2407.06886

  32. [40]

    Lopez-Paz, D.; and Ranzato, M. 2017. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30

  33. [41]

    Meta. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta-llama-3/

  34. [42]

    Metzinger, T. 2004. Being no one: The self-model theory of subjectivity. mit Press

  35. [43]

    R.; Sohl-Dickstein, J.; Fiedel, N.; Warkentin, T.; Dafoe, A.; Faust, A.; Farabet, C.; and Legg, S

    Morris, M. R.; Sohl-Dickstein, J.; Fiedel, N.; Warkentin, T.; Dafoe, A.; Faust, A.; Farabet, C.; and Legg, S. 2023. Levels of AGI for Operationalizing Progress on the Path to AGI. arXiv preprint arXiv:2311.02462

  36. [44]

    OpenAI. 2023. GPT-4 Technical Report. CoRR, abs/2303.08774

  37. [45]

    L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022. Training language mo...

  38. [46]

    Physical-Intelligence. 2025. _ 0.5 : a Vision-Language-Action Model with Open-World Generalization. arXiv preprint arXiv:2504.16054

  39. [47]

    Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. CoRR, abs/1707.06347

  40. [48]

    Sumers, T.; Yao, S.; Narasimhan, K.; and Griffiths, T. 2023. Cognitive architectures for language agents. Transactions on Machine Learning Research

  41. [49]

    Sun, Y.; Dong, L.; Patra, B.; Ma, S.; Huang, S.; Benhaim, A.; Chaudhary, V.; Song, X.; and Wei, F. 2023. A Length-Extrapolatable Transformer. In Rogers, A.; Boyd - Graber, J. L.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Computational...

  42. [50]

    Tan, C.; and Jaiswal, S. 2023. The path to AGI goes through embodiment. In Proceedings of the AAAI Symposium Series, volume 1, 104--108

  43. [51]

    Tong, P.; Brown, E.; Wu, P.; Woo, S.; IYER, A. J. V.; Akula, S. C.; Yang, S.; Yang, J.; Middepogu, M.; Wang, Z.; et al. 2024. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. Advances in Neural Information Processing Systems, 37: 87310--87356

  44. [52]

    Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Canton - Ferrer, C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goy...

  45. [53]

    N.; Kaiser, .; and Polosukhin, I

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, .; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30

  46. [54]

    Wang, C.; Wang, R.; Mandlekar, A.; Fei - Fei, L.; Savarese, S.; and Xu, D. 2021. Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2021...

  47. [55]

    Wang, H.; Shi, H.; Tan, S.; Qin, W.; Wang, W.; Zhang, T.; Nambi, A.; Ganu, T.; and Wang, H. 2024 a . Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models. arXiv preprint arXiv:2406.11230

  48. [56]

    Wang, P.; Lu, S.; Tang, Y.; Yan, S.; Xiong, Y.; and Xia, W. 2024 b . A Full-duplex Speech Dialogue Scheme Based On Large Language Models. CoRR, abs/2405.19487

  49. [57]

    Wang, S.; Zhu, Y.; Liu, H.; Zheng, Z.; Chen, C.; and Li, J. 2024 c . Knowledge editing for large language models: A survey. ACM Computing Surveys, 57(3): 1--37

  50. [58]

    Xu, J.; Guo, Z.; He, J.; Hu, H.; He, T.; Bai, S.; Chen, K.; Wang, J.; Fan, Y.; Dang, K.; et al. 2025. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215

  51. [59]

    Zeng, A.; Du, Z.; Liu, M.; Zhang, L.; Jiang, S.; Dong, Y.; and Tang, J. 2024. Scaling Speech-Text Pre-training with Synthetic Interleaved Data. CoRR, abs/2411.17607

  52. [60]

    Zhang, X.; Chen, Y.; Hu, S.; Han, X.; Xu, Z.; Xu, Y.; Zhao, W.; Sun, M.; and Liu, Z. 2024. Beyond the turn-based game: Enabling real-time conversations with duplex models. arXiv preprint arXiv:2406.15718

  53. [61]

    Z.; Tompson, J.; Driess, D.; Florence, P.; Ghasemipour, S

    Zhao, T. Z.; Tompson, J.; Driess, D.; Florence, P.; Ghasemipour, S. K. S.; Finn, C.; and Wahid, A. 2024. ALOHA Unleashed: A Simple Recipe for Robot Dexterity. In Agrawal, P.; Kroemer, O.; and Burgard, W., eds., Conference on Robot Learning, 6-9 November 2024, Munich, Germany, ...

  54. [62]

    Zheng, J.; Shi, C.; Cai, X.; Li, Q.; Zhang, D.; Li, C.; Yu, D.; and Ma, Q. 2025. Lifelong Learning of Large Language Model based Agents: A Roadmap. arXiv preprint arXiv:2501.07278

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.