REVIEW 3 major objections 6 minor 141 references
Embodied Intelligence: The Key to Unblocking Generalized Artificial Intelligence
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper argues that closed-loop embodied interaction is essential for artificial general intelligence, not merely an optional enhancement.
desk verdict Readable survey of embodied AI that overclaims a necessity result; useful overview with fixable but real reliability issues. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the closed-loop modular architecture of embodied intelligence, decomposed into four components: perception (multimodal sensor fusion), intelligent decision-making (environmental understanding, task planning, decision generation, and a learning-and-evolution framework), action (motion control and feedback adjustment), and feedback (perceptual, decision, and action feedback). The loop is the load-bearing mechanism: feedback re-enters perception and decision-making so that behavior and cognition are continuously reshaped by the environment. The paper maps these modules onto the six AGI principles adopted from the 'Levels of AGI' framework, using that mapping as the bridge between EAI and AGI.
What would settle it
Find an intelligent system that achieves the six AGI principles—as operationalized by Morris et al. (2023)—while having no physical body and no real-time interaction loop with an environment (for example, a purely offline-trained model scored on embodied benchmarks), or show that an embodied agent with the feedback module removed still reaches the same level of generalization; either result would falsify the claim that embodiment is essential.
Extended reading notes
Core claim
The paper's central claim is that embodiment is a necessary condition for AGI. It proposes a modular closed-loop architecture—perception gathers multimodal sensory data, decision-making plans and generates actions, action executes motion through the physical body, and feedback monitors outcomes and optimizes the loop—and argues that this loop is the mechanism by which a system can satisfy the six AGI principles formulated by DeepMind: focusing on capabilities, generality, cognitive/metacognitive tasks, potential, ecological validity, and a long-term development path. For each principle, the paper identifies which module or module interaction operationalizes it, concluding that the integration of dynamic learning and real-world interaction is what separates AGI from narrow AI.
Load-bearing premise
The argument rests on the assumption that the four-module perception–decision–action–feedback decomposition is a faithful and complete representation of embodied intelligence, and that DeepMind's six AGI principles are the right yardstick for AGI; if either fails, the mapping is one contingent framing rather than a systematic finding.
Editorial extensions
If this is right
- If embodiment is essential, AGI research should concentrate on real-time physical interaction and closed-loop learning rather than scaling static datasets alone.
- A system that lacks a feedback module—one that perceives and decides but does not monitor and correct its own actions—would fall short of AGI under the paper's criteria.
- Modular embodied architectures will remain a viable route to AGI, particularly where interpretability and independent module optimization matter, even as end-to-end systems push toward global optimization.
- Progress toward AGI should be evaluated on embodied, ecologically valid tasks that exercise the full loop, not only on text or image benchmarks.
Reading between the lines
- If the embodied thesis is correct, performance on physically interactive, closed-loop tasks should be a stronger predictor of AGI capability than performance on passive recognition or generation benchmarks—this is a testable prediction the authors imply but do not state.
- The paper's four-module taxonomy suggests a concrete design test: an architecture that removes or weakens any one module (for instance, an open-loop action module) should show a measurable ceiling in transfer and generalization, a comparison the field could run on existing robot benchmarks.
- The modular-versus-end-to-end framing implies a future hybrid—end-to-end perception-to-action cores augmented by explicit feedback pathways—rather than a winner-take-all outcome between the two paradigms.
- The DeepMind six-principle mapping could be operationalized into a checklist scoring embodied systems, turning a conceptual argument into an evaluation rubric.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a review-style position paper arguing that Embodied Artificial Intelligence (EAI) is the essential bridge from narrow AI to AGI. It proposes a technical taxonomy dividing EAI into end-to-end and modular architectures, then analyzes the modular architecture through four components (perception, decision-making, action, feedback). The paper maps each module onto DeepMind's six AGI principles, surveys recent techniques and industrial trends, and concludes that EAI's integration of dynamic learning and real-world interaction is essential for AGI. The paper contains no new experiments or derivations; its contribution is a conceptual framework and a broad literature survey.
Significance. If the central thesis were established, the paper would provide a useful organizing framework for AGI research: it offers a clear modular decomposition, a rich collection of recent references, and a concrete mapping between technical modules and AGI evaluation principles. The taxonomy of end-to-end versus modular architectures and the detailed module-by-module discussion are valuable reference material, and the paper is explicit about several open challenges. However, the load-bearing claim that physical embodiment is essential for AGI is not derived from the evidence presented; the paper's own analyses support only the weaker conclusion that EAI is one promising pathway. The significance is therefore conditional on a substantial reframing and additional argumentation.
major comments (3)
- [Abstract; Sections 4 and 6] The abstract's claim that EAI's integration of dynamic learning and real-world interaction is "essential" for AGI is not supported by the body of the paper. Sections 4.1-4.4 show at most that each modular component can contribute to or "align with" one or more of DeepMind's six principles; they do not rule out non-embodied systems that satisfy the same capability-oriented criteria. Indeed, Section 4 itself notes that DeepMind's principles focus on model capabilities, not processes, and physical embodiment is an implementation attribute. To support a necessity claim, the paper would need either a comparative analysis of non-embodied agents (e.g., tool-using language models operating in digital environments) or an explicit argument for why the capability criteria cannot be met without physical interaction. Without such an argument, the strongest conclusion available is that EAI is a promising pathway toward AGI, and the manuscript should be revised to state that conclusion rather than the current 'essential' claim.
- [Section 3.1 and Section 4] The paper defines two EAI paradigms, end-to-end and modular, in Section 3.1, but Section 4 analyzes only the modular four-component architecture in relation to the six AGI principles. End-to-end systems are described in Section 3.2 as major industrial approaches, yet their connection to AGI principles is never examined. Consequently, even the weaker claim that EAI contributes to AGI is incomplete: the analysis covers only one of the two paradigms that the paper itself identifies. Either the AGI mapping must be extended to end-to-end architectures, or the scope of the contribution claim should be explicitly limited to modular EAI.
- [Section 3.2.3 and Section 2.4] Several factual claims that support the narrative are unsupported or appear inaccurate. Section 3.2.3 states that Volkswagen's 'digital twin' pipelines achieve "78% cross-domain policy transferability" without any citation or methodological detail; as written, this is an unverifiable numerical assertion. Similarly, Section 2.4 attributes to reference [28] the development of "physics-informed neural controllers capable of adapting to environmental perturbations within 200ms latency," but reference [28] is a self-supervised correspondence paper for model-based reinforcement learning and does not appear to contain this claim. These unsupported numbers undermine the paper's reliability as a survey and should be either properly sourced and explained or removed.
minor comments (6)
- [Section 1 (Introduction)] The roadmap at the end of Section 1 is inconsistent with the actual structure: it states that Section 4 discusses future trends and challenges and Section 5 summarizes, whereas in the manuscript Section 4 covers the four modules, Section 5 covers future prospects and challenges, and Section 6 is the conclusion.
- [Section 4.1] The text describes the perceptual process as comprising "six critical steps," while the Figure 3 caption says "five steps" and lists only five items; the count and the caption should be reconciled.
- [Section 2.4] The three "fundamental advancements" listed in Section 2.4 are stated without supporting citations; given that the paper is a survey, each bullet should be accompanied by a specific reference.
- [Section 4.3] The phrase "The principle of autonomy is is demonstrated in this process" contains a duplicated word and a typo; the sentence should be rewritten.
- [Table 3] The column header "Innovative Industries" appears to be a mistranslation; it likely should read "Innovative Methods" or "Innovative Technologies." Additionally, the table lists year information in the same column as the method name, which is visually confusing.
- [Section 2.5] The prose in Section 2.5 (for example, "dialectical synthesis of symbolic priors and physical instantiation") is considerably more speculative and abstract than the rest of the survey, and would benefit from concrete examples or pointers to specific systems that instantiate these claims.
Circularity Check
The paper's strongest claim is self-confirming: it defines AGI by environmental interaction and EAI by environmental interaction, then concludes EAI is essential for AGI.
-
self definitional
[Section 1 (EAI definition), Section 2.2 / Table 1 (AGI operational feature), Abstract and Section 6 (conclusion)]
"Table 1 lists AGI's operational feature as 'Self-learning through interaction with the environment' while ANI 'Runs through a fixed programming framework'. Section 1 defines 'Embodied intelligence (EAI) ... a system in which an agent perceives, learns and makes decisions through the interaction between its body and its environment.' The abstract concludes: 'EAI's integration of dynamic learning and real-world interaction is essential for bridging the gap between narrow AI and AGI.'"
The necessity claim is obtained by construction from the paper's own definitions. If AGI's defining operational feature is 'self-learning through interaction with the environment' and EAI is defined as perceiving/learning/deciding through body-environment interaction, then 'EAI is essential for AGI' is already contained in the premises. Section 4's module-to-principle mapping does not supply independent evidence for this necessity: the adopted DeepMind criteria are capability-focused and process-agnostic, so they cannot by themselves force a physical-embodiment requirement. The survey organizes embodied AI research in a useful way, but the 'essential' verdict is a definitional consequence of the chosen AGI characterization rather than a result derived from the surveyed evidence.
full rationale
This manuscript is a review, not a quantitative derivation, and it contains no fitted-parameter predictions and no load-bearing self-citations: the references are external (e.g., DeepMind's 'Levels of AGI' [29], Brooks, Pfeifer and Scheier). The only significant circularity is definitional: AGI is characterized in Section 2.2 by 'self-learning through interaction with the environment,' and EAI is defined in Section 1 as perception, learning, and decision-making through body-environment interaction; the abstract then asserts interaction is essential for AGI. That makes the headline conclusion true by stipulation rather than by independent evidence. The four-module survey has genuinely independent descriptive content and is not manufactured, but the mapping in Section 4 does not establish a necessity claim because the adopted six-principle yardstick is process-agnostic and because only the modular architecture is analyzed, leaving end-to-end EAI unexamined in that mapping. These are assessment gaps that belong to correctness risk, not to circularity. On balance, the paper is mildly self-confirming in its framing but not systematically circular, so a low-to-moderate score of 3 is appropriate.
Assumptions & free parameters
assumptions (3)
- domain assumption The four-module decomposition (perception, decision-making, action, feedback) is a faithful and complete description of embodied intelligence systems.
- domain assumption DeepMind's six AGI principles, adopted from reference [29], are the appropriate criteria for evaluating progress toward AGI.
- domain assumption The cited references support the specific factual claims made in the text.
Cite this review
Pith. "Pith review of Embodied Intelligence: The Key to Unblocking Generalized Artificial Intelligence." pith.science (2026). https://pith.science/paper/4IEZWIB5
@misc{pith2026250506897,
author = {Pith},
title = {Pith review of: Embodied Intelligence: The Key to Unblocking Generalized Artificial Intelligence},
year = {2026},
howpublished = {\url{https://pith.science/paper/4IEZWIB5}},
note = {Machine review of arXiv:2505.06897}
}
read the original abstract
The ultimate goal of artificial intelligence (AI) is to achieve Artificial General Intelligence (AGI). Embodied Artificial Intelligence (EAI), which involves intelligent systems with physical presence and real-time interaction with the environment, has emerged as a key research direction in pursuit of AGI. While advancements in deep learning, reinforcement learning, large-scale language models, and multimodal technologies have significantly contributed to the progress of EAI, most existing reviews focus on specific technologies or applications. A systematic overview, particularly one that explores the direct connection between EAI and AGI, remains scarce. This paper examines EAI as a foundational approach to AGI, systematically analyzing its four core modules: perception, intelligent decision-making, action, and feedback. We provide a detailed discussion of how each module contributes to the six core principles of AGI. Additionally, we discuss future trends, challenges, and research directions in EAI, emphasizing its potential as a cornerstone for AGI development. Our findings suggest that EAI's integration of dynamic learning and real-world interaction is essential for bridging the gap between narrow AI and AGI.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[28]
Key- points into the future: Self-supervised correspondence in model- based reinforcement learning
Lucas Manuelli, Yunzhu Li, Pete Florence, and Russ Tedrake. Key- points into the future: Self-supervised correspondence in model- based reinforcement learning. arXiv preprint arXiv:2009.05085, 2020
arXiv 2009
-
[1]
FeiDou,JinYe,GengYuan,QinLu,WeiNiu,HaijianSun,LeGuan, Guoyu Lu, Gengchen Mai, Ninghao Liu, et al. Towards artificial generalintelligence(agi)intheinternetofthings(iot):Opportunities and challenges.arXiv preprint arXiv:2309.07438, 2023
arXiv 2023
-
[2]
N.Roy,I.Posner,T.Barfoot,P.Beaudoin,Y.Bengio,J.Bohg,etal. Frommachinelearningtorobotics:Challengesandopportunitiesfor embodied intelligence.arXiv preprint arXiv:2110.15245, 2021
arXiv 2021
-
[3]
A path toward explainable ai and autonomous adaptiveintelligence:deeplearning,adaptiveresonance,andmodels of perception, emotion, and action
Stephen Grossberg. A path toward explainable ai and autonomous adaptiveintelligence:deeplearning,adaptiveresonance,andmodels of perception, emotion, and action. Frontiers in neurorobotics, 14:36, 2020
2020
-
[4]
Embodied intelligence in soft robotics through hardwaremultifunctionality
Matteo Cianchetti. Embodied intelligence in soft robotics through hardwaremultifunctionality. FrontiersinRoboticsandAI ,8:724056, 2021
2021
-
[5]
Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li, and Jabbar Abdul. Recent advances and trends in multimodal deep learning: A review.arXiv preprint arXiv:2105.11087, 2021
arXiv 2021
-
[6]
Embodiedun- derstanding of driving scenarios.arXiv preprint arXiv:2403.04593, 2024
Y.Zhou,L.Huang,Q.Bu,J.Zeng,T.Li,H.Qiu,etal. Embodiedun- derstanding of driving scenarios.arXiv preprint arXiv:2403.04593, 2024
arXiv 2024
-
[7]
Y. Liu, W. Chen, Y. Bai, J. Luo, X. Song, K. Jiang, et al. Aligning cyber space with physical world: A comprehensive survey on em- bodied ai.arXiv preprint arXiv:2407.06886, 2024
arXiv 2024
Show all 141 references
-
[8]
Ai embodiment through 6g: Shaping the future of agi.IEEE Wireless Communications, 2024
Lina Bariah and Mérouane Debbah. Ai embodiment through 6g: Shaping the future of agi.IEEE Wireless Communications, 2024
2024
-
[9]
A. M. Turing. Computing machinery and intelligence. Mind, 59(236):433, 1950
1950
-
[10]
Artificial human intelligence: The role of hu- mans in the development of next generation ai
Suayb S Arslan. Artificial human intelligence: The role of hu- mans in the development of next generation ai. arXiv preprint arXiv:2409.16001, 2024
2024
-
[11]
MIT press, 2008
Dario Floreano and Claudio Mattiussi.Bio-inspired artificial intel- ligence: theories, methods, and technologies. MIT press, 2008
2008
-
[12]
Adams, I
S. Adams, I. Arel, J. Bach, R. Coop, R. Furlan, B. Goertzel, et al. Mappingthelandscapeofhuman-levelartificialgeneralintelligence. AI Magazine, 33(1):25–42, 2012
2012
-
[13]
Silver, S
D. Silver, S. Singh, D. Precup, and R. S. Sutton. Reward is enough. Artificial Intelligence, 299:103535, 2021
2021
-
[14]
generalproblemsolver
S.C.Garvey. The“generalproblemsolver”doesnotexist:Mortimer taube and the art of ai criticism.IEEE Annals of the History of Computing, 43(1):60–73, 2021
2021
-
[15]
Moto-Oka
T. Moto-Oka. Overview to the fifth generation computer system project. InProceedingsofthe10thAnnualInternationalSymposium on Computer Architecture, pages 417–422, 1983
1983
-
[16]
Roland and P
A. Roland and P. Shiman.Strategic Computing: DARPA and the Quest for Machine Intelligence, 1983-1993. MIT Press, 2002
1983
-
[17]
Goertzel
B. Goertzel. Artificial general intelligence: Concept, state of the art, and future prospects.Journal of Artificial General Intelligence, 5(1):1, 2014
2014
-
[18]
T. J. Sejnowski. The unreasonable effectiveness of deep learning in artificial intelligence. Proceedings of the National Academy of Sciences, 117(48):30033–30038, 2020
2020
-
[19]
Openagi: When llm meets domain experts.Advances in Neural Information Processing Systems, 36, 2024
Y.Ge,W.Hua,K.Mei,J.Tan,S.Xu,Z.Li,andY.Zhang. Openagi: When llm meets domain experts.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[20]
Pfeifer and F
R. Pfeifer and F. Iida. Embodied artificial intelligence: Trends and challenges. Lecture Notes in Computer Science, pages 1–26, 2004. : Preprint submitted to Elsevier Page 16 of 19
2004
-
[21]
R. A. Brooks. Intelligence without representation. Artificial Intelligence, 47(1-3):139–159, 1991
1991
-
[22]
Pfeifer and C
R. Pfeifer and C. Scheier.Understanding Intelligence. MIT Press, 2001
2001
-
[23]
L. B. Smith. Cognition as a dynamic system: Principles from embodiment. Developmental Review, 25(3-4):278–298, 2005
2005
-
[24]
Embodied cognition and learning environment design
John B Black, Ayelet Segal, Jonathan Vitale, and Cameron L Fadjo. Embodied cognition and learning environment design. In Theoretical foundations of learning environments, pages 198–223. Routledge, 2012
2012
-
[25]
Environ- mental cognition
Reginald G Golledge, Gary T Moore, Ronald Briggs, Martin T Cadwallader, Ann S Devlin, David L George, Georgia Zannaras, Aleira Kreimer, Alfred J Nigl, Harold D Fishbein, et al. Environ- mental cognition. In Environmental design research, pages 182–
-
[26]
Nikoleta Manakitsa, George S Maraslidis, Lazaros Moysis, and George F Fragulis. A review of machine learning and deep learn- ing for object detection, semantic segmentation, and human action recognition in machine and robotic vision.Technologies, 12(2):15, 2024
2024
-
[27]
In-sensor multisensory integrative perception.Available at SSRN 5128520
Tianrun Li, Zhimiao Yan, Yinghua Chen, and Ting Tan. In-sensor multisensory integrative perception.Available at SSRN 5128520
-
[29]
M. R. Morris, J. Sohl-Dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, et al. Levels of agi: Operationalizing progress on the path to agi.arXiv preprint arXiv:2311.02462, 2023
2023
-
[30]
H. M. Hegde. Autonomous Path Traversal and Object Avoidance in Cars-AirSim Simulation. PhD thesis, California State University, Northridge, 2021
2021
-
[31]
Nogueira
L. Nogueira. Comparative analysis between gazebo and v-rep robotic simulators.Seminario Interno de Cognicao Artificial-SICA, 2014(5):2, 2014
2014
-
[32]
F. Xia, W. B. Shen, C. Li, P. Kasimbeg, M. E. Tchapmi, A. Toshev, et al. Interactive gibson benchmark (igibson 0.5): A benchmark for interactive navigation in cluttered environments.arXiv preprint arXiv:2005.04307, 2020
2005 arXiv
-
[33]
Coumans and Y
E. Coumans and Y. Bai. Pybullet quickstart guide. 2021
2021
-
[34]
Fernandez-Chaves, J
D. Fernandez-Chaves, J. R. Ruiz-Sarmiento, A. Jaenal, N. Petkov, and J. Gonzalez-Jimenez. Robot@virtualhome, an ecosystem of virtualenvironmentsandtoolsforrealisticindoorroboticsimulation. Expert Systems with Applications, 208:117970, 2022
2022
-
[35]
Designandimplementationofa simulated robot using webots software
V.GanapathyandL.W.L.Dennis. Designandimplementationofa simulated robot using webots software. InProceedings of the Inter- nationalConference onControl,Instrumentationand Mechatronics Engineering, pages 28–29, 2007
2007
-
[36]
Pybullet quickstart guide.ed: Py- Bullet Quickstart Guide
Erwin Coumans and Yunfei Bai. Pybullet quickstart guide.ed: Py- Bullet Quickstart Guide. https://docs. google. com/document/u/1/d, 2021
2021
-
[37]
Shigemi, A
S. Shigemi, A. Goswami, and P. Vadakkepat. Asimo and humanoid robot research at honda.Humanoid Robotics: A Reference, 55:90, 2018
2018
-
[38]
H. Liu, D. Guo, F. Sun, W. Yang, S. Furber, and T. Sun. Embodied tactile perception and learning.Brain Science Advances, 6(2):132– 158, 2020
2020
-
[39]
Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, et al. Mp5: Amulti-modalopen-endedembodiedsysteminminecraftviaactive perception. In2024IEEE/CVFConferenceonComputerVisionand Pattern Recognition (CVPR), pages 16307–16316, 2024
2024
-
[40]
Surveyondeepmulti-modaldataanalytics:Collaboration, rivalry, and fusion.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 17(1s):1–25, 2021
Y.Wang. Surveyondeepmulti-modaldataanalytics:Collaboration, rivalry, and fusion.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 17(1s):1–25, 2021
2021
-
[41]
Winkler and B
T. Winkler and B. Rinner. Security and privacy protection in visual sensor networks: A survey. ACM Computing Surveys (CSUR), 47(1):1–42, 2014
2014
-
[42]
L. Zou, C. Ge, Z. J. Wang, E. Cretu, and X. Li. Novel tactile sensor technology and smart tactile sensing systems: A review.Sensors, 17(11):2653, 2017
2017
-
[43]
H. Guo, X. Pu, J. Chen, Y. Meng, M. H. Yeh, G. Liu, et al. A highlysensitive,self-poweredtriboelectricauditorysensorforsocial robotics and hearing aids.Science Robotics, 3(20):eaat2516, 2018
2018
-
[44]
Shorten and T
C. Shorten and T. M. Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6(1):1–48, 2019
2019
-
[45]
A. F. Elaksher and J. S. Bethel. Reconstructing 3d buildings from lidar data.International Archives of Photogrammetry Remote Sensing and Spatial Information Sciences, 34(3/A):102–107, 2002
2002
-
[46]
Big data preprocessing: Methods and prospects
S.García,S.Ramírez-Gallego,J.Luengo,J.M.Benítez,andF.Her- rera. Big data preprocessing: Methods and prospects. Big Data Analytics, 1:1–22, 2016
2016
-
[47]
J. Wu. Introduction to convolutional neural networks.National Key Lab for Novel Software Technology, 2017
2017
-
[48]
Lindemann, T
B. Lindemann, T. Müller, H. Vietz, N. Jazdi, and M. Weyrich. A survey on long short-term memory networks for time series prediction. Procedia CIRP, 99:650–655, 2021
2021
-
[49]
J. Luo, W. Wang, and H. Qi. Spatio-temporal feature extraction and representation for rgb-d human action recognition. Pattern Recognition Letters, 50:139–148, 2014
2014
-
[50]
Huang, B
K. Huang, B. Shi, X. Li, X. Li, S. Huang, and Y. Li. Multi-modal sensor fusion for auto driving perception: A survey.arXiv preprint arXiv:2202.02703, 2022
2022 arXiv
-
[51]
Evaluating ensemble learning methods for multi-modal emotion recognition using sensor data fusion.Sensors, 22(15):5611, 2022
E.M.Younis,S.M.Zaki,E.Kanjo,andE.H.Houssein. Evaluating ensemble learning methods for multi-modal emotion recognition using sensor data fusion.Sensors, 22(15):5611, 2022
2022
-
[52]
Algarni, et al
N.A.Almujally,A.A.Rafique,N.AlMudawi,A.Alazeb,M.Alon- azi, A. Algarni, et al. Multi-modal remote perception learning for object sensory data.Frontiers in Neurorobotics, 18:1427786, 2024
2024
-
[53]
Simanek, V
J. Simanek, V. Kubelka, and M. Reinstein. Improving multi-modal data fusion by anomaly detection.Autonomous Robots, 39(2):139– 154, 2015
2015
-
[54]
Onrobustnessofmulti-modal fusion—robotics perspective.Electronics, 9(7):1152, 2020
M.Bednarek,P.Kicki,andK.Walas. Onrobustnessofmulti-modal fusion—robotics perspective.Electronics, 9(7):1152, 2020
2020
-
[55]
Multi-sensordata fusionmethodbasedonself-attentionmechanism
X.Lin,S.Chao,D.Yan,L.Guo,Y.Liu,andL.Li. Multi-sensordata fusionmethodbasedonself-attentionmechanism. AppliedSciences, 13(21):11992, 2023
2023
-
[56]
Roheda, H
S. Roheda, H. Krim, and B. S. Riggan. Robust multi-modal sensor fusion: An adversarial approach. IEEE Sensors Journal, 21(2):1885–1896, 2020
2020
-
[57]
Nadon, A
F. Nadon, A. J. Valencia, and P. Payeur. Multi-modal sensing and robotic manipulation of non-rigid objects: A survey. Robotics, 7(4):74, 2018
2018
-
[58]
Noceti, B
N. Noceti, B. Caputo, C. Castellini, L. Baldassarre, A. Barla, L. Rosasco, et al. Towards a theoretical framework for learning multi-modal patterns for embodied agents. InInternational Con- ference on Image Analysis and Processing, pages 239–248, 2009
2009
-
[59]
Z. Jia, J. Wang, and R. Jin. Grnet: A graph reasoning network for enhanced multi-modal learning in scene text recognition.The Computer Journal, bxae085, 2024
2024
-
[60]
J. B. de la Cita.Multimodal Perception for Autonomous Driving. PhD thesis, Universidad Carlos III de Madrid, 2022
2022
-
[61]
Fritsch, M
J. Fritsch, M. Kleinehagenbrock, S. Lang, T. Plötz, G. A. Fink, and G. Sagerer. Multi-modal anchoring for human–robot interaction. Robotics and Autonomous Systems, 43(2-3):133–147, 2003
2003
-
[62]
Q. Tang, J. Liang, and F. Zhu. A comparative review on multi- modal sensors fusion based on deep learning.Signal Processing, page 109165, 2023
2023
-
[63]
Planninganddecision- making for autonomous vehicles
W.Schwarting,J.Alonso-Mora,andD.Rus. Planninganddecision- making for autonomous vehicles. Annual Review of Control, Robotics, and Autonomous Systems, 1(1):187–210, 2018
2018
-
[64]
J. Singh. Advancements in ai-driven autonomous robotics: Lever- agingdeeplearningforreal-timedecisionmakingandobjectrecog- nition. Journal of Artificial Intelligence Research and Applications, 3(1):657–697, 2023. : Preprint submitted to Elsevier Page 17 of 19
2023
-
[65]
Tsarouchi, S
P. Tsarouchi, S. Makris, and G. Chryssolouris. Human–robot interaction review and challenges on task planning and program- ming.InternationalJournalofComputerIntegratedManufacturing , 29(8):916–931, 2016
2016
-
[66]
Hierarchicalreinforcementlearningwiththemaxq value function decomposition
T.G.Dietterich. Hierarchicalreinforcementlearningwiththemaxq value function decomposition. Journal of Artificial Intelligence Research, 13:227–303, 2000
2000
-
[67]
O. Michel. Cyberbotics ltd. webots™: Professional mobile robot simulation. International Journal of Advanced Robotic Systems, 1(1):5, 2004
2004
-
[68]
L. Hohl, R. Tellez, O. Michel, and A. J. Ijspeert. Aibo and webots: Simulation,wirelessremotecontrolandcontrollertransfer. Robotics and Autonomous Systems, 54(6):472–485, 2006
2006
-
[69]
Robottaskplanningusingsemanticmaps
C.Galindo,J.A.Fernández-Madrigal,J.González,andA.Saffiotti. Robottaskplanningusingsemanticmaps. RoboticsandAutonomous Systems, 56(11):955–966, 2008
2008
-
[70]
To- wardhuman-awarerobottaskplanning
R.Alami,A.Clodic,V.Montreuil,E.A.Sisbot,andR.Chatila. To- wardhuman-awarerobottaskplanning. In AAAISpringSymposium: ToBoldlyGoWhereNoHuman-RobotTeamHasGoneBefore ,pages 39–46, 2006
2006
-
[71]
Paxton, Y
C. Paxton, Y. Barnoy, K. Katyal, R. Arora, and G. D. Hager. Visual robot task planning. In2019 International Conference on Robotics and Automation (ICRA), pages 8832–8838, 2019
2019
-
[72]
Gupta, S
A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei. Embodied intelligence via learning and evolution.Nature Communications, 12(1):5721, 2021
2021
-
[73]
Evolution of embodied intelligence
D.Floreano,F.Mondada,A.Perez-Uribe,andD.Roggen. Evolution of embodied intelligence. In Embodied Artificial Intelligence: International Seminar, Dagstuhl Castle, Germany, July 7-11, 2003. Revised Papers, pages 293–311, 2004
2003
-
[74]
InConference on Robot Learning, pages 1–10
Junning Huang, Sirui Xie, Jiankai Sun, Qiurui Ma, Chunxiao Liu, DahuaLin,andBoleiZhou.Learningadecisionmodulebyimitating driver’s control behaviors. InConference on Robot Learning, pages 1–10. PMLR, 2021
2021
-
[75]
Sadhu, T
A. Sadhu, T. Gupta, M. Yatskar, R. Nevatia, and A. Kembhavi. Visual semantic role labeling for video understanding. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5589–5600, 2021
2021
-
[76]
H. Guo, F. Wu, Y. Qin, R. Li, K. Li, and K. Li. Recent trends in task and motion planning for robotics: A survey.ACM Computing Surveys, 55(13s):1–36, 2023
2023
-
[77]
Policygradi- entmethodsforreinforcementlearningwithfunctionapproximation
R.S.Sutton,D.McAllester,S.Singh,andY.Mansour. Policygradi- entmethodsforreinforcementlearningwithfunctionapproximation. In Advances in Neural Information Processing Systems, volume 12, 1999
1999
-
[78]
Ho and S
J. Ho and S. Ermon. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems, volume 29, 2016
2016
-
[79]
Hospedales, A
T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey. Meta- learninginneuralnetworks:Asurvey. IEEETransactionsonPattern Analysis and Machine Intelligence, 44(9):5149–5169, 2021
2021
-
[80]
Salimans, J
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever. Evolution strategies as a scalable alternative to reinforcement learning.arXiv preprint arXiv:1703.03864, 2017
2017 arXiv
-
[81]
X. Hong, Y. Lan, L. Pang, J. Guo, and X. Cheng. Visual reasoning: Fromstatetotransformation. IEEETransactionsonPatternAnalysis and Machine Intelligence, 45(9):11352–11364, 2023
2023
-
[82]
Driess, J
D. Driess, J. S. Ha, and M. Toussaint. Deep visual reasoning: Learning to predict action sequences for task and motion planning fromaninitialsceneimage. arXivpreprintarXiv:2006.05398 ,2020
2006 arXiv
-
[83]
X. Li, F. Zhang, H. Diao, Y. Wang, X. Wang, and L. Y. Duan. Densefusion-1m: Merging vision experts for comprehensive multi- modal perception.arXiv preprint arXiv:2407.08303, 2024
2024 arXiv
-
[84]
H. Cai, Y. Wang, L. Liu, S. Zhu, and M. Chen. Dagnn: Deep autoencoder-based graph neural network for local anomaly detec- tion. In 2024 6th International Conference on Communications, Information System and Computer Engineering (CISCE), pages 988–994, 2024
2024
-
[85]
T. Zhou, T. Shi, H. Gao, and W. Rao. Learning to optimize state es- timation in multi-agent reinforcement learning-based collaborative detection. IEEE Transactions on Mobile Computing, 2024
2024
-
[86]
L. Li, T. Zhou, W. Wang, J. Li, and Y. Yang. Deep hierarchical se- mantic segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1246–1257, 2022
2022
-
[87]
L. Wang, X. Wu, Y. Zhang, X. Zhang, L. Xu, Z. Wu, and A. Fei. Deepadain-net: Deep adaptive device-edge collaborative inference for augmented reality. IEEE Journal of Selected Topics in Signal Processing, 2023
2023
-
[88]
Okubo and M
T. Okubo and M. Takahashi. Multi-agent action graph based task allocation and path planning considering changes in environment. IEEE Access, 11:21160–21175, 2023
2023
-
[89]
Multi-usvtaskplanning method based on improved deep reinforcement learning
J.Zhang,J.Ren,Y.Cui,D.Fu,andJ.Cong. Multi-usvtaskplanning method based on improved deep reinforcement learning. IEEE Internet of Things Journal, 2024
2024
-
[90]
Skilld- iffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution
Z.Liang,Y.Mu,H.Ma,M.Tomizuka,M.Ding,andP.Luo. Skilld- iffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16467–16476, 2024
2024
-
[91]
M. Klar, P. Ruediger, M. Schuermann, G. T. Gören, M. Glatt, B. Ravani, and J. C. Aurich. Explainable generative design in man- ufacturingforreinforcementlearningbasedfactorylayoutplanning. Journal of Manufacturing Systems, 72:74–92, 2024
2024
-
[92]
J. Sun, Q. Zhang, Y. Duan, X. Jiang, C. Cheng, and R. Xu. Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 16236–16242, 2024
2024
-
[93]
Liang, R
Z. Liang, R. Yang, J. Wang, L. Liu, X. Ma, and Z. Zhu. Dynamic constrained evolutionary optimization based on deep q-network. Expert Systems with Applications, 249:123592, 2024
2024
-
[94]
Afew-shot diseasediagnosisdecisionmakingmodelbasedonmeta-learningfor general practice
Q.Liu,Y.Tian,T.Zhou,K.Lyu,R.Xin,Y.Shang,etal. Afew-shot diseasediagnosisdecisionmakingmodelbasedonmeta-learningfor general practice. Artificial Intelligence in Medicine, 147:102718, 2024
2024
-
[95]
arXiv preprint arXiv:2402.08848, 2024
J.Ren,G.Swamy,Z.S.Wu,J.A.Bagnell,andS.Choudhury.Hybrid inverse reinforcement learning. arXiv preprint arXiv:2402.08848, 2024
2024 arXiv
-
[96]
J. Liu, X. Qi, P. Hang, and J. Sun. Enhancing social decision- making of autonomous vehicles: A mixed-strategy game approach with interaction orientation identification. IEEE Transactions on Vehicular Technology, 2024
2024
-
[97]
Jana, and N
G.S.Sahoo,R.H.J.Rani,M.K.Goyal,V.A.Mohammed,D.L.F. Jana, and N. N. Wasatkar. Leveraging adaptive algorithms for real- time data analysis. In2024 15th International Conference on Com- puting Communication and Networking Technologies (ICCCNT), pages 1–6, 2024
2024
-
[98]
K. Dong, Y. Luo, Y. Wang, Y. Liu, C. Qu, Q. Zhang, et al. Dyna- style model-based reinforcement learning with model-free policy optimization. Knowledge-Based Systems, 287:111428, 2024
2024
-
[99]
J. Shuford. Deep reinforcement learning unleashing the power of ai in decision-making. Journal of Artificial Intelligence General Science (JAIGS) ISSN: 3006-4023, 1(1), 2024
2024
-
[100]
Rimon, T
Z. Rimon, T. Jurgenson, O. Krupnik, G. Adler, and A. Tamar. Mamba:Aneffectiveworldmodelapproachformeta-reinforcement learning. arXiv preprint arXiv:2403.09859, 2024
2024 arXiv
-
[101]
A. Wachi. Failure-scenario maker for rule-based agent using multi- agent adversarial reinforcement learning and its application to au- tonomous driving.arXiv preprint arXiv:1903.10654, 2019
1903 arXiv
-
[102]
Y. Wu, Y. Chen, L. Wang, Y. Ye, Z. Liu, Y. Guo, and Y. Fu. Large scale incremental learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 374–382, 2019
2019
-
[103]
Astudy offorward-forwardalgorithm for self-supervised learning.arXiv preprint arXiv:2309.11955, 2023
J.Brenig andR.Timofte. Astudy offorward-forwardalgorithm for self-supervised learning.arXiv preprint arXiv:2309.11955, 2023. : Preprint submitted to Elsevier Page 18 of 19
2023 arXiv
-
[104]
Z. Wang, Y. Wu, and Q. Niu. Multi-sensor fusion in automated driving: A survey.IEEE Access, 8:2847–2868, 2019
2019
-
[105]
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age.IEEE Transactions on Robotics, 32(6):1309–1332, 2016
C.Cadena,L.Carlone,H.Carrillo,Y.Latif,D.Scaramuzza,J.Neira, et al. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age.IEEE Transactions on Robotics, 32(6):1309–1332, 2016
2016
-
[106]
T. Luo, B. Subagdja, D. Wang, and A. H. Tan. Multi-agent collabo- rativeexplorationthroughgraph-baseddeepreinforcementlearning. In2019IEEEInternationalConferenceonAgents(ICA) ,pages2–7. IEEE, October 2019
2019
-
[107]
Multi-usvcooperativechasing strategybasedonobstaclesassistanceanddeepreinforcementlearn- ing
W.Gan,X.Qu,D.Song,andP.Yao. Multi-usvcooperativechasing strategybasedonobstaclesassistanceanddeepreinforcementlearn- ing. IEEE Transactions on Automation Science and Engineering, 2023
2023
-
[108]
Cooperative and competitive multi-agent systems: From optimiza- tion to games.IEEE/CAA Journal of Automatica Sinica, 9(5):763– 783, 2022
J.Wang,Y.Hong,J.Wang,J.Xu,Y.Tang,Q.L.Han,andJ.Kurths. Cooperative and competitive multi-agent systems: From optimiza- tion to games.IEEE/CAA Journal of Automatica Sinica, 9(5):763– 783, 2022
2022
-
[109]
Atheoreticalanalysisofdeep q-learning
J.Fan,Z.Wang,Y.Xie,andZ.Yang. Atheoreticalanalysisofdeep q-learning. In Learning for dynamics and control, pages 486–489. PMLR, July 2020
2020
-
[110]
Fast context adaptation via meta-learning
L.Zintgraf,K.Shiarli,V.Kurin,K.Hofmann,andS.Whiteson. Fast context adaptation via meta-learning. InInternational Conference on Machine Learning, pages 7693–7702. PMLR, May 2019
2019
-
[111]
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015
2015
-
[112]
learning-to-communicate
H. Mao, Z. Gong, Y. Ni, and Z. Xiao. Accnet: Actor-coordinator- criticnetfor"learning-to-communicate"withdeepmulti-agentrein- forcement learning.arXiv preprint arXiv:1706.03235, 2017
2017 arXiv
-
[113]
L.JingandY.Tian.Self-supervisedvisualfeaturelearningwithdeep neural networks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4037–4058, 2020
2020
-
[114]
J. C. Yeo, H. K. Yap, W. Xi, Z. Wang, C. H. Yeow, and C. T. Lim. Flexible and stretchable strain sensing actuator for wear- able soft robotic applications. Advanced Materials Technologies, 1(3):1600018, 2016
2016
-
[115]
Areviewofshape memoryalloyresearch,applicationsandopportunities
J.M.Jani,M.Leary,A.Subic,andM.A.Gibson. Areviewofshape memoryalloyresearch,applicationsandopportunities. Materials& Design (1980-2015), 56:1078–1113, 2014
1980
-
[116]
Bar-Cohen
Y. Bar-Cohen. Electroactive polymers: current capabilities and challenges. In Smart Structures and Materials 2002: Electroactive PolymerActuatorsand Devices(EAPAD) ,volume4695, pages1–7. SPIE, July 2002
2002
-
[117]
L. Chen, Q. Hu, H. Zhang, B. Tong, X. Shi, C. Jiang, and L. Sun. Researchonunderwatermotionmodelingandclosed-loopcontrolof bionicundulatingfinrobot. OceanEngineering,299:117400,2024
2024
-
[118]
B. M. Wilamowski. Neural network architectures and learning algorithms. IEEE Industrial Electronics Magazine, 3(4):56–63, 2009
2009
-
[119]
Bhatti, H
G. Bhatti, H. Mohan, and R. R. Singh. Towards the future of smart electric vehicles: Digital twin technology. Renewable and Sustainable Energy Reviews, 141:110801, 2021
2021
-
[120]
Renewable and sustainable energy reviews
Chang-Hun Lee. Renewable and sustainable energy reviews. Coastal and Ocean, 5(2):62–70, 2012
2012
-
[121]
F. C. Gu, H. C. Chang, Y. M. Hsueh, C. C. Kuo, and B. R. Chen. Developmentofahigh-speeddataacquisitioncardforpartial discharge measurement.IEEE Access, 7:140312–140318, 2019
2019
-
[122]
Securityandprivacyprotec- tion in visual sensor networks: A survey.ACM Computing Surveys (CSUR), 47(1):1–42, 2014
ThomasWinklerandBernhardRinner. Securityandprivacyprotec- tion in visual sensor networks: A survey.ACM Computing Surveys (CSUR), 47(1):1–42, 2014
2014
-
[123]
George Thuruthel, Y
T. George Thuruthel, Y. Ansari, E. Falotico, and C. Laschi. Control strategies for soft robotic manipulators: A survey.Soft Robotics, 5(2):149–163, 2018
2018
-
[124]
Castelli, S
F. Castelli, S. Michieletto, S. Ghidoni, and E. Pagello. A machine learning-based visual servoing approach for fast robot control in in- dustrialsetting. InternationalJournalofAdvancedRoboticSystems , 14(6):1729881417738884, 2017
2017
-
[125]
G. S. Dordevic, M. Rasic, and R. Shadmehr. Parametric models for motion planning and control in biomimetic robotics. IEEE Transactions on Robotics, 21(1):80–92, 2005
2005
-
[126]
S. Duan, Q. Shi, and J. Wu. Multimodal sensors and ml-based data fusion for advanced robots. Advanced Intelligent Systems, 4(12):2200213, 2022
2022
-
[127]
Zhang, C
J. Zhang, C. Song, Y. Hu, and B. Yu. Improving robustness of roboticgraspingbyfusingmulti-sensor. In 2012IEEEInternational Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), pages 126–131. IEEE, September 2012
2012
-
[128]
Neuralnetworkmapping of industrial robots’ task times for real-time process optimization
P.Righettini,R.Strada,andF.Cortinovis. Neuralnetworkmapping of industrial robots’ task times for real-time process optimization. Robotics, 12(5):143, 2023
2023
-
[129]
D. M. Botín-Sanabria, A. S. Mihaita, R. E. Peimbert-García, M. A. Ramírez-Moreno, R. A. Ramírez-Mendoza, and J. D. J. Lozoya- Santos. Digital twin technology challenges and applications: A comprehensive review.Remote Sensing, 14(6):1335, 2022
2022
-
[130]
L. Ren, J. Dong, S. Liu, L. Zhang, and L. Wang. Embodied intelli- gence toward future smart manufacturing in the era of ai foundation model. IEEE/ASME Transactions on Mechatronics, 2024
2024
-
[131]
Yamakawa, Y
Y. Yamakawa, Y. Matsui, and M. Ishikawa. Development of a real-time human-robot collaborative system based on 1 khz visual feedback control and its application to a peg-in-hole task.Sensors, 21(2):663, 2021
2021
-
[132]
Vasilopoulos, G
V. Vasilopoulos, G. Pavlakos, S. L. Bowman, J. D. Caporale, K.Daniilidis,G.J.Pappas,andD.E.Koditschek. Reactivesemantic planning in unexplored semantic environments using deep percep- tual feedback. IEEE Robotics and Automation Letters, 5(3):4455– 4462, 2020
2020
-
[133]
R. C. Luo and M. G. Kay. Multisensor integration and fusion in intelligent systems. IEEE Transactions on Systems, Man, and Cybernetics, 19(5):901–931, 1989
1989
-
[134]
Asensorfordynamictactile informationwithapplicationsinhuman–robotinteractionandobject exploration
P.A.Schmidt,E.Maël,andR.P.Würtz. Asensorfordynamictactile informationwithapplicationsinhuman–robotinteractionandobject exploration. RoboticsandAutonomousSystems ,54(12):1005–1014, 2006
2006
-
[135]
L. A. Dennis, M. Fisher, N. K. Lincoln, A. Lisitsa, and S. M. Veres. Practicalverificationofdecision-makinginagent-basedautonomous systems. Automated Software Engineering, 23:305–359, 2016
2016
-
[136]
Integratinghumanandrobot decision-making dynamics with feedback: Models and convergence analysis
M.Cao,A.Stewart,andN.E.Leonard. Integratinghumanandrobot decision-making dynamics with feedback: Models and convergence analysis. In 2008 47th IEEE Conference on Decision and Control, pages 1127–1132. IEEE, December 2008
2008
-
[137]
Synthesisforrobots: Guarantees and feedback for robot behavior
H.Kress-Gazit,M.Lahijanian,andV.Raman. Synthesisforrobots: Guarantees and feedback for robot behavior. Annual Review of Control, Robotics, and Autonomous Systems, 1(1):211–236, 2018
2018
-
[138]
Merckaert, B
K. Merckaert, B. Convens, C. J. Wu, A. Roncone, M. M. Nicotra, and B. Vanderborght. Real-time motion control of robotic manip- ulators for safe human–robot coexistence.Robotics and Computer- Integrated Manufacturing, 73:102223, 2022
2022
-
[139]
Ferrucci and S
F. Ferrucci and S. Bock. Real-time control of express pickup and delivery processes in a dynamic environment.Transportation Research Part B: Methodological, 63:1–14, 2014
2014
-
[140]
Gheibi, D
O. Gheibi, D. Weyns, and F. Quin. Applying machine learning in self-adaptive systems: A systematic literature review.ACM Trans- actions on Autonomous and Adaptive Systems (TAAS), 15(3):1–37, 2021
2021
-
[141]
Huang, C
Z. Huang, C. Lv, Y. Xing, and J. Wu. Multi-modal sensor fusion- based deep neural network for end-to-end autonomous driving with scene understanding.IEEE Sensors Journal, 21(10):11781–11790, 2020. : Preprint submitted to Elsevier Page 19 of 19
2020
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.