REVIEW 21 cited by
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications (e.g., intelligent mechatronics systems, smart manufacturing) that bridge cyberspace and the physical world. Recently, the emergence of Multi-modal Large Models (MLMs) and World Models (WMs) have attracted significant attention due to their remarkable perception, interaction, and reasoning capabilities, making them a promising architecture for embodied agents. In this survey, we give a comprehensive exploration of the latest advancements in Embodied AI. Our analysis firstly navigates through the forefront of representative works of embodied robots and simulators, to fully understand the research focuses and their limitations. Then, we analyze four main research targets: 1) embodied perception, 2) embodied interaction, 3) embodied agent, and 4) sim-to-real adaptation, covering state-of-the-art methods, essential paradigms, and comprehensive datasets. Additionally, we explore the complexities of MLMs in virtual and real embodied agents, highlighting their significance in facilitating interactions in digital and physical environments. Finally, we summarize the challenges and limitations of embodied AI and discuss potential future directions. We hope this survey will serve as a foundational reference for the research community. The associated project can be found at https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List.
Forward citations
Cited by 21 Pith papers
-
SocietyBench: Forecasting Counterfactual Social-World Evolution
A new benchmark measures LLM social-world forecasting on anonymized real events, finding the best model reaches 75/100 and agent scaffolding does not help.
-
Experimental Evidence That AI-Managed Workers Tolerate Lower Pay Without Demotivation
Workers paid 40% less by a rule-based AI manager reported no loss of fairness or motivation, while a harsher decision-tree AI did trigger backlash.
-
RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
RFTF trains a value model on temporal state orderings to supply dense rewards for reinforcement fine-tuning of vision-language-action models, achieving an average success length of 4.296 on CALVIN ABC-D with the Seer-...
-
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
MTU3D unifies visual grounding and frontier-based exploration in a single transformer, achieving state-of-the-art success rates on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA after large-scale vision-language-exploration p...
-
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
A 7B LLM trained with sparse completion rewards and a GRPO-style algorithm reaches state-of-the-art on ALFWorld and ScienceWorld.
-
In-the-wild Audio Spatialization with Flexible Text-guided Localization
A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.
-
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning
A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.
-
Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety
A unified framework of secure prompting, state memory, and rule-based safety validation improves LLM-driven robot navigation under prompt injection attacks and obstacle-heavy environments, with modest real-robot verification.
-
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.
-
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
A modular pipeline generates navigation instructions from egocentric video by extracting actions, scenes, and objects and synthesizing them with LLMs, plus a metric suite that evaluates the result without human refere...
-
Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
A survey that groups graph-empowered AI agent research into planning, execution, memory, and multi-agent coordination, plus agents-for-graphs and applications.
-
AI Flow: Perspectives, Scenarios, and Approaches
AI Flow proposes to combine device-edge-cloud deployment, feature-aligned model families, and multi-model collaboration to make large AI models cheaper, faster, and more widely accessible.
-
RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...
-
Situating AI Agents in their World: Aspective Agentic AI for Dynamic Partially Observable Information Systems
Agents that perceive only a filtered 'aspect' of a shared environment leaked no secrets in the authors' test, while an unfiltered AutoGen baseline leaked in most runs.
-
FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field
FGO-SLAM combines feature-based global pose adjustment with a 3D Gaussian opacity field to improve tracking accuracy, mapping quality, and direct mesh extraction in Gaussian SLAM.
-
Embodied AI: Emerging Risks and Opportunities for Policy Action
A policy analysis arguing that embodied AI risks are real, under-covered by current US/EU/UK frameworks, and best handled through certification, benchmarks, clarified liability, and economic adaptation.
-
Reconstructing 4D Spatial Intelligence: A Survey
A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.
-
A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
-
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.
-
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.
Discussion (0). Sign in to comment.