Pith. sign in

REVIEW 21 cited by

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06886 v8 pith:PXVWI57N submitted 2024-07-09 cs.CV cs.AIcs.LGcs.MAcs.RO

classification cs.CVcs.AIcs.LGcs.MAcs.RO
keywords embodiedcomprehensivephysicalresearchsurveyworldagentsartificial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications (e.g., intelligent mechatronics systems, smart manufacturing) that bridge cyberspace and the physical world. Recently, the emergence of Multi-modal Large Models (MLMs) and World Models (WMs) have attracted significant attention due to their remarkable perception, interaction, and reasoning capabilities, making them a promising architecture for embodied agents. In this survey, we give a comprehensive exploration of the latest advancements in Embodied AI. Our analysis firstly navigates through the forefront of representative works of embodied robots and simulators, to fully understand the research focuses and their limitations. Then, we analyze four main research targets: 1) embodied perception, 2) embodied interaction, 3) embodied agent, and 4) sim-to-real adaptation, covering state-of-the-art methods, essential paradigms, and comprehensive datasets. Additionally, we explore the complexities of MLMs in virtual and real embodied agents, highlighting their significance in facilitating interactions in digital and physical environments. Finally, we summarize the challenges and limitations of embodied AI and discuss potential future directions. We hope this survey will serve as a foundational reference for the research community. The associated project can be found at https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SocietyBench: Forecasting Counterfactual Social-World Evolution

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A new benchmark measures LLM social-world forecasting on anonymized real events, finding the best model reaches 75/100 and agent scaffolding does not help.

  2. Experimental Evidence That AI-Managed Workers Tolerate Lower Pay Without Demotivation

    cs.CY 2025-05 conditional novelty 7.0 of 10

    Workers paid 40% less by a rule-based AI manager reported no loss of fairness or motivation, while a harsher decision-tree AI did trigger backlash.

  3. RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback

    cs.RO 2025-05 conditional novelty 7.0 of 10

    RFTF trains a value model on temporal state orderings to supply dense rewards for reinforcement fine-tuning of vision-language-action models, achieving an average success length of 4.296 on CALVIN ABC-D with the Seer-...

  4. Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MTU3D unifies visual grounding and frontier-based exploration in a single transformer, achieving state-of-the-art success rates on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA after large-scale vision-language-exploration p...

  5. Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 7B LLM trained with sparse completion rewards and a GRPO-style algorithm reaches state-of-the-art on ALFWorld and ScienceWorld.

  6. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  7. RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

    cs.AI 2025-10 conditional novelty 5.0 of 10

    A 3B VLM trained with SFT plus GRPO and an LCS-based reward reaches 55.3% on EmbodiedBench's EB-ALFRED, beating GPT-4o-mini and the 7B REBP planner.

  8. Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A unified framework of secure prompting, state memory, and rule-based safety validation improves LLM-driven robot navigation under prompt injection attacks and obstacle-heavy environments, with modest real-robot verification.

  9. SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.

  10. NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization

    cs.AI 2025-07 reject novelty 5.0 of 10

    A modular pipeline generates navigation instructions from egocentric video by extracting actions, scenes, and objects and synthesizing them with LLMs, plus a metric suite that evaluates the result without human refere...

  11. Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A survey that groups graph-empowered AI agent research into planning, execution, memory, and multi-agent coordination, plus agents-for-graphs and applications.

  12. AI Flow: Perspectives, Scenarios, and Approaches

    cs.AI 2025-06 conditional novelty 5.0 of 10

    AI Flow proposes to combine device-edge-cloud deployment, feature-aligned model families, and multi-model collaboration to make large AI models cheaper, faster, and more widely accessible.

  13. RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

    cs.AI 2025-06 conditional novelty 5.0 of 10

    RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...

  14. Situating AI Agents in their World: Aspective Agentic AI for Dynamic Partially Observable Information Systems

    cs.AI 2025-09 reject novelty 4.0 of 10

    Agents that perceive only a filtered 'aspect' of a shared environment leaked no secrets in the authors' test, while an unfiltered AutoGen baseline leaked in most runs.

  15. FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field

    cs.RO 2025-09 conditional novelty 4.0 of 10

    FGO-SLAM combines feature-based global pose adjustment with a 3D Gaussian opacity field to improve tracking accuracy, mapping quality, and direct mesh extraction in Gaussian SLAM.

  16. Embodied AI: Emerging Risks and Opportunities for Policy Action

    cs.CY 2025-08 conditional novelty 4.0 of 10

    A policy analysis arguing that embodied AI risks are real, under-covered by current US/EU/UK frameworks, and best handled through certification, benchmarks, clarified liability, and economic adaptation.

  17. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

  18. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

  19. Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications

    cs.MA 2025-07 conditional novelty 4.0 of 10

    The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.

  20. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

  21. Toward Embodied AGI: A Review of Embodied AI and the Road Ahead

    cs.AI 2025-05 accept novelty 4.0 of 10

    Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.

Pith tools