Pith. sign in

REVIEW 9 cited by

EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04449 v5 pith:BFKHX32C submitted 2024-08-08 cs.AI

classification cs.AI
keywords riskmodelssafetyphysicalagentsfoundationtaskawareness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embodied artificial intelligence (EAI) integrates advanced AI models into physical entities for real-world interaction. The emergence of foundation models as the "brain" of EAI agents for high-level task planning has shown promising results. However, the deployment of these agents in physical environments presents significant safety challenges. For instance, a housekeeping robot lacking sufficient risk awareness might place a metal container in a microwave, potentially causing a fire. To address these critical safety concerns, comprehensive pre-deployment risk assessments are imperative. This study introduces EARBench, a novel framework for automated physical risk assessment in EAI scenarios. EAIRiskBench employs a multi-agent cooperative system that leverages various foundation models to generate safety guidelines, create risk-prone scenarios, make task planning, and evaluate safety systematically. Utilizing this framework, we construct EARDataset, comprising diverse test cases across various domains, encompassing both textual and visual scenarios. Our comprehensive evaluation of state-of-the-art foundation models reveals alarming results: all models exhibit high task risk rates (TRR), with an average of 95.75% across all evaluated models. To address these challenges, we further propose two prompting-based risk mitigation strategies. While these strategies demonstrate some efficacy in reducing TRR, the improvements are limited, still indicating substantial safety concerns. This study provides the first large-scale assessment of physical risk awareness in EAI agents. Our findings underscore the critical need for enhanced safety measures in EAI systems and provide valuable insights for future research directions in developing safer embodied artificial intelligence system. Data and code are available at https://github.com/zihao-ai/EARBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Content danger and physical danger form separable hidden-state signals in LLMs, and a single-layer logistic probe (PRISM) detects both at lower false-positive rates than LLM judges or text guardrails.

  2. Self-Evolving Just-In-Time Memory for Proactive Embodied Safety

    cs.LG 2026-06 conditional novelty 6.0 of 10

    A Just-In-Time Memory framework with graph-based state tracking and self-evolving safety skills improves safe task completion in household robots by up to 30 percentage points on IS-Bench.

  3. Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Most of 72 tested LLMs complied with harmful medical-robot orders over half the time, with open-weight models far worse than proprietary ones.

  4. SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A three-level temporal-logic safety evaluator for embodied LLM agents that checks NL-to-LTL interpretation, plan compliance, and CTL over simulated execution trees.

  5. RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    LLM-generated robot policy code is unreliable, with failures clustering into four behavior types that grow with task complexity and shrink with instruction detail; a failure-feedback retry improves success up to 35%.

  6. HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices

    cs.CL 2025-05 conditional novelty 6.0 of 10

    HomeBench is a new smart home benchmark that exposes near-zero success rates for top LLMs on invalid multi-device instructions.

  7. Context-Aware Risk Estimation in Home Environments: A Probabilistic Framework for Service Robots

    cs.RO 2025-08 reject novelty 5.0 of 10

    A semantic graph framework propagates risk scores derived from a national accident database across spatial object relations, reporting 75% binary risk detection accuracy on 20 human-annotated NYU V2 home images.

  8. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

  9. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools