Pith. sign in

REVIEW 7 cited by

EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09604 v1 pith:5NZLPENV submitted 2024-10-12 cs.AI cs.RO

classification cs.AIcs.RO
keywords embodiedintelligenceenvironmentcityplatformrealabilitiesagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embodied artificial intelligence emphasizes the role of an agent's body in generating human-like behaviors. The recent efforts on EmbodiedAI pay a lot of attention to building up machine learning models to possess perceiving, planning, and acting abilities, thereby enabling real-time interaction with the world. However, most works focus on bounded indoor environments, such as navigation in a room or manipulating a device, with limited exploration of embodying the agents in open-world scenarios. That is, embodied intelligence in the open and outdoor environment is less explored, for which one potential reason is the lack of high-quality simulators, benchmarks, and datasets. To address it, in this paper, we construct a benchmark platform for embodied intelligence evaluation in real-world city environments. Specifically, we first construct a highly realistic 3D simulation environment based on the real buildings, roads, and other elements in a real city. In this environment, we combine historically collected data and simulation algorithms to conduct simulations of pedestrian and vehicle flows with high fidelity. Further, we designed a set of evaluation tasks covering different EmbodiedAI abilities. Moreover, we provide a complete set of input and output interfaces for access, enabling embodied agents to easily take task requirements and current environmental observations as input and then make decisions and obtain performance evaluations. On the one hand, it expands the capability of existing embodied intelligence to higher levels. On the other hand, it has a higher practical value in the real world and can support more potential applications for artificial general intelligence. Based on this platform, we evaluate some popular large language models for embodied intelligence capabilities of different dimensions and difficulties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

    cs.RO 2026-07 conditional novelty 7.0 of 10

    ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.

  2. 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A new benchmark evaluates embodied agents in a photorealistic 360-video reconstruction of Akihabara, and state-of-the-art LMM agents score far below local human experts.

  3. MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

    cs.MA 2026-07 conditional novelty 6.0 of 10

    Across 17 multimodal models on 3,024 protocol-conditioned UAV decision samples, the best semantic protocol-decision score is only 0.5141 and strict mean dimension accuracy only 0.1599.

  4. ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A robotic agent operating system with source-grounded graph memory and split-wise self-evolution improves long-horizon embodied task success and memory QA scores over baseline controllers.

  5. SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World

    cs.AI 2024-12 reject novelty 5.0 of 10

    SmartAgent is a GUI agent that adds user-preference reasoning through three thought steps, but its intermediate 'underlying requirement' step does not improve item recommendation over end-to-end training.

  6. InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction

    cs.RO 2024-12 conditional novelty 5.0 of 10

    InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.

  7. Fast SSVEP Detection Using a Calibration-Free EEG Decoding Framework

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A compact calibration-free EEG decoder with trial-remixing augmentation and adaptive spectral denoising beats CCA, FBCCA, TRCA, TFF, and EEGConformer on short SSVEP signals across three public datasets.

Pith tools