Pith. sign in

REVIEW 1 cited by

OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction Following

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03017 v1 pith:WWWGXGGN submitted 2024-03-05 cs.AI

classification cs.AI
keywords embodiedperformancetasklearningactionagentsanalysisfollowing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embodied Instruction Following (EIF) is a crucial task in embodied learning, requiring agents to interact with their environment through egocentric observations to fulfill natural language instructions. Recent advancements have seen a surge in employing large language models (LLMs) within a framework-centric approach to enhance performance in embodied learning tasks, including EIF. Despite these efforts, there exists a lack of a unified understanding regarding the impact of various components-ranging from visual perception to action execution-on task performance. To address this gap, we introduce OPEx, a comprehensive framework that delineates the core components essential for solving embodied learning tasks: Observer, Planner, and Executor. Through extensive evaluations, we provide a deep analysis of how each component influences EIF task performance. Furthermore, we innovate within this space by deploying a multi-agent dialogue strategy on a TextWorld counterpart, further enhancing task performance. Our findings reveal that LLM-centric design markedly improves EIF outcomes, identify visual perception and low-level action execution as critical bottlenecks, and demonstrate that augmenting LLMs with a multi-agent framework further elevates performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    GraphThink uses a task graph for LLM planning prompts, GRPO rewards, and plan verification, plus a scene-graph event-driven replanner, achieving SOTA ALFRED results and stronger long-horizon generalization than API LLMs.

Pith tools