Pith. sign in

REVIEW 15 cited by

Building Cooperative Embodied Agents Modularly with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.02485 v2 pith:VILIFK4K submitted 2023-07-05 cs.AI cs.CLcs.CVcs.RO

classification cs.AIcs.CLcs.CVcs.RO
keywords coelalanguagecommunicationembodiedresearchagentsbuildingcooperate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we address challenging multi-agent cooperation problems with decentralized control, raw sensory observations, costly communication, and multi-objective tasks instantiated in various embodied environments. While previous research either presupposes a cost-free communication channel or relies on a centralized controller with shared observations, we harness the commonsense knowledge, reasoning ability, language comprehension, and text generation prowess of LLMs and seamlessly incorporate them into a cognitive-inspired modular framework that integrates with perception, memory, and execution. Thus building a Cooperative Embodied Language Agent CoELA, who can plan, communicate, and cooperate with others to accomplish long-horizon tasks efficiently. Our experiments on C-WAH and TDW-MAT demonstrate that CoELA driven by GPT-4 can surpass strong planning-based methods and exhibit emergent effective communication. Though current Open LMs like LLAMA-2 still underperform, we fine-tune a CoELA with data collected with our agents and show how they can achieve promising performance. We also conducted a user study for human-agent interaction and discovered that CoELA communicating in natural language can earn more trust and cooperate more effectively with humans. Our research underscores the potential of LLMs for future research in multi-agent cooperation. Videos can be found on the project website https://vis-www.cs.umass.edu/Co-LLM-Agents/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 37 citations worldwide. Full citation record

  1. CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

    cs.CV 2026-07 accept novelty 6.5 of 10

    CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.

  2. A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A closed-loop multi-agent LLM framework enables heterogeneous robots to collaboratively manipulate objects by decomposing tasks, grounding actions via visual tools, and recovering from execution failures hierarchically.

  3. LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning

    cs.RO 2025-11 conditional novelty 6.0 of 10

    An LLM-based hierarchical system lets heterogeneous robot teams plan, execute, and autonomously replan in response to unexpected events, demonstrated on physical robots and in simulation.

  4. GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A four-agent LLM pipeline retrieves known Solidity gas-waste patterns, proposes new ones, and automatically verifies and applies the fixes, saving about 10% deployment gas on 82% of real contracts.

  5. DPMT: Dual Process Multi-scale Theory of Mind Framework for Real-time Human-AI Collaboration

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A dual-process LLM agent with a three-stage theory-of-mind module outperforms baselines in real-time Overcooked human-AI collaboration.

  6. LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.

  7. Your Agent Can Defend Itself against Backdoor Attacks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.

  8. Optimizing Temperature for Language Models with Multi-Sample Inference

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Selecting the temperature at the entropy turning point of a language model's generated text yields near-optimal multi-sample inference accuracy without labeled validation data.

  9. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  10. Multi-Actor Generative Artificial Intelligence as a Game Engine

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Generative multi-actor AI platforms can be built on the Entity-Component pattern, treating the environment (Game Master) as a composable entity, so that one library serves simulation, storytelling, and evaluation goals.

  11. CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CoNav lets a frozen 3D-text model pass spatial text hints to a lightly fine-tuned image-text navigation agent, improving path efficiency on several VLN benchmarks, though not all claimed state-of-the-art results hold.

  12. Communicative Agents for Slideshow Storytelling Video Generation based on LLMs

    cs.AI 2025-09 conditional novelty 4.0 of 10

    VGTeam uses communicating LLM agents plus commercial APIs to turn a text prompt into a slideshow video for about $0.10 per clip, with a self-reported 75.7% quality rate.

  13. Reinforced Language Models for Sequential Decision Making

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A 3B LLM post-trained with MS-GRPO, which gives every step the episode's total reward and samples high-advantage episodes, beats a 72B baseline on Frozen Lake but is inconsistent on Snake.

  14. Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A position paper arguing that Bayesian inference could become a key design principle for embodied AI in open physical worlds, using Sutton's search-and-learning lens to explain its current absence.

  15. HiBerNAC: Hierarchical Brain-emulated Robotic Neural Agent Collective for Disentangling Complex Manipulation

    cs.RO 2025-06 reject novelty 4.0 of 10

    HiBerNAC, a multi-agent 'brain-inspired' planner layered on a reactive VLA, is claimed to cut long-horizon task time by 23% and reach 12-31% success where VLA baselines fail, but the supporting data are inconsistent.

Pith tools