Pith. sign in

REVIEW 5 cited by

Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15836 v2 pith:AQO6VIEB submitted 2024-06-22 cs.LG cs.AIcs.MA

classification cs.LGcs.AIcs.MA
keywords agentsmodelworldcentralizedlearningmulti-agentaggregationarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning a world model for model-free Reinforcement Learning (RL) agents can significantly improve the sample efficiency by learning policies in imagination. However, building a world model for Multi-Agent RL (MARL) can be particularly challenging due to the scalability issue in a centralized architecture arising from a large number of agents, and also the non-stationarity issue in a decentralized architecture stemming from the inter-dependency among agents. To address both challenges, we propose a novel world model for MARL that learns decentralized local dynamics for scalability, combined with a centralized representation aggregation from all agents. We cast the dynamics learning as an auto-regressive sequence modeling problem over discrete tokens by leveraging the expressive Transformer architecture, in order to model complex local dynamics across different agents and provide accurate and consistent long-term imaginations. As the first pioneering Transformer-based world model for multi-agent systems, we introduce a Perceiver Transformer as an effective solution to enable centralized representation aggregation within this context. Results on Starcraft Multi-Agent Challenge (SMAC) show that it outperforms strong model-free approaches and existing model-based methods in both sample efficiency and overall performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A new transformer-based multi-agent world model with teammate prediction and prioritized replay achieves near-optimal performance on cooperative benchmarks in as few as 50,000 environment steps.

  2. Hume: Introducing System-2 Thinking in Visual-Language-Action Model

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.

  3. Pre-Trained Video Generative Models as World Simulators

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A lightweight action-conditioning module and a motion-reinforced loss convert pre-trained video generators into action-following world simulators that also speed up model-based reinforcement learning.

  4. DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets

    cs.DC 2025-11 conditional novelty 4.0 of 10

    A decentralized world-model MARL framework with uncertainty-gated edge offloading and latent-state mean subtraction improves simulated SWIPT beamforming and power-splitting performance.

  5. Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G

    cs.IT 2025-02 conditional novelty 1.0 of 10

    A comprehensive survey of multi-agent reinforcement learning for wireless distributed networks in 6G, covering structures, algorithms, enhanced techniques, and applications.

Pith tools