Pith. sign in

REVIEW 3 cited by

Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19391 v1 pith:5ZKOA2SO submitted 2024-09-28 cs.LG

Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training

classification cs.LG
keywords learningmarltrainingsparsedeepmulti-agentmastvalue-based
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep Multi-agent Reinforcement Learning (MARL) relies on neural networks with numerous parameters in multi-agent scenarios, often incurring substantial computational overhead. Consequently, there is an urgent need to expedite training and enable model compression in MARL. This paper proposes the utilization of dynamic sparse training (DST), a technique proven effective in deep supervised learning tasks, to alleviate the computational burdens in MARL training. However, a direct adoption of DST fails to yield satisfactory MARL agents, leading to breakdowns in value learning within deep sparse value-based MARL models. Motivated by this challenge, we introduce an innovative Multi-Agent Sparse Training (MAST) framework aimed at simultaneously enhancing the reliability of learning targets and the rationality of sample distribution to improve value learning in sparse models. Specifically, MAST incorporates the Soft Mellowmax Operator with a hybrid TD-($\lambda$) schema to establish dependable learning targets. Additionally, it employs a dual replay buffer mechanism to enhance the distribution of training samples. Building upon these aspects, MAST utilizes gradient-based topology evolution to exclusively train multiple MARL agents using sparse networks. Our comprehensive experimental investigation across various value-based MARL algorithms on multiple benchmarks demonstrates, for the first time, significant reductions in redundancy of up to $20\times$ in Floating Point Operations (FLOPs) for both training and inference, with less than $3\%$ performance degradation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

    cs.AI 2026-04 unverdicted novelty 6.0

    KD-MARL distills both actions and coordination structure from expert MARL policies into heterogeneous lightweight students, retaining over 90% performance while cutting FLOPs by up to 28.6 times on SMAC and MPE benchmarks.

  2. When Agent Markets Arrive

    cs.CE 2026-04 unverdicted novelty 6.0

    DIAGON simulation shows agent markets produce 3.2 times more wealth than isolated agents, but institutional choices like transparency and competitive selection can reduce rather than increase performance.

  3. When Agent Markets Arrive

    cs.CE 2026-04 unverdicted novelty 5.0

    Market exchange among AI agents can raise productivity over self-sufficient agents, but institutional rules such as identity transparency and stronger selection can degrade those gains.