Pith. sign in

REVIEW 2 major objections 5 minor 3 cited by

Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G

T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This paper argues that multi-agent reinforcement learning is an integral part of wireless distributed networks for 6G, and it provides a systematic taxonomy and enhancement roadmap to realize that integration.

desk verdict Useful, well-organized MARL-for-6G tutorial, but the inverted ratio in the mutual information definition (Eqs. 10–11) is a load-bearing correctness defect that must be fixed before the IB section can be trusted. read the letter →

arxiv 2502.05812 v1 pith:ZDN7J667 submitted 2025-02-09 cs.IT cs.SYeess.SYmath.IT

classification cs.ITcs.SYeess.SYmath.IT
keywords multi-agentreinforcementlearningwirelessdistributednetworks6Gmodel-basedMARLmodel-freeinformationbottleneckmirrorcell-freemassiveMIMO
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This tutorial sets out to establish that multi-agent reinforcement learning (MARL) is not just an optional add-on but an integral part of wireless distributed networks for 6G. It argues that the two fields share a decentralized, collaborative core, and that aligning them provides a systematic route to meet 6G's demands for massive capacity, ultra-low latency, and extreme reliability. The paper maps the space: homogeneous versus heterogeneous distributed networks, model-based versus model-free MARL, and four families of enhancement techniques (graph, collaboration, information bottleneck, and mirror learning). It then anchors the discussion in concrete application scenarios and reports quantitative gains from recent studies to show what the integration buys.

What carries the argument

The argument runs on the mapping between network structure and learning formalism. Agents and their interactions are represented as nodes and edges of a graph, which turns homogeneous and heterogeneous wireless distributed networks into instances of Markov decision processes (sequential decision problems with full observability), decentralized partially observable MDPs (agents see only local noisy observations), or networked MDPs (agents use local and neighbor states). On top of that base sit two algorithmic families: model-based MARL, which plans with a known or learned transition model, and model-free MARL, which learns policies directly from interaction. The paper adds four enhancement mechanisms: graph-enhanced MARL for relationship modeling, embedded learning, and hierarchical decoupling; collaboration-enhanced MARL organized around whom, when, what, and how to share information; information-bottleneck-enhanced MARL built on the minimal-sufficient-information objective of IB, GIB, and SubIB; and mirror learning-enhanced MARL with maximum-entropy and Transformer-based world models. These mechanisms are what carry the claimed improvements in robustness, communication efficiency, and sample efficiency.

What would settle it

Re-run the collaborative information-sharing protocol described in Section V-B on the same homogeneous and heterogeneous network models and check whether the reported sum spectral-efficiency improvements of 19.39% and 20.46% over the centralized-training baseline reproduce; a substantial shortfall would remove the paper's quantitative justification for collaboration-enhanced MARL.

Watch

Extended reading notes

Core claim

The central discovery, on the survey's own terms, is that MARL-assisted wireless distributed networks form a coherent design space rather than a collection of isolated tricks. The paper shows that wireless networks evolved from centralized to distributed structures, while reinforcement learning evolved from single-agent to multi-agent, and that both trajectories meet in the same decentralized decision-making paradigm. Given that meeting point, the authors claim, every 6G distributed network can be classified as homogeneous or heterogeneous and then paired with the appropriate MARL formulation—MDP, Dec-POMDP, or networked MDP—and the appropriate algorithmic family, model-based or model-free. The tutorial further claims that four enhancement techniques—graph modeling, collaborative communication protocols, information bottleneck compression, and mirror learning—resolve the main remaining obstacles such as partial observability, poor communication efficiency, and insufficient exploration.

Load-bearing premise

The tutorial's practical guidance assumes the reported quantitative results from the authors' own simulations—such as the 19.39% and 20.46% spectral-efficiency gains, the information-bottleneck error tolerances, and the mirror-learning advantage—are correct and representative of typical 6G settings.

Editorial extensions

If this is right

  • A 6G designer can treat network architecture choice (homogeneous versus heterogeneous) as the first step in selecting a MARL formalism and algorithm family, giving a systematic rather than ad hoc design process.
  • Collaboration protocols that specify whom, when, what, and how to communicate are claimed to lift spectral efficiency by 19.39% in homogeneous and 20.46% in heterogeneous settings over standard centralized-training baselines.
  • Applying information-bottleneck compression to MARL representations is claimed to approach centralized upper-bound performance under interference, with reported error tolerances near 72.23% for feature information and 87.29% for structural information.
  • Mirror-learning methods such as HATRPO provide a theoretical foundation for policy updates that generalize beyond generalized policy iteration and trust-region learning, and Transformer-based world models extend this to long-horizon multi-agent prediction.
  • Future directions point to large AI models, continuous learning, and emerging communications (semantic, ISAC, URLLC, and ultra-high dynamic) as the next expansion of MARL-assisted distributed networks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy holds, a reusable recipe emerges: choose the MARL formalism from the network graph, then add collaboration protocols for communication and IB or SubIB for representation robustness; this recipe could be applied to new distributed problems beyond the four applications surveyed.
  • The quantitative anchors in the survey all come from the authors' own simulation studies; a natural next step is third-party reproduction under different channel models and topologies to test how far the reported gains generalize.
  • The Transformer-based world-model direction suggests that future wireless MARL agents may exchange compact learned tokens instead of raw observations, which would shift communication efficiency from protocol design to representation design.
  • The paper's emphasis on minimal sufficient information connects naturally to rate-distortion theory, hinting that MARL communication in 6G could eventually be analyzed with the same tools used for source coding.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a tutorial-style survey of multi-agent reinforcement learning (MARL) for wireless distributed networks in 6G. It positions MARL and wireless distributed networks as two decentralized paradigms with a natural synergy; introduces taxonomies of network structures (homogeneous vs. heterogeneous) and MARL algorithms (model-based vs. model-free); surveys enhanced-MARL techniques (graph-enhanced, collaboration-enhanced, information-bottleneck-enhanced, mirror-learning-enhanced); and presents application scenarios (UAV-assisted communications, autonomous driving, cell-free massive MIMO, RIS-assisted MIMO) together with future directions. The contribution is primarily organizational: systematic tables, roadmaps, and modeling tutorials rather than new theoretical results.

Significance. If the mathematical content were correct, this would be a useful reference for researchers entering the area, thanks to its broad coverage, clear taxonomy tables, and concrete application tutorials. The IB-enhanced MARL subsection in particular attempts to provide a step-by-step modeling guide, which is a genuinely valuable tutorial feature. However, the tutorial's reliability is currently undermined by a sign error in the definition of mutual information and by related sign errors in the proposed IB-based objectives; because a tutorial is judged by whether readers can teach or implement from its equations, these errors are load-bearing and must be corrected before the paper can serve as the intended guideline.

major comments (2)
  1. [Section V-C, Eqs. (10), (11a), (11b)] The mutual information is defined as I(X;Y) = E_{p(x,y)}[log p(x)p(y)/p(x,y)], which equals the negative of the standard mutual information. The correct definition is I(X;Y) = E_{p(x,y)}[log p(x,y)/(p(x)p(y))], equivalently the KL divergence between p(x,y) and p(x)p(y). The same inverted ratio appears in the discrete and continuous forms in Eqs. (11a) and (11b). Because the IB-enhanced formulations in Eqs. (13)-(20) and the discussion around Fig. 13 all build on this quantity, the error inverts the sign of the compression and preservation terms: read literally, minimizing Eq. (13) would not implement the IB principle described in the text. The definition must be corrected and all downstream sign conventions re-checked.
  2. [Section V-C, Eqs. (17)-(20)] In the four 'Challenging Problem' updates, the objective is written as min_z L_MARL(G;Z) = -L_MARL(Z) + βI(G;Z) (and similarly in Eqs. (18)-(20)), where L_MARL(Z) is explicitly called the MARL loss function. Minimizing a negative loss is equivalent to maximizing the loss, so, as written, these objectives would degrade the MARL loss rather than optimize it. Either the sign should be changed to +L_MARL(Z) if L_MARL is indeed a loss to be minimized, or the text should state that L_MARL denotes a return to be maximized. This sign error affects all four tutorial templates and makes the enhanced-MARL guidance unreliable.
minor comments (5)
  1. [Section V, introductory paragraph] The text says 'graph-enhanced MARL in Section IV-A, collaboration-enhanced MARL in Section IV-B, information bottleneck-enhanced MARL in Section IV-C, and mirror learning-enhanced MARL in Section IV-D', but these subsections are actually in Section V (V-A through V-D); the cross-references should be corrected.
  2. [Section IV-C.2] The two application categories are both labeled 'a)', i.e., 'a) Direct Methods' and 'a) Communication Methods'; the second should be labeled 'b) Communication Methods'.
  3. [Section V-C, Challenging Problems 1 and 2] The text says 'update equations (4), (5), and (6)' and then 'equation (7) can be updated to', but those equation numbers refer to the value function, policy objective, GPI step, and Bellman update in Section IV, not to the IB formulations being revised; the references should point to the actual target equations (e.g., Eqs. (10) and (13)).
  4. [Section VII-A] The sentence 'This scalability unlocks the ability to process vast amounts of data in real-time, empowering empowering to make more informed decisions' contains a duplicated word 'empowering'.
  5. [Section V-D] The paragraph beginning 'To achieve this, a VQ-VAE is utilized...' appears immediately after the 'Lessons Learned' paragraph without a subsection heading or transitional text; it belongs to the Transformer-enhanced part and should be integrated into Subsection V-D.2 before the lessons learned.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation chain; self-citations are illustrative, while a sign error in Eq. (10) is a correctness issue, not a circularity.

full rationale

The paper is a tutorial/survey: its 'derivations' are definitions, taxonomies, and summaries of prior work, not new predictions built from fitted parameters. The quantitative results cited in Figs. 10-11, 13, and 15 come from prior papers (including some authored by the current group, e.g., [45], [50], [153], [214]), but they serve as illustrative examples of MARL enhancements rather than as inputs to a derivation that forces the tutorial's organizational conclusions. No parameter is fitted in this paper and then renamed a prediction; no uniqueness theorem is imported from the authors' prior work to forbid alternatives; no ansatz is smuggled in by citation. The only notable defect found in the mathematical foundation is in Section V-C, Eq. (10)-(11), where the mutual information ratio is inverted (log p(x)p(y)/p(x,y) instead of log p(x,y)/(p(x)p(y))), and Eq. (13) consequently has the wrong sign if read with the standard definition. This is a correctness error in the IB-enhanced MARL tutorial, not a circularity: the IB principle itself is adopted from the cited literature, not re-derived from the inverted definition. Because the self-citations are neither load-bearing nor used to define the paper's contributions, the circularity burden is low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters, new axioms beyond standard domain knowledge, or invented entities. All quantitative claims are inherited from the cited literature.

assumptions (3)
  • standard math Standard MDP, Dec-POMDP, and networked MDP definitions as used in Section IV are accepted as correct background.
    The tutorial's algorithm taxonomy is built on these formulations.
  • domain assumption The information bottleneck, graph information bottleneck, and subgraph information bottleneck objectives in Section V-C are faithful summaries of the cited works.
    The paper uses these IB variants as the basis for recommending enhanced MARL, so misrepresentation would invalidate the guidance.
  • domain assumption The homogeneous/heterogeneous classification of wireless distributed networks in Section III is a valid and useful dichotomy.
    The tutorial's structural organization depends on this taxonomy being non-arbitrary for 6G network design.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G." pith.science (2026). https://pith.science/paper/ZDN7J667

@misc{pith2026250205812,
  author       = {Pith},
  title        = {Pith review of: Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZDN7J667}},
  note         = {Machine review of arXiv:2502.05812}
}
read the original abstract

The introduction of intelligent interconnectivity between the physical and human worlds has attracted great attention for future sixth-generation (6G) networks, emphasizing massive capacity, ultra-low latency, and unparalleled reliability. Wireless distributed networks and multi-agent reinforcement learning (MARL), both of which have evolved from centralized paradigms, are two promising solutions for the great attention. Given their distinct capabilities, such as decentralization and collaborative mechanisms, integrating these two paradigms holds great promise for unleashing the full power of 6G, attracting significant research and development attention. This paper provides a comprehensive study on MARL-assisted wireless distributed networks for 6G. In particular, we introduce the basic mathematical background and evolution of wireless distributed networks and MARL, as well as demonstrate their interrelationships. Subsequently, we analyze different structures of wireless distributed networks from the perspectives of homogeneous and heterogeneous. Furthermore, we introduce the basic concepts of MARL and discuss two typical categories, including model-based and model-free. We then present critical challenges faced by MARL-assisted wireless distributed networks, providing important guidance and insights for actual implementation. We also explore an interplay between MARL-assisted wireless distributed networks and emerging techniques, such as information bottleneck and mirror learning, delivering in-depth analyses and application scenarios. Finally, we outline several compelling research directions for future MARL-assisted wireless distributed networks.

Figures

Figures reproduced from arXiv: 2502.05812 by the authors.

Figure 1
Figure 1. Meanwhile, Table I presents a detailed comparison [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 1
Figure 1. The development roadmap of wireless distributed networks and MARL from 2002 to Dec 2024. From the perspective of wireless distributed network [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The outline of this tutorial, where we introduce the provisioning of different MARL algorithms on wireless distributed networks for 6G, and highlight [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (17 more)
Figure 3
Figure 3. Figure 3: The connection between wireless distributed networks and MARL [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: The network architecture and connection of different wireless distributed networks, including homogeneous types (e.g., sharing the same communication [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Two typical heterogeneous distributed networks for 6G, including CF [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: The network architecture and connection of different MARL algorithms, including model-free types (e.g., any prior knowledge about wireless distributed [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Two typical network frameworks of MARL, including model-free and model-based MARL. Note that “V-Net” and “Q-Net” are defined as “Value [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Four types of application scenarios for model-based and model-free MARL. Please refer to [53], [60], [119], and [128] for more details. [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Four emerging techniques for enhancing MARL, including graph-enhanced, collaboration-enhanced, information bottleneck-enhanced, and mirror [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: Test reward (sum SE) curves of homogeneous distributed networks [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Test reward (sum SE) curves of heterogeneous distributed networks [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]
Figure 12
Figure 12. Figure 12: Three frameworks of information bottleneck-enhanced MARL, including basic IB, GIB, and SubIB, all reflect an innovative representation principle [PITH_FULL_IMAGE:figures/full_fig_p024_12.png]
Figure 13
Figure 13. Figure 13: Mutual information curves of MARL-assisted wireless distributed networks under the typical GIB theory, including [PITH_FULL_IMAGE:figures/full_fig_p025_13.png]
Figure 14
Figure 14. Figure 14: Two frameworks of mirror learning-enhanced MARL, including entropy and Transformer. Please refer to [151] and [201] for more details. [PITH_FULL_IMAGE:figures/full_fig_p027_14.png]
Figure 15
Figure 15. Figure 15: Test reward (average SE) curves of heterogeneous distributed [PITH_FULL_IMAGE:figures/full_fig_p027_15.png]
Figure 16
Figure 16. Figure 16: Key challenges faced and feasible techniques in autonomous driving for 6G. Challenge 1: Complex transportation areas caused by uncertain dynamics [PITH_FULL_IMAGE:figures/full_fig_p029_16.png]
Figure 17
Figure 17. Figure 17: Description of time-space diagram of a 3-lane arterial illustrating the effect of bus stops for 6G, including scenarios input, ground truth, and model [PITH_FULL_IMAGE:figures/full_fig_p030_17.png]
Figure 18
Figure 18. Figure 18: Key challenges faced and feasible techniques in CF massive MIMO networks for 6G. Challenge 1: Uneven service areas caused by severe interference [PITH_FULL_IMAGE:figures/full_fig_p031_18.png]
Figure 19
Figure 19. Figure 19: The evolution of future directions for MARL-assisted wireless distributed networks, including from traditional AI models to large AI models towards [PITH_FULL_IMAGE:figures/full_fig_p033_19.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cost and Accuracy of Long-Term Memory in Distributed Multi-Agent Systems Based on Large Language Models

    cs.IR 2026-01 reject novelty 5.0 of 10

    A two-framework testbed comparison claims mem0 is Pareto-optimal over Graphiti for distributed LLM agents because its lower cost is paired with accuracy that is not significantly different.

  2. Joint Power Control and Precoding for Cell-Free Massive MIMO Systems With Sparse Multi-Dimensional Graph Neural Networks

    cs.IT 2025-07 conditional novelty 4.0 of 10

    A sparse multi-dimensional graph neural network reduces computational complexity of joint power control and precoding in cell-free massive MIMO by about 48% with only a 2.11% spectral efficiency loss.

  3. AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need

    cs.CL 2025-06 reject novelty 4.0 of 10

    A divide-and-conquer multi-agent framework with task forests and specialized roles improves math and code benchmarks but not commonsense or domain QA, and the adaptive heterogeneous-LLM engine is never tested.

Reference graph

Works this paper leans on

262 extracted references · 60 canonical work pages · cited by 3 Pith papers

  1. [45]

    Graph neural network meets multi-agent reinforcement learning: Fundamen- tals, applications, and future directions,

    Z. Liu, J. Zhang, E. Shi, Z. Liu, D. Niyato, B. Ai, and X. Shen, “Graph neural network meets multi-agent reinforcement learning: Fundamen- tals, applications, and future directions,” IEEE Wireless Commun. , vol. 31, no. 6, pp. 39–47, Dec. 2024

  2. [153]

    Robust multidimen- sional graph neural networks for signal processing in wireless com- munications with edge-graph information bottleneck,

    Z. Liu, J. Zhang, Y . Zhu, E. Shi, and B. Ai, “Robust multidimen- sional graph neural networks for signal processing in wireless com- munications with edge-graph information bottleneck,” arXiv preprint arXiv:2502.15836, 2025

  3. [214]

    Joint precoding and phase shift design for RIS-aided cell-free massive MIMO with heterogeneous-agent trust region policy,

    Y . Zhu, J. Zhang, E. Shi, Z. Liu, C. Yuen, D. Niyato, and B. Ai, “Joint precoding and phase shift design for RIS-aided cell-free massive MIMO with heterogeneous-agent trust region policy,” IEEE Trans. Veh. Technol., vol. 74, no. 1, pp. 1794–1799, Jan. 2025

  4. [1]

    5G wireless network slicing for eMBB, URLLC, and mMTC: A communication-theoretic view,

    P. Popovski, K. F. Trillingsgaard, O. Simeone, and G. Durisi, “5G wireless network slicing for eMBB, URLLC, and mMTC: A communication-theoretic view,” IEEE Access , vol. 6, pp. 55 765– 55 779, Sep. 2018

  5. [2]

    Next generation 5G wireless networks: A comprehensive survey,

    M. Agiwal, A. Roy, and N. Saxena, “Next generation 5G wireless networks: A comprehensive survey,” IEEE Commun. Surveys Tuts. , vol. 18, no. 3, pp. 1617–1655, Thirdquarter 2016

  6. [3]

    5G: A tutorial overview of standards, trials, challenges, deployment, and practice,

    M. Shafi, A. F. Molisch, P. J. Smith, T. Haustein, P. Zhu, P. De Silva, F. Tufvesson, A. Benjebbour, and G. Wunder, “5G: A tutorial overview of standards, trials, challenges, deployment, and practice,” IEEE J. Sel. Areas Commun., vol. 35, no. 6, pp. 1201–1221, Jun. 2017

  7. [4]

    A tutorial on extremely large-scale MIMO for 6G: Fundamentals, signal processing, and applications,

    Z. Wang, J. Zhang, H. Du, D. Niyato, S. Cui, B. Ai, M. Debbah, K. B. Letaief, and H. V . Poor, “A tutorial on extremely large-scale MIMO for 6G: Fundamentals, signal processing, and applications,” IEEE Commun. Surveys Tuts. , vol. 26, no. 3, pp. 1560–1605, Thirdquarter 2024

  8. [5]

    A tutorial on near-field XL-MIMO communications toward 6G,

    H. Lu, Y . Zeng, C. You, Y . Han, J. Zhang, Z. Wang, Z. Dong, S. Jin, C.-X. Wang, T. Jiang, X. You, and R. Zhang, “A tutorial on near-field XL-MIMO communications toward 6G,” IEEE Commun. Surveys Tuts., vol. 26, no. 4, pp. 2213–2257, Fourthquarter 2024

Show all 262 references
  1. [6]

    Toward 6G networks: Use cases and technologies,

    M. Giordani, M. Polese, M. Mezzavilla, S. Rangan, and M. Zorzi, “Toward 6G networks: Use cases and technologies,” IEEE Commun. Mag., vol. 58, no. 3, pp. 55–61, Mar. 2020

  2. [7]

    A vision of 6G wireless systems: Applications, trends, technologies, and open research problems,

    W. Saad, M. Bennis, and M. Chen, “A vision of 6G wireless systems: Applications, trends, technologies, and open research problems,” IEEE Netw., vol. 34, no. 3, pp. 134–142, May 2020

  3. [8]

    A speculative study on 6G,

    F. Tariq, M. R. A. Khandaker, K.-K. Wong, M. A. Imran, M. Bennis, and M. Debbah, “A speculative study on 6G,”IEEE Wireless Commun., vol. 27, no. 4, pp. 118–125, Aug. 2020

  4. [9]

    Towards 6G wireless communica- tion networks: Vision, enabling technologies, and new paradigm shifts,

    X. You, C.-X. Wang, J. Huang, X. Gao, Z. Zhang, M. Wang, Y . Huang, C. Zhang, Y . Jiang, J. Wang et al., “Towards 6G wireless communica- tion networks: Vision, enabling technologies, and new paradigm shifts,” Sci. China Inf. Sci , vol. 64, pp. 1–74, Jan. 2021

  5. [10]

    6G wireless systems: Vision, requirements, challenges, insights, and opportunities,

    H. Tataria, M. Shafi, A. F. Molisch, M. Dohler, H. Sj ¨oland, and F. Tufvesson, “6G wireless systems: Vision, requirements, challenges, insights, and opportunities,” Proc. IEEE , vol. 109, no. 7, pp. 1166– 1199, Jul. 2021

  6. [11]

    Resource allocation for near-field communications: Fundamentals, tools, and outlooks,

    B. Xu, J. Zhang, H. Du, Z. Wang, Y . Liu, D. Niyato, B. Ai, and K. B. Letaief, “Resource allocation for near-field communications: Fundamentals, tools, and outlooks,” IEEE Wireless Commun., vol. 31, no. 5, pp. 42–50, Jul. 2024

  7. [12]

    Cell-free massive MIMO versus small cells,

    H. Q. Ngo, A. Ashikhmin, H. Yang, E. G. Larsson, and T. L. Marzetta, “Cell-free massive MIMO versus small cells,” IEEE Trans. Wireless Commun., vol. 16, no. 3, pp. 1834–1850, Mar. 2017

  8. [13]

    Cell- free massive MIMO: A new next-generation paradigm,

    J. Zhang, S. Chen, Y . Lin, J. Zheng, B. Ai, and L. Hanzo, “Cell- free massive MIMO: A new next-generation paradigm,” IEEE Access, vol. 7, pp. 99 878–99 888, Jul. 2019

  9. [14]

    Ultradense cell-free massive MIMO for 6G: Technical overview and open questions,

    H. Q. Ngo, G. Interdonato, E. G. Larsson, G. Caire, and J. G. Andrews, “Ultradense cell-free massive MIMO for 6G: Technical overview and open questions,” Proc. IEEE, vol. 112, no. 7, pp. 805–831, Jul. 2024

  10. [15]

    Next- generation multiple access with cell-free massive MIMO,

    M. Mohammadi, Z. Mobini, H. Quoc Ngo, and M. Matthaiou, “Next- generation multiple access with cell-free massive MIMO,” Proc. IEEE, vol. 112, no. 9, pp. 1372–1420, Sep. 2024

  11. [16]

    Foundations of user- centric cell-free massive MIMO,

    ¨O. T. Demir, E. Bj ¨ornson, L. Sanguinetti et al., “Foundations of user- centric cell-free massive MIMO,” Found. Trends® in Signal Process. , vol. 14, no. 3-4, pp. 162–472, 2021

  12. [17]

    User-centric cell-free massive MIMO networks: A survey of opportunities, challenges and solutions,

    H. A. Ammar, R. Adve, S. Shahbazpanahi, G. Boudreau, and K. V . Srinivas, “User-centric cell-free massive MIMO networks: A survey of opportunities, challenges and solutions,” IEEE Commun. Surveys Tuts., vol. 24, no. 1, pp. 611–652, Firstquarter 2022

  13. [18]

    Intelligent reflecting surface enhanced wireless network via joint active and passive beamforming,

    Q. Wu and R. Zhang, “Intelligent reflecting surface enhanced wireless network via joint active and passive beamforming,” IEEE Trans. Wireless Commun., vol. 18, no. 11, pp. 5394–5409, Nov. 2019

  14. [19]

    Reconfigurable intelligent surfaces: Principles and opportu- nities,

    Y . Liu, X. Liu, X. Mu, T. Hou, J. Xu, M. Di Renzo, and N. Al- Dhahir, “Reconfigurable intelligent surfaces: Principles and opportu- nities,” IEEE Commun. Surveys Tuts. , vol. 23, no. 3, pp. 1546–1577, Thirdquarter 2021

  15. [20]

    Reconfigurable intelligent surfaces for 6G systems: Prin- ciples, applications, and research directions,

    C. Pan, H. Ren, K. Wang, J. F. Kolb, M. Elkashlan, M. Chen, M. Di Renzo, Y . Hao, J. Wang, A. L. Swindlehurst, X. You, and L. Hanzo, “Reconfigurable intelligent surfaces for 6G systems: Prin- ciples, applications, and research directions,” IEEE Commun. Mag. , vol. 59, no. 6, p...

  16. [21]

    Reconfigurable intelligent surfaces for energy efficiency in wireless communication,

    C. Huang, A. Zappone, G. C. Alexandropoulos, M. Debbah, and C. Yuen, “Reconfigurable intelligent surfaces for energy efficiency in wireless communication,” IEEE Trans. Wireless Commun, , vol. 18, no. 8, pp. 4157–4170, Aug. 2019

  17. [22]

    RIS-aided cell-free massive MIMO systems for 6G: Fundamentals, system design, and applications,

    E. Shi, J. Zhang, H. Du, B. Ai, C. Yuen, D. Niyato, K. B. Letaief, and X. Shen, “RIS-aided cell-free massive MIMO systems for 6G: Fundamentals, system design, and applications,” Proc. IEEE, vol. 112, no. 4, pp. 331–364, Apr. 2024

  18. [23]

    Dynamic computation offloading for mobile-edge computing with energy harvesting devices,

    Y . Mao, J. Zhang, and K. B. Letaief, “Dynamic computation offloading for mobile-edge computing with energy harvesting devices,” IEEE J. Sel. Areas Commun. , vol. 34, no. 12, pp. 3590–3605, Dec. 2016

  19. [24]

    A survey on mobile edge computing: The communication perspective,

    Y . Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief, “A survey on mobile edge computing: The communication perspective,” IEEE Commun. Surveys Tuts. , vol. 19, no. 4, pp. 2322–2358, Fourthquarter 2017

  20. [25]

    Present and future of terahertz commu- nications,

    H.-J. Song and T. Nagatsuma, “Present and future of terahertz commu- nications,” IEEE Trans. THz Sci. Technol. , vol. 1, no. 1, pp. 256–263, Sep. 2011

  21. [26]

    Tera- hertz channel propagation phenomena, measurement techniques and modeling for 6G wireless communication applications: A survey, open challenges and future research directions,

    D. Serghiou, M. Khalily, T. W. C. Brown, and R. Tafazolli, “Tera- hertz channel propagation phenomena, measurement techniques and modeling for 6G wireless communication applications: A survey, open challenges and future research directions,” IEEE Commun. Surveys Tuts., vol. 24...

  22. [27]

    Survey of important issues in UA V communication networks,

    L. Gupta, R. Jain, and G. Vaszkun, “Survey of important issues in UA V communication networks,” IEEE Commun. Surveys Tuts., vol. 18, no. 2, pp. 1123–1152, Secondquarter 2016

  23. [28]

    Toward autonomous multi-UA V wireless network: A survey of reinforcement learning-based approaches,

    Y . Bai, H. Zhao, X. Zhang, Z. Chang, R. J ¨antti, and K. Yang, “Toward autonomous multi-UA V wireless network: A survey of reinforcement learning-based approaches,” IEEE Commun. Surveys Tuts. , vol. 25, no. 4, pp. 3038–3067, Fourthquarter 2023

  24. [29]

    A multi-agent collaborative environment learning method for UA V deployment and resource allocation,

    Z. Dai, Y . Zhang, W. Zhang, X. Luo, and Z. He, “A multi-agent collaborative environment learning method for UA V deployment and resource allocation,” IEEE Trans. Signal Inf. Process. Netw., vol. 8, pp. 120–130, Feb. 2022

  25. [30]

    Communication-efficient and distributed learning over wireless networks: Principles and applications,

    J. Park, S. Samarakoon, A. Elgabli, J. Kim, M. Bennis, S.-L. Kim, and M. Debbah, “Communication-efficient and distributed learning over wireless networks: Principles and applications,” Proc. IEEE, vol. 109, no. 5, pp. 796–819, May 2021

  26. [31]

    Impact of adaptive consistency on distributed SDN applications: An empirical study,

    E. Sakic and W. Kellerer, “Impact of adaptive consistency on distributed SDN applications: An empirical study,” IEEE J. Sel. Areas Commun. , vol. 36, no. 12, pp. 2702–2715, Dec. 2018

  27. [32]

    Distributed learning in wireless networks: Recent progress and future challenges,

    M. Chen, D. G ¨und¨uz, K. Huang, W. Saad, M. Bennis, A. V . Feljan, and H. V . Poor, “Distributed learning in wireless networks: Recent progress and future challenges,” IEEE J. Sel. Areas Commun. , vol. 39, no. 12, pp. 3579–3605, Dec. 2021

  28. [33]

    A survey on distributed machine learning,

    J. Verbraeken, M. Wolting, J. Katzy, J. Kloppenburg, T. Verbelen, and J. S. Rellermeyer, “A survey on distributed machine learning,” ACM Comput. Surveys, vol. 53, no. 2, pp. 1–33, Mar. 2020

  29. [34]

    Communication-efficient distributed learning: An overview,

    X. Cao, T. Bas ¸ar, S. Diggavi, Y . C. Eldar, K. B. Letaief, H. V . Poor, and J. Zhang, “Communication-efficient distributed learning: An overview,” IEEE J. Sel. Areas Commun. , vol. 41, no. 4, pp. 851–873, Apr. 2023

  30. [35]

    Envisioning device-to- device communications in 6G,

    S. Zhang, J. Liu, H. Guo, M. Qi, and N. Kato, “Envisioning device-to- device communications in 6G,” IEEE Netw., vol. 34, no. 3, pp. 86–91, May 2020

  31. [36]

    Survey on device to device (D2D) communication for 5GB/6G networks: Concept, applications, challenges, and future directions,

    M. S. M. Gismalla, A. I. Azmi, M. R. B. Salim, M. F. L. Abdullah, F. Iqbal, W. A. Mabrouk, M. B. Othman, A. Y . I. Ashyap, and A. S. M. Supa’at, “Survey on device to device (D2D) communication for 5GB/6G networks: Concept, applications, challenges, and future directions,” IEEE...

  32. [37]

    From federated to fog learning: Distributed machine learning over heterogeneous wireless networks,

    S. Hosseinalipour, C. G. Brinton, V . Aggarwal, H. Dai, and M. Chiang, “From federated to fog learning: Distributed machine learning over heterogeneous wireless networks,” IEEE Commun. Mag. , vol. 58, no. 12, pp. 41–47, Dec. 2020. 36

  33. [38]

    Single and multi-agent deep reinforcement learning for AI-enabled wireless networks: A tutorial,

    A. Feriani and E. Hossain, “Single and multi-agent deep reinforcement learning for AI-enabled wireless networks: A tutorial,” IEEE Commun. Surveys Tuts., vol. 23, no. 2, pp. 1226–1252, Secondquarter 2021

  34. [39]

    Multi-agent reinforcement learning: A review of challenges and applications,

    L. Canese, G. C. Cardarilli, L. Di Nunzio, R. Fazzolari, D. Giardino, M. Re, and S. Span `o, “Multi-agent reinforcement learning: A review of challenges and applications,” Appl. Sci. , vol. 11, no. 11, p. 4948, Apr. 2021

  35. [40]

    Learning to communicate with deep multi-agent reinforcement learning,

    J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” Proc. Adv. Neural Inf. Process. Syst. , vol. 29, 2016

  36. [41]

    Multiagent cooperation and competition with deep reinforcement learning,

    A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” PloS one, vol. 12, no. 4, Apr. 2017

  37. [42]

    Optimization for reinforcement learning: From a single agent to cooperative agents,

    D. Lee, N. He, P. Kamalaruban, and V . Cevher, “Optimization for reinforcement learning: From a single agent to cooperative agents,” IEEE Signal Process. Mag. , vol. 37, no. 3, pp. 123–135, May 2020

  38. [43]

    Applications of multi-agent reinforcement learning in future internet: A comprehensive survey,

    T. Li, K. Zhu, N. C. Luong, D. Niyato, Q. Wu, Y . Zhang, and B. Chen, “Applications of multi-agent reinforcement learning in future internet: A comprehensive survey,”IEEE Commun. Surveys Tuts., vol. 24, no. 2, pp. 1240–1279, Secondquarter 2022

  39. [44]

    Emergent communication in multi-agent reinforcement learning for future wireless networks,

    M. Chafii, S. Naoumi, R. Alami, E. Almazrouei, M. Bennis, and M. Debbah, “Emergent communication in multi-agent reinforcement learning for future wireless networks,” IEEE Internet Things Mag. , vol. 6, no. 4, pp. 18–24, Dec. 2023

  40. [46]

    Deep reinforcement learning: A brief survey,

    K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Process. Mag., vol. 34, no. 6, pp. 26–38, Nov. 2017

  41. [47]

    Applications of deep reinforcement learning in communications and networking: A survey,

    N. C. Luong, D. T. Hoang, S. Gong, D. Niyato, P. Wang, Y .-C. Liang, and D. I. Kim, “Applications of deep reinforcement learning in communications and networking: A survey,” IEEE Commun. Surveys Tuts., vol. 21, no. 4, pp. 3133–3174, Fourthquarter 2019

  42. [48]

    A review of safe reinforcement learning: Methods, theories, and applications,

    S. Gu, L. Yang, Y . Du, G. Chen, F. Walter, J. Wang, and A. Knoll, “A review of safe reinforcement learning: Methods, theories, and applications,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 12, pp. 11 216–11 235, Dec. 2024

  43. [49]

    Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,

    Y . S. Nasir and D. Guo, “Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,” IEEE J. Sel. Areas Commun., vol. 37, no. 10, pp. 2239–2250, Oct. 2019

  44. [50]

    Cell-free XL-MIMO meets multi-agent reinforcement learn- ing: Architectures, challenges, and future directions,

    Z. Liu, J. Zhang, Z. Liu, H. Du, Z. Wang, D. Niyato, M. Guizani, and B. Ai, “Cell-free XL-MIMO meets multi-agent reinforcement learn- ing: Architectures, challenges, and future directions,” IEEE Wireless Commun., vol. 31, no. 4, pp. 155–162, Aug. 2024

  45. [51]

    Reinforcement learning: A survey,

    L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” J. Artif. Intell. Res. , vol. 4, pp. 237–285, May 1996

  46. [52]

    Pilco: A model-based and data- efficient approach to policy search,

    M. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data- efficient approach to policy search,” in Proc. Int. Conf. Mach. Learn. , 2011

  47. [53]

    Model- based reinforcement learning: A survey,

    T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker et al. , “Model- based reinforcement learning: A survey,” Found. Trends Mach. Learn., vol. 16, no. 1, pp. 1–118, Jun. 2023

  48. [54]

    TRGP: Trust region gradient projection for continual learning,

    S. Lin, L. Yang, D. Fan, and J. Zhang, “TRGP: Trust region gradient projection for continual learning,” Proc. Int. Conf. Learn. Represent. , 2022

  49. [55]

    Multi-agent actor-critic for mixed cooperative-competitive en- vironments,

    R. Lowe, Y . I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mor- datch, “Multi-agent actor-critic for mixed cooperative-competitive en- vironments,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017

  50. [56]

    Is independent learning all you need in the starcraft multi-agent challenge?

    C. S. De Witt, T. Gupta, D. Makoviichuk, V . Makoviychuk, P. H. Torr, M. Sun, and S. Whiteson, “Is independent learning all you need in the starcraft multi-agent challenge?” arXiv preprint arXiv:2011.09533, 2020

  51. [57]

    When to trust your model: Model-based policy optimization,

    M. Janner, J. Fu, M. Zhang, and S. Levine, “When to trust your model: Model-based policy optimization,”Proc. Adv. Neural Inf. Process. Syst., vol. 32, 2019

  52. [58]

    Mastering atari, go, chess and shogi by planning with a learned model,

    J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel et al. , “Mastering atari, go, chess and shogi by planning with a learned model,” Nature, vol. 588, no. 7839, pp. 604–609, Dec. 2020

  53. [59]

    Efficient and scalable reinforcement learning for large-scale network control,

    C. Ma, A. Li, Y . Du, H. Dong, and Y . Yang, “Efficient and scalable reinforcement learning for large-scale network control,” Nat. Mach. Intell., vol. 6, no. 9, pp. 1006–1020, Sep. 2024

  54. [60]

    Multi-agent reinforcement learn- ing: A selective overview of theories and algorithms,

    K. Zhang, Z. Yang, and T. Bas ¸ar, “Multi-agent reinforcement learn- ing: A selective overview of theories and algorithms,” Handbook of reinforcement learning and control , pp. 321–384, Jun. 2021

  55. [61]

    Monotonic value function factorisation for deep multi- agent reinforcement learning,

    T. Rashid, M. Samvelyan, C. S. De Witt, G. Farquhar, J. Foerster, and S. Whiteson, “Monotonic value function factorisation for deep multi- agent reinforcement learning,” J. Mach. Learn. Res. , vol. 21, no. 178, pp. 1–51, Aug. 2020

  56. [62]

    Reducing overestimation bias in multi-agent domains using double centralized critics,

    J. Ackermann, V . Gabler, T. Osa, and M. Sugiyama, “Reducing overestimation bias in multi-agent domains using double centralized critics,” arXiv preprint arXiv:1910.01465 , 2019

  57. [63]

    Multi-agent reinforcement learning as a rehearsal for decentralized planning,

    L. Kraemer and B. Banerjee, “Multi-agent reinforcement learning as a rehearsal for decentralized planning,” Neurocomputing, vol. 190, pp. 82–94, May 2016

  58. [64]

    Multi-agent deep reinforcement learning: a survey,

    S. Gronauer and K. Diepold, “Multi-agent deep reinforcement learning: a survey,” Artif. Intell. Rev., vol. 55, no. 2, pp. 895–943, Apr. 2021

  59. [65]

    A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications,

    W. Du and S. Ding, “A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications,” Artif. Intell. Rev., vol. 54, no. 5, pp. 3215–3238, Nov. 2021

  60. [66]

    (a partial survey of) decentralized, cooperative multi-agent reinforcement learning,

    C. Amato, “(a partial survey of) decentralized, cooperative multi-agent reinforcement learning,” arXiv preprint arXiv:2405.06161 , 2024

  61. [67]

    A software-defined MARL-based ar- chitecture for AUV cluster network to enable cooperative and smart underwater target tracking,

    S. Zhu, G. Han, and C. Lin, “A software-defined MARL-based ar- chitecture for AUV cluster network to enable cooperative and smart underwater target tracking,” IEEE Wireless Commun. , vol. 31, no. 6, pp. 56–62, Dec. 2024

  62. [68]

    An introduction to centralized training for decentralized execution in cooperative multi-agent reinforcement learning,

    C. Amato, “An introduction to centralized training for decentralized execution in cooperative multi-agent reinforcement learning,” arXiv preprint arXiv:2409.03052, 2024

  63. [70]

    Fully decentralized cooperative multi-agent reinforcement learning: A survey,

    J. Jiang, K. Su, and Z. Lu, “Fully decentralized cooperative multi-agent reinforcement learning: A survey,” arXiv preprint arXiv:2401.04934 , 2024

  64. [71]

    DTDE: A new cooperative multi- agent reinforcement learning framework,

    G. Wen, J. Fu, P. Dai, and J. Zhou, “DTDE: A new cooperative multi- agent reinforcement learning framework,” The Innovation, vol. 2, no. 4, Nov. 2021

  65. [72]

    Survey on recent advances in multiagent reinforcement learning focusing on decentralized training with decentralized execution framework,

    Y . Shin, S. Seo, B. Yoo, H. Kim, H. Song, and S. Yi, “Survey on recent advances in multiagent reinforcement learning focusing on decentralized training with decentralized execution framework,” Electronics and Telecommunications Trends , vol. 38, no. 4, pp. 95– 103, Aug. 2023

  66. [73]

    Differential evolution for discrete-time large dynamic games,

    M. Suciu, R. I. Lung, N. Gask ´o, and D. Dumitrescu, “Differential evolution for discrete-time large dynamic games,” in IEEE Congress on Evolutionary Computation , 2013, pp. 2108–2113

  67. [74]

    F2a2: Flexible fully- decentralized approximate actor-critic for cooperative multi-agent re- inforcement learning,

    W. Li, B. Jin, X. Wang, J. Yan, and H. Zha, “F2a2: Flexible fully- decentralized approximate actor-critic for cooperative multi-agent re- inforcement learning,” J. Mach. Learn. Res., vol. 24, no. 178, pp. 1–75, Jun. 2023

  68. [75]

    Fully decentralized multi-agent reinforcement learning with networked agents,

    K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proc. Int. Conf. Mach. Learn. , 2018

  69. [76]

    Deep de- centralized multi-task multi-agent reinforcement learning under partial observability,

    S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep de- centralized multi-task multi-agent reinforcement learning under partial observability,” in Proc. Int. Conf. Mach. Learn. , 2017

  70. [77]

    Accelerating multiagent reinforcement learning by equilibrium transfer,

    Y . Hu, Y . Gao, and B. An, “Accelerating multiagent reinforcement learning by equilibrium transfer,” IEEE Trans. Cybern. , Jul

  71. [78]

    Multiagent adversarial collaborative learning via mean-field theory,

    G. Luo, H. Zhang, H. He, J. Li, and F.-Y . Wang, “Multiagent adversarial collaborative learning via mean-field theory,” IEEE Trans. Cybern. , vol. 51, no. 10, pp. 4994–5007, Oct. 2021

  72. [79]

    Deep learning and the information bottleneck principle,

    N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in IEEE Inf. Theory Workshop , 2015, pp. 1–5

  73. [80]

    Deep variational information bottleneck,

    A. A. Alemi, I. Fischer, J. V . Dillon, and K. Murphy, “Deep variational information bottleneck,” Proc. Int. Conf. Learn. Represent. , 2017

  74. [81]

    Graph information bottleneck,

    T. Wu, H. Ren, P. Li, and J. Leskovec, “Graph information bottleneck,” Proc. Adv. Neural Inf. Process. Syst., vol. 33, pp. 20 437–20 448, 2020

  75. [82]

    Graph structure learning with variational information bottleneck,

    Q. Sun, J. Li, H. Peng, J. Wu, X. Fu, C. Ji, and S. Y . Philip, “Graph structure learning with variational information bottleneck,” in Proc. AAAI Conf. Artif. Intell. , vol. 36, no. 4, 2022, pp. 4165–4174

  76. [83]

    Recognizing predictive substructures with subgraph information bottleneck,

    J. Yu, T. Xu, Y . Rong, Y . Bian, J. Huang, and R. He, “Recognizing predictive substructures with subgraph information bottleneck,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 3, pp. 1650–1663, Mar. 2024

  77. [84]

    Improving subgraph recognition with variational graph information bottleneck,

    J. Yu, J. Cao, and R. He, “Improving subgraph recognition with variational graph information bottleneck,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 19 396–19 405

  78. [85]

    Heterogeneous-agent reinforcement learning,

    Y . Zhong, J. G. Kuba, X. Feng, S. Hu, J. Ji, and Y . Yang, “Heterogeneous-agent reinforcement learning,” J. Mach. Learn. Res. , vol. 25, no. 1-67, p. 1, Jan. 2024. 37

  79. [86]

    Max- imum entropy heterogeneous-agent reinforcement learning,

    J. Liu, Y . Zhong, S. Hu, H. Fu, Q. Fu, X. Chang, and Y . Yang, “Max- imum entropy heterogeneous-agent reinforcement learning,” Proc. Int. Conf. Learn. Represent. , 2024

  80. [87]

    Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning,

    Q. Hao, W. Huang, T. Feng, J. Yuan, and Y . Li, “Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning,” in Proc. ACM SIGKDD Intl. Conf. Knowledge Discovery & Data Mining , 2023, pp. 685–697

  81. [88]

    Multi- cluster cooperative offloading for VR task: A MARL approach with graph embedding,

    Y . Yang, L. Feng, Y . Sun, Y . Li, W. Li, and M. A. Imran, “Multi- cluster cooperative offloading for VR task: A MARL approach with graph embedding,” IEEE Trans. Mobile Comput. , vol. 23, no. 9, pp. 8773–8788, Sep. 2024

  82. [89]

    Subgoal-based hierarchical reinforcement learning for multi-agent collaboration,

    C. Xu, C. Zhang, Y . Shi, R. Wang, S. Duan, Y . Wan, and X. Zhang, “Subgoal-based hierarchical reinforcement learning for multi-agent collaboration,” arXiv preprint arXiv:2408.11416 , 2024

  83. [90]

    6G internet of things: A comprehensive survey,

    D. C. Nguyen, M. Ding, P. N. Pathirana, A. Seneviratne, J. Li, D. Niyato, O. Dobre, and H. V . Poor, “6G internet of things: A comprehensive survey,”IEEE Internet Things J., vol. 9, no. 1, pp. 359– 383, Jan. 2022

  84. [91]

    Convergent communication, sensing and localization in 6G systems: An overview of technologies, opportunities and chal- lenges,

    C. De Lima, D. Belot, R. Berkvens, A. Bourdoux, D. Dardari, M. Guillaud, M. Isomursu, E.-S. Lohan, Y . Miao, A. N. Barreto, M. R. K. Aziz, J. Saloranta, T. Sanguanpuak, H. Sarieddeen, G. Seco- Granados, J. Suutala, T. Svensson, M. Valkama, B. Van Liempd, and H. Wymeersch, “Con...

  85. [92]

    Multiagent deep reinforcement learning for cost- and delay-sensitive virtual network function placement and routing,

    S. Wang, C. Yuen, W. Ni, Y . L. Guan, and T. Lv, “Multiagent deep reinforcement learning for cost- and delay-sensitive virtual network function placement and routing,” IEEE Trans. Commun., vol. 70, no. 8, pp. 5208–5224, Aug. 2022

  86. [93]

    6G: A comprehensive survey on technologies, applications, challenges, and research prob- lems,

    H. H. H. Mahmoud, A. A. Amer, and T. Ismail, “6G: A comprehensive survey on technologies, applications, challenges, and research prob- lems,” Trans. Emerg. Telecommun. Technol., vol. 32, no. 4, p. e4233, Apr. 2021

  87. [94]

    Are we ready for autonomous driving? the kitti vision benchmark suite,

    A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2012, pp. 3354–3361

  88. [95]

    Security and privacy in device-to-device (D2D) communication: A review,

    M. Haus, M. Waqas, A. Y . Ding, Y . Li, S. Tarkoma, and J. Ott, “Security and privacy in device-to-device (D2D) communication: A review,” IEEE Commun. Surveys Tuts. , vol. 19, no. 2, pp. 1054–1079, Secondquarter 2017

  89. [96]

    MIMO channel capacity for the distributed antenna,

    W. Roh and A. Paulraj, “MIMO channel capacity for the distributed antenna,” in Proc. IEEE Veh. Technol. Conf., vol. 2, 2002, pp. 706–709 vol.2

  90. [97]

    Cloud radio access network (C-RAN): a primer,

    J. Wu, Z. Zhang, Y . Hong, and Y . Wen, “Cloud radio access network (C-RAN): a primer,” IEEE Netw., vol. 29, no. 1, pp. 35–41, Jan. 2015

  91. [98]

    Downlink MIMO in LTE-advanced: SU-MIMO vs. MU-MIMO,

    L. Liu, R. Chen, S. Geirhofer, K. Sayana, Z. Shi, and Y . Zhou, “Downlink MIMO in LTE-advanced: SU-MIMO vs. MU-MIMO,” IEEE Commun. Mag. , vol. 50, no. 2, pp. 140–147, Feb. 2012

  92. [99]

    Stacked intelligent metasurfaces for efficient holo- graphic MIMO communications in 6G,

    J. An, C. Xu, D. W. K. Ng, G. C. Alexandropoulos, C. Huang, C. Yuen, and L. Hanzo, “Stacked intelligent metasurfaces for efficient holo- graphic MIMO communications in 6G,” IEEE J. Sel. Areas Commun. , vol. 41, no. 8, pp. 2380–2396, Aug. 2023

  93. [100]

    Fluid antenna systems,

    K.-K. Wong, A. Shojaeifard, K.-F. Tong, and Y . Zhang, “Fluid antenna systems,” IEEE Trans. Wireless Commun. , vol. 20, no. 3, pp. 1950– 1962, Mar. 2021

  94. [101]

    Movable antennas for wireless communication: Opportunities and challenges,

    L. Zhu, W. Ma, and R. Zhang, “Movable antennas for wireless communication: Opportunities and challenges,” IEEE Commun. Mag. , vol. 62, no. 6, pp. 114–120, Jun. 2024

  95. [102]

    The space-terrestrial integrated network: An overview,

    H. Yao, L. Wang, X. Wang, Z. Lu, and Y . Liu, “The space-terrestrial integrated network: An overview,”IEEE Commun. Mag., vol. 56, no. 9, pp. 178–185, 2018

  96. [103]

    Integrated sensing and communications: Toward dual- functional wireless networks for 6G and beyond,

    F. Liu, Y . Cui, C. Masouros, J. Xu, T. X. Han, Y . C. Eldar, and S. Buzzi, “Integrated sensing and communications: Toward dual- functional wireless networks for 6G and beyond,” IEEE J. Sel. Areas Commun., vol. 40, no. 6, pp. 1728–1767, Jun. 2022

  97. [104]

    MIMO transmission through reconfigurable intelligent surface: System design, analysis, and implementation,

    W. Tang, J. Y . Dai, M. Z. Chen, K.-K. Wong, X. Li, X. Zhao, S. Jin, Q. Cheng, and T. J. Cui, “MIMO transmission through reconfigurable intelligent surface: System design, analysis, and implementation,” IEEE J. Sel. Areas Commun. , vol. 38, no. 11, pp. 2683–2699, Nov. 2020

  98. [105]

    Identity and search in social networks,

    D. J. Watts, P. S. Dodds, and M. E. Newman, “Identity and search in social networks,” Sci., vol. 296, no. 5571, pp. 1302–1305, May 2002

  99. [106]

    Deep learning enabled se- mantic communication systems,

    H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled se- mantic communication systems,” IEEE Trans. Signal Process., vol. 69, pp. 2663–2675, Apr. 2021

  100. [107]

    Joint AP-UE association and precoding for SIM-aided cell-free massive MIMO systems,

    E. Shi, J. Zhang, J. An, G. Zhang, Z. Liu, C. Yuen, and B. Ai, “Joint AP-UE association and precoding for SIM-aided cell-free massive MIMO systems,” arXiv preprint arXiv:2409.12870 , 2024

  101. [108]

    UA V-enabled ultra-reliable low- latency communications for 6G: A comprehensive survey,

    A. Masaracchia, Y . Li, K. K. Nguyen, C. Yin, S. R. Khosravirad, D. B. D. Costa, and T. Q. Duong, “UA V-enabled ultra-reliable low- latency communications for 6G: A comprehensive survey,” IEEE Access, vol. 9, pp. 137 338–137 352, Oct. 2021

  102. [109]

    Covert communication in UA V-assisted air-ground networks,

    X. Jiang, X. Chen, J. Tang, N. Zhao, X. Y . Zhang, D. Niyato, and K.-K. Wong, “Covert communication in UA V-assisted air-ground networks,” IEEE Wireless Commun., vol. 28, no. 4, pp. 190–197, Aug. 2021

  103. [110]

    UA V-assisted emergency networks in disasters,

    N. Zhao, W. Lu, M. Sheng, Y . Chen, J. Tang, F. R. Yu, and K.-K. Wong, “UA V-assisted emergency networks in disasters,”IEEE Wireless Commun., vol. 26, no. 1, pp. 45–51, Feb. 2019

  104. [111]

    A survey of autonomous driving: Common practices and emerging technologies,

    E. Yurtsever, J. Lambert, A. Carballo, and K. Takeda, “A survey of autonomous driving: Common practices and emerging technologies,” IEEE Access, vol. 8, pp. 58 443–58 469, Mar. 2020

  105. [112]

    Networking and communications in autonomous driving: A survey,

    J. Wang, J. Liu, and N. Kato, “Networking and communications in autonomous driving: A survey,” IEEE Commun. Surveys Tuts., vol. 21, no. 2, pp. 1243–1274, Secondquarter 2019

  106. [113]

    An end-to-end curriculum learning approach for autonomous driving scenarios,

    L. Anzalone, P. Barra, S. Barra, A. Castiglione, and M. Nappi, “An end-to-end curriculum learning approach for autonomous driving scenarios,” IEEE Trans. Intell. Transp. Syst. , vol. 23, no. 10, pp. 19 817–19 826, Oct. 2022

  107. [114]

    End-to-end autonomous driving: Challenges and frontiers,

    L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-end autonomous driving: Challenges and frontiers,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 12, pp. 10 164–10 183, Dec. 2024

  108. [115]

    Computing systems for autonomous driving: State of the art and challenges,

    L. Liu, S. Lu, R. Zhong, B. Wu, Y . Yao, Q. Zhang, and W. Shi, “Computing systems for autonomous driving: State of the art and challenges,” IEEE Internet Things J. , vol. 8, no. 8, pp. 6469–6486, Apr. 2021

  109. [116]

    Asynchronous cell-free massive MIMO with rate-splitting,

    J. Zheng, J. Zhang, J. Cheng, V . C. M. Leung, D. W. K. Ng, and B. Ai, “Asynchronous cell-free massive MIMO with rate-splitting,” IEEE J. Sel. Areas Commun. , vol. 41, no. 5, pp. 1366–1382, May 2023

  110. [117]

    Wireless energy transfer in RIS-aided cell-free massive MIMO systems: Opportunities and challenges,

    E. Shi, J. Zhang, S. Chen, J. Zheng, Y . Zhang, D. W. Kwan Ng, and B. Ai, “Wireless energy transfer in RIS-aided cell-free massive MIMO systems: Opportunities and challenges,” IEEE Commun. Mag., vol. 60, no. 3, pp. 26–32, Mar. 2022

  111. [118]

    Uplink performance of RIS-aided cell-free massive MIMO system with electromagnetic interference,

    E. Shi, J. Zhang, D. W. K. Ng, and B. Ai, “Uplink performance of RIS-aided cell-free massive MIMO system with electromagnetic interference,” IEEE J. Sel. Areas Commun. , vol. 41, no. 8, pp. 2431– 2445, Aug. 2023

  112. [119]

    Playing atari with deep reinforcement learning,

    V . Mnih, “Playing atari with deep reinforcement learning,” arXiv preprint arXiv:1312.5602, 2013

  113. [120]

    Where does alphago go: From church-turing thesis to alphago thesis and beyond,

    F.-Y . Wang, J. J. Zhang, X. Zheng, X. Wang, Y . Yuan, X. Dai, J. Zhang, and L. Yang, “Where does alphago go: From church-turing thesis to alphago thesis and beyond,” IEEE/CAA J. Autom. Sin. , vol. 3, no. 2, pp. 113–120, May 2016

  114. [121]

    Recurrent world models facilitate policy evolution,

    D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” Proc. Adv. Neural Inf. Process. Syst. , vol. 31, 2018

  115. [122]

    The ubiquity of model- based reinforcement learning,

    B. B. Doll, D. A. Simon, and N. D. Daw, “The ubiquity of model- based reinforcement learning,” Curr. Opin. Neurobiol., vol. 22, no. 6, pp. 1075–1081, Aug. 2012

  116. [123]

    Model-based reinforcement learning for atari,

    L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine et al. , “Model-based reinforcement learning for atari,” Proc. Int. Conf. Learn. Represent., 2019

  117. [124]

    Model-based multi-agent policy optimization with adaptive opponent-wise rollouts,

    W. Zhang, X. Wang, J. Shen, and M. Zhou, “Model-based multi-agent policy optimization with adaptive opponent-wise rollouts,” Proc. Int. Joint Conf. Artif. Intell. , 2021

  118. [125]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752 , 2023

  119. [126]

    Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity,

    K. Zhang, S. Kakade, T. Basar, and L. Yang, “Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity,” Proc. Adv. Neural Inf. Process. Syst. , vol. 33, pp. 1166–1178, 2020

  120. [127]

    Large language model based multi-agents: A survey of progress and challenges,

    T. Guo, X. Chen, Y . Wang, R. Chang, S. Pei, N. V . Chawla, O. Wiest, and X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” Int. Joint Conf. Artif. Intell. , 2024

  121. [128]

    A survey of multi-agent reinforce- ment learning with communication,

    C. Zhu, M. Dastani, and S. Wang, “A survey of multi-agent reinforce- ment learning with communication,” Auton. Agents Multi-Agent Syst. , vol. 1, Mar. 2022

  122. [129]

    Learning multiagent communication with back propagation,

    S. Sukhbaatar, R. Fergus et al. , “Learning multiagent communication with back propagation,” Proc. Adv. Neural Inf. Process. Syst. , vol. 29, 2016

  123. [130]

    Value-decomposition networks for cooperative multi-agent learning,

    P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V . Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls et al. , “Value-decomposition networks for cooperative multi-agent learning,” arXiv preprint arXiv:1706.05296 , 2017

  124. [131]

    Learning to communicate with deep multi-agent reinforcement learning,

    J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” Proc. Adv. Neural Inf. Process. Syst. , vol. 29, 2016. 38

  125. [132]

    Counterfactual multi-agent policy gradients,

    J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson, “Counterfactual multi-agent policy gradients,” in AAAI Conf. Artif. Intell., vol. 32, no. 1, 2018

  126. [133]

    The surprising effectiveness of PPO in cooperative multi-agent games,

    C. Yu, A. Velu, E. Vinitsky, J. Gao, Y . Wang, A. Bayen, and Y . Wu, “The surprising effectiveness of PPO in cooperative multi-agent games,” Proc. Adv. Neural Inf. Process. Syst. , vol. 35, pp. 24 611– 24 624, 2022

  127. [134]

    QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning,

    K. Son, D. Kim, W. J. Kang, D. E. Hostallero, and Y . Yi, “QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning,” in Proc. Int. Conf. Mach. Learn. , 2019

  128. [135]

    Learning attentional communication for multi- agent cooperation,

    J. Jiang and Z. Lu, “Learning attentional communication for multi- agent cooperation,” Proc. Adv. Neural Inf. Process. Syst., vol. 31, 2018

  129. [136]

    Actor-attention-critic for multi-agent reinforce- ment learning,

    S. Iqbal and F. Sha, “Actor-attention-critic for multi-agent reinforce- ment learning,” in Proc. Int. Conf. Mach. Learn. , 2019

  130. [137]

    Mean-field approx- imation of cooperative constrained multi-agent reinforcement learning (CMARL),

    W. U. Mondal, V . Aggarwal, and S. V . Ukkusuri, “Mean-field approx- imation of cooperative constrained multi-agent reinforcement learning (CMARL),” J. Mach. Learn. Res. , vol. 25, no. 260, pp. 1–33, Mar. 2024

  131. [138]

    Multi-agent reinforcement learning-based joint precoding and phase shift optimization for RIS- aided cell-free massive MIMO systems,

    Y . Zhu, E. Shi, Z. Liu, J. Zhang, and B. Ai, “Multi-agent reinforcement learning-based joint precoding and phase shift optimization for RIS- aided cell-free massive MIMO systems,” IEEE Trans. Veh. Technol. , vol. 73, no. 9, pp. 14 015–14 020, Sep. 2024

  132. [139]

    Algorithmic framework for model-based deep reinforcement learning with theoret- ical guarantees,

    Y . Luo, H. Xu, Y . Li, Y . Tian, T. Darrell, and T. Ma, “Algorithmic framework for model-based deep reinforcement learning with theoret- ical guarantees,” Proc. Int. Conf. Learn. Represent. , 2018

  133. [140]

    Coordinating multi-agent reinforcement learning with limited communication,

    C. Zhang and V . Lesser, “Coordinating multi-agent reinforcement learning with limited communication,” in Proc. AAMAS, 2013

  134. [141]

    Multi-agent reinforcement learning is a sequence modeling problem,

    M. Wen, J. Kuba, R. Lin, W. Zhang, Y . Wen, J. Wang, and Y . Yang, “Multi-agent reinforcement learning is a sequence modeling problem,” Proc. Adv. Neural Inf. Process. Syst., vol. 35, pp. 16 509–16 521, 2022

  135. [142]

    Winder, Reinforcement learning

    P. Winder, Reinforcement learning. O’Reilly Media, 2020

  136. [143]

    Decentralized multi-agent reinforce- ment learning with networked agents: Recent advances,

    K. Zhang, Z. Yang, and T. Bas ¸ar, “Decentralized multi-agent reinforce- ment learning with networked agents: Recent advances,” Front. Inf. Technol. Electron. Eng., vol. 22, no. 6, pp. 802–814, Jun. 2021

  137. [144]

    Efficient communication in multi- agent reinforcement learning via variance based control,

    S. Q. Zhang, Q. Zhang, and J. Lin, “Efficient communication in multi- agent reinforcement learning via variance based control,” Proc. Adv. Neural Inf. Process. Syst. , vol. 32, 2019

  138. [145]

    Multi-agent reinforcement learning: A comprehensive survey,

    D. Huh and P. Mohapatra, “Multi-agent reinforcement learning: A comprehensive survey,” arXiv preprint arXiv:2312.10256 , 2023

  139. [146]

    Model predictive actor-critic: Accelerating robot skill acquisition with deep reinforcement learning,

    A. S. Morgan, D. Nandha, G. Chalvatzaki, C. D’Eramo, A. M. Dollar, and J. Peters, “Model predictive actor-critic: Accelerating robot skill acquisition with deep reinforcement learning,” in IEEE Int. Conf. Robot. Auton., 2021, pp. 6672–6678

  140. [147]

    Generalized policy iteration adaptive dynamic programming for discrete-time nonlinear systems,

    D. Liu, Q. Wei, and P. Yan, “Generalized policy iteration adaptive dynamic programming for discrete-time nonlinear systems,” IEEE Trans. Syst. Man Cybern. , vol. 45, no. 12, pp. 1577–1591, May 2015

  141. [148]

    Scalable trust-region method for deep reinforcement learning using kronecker- factored approximation,

    Y . Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba, “Scalable trust-region method for deep reinforcement learning using kronecker- factored approximation,” Proc. Adv. Neural Inf. Process. Syst., vol. 30, 2017

  142. [150]

    Mobile cell-free massive MIMO with multi-agent reinforcement learning: A scalable framework,

    Z. Liu, J. Zhang, Y . Zhu, E. Shi, and B. Ai, “Mobile cell-free massive MIMO with multi-agent reinforcement learning: A scalable framework,” IEEE Trans. Wireless Commun. , early access, 2024

  143. [151]

    Performant, memory efficient and scalable multi-agent reinforcement learning,

    O. Mahjoub, S. Abramowitz, R. de Kock, W. Khlifi, S. d. Toit, J. Daniel, L. B. Nessir, L. Beyers, C. Formanek, L. Clark et al. , “Performant, memory efficient and scalable multi-agent reinforcement learning,” arXiv preprint arXiv:2410.01706 , 2024

  144. [152]

    Combating bilateral edge noise for robust link prediction,

    Z. Zhou, J. Yao, J. Liu, X. Guo, Q. Yao, L. He, L. Wang, B. Zheng, and B. Han, “Combating bilateral edge noise for robust link prediction,” Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2024

  145. [154]

    Lateral transfer learning for multiagent reinforcement learning,

    H. Shi, J. Li, J. Mao, and K.-S. Hwang, “Lateral transfer learning for multiagent reinforcement learning,” IEEE Trans. Cybern. , vol. 53, no. 3, pp. 1699–1711, Mar. 2023

  146. [155]

    Enhancing collaboration in het- erogeneous multiagent systems through communication complementary graph,

    K. Peng, T. Ma, L. Jia, and H. Rong, “Enhancing collaboration in het- erogeneous multiagent systems through communication complementary graph,” IEEE Trans. Cybern. , vol. 54, no. 11, pp. 6881–6894, Nov. 2024

  147. [156]

    Graph neural networks for wireless communications: From theory to practice,

    Y . Shen, J. Zhang, S. H. Song, and K. B. Letaief, “Graph neural networks for wireless communications: From theory to practice,” IEEE Trans. Wireless Commun., vol. 22, no. 5, pp. 3554–3569, May 2023

  148. [157]

    Deep multi-agent reinforcement learning with relevance graphs,

    A. Malysheva, T. T. Sung, C.-B. Sohn, D. Kudenko, and A. Shpilman, “Deep multi-agent reinforcement learning with relevance graphs,”Proc. Adv. Neural Inf. Process. Syst. , 2018

  149. [158]

    Learning multi-agent com- munication from graph modeling perspective,

    S. Hu, L. Shen, Y . Zhang, and D. Tao, “Learning multi-agent com- munication from graph modeling perspective,” Proc. Int. Conf. Learn. Represent., 2024

  150. [159]

    Learning correlated communication topology in multi-agent reinforcement learning,

    Y . Du, B. Liu, V . Moens, Z. Liu, Z. Ren, J. Wang, X. Chen, and H. Zhang, “Learning correlated communication topology in multi-agent reinforcement learning,” in Proc. AAMAS, 2021, pp. 456–464

  151. [160]

    Path reasoning over knowledge graph: A multi-agent and reinforcement learning based method,

    Z. Li, X. Jin, S. Guan, Y . Wang, and X. Cheng, “Path reasoning over knowledge graph: A multi-agent and reinforcement learning based method,” in IEEE Int. Conf. Data Mining Workshops , 2018, pp. 929– 936

  152. [161]

    Relational reasoning via set transformers: Provable efficiency and applications to MARL,

    F. Zhang, B. Liu, K. Wang, V . Tan, Z. Yang, and Z. Wang, “Relational reasoning via set transformers: Provable efficiency and applications to MARL,” Proc. Adv. Neural Inf. Process. Syst. , vol. 35, pp. 35 825– 35 838, 2022

  153. [162]

    Graph convolutional network-based topology embedded deep reinforcement learning for voltage stability control,

    R. R. Hossain, Q. Huang, and R. Huang, “Graph convolutional network-based topology embedded deep reinforcement learning for voltage stability control,” IEEE Trans. Power Syst. , vol. 36, no. 5, pp. 4848–4851, Sep. 2021

  154. [163]

    Leveraging joint-action embedding in multiagent reinforcement learning for coop- erative games,

    X. Lou, J. Zhang, Y . Du, C. Yu, Z. He, and K. Huang, “Leveraging joint-action embedding in multiagent reinforcement learning for coop- erative games,” IEEE Trans. Games, vol. 16, no. 2, pp. 470–482, Jun. 2024

  155. [164]

    State augmentation via self- supervision in offline multiagent reinforcement learning,

    S. Wang, X. Li, H. Qu, and W. Chen, “State augmentation via self- supervision in offline multiagent reinforcement learning,” IEEE Trans. Cogn. Develop. Syst. , vol. 16, no. 3, pp. 1051–1062, Jun. 2024

  156. [165]

    Virtual network embedding based on hierarchical cooperative multiagent reinforcement learning,

    H.-K. Lim, I. Ullah, J.-B. Kim, and Y .-H. Han, “Virtual network embedding based on hierarchical cooperative multiagent reinforcement learning,” IEEE Internet Things J., vol. 11, no. 5, pp. 8552–8568, Mar. 2024

  157. [166]

    AHAC: Actor hierarchical attention critic for multi-agent reinforcement learn- ing,

    Y . Wang, D. Shi, C. Xue, H. Jiang, G. Wang, and P. Gong, “AHAC: Actor hierarchical attention critic for multi-agent reinforcement learn- ing,” in IEEE Int. Conf. Syst., Man, Cybern. , 2020, pp. 3013–3020

  158. [167]

    Hierarchical learning with heuristic guidance for multi-task assignment and distributed planning in inter- active scenarios,

    S. Chen, M. Wang, and W. Song, “Hierarchical learning with heuristic guidance for multi-task assignment and distributed planning in inter- active scenarios,” IEEE Trans. Intell. Veh., pp. 1–13, Feb. 2024

  159. [168]

    Hierarchical multi-agent reinforcement learning for repair crews dispatch control towards multi-energy microgrid resilience,

    D. Qiu, Y . Wang, T. Zhang, M. Sun, and G. Strbac, “Hierarchical multi-agent reinforcement learning for repair crews dispatch control towards multi-energy microgrid resilience,” Applied Energy, vol. 336, p. 120826, 2023

  160. [169]

    D. Qiu, Y . Wang, M. Sun, and G. Strbac, “Multi-service provision for electric vehicles in power-transportation networks towards a low- carbon transition: A hierarchical and hybrid multi-agent reinforcement learning approach,” Applied Energy, vol. 313, p. 118790, 2022

  161. [170]

    Effective commu- nications: A joint learning and communication framework for multi- agent reinforcement learning over noisy channels,

    T.-Y . Tung, S. Kobus, J. P. Roig, and D. G ¨und¨uz, “Effective commu- nications: A joint learning and communication framework for multi- agent reinforcement learning over noisy channels,” IEEE J. Sel. Areas Commun., vol. 39, no. 8, pp. 2590–2603, Aug. 2021

  162. [171]

    Learning structured communication for multi-agent reinforcement learning,

    J. Sheng, X. Wang, B. Jin, J. Yan, W. Li, T.-H. Chang, J. Wang, and H. Zha, “Learning structured communication for multi-agent reinforcement learning,” Auton. Agent Multi-Agent Syst., vol. 36, no. 2, Aug. 2022

  163. [172]

    Learning when to communicate at scale in multiagent cooperative and competitive tasks,

    A. Singh, T. Jain, and S. Sukhbaatar, “Learning when to communicate at scale in multiagent cooperative and competitive tasks,” Proc. Int. Conf. Learn. Represent. , 2018

  164. [173]

    Learning to schedule communication in multi-agent reinforcement learning,

    D. Kim, S. Moon, D. Hostallero, W. J. Kang, T. Lee, K. Son, and Y . Yi, “Learning to schedule communication in multi-agent reinforcement learning,” Proc. Int. Conf. Learn. Represent. , 2019

  165. [174]

    Multi-agent graph-attention communication and teaming,

    Y . Niu, R. R. Paleja, and M. C. Gombolay, “Multi-agent graph-attention communication and teaming,” in Proc. AAMAS, 2021

  166. [175]

    Tarmac: Targeted multi-agent communication,

    A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, and J. Pineau, “Tarmac: Targeted multi-agent communication,” in Proc. Int. Conf. Mach. Learn. , 2019

  167. [176]

    A survey of multi-agent reinforce- ment learning with communication,

    C. Zhu, M. Dastani, and S. Wang, “A survey of multi-agent reinforce- ment learning with communication,” Auton. Agent Multi-Agent Syst. , vol. 38, no. 4, Jan. 2024

  168. [177]

    An efficient transfer learning framework for multiagent reinforcement learning,

    T. Yang, W. Wang, H. Tang, J. Hao, Z. Meng, H. Mao, D. Li, W. Liu, Y . Chen, Y . Hu et al. , “An efficient transfer learning framework for multiagent reinforcement learning,” Proc. Adv. Neural Inf. Process. Syst., vol. 34, pp. 17 037–17 048, 2021

  169. [178]

    Deep multitask multiagent reinforcement learning with knowledge transfer,

    Y . Mai, Y . Zang, Q. Yin, W. Ni, and K. Huang, “Deep multitask multiagent reinforcement learning with knowledge transfer,” IEEE Trans. Games, vol. 16, no. 3, pp. 566–576, Sep. 2024. 39

  170. [179]

    Communication in multi-agent re- inforcement learning: Intention sharing,

    W. Kim, J. Park, and Y . Sung, “Communication in multi-agent re- inforcement learning: Intention sharing,” in Proc. Int. Conf. Learn. Represent., 2021

  171. [180]

    Learning task-oriented communication for edge inference: An information bottleneck approach,

    J. Shao, Y . Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE J. Sel. Areas Commun., vol. 40, no. 1, pp. 197–211, Jan. 2022

  172. [181]

    Self-supervised information bottleneck for deep multi-view subspace clustering,

    S. Wang, C. Li, Y . Li, Y . Yuan, and G. Wang, “Self-supervised information bottleneck for deep multi-view subspace clustering,” IEEE Trans. Image Process., vol. 32, pp. 1555–1567, Feb. 2023

  173. [182]

    CCGIB: A cross-channel graph information bottleneck principle,

    X. Fan, M. Gong, Y . Wu, M. Zhang, H. Li, and X. Jiang, “CCGIB: A cross-channel graph information bottleneck principle,” Trans. Neural Netw. Learn. Syst. , early access, 2024

  174. [183]

    Self-supervised graph information bottleneck for multiview molecular embedding learning,

    C. Li, K. Mao, S. Wang, Y . Yuan, and G. Wang, “Self-supervised graph information bottleneck for multiview molecular embedding learning,” IEEE Trans. Artif. Intell. , vol. 5, no. 4, pp. 1554–1562, Jul. 2024

  175. [184]

    Dynamic graph information bottleneck,

    H. Yuan, Q. Sun, X. Fu, C. Ji, and J. Li, “Dynamic graph information bottleneck,” in Proc. Web Conf., 2024, pp. 469–480

  176. [185]

    Robust multi- agent communication with graph information bottleneck optimization,

    S. Ding, W. Du, L. Ding, J. Zhang, L. Guo, and B. An, “Robust multi- agent communication with graph information bottleneck optimization,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 5, pp. 3096–3107, May 2024

  177. [186]

    Conditional graph information bottleneck for molecular relational learning,

    N. Lee, D. Hyun, G. S. Na, S. Kim, J. Lee, and C. Park, “Conditional graph information bottleneck for molecular relational learning,” in Proc. Int. Conf. Mach. Learn. , 2023

  178. [187]

    Heterogeneous graph information bottleneck

    L. Yang, F. Wu, Z. Zheng, B. Niu, J. Gu, C. Wang, X. Cao, and Y . Guo, “Heterogeneous graph information bottleneck.” in Proc. Int. Joint Conf. Artif. Intell., 2021, pp. 1638–1645

  179. [188]

    Sugar: Subgraph neural network with reinforcement pooling and self- supervised mutual information mechanism,

    Q. Sun, J. Li, H. Peng, J. Wu, Y . Ning, P. S. Yu, and L. He, “Sugar: Subgraph neural network with reinforcement pooling and self- supervised mutual information mechanism,” in Proc. Web Conf., 2021, pp. 2081–2091

  180. [189]

    Adaptive subgraph neural network with reinforced critical structure mining,

    J. Li, Q. Sun, H. Peng, B. Yang, J. Wu, and P. S. Yu, “Adaptive subgraph neural network with reinforced critical structure mining,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 7, pp. 8063–8080, 2023

  181. [190]

    The information bottleneck method,

    N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” Proc. Annu. Allerton Conf. Commun. Control Comput. , pp. 368–377, 1999

  182. [191]

    The information bottleneck: Theory and applications,

    N. Slonim, “The information bottleneck: Theory and applications,” Ph.D. dissertation, Ph.D Thesis, Hebrew, Hebrew Univ., Israel, 2002

  183. [192]

    On the information bottleneck theory of deep learning,

    A. M. Saxe, Y . Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the information bottleneck theory of deep learning,” Proc. Int. Conf. Learn. Represent. , 2018

  184. [193]

    On variational bounds of mutual information,

    B. Poole, S. Ozair, A. Van Den Oord, A. Alemi, and G. Tucker, “On variational bounds of mutual information,” in Proc. Int. Conf. Learn. Represent., 2019, pp. 5171–5180

  185. [194]

    A survey on information bottleneck,

    S. Hu, Z. Lou, X. Yan, and Y . Ye, “A survey on information bottleneck,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 8, pp. 5325–5344, Aug. 2024

  186. [195]

    Graph information bottleneck for subgraph recognition,

    J. Yu, T. Xu, Y . Rong, Y . Bian, J. Huang, and R. He, “Graph information bottleneck for subgraph recognition,” Proc. Int. Conf. Learn. Represent., 2021

  187. [196]

    Learnable graph convolutional network with semisupervised graph information bottleneck,

    L. Zhong, Z. Chen, Z. Wu, S. Du, Z. Chen, and S. Wang, “Learnable graph convolutional network with semisupervised graph information bottleneck,” Trans. Neural Netw. Learn. Syst. , pp. 1–14, Oct. 2023

  188. [197]

    Toward enhanced robustness in unsupervised graph representation learning: A graph information bottleneck perspective,

    J. Wang, M. Luo, J. Li, Z. Liu, J. Zhou, and Q. Zheng, “Toward enhanced robustness in unsupervised graph representation learning: A graph information bottleneck perspective,” IEEE Trans. Knowl. Data Eng., Aug. 2024

  189. [198]

    A novel framework for policy mirror descent with general parameterization and linear convergence,

    C. Alfano, R. Yuan, and P. Rebeschini, “A novel framework for policy mirror descent with general parameterization and linear convergence,” Proc. Adv. Neural Inf. Process. Syst., vol. 36, pp. 30 681–30 725, 2024

  190. [199]

    Linear con- vergence of natural policy gradient methods with log-linear policies,

    R. Yuan, S. S. Du, R. M. Gower, A. Lazaric, and L. Xiao, “Linear con- vergence of natural policy gradient methods with log-linear policies,” Proc. Int. Conf. Learn. Represent. , 2023

  191. [200]

    A general class of surrogate functions for stable and efficient reinforcement learning,

    S. Vaswani, O. Bachem, S. Totaro, R. M ¨uller, S. Garg, M. Geist, M. C. Machado, P. S. Castro, and N. L. Roux, “A general class of surrogate functions for stable and efficient reinforcement learning,” arXiv preprint arXiv:2108.05828 , 2021

  192. [201]

    Scalable multi-agent model-based rein- forcement learning,

    V . Egorov and A. Shpilman, “Scalable multi-agent model-based rein- forcement learning,” arXiv preprint arXiv:2205.15023 , 2022

  193. [202]

    Large language model based multi-agents: A survey of progress and challenges,

    T. Guo, X. Chen, Y . Wang, R. Chang, S. Pei, N. V . Chawla, O. Wiest, and X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” arXiv preprint arXiv:2402.01680 , 2024

  194. [203]

    Robust multi-agent reinforcement learning with state uncertainty,

    S. He, S. Han, S. Su, S. Han, S. Zou, and F. Miao, “Robust multi-agent reinforcement learning with state uncertainty,” Proc. Int. Conf. Learn. Represent., 2023

  195. [204]

    Asp: Learn a universal neural solver!

    C. Wang, Z. Yu, S. McAleer, T. Yu, and Y . Yang, “Asp: Learn a universal neural solver!” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 6, pp. 4102–4114, Jan. 2024

  196. [205]

    Decentralized transformers with centralized aggregation are sample-efficient multi- agent world models,

    Y . Zhang, C. Bai, B. Zhao, J. Yan, X. Li, and X. Li, “Decentralized transformers with centralized aggregation are sample-efficient multi- agent world models,” arXiv preprint arXiv:2406.15836 , 2024

  197. [206]

    RMIO: A model-based MARL framework for scenarios with observation loss in some agents,

    S. Zifeng, L. Meiqin, Z. Senlin, Z. Ronghao, and D. Shanling, “RMIO: A model-based MARL framework for scenarios with observation loss in some agents,” arXiv preprint arXiv:2411.19639 , 2024

  198. [207]

    Mirror learning: A unifying framework of policy optimisation,

    J. Grudzien, C. A. S. De Witt, and J. Foerster, “Mirror learning: A unifying framework of policy optimisation,” in Proc. Int. Conf. Mach. Learn., 2022

  199. [208]

    A novel framework for policy mirror descent with general parameterization and linear convergence,

    C. Alfano, R. Yuan, and P. Rebeschini, “A novel framework for policy mirror descent with general parameterization and linear convergence,” Proc. Adv. Neural Inf. Process. Syst., vol. 36, pp. 30 681–30 725, 2023

  200. [209]

    Bridging evolutionary algorithms and reinforcement learning: A comprehensive survey on hybrid algorithms,

    P. Li, J. Hao, H. Tang, X. Fu, Y . Zhen, and K. Tang, “Bridging evolutionary algorithms and reinforcement learning: A comprehensive survey on hybrid algorithms,” IEEE Trans. Evol. Comput. , Aug. 2024

  201. [210]

    Multi-task multi-agent reinforcement learning with interaction and task representations,

    C. Li, S. Dong, S. Yang, Y . Hu, T. Ding, W. Li, and Y . Gao, “Multi-task multi-agent reinforcement learning with interaction and task representations,” IEEE Trans. Neural Netw. Learn. Syst. , Oct. 2024

  202. [211]

    A survey on offline reinforcement learning: Taxonomy, review, and open problems,

    R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,” IEEE Trans. Neural Netw. Learn. Syst. , Mar. 2023

  203. [212]

    Towards efficient collaboration via graph modeling in reinforcement learning,

    W. Fan, Z. Yu, C. Ma, C. Li, Y . Yang, and X. Zhang, “Towards efficient collaboration via graph modeling in reinforcement learning,” arXiv preprint arXiv:2410.15841, 2024

  204. [213]

    Reinforcement learning and control as probabilistic infer- ence: Tutorial and review,

    S. Levine, “Reinforcement learning and control as probabilistic infer- ence: Tutorial and review,” arXiv preprint arXiv:1805.00909 , 2018

  205. [215]

    Reward-free curricula for training robust world models,

    M. Rigter, M. Jiang, and I. Posner, “Reward-free curricula for training robust world models,” Proc. Int. Conf. Learn. Represent. , 2024

  206. [216]

    Offline reinforcement learning as one big sequence modeling problem,

    M. Janner, Q. Li, and S. Levine, “Offline reinforcement learning as one big sequence modeling problem,” Proc. Adv. Neural Inf. Process. Syst., vol. 34, pp. 1273–1286, 2021

  207. [217]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,” Proc. NAACL-HLT, 2018

  208. [218]

    Neural discrete representation learning,

    A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017

  209. [219]

    A multi-agent reinforcement learning approach for safe and efficient behavior planning of connected autonomous vehicles,

    S. Han, S. Zhou, J. Wang, L. Pepin, C. Ding, J. Fu, and F. Miao, “A multi-agent reinforcement learning approach for safe and efficient behavior planning of connected autonomous vehicles,” IEEE Trans. Intell. Transp. Syst. , vol. 25, no. 5, pp. 3654–3670, May 2024

  210. [220]

    Multi-agent reinforcement learning with policy clipping and average evaluation for UA V-assisted communication markov game,

    Z. Feng, M. Huang, D. Wu, E. Q. Wu, and C. Yuen, “Multi-agent reinforcement learning with policy clipping and average evaluation for UA V-assisted communication markov game,” IEEE Trans. Intell. Transp. Syst., vol. 24, no. 12, pp. 14 281–14 293, Dec. 2023

  211. [221]

    Graph-attention-based reinforcement learning for trajectory design and resource assignment in multi-UA V-assisted communication,

    Z. Feng, D. Wu, M. Huang, and C. Yuen, “Graph-attention-based reinforcement learning for trajectory design and resource assignment in multi-UA V-assisted communication,”IEEE Internet Things J. , vol. 11, no. 16, pp. 27 421–27 434, Aug. 2024

  212. [222]

    Resource allocation in UA V-assisted networks: A clustering-aided reinforcement learning approach,

    S. Zhou, Y . Cheng, X. Lei, Q. Peng, J. Wang, and S. Li, “Resource allocation in UA V-assisted networks: A clustering-aided reinforcement learning approach,” IEEE Trans. Veh. Technol. , vol. 71, no. 11, pp. 12 088–12 103, Nov. 2022

  213. [223]

    UA V-enabled secure communications by multi-agent deep reinforcement learning,

    Y . Zhang, Z. Mou, F. Gao, J. Jiang, R. Ding, and Z. Han, “UA V-enabled secure communications by multi-agent deep reinforcement learning,” IEEE Trans. Veh. Technol. , vol. 69, no. 10, pp. 11 599–11 611, Oct. 2020

  214. [224]

    Secure offloading with adversarial multi-agent reinforcement learning against intelligent eavesdroppers in UA V-enabled mobile edge computing,

    X. Li, W. Huangfu, X. Xu, J. Huo, and K. Long, “Secure offloading with adversarial multi-agent reinforcement learning against intelligent eavesdroppers in UA V-enabled mobile edge computing,” IEEE Trans. Mobile Comput., vol. 23, no. 12, pp. 13 914–13 928, Dec. 2024

  215. [225]

    Dg-trans: Dual-level graph transformer for spatiotemporal incident impact prediction on traffic networks,

    Y . Sun, K. Fu, and C.-T. Lu, “Dg-trans: Dual-level graph transformer for spatiotemporal incident impact prediction on traffic networks,” arXiv preprint arXiv:2303.12238 , 2023

  216. [226]

    Multi- agent DRL for task offloading and resource allocation in multi-UA V enabled IoT edge network,

    A. M. Seid, G. O. Boateng, B. Mareri, G. Sun, and W. Jiang, “Multi- agent DRL for task offloading and resource allocation in multi-UA V enabled IoT edge network,” IEEE Trans. Netw. Service Manag., vol. 18, no. 4, pp. 4531–4547, Dec. 2021

  217. [227]

    Computing over space-air-ground inte- grated networks: Challenges and opportunities,

    B. Shang, Y . Yi, and L. Liu, “Computing over space-air-ground inte- grated networks: Challenges and opportunities,” IEEE Netw., vol. 35, no. 4, pp. 302–309, Aug. 2021. 40

  218. [228]

    Cooperative multi-target positioning for cell-free massive MIMO with multi-agent reinforcement learning,

    Z. Liu, J. Zhang, E. Shi, Y . Zhu, D. W. K. Ng, and B. Ai, “Cooperative multi-target positioning for cell-free massive MIMO with multi-agent reinforcement learning,” IEEE Trans. Wireless Commun. , vol. 23, no. 12, pp. 19 034–19 049, Dec. 2024

  219. [229]

    Joint cooperative clustering and power control for energy-efficient cell-free XL-MIMO with multi-agent reinforcement learning,

    Z. Liu, J. Zhang, Z. Liu, D. W. K. Ng, and B. Ai, “Joint cooperative clustering and power control for energy-efficient cell-free XL-MIMO with multi-agent reinforcement learning,” IEEE Trans. Commun. , vol. 72, no. 12, pp. 7772–7786, Dec. 2024

  220. [230]

    Multi-agent reinforcement learning for autonomous driving: A survey,

    R. Zhang, J. Hou, F. Walter, S. Gu, J. Guan, F. R ¨ohrbein, Y . Du, P. Cai, G. Chen, and A. Knoll, “Multi-agent reinforcement learning for autonomous driving: A survey,” arXiv preprint arXiv:2408.09675 , 2024

  221. [231]

    Deep reinforcement learning for autonomous driving: A survey,

    B. R. Kiran, I. Sobh, V . Talpaert, P. Mannion, A. A. A. Sallab, S. Yo- gamani, and P. P ´erez, “Deep reinforcement learning for autonomous driving: A survey,” IEEE Trans. Intell. Transp. Syst. , vol. 23, no. 6, pp. 4909–4926, Jun. 2022

  222. [232]

    Survey of deep reinforcement learning for motion planning of autonomous vehicles,

    S. Aradi, “Survey of deep reinforcement learning for motion planning of autonomous vehicles,” IEEE Trans. Intell. Transp. Syst. , vol. 23, no. 2, pp. 740–759, Feb. 2022

  223. [233]

    A review of motion planning for highway autonomous driving,

    L. Claussmann, M. Revilloud, D. Gruyer, and S. Glaser, “A review of motion planning for highway autonomous driving,” IEEE Trans. Intell. Transp. Syst., vol. 21, no. 5, pp. 1826–1848, May 2020

  224. [234]

    Differentiable integrated motion prediction and planning with learnable cost function for autonomous driving,

    Z. Huang, H. Liu, J. Wu, and C. Lv, “Differentiable integrated motion prediction and planning with learnable cost function for autonomous driving,” IEEE Trans. Neural Netw. Learn. Syst. , vol. 35, no. 11, pp. 15 222–15 236, Nov. 2024

  225. [235]

    Learning decentralized traffic signal controllers with multi-agent graph reinforcement learning,

    Y . Zhang, Z. Yu, J. Zhang, L. Wang, T. H. Luan, B. Guo, and C. Yuen, “Learning decentralized traffic signal controllers with multi-agent graph reinforcement learning,” IEEE Trans. Mobile Comput. , vol. 23, no. 6, pp. 7180–7195, Jun. 2024

  226. [236]

    Uplink performance of cell-free massive MIMO with multi-antenna users over jointly-correlated rayleigh fading channels,

    Z. Wang, J. Zhang, B. Ai, C. Yuen, and M. Debbah, “Uplink performance of cell-free massive MIMO with multi-antenna users over jointly-correlated rayleigh fading channels,” IEEE Trans. Wireless Commun., vol. 21, no. 9, pp. 7391–7406, Sep. 2022

  227. [237]

    Distributed collaborative user positioning for cell-free massive MIMO with multi-agent reinforcement learning,

    Z. Liu, J. Zhang, E. Shi, Y . Zhu, D. W. Kwan Ng, and B. Ai, “Distributed collaborative user positioning for cell-free massive MIMO with multi-agent reinforcement learning,” in IEEE SPAWC, 2024, pp. 466–470

  228. [238]

    Access point clustering in cell-free massive MIMO using conventional and federated multi-agent reinforcement learning,

    B. Banerjee, R. C. Elliott, W. A. Krzymie ˜n, and M. Medra, “Access point clustering in cell-free massive MIMO using conventional and federated multi-agent reinforcement learning,” IEEE Trans. Mach. Learn. Commun. Netw. , vol. 1, pp. 107–123, jun. 2023

  229. [239]

    Deep reinforce- ment learning for energy efficiency maximization in cache-enabled cell- free massive MIMO networks: Single- and multi-agent approaches,

    Y .-C. Chuang, W.-Y . Chiu, R. Y . Chang, and Y .-C. Lai, “Deep reinforce- ment learning for energy efficiency maximization in cache-enabled cell- free massive MIMO networks: Single- and multi-agent approaches,” IEEE Trans. Veh. Technol. , vol. 72, no. 8, pp. 10 826–10 839, Aug. 2023

  230. [240]

    Scalable multi-agent reinforcement learning for dynamic coordinated multipoint clustering,

    F. Hu, Y . Deng, and A. Hamid Aghvami, “Scalable multi-agent reinforcement learning for dynamic coordinated multipoint clustering,” IEEE Trans. Commun. , vol. 71, no. 1, pp. 101–114, Jan. 2023

  231. [241]

    Antenna se- lection of cell-free XL-MIMO systems with multi-agent reinforcement learning,

    Z. Liu, Z. Liu, X. Wang, J. Li, J. Zhang, and B. Ai, “Antenna se- lection of cell-free XL-MIMO systems with multi-agent reinforcement learning,” in IEEE Ucom, 2023, pp. 379–383

  232. [242]

    Joint precoding and AP selection for energy efficient RIS-aided cell-free massive MIMO using multi-agent reinforcement learning,

    E. Shi, J. Zhang, Z. Liu, Y . Zhu, C. Yuen, D. W. K. Ng, M. Di Renzo, and B. Ai, “Joint precoding and AP selection for energy efficient RIS-aided cell-free massive MIMO using multi-agent reinforcement learning,” arXiv preprint arXiv:2411.11070 , 2024

  233. [243]

    Multi-agent reinforcement learning for distributed resource allocation in cell-free massive MIMO- enabled mobile edge computing network,

    F. D. Tilahun, A. T. Abebe, and C. G. Kang, “Multi-agent reinforcement learning for distributed resource allocation in cell-free massive MIMO- enabled mobile edge computing network,” IEEE Trans. Veh. Technol., vol. 72, no. 12, pp. 16 454–16 468, Dec. 2023

  234. [244]

    Uplink power control for extremely large-scale MIMO with multi- agent reinforcement learning and fuzzy logic,

    Z. Liu, Z. Liu, J. Zhang, H. Xiao, B. Ai, and D. W. K. Ng, “Uplink power control for extremely large-scale MIMO with multi- agent reinforcement learning and fuzzy logic,” in IEEE INFOCOM Workshops, 2023, pp. 1–6

  235. [245]

    Double-layer power control for mobile cell-free XL-MIMO with multi-agent reinforcement learning,

    Z. Liu, J. Zhang, Z. Liu, H. Xiao, and B. Ai, “Double-layer power control for mobile cell-free XL-MIMO with multi-agent reinforcement learning,” IEEE Trans. Wireless Commun. , vol. 23, no. 5, pp. 4658– 4674, May 2024

  236. [246]

    Secure transmission scheme based on fingerprint positioning in cell- free massive MIMO systems,

    J. Qiu, K. Xu, X. Xia, Z. Shen, W. Xie, D. Zhang, and M. Wang, “Secure transmission scheme based on fingerprint positioning in cell- free massive MIMO systems,” IEEE Trans. Signal Inf. Process. Netw. , vol. 8, pp. 92–105, Feb. 2022

  237. [247]

    Fingerprint-based localization and channel estimation integration for cell-free massive MIMO IoT systems,

    C. Wei, K. Xu, Z. Shen, X. Xia, C. Li, W. Xie, D. Zhang, and H. Liang, “Fingerprint-based localization and channel estimation integration for cell-free massive MIMO IoT systems,” IEEE Internet of Things Jour- nal, vol. 9, no. 24, pp. 25 237–25 252, Dec. 2022

  238. [248]

    Reconfigurable intelligent surface-based index modulation: A new beyond MIMO paradigm for 6G,

    E. Basar, “Reconfigurable intelligent surface-based index modulation: A new beyond MIMO paradigm for 6G,” IEEE Trans. Commun. , vol. 68, no. 5, pp. 3187–3196, May 2020

  239. [249]

    Collaborative intelligent reflecting surface networks with multi-agent reinforcement learning,

    J. Zhang, J. Li, Y . Zhang, Q. Wu, X. Wu, F. Shu, S. Jin, and W. Chen, “Collaborative intelligent reflecting surface networks with multi-agent reinforcement learning,” IEEE J. Sel. Top. Signal Process. , vol. 16, no. 3, pp. 532–545, Apr. 2022

  240. [250]

    Reconfigurable intelligent surface aided vehicular edge computing: Joint phase-shift optimization and multi-user power allocation,

    K. Qi, Q. Wu, P. Fan, N. Cheng, W. Chen, and K. B. Letaief, “Reconfigurable intelligent surface aided vehicular edge computing: Joint phase-shift optimization and multi-user power allocation,” IEEE Internet Things J. , vol. 12, no. 1, pp. 764–777, Jan. 2025

  241. [251]

    Big AI models for 6G wireless networks: Opportunities, challenges, and research directions,

    Z. Chen, Z. Zhang, and Z. Yang, “Big AI models for 6G wireless networks: Opportunities, challenges, and research directions,” IEEE Wireless Commun., vol. 31, no. 5, pp. 164–172, Oct. 2024

  242. [252]

    AI models for green communications towards 6G,

    B. Mao, F. Tang, Y . Kawamoto, and N. Kato, “AI models for green communications towards 6G,” IEEE Commun. Surveys Tuts. , vol. 24, no. 1, pp. 210–247, Firstquarter 2022

  243. [253]

    Nonparametric teaching of implicit neural representations,

    C. Zhang, S. T. S. Luo, J. C. L. Li, Y .-C. Wu, and N. Wong, “Nonparametric teaching of implicit neural representations,” Proc. Int. Conf. Mach. Learn. , 2024

  244. [254]

    Implementation of big AI models for wireless networks with collaborative edge computing,

    L. Zeng, S. Ye, X. Chen, and Y . Yang, “Implementation of big AI models for wireless networks with collaborative edge computing,” IEEE Wireless Commun., vol. 31, no. 3, pp. 50–58, Jun. 2024

  245. [255]

    Large generative AI models for telecom: The next big thing?

    L. Bariah, Q. Zhao, H. Zou, Y . Tian, F. Bader, and M. Debbah, “Large generative AI models for telecom: The next big thing?” IEEE Commun. Mag., vol. 62, no. 11, pp. 84–90, Nov. 2024

  246. [256]

    Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services,

    M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, V . C. M. Leung, and H. V . Poor, “Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services,” IEEE Commun. Surveys Tuts. , vol. 26, no. 2, pp. 1...

  247. [257]

    Towards continual reinforcement learning: A review and perspectives,

    K. Khetarpal, M. Riemer, I. Rish, and D. Precup, “Towards continual reinforcement learning: A review and perspectives,”J. Artif. Intell. Res., vol. 75, pp. 1401–1476, Dec. 2022

  248. [258]

    Loss of plasticity in continual deep reinforcement learning,

    Z. Abbas, R. Zhao, J. Modayil, A. White, and M. C. Machado, “Loss of plasticity in continual deep reinforcement learning,” in Conference on Lifelong Learning Agents (CoLLAs) , 2023, pp. 620–636

  249. [259]

    Continual learning through synaptic intelligence,

    F. Zenke, B. Poole, and S. Ganguli, “Continual learning through synaptic intelligence,” in Proc. Int. Conf. Mach. Learn. , 2017

  250. [260]

    Loss of plasticity in deep continual learning,

    S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mah- mood, and R. S. Sutton, “Loss of plasticity in deep continual learning,” Nature, vol. 632, no. 8026, pp. 768–774, Aug. 2024

  251. [261]

    Semantic communications for future internet: Fundamentals, applications, and challenges,

    W. Yang, H. Du, Z. Q. Liew, W. Y . B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Shen, and C. Miao, “Semantic communications for future internet: Fundamentals, applications, and challenges,” IEEE Commun. Surv. Tuts., vol. 25, no. 1, pp. 213–250, Firstquarter 2023

  252. [262]

    Cell-free massive MIMO for URLLC: A finite-blocklength analysis,

    A. Lancho, G. Durisi, and L. Sanguinetti, “Cell-free massive MIMO for URLLC: A finite-blocklength analysis,” IEEE Trans. Wireless Commun., vol. 22, no. 12, pp. 8723–8735, Dec. 2023

  253. [263]

    Service multiplexing and revenue maximization in sliced C-RAN incorporated with URLLC and multicast eMBB,

    J. Tang, B. Shim, and T. Q. S. Quek, “Service multiplexing and revenue maximization in sliced C-RAN incorporated with URLLC and multicast eMBB,” IEEE J. Sel. Areas Commun., vol. 37, no. 4, pp. 881– 895, Apr. 2019

  254. [264]

    Mobile cell-free massive MIMO: Challenges, solutions, and future directions,

    J. Zheng, J. Zhang, H. Du, D. Niyato, B. Ai, M. Debbah, and K. B. Letaief, “Mobile cell-free massive MIMO: Challenges, solutions, and future directions,” IEEE Wireless Commun. , vol. 31, no. 3, pp. 140– 147, Jun. 2024

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.