REVIEW 3 major objections 4 minor 1 cited by
Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A sliding-window UCB bandit that picks each agent's sight range episode by episode can replace manual range selection in cooperative MARL, lifting performance in LBF, RWARE, and SMAC.
desk verdict A practical drop-in sight-range selector with broad experiments, but the 'consistent improvement' and 'optimal range' claims overreach given the confounded UCB rewards and several counter rows in Table 1. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the sliding-window UCB meta-controller of equation (3): each candidate sight range $d_i$ is an arm, the reward is the episode return collected while that range was active, and the controller scores arms by the windowed mean return plus an exploration bonus $c\sqrt{\log\min(e,w)/N_e(d_i,w)}$, with window size $w=5000$ episodes and exploration constant $c=2$. It sits on top of a modified observation function $Z(s,n_i,d)$ that masks the observation to the selected range, cropping grid views in LBF and RWARE and limiting the visibility radius in SMAC, with out-of-range entries set to a default or zero. The meta-controller operates at episode granularity while the underlying MARL algorithm trains on a single replay buffer fed by all ranges, and the paper argues this lets the controller steadily concentrate episodes on whatever range yields the best recent returns.
What would settle it
Take an environment where exhaustive fixed-range training has already identified the best sight range, run DSR with matched total compute, and compare the range DSR converges to against that best range; if the converged range is not the empirically best one, the UCB is selecting the easiest range during policy learning rather than the best range. A sharper version freezes a fully trained policy and runs the meta-controller without further policy updates: DSR's design gives no signal in that regime, which would show that it tracks learning progress under each range, not the range's intrinsic value.
Extended reading notes
Core claim
The central discovery claimed is that the sight range dilemma can be handled by selection alone: keep the MARL algorithm untouched, modify only the observation function $Z(s,n_i,d)$ to crop each agent's local view to the chosen range $d$, and let a meta-controller decide $d$ at the start of each episode from the recent episode returns of each candidate range. The meta-controller uses the sliding-window UCB rule (equation 3), with a window of 5000 episodes and an exploration constant $c=2$, treating each range as a bandit arm. Across their experiments the authors report that DSR matches or beats the fixed-range baseline in nearly every setting, with the largest improvements exactly where large fixed ranges are known to hurt: for example, LBF 10x10-4p-4f-coop-10s improves from 0.338 to 0.798 and SMAC MMM2-21s from 0.190 to 0.714. They further report that the selected range typically starts small and grows during training, and that a hand-designed fixed expansion schedule underperforms DSR, which they read as evidence that dynamic selection both accelerates learning and reveals how much information the task actually needs.
Load-bearing premise
The load-bearing premise is that the recent episode returns collected under each sight range honestly measure that range's quality, even though those returns come from a single policy that is still learning and is shared across all ranges through one replay buffer.
Editorial extensions
If this is right
- Practitioners no longer need to sweep sight ranges by hand: DSR wraps an existing MARL algorithm and converges to a range on its own, so the observation width becomes an output of training rather than an input.
- Training is accelerated because the controller tends to start agents on small, simple observations and widen them as the policy matures; the reported curves rise faster than fixed-range baselines while reaching equal or better final scores.
- The converged range is an interpretable design signal: it states how much of the environment the agents actually rely on, which the authors propose as guidance for sensor design in applications such as autonomous driving.
- Because DSR relies only on per-agent observation cropping, it applies where no global state or communication channel exists, which is the regime the authors argue prior communication-based solutions cannot serve.
Reading between the lines
- My read is that a large share of the reported gain may come from a curriculum effect rather than from discovering the truly optimal range: small ranges pay off early in training, so the UCB naturally lingers there before widening, and a well-timed manual schedule might capture part of the same benefit.
- A testable extension the paper invites but does not pursue: feed the selected range $d$ to the policy as an explicit conditioning input, so the network can specialize per range instead of inferring the range from observation statistics; this would also make the UCB's reward signal cleaner.
- The authors explicitly leave per-agent, heterogeneous ranges and continuous range spaces to future work, so the current claim is bounded to one shared, discrete range per environment; whether the selection signal survives finer granularity is open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Dynamic Sight Range Selection (DSR), a meta-controller that uses a sliding-window UCB algorithm to select, at the beginning of each episode, one sight range from a hand-chosen set D. The selected range is applied to all agents' observations, while the underlying MARL algorithm (QMIX, IQL, VDN, IPPO, MAPPO, IM-Qatten) trains as usual on the resulting episodes. The authors evaluate DSR on LBF, RWARE, SMAC, and SMAC-DT, reporting that DSR outperforms fixed-range baselines, accelerates training, and 'automatically discovers the optimal sight range.' The paper includes an extensive appendix with fixed-range comparisons, hyperparameter sensitivity for c and w, and a comparison against the communication-based method CAMA.
Significance. If the claims hold, DSR is an attractive drop-in wrapper: it removes the need to manually select an observation range, requires no global information or communication, and could guide sensor design in real-world applications. The experimental breadth is a genuine strength: three environments, five MARL algorithms, many maps, fixed-range baselines for every setting, five seeds for the main QMIX experiments, and an integration with the CAMA codebase. However, the central claims of 'consistently improves' and 'discovers the optimal sight range' need additional support: several Table 1 rows contradict the consistency claim, no significance testing is reported, and the UCB meta-controller's reward signal is confounded with the shared policy's ongoing learning. The paper's practical value is real but currently overstated.
major comments (3)
- [Section 4.2, Table 1] The abstract and §4.2 claim that DSR 'consistently improves performance across three common MARL environments,' but Table 1 contains multiple rows where the DSR mean is below the baseline mean: 10x10-4p-2f-coop-6s (0.957 vs 0.972), 10x10-4p-2f-6s (0.987 vs 0.998), small-2ag-5s (0.050 vs 0.074), small-2ag-3s (0.036 vs 0.182), 3s_vs_5z-9s (0.676 vs 0.716), 3s5z-15s (0.770 vs 0.808), and 3s5z-21s (0.736 vs 0.784). Additionally, the standard deviations are large relative to the differences (e.g., tiny-2ag-5s: 4.762 ± 4.702 vs 1.486 ± 1.361), and no significance tests or confidence intervals are provided. To support the consistency claim, the authors should either weaken the wording to 'often improves' or provide paired significance tests (e.g., bootstrap or Wilcoxon across the five seeds) and report how many of the 22 settings show a statistically significant improvement.
- [Section 3.2, Eq. (3) and Algorithm 1] The meta-controller's UCB scores are computed from episode returns that are generated by a single policy shared across all sight ranges and trained on one replay buffer containing data from all ranges. Consequently, the return r_j(d_i) in Eq. (3) is not a sample of a stationary arm quality; it is a function of the current policy, which has itself been shaped by previous sight-range selections and by data collected under other ranges. The paper therefore does not establish that DSR 'automatically discovers the optimal sight range,' because 'optimal' is defined operationally as the argmax of recent episode returns within a hand-chosen set D, with no evaluation of each candidate range under a fixed, converged policy. This confounding is a load-bearing issue: the observed benefits could arise from training on a mixture of sight ranges rather than from UCB's selection mechanism. I recommend adding an ablation that replaces the UCB meta-controller with random selection over the same set D, and an additional evaluation where the final policy is tested under each fixed sight range to check whether the range selected by DSR is actually the best at convergence.
- [Section 4.1 and Appendix B.1] The SMAC experiments use a non-standard state construction: the global state is formed by concatenating all agents' observations, rather than using the environment-provided global state. This choice likely accentuates the sight range dilemma and may make the SMAC results not directly comparable to standard SMAC benchmarks. The appendix reports 'w/ Given State' comparisons (Figures 10 and 26–31), but the main text and abstract do not qualify the SMAC claim accordingly. The authors should clearly state in §4.1 and the abstract that the SMAC results use this observation-concatenation variant, or present the standard-state results as the primary SMAC evidence.
minor comments (4)
- [Algorithm 1, line 8] Line 8 contains the condition 'if d_{e-t} ≠ d*_e'; the index 'e-t' appears to be a typo (likely 'e-1' or a comparison of the previous selection). The intended semantics for updating N_e(d_i,w) are unclear, especially in relation to the definition in Eq. (3). Please clarify.
- [Table 2] The row for '10m_vs_11m' lists '3 Stalkers + 5 Zealots' for both sides, which appears to be a copy-paste error; 10m_vs_11m in SMAC consists of Marines. Please correct the table.
- [Section 4.6 and Appendix B.2] The SMAC-DT comparison with CAMA reports no standard deviations, seeds, or number of runs for Figures 9, 36, and 37. Given that the main experiments use five seeds, the dynamic team composition results should report the same level of uncertainty to support the claim that DSR 'outperforms' CAMA.
- [Section 4.2, Figure 4] The training acceleration claim is supported only by visual inspection of two LBF curves. Please provide a quantitative summary, such as the number of steps to reach a given return threshold or the area under the training curve, for all settings where final performance is comparable.
Circularity Check
'Optimal sight range' is defined as the windowed-return argmax, so its 'discovery' is the selection rule; empirical gains vs fixed ranges remain independent.
-
self definitional
[Section 3.2 (Meta-Controller) and Abstract; see also Algorithm 1, line 4]
"After training, the meta-controller converges to an optimal sight range. During execution, we simply choose the sight range d with the maximum average return in the window, arg max(r(d_i)), where r(d_i) is the average return for each sight range d_i. ... DSR provides additional interpretability by indicating the optimal sight range used during training."
The paper never defines 'optimal sight range' by an external criterion; the only definition supplied is the meta-controller's windowed average-return estimator. DSR's output is exactly that argmax: Algorithm 1 line 4 selects d*_e = argmax(Xhat_e(d_i) + c*U_e(d_i)), and the execution rule in Section 3.2 is the same argmax of recent returns. Thus the claim that DSR 'automatically discovers the optimal sight range' or 'indicates the optimal sight range' is a restatement of the selection rule, not a prediction from independent evidence. The non-circular part is the empirical comparison against fixed-sight-range baselines, which is genuine external evidence and keeps the score below 6.
full rationale
The only clearly circular step is the 'optimal sight range' claim. The meta-controller's policy is to select the sight range with the largest windowed recent return (Eq. 3, Algorithm 1), and the paper's closing claim that DSR 'automatically identifies the optimal sight range' is definitionally the same as running that selection rule. No separate ground-truth 'optimal' range is measured or predicted. This is a self-definitional reduction of the discovery claim. However, the paper's main empirical contribution is not circular: DSR is compared against fixed-sight-range baselines on LBF, RWARE, and SMAC (Table 1, Figures 4-9), and those comparisons are external to the meta-controller's own objective. The fixed-schedule ablation (Figure 6) and the CAMA comparison (Figure 9) also provide independent evidence that the mechanism has value beyond its own definition. There is no load-bearing self-citation: the UCB machinery cites external work [3, 8], and no uniqueness theorem or prior result by the same authors is invoked to force the approach. The UCB confounding with the shared replay buffer is a validity risk rather than a circularity: it concerns whether the selected range reflects true quality or policy-learning feedback, but it is not a case where the paper's output is equivalent to its input by construction. Overall, the circularity is partial and localized to the 'optimal discovery' framing; the headline performance improvements stand independently, so a moderate score of 4 is appropriate.
Assumptions & free parameters
free parameters (4)
- UCB exploration coefficient c =
2.0
- Sliding window size w =
5000
- Candidate sight range set D =
Varies per environment, e.g., {2,4,6} or {2,4,6,8,10} in LBF
- SMAC return normalization divisor =
20
assumptions (4)
- domain assumption Episode return obtained under a concurrently learning shared policy is a valid reward for comparing sight ranges.
- domain assumption SW-UCB non-stationary bandit guarantees transfer to the stratified MARL setting.
- domain assumption Concatenated agent observations form a valid global state proxy in SMAC.
- standard math Standard UCB regret bounds hold for stationary bandits.
Cite this review
Pith. "Pith review of Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning." pith.science (2026). https://pith.science/paper/EQNVNSI5
@misc{pith2026250512811,
author = {Pith},
title = {Pith review of: Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/EQNVNSI5}},
note = {Machine review of arXiv:2505.12811}
}
read the original abstract
Multi-agent reinforcement Learning (MARL) is often challenged by the sight range dilemma, where agents either receive insufficient or excessive information from their environment. In this paper, we propose a novel method, called Dynamic Sight Range Selection (DSR), to address this issue. DSR utilizes an Upper Confidence Bound (UCB) algorithm and dynamically adjusts the sight range during training. Experiment results show several advantages of using DSR. First, we demonstrate using DSR achieves better performance in three common MARL environments, including Level-Based Foraging (LBF), Multi-Robot Warehouse (RWARE), and StarCraft Multi-Agent Challenge (SMAC). Second, our results show that DSR consistently improves performance across multiple MARL algorithms, including QMIX and MAPPO. Third, DSR offers suitable sight ranges for different training steps, thereby accelerating the training process. Finally, DSR provides additional interpretability by indicating the optimal sight range used during training. Unlike existing methods that rely on global information or communication mechanisms, our approach operates solely based on the individual sight ranges of agents. This approach offers a practical and efficient solution to the sight range dilemma, making it broadly applicable to real-world complex environments.
Figures
Figures from the paper (33 more)
Forward citations
Cited by 1 Pith paper
-
PLATO: Pointer Learner for Agent and Task Openness
A pointer-network actor plus GNN critic jointly handles agent and task openness in MARL without fixed bounds, with proofs of well-definedness and strong wildfire results.
Reference graph
Works this paper leans on
-
[1]
Mehdi Afsar, Trafford Crump, and Behrouz Far
M. Mehdi Afsar, Trafford Crump, and Behrouz Far. 2022. Reinforcement Learning Based Recommender Systems: A Survey. ACM Comput. Surv. 55, 7 (2022), 145:1– 145:38. https://doi.org/10.1145/3543846
doi:10.1145/3543846 2022
-
[2]
Albrecht and Subramanian Ramamoorthy
Stefano V. Albrecht and Subramanian Ramamoorthy. 2013. A Game-Theoretic Model and Best-Response Learning Method for Ad Hoc Coordination in Multia- gent Systems. In Proceedings of the 2013 International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS ’13) . International Foundation for Au- tonomous Agents and Multiagent Systems, Richland, SC...
work page 2013
-
[3]
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer. 2002. Finite-Time Analysis of the Multiarmed Bandit Problem. Machine Learning 47, 2 (May 2002), 235–256. https://doi.org/10.1023/A:1013689704352
-
[4]
Bernstein, Shlomo Zilberstein, and Neil Immerman
Daniel S. Bernstein, Shlomo Zilberstein, and Neil Immerman. 2000. The Com- plexity of Decentralized Control of Markov Decision Processes. In Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence (UAI’00) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 32–37
work page 2000
-
[5]
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. https://doi.org/10.48550/arXiv.1412.3555 arXiv:1412.3555
-
[6]
Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviy- chuk, Philip H. S. Torr, Mingfei Sun, and Shimon Whiteson. 2020. Is In- dependent Learning All You Need in the StarCraft Multi-Agent Challenge? https://doi.org/10.48550/arXiv.2011.09533 arXiv:2011.09533 [cs]
-
[7]
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018. Counterfactual Multi-Agent Policy Gradients. Pro- ceedings of the AAAI Conference on Artificial Intelligence 32, 1 (April 2018). https://doi.org/10.1609/aaai.v32i1.11794
-
[8]
Aurélien Garivier and Eric Moulines. 2008. On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems. https://doi.org/10.48550/arXiv.0805.3415 arXiv:0805.3415 [math, stat]
Show all 42 references
-
[9]
Cong Guan, Feng Chen, Lei Yuan, Chenghe Wang, Hao Yin, Zongzhang Zhang, and Yang Yu. 2022. Efficient Multi-agent Communication via Self-supervised Information Aggregation. Advances in Neural Information Processing Systems 35 (Dec. 2022), 1020–1033. https://proceedings.neurips....
2022
-
[10]
Siyi Hu, Yifan Zhong, Minquan Gao, Weixun Wang, Hao Dong, Xiaodan Liang, Zhihui Li, Xiaojun Chang, and Yaodong Yang. 2023. MARLlib: A Scalable and Ef- ficient Multi-agent Reinforcement Learning Library. Journal of Machine Learning Research 24, 315 (2023), 1–23. http://jmlr.org...
2023
-
[11]
Schroeder De Witt, Bei Peng, Wendelin Boehmer, Shimon Whiteson, and Fei Sha
Shariq Iqbal, Christian A. Schroeder De Witt, Bei Peng, Wendelin Boehmer, Shimon Whiteson, and Fei Sha. 2021. Randomized Entity-wise Factorization for Multi-Agent Reinforcement Learning. In Proceedings of the 38th International Conference on Machine Learning . PMLR, 4596–4606....
2021
-
[12]
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. 2022. Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning. In International Conference on Learning Representations . https://openreview.net/forum?id=EcGGFkNTxdJ
2022
-
[13]
Ryan Lowe, YI WU, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc. https://proceedings.neurips....
2017
-
[14]
Le, James Laudon, Richard Ho, Roger Carpenter, and Jeff Dean
Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Wenjie Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Azade Nova, Jiwoo Pak, Andy Tong, Kavya Srinivasa, William Hang, Emre Tuncer, Quoc V. Le, James Laudon, Richard Ho, Roger Carpenter, and J...
2021 doi
-
[15]
Zepeng Ning and Lihua Xie. 2024. A Survey on Multi-Agent Reinforcement Learning and Its Application. Journal of Automation and Intelligence 3, 2 (June 2024), 73–91. https://doi.org/10.1016/j.jai.2024.02.003
2024 doi
-
[16]
Mohammad Noaeen, Atharva Naik, Liana Goodman, Jared Crebo, Taimoor Abrar, Zahra Shakeri Hossein Abad, Ana L. C. Bazzan, and Behrouz Far. 2022. Re- inforcement Learning in Urban Network Traffic Signal Control: A Systematic Literature Review. Expert Systems with Applications 199...
2022
-
[17]
Oliehoek and Christopher Amato
Frans A. Oliehoek and Christopher Amato. 2016. A Concise Introduction to Decentralized POMDPs. Springer International Publishing, Cham. https://doi. org/10.1007/978-3-319-28929-8
2016 doi
-
[18]
Albrecht
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht
-
[19]
Efros, and Trevor Darrell
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. 2017. Curiosity-Driven Exploration by Self-supervised Prediction. In Proceedings of the 34th International Conference on Machine Learning . PMLR, 2778–2787. https://proceedings.mlr.press/v70/pathak17a.html
2017
-
[20]
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic Value Func- tion Factorisation for Deep Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learni...
2018
-
[21]
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Far- quhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson. 2019. The StarCraft Multi-Agent Challenge. In Proceedings of the 18th International Conf...
2019
-
[22]
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver. 2020. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model....
2020 doi
-
[23]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
-
[24]
Jianzhun Shao, Hongchang Zhang, Yun Qu, Chang Liu, Shuncheng He, Yuhang Jiang, and Xiangyang Ji. 2023. Complementary Attention for Multi-Agent Rein- forcement Learning. InProceedings of the 40th International Conference on Machine Learning. PMLR, 30776–30793. https://proceedin...
2023
-
[25]
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019. QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning . PMLR, 5887–5896. htt...
2019
-
[26]
Sainbayar Sukhbaatar, arthur szlam, and Rob Fergus. 2016. Learning Multiagent Communication with Backpropagation. In Advances in Neural Information Pro- cessing Systems, Vol. 29. Curran Associates, Inc. https://proceedings.neurips.cc/ paper/2016/hash/55b1927fdafef39c48e5b73b5d...
2016
-
[27]
Leibo, Karl Tuyls, and Thore Graepel
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vini- cius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2018. Value-Decomposition Networks For Cooper- ative Multi-Agent Learning Based On Team Rewa...
2018
-
[28]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Intro- duction. A Bradford Book, Cambridge, MA, USA
2018
-
[29]
J Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis S Santos, Clemens Dieffendahl, Caroline Horsch, Rodrigo Perez-Vicente, et al . 2021. Pettingzoo: Gym for Multi-Agent Reinforcement Learning. Advances in Neural Information Processing ...
2021
-
[30]
Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max...
2019
-
[31]
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang. 2021. QPLEX: Duplex Dueling Multi-Agent Q-Learning. In International Conference on Learning Representations. https://openreview.net/forum?id=Rcmk0xxIQV
2021
-
[32]
Tonghan Wang, Heng Dong, Victor Lesser, and Chongjie Zhang. 2020. ROMA: Multi-Agent Reinforcement Learning with Emergent Roles. In Proceedings of the 37th International Conference on Machine Learning . PMLR, 9876–9886. https: //proceedings.mlr.press/v119/wang20f.html
2020
-
[33]
Yuchen Xiao, Weihao Tan, and Christopher Amato. 2022. Asyn- chronous Actor-Critic for Multi-Agent Reinforcement Learning. Ad- vances in Neural Information Processing Systems 35 (Dec. 2022), 4385–
2022
- [34]
-
[35]
Yaodong Yang and Jun Wang. 2020. An Overview of Multi-Agent Reinforcement Learning from Game Theoretical Perspective. https://arxiv.org/abs/2011.00583v3
2020 arXiv
-
[36]
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The Surprising Effectiveness of PPO in Cooperative Multi- Agent Games. Advances in Neural Information Processing Systems 35 (Dec. 2022), 24611–24624. https://proceedings.neurips.cc/pa...
2022
-
[37]
Peihong Yu, Bhoram Lee, Aswin Raghavan, Supun Samarasekera, Pratap Tokekar, and James Zachary Hare. 2023. Enhancing Multi-Agent Coordination through Common Operating Picture Integration. In First Workshop on Out-of-Distribution Generalization in Robotics at CoRL 2023 . https:/...
2023
- [38]
-
[39]
NxN, ” indicating an N by N grid. For example, 8x8 refers to a grid of size 8x8. The number of agents (p) is indicated after the letter “p,
Haiyan Zhao, Chengcheng Dong, Jian Cao, and Qingkui Chen. 2024. A Survey on Deep Reinforcement Learning Approaches for Traffic Signal Control.Engineering Applications of Artificial Intelligence 133 (July 2024), 108100. https://doi.org/10. 1016/j.engappai.2024.108100 A ENVIRONM...
2024
-
[2017]
https://arxiv.org/abs/1707
Proximal Policy Optimization Algorithms. https://arxiv.org/abs/1707. 06347v2
-
[2021]
In Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)
Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks. In Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1). https://openreview.net/forum? id=cIrPX-Sn5n
-
[4400]
https://proceedings.neurips.cc/paper_files/paper/2022/hash/ 1c153788756d35559c22d105d1182c30-Abstract-Conference.html
2022
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.