Pith. sign in

REVIEW 3 major objections 6 minor 149 references

Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning

T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper argues that post hoc interpretability tools—relevance backpropagation, sparse autoencoders, activation steering, and circuit discovery—are a scalable alternative to designing intrinsically interpretable multi-agent…

desk verdict A clear, honest perspective piece that maps direct interpretability tools to MADRL challenges, but its central transfer assumption is asserted rather than argued, and the paper itself concedes the main weakness. read the letter →

arxiv 2502.00726 v1 pith:5EBVCYWJ submitted 2025-02-02 cs.AI

classification cs.AI
keywords multi-agentdeepreinforcementlearninginterpretabilityposthocexplanationsexplainablemechanisticsparseautoencodersrelevancebackpropagationrepresentationengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multi-agent deep reinforcement learning trains powerful but opaque policies, and the usual remedy—building interpretable architectures from the start—trades away performance and struggles to scale. This paper argues for a different route: direct interpretability, meaning post hoc methods that extract explanations from an already-trained network without changing its design. It surveys a modern toolkit (relevance backpropagation, sparse autoencoders, activation patching and steering, circuit discovery) and maps each tool to concrete MADRL problems: finding biases, decomposing decisions, identifying team roles, monitoring communication, coordinating swarms, and making training more sample-efficient. The payoff, if the transfer works, is that transparency in multi-agent systems would no longer require sacrificing performance or redesigning architectures.

What carries the argument

The central object is the trained deep network of an agent, treated as a directly manipulable artifact rather than a black box to be replaced. The mechanisms carried by the argument are the interpretability tools themselves: layer-wise relevance propagation for attributing decisions to inputs, sparse autoencoders for eliciting interpretable features (prototypes), activation steering and causal tracing for editing or controlling internal representations, and automated circuit discovery for isolating functional sub-networks. The paper's contribution is to connect these mechanisms to specific MADRL challenges—for instance, partitioning the positive weights of a QMIX mixing network with non-negative matrix factorization to identify teams, or using steering vectors to shift swarm behaviour toward cooperation.

What would settle it

A controlled study in a cooperative multi-agent environment with a known ground-truth attribution (for instance, a reward decomposition): if relevance maps or activation-patching attributions from a trained policy do not predict the effect of intervening on the highlighted components—ablating the 'important' neurons should change behaviour more than ablating random ones—the transfer claim would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that direct interpretability is a versatile and scalable alternative to intrinsically interpretable models for MADRL. It posits that because trained agents are deep networks, the same post hoc methods that have opened up vision and language models—attributing decisions to inputs, locating and editing internal knowledge, steering activations, and discovering computational circuits—can be applied to multi-agent policies to answer questions that matter in the field: what spurious cues does an agent rely on, which agents share a role, what is each agent contributing, and what are agents communicating. The paper organizes these questions into a taxonomy spanning single-agent, multi-agent, and training-process levels, and argues that this route is more flexible than intrinsic interpretability, particularly for large or already-trained systems. It closes by noting that the approach will only be trustworthy once evaluation protocols exist to check post hoc explanations, which currently can produce metrics with limited predictive power.

Load-bearing premise

The load-bearing premise is that post hoc interpretability methods, developed and tested mostly on vision and language models, will transfer to multi-agent deep reinforcement learning and yield explanations that faithfully reflect what the trained policies actually compute.

Editorial extensions

If this is right

  • Direct interpretability could let developers remove biases or dangerous behaviors from a trained policy without retraining, by editing internal representations.
  • Interpretability of latent spaces could automate team identification in large agent populations, reducing the number of policies needing training.
  • Activation steering could make swarms coordinate by shifting agents toward cooperative goals without architecture changes.
  • Interpretability of the training process could improve sample efficiency, for example by prioritizing samples based on learned importance.
  • Reliable evaluation protocols for post hoc explanations would need to be built before these methods are trusted in practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If direct interpretability works for MADRL, it could be combined with intrinsic methods—for example, using circuit discovery to verify the mechanisms inside a concept-bottleneck agent—blurring the line between the two paradigms.
  • The same toolkit could be applied to human-AI teams, where explanations of agent internals could help humans build accurate mental models of autonomous teammates.
  • The evaluation bottleneck could be attacked with interventions: instead of trusting saliency maps, one could ablate the 'important' features or neurons and measure whether behavior changes more than with random ablations.
  • A testable bet is that LRP-style relevance propagation can match Shapley-based credit assignment on small cooperative tasks while being far cheaper to compute.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This position paper argues that post hoc ``direct interpretability'' methods (feature attribution, prototypes/sparse autoencoders, latent manipulation, circuit analysis) should be adopted as a scalable and flexible alternative to intrinsically interpretable models for multi-agent deep reinforcement learning (MADRL). It proposes a taxonomy of MADRL challenges in three categories (single-agent, multi-agent, and training process), sketches concrete candidate applications such as bias identification via LRP, team identification via sparse autoencoders, and swarm coordination via activation steering, and discusses related work in post hoc interpretability for RL. The paper also acknowledges limitations of post hoc methods in Section 5.2, including saliency map unreliability and ``interpretability illusions,'' and calls for robust evaluation protocols.

Significance. The paper's value lies in its broad synthesis of modern interpretability tools and a clear organizational taxonomy that connects specific MADRL challenges to candidate post hoc methods. It is honest about known weaknesses of these methods and cites a wide body of recent literature, including work on chess agents and sparse autoencoders. If the underlying premise that CV/NLP interpretability tools transfer to multi-agent neural policies is correct, the proposed research directions could stimulate productive empirical work. However, the paper provides no evidence for transferability, and the central advocacy claim rests on an assumption that the authors themselves concede is fragile. The contribution is therefore a speculative research agenda rather than a demonstrated result; its significance will depend on whether future work can validate the proposed applications.

major comments (3)
  1. [§3.1–§3.3 vs §5.2] The concrete proposals in Section 3 (e.g., LRP for bias identification, SAE for team identification, activation steering for swarm coordination) require that post hoc methods yield causally faithful explanations in MADRL settings, yet Section 5.2 concedes that these methods ``often generate metrics with limited predictive power'' and are subject to ``interpretability illusions.'' This creates a direct tension: policy edition and activation steering are causal operations that require faithfulness, but the paper does not resolve how the conceded unreliability is compatible with those applications. For the central claim of the abstract to be load-bearing, the authors should either explicitly label each Section 3 proposal as an untested hypothesis or provide a concrete account of how the proposed evaluation protocols would validate or falsify each application. As written, the advocacy is internally inconsistent in its level of commitment.
  2. [§5.2] ``Robust Evaluation Protocols'' is too underspecified to support the advocacy. The section cites existing evaluation work (e.g., [7, 18, 36, 45, 51, 73, 129]) and states that reliable metrics must be established, but does not state what would count as evidence for transferability of the Section 3 applications. Because the absence of ground-truth explanations is acknowledged as a central difficulty, the paper needs at least one concrete evaluation idea (for instance, controlled interventions on policies, behavioral tests with perturbed observations, or ablation studies that compare explanation-predicted importance with actual policy changes). Without such a proposal, the call for direct interpretability is unfalsifiable, and the limitation statement in Section 5.2 does not resolve the burden placed on the earlier sections.
  3. [§4.3] The comparison with intrinsically interpretable models is one-sided. The section argues that intrinsic models face scalability and flexibility challenges, but it does not engage with the well-known argument that high-stakes applications may require intrinsically interpretable models precisely because post hoc explanations can be misleading (the paper cites [105] but does not respond to its central challenge). Since the paper's thesis is that direct interpretability is a viable alternative, Section 4.3 should at least acknowledge the conditions under which intrinsic interpretability remains preferable, and explain why the proposed direct methods overcome those conditions in MADRL. This would strengthen the advocacy rather than weaken it.
minor comments (6)
  1. [Abstract vs §6] The abstract states that direct interpretability ``offering insights into agents' behaviour, emergent phenomena, and biases,'' which is assertive, while the conclusion (Section 6) says ``direct interpretability might be vital.'' The framing should be aligned so that the abstract reflects the appropriately hedged nature of a position paper.
  2. [Figure 1] The caption refers to green, blue, and red colors for different challenge categories, but the figure itself is not included in the text and the color coding is not explained. If the figure is published, the caption should be self-contained.
  3. [§2.2] In the ``Prototypes'' paragraph, sparse autoencoders are mentioned in one sentence; given their prominent role in later sections (e.g., §3.2), a brief explanation of how SAEs elicit features would help readers unfamiliar with this method.
  4. [Figure 1] There is a typo in the figure: ``Swarn Coordination'' should be ``Swarm Coordination.''
  5. [§5.2] Reference [33] appears twice in the list ``[15, 33, 33]''; the duplicate should be removed.
  6. [§1] The phrase ``these approaches often need to be revised for large and performant systems'' is unclear; it likely means ``these approaches often need to be revisited'' or ``are often inadequate for.'' Please rephrase.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is a position paper with no derivations, fitted parameters, or empirical predictions, and its only self-citation is peripheral background.

full rationale

The paper is a vision/position piece advocating direct interpretability methods for MADRL. It contains no mathematical derivations, no fitted parameters, no experimental results, and no equation-level dependencies. The central claim is an advocacy claim: post hoc interpretability methods from CV/NLP and XRL could be usefully transferred to MADRL challenges. This is supported by citing prior work on such methods and by proposing speculative applications (e.g., LRP for bias identification, SAEs for team identification, activation steering for swarm coordination). None of these proposals is presented as a derived prediction from the paper's own inputs. The only self-citation is reference [92], Yoann Poupart's 'Contrastive Sparse Autoencoders for Interpreting Planning of Chess-Playing Agents,' cited in Section 4.1 as one of several efforts interpreting chess engines like AlphaZero. That citation is descriptive background, not a load-bearing premise: the paper's advocacy does not reduce to it, and the argument would stand equally without it. Section 5.2 even concedes limitations of post hoc methods (saliency map shortcomings, interpretability illusions, metrics with limited predictive power), which is the opposite of circularly assuming their validity. Therefore the circularity burden is zero and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no new free parameters or invented entities. Its central argument depends on two unproven domain assumptions: that methods imported from CV and NLP carry over to MADRL, and that explanations produced by these methods are meaningful despite known faithfulness problems. The authors themselves flag these concerns in Section 5.2.

assumptions (2)
  • domain assumption Post hoc interpretability methods developed for CV and NLP are applicable to MADRL policies and training processes.
    Section 3 maps LRP, SAEs, activation steering, and circuit discovery onto single-agent, multi-agent, and training challenges without empirical validation.
  • domain assumption Explanations from direct methods are faithful enough to support debugging, steering, and evaluation.
    Section 5.2 acknowledges interpretability illusions and the lack of ground truth, yet the central advocacy depends on the usefulness of these explanations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning." pith.science (2026). https://pith.science/paper/5EBVCYWJ

@misc{pith2026250200726,
  author       = {Pith},
  title        = {Pith review of: Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5EBVCYWJ}},
  note         = {Machine review of arXiv:2502.00726}
}
read the original abstract

Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While learning intrinsically interpretable models remains a prominent approach, its scalability and flexibility are limited in handling complex tasks or multi-agent dynamics. This paper advocates for direct interpretability, generating post hoc explanations directly from trained models, as a versatile and scalable alternative, offering insights into agents' behaviour, emergent phenomena, and biases without altering models' architectures. We explore modern methods, including relevance backpropagation, knowledge edition, model steering, activation patching, sparse autoencoders and circuit discovery, to highlight their applicability to single-agent, multi-agent, and training process challenges. By addressing MADRL interpretability, we propose directions aiming to advance active topics such as team identification, swarm coordination and sample efficiency.

Figures

Figures reproduced from arXiv: 2502.00726 by the authors.

Figure 1
Figure 1. Visual taxonomy of MADRL challenges that could [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Schema of a simplified view of MADRL systems. At [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

149 extracted references · 21 canonical work pages

  1. [105]

    Tabish Rashid, Mikayel Samvelyan, C. S. D. Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic Value Function Fac- torisation for Deep Multi-Agent Reinforcement Learning.ArXiv abs/1803.11485 (2018)

  2. [1]

    Reduan Achtibat, Maximilian Dreyer, Ilona Eisenbraun, Sebastian Bosse, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin. 2022. From attribution maps to human-understandable explanations through Concept Rel- evance Propagation. Nature Machine Intelligence 5 (2022), 1006 – 1019

  3. [2]

    Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. 2024. AttnLRP: Attention-Aware Layer-wise Relevance Propagation for Transformers. ArXiv abs/2402.05602 (2024)

  4. [3]

    Goodfellow, Moritz Hardt, and Been Kim

    Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim. 2018. Sanity Checks for Saliency Maps. In Neural Information Processing Systems

  5. [4]

    Guillaume Alain and Yoshua Bengio. 2018. Understanding intermediate layers using linear classifier probes. arXiv:1610.01644 [stat.ML]

  6. [5]

    Gülsüm Alicioğlu and Bo Sun. 2024. Use Bag-of-Patterns Approach to Explore Learned Behaviors of Reinforcement Learning. In xAI

  7. [6]

    Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, and François Fleuret. [n.d.]. Diffusion for World Modeling: Visual Details Matter in Atari. In Thirty-eighth Conference on Neural Information Processing Systems

  8. [7]

    José Pereira Amorim, Pedro Henriques Abreu, João A. M. Santos, and Henning Müller. 2023. Evaluating Post-hoc Interpretability with Intrinsic Interpretability. ArXiv abs/2305.03002 (2023)

Show all 149 references
  1. [8]

    Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015. On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation.PLoS ONE 10 (2015)

  2. [9]

    Osbert Bastani, Yewen Pu, and Armando Solar-Lezama. 2018. Verifiable Re- inforcement Learning via Policy Extraction. In Neural Information Processing Systems

  3. [10]

    Yanzhe Bekkemoen. 2023. Explainable reinforcement learning (XRL): a system- atic literature review and taxonomy. Machine Learning 113 (2023), 355–441

  4. [11]

    Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt

    Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor V. Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting Latent Predictions from Transformers with the Tuned Lens. ArXiv abs/2303.08112 (2023)

  5. [12]

    Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. LEACE: Perfect linear concept erasure in closed form. arXiv:2306.03819 [cs.LG]

  6. [13]

    David Bertoin, Adil Zouitine, Mehdi Zouitine, and Emmanuel Rachelson. 2022. Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement Learning. In Neural Information Processing Systems

  7. [14]

    Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim. 2022. Impossibility theorems for feature attribution. Proceedings of the National Academy of Sciences of the United States of America 121 (2022)

  8. [15]

    Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Vi’egas, and Martin Wattenberg. 2021. An Interpretability Illusion for BERT. ArXiv abs/2104.07143 (2021)

  9. [16]

    Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M

    Jake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie Chan, Nicolas Manfred Otto Heess, Lucy Gonzalez, Simon Osind...

  10. [17]

    Aditya Chattopadhyay, Stewart Slocum, Benjamin David Haeffele, René Vidal, and Donald Geman. 2022. Interpretable by Design: Learning Predictors by Composing Interpretable Queries. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2022), 7430–7443

  11. [18]

    Maheep Chaudhary and Atticus Geiger. 2024. Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small. ArXiv abs/2409.04478 (2024)

  12. [19]

    Paul Constantin Chelarescu. 2021. Deception in Social Learning: A Multi-Agent Reinforcement Learning Perspective. ArXiv abs/2106.05402 (2021)

  13. [20]

    Zhi Chen, Yijie Bei, and Cynthia Rudin. 2020. Concept whitening for inter- pretable image recognition. Nature Machine Intelligence 2 (2020), 772 – 782

  14. [21]

    Albrecht

    Filippos Christianos, Georgios Papoudakis, Arrasy Rahman, and Stefano V. Albrecht. 2021. Scaling Multi-Agent Reinforcement Learning with Selective Parameter Sharing. ArXiv abs/2102.07475 (2021)

  15. [22]

    Albrecht

    Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2020. Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning. ArXiv abs/2006.07169 (2020)

  16. [23]

    Xiangxiang Chu and Hangjun Ye. 2017. Parameter Sharing Deep Deterministic Policy Gradient for Cooperative Multi-agent Reinforcement Learning. ArXiv abs/1710.00336 (2017)

  17. [24]

    Stephen Chung, Scott Niekum, and David Krueger. 2024. Predicting Future Actions of Reinforcement Learning Agents. ArXiv abs/2410.22459 (2024)

  18. [25]

    Alex Cloud, Jacob Goldman-Wetzler, Evzen Wybitul, Joseph Miller, and Alexan- der Matt Turner. 2024. Gradient Routing: Masking Gradients to Localize Com- putation in Neural Networks. ArXiv abs/2410.04332 (2024)

  19. [26]

    Mavor-Parker, Aengus Lynch, Stefan Heimer- sheim, and Adrià Garriga-Alonso

    Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimer- sheim, and Adrià Garriga-Alonso. 2023. Towards Automated Circuit Discovery for Mechanistic Interpretability. arXiv:2304.14997 [cs.LG]

  20. [27]

    Lundberg, and Su-In Lee

    Ian Covert, Scott M. Lundberg, and Su-In Lee. 2020. Explaining by Removing: A Unified Framework for Model Explanation. J. Mach. Learn. Res. 22 (2020), 209:1–209:90

  21. [28]

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey

  22. [29]

    Maximilian Dreyer, Reduan Achtibat, Wojciech Samek, and Sebastian La- puschkin. 2023. Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops...

  23. [30]

    Anders, Wojciech Samek, and Sebastian Lapuschkin

    Maximilian Dreyer, Frederik Pahde, Christopher J. Anders, Wojciech Samek, and Sebastian Lapuschkin. 2023. From Hope to Safety: Unlearning Biases of Deep Models via Gradient Penalization in Latent Space. In AAAI Conference on Artificial Intelligence

  24. [31]

    Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. 2024. Transcoders Find Interpretable LLM Feature Circuits. ArXiv abs/2406.11944 (2024)

  25. [32]

    Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits.Transformer Circuits Thread 1, 1 (2021), 12

  26. [33]

    Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun

    Dan Friedman, Andrew K. Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun. 2023. Interpretability Illusions in the Generalization of Simpli- fied Models. ArXiv abs/2312.03656 (2023)

  27. [34]

    Matthias Gerstgrasser, Tom Danino, and Sarah Keren. 2023. Selectively Sharing Experiences Improves Multi-Agent Reinforcement Learning. ArXiv abs/2311.00865 (2023)

  28. [35]

    Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva

  29. [36]

    Navdeep Gill, Patrick Hall, Kim Montgomery, and Nicholas Schmidt. 2020. A Responsible Machine Learning Workflow with Focus on Interpretable Models, Post-hoc Explanation, and Discrimination Testing. Inf. 11 (2020), 137

  30. [37]

    Sam Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern. 2017. Visualizing and Understanding Atari Agents. ArXiv abs/1711.00138 (2017)

  31. [38]

    Sven Gronauer and Klaus Diepold. 2021. Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review 55 (2021), 895 – 943

  32. [39]

    Hung Guei, Yan-Ru Ju, Wei-Yu Chen, and Ti-Rong Wu. 2024. Interpreting the Learned Model in MuZero Planning

  33. [40]

    Gupta, Maxim Egorov, and Mykel J

    Jayesh K. Gupta, Maxim Egorov, and Mykel J. Kochenderfer. 2017. Cooper- ative Multi-agent Control Using Deep Reinforcement Learning. In AAMAS Workshops

  34. [41]

    Pazukonis, Jimmy Ba, and Timothy P

    Danijar Hafner, J. Pazukonis, Jimmy Ba, and Timothy P. Lillicrap. 2023. Master- ing Diverse Domains through World Models. ArXiv abs/2301.04104 (2023)

  35. [42]

    Patrik Hammersborg and Inga Strümke. 2023. Information based explanation methods for deep learning agents–with applications on large open-source chess models. arXiv preprint arXiv:2309.09702 (2023)

  36. [43]

    Shanshan Han, Qifan Zhang, Yuhang Yao, Weizhao Jin, Zhaozhuo Xu, and Chaoyang He. 2024. LLM Multi-Agent Systems: Challenges and Open Problems. ArXiv abs/2402.03578 (2024)

  37. [44]

    Zhang, Shaoqing Ren, and Jian Sun

    Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), 770–778

  38. [45]

    Anna Hedström, Leander Weber, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin, and Marina M.-C. Höhne. 2022. Quantus: An Ex- plainable AI Toolkit for Responsible Evaluation of Neural Network Explanations. ArXiv abs/2202.06861 (2022)

  39. [46]

    Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor. 2018. A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 33 (2018), 750 – 797

  40. [47]

    Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz Rodríguez. 2020. Ex- plainability in Deep Reinforcement Learning. Knowl. Based Syst. 214 (2020), 106685

  41. [48]

    Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz Rodríguez. 2021. Col- lective eXplainable AI: Explaining Cooperative Strategies and Agent Contri- bution in Multiagent Reinforcement Learning With Shapley Values. IEEE Computational Intelligence Magazine 17 (2021), 59–71

  42. [49]

    Tom Hickling, Abdelhafid Zenati, Nabil Aouf, and Phillippa Spencer. 2022. Explainability in Deep Reinforcement Learning: A Review into Current Methods and Applications. Comput. Surveys 56 (2022), 1 – 35

  43. [50]

    Jacob Hilton, Nick Cammarata, Shan Carter, Gabriel Goh, and Christopher Olah

  44. [51]

    Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger

  45. [52]

    Ivanitskiy, Alex F Spies, Tilman Rauker, Guillaume Corlouer, Chris Mathwin, Lucia Quirke, Can Rager, Rusheb Shah, Dan Valentine, Cecilia G

    Michael I. Ivanitskiy, Alex F Spies, Tilman Rauker, Guillaume Corlouer, Chris Mathwin, Lucia Quirke, Can Rager, Rusheb Shah, Dan Valentine, Cecilia G. Diniz Behn, Katsumi Inoue, and Samy Wu Fung. 2023. Structured World Representa- tions in Maze-Solving Transformers. ArXiv abs/...

  46. [53]

    Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio García Castañeda, Charlie Beattie, Neil C

    Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio García Castañeda, Charlie Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuogl...

  47. [55]

    ArXiv abs/2402.17700 (2024)

    RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. ArXiv abs/2402.17700 (2024)

  48. [56]

    Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel. 2020. Prioritized Level Replay. In International Conference on Machine Learning

  49. [57]

    Zoe Juozapaitis, Anurag Koul, Alan Fern, Martin Erwig, and Finale Doshi-Velez

  50. [58]

    Shahar Katz, Yonatan Belinkov, Mor Geva, and Lior Wolf. 2024. Backward Lens: Projecting Language Model Gradients into the Vocabulary Space. ArXiv abs/2402.12865 (2024)

  51. [59]

    Erik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen, Scott Emmons, and Stuart Russell. 2024. Evidence of Learned Look-Ahead in a Chess-Playing Neural Network. ArXiv abs/2406.00877 (2024)

  52. [60]

    Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fer- nanda Viegas, and Rory Sayres. 2018. Interpretability Beyond Feature At- tribution: Quantitative Testing with Concept Activation Vectors (TCAV). arXiv:1711.11279 [stat.ML]

  53. [61]

    Hector Kohler, Quentin Delfosse, Paul Festor, and Philippe Preux. 2024. Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop. ArXiv abs/2404.10906 (2024)

  54. [62]

    J’anos Kram’ar, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024. AtP*: An efficient and scalable method for localizing LLM behaviour to components. ArXiv abs/2403.00745 (2024)

  55. [63]

    Gaspard Lambrechts, Adrien Bolland, and Damien Ernst. 2023. Informed POMDP: Leveraging Additional Information in Model-Based RL. In RLC

  56. [64]

    Gorsane, and Arnu Pretorius

    Wiem Khlifi, Siddarth Singh, Omayma Mahjoub, Ruan de Kock, Abidine Vall, R. Gorsane, and Arnu Pretorius. 2023. On Diagnostics for Understanding Agent Training Behaviour in Cooperative MARL. ArXiv abs/2312.08468 (2023)

  57. [65]

    Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Mon- tavon, Wojciech Samek, and Klaus-Robert Müller. 2019. Unmasking Clever Hans predictors and assessing what machines really learn. Nature Communications 10 (2019)

  58. [66]

    Renard, and Marcin Detyniecki

    Thibault Laugel, Marie-Jeanne Lesot, Christophe Marsala, X. Renard, and Marcin Detyniecki. 2019. The Dangers of Post-hoc Interpretability: Unjustified Counterfactual Explanations. In International Joint Conference on Artificial Intelligence

  59. [67]

    Mark Levin and Hana Chockler. 2023. Clustered Policy Decision Ranking.ArXiv abs/2311.12970 (2023)

  60. [68]

    Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, and Yi Yang. 2024. A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth (2024)

  61. [69]

    Engelhardt, Wolfgang Konen, and Laurenz Wiskott

    Moritz Lange, Raphael C. Engelhardt, Wolfgang Konen, and Laurenz Wiskott

  62. [70]

    ArXiv abs/2402.12067 (2024)

    Interpretable Brain-Inspired Representations Improve RL Performance on Visual Navigation Tasks. ArXiv abs/2402.12067 (2024)

  63. [71]

    Lundberg and Su-In Lee

    Scott M. Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. In Neural Information Processing Systems

  64. [72]

    Andreas Madsen, Himabindu Lakkaraju, Siva Reddy, and Sarath Chandar. 2024. Interpretability Needs a New Paradigm. ArXiv abs/2405.05386 (2024)

  65. [73]

    Andreas Madsen, Siva Reddy, and A. P. Sarath Chandar. 2021. Post-hoc Inter- pretability for Neural NLP: A Survey. Comput. Surveys 55 (2021), 1 – 42

  66. [74]

    Aravindh Mahendran and Andrea Vedaldi. 2015. Visualizing Deep Convolu- tional Neural Networks Using Natural Pre-images. International Journal of Computer Vision 120 (2015), 233–255

  67. [75]

    Charles Lovering, Jessica Forde, George Konidaris, Ellie Pavlick, and Michael Littman. 2022. Evaluation Beyond Task Performance: Analyzing Concepts in AlphaZero in Hex. Advances in Neural Information Processing Systems 35 (2022), 25992–26006

  68. [76]

    Abbeel, and Igor Mordatch

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, P. Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. ArXiv abs/1706.02275 (2017)

  69. [77]

    Thomas McGrath, Andrei Kapishnikov, Nenad Tomaš ev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kram- nik. 2022. Acquisition of chess knowledge in AlphaZero. Proceedings of the National Academy of Sciences 119, 47 (nov 2022). https://d...

  70. [78]

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and Editing Factual Associations in GPT. In Neural Information Processing Systems

  71. [79]

    Stephanie Milani, Nicholay Topin, Manuela Veloso, and Fei Fang. 2023. Explain- able Reinforcement Learning: A Survey and Comparative Review. Comput. Surveys 56 (2023), 1 – 36

  72. [80]

    Kamhoua, Evangelos E

    Stephanie Milani, Zhicheng Zhang, Nicholay Topin, Zheyuan Ryan Shi, Charles A. Kamhoua, Evangelos E. Papalexakis, and Fei Fang. 2022. MAVIPER: Learning Decision Tree Policies for Interpretable Multi-Agent Reinforcement Learning. In ECML/PKDD

  73. [81]

    Omayma Mahjoub, Ruan de Kock, Siddarth Singh, Wiem Khlifi, Abidine Vall, Kale ab Tessera, and Arnu Pretorius. 2023. Efficiently Quantifying Individual Agent Importance in Cooperative MARL. ArXiv abs/2312.08466 (2023)

  74. [82]

    Czarnecki, Nando de Freitas, and Oriol Vinyals

    Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan, Caglar Gulcehre, Shang- tong Zhang, Ray Jiang, Tom Le Paine, Richard Powell, Konrad Zolna, Julian Schrittwieser, David Choi, Petko Georgiev, Daniel Toyama, Aja Huang, Roman Ring, Igor Babuschkin, Timo Ewalds, Mahyar Bordbar...

  75. [83]

    Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2015. Explaining nonlinear classification decisions with deep Taylor decomposition. ArXiv abs/1512.02479 (2015)

  76. [84]

    Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt

  77. [85]

    Christopher Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom In: An Introduction to Circuits

  78. [86]

    Brown, Jack Clark, Jared Kaplan, Sam McCan- dlish, and Christopher Olah

    Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova Das- sarma, Tom Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt,...

  79. [87]

    Ulisse Mini, Peli Grietzer, Mrinank Sharma, Austin Meek, Monte Stuart Mac- Diarmid, and Alexander Matt Turner. 2023. Understanding and Controlling a Maze-Solving Policy Network. ArXiv abs/2310.08043 (2023)

  80. [88]

    Catalin Mitelut, Ben Smith, and Peter Vamplew. 2023. Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety. arXiv:2305.19223 [cs.AI]

  81. [89]

    Aðalsteinn Pálsson and Yngvi Björnsson. [n.d.]. Unveiling concepts learned by a world-class chess-playing agent

  82. [90]

    Nicholas Pochinkov and Nandi Schoots. 2024. Dissecting Language Models: Machine Unlearning via Selective Pruning. ArXiv abs/2403.01267 (2024)

  83. [91]

    ArXiv abs/2301.05217 (2023)

    Progress measures for grokking via mechanistic interpretability. ArXiv abs/2301.05217 (2023)

  84. [92]

    Yoann Poupart. 2024. Contrastive Sparse Autoencoders for Interpreting Plan- ning of Chess-Playing Agents. ArXiv abs/2406.04028 (2024)

  85. [93]

    Yunpeng Qing, Shunyu Liu, Jie Song, and Mingli Song. 2022. A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges. ArXiv abs/2211.06665 (2022)

  86. [94]

    James Orr and Ayan Dutta. 2023. Multi-Agent Deep Reinforcement Learning for Multi-Robot Applications: A Survey.Sensors (Basel, Switzerland) 23 (2023)

  87. [95]

    Pentti Paatero and Unto Tapper. 1994. Positive matrix factorization: A non- negative factor model with optimal utilization of error estimates of data values†. Environmetrics 5 (1994), 111–126

  88. [96]

    Alec Radford and Karthik Narasimhan. 2018. Improving Language Understand- ing by Generative Pre-Training

  89. [97]

    Ragodos, Tong Wang, Qihang Lin, and Xun Zhou

    Ronilo J. Ragodos, Tong Wang, Qihang Lin, and Xun Zhou. 2022. Pro- toX: Explaining a Reinforcement Learning Agent via Prototyping. ArXiv abs/2211.03162 (2022)

  90. [98]

    Eleonora Poeta, Gabriele Ciravegna, Eliana Pastor, Tania Cerquitelli, and Elena Baralis. 2023. Concept-based Explainable Artificial Intelligence: A Survey.ArXiv abs/2312.12936 (2023)

  91. [99]

    Ed- wards, Nicolas Manfred Otto Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas

    Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexan- der Novikov, Gabriel Barth-Maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley D. Ed- wards, Nicolas Manfred Otto Heess, Yutian Chen, Rai...

  92. [100]

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Anchors: High-Precision Model-Agnostic Explanations. InAAAI Conference on Artificial Intelligence

  93. [101]

    Philip Quirke and Fazl Barez. 2023. Understanding Addition in Transformers. ArXiv abs/2310.13121 (2023)

  94. [102]

    Alec Radford, Luke Metz, and Soumith Chintala. 2015. Unsupervised Represen- tation Learning with Deep Convolutional Generative Adversarial Networks. CoRR abs/1511.06434 (2015)

  95. [103]

    Sebastian Rodriguez, John Thangarajah, and Andrew Davey. 2024. Design Patterns for Explainable Agents (XAg). In Adaptive Agents and Multi-Agent Systems

  96. [104]

    Gordon, and J

    Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. 2010. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In International Conference on Artificial Intelligence and Statistics

  97. [106]

    Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell

    Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell. 2015. Policy Distillation. CoRR abs/1511.06295 (2015)

  98. [107]

    Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2015. Prioritized Experience Replay. CoRR abs/1511.05952 (2015)

  99. [108]

    Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023. Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681 [cs.CL]

  100. [109]

    Sebastian Rodriguez and John Thangarajah. 2024. Explainable Agents (XAg) by Design. In Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems (Auckland, New Zealand) (AAMAS ’24). Inter- national Foundation for Autonomous Agents and Multiag...

  101. [110]

    Hyunki Seong and David Hyunchul Shim. 2024. Self-Supervised Inter- pretable End-to-End Learning via Latent Functional Modularity. InInternational Conference on Machine Learning

  102. [111]

    Gervasio

    Pedro Sequeira, Eric Yeh, and Melinda T. Gervasio. 2019. Interestingness Ele- ments for Explainable Reinforcement Learning through Introspection. In IUI Workshops

  103. [112]

    Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2021. Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges. ArXiv abs/2103.11251 (2021)

  104. [113]

    Lloyd S. Shapley. 1988. A Value for n-person Games

  105. [114]

    Robinson

    Yonadav Shavit, Sandhini Agarwal, Miles Brundage, Contributing authors, Steven Adler, Cullen O’Keefe, Rosie Campbell, Teddy Lee, Pamela Mishkin, Tyna Eloundou, Alan Hickey, Katarina Slama, Lama Ahmad, Paul McMillan, Alex Beutel, Alexandre Passos, and David G. Robinson. 2023. P...

  106. [115]

    Lisa Schut, Nenad Tomasev, Tom McGrath, Demis Hassabis, Ulrich Paquet, and Been Kim. 2023. Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero. arXiv:2310.16410 [cs.AI]

  107. [116]

    Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra

    Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision 128 (2016), 336 – 359

  108. [117]

    Viégas, and Martin Wattenberg

    Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017. SmoothGrad: removing noise by adding noise. ArXiv abs/1706.03825 (2017)

  109. [118]

    Yiheng Su, Juni Jessy Li, and Matthew Lease. 2023. Interpretable by Design: Wrapper Boxes Combine Neural Performance with Faithful Explanations.ArXiv abs/2311.08644 (2023)

  110. [119]

    Thanveer Basha Shaik, Xiaohui Tao, Haoran Xie, Lin Li, Jianming Yong, and Hongning Dai. 2023. Adaptive Multi-Agent Deep Reinforcement Learning for Timely Healthcare Interventions

  111. [120]

    Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente. 2015. Multiagent cooperation and competition with deep reinforcement learning. PLoS ONE 12 (2015)

  112. [121]

    Mohammad Taufeeque, Philip Quirke, Maximilian Li, Chris Cundy, Aaron David Tucker, Adam Gleave, and Adrià Garriga-Alonso. 2024. Planning in a recurrent neural network that plays Sokoban

  113. [122]

    Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje

  114. [123]

    Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

    Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Neural Information Processing Systems

  115. [124]

    Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs/1409.1556 (2014)

  116. [125]

    Dennis, and Marija Slavkovik

    Ajay Vishwanath, Louise A. Dennis, and Marija Slavkovik. 2024. Reinforcement Learning and Machine ethics:a systematic review.ArXiv abs/2407.02425 (2024)

  117. [126]

    Ulrike von Luxburg. 2007. A tutorial on spectral clustering. Statistics and Computing 17 (2007), 395–416

  118. [127]

    Czarnecki, Viní- cius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z

    Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech M. Czarnecki, Viní- cius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2017. Value-Decomposition Networks For Cooperative Multi-Agent Learning. ArXiv abs/1706.0...

  119. [128]

    Lei Wang, Chengbang Ma, Xueyang Feng, Zeyu Zhang, Hao ran Yang, Jingsen Zhang, Zhi-Yang Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji rong Wen. 2023. A Survey on Large Language Model based Autonomous Agents. ArXiv abs/2308.11432 (2023)

  120. [129]

    Jiawen Wei, Hugues Turb’e, and Gianmarco Mengaldo. 2024. Revisiting the robustness of post-hoc interpretability methods. ArXiv abs/2407.19683 (2024)

  121. [130]

    Harm van Seijen, Mehdi Fatemi, Romain Laroche, Joshua Romoff, Tavian Barnes, and Jeffrey Tsang. 2017. Hybrid Reward Architecture for Reinforcement Learn- ing. ArXiv abs/1706.04208 (2017)

  122. [131]

    Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not Explanation. In Conference on Empirical Methods in Natural Language Processing

  123. [132]

    Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, L. Sifre, Trevor Cai, John P. Agapiou, Max Jaderbe...

  124. [133]

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:230...

  125. [134]

    Abbeel, and Dale Schuur- mans

    Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, P. Abbeel, and Dale Schuur- mans. 2023. Foundation Models for Decision Making: Problems, Methods, and Opportunities. ArXiv abs/2303.04129 (2023)

  126. [135]

    Jianhong Wang, Yuan Zhang, Yunjie Gu, and Tae-Kyun Kim. 2021. SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning. In Neural Information Processing Systems

  127. [136]

    Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Simon Wiedemann, Klaus-Robert Müller, and Wojciech Samek. 2019. Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning. ArXiv abs/1912.08881 (2019)

  128. [137]

    Bayen, and Yi Wu

    Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu. 2021. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Neural Information Processing Systems

  129. [138]

    Yongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan, Hangyu Mao, and Peng Liu. 2024. SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks

  130. [139]

    Zeiler and Rob Fergus

    Matthew D. Zeiler and Rob Fergus. 2013. Visualizing and Understanding Con- volutional Networks. ArXiv abs/1311.2901 (2013)

  131. [140]

    Kononova, and Aske Plaat

    Annie Wong, Thomas Bäck, Anna V. Kononova, and Aske Plaat. 2021. Mul- tiagent Deep Reinforcement Learning: Challenges and Directions Towards Human-Like Approaches. ArXiv abs/2106.15691 (2021)

  132. [141]

    Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico...

  133. [143]

    Zhou, Weinan Zhang, and Jun Wang

    Yaodong Yang, Rui Luo, Minne Li, M. Zhou, Weinan Zhang, and Jun Wang

  134. [147]

    Renos Zabounidis, Joseph Campbell, Simon Stepputtis, Dana Hughes, and Ka- tia P. Sycara. 2023. Concept Learning for Interpretable Multi-Agent Reinforce- ment Learning. ArXiv abs/2302.12232 (2023)

  135. [149]

    Dastani, and Shihan Wang

    Changxi Zhu, Mehdi M. Dastani, and Shihan Wang. 2022. A survey of multi- agent deep reinforcement learning with communication. Autonomous Agents and Multi-Agent Systems 38 (2022), 1–48

  136. [2016]

    ArXiv abs/1605.01713 (2016)

    Not Just a Black Box: Learning Important Features Through Propagating Activation Differences. ArXiv abs/1605.01713 (2016)

  137. [2018]

    ArXiv abs/1802.05438 (2018)

    Mean Field Multi-Agent Reinforcement Learning. ArXiv abs/1802.05438 (2018)

  138. [2019]

    Explainable Reinforcement Learning via Reward Decomposition

  139. [2020]

    Understanding RL vision

  140. [2023]

    ArXiv abs/2309.08600 (2023)

    Sparse Autoencoders Find Highly Interpretable Features in Language Models. ArXiv abs/2309.08600 (2023)

  141. [2024]

    ArXiv abs/2401.06102 (2024)

    Patchscopes: A Unifying Framework for Inspecting Hidden Representa- tions of Language Models. ArXiv abs/2401.06102 (2024)

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.