REVIEW 3 major objections 6 minor 149 references
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper argues that post hoc interpretability tools—relevance backpropagation, sparse autoencoders, activation steering, and circuit discovery—are a scalable alternative to designing intrinsically interpretable multi-agent…
desk verdict A clear, honest perspective piece that maps direct interpretability tools to MADRL challenges, but its central transfer assumption is asserted rather than argued, and the paper itself concedes the main weakness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the trained deep network of an agent, treated as a directly manipulable artifact rather than a black box to be replaced. The mechanisms carried by the argument are the interpretability tools themselves: layer-wise relevance propagation for attributing decisions to inputs, sparse autoencoders for eliciting interpretable features (prototypes), activation steering and causal tracing for editing or controlling internal representations, and automated circuit discovery for isolating functional sub-networks. The paper's contribution is to connect these mechanisms to specific MADRL challenges—for instance, partitioning the positive weights of a QMIX mixing network with non-negative matrix factorization to identify teams, or using steering vectors to shift swarm behaviour toward cooperation.
What would settle it
A controlled study in a cooperative multi-agent environment with a known ground-truth attribution (for instance, a reward decomposition): if relevance maps or activation-patching attributions from a trained policy do not predict the effect of intervening on the highlighted components—ablating the 'important' neurons should change behaviour more than ablating random ones—the transfer claim would be falsified.
Extended reading notes
Core claim
The paper's central claim is that direct interpretability is a versatile and scalable alternative to intrinsically interpretable models for MADRL. It posits that because trained agents are deep networks, the same post hoc methods that have opened up vision and language models—attributing decisions to inputs, locating and editing internal knowledge, steering activations, and discovering computational circuits—can be applied to multi-agent policies to answer questions that matter in the field: what spurious cues does an agent rely on, which agents share a role, what is each agent contributing, and what are agents communicating. The paper organizes these questions into a taxonomy spanning single-agent, multi-agent, and training-process levels, and argues that this route is more flexible than intrinsic interpretability, particularly for large or already-trained systems. It closes by noting that the approach will only be trustworthy once evaluation protocols exist to check post hoc explanations, which currently can produce metrics with limited predictive power.
Load-bearing premise
The load-bearing premise is that post hoc interpretability methods, developed and tested mostly on vision and language models, will transfer to multi-agent deep reinforcement learning and yield explanations that faithfully reflect what the trained policies actually compute.
Editorial extensions
If this is right
- Direct interpretability could let developers remove biases or dangerous behaviors from a trained policy without retraining, by editing internal representations.
- Interpretability of latent spaces could automate team identification in large agent populations, reducing the number of policies needing training.
- Activation steering could make swarms coordinate by shifting agents toward cooperative goals without architecture changes.
- Interpretability of the training process could improve sample efficiency, for example by prioritizing samples based on learned importance.
- Reliable evaluation protocols for post hoc explanations would need to be built before these methods are trusted in practice.
Reading between the lines
- If direct interpretability works for MADRL, it could be combined with intrinsic methods—for example, using circuit discovery to verify the mechanisms inside a concept-bottleneck agent—blurring the line between the two paradigms.
- The same toolkit could be applied to human-AI teams, where explanations of agent internals could help humans build accurate mental models of autonomous teammates.
- The evaluation bottleneck could be attacked with interventions: instead of trusting saliency maps, one could ablate the 'important' features or neurons and measure whether behavior changes more than with random ablations.
- A testable bet is that LRP-style relevance propagation can match Shapley-based credit assignment on small cooperative tasks while being far cheaper to compute.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that post hoc ``direct interpretability'' methods (feature attribution, prototypes/sparse autoencoders, latent manipulation, circuit analysis) should be adopted as a scalable and flexible alternative to intrinsically interpretable models for multi-agent deep reinforcement learning (MADRL). It proposes a taxonomy of MADRL challenges in three categories (single-agent, multi-agent, and training process), sketches concrete candidate applications such as bias identification via LRP, team identification via sparse autoencoders, and swarm coordination via activation steering, and discusses related work in post hoc interpretability for RL. The paper also acknowledges limitations of post hoc methods in Section 5.2, including saliency map unreliability and ``interpretability illusions,'' and calls for robust evaluation protocols.
Significance. The paper's value lies in its broad synthesis of modern interpretability tools and a clear organizational taxonomy that connects specific MADRL challenges to candidate post hoc methods. It is honest about known weaknesses of these methods and cites a wide body of recent literature, including work on chess agents and sparse autoencoders. If the underlying premise that CV/NLP interpretability tools transfer to multi-agent neural policies is correct, the proposed research directions could stimulate productive empirical work. However, the paper provides no evidence for transferability, and the central advocacy claim rests on an assumption that the authors themselves concede is fragile. The contribution is therefore a speculative research agenda rather than a demonstrated result; its significance will depend on whether future work can validate the proposed applications.
major comments (3)
- [§3.1–§3.3 vs §5.2] The concrete proposals in Section 3 (e.g., LRP for bias identification, SAE for team identification, activation steering for swarm coordination) require that post hoc methods yield causally faithful explanations in MADRL settings, yet Section 5.2 concedes that these methods ``often generate metrics with limited predictive power'' and are subject to ``interpretability illusions.'' This creates a direct tension: policy edition and activation steering are causal operations that require faithfulness, but the paper does not resolve how the conceded unreliability is compatible with those applications. For the central claim of the abstract to be load-bearing, the authors should either explicitly label each Section 3 proposal as an untested hypothesis or provide a concrete account of how the proposed evaluation protocols would validate or falsify each application. As written, the advocacy is internally inconsistent in its level of commitment.
- [§5.2] ``Robust Evaluation Protocols'' is too underspecified to support the advocacy. The section cites existing evaluation work (e.g., [7, 18, 36, 45, 51, 73, 129]) and states that reliable metrics must be established, but does not state what would count as evidence for transferability of the Section 3 applications. Because the absence of ground-truth explanations is acknowledged as a central difficulty, the paper needs at least one concrete evaluation idea (for instance, controlled interventions on policies, behavioral tests with perturbed observations, or ablation studies that compare explanation-predicted importance with actual policy changes). Without such a proposal, the call for direct interpretability is unfalsifiable, and the limitation statement in Section 5.2 does not resolve the burden placed on the earlier sections.
- [§4.3] The comparison with intrinsically interpretable models is one-sided. The section argues that intrinsic models face scalability and flexibility challenges, but it does not engage with the well-known argument that high-stakes applications may require intrinsically interpretable models precisely because post hoc explanations can be misleading (the paper cites [105] but does not respond to its central challenge). Since the paper's thesis is that direct interpretability is a viable alternative, Section 4.3 should at least acknowledge the conditions under which intrinsic interpretability remains preferable, and explain why the proposed direct methods overcome those conditions in MADRL. This would strengthen the advocacy rather than weaken it.
minor comments (6)
- [Abstract vs §6] The abstract states that direct interpretability ``offering insights into agents' behaviour, emergent phenomena, and biases,'' which is assertive, while the conclusion (Section 6) says ``direct interpretability might be vital.'' The framing should be aligned so that the abstract reflects the appropriately hedged nature of a position paper.
- [Figure 1] The caption refers to green, blue, and red colors for different challenge categories, but the figure itself is not included in the text and the color coding is not explained. If the figure is published, the caption should be self-contained.
- [§2.2] In the ``Prototypes'' paragraph, sparse autoencoders are mentioned in one sentence; given their prominent role in later sections (e.g., §3.2), a brief explanation of how SAEs elicit features would help readers unfamiliar with this method.
- [Figure 1] There is a typo in the figure: ``Swarn Coordination'' should be ``Swarm Coordination.''
- [§5.2] Reference [33] appears twice in the list ``[15, 33, 33]''; the duplicate should be removed.
- [§1] The phrase ``these approaches often need to be revised for large and performant systems'' is unclear; it likely means ``these approaches often need to be revisited'' or ``are often inadequate for.'' Please rephrase.
Circularity Check
No significant circularity: this is a position paper with no derivations, fitted parameters, or empirical predictions, and its only self-citation is peripheral background.
full rationale
The paper is a vision/position piece advocating direct interpretability methods for MADRL. It contains no mathematical derivations, no fitted parameters, no experimental results, and no equation-level dependencies. The central claim is an advocacy claim: post hoc interpretability methods from CV/NLP and XRL could be usefully transferred to MADRL challenges. This is supported by citing prior work on such methods and by proposing speculative applications (e.g., LRP for bias identification, SAEs for team identification, activation steering for swarm coordination). None of these proposals is presented as a derived prediction from the paper's own inputs. The only self-citation is reference [92], Yoann Poupart's 'Contrastive Sparse Autoencoders for Interpreting Planning of Chess-Playing Agents,' cited in Section 4.1 as one of several efforts interpreting chess engines like AlphaZero. That citation is descriptive background, not a load-bearing premise: the paper's advocacy does not reduce to it, and the argument would stand equally without it. Section 5.2 even concedes limitations of post hoc methods (saliency map shortcomings, interpretability illusions, metrics with limited predictive power), which is the opposite of circularly assuming their validity. Therefore the circularity burden is zero and the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Post hoc interpretability methods developed for CV and NLP are applicable to MADRL policies and training processes.
- domain assumption Explanations from direct methods are faithful enough to support debugging, steering, and evaluation.
Cite this review
Pith. "Pith review of Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning." pith.science (2026). https://pith.science/paper/5EBVCYWJ
@misc{pith2026250200726,
author = {Pith},
title = {Pith review of: Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/5EBVCYWJ}},
note = {Machine review of arXiv:2502.00726}
}
read the original abstract
Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While learning intrinsically interpretable models remains a prominent approach, its scalability and flexibility are limited in handling complex tasks or multi-agent dynamics. This paper advocates for direct interpretability, generating post hoc explanations directly from trained models, as a versatile and scalable alternative, offering insights into agents' behaviour, emergent phenomena, and biases without altering models' architectures. We explore modern methods, including relevance backpropagation, knowledge edition, model steering, activation patching, sparse autoencoders and circuit discovery, to highlight their applicability to single-agent, multi-agent, and training process challenges. By addressing MADRL interpretability, we propose directions aiming to advance active topics such as team identification, swarm coordination and sample efficiency.
Figures
Reference graph
Works this paper leans on
-
[105]
Tabish Rashid, Mikayel Samvelyan, C. S. D. Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic Value Function Fac- torisation for Deep Multi-Agent Reinforcement Learning.ArXiv abs/1803.11485 (2018)
arXiv 2018
-
[1]
Reduan Achtibat, Maximilian Dreyer, Ilona Eisenbraun, Sebastian Bosse, Thomas Wiegand, Wojciech Samek, and Sebastian Lapuschkin. 2022. From attribution maps to human-understandable explanations through Concept Rel- evance Propagation. Nature Machine Intelligence 5 (2022), 1006 – 1019
2022
-
[2]
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. 2024. AttnLRP: Attention-Aware Layer-wise Relevance Propagation for Transformers. ArXiv abs/2402.05602 (2024)
arXiv 2024
-
[3]
Goodfellow, Moritz Hardt, and Been Kim
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim. 2018. Sanity Checks for Saliency Maps. In Neural Information Processing Systems
2018
-
[4]
Guillaume Alain and Yoshua Bengio. 2018. Understanding intermediate layers using linear classifier probes. arXiv:1610.01644 [stat.ML]
arXiv 2018
-
[5]
Gülsüm Alicioğlu and Bo Sun. 2024. Use Bag-of-Patterns Approach to Explore Learned Behaviors of Reinforcement Learning. In xAI
2024
-
[6]
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, and François Fleuret. [n.d.]. Diffusion for World Modeling: Visual Details Matter in Atari. In Thirty-eighth Conference on Neural Information Processing Systems
-
[7]
José Pereira Amorim, Pedro Henriques Abreu, João A. M. Santos, and Henning Müller. 2023. Evaluating Post-hoc Interpretability with Intrinsic Interpretability. ArXiv abs/2305.03002 (2023)
work page Pith review arXiv 2023
Show all 149 references
-
[8]
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015. On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation.PLoS ONE 10 (2015)
2015
-
[9]
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama. 2018. Verifiable Re- inforcement Learning via Policy Extraction. In Neural Information Processing Systems
2018
-
[10]
Yanzhe Bekkemoen. 2023. Explainable reinforcement learning (XRL): a system- atic literature review and taxonomy. Machine Learning 113 (2023), 355–441
2023
-
[11]
Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor V. Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting Latent Predictions from Transformers with the Tuned Lens. ArXiv abs/2303.08112 (2023)
2023 arXiv
-
[12]
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. LEACE: Perfect linear concept erasure in closed form. arXiv:2306.03819 [cs.LG]
2023 arXiv
-
[13]
David Bertoin, Adil Zouitine, Mehdi Zouitine, and Emmanuel Rachelson. 2022. Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement Learning. In Neural Information Processing Systems
2022
-
[14]
Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim. 2022. Impossibility theorems for feature attribution. Proceedings of the National Academy of Sciences of the United States of America 121 (2022)
2022
-
[15]
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Vi’egas, and Martin Wattenberg. 2021. An Interpretability Illusion for BERT. ArXiv abs/2104.07143 (2021)
2021 arXiv
-
[16]
Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M
Jake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie Chan, Nicolas Manfred Otto Heess, Lucy Gonzalez, Simon Osind...
2024 arXiv
-
[17]
Aditya Chattopadhyay, Stewart Slocum, Benjamin David Haeffele, René Vidal, and Donald Geman. 2022. Interpretable by Design: Learning Predictors by Composing Interpretable Queries. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2022), 7430–7443
2022
-
[18]
Maheep Chaudhary and Atticus Geiger. 2024. Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small. ArXiv abs/2409.04478 (2024)
2024 arXiv
-
[19]
Paul Constantin Chelarescu. 2021. Deception in Social Learning: A Multi-Agent Reinforcement Learning Perspective. ArXiv abs/2106.05402 (2021)
2021 arXiv
-
[20]
Zhi Chen, Yijie Bei, and Cynthia Rudin. 2020. Concept whitening for inter- pretable image recognition. Nature Machine Intelligence 2 (2020), 772 – 782
2020
-
[21]
Albrecht
Filippos Christianos, Georgios Papoudakis, Arrasy Rahman, and Stefano V. Albrecht. 2021. Scaling Multi-Agent Reinforcement Learning with Selective Parameter Sharing. ArXiv abs/2102.07475 (2021)
2021 arXiv
-
[22]
Albrecht
Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2020. Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning. ArXiv abs/2006.07169 (2020)
2020 arXiv
-
[23]
Xiangxiang Chu and Hangjun Ye. 2017. Parameter Sharing Deep Deterministic Policy Gradient for Cooperative Multi-agent Reinforcement Learning. ArXiv abs/1710.00336 (2017)
2017 arXiv
-
[24]
Stephen Chung, Scott Niekum, and David Krueger. 2024. Predicting Future Actions of Reinforcement Learning Agents. ArXiv abs/2410.22459 (2024)
2024 arXiv
-
[25]
Alex Cloud, Jacob Goldman-Wetzler, Evzen Wybitul, Joseph Miller, and Alexan- der Matt Turner. 2024. Gradient Routing: Masking Gradients to Localize Com- putation in Neural Networks. ArXiv abs/2410.04332 (2024)
2024 arXiv
-
[26]
Mavor-Parker, Aengus Lynch, Stefan Heimer- sheim, and Adrià Garriga-Alonso
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimer- sheim, and Adrià Garriga-Alonso. 2023. Towards Automated Circuit Discovery for Mechanistic Interpretability. arXiv:2304.14997 [cs.LG]
2023 arXiv
-
[27]
Lundberg, and Su-In Lee
Ian Covert, Scott M. Lundberg, and Su-In Lee. 2020. Explaining by Removing: A Unified Framework for Model Explanation. J. Mach. Learn. Res. 22 (2020), 209:1–209:90
2020
-
[28]
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey
-
[29]
Maximilian Dreyer, Reduan Achtibat, Wojciech Samek, and Sebastian La- puschkin. 2023. Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops...
2023
-
[30]
Anders, Wojciech Samek, and Sebastian Lapuschkin
Maximilian Dreyer, Frederik Pahde, Christopher J. Anders, Wojciech Samek, and Sebastian Lapuschkin. 2023. From Hope to Safety: Unlearning Biases of Deep Models via Gradient Penalization in Latent Space. In AAAI Conference on Artificial Intelligence
2023
-
[31]
Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. 2024. Transcoders Find Interpretable LLM Feature Circuits. ArXiv abs/2406.11944 (2024)
2024 arXiv
-
[32]
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits.Transformer Circuits Thread 1, 1 (2021), 12
2021
-
[33]
Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun
Dan Friedman, Andrew K. Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun. 2023. Interpretability Illusions in the Generalization of Simpli- fied Models. ArXiv abs/2312.03656 (2023)
2023 arXiv
-
[34]
Matthias Gerstgrasser, Tom Danino, and Sarah Keren. 2023. Selectively Sharing Experiences Improves Multi-Agent Reinforcement Learning. ArXiv abs/2311.00865 (2023)
2023 arXiv
-
[35]
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva
-
[36]
Navdeep Gill, Patrick Hall, Kim Montgomery, and Nicholas Schmidt. 2020. A Responsible Machine Learning Workflow with Focus on Interpretable Models, Post-hoc Explanation, and Discrimination Testing. Inf. 11 (2020), 137
2020
-
[37]
Sam Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern. 2017. Visualizing and Understanding Atari Agents. ArXiv abs/1711.00138 (2017)
2017 arXiv
-
[38]
Sven Gronauer and Klaus Diepold. 2021. Multi-agent deep reinforcement learning: a survey. Artificial Intelligence Review 55 (2021), 895 – 943
2021
-
[39]
Hung Guei, Yan-Ru Ju, Wei-Yu Chen, and Ti-Rong Wu. 2024. Interpreting the Learned Model in MuZero Planning
2024
-
[40]
Gupta, Maxim Egorov, and Mykel J
Jayesh K. Gupta, Maxim Egorov, and Mykel J. Kochenderfer. 2017. Cooper- ative Multi-agent Control Using Deep Reinforcement Learning. In AAMAS Workshops
2017
-
[41]
Pazukonis, Jimmy Ba, and Timothy P
Danijar Hafner, J. Pazukonis, Jimmy Ba, and Timothy P. Lillicrap. 2023. Master- ing Diverse Domains through World Models. ArXiv abs/2301.04104 (2023)
2023 arXiv
-
[42]
Patrik Hammersborg and Inga Strümke. 2023. Information based explanation methods for deep learning agents–with applications on large open-source chess models. arXiv preprint arXiv:2309.09702 (2023)
2023 arXiv
-
[43]
Shanshan Han, Qifan Zhang, Yuhang Yao, Weizhao Jin, Zhaozhuo Xu, and Chaoyang He. 2024. LLM Multi-Agent Systems: Challenges and Open Problems. ArXiv abs/2402.03578 (2024)
2024 arXiv
-
[44]
Zhang, Shaoqing Ren, and Jian Sun
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), 770–778
2015
-
[45]
Anna Hedström, Leander Weber, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin, and Marina M.-C. Höhne. 2022. Quantus: An Ex- plainable AI Toolkit for Responsible Evaluation of Neural Network Explanations. ArXiv abs/2202.06861 (2022)
2022 arXiv
-
[46]
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor. 2018. A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 33 (2018), 750 – 797
2018
-
[47]
Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz Rodríguez. 2020. Ex- plainability in Deep Reinforcement Learning. Knowl. Based Syst. 214 (2020), 106685
2020
-
[48]
Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz Rodríguez. 2021. Col- lective eXplainable AI: Explaining Cooperative Strategies and Agent Contri- bution in Multiagent Reinforcement Learning With Shapley Values. IEEE Computational Intelligence Magazine 17 (2021), 59–71
2021
-
[49]
Tom Hickling, Abdelhafid Zenati, Nabil Aouf, and Phillippa Spencer. 2022. Explainability in Deep Reinforcement Learning: A Review into Current Methods and Applications. Comput. Surveys 56 (2022), 1 – 35
2022
-
[50]
Jacob Hilton, Nick Cammarata, Shan Carter, Gabriel Goh, and Christopher Olah
-
[51]
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger
-
[52]
Ivanitskiy, Alex F Spies, Tilman Rauker, Guillaume Corlouer, Chris Mathwin, Lucia Quirke, Can Rager, Rusheb Shah, Dan Valentine, Cecilia G
Michael I. Ivanitskiy, Alex F Spies, Tilman Rauker, Guillaume Corlouer, Chris Mathwin, Lucia Quirke, Can Rager, Rusheb Shah, Dan Valentine, Cecilia G. Diniz Behn, Katsumi Inoue, and Samy Wu Fung. 2023. Structured World Representa- tions in Maze-Solving Transformers. ArXiv abs/...
2023 arXiv
-
[53]
Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio García Castañeda, Charlie Beattie, Neil C
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio García Castañeda, Charlie Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuogl...
2018
-
[55]
ArXiv abs/2402.17700 (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. ArXiv abs/2402.17700 (2024)
2024 arXiv
-
[56]
Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel. 2020. Prioritized Level Replay. In International Conference on Machine Learning
2020
-
[57]
Zoe Juozapaitis, Anurag Koul, Alan Fern, Martin Erwig, and Finale Doshi-Velez
-
[58]
Shahar Katz, Yonatan Belinkov, Mor Geva, and Lior Wolf. 2024. Backward Lens: Projecting Language Model Gradients into the Vocabulary Space. ArXiv abs/2402.12865 (2024)
2024 arXiv
-
[59]
Erik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen, Scott Emmons, and Stuart Russell. 2024. Evidence of Learned Look-Ahead in a Chess-Playing Neural Network. ArXiv abs/2406.00877 (2024)
2024 arXiv
-
[60]
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fer- nanda Viegas, and Rory Sayres. 2018. Interpretability Beyond Feature At- tribution: Quantitative Testing with Concept Activation Vectors (TCAV). arXiv:1711.11279 [stat.ML]
2018 arXiv
-
[61]
Hector Kohler, Quentin Delfosse, Paul Festor, and Philippe Preux. 2024. Towards a Research Community in Interpretable Reinforcement Learning: the InterpPol Workshop. ArXiv abs/2404.10906 (2024)
2024 arXiv
-
[62]
J’anos Kram’ar, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024. AtP*: An efficient and scalable method for localizing LLM behaviour to components. ArXiv abs/2403.00745 (2024)
2024 arXiv
-
[63]
Gaspard Lambrechts, Adrien Bolland, and Damien Ernst. 2023. Informed POMDP: Leveraging Additional Information in Model-Based RL. In RLC
2023
-
[64]
Gorsane, and Arnu Pretorius
Wiem Khlifi, Siddarth Singh, Omayma Mahjoub, Ruan de Kock, Abidine Vall, R. Gorsane, and Arnu Pretorius. 2023. On Diagnostics for Understanding Agent Training Behaviour in Cooperative MARL. ArXiv abs/2312.08468 (2023)
2023 arXiv
-
[65]
Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Mon- tavon, Wojciech Samek, and Klaus-Robert Müller. 2019. Unmasking Clever Hans predictors and assessing what machines really learn. Nature Communications 10 (2019)
2019
-
[66]
Renard, and Marcin Detyniecki
Thibault Laugel, Marie-Jeanne Lesot, Christophe Marsala, X. Renard, and Marcin Detyniecki. 2019. The Dangers of Post-hoc Interpretability: Unjustified Counterfactual Explanations. In International Joint Conference on Artificial Intelligence
2019
-
[67]
Mark Levin and Hana Chockler. 2023. Clustered Policy Decision Ranking.ArXiv abs/2311.12970 (2023)
2023 arXiv
-
[68]
Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, and Yi Yang. 2024. A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth (2024)
2024
-
[69]
Engelhardt, Wolfgang Konen, and Laurenz Wiskott
Moritz Lange, Raphael C. Engelhardt, Wolfgang Konen, and Laurenz Wiskott
-
[70]
ArXiv abs/2402.12067 (2024)
Interpretable Brain-Inspired Representations Improve RL Performance on Visual Navigation Tasks. ArXiv abs/2402.12067 (2024)
2024 arXiv
-
[71]
Lundberg and Su-In Lee
Scott M. Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. In Neural Information Processing Systems
2017
-
[72]
Andreas Madsen, Himabindu Lakkaraju, Siva Reddy, and Sarath Chandar. 2024. Interpretability Needs a New Paradigm. ArXiv abs/2405.05386 (2024)
2024 arXiv
-
[73]
Andreas Madsen, Siva Reddy, and A. P. Sarath Chandar. 2021. Post-hoc Inter- pretability for Neural NLP: A Survey. Comput. Surveys 55 (2021), 1 – 42
2021
-
[74]
Aravindh Mahendran and Andrea Vedaldi. 2015. Visualizing Deep Convolu- tional Neural Networks Using Natural Pre-images. International Journal of Computer Vision 120 (2015), 233–255
2015
-
[75]
Charles Lovering, Jessica Forde, George Konidaris, Ellie Pavlick, and Michael Littman. 2022. Evaluation Beyond Task Performance: Analyzing Concepts in AlphaZero in Hex. Advances in Neural Information Processing Systems 35 (2022), 25992–26006
2022
-
[76]
Abbeel, and Igor Mordatch
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, P. Abbeel, and Igor Mordatch. 2017. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. ArXiv abs/1706.02275 (2017)
2017 arXiv
-
[77]
Thomas McGrath, Andrei Kapishnikov, Nenad Tomaš ev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kram- nik. 2022. Acquisition of chess knowledge in AlphaZero. Proceedings of the National Academy of Sciences 119, 47 (nov 2022). https://d...
2022 doi
-
[78]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and Editing Factual Associations in GPT. In Neural Information Processing Systems
2022
-
[79]
Stephanie Milani, Nicholay Topin, Manuela Veloso, and Fei Fang. 2023. Explain- able Reinforcement Learning: A Survey and Comparative Review. Comput. Surveys 56 (2023), 1 – 36
2023
-
[80]
Kamhoua, Evangelos E
Stephanie Milani, Zhicheng Zhang, Nicholay Topin, Zheyuan Ryan Shi, Charles A. Kamhoua, Evangelos E. Papalexakis, and Fei Fang. 2022. MAVIPER: Learning Decision Tree Policies for Interpretable Multi-Agent Reinforcement Learning. In ECML/PKDD
2022
-
[81]
Omayma Mahjoub, Ruan de Kock, Siddarth Singh, Wiem Khlifi, Abidine Vall, Kale ab Tessera, and Arnu Pretorius. 2023. Efficiently Quantifying Individual Agent Importance in Cooperative MARL. ArXiv abs/2312.08466 (2023)
2023 arXiv
-
[82]
Czarnecki, Nando de Freitas, and Oriol Vinyals
Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan, Caglar Gulcehre, Shang- tong Zhang, Ray Jiang, Tom Le Paine, Richard Powell, Konrad Zolna, Julian Schrittwieser, David Choi, Petko Georgiev, Daniel Toyama, Aja Huang, Roman Ring, Igor Babuschkin, Timo Ewalds, Mahyar Bordbar...
2023 arXiv
-
[83]
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2015. Explaining nonlinear classification decisions with deep Taylor decomposition. ArXiv abs/1512.02479 (2015)
2015 arXiv
-
[84]
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt
-
[85]
Christopher Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom In: An Introduction to Circuits
2020
-
[86]
Brown, Jack Clark, Jared Kaplan, Sam McCan- dlish, and Christopher Olah
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova Das- sarma, Tom Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt,...
2022 arXiv
-
[87]
Ulisse Mini, Peli Grietzer, Mrinank Sharma, Austin Meek, Monte Stuart Mac- Diarmid, and Alexander Matt Turner. 2023. Understanding and Controlling a Maze-Solving Policy Network. ArXiv abs/2310.08043 (2023)
2023 arXiv
-
[88]
Catalin Mitelut, Ben Smith, and Peter Vamplew. 2023. Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety. arXiv:2305.19223 [cs.AI]
2023 arXiv
-
[89]
Aðalsteinn Pálsson and Yngvi Björnsson. [n.d.]. Unveiling concepts learned by a world-class chess-playing agent
-
[90]
Nicholas Pochinkov and Nandi Schoots. 2024. Dissecting Language Models: Machine Unlearning via Selective Pruning. ArXiv abs/2403.01267 (2024)
2024 arXiv
-
[91]
ArXiv abs/2301.05217 (2023)
Progress measures for grokking via mechanistic interpretability. ArXiv abs/2301.05217 (2023)
2023 arXiv
-
[92]
Yoann Poupart. 2024. Contrastive Sparse Autoencoders for Interpreting Plan- ning of Chess-Playing Agents. ArXiv abs/2406.04028 (2024)
2024 arXiv
-
[93]
Yunpeng Qing, Shunyu Liu, Jie Song, and Mingli Song. 2022. A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges. ArXiv abs/2211.06665 (2022)
2022 arXiv
-
[94]
James Orr and Ayan Dutta. 2023. Multi-Agent Deep Reinforcement Learning for Multi-Robot Applications: A Survey.Sensors (Basel, Switzerland) 23 (2023)
2023
-
[95]
Pentti Paatero and Unto Tapper. 1994. Positive matrix factorization: A non- negative factor model with optimal utilization of error estimates of data values†. Environmetrics 5 (1994), 111–126
1994
-
[96]
Alec Radford and Karthik Narasimhan. 2018. Improving Language Understand- ing by Generative Pre-Training
2018
-
[97]
Ragodos, Tong Wang, Qihang Lin, and Xun Zhou
Ronilo J. Ragodos, Tong Wang, Qihang Lin, and Xun Zhou. 2022. Pro- toX: Explaining a Reinforcement Learning Agent via Prototyping. ArXiv abs/2211.03162 (2022)
2022 arXiv
-
[98]
Eleonora Poeta, Gabriele Ciravegna, Eliana Pastor, Tania Cerquitelli, and Elena Baralis. 2023. Concept-based Explainable Artificial Intelligence: A Survey.ArXiv abs/2312.12936 (2023)
2023
-
[99]
Ed- wards, Nicolas Manfred Otto Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexan- der Novikov, Gabriel Barth-Maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley D. Ed- wards, Nicolas Manfred Otto Heess, Yutian Chen, Rai...
2022 arXiv
-
[100]
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Anchors: High-Precision Model-Agnostic Explanations. InAAAI Conference on Artificial Intelligence
2018
-
[101]
Philip Quirke and Fazl Barez. 2023. Understanding Addition in Transformers. ArXiv abs/2310.13121 (2023)
2023 arXiv
-
[102]
Alec Radford, Luke Metz, and Soumith Chintala. 2015. Unsupervised Represen- tation Learning with Deep Convolutional Generative Adversarial Networks. CoRR abs/1511.06434 (2015)
2015 arXiv
-
[103]
Sebastian Rodriguez, John Thangarajah, and Andrew Davey. 2024. Design Patterns for Explainable Agents (XAg). In Adaptive Agents and Multi-Agent Systems
2024
-
[104]
Gordon, and J
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. 2010. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In International Conference on Artificial Intelligence and Statistics
2010
-
[106]
Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell
Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell. 2015. Policy Distillation. CoRR abs/1511.06295 (2015)
2015 arXiv
-
[107]
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2015. Prioritized Experience Replay. CoRR abs/1511.05952 (2015)
2015 arXiv
-
[108]
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023. Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681 [cs.CL]
2023 arXiv
-
[109]
Sebastian Rodriguez and John Thangarajah. 2024. Explainable Agents (XAg) by Design. In Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems (Auckland, New Zealand) (AAMAS ’24). Inter- national Foundation for Autonomous Agents and Multiag...
2024
-
[110]
Hyunki Seong and David Hyunchul Shim. 2024. Self-Supervised Inter- pretable End-to-End Learning via Latent Functional Modularity. InInternational Conference on Machine Learning
2024
-
[111]
Gervasio
Pedro Sequeira, Eric Yeh, and Melinda T. Gervasio. 2019. Interestingness Ele- ments for Explainable Reinforcement Learning through Introspection. In IUI Workshops
2019
-
[112]
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2021. Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges. ArXiv abs/2103.11251 (2021)
2021 arXiv
-
[113]
Lloyd S. Shapley. 1988. A Value for n-person Games
1988
-
[114]
Robinson
Yonadav Shavit, Sandhini Agarwal, Miles Brundage, Contributing authors, Steven Adler, Cullen O’Keefe, Rosie Campbell, Teddy Lee, Pamela Mishkin, Tyna Eloundou, Alan Hickey, Katarina Slama, Lama Ahmad, Paul McMillan, Alex Beutel, Alexandre Passos, and David G. Robinson. 2023. P...
2023
-
[115]
Lisa Schut, Nenad Tomasev, Tom McGrath, Demis Hassabis, Ulrich Paquet, and Been Kim. 2023. Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero. arXiv:2310.16410 [cs.AI]
2023 arXiv
-
[116]
Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. International Journal of Computer Vision 128 (2016), 336 – 359
2016
-
[117]
Viégas, and Martin Wattenberg
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017. SmoothGrad: removing noise by adding noise. ArXiv abs/1706.03825 (2017)
2017 arXiv
-
[118]
Yiheng Su, Juni Jessy Li, and Matthew Lease. 2023. Interpretable by Design: Wrapper Boxes Combine Neural Performance with Faithful Explanations.ArXiv abs/2311.08644 (2023)
2023 arXiv
-
[119]
Thanveer Basha Shaik, Xiaohui Tao, Haoran Xie, Lin Li, Jianming Yong, and Hongning Dai. 2023. Adaptive Multi-Agent Deep Reinforcement Learning for Timely Healthcare Interventions
2023
-
[120]
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente. 2015. Multiagent cooperation and competition with deep reinforcement learning. PLoS ONE 12 (2015)
2015
-
[121]
Mohammad Taufeeque, Philip Quirke, Maximilian Li, Chris Cundy, Aaron David Tucker, Adam Gleave, and Adrià Garriga-Alonso. 2024. Planning in a recurrent neural network that plays Sokoban
2024
-
[122]
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje
-
[123]
Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Neural Information Processing Systems
2017
-
[124]
Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs/1409.1556 (2014)
2014 arXiv
-
[125]
Dennis, and Marija Slavkovik
Ajay Vishwanath, Louise A. Dennis, and Marija Slavkovik. 2024. Reinforcement Learning and Machine ethics:a systematic review.ArXiv abs/2407.02425 (2024)
2024
-
[126]
Ulrike von Luxburg. 2007. A tutorial on spectral clustering. Statistics and Computing 17 (2007), 395–416
2007
-
[127]
Czarnecki, Viní- cius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech M. Czarnecki, Viní- cius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2017. Value-Decomposition Networks For Cooperative Multi-Agent Learning. ArXiv abs/1706.0...
2017 arXiv
-
[128]
Lei Wang, Chengbang Ma, Xueyang Feng, Zeyu Zhang, Hao ran Yang, Jingsen Zhang, Zhi-Yang Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji rong Wen. 2023. A Survey on Large Language Model based Autonomous Agents. ArXiv abs/2308.11432 (2023)
2023 arXiv
-
[129]
Jiawen Wei, Hugues Turb’e, and Gianmarco Mengaldo. 2024. Revisiting the robustness of post-hoc interpretability methods. ArXiv abs/2407.19683 (2024)
2024 arXiv
-
[130]
Harm van Seijen, Mehdi Fatemi, Romain Laroche, Joshua Romoff, Tavian Barnes, and Jeffrey Tsang. 2017. Hybrid Reward Architecture for Reinforcement Learn- ing. ArXiv abs/1706.04208 (2017)
2017 arXiv
-
[131]
Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not Explanation. In Conference on Empirical Methods in Natural Language Processing
2019
-
[132]
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, L. Sifre, Trevor Cai, John P. Agapiou, Max Jaderbe...
2019
-
[133]
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:230...
2023 arXiv
-
[134]
Abbeel, and Dale Schuur- mans
Sherry Yang, Ofir Nachum, Yilun Du, Jason Wei, P. Abbeel, and Dale Schuur- mans. 2023. Foundation Models for Decision Making: Problems, Methods, and Opportunities. ArXiv abs/2303.04129 (2023)
2023 arXiv
-
[135]
Jianhong Wang, Yuan Zhang, Yunjie Gu, and Tae-Kyun Kim. 2021. SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning. In Neural Information Processing Systems
2021
-
[136]
Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Simon Wiedemann, Klaus-Robert Müller, and Wojciech Samek. 2019. Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning. ArXiv abs/1912.08881 (2019)
2019 arXiv
-
[137]
Bayen, and Yi Wu
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu. 2021. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Neural Information Processing Systems
2021
-
[138]
Yongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan, Hangyu Mao, and Peng Liu. 2024. SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks
2024
-
[139]
Zeiler and Rob Fergus
Matthew D. Zeiler and Rob Fergus. 2013. Visualizing and Understanding Con- volutional Networks. ArXiv abs/1311.2901 (2013)
2013 arXiv
-
[140]
Kononova, and Aske Plaat
Annie Wong, Thomas Bäck, Anna V. Kononova, and Aske Plaat. 2021. Mul- tiagent Deep Reinforcement Learning: Challenges and Directions Towards Human-Like Approaches. ArXiv abs/2106.15691 (2021)
2021 arXiv
-
[141]
Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico...
2023 arXiv
-
[143]
Zhou, Weinan Zhang, and Jun Wang
Yaodong Yang, Rui Luo, Minne Li, M. Zhou, Weinan Zhang, and Jun Wang
-
[147]
Renos Zabounidis, Joseph Campbell, Simon Stepputtis, Dana Hughes, and Ka- tia P. Sycara. 2023. Concept Learning for Interpretable Multi-Agent Reinforce- ment Learning. ArXiv abs/2302.12232 (2023)
2023 arXiv
-
[149]
Dastani, and Shihan Wang
Changxi Zhu, Mehdi M. Dastani, and Shihan Wang. 2022. A survey of multi- agent deep reinforcement learning with communication. Autonomous Agents and Multi-Agent Systems 38 (2022), 1–48
2022
-
[2016]
ArXiv abs/1605.01713 (2016)
Not Just a Black Box: Learning Important Features Through Propagating Activation Differences. ArXiv abs/1605.01713 (2016)
2016 arXiv
-
[2018]
ArXiv abs/1802.05438 (2018)
Mean Field Multi-Agent Reinforcement Learning. ArXiv abs/1802.05438 (2018)
2018 arXiv
-
[2019]
Explainable Reinforcement Learning via Reward Decomposition
-
[2020]
Understanding RL vision
-
[2023]
ArXiv abs/2309.08600 (2023)
Sparse Autoencoders Find Highly Interpretable Features in Language Models. ArXiv abs/2309.08600 (2023)
2023 arXiv
-
[2024]
ArXiv abs/2401.06102 (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representa- tions of Language Models. ArXiv abs/2401.06102 (2024)
2024 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.