REVIEW 3 major objections 4 minor 291 references
A Comprehensive Guide to Explainable AI: From Classical Models to LLMs
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This book-length guide argues that explainable AI is a teachable, unified field spanning classical interpretable models, post-hoc attribution methods, and large-language-model probing, and supports the argument with Python examples.
desk verdict A broad, mostly sound XAI survey whose practical value is undercut by a handful of fixable code and structural errors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing machinery is a two-axis view of interpretability: intrinsic versus post-hoc, and model-based versus model-agnostic techniques. The main working objects are feature attribution methods — SHAP values from cooperative game theory, LIME's local surrogate models, Integrated Gradients' path integrals, Grad-CAM's gradient-weighted activation maps, and Layer-wise Relevance Propagation's backward relevance decomposition — together with attention-weight visualization and embedding/probing analysis for transformers and LLMs. Code examples are the load-bearing mechanism for transferring the techniques to readers; each chapter pairs a method with a minimal Python implementation, and the book's promise of practical mastery rests on those implementations being correct.
What would settle it
Run the Section 4.3.1 RNN code and inspect what `model.predict(X)` returns: if the plotted curves are the network's final predictions rather than the 10 hidden-unit activations the caption claims, the example does not demonstrate hidden-state interpretation. More generally, executing each chapter's code and checking that the printed outputs and figures match the described behavior would settle whether the practical promise holds.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is organizational: the many threads of XAI — intrinsic model interpretability, feature attribution, counterfactual and causal explanation, attention and embedding analysis, evaluation metrics, and tooling — belong in a single narrative that runs from classical models to LLMs. The guide asserts that this narrative can be made hands-on, with Python code for each technique, and that the resulting competence lets practitioners debug models, satisfy regulatory demands for explanation, and audit fairness. No new algorithm or empirical result is claimed; the contribution is the synthesis and its pedagogical packaging.
Load-bearing premise
The guide's value as a practical resource depends on its code examples actually implementing the techniques they describe, and the RNN hidden-state example in Section 4.3.1 already plots the model's output as though it were hidden states, so this premise is not fully met.
Editorial extensions
If this is right
- A reader who works through the guide is expected to be able to apply SHAP, LIME, and related methods to tabular, image, and text models, and to distinguish local from global explanations.
- The book implies that interpretability is not one property: transparency, interpretability, explainability, and fairness are related but distinct, so evaluation should use multiple metrics.
- LLM interpretability is treated as an extension of existing machinery (gradient attribution, probing classifiers, attention analysis) rather than a separate field, so skills transfer from classical models to BERT, GPT, and T5.
- Case studies in healthcare, finance, and policy are presented as evidence that XAI methods have practical decision-support value beyond model debugging.
Reading between the lines
- The guide's own Section 4.3.1 example plots the model's prediction output as if it were hidden-state activations; if similar mismatches occur elsewhere, readers following the code would internalize a misleading picture of what the code computes.
- A natural extension the author leaves implicit is a curated, testable companion suite that verifies each code example against the method it claims to illustrate, especially the RNN and attention examples.
- The synthesis suggests a research hypothesis: a practitioner trained on classical post-hoc methods can transfer that training to LLM interpretability faster than one starting directly with LLM-specific tools; this could be tested with a controlled learning study.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a book-length survey/guide to Explainable AI, covering theoretical foundations, interpretability of classical models (decision trees, linear models, SVMs, rule-based systems, GAMs, Bayesian models), deep learning models (CNNs, RNNs, Transformers), large language models (BERT, GPT, T5, LLaMA), a wide range of XAI techniques (SHAP, LIME, Integrated Gradients, Grad-CAM, LRP, counterfactuals, causal methods, graph and multimodal methods), applications, evaluation metrics, tools, and future directions. The stated goal, in the Abstract and §1.4, is to provide a comprehensive and practical guide that 'bridges theory and practice' through Python code examples and a companion GitHub repository.
Significance. If the code examples and organizational structure were reliable, this would be a useful educational resource: the breadth is substantial, and the practical emphasis with runnable Python snippets for SHAP, LIME, Integrated Gradients, DeepLIFT, Grad-CAM, LRP, and several counterfactual methods is a genuine strength. The book also gives credit to many standard references and includes a counterpart GitHub repository, which supports reproducibility in principle. However, the central value proposition—a hands-on, trustworthy guide—depends critically on the correctness of the code and the coherence of the presentation, and both are currently undermined by the errors described below.
major comments (3)
- [§4.3.1, lines 24–25] The code defines a SimpleRNN followed by a Dense layer (lines 15–18) and then computes both `y_pred` and `hidden_states` as `model.predict(X)` (lines 24–25). These are identical calls, so `hidden_states` is the Dense output, not the RNN hidden-state activations. Concretely, `hidden_states.shape` is (1, 100, 1), so the loop `for i in range(hidden_states.shape[-1])` draws only one trace, not the activations of all 10 hidden units as the figure caption claims. A reader following the code would not learn how to visualize hidden states; they would see the model's prediction plotted under a misleading label. This error directly contradicts the promise in §1.4 of providing 'hands-on understanding' and must be corrected, for example by building a separate model that outputs `model.layers[0].output`.
- [§3.2] The subsection titled 'Advantages and Disadvantages of Logistic Regression' appears in the middle of the Decision Trees section, immediately after the discussion of pruning and before the decision-tree feature-importance code example. This is a structural error: the content belongs in §3.3, which already contains a logistic regression discussion, and its placement here is confusing for a reader using the book as a reference. The guide's claim to be comprehensive and reliable is weakened by such organization mistakes, and this passage should be relocated or removed as a duplicate.
- [§1.2 vs. §5.4.3] The book gives inconsistent definitions of its two central terms. In §1.2, 'interpretability' is defined as 'the degree to which a human can understand the cause of a decision,' and 'explainability' is defined as 'the extent to which the internal mechanics of a machine learning model can be understood,' with explainability said to go further by focusing on 'why.' In §5.4.3, however, 'interpretability' is described as understanding internal workings (e.g., neurons and layers), while 'explainability' is described as providing human-understandable reasons such as feature importance or visualizations. These are essentially swapped. Because the book explicitly introduces these terms as foundational concepts, the inconsistency is pedagogically misleading and should be reconciled in a revision.
minor comments (4)
- [§2.4.1, code listing] Lines 8 and 14 of the SHAP example appear in the manuscript text as bare sentences ('Load the dataset', 'Split the dataset...') without the '#' comment prefix; if reproduced as written, the code is syntactically invalid. Please ensure all comment lines are properly prefixed in the printed code.
- [§5.6.2] The text states that the layer-activation plot shows 'all 13 layers of BERT,' but BERT-base has 12 transformer layers; the 13 hidden states returned by `output_hidden_states=True` include the embedding layer. The wording should be clarified to avoid confusion.
- [§4.4.3] The attention-heatmap example uses a synthetic 3×3 weight matrix, but the preceding text discusses the phrase 'The cat sat' and alignment in machine translation; the figure caption and the result explanation should make clear that the heatmap is illustrative and not computed from a real model on that phrase.
- [§5.2, GPT-4 bullet] The claim that GPT-4 supports multimodal inputs is referenced to a generic citation [102]; a primary source or clearer specification of the model version would be more helpful to readers.
Circularity Check
No circular reasoning found; the guide is a survey-style resource with no derivation chain whose outputs reduce to its inputs.
full rationale
This manuscript is a comprehensive survey and educational guide to explainable AI, not a research paper making novel predictions or deriving results from fitted parameters. The central claims are descriptive promises about coverage and pedagogy, and the content consists of textbook explanations, literature summaries, and illustrative code examples. I checked for the seven circularity patterns: there is no parameter fitted to a subset and then presented as a prediction, no equation defined in terms of the quantity it is said to explain, no load-bearing self-citation chain, and no imported uniqueness theorem forcing a particular choice. The cited works are standard external references (SHAP, LIME, Grad-CAM, BERT, etc.) that support the survey independently. I also noted that Section 4.3.1 contains a concrete implementation flaw: the code sets both `y_pred` and `hidden_states` to `model.predict(X)`, so the figure labeled "Activations of All Hidden Units Over Time" actually plots the Dense-layer output rather than the RNN hidden-state activations. This is a correctness and pedagogical quality defect that undermines the book's practical usefulness, but it is not a circularity: the mistake is a mismatch between code and narrative, not a reduction of a claimed result to its own input. Therefore, the appropriate circularity score is 0.
Assumptions & free parameters
Cite this review
Pith. "Pith review of A Comprehensive Guide to Explainable AI: From Classical Models to LLMs." pith.science (2026). https://pith.science/paper/5ZJ4UUMV
@misc{pith2026241200800,
author = {Pith},
title = {Pith review of: A Comprehensive Guide to Explainable AI: From Classical Models to LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/5ZJ4UUMV}},
note = {Machine review of arXiv:2412.00800}
}
read the original abstract
Explainable Artificial Intelligence (XAI) addresses the growing need for transparency and interpretability in AI systems, enabling trust and accountability in decision-making processes. This book offers a comprehensive guide to XAI, bridging foundational concepts with advanced methodologies. It explores interpretability in traditional models such as Decision Trees, Linear Regression, and Support Vector Machines, alongside the challenges of explaining deep learning architectures like CNNs, RNNs, and Large Language Models (LLMs), including BERT, GPT, and T5. The book presents practical techniques such as SHAP, LIME, Grad-CAM, counterfactual explanations, and causal inference, supported by Python code examples for real-world applications. Case studies illustrate XAI's role in healthcare, finance, and policymaking, demonstrating its impact on fairness and decision support. The book also covers evaluation metrics for explanation quality, an overview of cutting-edge XAI tools and frameworks, and emerging research directions, such as interpretability in federated learning and ethical AI considerations. Designed for a broad audience, this resource equips readers with the theoretical insights and practical skills needed to master XAI. Hands-on examples and additional resources are available at the companion GitHub repository: https://github.com/Echoslayer/XAI_From_Classical_Models_to_LLMs.
Figures
Figures from the paper (47 more)
Reference graph
Works this paper leans on
-
[1]
Jordan and Tom M
Michael I. Jordan and Tom M. Mitchell. Machine learning: Trends, perspectives, and prospects. Science, 349(6245):255–260, 2015
2015
-
[5]
Gilpin, David Bau, Ben Z
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA), pages 80–89. IEEE, 2018
2018
-
[6]
General data protection regulation (gdpr)
European Union. General data protection regulation (gdpr). Official Journal of the European Union, L119:1–88, 2016
2016
-
[7]
European union regulations on algorithmic decision- making and a â ˘AIJright to explanationâ ˘A˙I
Bryce Goodman and Seth Flaxman. European union regulations on algorithmic decision- making and a â ˘AIJright to explanationâ ˘A˙I. AI Magazine, 38(3):50–57, 2017
2017
-
[10]
Solon Barocas and Andrew D. Selbst. Big data’s disparate impact. California Law Review , 104:671–732, 2016
2016
-
[11]
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta et al. Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion, 58:82–115, 2020
2020
-
[12]
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 267:1–38, 2019
2019
-
[13]
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning. Springer, 2nd edition, 2009
2009
Show all 291 references
-
[14]
Ross Quinlan
J. Ross Quinlan. Induction of decision trees. Machine Learning, 1(1):81–106, 1986. 233 234 BIBLIOGRAPHY
1986
-
[15]
To predict and serve? Significance, 13(5):14–19, 2016
Kristian Lum and William Isaac. To predict and serve? Significance, 13(5):14–19, 2016
2016
-
[16]
Support-vector networks
Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning, 20(3):273– 297, 1995
1995
-
[17]
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In European Conference on Computer Vision, pages 818–833. Springer, 2014
2014
-
[18]
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Y oshua Bengio. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations (ICLR), 2015
2015
-
[19]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186, 2019
2019
-
[20]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. OpenAI Preprint, 2018
2018
-
[21]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Y anqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020
2020
-
[22]
Unified framework for interpretable methods
Scott M Lundberg and Su-In Lee. Unified framework for interpretable methods. In Advances in Neural Information Processing Systems, pages 4768–4777, 2017
2017
-
[23]
â ˘AIJwhy should i trust you?â ˘A˙I: Ex- plaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. â ˘AIJwhy should i trust you?â ˘A˙I: Ex- plaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144, 2016
2016
-
[24]
Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient- based localization. In Proceedings of the IEEE International Conference on Computer Vision , pages 618–626, 2017
2017
-
[25]
A survey on explainable artificial intelligence (xai): Toward medical xai
Ernesta Tjoa and Chuan Guan. A survey on explainable artificial intelligence (xai): Toward medical xai. IEEE Transactions on Neural Networks and Learning Systems, 2020
2020
-
[26]
Hager, and Kaitlyn Confundus
Narine Kokhlikyan, Vivek Miglani, Michael Martin, Edward Wang, Jacek Reynolds, Alessandro Meloni, Natalia Dotta, Babak Schreiber, Diego Garcia, Wenbing Tang, Ariel Ramos, Meghana Rege, Patrick R. Hager, and Kaitlyn Confundus. Captum: A unified and generic model inter- pretabil...
2009 arXiv
-
[28]
MIT press, 2016
Ian Goodfellow, Y oshua Bengio, and Aaron Courville.Deep learning. MIT press, 2016
2016
-
[29]
A survey of methods for explaining black box models.ACM computing surveys (CSUR), 51(5):1–42, 2018
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. A survey of methods for explaining black box models.ACM computing surveys (CSUR), 51(5):1–42, 2018. BIBLIOGRAPHY 235
2018
-
[30]
Novoa, Justin Ko, Susan M
Andre Esteva, Brett Kuprel, Roberto A. Novoa, Justin Ko, Susan M. Swetter, Helen M. Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639):115–118, 2017
2017
-
[31]
Khandani, Adlar J
Amir E. Khandani, Adlar J. Kim, and Andrew W. Lo. Consumer credit-risk models via machine- learning algorithms. Journal of Banking & Finance, 34(11):2767–2787, 2010
2010
-
[32]
Frey, Joshua P
Anand Chandrasekaran, Nathan C. Frey, Joshua P . Duncan, David Ha, Scott E. Hudson, and Jennifer Mankoff. Explainable ai for designers: A human-centered perspective on mixed- initiative co-creation. In 2018 IEEE Conference on Artificial Intelligence and Virtual Reality (AIVR),...
2018
-
[33]
Why a right to explanation of automated decision-making does not exist in the general data protection regulation
Sandra Wachter, Brent Mittelstadt, and Luciano Floridi. Why a right to explanation of automated decision-making does not exist in the general data protection regulation. International Data Privacy Law, 7(2):76–99, 2017
2017
-
[34]
Interpretable Machine Learning
Christoph Molnar. Interpretable Machine Learning. Lulu. com, 2020
2020
-
[35]
Imagenet classification with deep con- volutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep con- volutional neural networks. In Advances in Neural Information Processing Systems , pages 1097–1105, 2012
2012
-
[36]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, pages 5998–6008, 2017
2017
-
[37]
Methods for interpreting and understanding deep neural networks
Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. Methods for interpreting and understanding deep neural networks. Digital Signal Processing, 73:1–15, 2018
2018
-
[38]
Visual analytics for explainable deep learning
Jaegul Choo and Shixia Liu. Visual analytics for explainable deep learning. IEEE Computer Graphics and Applications, 38(4):84–92, 2018
2018
-
[39]
Ronald A. Fisher. The use of multiple measurements in taxonomic problems. Annals of Eugen- ics, 7(2):179–188, 1936
1936
-
[40]
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. A survey of methods for explaining black box models. ACM Computing Surveys (CSUR), 51(5):1–42, 2018
2018
-
[41]
Zachary C. Lipton. The mythos of model interpretability. arXiv preprint arXiv:1606.03490, 2016
2016 arXiv
-
[42]
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5):206–215, 2019
2019
-
[43]
Ross Quinlan
J. Ross Quinlan. Learning decision trees. Machine Learning, 1(1), 1996
1996
-
[44]
Regression shrinkage and selection via the lasso
Robert Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58(1):267–288, 1996
1996
-
[45]
Random forests
Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001
2001
-
[46]
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , pages 785–794, 2016. 236 BIBLIOGRAPHY
2016
-
[47]
Deep learning.Nature, 521(7553):436–444, 2015
Y ann LeCun, Y oshua Bengio, and Geoffrey Hinton. Deep learning.Nature, 521(7553):436–444, 2015
2015
-
[48]
Decision trees and multivariate analysis, volume 1
J Ross Quinlan. Decision trees and multivariate analysis, volume 1. Springer, 1986
1986
-
[49]
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,...
2015
-
[50]
Explainable artificial intel- ligence: Understanding, visualizing and interpreting deep learning models
Wojciech Samek, Thomas Wiegand, and Klaus-Robert Müller. Explainable artificial intel- ligence: Understanding, visualizing and interpreting deep learning models. arXiv preprint arXiv:1708.08296, 2017
2017 arXiv
-
[51]
Logistic Regression Explained
Sachin Narkhede. Logistic Regression Explained. Self-published, 2021
2021
-
[52]
Applied Logistic Regression
David W Hosmer et al. Applied Logistic Regression. Wiley, 2013
2013
-
[53]
Pattern recognition and machine learning
Christopher M Bishop. Pattern recognition and machine learning. Springer, 2013
2013
-
[54]
An Introduction to Statis- tical Learning
Gareth James, Daniela Witten, Trevor Hastie, and Robert Tibshirani. An Introduction to Statis- tical Learning. Springer, 2013
2013
-
[55]
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014
2014
-
[56]
Scikit- learn: Machine learning in python
Fabian Pedregosa, Gael Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit- learn: Machine learning in python. Journal of Machine Learning Research , 12:2825–2830, 2011
2011
-
[57]
James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu
W. James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. Defini- tions, methods, and applications in interpretable machine learning.Proceedings of the National Academy of Sciences, 116(44):22071–22080, 2019
2019
-
[58]
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach . Prentice Hall, 3rd edition, 2010
2010
-
[59]
McCormick, and David Madigan
Benjamin Letham, Cynthia Rudin, Tyler H. McCormick, and David Madigan. Interpretable clas- sifiers using rules and bayesian analysis: Building a better stroke prediction model.The Annals of Applied Statistics, 9(3):1350–1371, 2015
2015
-
[60]
Rule-based machine learning and knowledge extraction for model interpretability
Rok Piltaver, Mitja LuÅ ˛ atrek, MatjaÅ¿ Gams, and Denis Ä ˇRonlagiÄ ˘G. Rule-based machine learning and knowledge extraction for model interpretability. In 2016 IEEE 14th International Symposium on Applied Machine Intelligence and Informatics (SAMI) , pages 137–142. IEEE, 2016
2016
-
[61]
Rule-based systems
Lukasz Kurgan and Petr Musilek. Rule-based systems. IEEE Potentials, 20(2):10–15, 2011
2011
-
[62]
Emerging trends and challenges in rule-based machine learning
Tom M Mitchell. Emerging trends and challenges in rule-based machine learning. Frontiers of Computer Science, 12(4):488–501, 2018
2018
-
[63]
Biomedical Informatics: Computer Applications in Health Care and Biomedicine
Edward H Shortliffe and James J Cimino. Biomedical Informatics: Computer Applications in Health Care and Biomedicine. Springer, 2014. BIBLIOGRAPHY 237
2014
-
[64]
Expert systems in medical applications: a review
M Durairaj and V Ranjani. Expert systems in medical applications: a review. International Journal of Advanced Research in Computer Science and Software Engineering, 5(4):450–454, 2015
2015
-
[65]
Legal reasoning and legal argumentation.The Knowl- edge Engineering Review, 27(1):1–5, 2012
Trevor Bench-Capon and Henry Prakken. Legal reasoning and legal argumentation.The Knowl- edge Engineering Review, 27(1):1–5, 2012
2012
-
[66]
The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature
Eric WT Ngai, Yunan Hu, Yijun Wong, Yimin Chen, and Xin Sun. The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature. Decision Support Systems, 50(3):559–569, 2011
2011
-
[67]
Guidelines for a knowledge-based systems paper
Tomáš Kliegr, Št ˇepán Bahnik, and Jürgen Förster. Guidelines for a knowledge-based systems paper. Knowledge-Based Systems, 154:136–146, 2018
2018
-
[68]
Generalized Additive Models
Trevor J Hastie and Robert J Tibshirani. Generalized Additive Models. Routledge, 2017
2017
-
[69]
Generalized Additive Models: An Introduction with R
Simon N Wood. Generalized Additive Models: An Introduction with R. CRC Press, 2017
2017
-
[70]
Intelligible models for classification and re- gression
Yin Lou, Rich Caruana, and Johannes Gehrke. Intelligible models for classification and re- gression. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 150–158, 2012
2012
-
[71]
pygam: Generalized additive models in python
Daniel Servén and Charles D Brummitt. pygam: Generalized additive models in python. Zen- odo, 2018
2018
-
[72]
Intel- ligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noga Elhadad. Intel- ligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mini...
2015
-
[73]
Interpretable machine learning: The game perspective
Ruijiang Chen, Ying Shi, Hanyu Li, Y ongyi Y ang, and Xin Zhou. Interpretable machine learning: The game perspective. arXiv preprint arXiv:1809.06512, 2018
2018 arXiv
-
[74]
Gam: The predictive modeling silver bullet
Thomas Lin Pedersen. Gam: The predictive modeling silver bullet. Blog Post, 2019
2019
-
[75]
Frequentist model averaging of generalized additive models
Anne H Petersen et al. Frequentist model averaging of generalized additive models. arXiv preprint arXiv:1903.06066, 2019
1903 arXiv
-
[76]
Machine learning: a probabilistic perspective
Kevin P Murphy. Machine learning: a probabilistic perspective. MIT press, 2012
2012
-
[77]
Bayesian data analysis
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, and Donald B Rubin. Bayesian data analysis. Chapman and Hall/CRC, 2013
2013
-
[78]
Pattern Recognition and Machine Learning
Christopher M Bishop. Pattern Recognition and Machine Learning. Springer, 2013
2013
-
[79]
Tensorflow distributions
Joshua V Dillon, Ian Langmore, Dustin Tran, Eugene Brevdo, Sundaram Vasudevan, David Moore, Andrew Patton, Alexander Alemi, Matthew D Hoffman, and Rif A Saurous. Tensorflow distributions. arXiv preprint arXiv:1711.10604, 2017
2017 arXiv
-
[80]
Handbook of Markov Chain Monte Carlo
Steve Brooks, Andrew Gelman, Galin Jones, and Xiao-Li Meng. Handbook of Markov Chain Monte Carlo. CRC press, 2011
2011
-
[81]
Bayesian portfolio analysis
Doron Avramov and Guofu Zhou. Bayesian portfolio analysis. Annual Review of Financial Economics, 2:25–47, 2010. 238 BIBLIOGRAPHY
2010
-
[82]
Online controlled experiments and a/b testing
Ron Kohavi and Stefan Thomke. Online controlled experiments and a/b testing. InEncyclopedia of Machine Learning and Data Science, pages 922–929. Springer, 2017
2017
-
[83]
Variational inference: A review for statisti- cians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe. Variational inference: A review for statisti- cians. Journal of the American Statistical Association, 112(518):859–877, 2017
2017
-
[84]
Streaming variational bayes
Tamara Broderick, Nicholas Boyd, Andre Wibisono, Emmanuel Candes, and Michael I Jordan. Streaming variational bayes. In Advances in Neural Information Processing Systems , pages 1727–1735, 2013
2013
-
[85]
Representation learning: A review and new perspectives
Y oshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence , 35(8):1798–1828, 2013
2013
-
[86]
Understanding neu- ral networks through deep visualization
Jason Y osinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson. Understanding neu- ral networks through deep visualization. In Deep Learning Workshop, International Conference on Machine Learning, 2015
2015
-
[87]
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015
2015
-
[88]
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. In Workshop at International Con- ference on Learning Representations, 2014
2014
-
[89]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Y an. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning , pages 3319–3328, 2017
2017
-
[90]
Critical review of recurrent neural net- works for sequence learning
Zachary C Lipton, John Berkowitz, and Charles Elkan. Critical review of recurrent neural net- works for sequence learning. arXiv preprint arXiv:1506.00019, 2015
2015 arXiv
-
[91]
Visualizing and understanding recurrent net- works
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. Visualizing and understanding recurrent net- works. arXiv preprint arXiv:1506.02078, 2015
2015 arXiv
-
[92]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018
2018 arXiv
-
[93]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Y anqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019
1910 arXiv
-
[94]
A multiscale visualization of attention in the transformer model
Jesse Vig. A multiscale visualization of attention in the transformer model. arXiv preprint arXiv:1906.05714, 2019
1906 arXiv
-
[95]
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020
1901
-
[96]
Large Lan- guage Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges
Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, Keyu Chen, and Ming Li. Large Lan- guage Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges. arXiv, Online, 2024. Preprint available on arXiv. BIBLIOGRAPHY 239
2024
-
[97]
Design of intelligent customer service system based on deep learning
Peng Zhou, Guoxin Jin, and Hua Liu. Design of intelligent customer service system based on deep learning. Journal of Physics: Conference Series, 1486(3):032036, 2020
2020
-
[98]
Ctrl: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. Ctrl: A conditional transformer language model for controllable generation. In arXiv preprint arXiv:1909.05858, 2019
1909 arXiv
-
[99]
Github copilot: Y our ai pair programmer
GitHub. Github copilot: Y our ai pair programmer. https://copilot.github.com/, 2021
2021
-
[100]
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Y oon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240, 2020
2020
-
[101]
Squad: 100,000+ ques- tions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ ques- tions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016
2016 arXiv
-
[102]
Gpt-4 technical report
OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[103]
Llama: Open and efficient foundation language models
Hugo Touvron et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[104]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[105]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[106]
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645–3650, 2019
2019
-
[107]
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S M Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212, 2021
2021
-
[108]
Green ai
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai. Communications of the ACM, 63(12):54–63, 2020
2020
-
[109]
On the dangers of stochastic parrots: Can language models be too big?Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big?Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021
2021
-
[110]
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
Benji Peng, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Junyu Liu, and Qian Niu. Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks. arXiv, Online,
-
[111]
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Peter Clark, Lianhui Qin, Ali Farhadi, and Y ejin Choi. Defending against neural fake news. InAdvances in Neural Information Processing Systems, pages 9051– 9062, 2019. 240 BIBLIOGRAPHY
2019
-
[112]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016
2016
-
[113]
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Connor Leahy, and Kyle McDonell. Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow. arXiv preprint arXiv:2108.12409, 2021
2021 arXiv
-
[114]
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022
2022 arXiv
-
[115]
Analysis methods in neural language processing: A survey
Y onatan Belinkov and James Glass. Analysis methods in neural language processing: A survey. In Transactions of the Association for Computational Linguistics, volume 7, pages 49–72, 2019
2019
-
[116]
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau and Douwe Kiela. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, pages 2126–2136, 2018
2018
-
[117]
What does bert learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651–3657, 2019
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. What does bert learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651–3657, 2019
2019
-
[118]
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 5797–...
2019
-
[119]
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceed- ings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages 2733–2743, 2019
2019
-
[120]
Prompt programming for large language models: Beyond the few-shot paradigm
Lucas Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. arXiv preprint arXiv:2102.07350, 2021
2021 arXiv
-
[121]
How can we know what language models know? In Transactions of the Association for Computational Linguistics , volume 8, pages 423–438, 2020
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know? In Transactions of the Association for Computational Linguistics , volume 8, pages 423–438, 2020
2020
-
[122]
Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pages 2463–2473, 2019
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pages 2463–2473, 2019
2019
-
[123]
Calibration of pre-trained transformers
Saahil Desai and Greg Durrett. Calibration of pre-trained transformers. arXiv preprint arXiv:2003.07892, 2020
2003 arXiv
-
[124]
OpenAI. Chatgpt. https://chat.openai.com/, 2023
2023
-
[125]
Claude: An ai assistant for your tasks
Anthropic. Claude: An ai assistant for your tasks. https://www.anthropic.com/, 2023
2023
-
[126]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022. BIBLIOGRAPHY 241
2022 arXiv
-
[127]
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Ali Najafi, David Bau, Y onatan Belinkov, and James Glass. What is one grain of sand in the desert? analyzing individual neurons in deep nlp models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6309–6317, 2019
2019
-
[128]
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings
Kawin Ethayarajh. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pages 55–65, 2019
2019
-
[129]
The mythos of model interpretability
Zachary C Lipton. The mythos of model interpretability. Communications of the ACM , 61(10):36–43, 2018
2018
-
[130]
Applied linear regression
Sanford Weisberg. Applied linear regression. John Wiley & Sons, 2014
2014
-
[131]
Applied logistic regression
David W Hosmer Jr, Stanley Lemeshow, and Rodney X Sturdivant. Applied logistic regression. John Wiley & Sons, 2013
2013
-
[132]
Data mining with decision trees: theory and applications
Lior Rokach and Oded Maimon. Data mining with decision trees: theory and applications . World scientific, 2014
2014
-
[133]
Research on rule-based expert system and application in fault diagnosis
Jing Wang and Wei Wang. Research on rule-based expert system and application in fault diagnosis. International Journal of Computer and Electrical Engineering, 9(5):401–407, 2017
2017
-
[134]
Introduction to machine learning: k-nearest neighbors
Zhi Zhang. Introduction to machine learning: k-nearest neighbors. Annals of translational medicine, 5(10):101, 2017
2017
-
[135]
A user’s guide to support vector machines
Asa Ben-Hur and Jason Weston. A user’s guide to support vector machines. In Data mining techniques for the life sciences, pages 223–239. Springer, 2010
2010
-
[136]
Ensemble methods: foundations and algorithms
Zhi-Hua Zhou. Ensemble methods: foundations and algorithms. CRC press, 2012
2012
-
[137]
A comprehensive survey on graph neural networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems, 32(1):4–24, 2021
2021
-
[138]
Making tree ensembles interpretable: A bayesian model selection approach
Satoshi Hara and Kohei Hayashi. Making tree ensembles interpretable: A bayesian model selection approach. In International Conference on Artificial Intelligence and Statistics , pages 77–85. PMLR, 2018
2018
-
[139]
Regression shrinkage and selection via the lasso: a retrospective
Robert Tibshirani. Regression shrinkage and selection via the lasso: a retrospective. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(3):273–282, 2011
2011
-
[140]
Towards robust interpretability with self-explaining neural networks
David Alvarez-Melis and Tommi S Jaakkola. Towards robust interpretability with self-explaining neural networks. In Advances in Neural Information Processing Systems , pages 7786–7795, 2018
2018
-
[141]
Concept bottleneck models
Pang Wei Koh, Ankur Suresh, Andrew Angus, Thao Nguyen, Y ew Siang Basu, Jure Leskovec Tatsunori B Hashimoto, and Kai Li. Concept bottleneck models. In International Conference on Machine Learning, pages 5338–5348. PMLR, 2020
2020
-
[142]
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Glorot-Xavier, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-vae: Learning basic visual concepts with a constrained variational framework. In 5th International Conference on Learning Representa- tion...
2017
-
[143]
Understanding variable importances in forests of randomized trees
Gilles Louppe, Louis Wehenkel, Antonio Sutera, and Pierre Geurts. Understanding variable importances in forests of randomized trees. In Advances in Neural Information Processing Systems, pages 431–439, 2013
2013
-
[145]
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus- Robert Müller. Explaining nonlinear classification decisions with deep taylor decomposition. In Pattern Recognition, pages 211–230. Springer, 2017
2017
-
[146]
Uci machine learning repository
Dheeru Dua and Casey Graff. Uci machine learning repository. http://archive.ics.uci. edu/ml, 2019
2019
-
[147]
An analysis of the iris dataset using decision tree classifier
Li Zhao and Ming Jin. An analysis of the iris dataset using decision tree classifier. Journal of Machine Learning, 9(2):85–94, 2018
2018
-
[148]
Permutation importance: a corrected feature importance measure
Andre Altmann, Laura Tolo¸ si, Oliver Sander, and Thomas Lengauer. Permutation importance: a corrected feature importance measure. Bioinformatics, 26(10):1340–1347, 2010
2010
-
[149]
Correlation and variable impor- tance in random forests
Baptiste Gregorutti, Bertrand Michel, and Philippe Saint-Pierre. Correlation and variable impor- tance in random forests. Statistics and Computing, 27(3):659–678, 2017
2017
-
[150]
Towards explainable artificial intelligence , pages 5–22
Wojciech Samek and Klaus-Robert Müller. Towards explainable artificial intelligence , pages 5–22. Springer, 2019
2019
-
[151]
Lundberg, Bala Nair, Monica S
Scott M. Lundberg, Bala Nair, Monica S. Vavilala, Mayumi Horibe, Matthew J. Eisses, Trevor Adams, David E. Liston, Daniel K.-W. Low, Soren-Frederik Newman, Jason Kim, and Su-In Lee. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Natur...
2018
-
[152]
Gradient-based attribution methods
Marco Ancona, Cengiz Oztireli, and Markus Gross. Gradient-based attribution methods. In Ex- plainable AI: Interpreting, Explaining and Visualizing Deep Learning, pages 169–191. Springer, 2019
2019
-
[153]
Evaluating the explainability of attention-based lstm models using shap values
Tong Li, Ming Ding, and Zhanyu Sun. Evaluating the explainability of attention-based lstm models using shap values. arXiv preprint arXiv:2004.05587, 2020
2004 arXiv
-
[154]
50 Y ears of Game Theory
Lloyd S Shapley. 50 Y ears of Game Theory. World Scientific, 2016
2016
-
[155]
Prob- lems with shapley-value-based explanations as feature importance measures
Ira Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, and Sorelle A Friedler. Prob- lems with shapley-value-based explanations as feature importance measures. arXiv preprint arXiv:2002.11097, 2020
2002 arXiv
-
[156]
A practical guide to explainable ai and lime
Marcos Garcia et al. A practical guide to explainable ai and lime. In Proceedings of the Twenty- Eighth International Joint Conference on Artificial Intelligence, pages 5735–5737, 2020
2020
-
[157]
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142– 150, 2011....
2011
-
[158]
On the limitations of local interpretability in explaining feature importance
Muhammad Bilal Zafar, , et al. On the limitations of local interpretability in explaining feature importance. arXiv preprint arXiv:1906.11197, 2019
1906 arXiv
-
[159]
Allennlp interpret: A framework for explaining predictions of nlp models
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Allennlp interpret: A framework for explaining predictions of nlp models. arXiv preprint arXiv:1909.09251, 2019
1909 arXiv
-
[160]
Mnist handwritten digit database
Y ann LeCun, Corinna Cortes, and CJ Burges. Mnist handwritten digit database. In ATT Labs [Online], 2010
2010
-
[161]
Smooth- grad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. Smooth- grad: removing noise by adding noise. In ICML Workshop on Visualization for Deep Learning, 2017
2017
-
[162]
Grad-cam: Visual explanations from deep networks via gradient- based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient- based localization. In Proceedings of the IEEE International Conference on Computer Vision , pages 618–626, 2017
2017
-
[163]
Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks
Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth N Balasubramanian. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 839–847. IEEE, 2018
2018
-
[164]
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014
2014 arXiv
-
[165]
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf. Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 782–791, 2021
2021
-
[166]
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS one, 10(7):e0130140, 2015
2015
-
[167]
Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation
Alex Goldstein, Adam Kapelner, Justin Bleich, and Emil Pitkin. Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation. Journal of Com- putational and Graphical Statistics, 24(1):44–65, 2015
2015
-
[168]
Visualizing the effects of predictor variables in black box su- pervised learning models
Daniel W Apley and Jingyu Zhu. Visualizing the effects of predictor variables in black box su- pervised learning models. Journal of the Royal Statistical Society: Series B (Statistical Method- ology), 82(4):1059–1086, 2020
2020
-
[169]
All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simul- taneously
Aaron Fisher, Cynthia Rudin, and Francesca Dominici. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simul- taneously. Journal of Machine Learning Research, 20(177):1–81, 2019
2019
-
[171]
pdp: An r package for constructing partial dependence plots
Brandon M Greenwell. pdp: An r package for constructing partial dependence plots. The R Journal, 9(1):421–436, 2017. 244 BIBLIOGRAPHY
2017
-
[172]
Deep learning for time series classification: a review
Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre- Alain Muller. Deep learning for time series classification: a review. Data Mining and Knowledge Discovery, 33(4):917–963, 2019
2019
-
[173]
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alejandro Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Ben- jamins, et al. Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and ...
2020
-
[174]
Timeshap: Explaining recurrent models through sequence perturbations
Eduardo Arango, Artur Luczak, and Peter Willett. Timeshap: Explaining recurrent models through sequence perturbations. arXiv preprint arXiv:2006.12019, 2020
2006 arXiv
-
[175]
A dual- stage attention-based recurrent neural network for time series prediction
Y ao Qin, Dongjin Song, Haifeng Chen, Wei Cheng, Guofei Jiang, and Garrison Cottrell. A dual- stage attention-based recurrent neural network for time series prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI-17) , pages 2627...
2017
-
[176]
Explaining recur- rent neural network predictions in sentiment analysis
Leila Arras, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. Explaining recur- rent neural network predictions in sentiment analysis. In Proceedings of the EMNLP Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis , pages 159–168, 2017
2017
-
[177]
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Y oshua Bengio. Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), 2015
2015
-
[178]
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 , 2013
2013 arXiv
-
[179]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Y an. Axiomatic attribution for deep networks. In International Conference on Machine Learning, pages 3319–3328, 2017
2017
-
[180]
Causal inference in statistics: A primer
Judea Pearl, Madelyn Glymour, and Nicholas P Jewell. Causal inference in statistics: A primer. John Wiley & Sons, 2016
2016
-
[181]
General approach to causal mediation analysis
Kosuke Imai, Luke Keele, and Dustin Tingley. General approach to causal mediation analysis. Psychological Methods, 15(4):309–334, 2010
2010
-
[182]
Invariant risk minimization
Martin Arjovsky et al. Invariant risk minimization. In arXiv preprint arXiv:1907.02893, 2019
1907 arXiv
-
[183]
Causal inference in statistics, social, and biomedical sciences
Peter Spirtes, Clark Glymour, and Richard Scheines. Causal inference in statistics, social, and biomedical sciences. Cambridge University Press, 2016
2016
-
[184]
Elements of causal inference: Foun- dations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. Elements of causal inference: Foun- dations and learning algorithms. MIT Press, 2017
2017
-
[185]
Explanation in causal inference: Methods for mediation and interaction
Tyler J VanderWeele. Explanation in causal inference: Methods for mediation and interaction . Oxford University Press, 2015
2015
-
[186]
Identification, inference, and sensitivity anal- ysis for causal mediation effects
Kosuke Imai, Luke Keele, and Teppei Y amamoto. Identification, inference, and sensitivity anal- ysis for causal mediation effects. Statistical Science, 25(1):51–71, 2010. BIBLIOGRAPHY 245
2010
-
[187]
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual explanations without opening the black box: Automated decisions and the gdpr. Harvard Journal of Law & Technol- ogy, 31:841, 2017
2017
-
[188]
A survey of algorithmic recourse: Contrastive explanations and consequential recommendations
Amir-Hossein Karimi, Bernhard Schölkopf, and Isabel Valera. A survey of algorithmic recourse: Contrastive explanations and consequential recommendations. arXiv preprint arXiv:2010.04050, 2020
2010 arXiv
-
[189]
Face: fea- sible and actionable counterfactual explanations
Rafael Poyiadzi, Kacper Sokol, Raul Santos-Rodriguez, Tijl De Bie, and Peter Flach. Face: fea- sible and actionable counterfactual explanations. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 344–350, 2020
2020
-
[190]
Generating counterfactual explanations with generative models
Lisa Schut, Frank Pijpers, Pieter van den Broek, and Peter Flach. Generating counterfactual explanations with generative models. Proceedings of the AAAI Conference on Artificial Intelli- gence, 35(8):8766–8774, 2021
2021
-
[191]
Efficient counterfactual explanations in kernel methods via mixed-integer programming
Théo Laugel, Marie-Jeanne Lesot, Christophe Marsala, Xavier Renard, and Marcin Detyniecki. Efficient counterfactual explanations in kernel methods via mixed-integer programming. In In- ternational Joint Conference on Artificial Intelligence (IJCAI), pages 136–142, 2018
2018
-
[192]
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Xiaojian Tao, Alice Barnett, Cynthia Rudin, and Zhi Su. This looks like that: deep learning for interpretable image recognition. In Advances in Neural Information Processing Systems, volume 32, pages 8930–8941, 2019
2019
-
[193]
Explaining machine learning classi- fiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan. Explaining machine learning classi- fiers through diverse counterfactual explanations. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 607–617, 2020
2020
-
[194]
Actionable recourse in linear classification
Berk Ustun, Alexander Spangher, and Y ang Liu. Actionable recourse in linear classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency , pages 10–19, 2019
2019
-
[195]
Counterfactual visual explanations
Y ash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. Counterfactual visual explanations. In Proceedings of the 36th International Conference on Machine Learning, pages 2376–2384, 2019
2019
-
[196]
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. Learning the difference that makes a difference with counterfactually-augmented data. In International Conference on Learning Representations, 2020
2020
-
[197]
Game theoretic k-nearest neighbor explanation
Jia-Huei Chang, Tsung-Hsien Chen, and Chung-Chi Wei. Game theoretic k-nearest neighbor explanation. In International Conference on Machine Learning, pages 1135–1144, 2018
2018
-
[198]
Conditional text generation for coun- terfactual explanations
Mukund Kumar, Alexander Levine, and Soheil Feizi Shah. Conditional text generation for coun- terfactual explanations. In Proceedings of the 58th Annual Meeting of the Association for Com- putational Linguistics, pages 79–93, 2020
2020
-
[199]
Language gans falling short
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin. Language gans falling short. In Advances in Neural Information Processing Systems, volume 31, pages 1–10, 2018. 246 BIBLIOGRAPHY
2018
-
[200]
Beyond backprop: Counterfactual training for interpretability
Roy Rodriguez, John Wohlwend, and Honglak Lee. Beyond backprop: Counterfactual training for interpretability. In Advances in Neural Information Processing Systems , volume 32, pages 10055–10065, 2019
2019
-
[201]
Generating natural language counter- factuals with minimal edits
Wenpeng Yin, Graham Neubig, and Emanuele Sachs. Generating natural language counter- factuals with minimal edits. arXiv preprint arXiv:2106.10112, 2021
2021 arXiv
-
[202]
Interpretable counterfactual explanations guided by prototypes
Arnaud Van Looveren and Jan Klaise. Interpretable counterfactual explanations guided by prototypes. arXiv preprint arXiv:1907.02584, 2019
1907 arXiv
-
[203]
Efficient search for diverse coherent explanations
Chris Russell. Efficient search for diverse coherent explanations. In Proceedings of the Con- ference on Fairness, Accountability, and Transparency, pages 20–28, 2019
2019
-
[204]
A survey of methods for explaining black box models
Riccardo Guidotti et al. A survey of methods for explaining black box models. ACM Computing Surveys, 51(5):93, 2019
2019
-
[205]
Explaining machine learning classifiers through diverse counterfactual explanations
Rudini K Mothilal, Amit Sharma, and Chenhao Tan. Explaining machine learning classifiers through diverse counterfactual explanations. In Proceedings of the 2020 Conference on Fair- ness, Accountability, and Transparency, pages 607–617, 2020
2020
-
[206]
Algorithmic recourse: from coun- terfactual explanations to interventions
Amir-Hossein Karimi, Bernhard Schölkopf, and Isabel Valera. Algorithmic recourse: from coun- terfactual explanations to interventions. arXiv preprint arXiv:2002.06278, 2020
2002 arXiv
-
[207]
Anchors: High-precision model- agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. Anchors: High-precision model- agnostic explanations. Proceedings of the AAAI Conference on Artificial Intelligence , 32(1), 2018
2018
-
[208]
Alon Jacovi and Y oav Goldberg. Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4198–4205, 2020
2020
-
[209]
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace et al. Universal adversarial triggers for attacking and analyzing nlp. In Empirical Methods in Natural Language Processing, pages 2153–2162, 2019
2019
-
[210]
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Vijil Iyengar, Ronny Luss, and Karthikeyan Shanmugam. Explanations based on the missing: Towards contrastive explanations with pertinent negatives. In Advances in Neural Information Processing Systems, pages 592–603, 2018
2018
-
[211]
Tabnet: Attentive interpretable tabular learning
Sercan Ö Arik and Tomas Pfister. Tabnet: Attentive interpretable tabular learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(8):6679–6687, 2021
2021
-
[212]
Semi-supervised classification with graph convolutional net- works
Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional net- works. International Conference on Learning Representations (ICLR), 2017
2017
-
[214]
Gnnexplainer: Generating explanations for graph neural networks
Rex Ying, Dylan Bourgeois, Jiaxuan Y ou, Marinka Zitnik, and Jure Leskovec. Gnnexplainer: Generating explanations for graph neural networks. In Advances in Neural Information Pro- cessing Systems (NeurIPS), pages 9240–9251, 2019. BIBLIOGRAPHY 247
2019
-
[215]
Graphsvx: Shapley value explanations for graph neural networks
Antoine Duval and Fragkiskos D Malliaros. Graphsvx: Shapley value explanations for graph neural networks. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 2823–2832, 2021
2021
-
[216]
Explainability methods for graph neural networks
Hao Yuan, Jiliang Tang, Xia Hu, and Shuiwang Ji. Explainability methods for graph neural networks. In Proceedings of the 30th International Joint Conference on Artificial Intelligence (IJCAI), pages 4561–4567, 2021
2021
-
[217]
Explainability techniques for graph convolutional networks
Francesco Baldassarre and Hossein Azizpour. Explainability techniques for graph convolutional networks. arXiv preprint arXiv:1905.13686, 2019
1905 arXiv
-
[218]
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Y oshua Bengio. Graph attention networks. In International Conference on Learning Representations (ICLR), 2018
2018
-
[219]
Inductive representation learning on large graphs
William L Hamilton, Rex Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems (NeurIPS), pages 1024–1034, 2017
2017
-
[220]
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. Proceedings of the 34th International Conference on Machine Learning (ICML), pages 1263–1272, 2017
2017
-
[221]
Knowledge graph embedding: A sur- vey of approaches and applications
Quan Wang, Zhendong Mao, Bin Wang, and Li Guo. Knowledge graph embedding: A sur- vey of approaches and applications. IEEE Transactions on Knowledge and Data Engineering, 29(12):2724–2743, 2017
2017
-
[222]
Graph neural networks for natural language processing: A survey
Wenxuan Huang and Minlie Huang. Graph neural networks for natural language processing: A survey. arXiv preprint arXiv:2012.15445, 2022
2012 arXiv
-
[223]
Fast graph representation learning with pytorch geometric
Matthias Fey and Jan Eric Lenssen. Fast graph representation learning with pytorch geometric. ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019
2019
-
[224]
Collective classification in network data
Prithviraj Sen, Galen Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi- Rad. Collective classification in network data. AI Magazine, 29(3):93–106, 2008
2008
-
[225]
Multimodal machine learn- ing: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learn- ing: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence , 41(2):423–443, 2018
2018
-
[226]
Visualbert: A simple and perfor- mant baseline for vision and language
Liunian Harold Li, Yunyi Su, Chunyuan Li Noah Sheng, Raghav Kulkarni, Katrina Szeto, Lugui Yu, Jianwei Feng, Devi Parikh, Y ejin Choi, and Jianfeng Gao. Visualbert: A simple and perfor- mant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019
1908 arXiv
-
[227]
Attention-based multimodal fusion for video description
Chiori Hori, Takaaki Hori, Tae-Hyun Lee, Zhuo Zhang, Benjamin Harsham, John R Hershey, Tim K Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In Proceedings of the IEEE international conference on computer vision, pages 4193–4202, 2017
2017
-
[228]
Vilbert: Pretraining task-agnostic visi- olinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visi- olinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems, volume 32, 2019. 248 BIBLIOGRAPHY
2019
-
[229]
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Y oshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014
2014 arXiv
-
[230]
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015
2015
-
[231]
Show, attend and tell: Neural image caption genera- tion with visual attention
Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Richard Zemel, and Y oshua Bengio. Show, attend and tell: Neural image caption genera- tion with visual attention. In International conference on machine learning, pages 2048–2057. PMLR, 2015
2015
-
[232]
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 conference on empirical methods in natural language processing, pages 1103–1114, 2017
2017
-
[233]
Multi- modal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. Multi- modal deep learning. In Proceedings of the 28th international conference on machine learning (ICML-11), pages 689–696, 2011
2011
-
[234]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Paul Luc, Antoine Miech, Ian Barr, Y ana Hasson, Sander Dieleman, Arthur Mensch, Katie Millican, Jack W Rae, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022
2022 arXiv
-
[235]
A multiscale visualization of attention in the transformer model
Jesse Vig and Y onatan Belinkov. A multiscale visualization of attention in the transformer model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 37–42, 2019
2019
-
[236]
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019
1904 arXiv
-
[237]
Attention is not explanation
Sarthak Jain and Byron C Wallace. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , pages 3543–3556, 2019
2019
-
[238]
What makes training multi-modal classifi- cation networks hard? arXiv preprint arXiv:2011.12558, 2020
Xingyi Wang, Huaishao Hu, Jinyu Li, and Dong Yu. What makes training multi-modal classifi- cation networks hard? arXiv preprint arXiv:2011.12558, 2020
2011 arXiv
-
[239]
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. Towards better understanding of gradient-based attribution methods for deep neural networks. InInternational Conference on Learning Representations, 2018
2018
-
[240]
Multi-modal graph neural networks for molecular property prediction
Xiong Wang, Ruocheng Guo, Jing Ma, Yunan Hu, Yu Gai, and Y anjun Qi. Multi-modal graph neural networks for molecular property prediction. Proceedings of the AAAI Conference on Artificial Intelligence, 34(04):10524–10531, 2020
2020
-
[241]
Multimodal routing: Improving information flow in multimodal language analysis
Y ao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. Multimodal routing: Improving information flow in multimodal language analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pages ...
2019
-
[242]
Visualizing the impact of feature attribution baselines
Payton Sturmfels, Scott Lundberg, and Su-In Lee. Visualizing the impact of feature attribution baselines. Distill, 5(1):e22, 2020
2020
-
[243]
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Lan- guage Processing, pages 5100–5111, 2019
2019
-
[244]
Vilt: Vision-and-language transformer without con- volution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without con- volution or region supervision. In International Conference on Machine Learning, pages 5583–
-
[245]
Heterogeneous graph transformer
Ziniu Hu, Yuxiao Dong, Kuansan Wang, and Kai Chang. Heterogeneous graph transformer. In Proceedings of the Web Conference 2020, pages 2704–2710, 2020
2020
-
[246]
Towards a rigorous science of interpretable machine learn- ing
Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learn- ing. arXiv preprint arXiv:1702.08608, 2017
2017 arXiv
-
[247]
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. The (un)reliability of saliency methods. In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, pages 267–280. Springer, 2019
2019
-
[248]
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. arXiv preprint arXiv:1908.09635, 2019
1908 arXiv
-
[249]
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference, pages 214–226, 2012
2012
-
[250]
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. Interpretation of neural networks is fragile. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3681–3688, 2019
2019
-
[251]
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nathan Srebro. Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems, pages 3315–3323, 2016
2016
-
[252]
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity checks for saliency maps. InAdvances in Neural Information Processing Systems, pages 9505–9515, 2018
2018
-
[253]
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In2016 IEEE Symposium on Security and Privacy, pages 582–597. IEEE, 2016
2016
-
[254]
Towards robust interpretability with self-explaining neural networks
Diego Alvarez-Melis and Tommi S Jaakkola. Towards robust interpretability with self-explaining neural networks. In Advances in Neural Information Processing Systems , pages 7786–7795, 2018
2018
-
[255]
Counterfactual fairness
Matt J Kusner et al. Counterfactual fairness. In Advances in Neural Information Processing Systems, pages 4066–4076, 2017
2017
-
[256]
On the sensitivity of adversarial robustness to input data distributions
Cheng He, Hailiang Huang, Bingsheng He, and Xiaokui Xiao. On the sensitivity of adversarial robustness to input data distributions. In Proceedings of the IEEE International Conference on Data Mining, pages 1050–1055. IEEE, 2019. 250 BIBLIOGRAPHY
2019
-
[257]
Data preprocessing techniques for classification without discrimination
Faisal Kamiran and Toon Calders. Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 33(1):1–33, 2012
2012
-
[258]
Fairness in machine learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness in machine learning. NIPS Tutorial, 2017
2017
-
[259]
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning, pages 3145–3153, 2017
2017
-
[260]
The woman worked as a babysitter: On biases in language generation
Emily Sheng et al. The woman worked as a babysitter: On biases in language generation. In Empirical Methods in Natural Language Processing, pages 3407–3412, 2019
2019
-
[261]
Learning fair repre- sentations
Richard Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork. Learning fair repre- sentations. In Proceedings of the 30th International Conference on Machine Learning , pages 325–333, 2013
2013
-
[262]
Peeking inside the black-box: A survey on explainable artificial intelligence (xai)
Amina Adadi and Mohammed Berrada. Peeking inside the black-box: A survey on explainable artificial intelligence (xai). IEEE Access, 6:52138–52160, 2018
2018
-
[263]
Inherent trade-offs in learning fair representations
Hongyu Zhao and Houtao Deng. Inherent trade-offs in learning fair representations. In Ad- vances in Neural Information Processing Systems, volume 32, pages 15675–15686, 2019
2019
-
[264]
Perturbation sensitivity anal- ysis to detect unintended model biases
Vinodkumar Prabhakaran, Kelly Ann Blount, and Karen Livescu. Perturbation sensitivity anal- ysis to detect unintended model biases. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 4722–4729, 2019
2019
-
[265]
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019
2019
-
[266]
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S Jaakkola. On the robustness of interpretability methods. arXiv preprint arXiv:1806.08049, 2018
2018 arXiv
-
[267]
Robustness to paraphrase in text classification
Kai Sun, Yiyan Zhang, Xiaojun Ren, and Xiaodan Wang. Robustness to paraphrase in text classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4455–4462, 2019
2019
-
[268]
Explaining and harnessing adver- sarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adver- sarial examples. In International Conference on Learning Representations, 2015
2015
-
[269]
The limitations of deep learning in adversarial settings
Nicolas Papernot, Patrick McDaniel, and Ananthram Swami. The limitations of deep learning in adversarial settings. In IEEE European Symposium on Security and Privacy , pages 372–387, 2016
2016
-
[270]
Evasion attacks against machine learning at test time
Battista Biggio et al. Evasion attacks against machine learning at test time. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases , pages 387–402, 2013
2013
-
[271]
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal Mian. Threat of adversarial attacks on deep learning in computer vision: A survey. IEEE Access, 6:14410–14430, 2018
2018
-
[272]
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy, pages 39–57, 2017. BIBLIOGRAPHY 251
2017
-
[273]
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International Conference on Machine Learning, pages 274–283, 2018
2018
-
[274]
Robustness may be at odds with accuracy
Dimitris Tsipras et al. Robustness may be at odds with accuracy. In International Conference on Learning Representations, 2019
2019
-
[275]
Is bert really robust? a strong baseline for natural language attack on text classifi- cation and entailment
Di Jin et al. Is bert really robust? a strong baseline for natural language attack on text classifi- cation and entailment. In AAAI Conference on Artificial Intelligence, pages 8018–8025, 2020
2020
-
[276]
Environment inference for invariant learning
Elliot Creager et al. Environment inference for invariant learning. In International Conference on Machine Learning, pages 2189–2200, 2021
2021
-
[277]
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks et al. Women also snowboard: Overcoming bias in captioning models. In European Conference on Computer Vision, pages 793–811, 2018
2018
-
[278]
The (un)reliability of saliency methods
Pieter-Jan Kindermans et al. The (un)reliability of saliency methods. arXiv preprint arXiv:1711.00867, 2019
2019 arXiv
-
[279]
Faithful and customizable explanations of black box models
Himabindu Lakkaraju et al. Faithful and customizable explanations of black box models. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society , pages 131–138, 2019
2019
-
[280]
Not just black and white: Understanding racial disparities in natural language processing
Harini Suresh and John Guttag. Not just black and white: Understanding racial disparities in natural language processing. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 146–156, 2021
2021
-
[281]
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Victor Veitch et al. Counterfactual invariance to spurious correlations: Why and how to pass stress tests. In Advances in Neural Information Processing Systems , pages 10387–10399, 2021
2021
-
[282]
Gradient starvation: A learning proclivity in neural networks
Mohammad Pezeshki et al. Gradient starvation: A learning proclivity in neural networks. In Advances in Neural Information Processing Systems, pages 1256–1268, 2020
2020
-
[283]
Explanation invariance: A measure of explanations’ consistency with domain knowledge
Luke Daniel et al. Explanation invariance: A measure of explanations’ consistency with domain knowledge. In AAAI Conference on Artificial Intelligence, pages 11453–11461, 2020
2020
-
[284]
Fairness constraints: Mechanisms for fair classification
Brian Hu Zhang et al. Fairness constraints: Mechanisms for fair classification. In Artificial Intelligence, Ethics, and Society, pages 77–83, 2018
2018
-
[285]
Fairness through causal awareness: Learning causal latent-variable mod- els for biased data
David Madras et al. Fairness through causal awareness: Learning causal latent-variable mod- els for biased data. In AAAI Conference on Artificial Intelligence, pages 6572–6579, 2019
2019
-
[286]
Avoiding discrimination through causal reasoning
Niki Kilbertus et al. Avoiding discrimination through causal reasoning. In Advances in Neural Information Processing Systems, pages 656–666, 2017
2017
-
[287]
Estimating individual treatment effect: Generalization bounds and algorithms
Uri Shalit et al. Estimating individual treatment effect: Generalization bounds and algorithms. In International Conference on Machine Learning, pages 3076–3085, 2017
2017
-
[288]
Pc-fairness: A unified framework for measuring causal fairness
Zhiqi Bu Wu et al. Pc-fairness: A unified framework for measuring causal fairness. In Advances in Neural Information Processing Systems, pages 3391–3401, 2019
2019
-
[289]
Ai fairness 360: An extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias
Rachel KE Bellamy et al. Ai fairness 360: An extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias. In arXiv preprint arXiv:1810.01943, 2018. 252 BIBLIOGRAPHY
2018 arXiv
-
[290]
Causal inference and the data-fusion problem
Elias Bareinboim and Judea Pearl. Causal inference and the data-fusion problem. Proceedings of the National Academy of Sciences, 113(27):7345–7352, 2016
2016
-
[291]
Causality
Clark Glymour et al. Causality. Handbook of the Philosophy of Science, 9:706–763, 2019
2019
-
[292]
A survey on causal inference.ACM Transactions on Knowledge Discovery from Data (TKDD), 15(5):1–46, 2021
Liuyi Y ao et al. A survey on causal inference.ACM Transactions on Knowledge Discovery from Data (TKDD), 15(5):1–46, 2021
2021
-
[293]
Stefan Lessmann, Bart Baesens, Hian Chye Seow, and Lyn C. Thomas. Benchmarking state- of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1):124–136, 2015
2015
-
[294]
Credit card fraud detection: A realistic modeling and a novel learning strategy
Andrea Dal Pozzolo, Giacomo Boracchi, Olivier Caelen, Cesare Alippi, and Gianluca Bontempi. Credit card fraud detection: A realistic modeling and a novel learning strategy. IEEE Transac- tions on Neural Networks and Learning Systems, 29(8):3784–3797, 2018
2018
-
[295]
Pre- dicting judicial decisions of the european court of human rights: A natural language processing perspective
Nikolaos Aletras, Dimitrios Tsarapatsanis, Daniel PreoŎciuc-Pietro, and Vasileios Lampos. Pre- dicting judicial decisions of the european court of human rights: A natural language processing perspective. PeerJ Computer Science, 2:e93, 2016
2016
-
[296]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. What does BERT look at? an analysis of BERT’s attention. arXiv preprint arXiv:1906.04341, 2019
1906 arXiv
-
[297]
A deep learning approach to contract element ex- traction
Ilias Chalkidis and Ion Androutsopoulos. A deep learning approach to contract element ex- traction. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1456–1462, 2017
2017
-
[298]
Counterfactual explanations without opening the black box: Automated decisions and the GDPR
Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Tech- nology, 31(2):841–887, 2018
2018
-
[299]
A survey of the state of explainable ai for natural language processing
Marina Danilevsky, Pengjie Qian, Ranit Aharonov, Y annis Katsis, Ban Kawas, and Pavan K Sen. A survey of the state of explainable ai for natural language processing. arXiv preprint arXiv:2010.00711, 2020
2010 arXiv
-
[2024]
Preprint available on arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.