Pith. sign in

REVIEW 3 major objections 5 minor 300 references

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This Ph.D. thesis claims that faithfulness-measurable models, built by randomly masking tokens during fine-tuning, yield token-importance explanations that are near-theoretically-optimal faithful under erasure, without architectural…

desk verdict Solid thesis packaging solid prior work into a provocative but imperfect paradigm; the 'near-optimal faithfulness' claim needs a metric-relative disclaimer. read the letter →

arxiv 2411.17992 v1 pith:4ETZ3SR4 submitted 2024-11-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords faithfulnessinterpretabilityimportancemeasuresmaskedfine-tuningerasuremetricself-explanationsnaturallanguageprocessingpost-hocexplanations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis tries to establish that faithfulness—whether an explanation reflects the model's actual reasoning—can be engineered rather than hoped for. It argues that the two standard paradigms, intrinsic and post-hoc, are both unproductive: post-hoc explanations are often no better than random, and intrinsic models sacrifice generality or overstate their interpretability. The proposed alternative, faithfulness measurable models (FMMs), reframes the goal from "design a model that can be explained" to "design a model for which faithfulness can be measured cheaply and reliably." With masked fine-tuning, removing tokens becomes an in-distribution operation, so the erasure metric can be applied directly; explanations can then be optimized toward maximum faithfulness. The central empirical claim is that FMMs give near-theoretical-optimal faithfulness on synthetic tasks and consistently faithful explanations across real tasks, whereas the same post-hoc methods on plain models are model- and task-dependent.

What carries the argument

The central mechanism is masked fine-tuning: during training, random input tokens are replaced with a mask token, so the model learns to treat masked inputs as ordinary in-distribution inputs. This is what unlocks the erasure metric, because removing allegedly important tokens no longer sends the input out of distribution. On top of this, beam search over token subsets optimizes an importance measure toward maximal faithfulness, turning the FMM into an indirectly self-explaining model without any architectural constraint. The faithfulness score is the relative area between curves, which compares the performance loss from removing the explanation's top tokens against removing tokens at random.

What would settle it

Take a trained FMM on a task with known ground-truth important tokens, for instance a synthetic dataset where only certain tokens determine the label. If the faithfulness-optimized explanation fails to recover those tokens, or if the model's performance after masking its top tokens is not below the random-masking baseline, the central claim of near-optimal faithfulness would be refuted. A simpler check: find any dataset where masked fine-tuning leaves out-of-distribution p-values below the 5% threshold, showing masked inputs are still out-of-distribution.

Watch

Extended reading notes

Core claim

The discovery is that a small, architecture-free change to training—randomly masking input tokens during fine-tuning—makes the faithfulness of token-importance explanations measurable and optimizable. The thesis defines an FMM as a model built so that the erasure faithfulness metric ("if a token is truly important, removing it should hurt the prediction") is cheap and reliable, by making masked inputs in-distribution. It validates this with an out-of-distribution test and then optimizes explanations with beam search. Measured by the relative area between curves, FMMs reach near theoretical optimal faithfulness on synthetic problems and, unlike plain fine-tuning, produce consistently faithful explanations across models, tasks, and explanation methods. The thesis also examines self-explanations from large language models, proposing self-consistency checks, and finds those explanations remain model- and task-dependent.

Load-bearing premise

The entire faithfulness guarantee rests on the assumption that "faithful" means the erasure test: an explanation is faithful if removing its top tokens hurts the model more than removing random tokens, and that masked fine-tuning indeed makes all masked inputs in-distribution.

Editorial extensions

If this is right

  • Practitioners can get consistently faithful token-importance explanations from a standard transformer by adding random masking to fine-tuning, with no loss in predictive performance.
  • Faithfulness of post-hoc and intrinsic explanations should no longer be assumed transferable: it is model- and task-dependent unless the model is designed to be faithfulness-measurable.
  • Because faithfulness is cheap to measure on an FMM, explanations can be audited per model instance before deployment, and optimized toward the theoretical optimum.
  • Self-explanations from large language models are not generally trustworthy; their faithfulness depends on model, task, and explanation type, so they need per-use validation.
  • The same base model, datasets, and post-hoc methods become consistently faithful when switched from plain fine-tuning to masked fine-tuning, isolating the training modification as the cause.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The FMM principle is not tied to token masking in principle; if faithfulness can be made cheap to measure for other explanation types, the same optimize-toward-faithfulness recipe could apply, and the thesis's future-work sketch points at causal language models.
  • If the erasure metric is accepted as the definition of faithfulness, then FMMs provide a practical audit protocol: before relying on an explanation in a high-stakes application, one can compute its faithfulness score on the deployed model and reject explanations that fall below the random baseline.
  • The result suggests that interpretability research may be better served by designing for measurability rather than for explainability, since measurability is what makes optimization and guarantees possible without constraining the architecture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This PhD thesis proposes two new interpretability paradigms: faithfulness measurable models (FMMs) and self-explanations. Chapter 3 develops Recursive ROAR and the RACU metric for erasure-based faithfulness of importance measures. Chapter 4 introduces masked fine-tuning to make erasure evaluation in-distribution and a Beam search that optimizes RACU, reporting that FMMs achieve near-optimal RACU and consistent faithfulness across tasks and models. Chapter 5 evaluates self-explanations from instruction-tuned LLMs via self-consistency checks and finds faithfulness to be model-, task-, and explanation-dependent. The thesis concludes that post-hoc and intrinsic explanations are model- and task-dependent by default, while FMMs yield consistently faithful explanations.

Significance. The FMM idea—reformulating interpretability as designing models for cheap and reliable faithfulness measurement rather than architectural explainability—is a valuable contribution. The thesis benefits from extensive experiments (eight datasets, RoBERTa-base/large and BiLSTM-Attention, confidence intervals, MaSF out-of-distribution checks, and synthetic ground-truth validation for Recursive ROAR) and is honest about limitations. If the central claim is accepted, it would give practitioners a recipe for obtaining post-hoc token-importance explanations that are consistently faithful under the erasure metric without sacrificing predictive performance. However, the headline claim goes beyond what is demonstrated, because faithfulness is operationalized solely through RACU and the Beam optimizer targets that same metric.

major comments (3)
  1. [Abstract; §4.1.5; Table 4.3] The abstract's claim of 'near theoretical optimal in terms of faithfulness' is stronger than the evidence. The Beam method in §4.1.5 explicitly optimizes an explanation to maximize the RACU score defined in Eq. (3.4), and Table 4.3 reports that same RACU score; therefore the near-optimal result is partly a consequence of optimizing and evaluating on the same objective. Please rephrase the claim as 'near-optimal under the RACU erasure metric,' or provide evidence that high-RACU explanations correspond to the model's actual reasoning independently of this metric.
  2. [§3.2.3; §4.2.3; §3.6; §4.4] The synthetic ground-truth validation in §3.2.3 uses a linear problem, while the FMM experiments use non-linear RoBERTa models; the thesis acknowledges in §3.6 and §4.4 that the model's true reasoning is unknown and that erasure is a proxy. Without an independent check that high-RACU explanations track true token importance for these non-linear models, the statement that FMMs yield 'consistently faithful explanations' is established only with respect to the erasure proxy, not with respect to faithfulness as 'reflecting the model's reasoning' as defined in Chapter 2.
  3. [§4.1.3; §4.1.5] The MaSF in-distribution validation is performed for masked inputs at evaluation time, but the Beam search in §4.1.5 generates many partially masked inputs along its trajectory that are not all certified by the MaSF check. Because RACU's validity rests on masked inputs being in-distribution, an out-of-distribution response during the search could inflate RACU even for tokens that are not causally important. The authors should either run the MaSF check on the full set of inputs explored by Beam, or restrict the claim to the subset of masked inputs verified to be in-distribution.
minor comments (5)
  1. [Abstract] The word 'intrisic' is a typo and should read 'intrinsic'.
  2. [Table 1.1] The column header 'defintion' should be 'definition'.
  3. [§1.4.1] The sentence 'this paradigm archives the goal of taking the best part from both paradigms' should use 'achieves' instead of 'archives'.
  4. [Abstract; §4.1.2] The phrase 'randomly masking the training dataset' is imprecise: the method is a two-stage masked fine-tuning procedure with specific masking ratios and a mixed-strategy schedule, which matters for reproducibility.
  5. [§4.2.1] The masked fine-tuning hyperparameters (masking ratio and schedule) are not analyzed for sensitivity; documenting this would strengthen the practical claim that 'simple modifications' to the model suffice.

Circularity Check

1 steps flagged · score 6.0 of 10

Near-optimal faithfulness claim reduces to the RACU objective: Beam optimizes the same metric that defines faithfulness, so high RACU is partly by construction; the FMM training contribution remains independent.

  1. self definitional [Chapter 4, Figure 4.2 caption; Section 4.1.5; metric defined in Chapter 3, eq. (3.4)]
    "Figure 4.2 Visualization of the faithfulness calculation. AUC is the faithfulness area, and RACU is the AUC normalized by the theoretical best explanation. See the definition for AUC and RACU in (3.4)."

    RACU (eq. 3.4) is the paper's operational definition of faithfulness: the area between the explanation's erasure curve and a random baseline, normalized by the theoretical best explanation. Beam (§4.1.5) searches token subsets to maximize exactly this faithfulness area. The abstract then reports 'near theoretical optimal in terms of faithfulness' for FMMs. Because the metric defines the optimum and Beam optimizes that metric, near-optimal RACU is expected if the search succeeds; it is not independent evidence that explanations track the model's reasoning. The non-circular content is the empirical demonstration that masked fine-tuning makes erasure in-distribution and that the optimization works; the 'near-optimal faithfulness' claim is the optimized objective itself.

full rationale

The thesis is largely self-contained: Chapter 3 develops Recursive ROAR/RACU with a synthetic validation showing the metric recovers known features for a linear problem, and Chapter 4 validates masked fine-tuning with the MaSF out-of-distribution check. No load-bearing uniqueness theorem or self-citation chain is invoked. The circularity is narrower but central: the headline result 'FMMs yield explanations that are near theoretical optimal in terms of faithfulness' is measured with RACU, and the Beam explanation method is explicitly designed to maximize that same RACU. Thus the high faithfulness scores reported in Table 4.3 and the abstract are partly guaranteed by construction. The thesis honestly identifies erasure-based faithfulness as a proxy because the model's true reasoning is unknown (Ch. 3), which is a limitation rather than circularity, but it does not undo the metric-objective identity. The FMM paradigm's independent contribution—training-time masking so that erasure is in-distribution, enabling cheap faithfulness measurement—stands on its own; the overreach is presenting the optimized RACU outcome as a discovery about faithfulness rather than as the value of the optimized objective. Score 6 reflects partial, not total, circularity.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim depends on two hand-chosen training and optimization components (masking schedule, Beam budget) and on domain assumptions about what faithfulness means and when masked inputs are in-distribution. No new physical entities are introduced.

free parameters (2)
  • Masking ratio and mixed-strategy schedule in masked fine-tuning = Selected via validation: 'Use 50/50' (half masked, half unmasked) works best
    The FMM training recipe is hand-chosen among several strategies; the choice affects whether masked inputs are in-distribution and, therefore, the measured faithfulness.
  • Beam search optimization budget for explanation search = Not reported in visible text
    The Beam method in Section 4.1.5 optimizes an explanation toward the RACU metric; its budget determines how close to the metric optimum the final explanation can be, so it directly influences the 'near theoretical optimal' result.
assumptions (3)
  • domain assumption Erasure-metric definition of faithfulness: a feature is important iff removing it degrades model performance more than removing random features.
    Adopted from Samek et al. and Hooker et al. in Chapter 3; it is a proxy because the true reasoning of the model is unknown, as stated in Sections 1.2 and 3.1.
  • domain assumption Masked fine-tuning makes masked inputs in-distribution, so erasure evaluations reflect the deployed model's behavior.
    Chapter 4.1.3 validates this with MaSF out-of-distribution tests, but this is a trained property rather than an architectural guarantee.
  • domain assumption Self-consistency of an LLM's explanation is a necessary condition for faithfulness of self-explanations.
    Chapter 5 uses self-consistency checks as faithfulness metrics; consistency alone is not sufficient, which the thesis acknowledges in Section 5.7 and in the Chapter 2 discussion of Parcalabescu and Frank.

how reviews work

0 comments
Cite this review

Pith. "Pith review of New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing." pith.science (2026). https://pith.science/paper/4ETZ3SR4

@misc{pith2026241117992,
  author       = {Pith},
  title        = {Pith review of: New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4ETZ3SR4}},
  note         = {Machine review of arXiv:2411.17992}
}
read the original abstract

As machine learning becomes more widespread and is used in more critical applications, it's important to provide explanations for these models, to prevent unintended behavior. Unfortunately, many current interpretability methods struggle with faithfulness. Therefore, this Ph.D. thesis investigates the question "How to provide and ensure faithful explanations for complex general-purpose neural NLP models?" The main thesis is that we should develop new paradigms in interpretability. This is achieved by first developing solid faithfulness metrics and then applying the lessons learned from this investigation to develop new paradigms. The two new paradigms explored are faithfulness measurable models (FMMs) and self-explanations. The idea in self-explanations is to have large language models explain themselves, we identify that current models are not capable of doing this consistently. However, we suggest how this could be achieved. The idea of FMMs is to create models that are designed such that measuring faithfulness is cheap and precise. This makes it possible to optimize an explanation towards maximum faithfulness, which makes FMMs designed to be explained. We find that FMMs yield explanations that are near theoretical optimal in terms of faithfulness. Overall, from all investigations of faithfulness, results show that post-hoc and intrinsic explanations are by default model and task-dependent. However, this was not the case when using FMMs, even with the same post-hoc explanation methods. This shows, that even simple modifications to the model, such as randomly masking the training dataset, as was done in FMMs, can drastically change the situation and result in consistently faithful explanations. This answers the question of how to provide and ensure faithful explanations.

Figures

Figures reproduced from arXiv: 2411.17992 by the authors.

Figure 2
Figure 2. [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figure 2
Figure 2. Fictive visualization of an [PITH_FULL_IMAGE:figures/full_fig_p058_2.png] view at source ↗
Figures from the paper (25 more)
Figure 2
Figure 2. Figure 2: Three examples from the SST dataset [ [PITH_FULL_IMAGE:figures/full_fig_p059_2.png]
Figure 2
Figure 2. Figure 2: Hypothetical visualization of applying [PITH_FULL_IMAGE:figures/full_fig_p063_2.png]
Figure 2
Figure 2. Figure 2: A fictive visualization of LIME, where the weights of the logistic regression [PITH_FULL_IMAGE:figures/full_fig_p066_2.png]
Figure 2
Figure 2. Figure 2: Fictive visualization of [PITH_FULL_IMAGE:figures/full_fig_p067_2.png]
Figure 2
Figure 2. Figure 2: Hypothetical results of [PITH_FULL_IMAGE:figures/full_fig_p071_2.png]
Figure 2
Figure 2. Figure 2: Hypothetical visualization of how [PITH_FULL_IMAGE:figures/full_fig_p072_2.png]
Figure 3
Figure 3. Figure 3: Example of how a redundancy can be removed in [PITH_FULL_IMAGE:figures/full_fig_p082_3.png]
Figure 3
Figure 3. Figure 3: Using the weights of a linear model as the explanation, ROAR and Recursive [PITH_FULL_IMAGE:figures/full_fig_p083_3.png]
Figure 3
Figure 3. Figure 3: shows that Recursive ROAR is identical to the ground truth, while ROAR is worse, [PITH_FULL_IMAGE:figures/full_fig_p084_3.png]
Figure 3
Figure 3. Figure 3: also presents the model performance at 100% masking, which provides a lower [PITH_FULL_IMAGE:figures/full_fig_p088_3.png]
Figure 3
Figure 3. Figure 3: Recursive ROAR results, showing model performance at x% of tokens masked. A [PITH_FULL_IMAGE:figures/full_fig_p089_3.png]
Figure 3
Figure 3. Figure 3: Visualization of the faithfulness cal [PITH_FULL_IMAGE:figures/full_fig_p090_3.png]
Figure 3
Figure 3. Figure 3: Recursive ROAR results, showing model performance at up to 10 tokens masked. [PITH_FULL_IMAGE:figures/full_fig_p093_3.png]
Figure 3
Figure 3. Figure 3: The accumulative importance score relative to the total importance score for the [PITH_FULL_IMAGE:figures/full_fig_p094_3.png]
Figure 4
Figure 4. Figure 4: To measure faithfulness, a [PITH_FULL_IMAGE:figures/full_fig_p099_4.png]
Figure 4
Figure 4. Figure 4: Visualization of the faithfulness [PITH_FULL_IMAGE:figures/full_fig_p106_4.png]
Figure 4
Figure 4. Figure 4: The unmasked performance for [PITH_FULL_IMAGE:figures/full_fig_p110_4.png]
Figure 4
Figure 4. Figure 4: The performance given the masked [PITH_FULL_IMAGE:figures/full_fig_p111_4.png]
Figure 5
Figure 5. Figure 5: Example of an LLM providing a [PITH_FULL_IMAGE:figures/full_fig_p120_5.png]
Figure 5
Figure 5. Figure 5: we explicitly express the target sentiment in the prompt. To evaluate robust [PITH_FULL_IMAGE:figures/full_fig_p123_5.png]
Figure 5
Figure 5. Figure 5: The explicit input-template prompt [PITH_FULL_IMAGE:figures/full_fig_p124_5.png]
Figure 5
Figure 5. Figure 5: Prompt-template for classification [PITH_FULL_IMAGE:figures/full_fig_p125_5.png]
Figure 5
Figure 5. Figure 5: shows that neither the redaction-instruction nor the persona affects the results [PITH_FULL_IMAGE:figures/full_fig_p128_5.png]
Figure 5
Figure 5. Figure 5: shows the faithfulness, for each prompt-variation for Llama2-70B. Figure 5.9 shows [PITH_FULL_IMAGE:figures/full_fig_p129_5.png]
Figure 5
Figure 5. Figure 5: Faithfulness evaluation using self [PITH_FULL_IMAGE:figures/full_fig_p130_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 16 canonical work pages

  1. [1]

    Visualizing memorization in RNNs,

    A. Madsen, “Visualizing memorization in RNNs,”Distill, vol. 4, no. 3, 3 2019. [Online]. Available: https://distill.pub/2019/memorization-in-rnns

  2. [2]

    Attention is not Explanation,

    S. Jain and B. C. Wallace, “Attention is not Explanation,” in Proceedings of the 2019 Conference of the North, vol. 1. Stroudsburg, PA, USA: Association for Computational Linguistics, 2 2019, pp. 3543–3556. [Online]. Available: http://aclweb.org/anthology/N19-1357

  3. [3]

    On the convergence of Adam and beyond,

    S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of Adam and beyond,”6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings, pp. 1–23, 4 2018. [Online]. Available: http://arxiv.org/abs/1904.09237

  4. [4]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach,

    Y. Liuet al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,”arXiv, 7 2019. [Online]. Available: http://arxiv.org/abs/1907.11692

  5. [5]

    Transformers: State-of-the-Art Natural Language Processing,

    T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Stroudsburg, PA, USA: Association for Computational Linguistics, 10 2020, pp. 38–45. [Online]. Available: http: //arxiv.org/abs/1910.03771https://www.aclweb.org/ant...

  6. [6]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”7th International Conference on Learning Representations, ICLR 2019, 2019

  7. [7]

    GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,

    A. Wanget al., “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” inInternational Conference on Learning Representations,

  8. [8]

    SuperGLUE: A stickier benchmark for general-purpose language understanding systems,

    ——, “SuperGLUE: A stickier benchmark for general-purpose language understanding systems,” Advances in Neural Information Processing Systems, vol. 32, no. July, pp. 1–30, 2019

Show all 300 references
  1. [9]

    MIMIC-III, a freely accessible critical care database,

    A. E. Johnson et al., “MIMIC-III, a freely accessible critical care database,” Scientific Data , vol. 3, no. 1, p. 160035, 12 2016. [Online]. Available: http://www.nature.com/articles/sdata201635

  2. [10]

    Towards AI-complete question answering: A set of prerequisite toy tasks,

    J. Weston et al. , “Towards AI-complete question answering: A set of prerequisite toy tasks,” 4th International Conference on Learning Representations, 102 ICLR 2016 - Conference Track Proceedings , 2 2016. [Online]. Available: http://arxiv.org/abs/1502.05698

  3. [11]

    Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference,

    T. McCoy, E. Pavlick, and T. Linzen, “Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Lin...

  4. [12]

    A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,

    A. Williams, N. Nangia, and S. Bowman, “A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference,” inProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (...

  5. [13]

    Parsing with compositional vector grammars,

    R. Socheret al., “Parsing with compositional vector grammars,” inACL 2013 - 51st Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference. Association for Computational Linguistics, 2013, vol. 1, pp. 455–465. [Online]. Available: https://a...

  6. [14]

    A benchmark for interpretability methods in deep neural networks,

    S. Hookeret al., “A benchmark for interpretability methods in deep neural networks,” in Advances in Neural Information Processing Systems, vol. 32, 6 2019. [Online]. Available: http://arxiv.org/abs/1806.10758

  7. [15]

    Semantically Equivalent Adversarial Rules for Debugging NLP models,

    M. T. Ribeiro, S. Singh, and C. Guestrin, “Semantically Equivalent Adversarial Rules for Debugging NLP models,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , vol. 1. Stroudsburg, PA, USA: Association for Co...

  8. [16]

    Investigating Gender Bias in Language Models Using Causal Mediation Analysis,

    J. Viget al., “Investigating Gender Bias in Language Models Using Causal Mediation Analysis,” inAdvances in Neural Information Processing Systems, H. Larochelleet al., Eds., vol. 33. Curran Associates, Inc., 2020, pp. 12388–12401. [Online]. Available: https://proceedings.neuri...

  9. [17]

    LIII. On lines and planes of closest fit to systems of points in space,

    K. Pearson, “LIII. On lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin Philosophical Magazine and 103 Journal of Science, vol. 2, no. 11, pp. 559–572, 11 1901. [Online]. Available: https://www.tandfonline.com/doi/full/10.1080/14...

  10. [18]

    Visualizing data using t-SNE,

    L. Van Der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research, vol. 9, pp. 2579–2605, 2008. [Online]. Available: https://www.jmlr.org/papers/v9/vandermaaten08a.html

  11. [19]

    Glove: Global Vectors for Word Representation,

    J. Pennington, R. Socher, and C. Manning, “Glove: Global Vectors for Word Representation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Stroudsburg, PA, USA: Association for Computational Linguistics, 2014, pp. 1532–1543. ...

  12. [20]

    BERT Rediscovers the Classical NLP Pipeline,

    I. Tenney, D. Das, and E. Pavlick, “BERT Rediscovers the Classical NLP Pipeline,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 5 2019, pp. 4593–4601. [Online]. Avail...

  13. [21]

    BERT: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlinet al., “BERT: Pre-training of deep bidirectional transformers for language understanding,” inNAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, ...

  14. [22]

    Boolq: Exploringthesurprisingdifficultyofnaturalyes/noquestions,

    C.Clark et al., “Boolq: Exploringthesurprisingdifficultyofnaturalyes/noquestions,” in NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, vol. 1, 2019, pp....

  15. [23]

    The CommitmentBank: Investigating projection in naturally occurring discourse,

    M.-C. d. Marneffe, M. Simons, and J. Tonhauser, “The CommitmentBank: Investigating projection in naturally occurring discourse,” Proceedings of Sinn und Bedeutung , vol. 23, no. 2, pp. 107–124, 2019. [Online]. Available: https://ojs.ub.uni-konstanz.de/sub/index.php/sub/article...

  16. [24]

    Neural Network Acceptability Judgments,

    A. Warstadt, A. Singh, and S. R. Bowman, “Neural Network Acceptability Judgments,” Transactions of the Association for Computational Linguistics, vol. 7, pp. 625–641, 11

  17. [25]

    CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,

    A. Talmor et al., “CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,” in Proceedings of the 2019 Conference of the North. 104 Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4149–4158. [Online]. Available: http://aclweb.o...

  18. [26]

    Available: https://direct.mit.edu/tacl/article/43528

    [Online]. Available: https://direct.mit.edu/tacl/article/43528

  19. [27]

    Automatically Constructing a Corpus of Sentential Paraphrases,

    W. B. Dolan and C. Brockett, “Automatically Constructing a Corpus of Sentential Paraphrases,” in Proceedings of the Third International Workshop on Paraphrasing (IWP2005) , 2005, pp. 9–16. [Online]. Available: https: //research.microsoft.com/apps/pubs/default.aspx?id=101076

  20. [28]

    Learning word vectors for sentiment analysis,

    A. L. Maas et al., “Learning word vectors for sentiment analysis,” in ACL-HLT 2011 - Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies , vol. 1. Portland, Oregon, USA: Association for Computational Linguistics,...

  21. [29]

    The PASCAL Recognising Textual Entailment Challenge,

    I. Dagan, O. Glickman, and B. Magnini, “The PASCAL Recognising Textual Entailment Challenge,” inMachine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, J. Quiñonero-Candelaet al., Eds. Berlin, Heidelberg...

  22. [30]

    MCTest: A challenge dataset for the open-domain machine comprehension of text,

    M. Richardson, C. J. Burges, and E. Renshaw, “MCTest: A challenge dataset for the open-domain machine comprehension of text,”EMNLP 2013 - 2013 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, vol. D13-1020, no. October, pp. 193–203, 2013

  23. [31]

    SQuad: 100,000+ questions for machine comprehension of text,

    P. Rajpurkaret al., “SQuad: 100,000+ questions for machine comprehension of text,” EMNLP 2016 - Conference on Empirical Methods in Natural Language Processing, Proceedings, pp. 2383–2392, 2016

  24. [32]

    A large annotated corpus for learning natural language inference,

    S. R. Bowmanet al., “A large annotated corpus for learning natural language inference,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2015, pp. 632–642. [Online]. Avai...

  25. [33]

    Long Short-Term Memory,

    S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 11 1997. [Online]. Available: https://www.mitpressjournals.org/doi/abs/10.1162/neco.1997.9.8.1735 105

  26. [34]

    First Quora Dataset Release: Question Pairs,

    S. Iyer, N. Dandekar, and K. Csernai, “First Quora Dataset Release: Question Pairs,”

  27. [35]

    The Solvability of Interpretability Evaluation Metrics,

    Y. Zhou and J. Shah, “The Solvability of Interpretability Evaluation Metrics,” in Findings of the Association for Computational Linguistics: EACL, 2023. [Online]. Available: http://arxiv.org/abs/2205.08696

  28. [36]

    Explain Yourself! Leveraging Language Models for Commonsense Reasoning,

    N. F. Rajani et al. , “Explain Yourself! Leveraging Language Models for Commonsense Reasoning,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4932–4942. [O...

  29. [37]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffelet al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research, vol. 21, pp. 1–67, 2020. [Online]. Available: https://jmlr.org/papers/v21/20-074.html

  30. [38]

    Visualizing and Understanding Neural Models in NLP,

    J. Liet al., “Visualizing and Understanding Neural Models in NLP,” inProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg, PA, USA: Association for Computational Linguistics,...

  31. [39]

    Axiomatic attribution for deep networks,

    M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in 34th International Conference on Machine Learning, ICML 2017, vol. 7, 3 2017, pp. 5109–5118. [Online]. Available: http://arxiv.org/abs/1703.01365

  32. [40]

    How to explain individual classification decisions,

    D. Baehrens et al., “How to explain individual classification decisions,”Journal of Machine Learning Research, vol. 11, pp. 1803–1831, 12 2010. [Online]. Available: http://arxiv.org/abs/0912.1128

  33. [41]

    "Why should i trust you?

    M. T. Ribeiro, S. Singh, and C. Guestrin, “"Why should i trust you?" Explaining the predictions of any classifier,” inProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , vol. 13-17- Augu. New York, NY, USA: ACM, 8 2016, pp. 1135–1144...

  34. [42]

    Explaining NLP Models via Minimal Contrastive Editing (MiCE),

    A. Ross, A. Marasović, and M. Peters, “Explaining NLP Models via Minimal Contrastive Editing (MiCE),” inFindings of the Association for Computational Linguistics: ACL- 106 IJCNLP 2021. Stroudsburg, PA, USA: Association for Computational Linguistics, 12 2021, pp. 3840–3852. [On...

  35. [43]

    Understanding Neural Networks through Representation Erasure,

    J. Li, W. Monroe, and D. Jurafsky, “Understanding Neural Networks through Representation Erasure,”arXiv, 2016. [Online]. Available: http://arxiv.org/abs/1612.0 8220

  36. [44]

    A Unified Approach to Interpreting Model Predictions,

    S. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in Advances in Neural Information Processing Systems, 5 2017, pp. 4766–4775. [Online]. Available: http://arxiv.org/abs/1705.07874

  37. [45]

    Sanity checks for saliency maps,

    J. Adebayoet al., “Sanity checks for saliency maps,” inAdvances in Neural Information Processing Systems, vol. 2018-Decem. Curran Associates, Inc., 10 2018, pp. 9505–9515. [Online]. Available: http://arxiv.org/abs/1810.03292

  38. [46]

    NILE : Natural Language Inference with Faithful Natural Language Explanations,

    S. Kumar and P. Talukdar, “NILE : Natural Language Inference with Faithful Natural Language Explanations,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 5 2020, pp. 8...

  39. [47]

    Explainable Machine Learning in Deployment,

    U. Bhattet al., “Explainable Machine Learning in Deployment,”Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp. 648–657, 9 2019. [Online]. Available: https://dl.acm.org/doi/10.1145/3351095.3375624

  40. [48]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,

    C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,”Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, 2019. [Online]. Available: http://www.nature.com/articles/s42256-019-0048-x

  41. [49]

    A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural Networks,

    H. Matanet al., “A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural Networks,”International Conference on Learning Representations, 2022. [Online]. Available: https://openreview.net/forum?id=Oy9WeuZD51

  42. [50]

    How bad is Sacramento’s air, exactly? Google results appear at odds with reality, some say,

    M. McGough, “How bad is Sacramento’s air, exactly? Google results appear at odds with reality, some say,” 2018. [Online]. Available: https: //www.sacbee.com/news/california/fires/article216227775.html

  43. [51]

    On the Safety of Machine Learning: Cyber-Physical Systems, Decision Sciences, and Data Products,

    K. R. Varshney and H. Alemzadeh, “On the Safety of Machine Learning: Cyber-Physical Systems, Decision Sciences, and Data Products,”Big Data, vol. 5, no. 3, pp. 246–255, 9

  44. [52]

    When a computer program keeps you in jail: How computers are harming criminal justice,

    R. Wexler, “When a computer program keeps you in jail: How computers are harming criminal justice,” 2017. [Online]. Available: https://www.nytimes.com/2017/06/13/opi nion/how-computers-are-harming-criminal-justice.html

  45. [53]

    Dissecting racial bias in an algorithm used to manage the health of populations,

    Z. Obermeyeret al., “Dissecting racial bias in an algorithm used to manage the health of populations,”Science, vol. 366, no. 6464, pp. 447–453, 10 2019. [Online]. Available: https://science.sciencemag.org/content/366/6464/447

  46. [54]

    Language Models are Few-Shot Learners,

    T. B. Brownet al., “Language Models are Few-Shot Learners,” inAdvances in Neural Information Processing Systems, H. Larochelleet al., Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: https://proceedings.neurips.cc/paper/2 020/file/1457c0d6bfcb4...

  47. [55]

    Available: http://www.ncbi.nlm.nih.gov/pubmed/28933947 107

    [Online]. Available: http://www.ncbi.nlm.nih.gov/pubmed/28933947 107

  48. [56]

    Accountability of AI Under the Law: The Role of Explanation,

    F. Doshi-Velez et al., “Accountability of AI Under the Law: The Role of Explanation,” SSRN Electronic Journal, vol. Online, 11 2017. [Online]. Available: https://www.ssrn.com/abstract=3064761

  49. [57]

    A Survey on Bias in Deep NLP,

    I. Garrido-Muñozet al., “A Survey on Bias in Deep NLP,”Applied Sciences, vol. 11, no. 7, p. 3184, 4 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/7/3184

  50. [58]

    Towards A Rigorous Science of Interpretable Machine Learning,

    F. Doshi-Velez and B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” arXiv, 2 2017. [Online]. Available: http://arxiv.org/abs/1702.08608

  51. [59]

    On the Dangers of Stochastic Parrots,

    E. M. Bender et al., “On the Dangers of Stochastic Parrots,” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . New York, NY, USA: ACM, 3 2021, pp. 610–623. [Online]. Available: https://dl.acm.org/doi/10.1145/3442188.3445922

  52. [60]

    A Survey on Bias and Fairness in Machine Learning,

    N. Mehrabi et al., “A Survey on Bias and Fairness in Machine Learning,” ACM Computing Surveys, vol. 54, no. 6, pp. 1–35, 2021

  53. [61]

    The mythos of model interpretability,

    Z. C. Lipton, “The mythos of model interpretability,” Communications of the ACM, vol. 61, no. 10, pp. 36–43, 9 2018. [Online]. Available: https: //dl.acm.org/doi/10.1145/3233231 108

  54. [62]

    T. S. Kuhn,The Structure of Scientific Revolutions, 3rd ed. University of Chicago Press, 1996

  55. [63]

    Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?

    A. Jacovi and Y. Goldberg, “Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics...

  56. [64]

    Evaluating the Visualization of What a Deep Neural Network Has Learned,

    W. Samek et al. , “Evaluating the Visualization of What a Deep Neural Network Has Learned,” IEEE Transactions on Neural Networks and Learning Systems, vol. 28, no. 11, pp. 2660–2673, 11 2017. [Online]. Available: https: //ieeexplore.ieee.org/document/7552539/

  57. [65]

    What we can’t measure, We can’t understand: Challenges to demographic data procurement in the pursuit of fairness,

    M. Andrus et al., “What we can’t measure, We can’t understand: Challenges to demographic data procurement in the pursuit of fairness,”FAccT 2021 - Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 249–260, 2021

  58. [66]

    An overview of ethical issues in using AI systems in hiring with a case study of Amazon’s AI based hiring tool,

    A. A. Kodiyan, “An overview of ethical issues in using AI systems in hiring with a case study of Amazon’s AI based hiring tool,”Researchgate Preprint, pp. 1–19, 2019

  59. [67]

    Barocas, M

    S. Barocas, M. Hardt, and A. Narayanan,Fairness and Machine Learning: Limitations and Opportunities. fairmlbook.org, 2019. [Online]. Available: https://fairmlbook.org/

  60. [68]

    On the Legal Compatibility of Fairness Definitions,

    A. Xiang and I. D. Raji, “On the Legal Compatibility of Fairness Definitions,”Workshop on Human-Centric Machine Learning at the 33rd Conference on Neural Information Processing Systems, 2019. [Online]. Available: http://arxiv.org/abs/1912.00761

  61. [69]

    Interpretable deep learning in drug discovery,

    K. Preuer et al., “Interpretable deep learning in drug discovery,” Explainable AI: interpreting, explaining and visualizing deep learning, pp. 331–345, 2019

  62. [70]

    Drug discovery with explainable artificial intelligence,

    J. Jiménez-Luna, F. Grisoni, and G. Schneider, “Drug discovery with explainable artificial intelligence,”Nature Machine Intelligence, vol. 2, no. 10, pp. 573–584, 2020

  63. [71]

    Hidden Workers: Untapped Talent,

    J. B. Fulleret al., “Hidden Workers: Untapped Talent,”Harvard Business School Project on Managing the Future of Work and Accenture, 2021. [Online]. Available: https://www.pw.hks.harvard.edu/post/hidden-workers-untapped-talent

  64. [72]

    Companies Need More Workers. Why Do They Reject Millions of Résumés?

    J. Fuller, “Companies Need More Workers. Why Do They Reject Millions of Résumés?” The project on workforce , 2021. [Online]. Available: https: //www.pw.hks.harvard.edu/post/companies-need-more-workers-wsj

  65. [73]

    A Mathematical Framework for Transformer Circuits,

    N. Elhageet al., “A Mathematical Framework for Transformer Circuits,”Anthropic,

  66. [74]

    AIintheUK:Ready, WillingandAble?

    U.G.HouseofLords, “AIintheUK:Ready, WillingandAble?” 2017.[Online].Available: https://publications.parliament.uk/pa/ld201719/ldselect/ldai/100/10007.htm

  67. [75]

    Machine learning in drug discovery: a review,

    S. Daraet al., “Machine learning in drug discovery: a review,”Artificial Intelligence Review, vol. 55, no. 3, pp. 1947–1999, 2022

  68. [76]

    Thread: Circuits,

    N. Cammarata et al., “Thread: Circuits,” Distill, vol. 5, no. 3, 3 2020. [Online]. Available: https://distill.pub/2020/circuits

  69. [77]

    One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques,

    V. Arya et al. , “One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques,” arXiv, 9 2019. [Online]. Available: http://arxiv.org/abs/1909.03012

  70. [78]

    Definitions, methods, and applications in interpretable machine learning,

    W. J. Murdoch et al., “Definitions, methods, and applications in interpretable machine learning,” Proceedings of the National Academy of Sciences of the United States of America, vol. 116, no. 44, pp. 22071–22080, 10 2019. [Online]. Available: http://www.pnas.org/lookup/doi/10...

  71. [79]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. H. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings. International Conference on Learning Representations, ICLR, 9 ...

  72. [80]

    Machine Learning Interpretability: A Survey on Methods and Metrics,

    D. V. Carvalho, E. M. Pereira, and J. S. Cardoso, “Machine Learning Interpretability: A Survey on Methods and Metrics,”Electronics, vol. 8, no. 8, p. 832, 7 2019. [Online]. Available: https://www.mdpi.com/2079-9292/8/8/832

  73. [81]

    Comparing Explanation Methods for Traditional Machine Learning Models Part 1: An Overview of Current Methods and Quantifying Their Disagreement,

    M. Floraet al., “Comparing Explanation Methods for Traditional Machine Learning Models Part 1: An Overview of Current Methods and Quantifying Their Disagreement,” arXiv, pp. 1–22, 2022. [Online]. Available: http://arxiv.org/abs/2211.08943

  74. [82]

    Neural module networks: A review,

    H. Fashandi, “Neural module networks: A review,”Neurocomputing, vol. 552, p. 126518,

  75. [83]

    Classification by Set Cover: The Prototype Vector Machine,

    J. Bien and R. Tibshirani, “Classification by Set Cover: The Prototype Vector Machine,” arXiv, pp. 1–24, 2009. [Online]. Available: http://arxiv.org/abs/0908.2284 110

  76. [84]

    The Bayesian case model: A generative approach for case-based reasoning and prototype classification,

    B. Kim, C. Rudin, and J. Shah, “The Bayesian case model: A generative approach for case-based reasoning and prototype classification,”Advances in Neural Information Processing Systems, vol. 3, no. January, pp. 1952–1960, 2014

  77. [85]

    Neural Module Networks,

    J. Andreaset al., “Neural Module Networks,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 6 2016, pp. 39–48. [Online]. Available: http://ieeexplore.ieee.org/document/7780381/

  78. [86]

    Neural Module Networks for Reasoning over Text,

    N. Guptaet al., “Neural Module Networks for Reasoning over Text,” inInternational Conference on Learning Representations (ICLR) , 12 2020. [Online]. Available: https://openreview.net/forum?id=SygWvAVFPr

  79. [87]

    Visualizing and Understanding Recurrent Networks,

    A. Karpathy, J. Johnson, and L. Fei-Fei, “Visualizing and Understanding Recurrent Networks,”arXiv, pp. 1–12, 6 2015. [Online]. Available: http://arxiv.org/abs/1506.02078

  80. [88]

    Is Attention Interpretable?

    S. Serrano and N. A. Smith, “Is Attention Interpretable?” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 6 2019, pp. 2931–2951. [Online]. Available: https://www.aclweb....

  81. [89]

    Attention Interpretability Across NLP Tasks,

    S. Vashishth et al., “Attention Interpretability Across NLP Tasks,”arXiv, 9 2019. [Online]. Available: http://arxiv.org/abs/1909.11218

  82. [90]

    Is Sparse Attention more Interpretable?

    C. Meister et al., “Is Sparse Attention more Interpretable?” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Stroudsburg, PA, USA: As...

  83. [91]

    This looks like that: Deep learning for interpretable image recognition,

    C. Chenet al., “This looks like that: Deep learning for interpretable image recognition,” Advances in Neural Information Processing Systems, vol. 32, 6 2019. [Online]. Available: http://arxiv.org/abs/1806.10574

  84. [92]

    Noise-adding Methods of Saliency Map as Series of Higher Order Partial Derivative,

    J. Seoet al., “Noise-adding Methods of Saliency Map as Series of Higher Order Partial Derivative,” in2018 ICML Workshop on Human Interpretability in Machine Learning, 6

  85. [93]

    European union regulations on algorithmic decision making and a

    B. Goodman and S. Flaxman, “European union regulations on algorithmic decision making and a "right to explanation",”AI Magazine, vol. 38, no. 3, pp. 50–57, 2017

  86. [94]

    The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective,

    S. Krishnaet al., “The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective,”arXiv, 2022. [Online]. Available: http://arxiv.org/abs/2202 .01602

  87. [95]

    The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?

    J. Bastings and K. Filippova, “The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?” in Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP. Stroudsburg, PA, USA: Association ...

  88. [96]

    A review of modularization techniques in artificial neural networks,

    M. Amer and T. Maul, “A review of modularization techniques in artificial neural networks,” Artificial Intelligence Review, vol. 52, no. 1, pp. 527–561, 6 2019. [Online]. Available: http://link.springer.com/10.1007/s10462-019-09706-7

  89. [97]

    Obtaining faithful interpretations from compositional neural networks,

    S. Subramanian et al., “Obtaining faithful interpretations from compositional neural networks,” Proceedings of the Annual Meeting of the Association for Computational Linguistics , pp. 5594–5608, 2020. [Online]. Available: https: //www.aclweb.org/anthology/2020.acl-main.495

  90. [98]

    “Will You Find These Shortcuts?

    J. Bastingset al., ““Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for C...

  91. [99]

    Explainable Artificial Intelligence (XAI) DARPA-BAA-16-53,

    DARPA, “Explainable Artificial Intelligence (XAI) DARPA-BAA-16-53,” Defense Advanced Research Projects Agency (DARPA), pp. 1–52, 2016. [Online]. Available: https://www.darpa.mil/attachments/DARPA-BAA-16-53.pdf 111

  92. [100]

    Learning important features through propagating activation differences,

    A. Shrikumar, P. Greenside, and A. Kundaje, “Learning important features through propagating activation differences,” in 34th International Conference on Machine Learning, ICML 2017, vol. 7, 2017, pp. 4844–4866. [Online]. Available: https://arxiv.org/

  93. [101]

    SmoothGrad: removing noise by adding noise,

    D. Smilkovet al., “SmoothGrad: removing noise by adding noise,”ICML workshop on visualization for deep learning, 2017. [Online]. Available: https://goo.gl/EfVzEE. 112

  94. [102]

    Normlime: A new feature importance metric for explaining deep neural networks,

    I. Ahern et al., “Normlime: A new feature importance metric for explaining deep neural networks,”arXiv, 9 2019. [Online]. Available: http://arxiv.org/abs/1909.04200

  95. [103]

    Generating Token-Level Explanations for Natural Language Inference,

    J. Thorneet al., “Generating Token-Level Explanations for Natural Language Inference,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), vol. 1. S...

  96. [104]

    ILIME: Local and Global Interpretable Model-Agnostic Explainer of Black-Box Decision,

    R. ElShawiet al., “ILIME: Local and Global Interpretable Model-Agnostic Explainer of Black-Box Decision,” inAdvances in Databases and Information Systems, T. Welzer et al., Eds. Cham: Springer International Publishing, 2019, pp. 53–68. [Online]. Available: http://link.springer...

  97. [105]

    Towards Faithful Model Explanation in NLP: A Survey,

    Q. Lyu, M. Apidianaki, and C. Callison-Burch, “Towards Faithful Model Explanation in NLP: A Survey,”Computational Linguistics, vol. 50, no. 2, pp. 657–723, 6 2024. [Online]. Available: http://arxiv.org/abs/2209.11326https://direct.mit.edu/coli/article/ 50/2/657/119158/Towards-...

  98. [106]

    Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers,

    A. Binder et al. , “Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers,” in Artificial Neural Networks and Machine Learning – ICANN 2016, vol. 9887 LNCS, 2016, pp. 63–71. [Online]. Available: http://link.springer.com/10.1007/978-3-319-44781-0_8

  99. [107]

    The (Un)reliability of Saliency Methods,

    P.-J. Kindermanset al., “The (Un)reliability of Saliency Methods,” inLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). Springer, 11 2019, vol. 11700 LNCS, pp. 267–280. [Online]. Available: http...

  100. [108]

    Fooling LIME and SHAP,

    D. Slack et al., “Fooling LIME and SHAP,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. New York, NY, USA: ACM, 2 2020, pp. 180–186. [Online]. Available: https://dl.acm.org/doi/10.1145/3375627.3375830

  101. [109]

    On the (In)fidelity and Sensitivity of Explanations,

    C.-K. Yehet al., “On the (In)fidelity and Sensitivity of Explanations,” inAdvances in Neural Information Processing Systems 32, H. Wallachet al., Eds. Vancouver, Canada: Curran Associates, Inc., 2019, pp. 10967–10978. [Online]. Available: https://arxiv.org/abs/1901.09392

  102. [110]

    Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations,

    T. Han, S. Srinivas, and H. Lakkaraju, “Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations,” Advances in Neural Information Processing Systems, vol. 35, no. NeurIPS, 2022. [Online]. Available: http://arxiv.org/abs/22...

  103. [111]

    Impossibility theorems for feature attribution,

    B. Bilodeauet al., “Impossibility theorems for feature attribution,”Proceedings of the National Academy of Sciences, vol. 121, no. 2, pp. 1–38, 1 2024. [Online]. Available: https://pnas.org/doi/10.1073/pnas.2304406120http://arxiv.org/abs/2212.11870

  104. [112]

    Guided-LIME: Structured sampling based hybrid approach towards explaining blackbox machine learning models,

    A. Sangroyaet al., “Guided-LIME: Structured sampling based hybrid approach towards explaining blackbox machine learning models,” inCEUR Workshop Proceedings, vol. 2699, 2020

  105. [113]

    Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation,

    J. Adebayoet al., “Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation,” inInternational Conference on Learning Representations, 2021, pp. 1–13. [Online]. Available: https://openreview.net/forum?id=xNOVfCCvDpM

  106. [114]

    Understanding Neural Networks Through Deep Visualization,

    J. Yosinskiet al., “Understanding Neural Networks Through Deep Visualization,” in Deep Learning Workshop at 31st International Conference on Machine Learning, 2015. [Online]. Available: http://arxiv.org/abs/1506.06579

  107. [115]

    Don’t trust your eyes: on the (un)reliability of feature visualizations,

    R. Geirhoset al., “Don’t trust your eyes: on the (un)reliability of feature visualizations,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2306.04719

  108. [116]

    Exemplary Natural Images Explain Cnn Activations Better Than State-of-the-Art Feature Visualization,

    J. Borowskiet al., “Exemplary Natural Images Explain Cnn Activations Better Than State-of-the-Art Feature Visualization,”ICLR 2021 - 9th International Conference on Learning Representations, pp. 1–41, 2021

  109. [117]

    How Well do Feature Visualizations Support Causal Under- standing of CNN Activations?

    R. S. Zimmermannet al., “How Well do Feature Visualizations Support Causal Under- standing of CNN Activations?”Advances in Neural Information Processing Systems, vol. 14, no. NeurIPS, pp. 11730–11744, 2021

  110. [118]

    Analysis Methods in Neural Language Processing: A Survey,

    Y. Belinkov and J. Glass, “Analysis Methods in Neural Language Processing: A Survey,” Transactions of the Association for Computational Linguistics, vol. 7, pp. 49–72, 4 2019. [Online]. Available: https://doi.org/10.1162/tacl_a_00254

  111. [119]

    Feature Visualization,

    C. Olah, A. Mordvintsev, and L. Schubert, “Feature Visualization,”Distill, vol. 2, no. 11, 11 2017. [Online]. Available: https://distill.pub/2017/feature-visualization

  112. [120]

    Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks,

    A. Nguyen, J. Yosinski, and J. Clune, “Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks,” Visualization for Deep Learning workshop at ICML , 2016. [Online]. Available: http://arxiv.org/abs/1602.03616

  113. [121]

    Visualizing and Measuring the Geometry of BERT,

    A. Coenenet al., “Visualizing and Measuring the Geometry of BERT,” inAdvances in Neural Information Processing Systems, H. Wallachet al., Eds., vol. 32. Curran Associates, Inc., 6 2019, pp. 8594–8603. [Online]. Available: https://proceedings.neurip s.cc/paper/2019/file/159c1ff...

  114. [122]

    What Does BERT Look at? An Analysis of BERT’s Attention,

    K. Clarket al., “What Does BERT Look at? An Analysis of BERT’s Attention,” in Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 276–286. [Online]. Ava...

  115. [123]

    Local Structure Matters Most: Perturbation Study in NLU,

    L. Clouatreet al., “Local Structure Matters Most: Perturbation Study in NLU,” in Findings of the Association for Computational Linguistics: ACL 2022. Stroudsburg, PA, USA: Association for Computational Linguistics, 7 2022, pp. 3712–3731. [Online]. Available: https://aclantholo...

  116. [124]

    What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties,

    A. Conneauet al., “What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: Association for Com...

  117. [125]

    Probing Classifiers: Promises, Shortcomings, and Advances,

    Y. Belinkov, “Probing Classifiers: Promises, Shortcomings, and Advances,”arXiv, pp. 1–12, 2 2021. [Online]. Available: http://arxiv.org/abs/2102.12452

  118. [126]

    Interpretability and Analysis in Neural NLP,

    Y. Belinkov, S. Gehrmann, and E. Pavlick, “Interpretability and Analysis in Neural NLP,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts . Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. ...

  119. [127]

    A Primer in BERTology: What We Know About How BERT Works,

    A. Rogers, O. Kovaleva, and A. Rumshisky, “A Primer in BERTology: What We Know About How BERT Works,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 842–866, 12 2020. [Online]. Available: https://direct.mit.edu/tacl/article/96482 114

  120. [128]

    Information-Theoretic Probing with Minimum Description Length,

    E. Voita and I. Titov, “Information-Theoretic Probing with Minimum Description Length,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Stroudsburg, PA, USA: Association 115 for Computational Linguistics, 3 2020, pp. 183–196....

  121. [129]

    Interpretability Needs a New Paradigm,

    A. Madsenet al., “Interpretability Needs a New Paradigm,”arXiv, 5 2024. [Online]. Available: http://arxiv.org/abs/2405.05386

  122. [130]

    Post-hoc Interpretability for Neural NLP: A Survey,

    A. Madsen, S. Reddy, and S. Chandar, “Post-hoc Interpretability for Neural NLP: A Survey,” ACM Computing Surveys, vol. 55, no. 8, pp. 1–42, 8 2022. [Online]. Available: https://dl.acm.org/doi/10.1145/3546577

  123. [131]

    Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining,

    A. Madsen et al., “Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining,” inFindings of the Association for Computational Linguistics: EMNLP 2022. Abu Dhabi, United Arab Emirates: Association for Computation...

  124. [132]

    Faithfulness Measurable Masked Language Models,

    ——, “Faithfulness Measurable Masked Language Models,” inForty-first International Conference on Machine Learning, 2024. [Online]. Available: https://openreview.net/for um?id=tw1PwpuAuNhttp://arxiv.org/abs/2310.07819

  125. [133]

    Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis,

    K. Zhang and S. Bowman, “Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis,” inProceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. Stroudsburg, PA, USA: Associ...

  126. [134]

    Designing and Interpreting Probes with Control Tasks,

    J. Hewitt and P. Liang, “Designing and Interpreting Probes with Control Tasks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Stroudsburg, PA, ...

  127. [135]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville,Deep Learning. MIT Press, 2016

  128. [136]

    Attention is all you need,

    A. Vaswaniet al., “Attention is all you need,” inAdvances in Neural Information Processing Systems, vol. 2017-Decem. Association for Computational Linguistics (ACL), 6 2017, pp. 5999–6009. [Online]. Available: http://arxiv.org/abs/1706.03762

  129. [137]

    Graves,Supervised Sequence Labelling with Recurrent Neural Networks, ser

    A. Graves,Supervised Sequence Labelling with Recurrent Neural Networks, ser. Studies in Computational Intelligence. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, vol. 385. [Online]. Available: https://link.springer.com/10.1007/978-3-642-24797-2

  130. [138]

    Speech and Language Processing,

    D. Jurafsky and J. Martin, “Speech and Language Processing,”Speech and Language Processing., vol. 3, pp. 441–458, 2014. 116

  131. [139]

    Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),

    A. Adadi and M. Berrada, “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),”IEEE Access, vol. 6, pp. 52138–52160, 2018. [Online]. Available: https://ieeexplore.ieee.org/document/8466590/

  132. [140]

    Are self-explanations from Large Language Models faithful?

    A. Madsen, S. Chandar, and S. Reddy, “Are self-explanations from Large Language Models faithful?” The 62nd Annual Meeting of the Association for Computational Linguistics, 12024.[Online].Available: https://openreview.net/forum?id=0fB5OROAIq

  133. [141]

    Explanation in artificial intelligence: Insights from the social sciences,

    T. Miller, “Explanation in artificial intelligence: Insights from the social sciences,” Artificial Intelligence, vol. 267, pp. 1–38, 2 2019. [Online]. Available: http://arxiv.org/ abs/1706.07269https://linkinghub.elsevier.com/retrieve/pii/S0004370218305988

  134. [142]

    Attention is not not Explanation,

    S. Wiegreffe and Y. Pinter, “Attention is not not Explanation,”Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 11–20, 8 2019. [Online]. Availabl...

  135. [143]

    AXIS: Generating Explanations at Scale with Learnersourcing and Machine Learning,

    J. J. Williamset al., “AXIS: Generating Explanations at Scale with Learnersourcing and Machine Learning,” in Proceedings of the Third (2016) ACM Conference on Learning @ Scale. New York, NY, USA: ACM, 4 2016, pp. 379–388. [Online]. Available: https://dl.acm.org/doi/10.1145/287...

  136. [144]

    Robnik-Šikonja and M

    M. Robnik-Šikonja and M. Bohanec,Perturbation-Based Explanations of Prediction Models. Springer International Publishing, 2018. [Online]. Available: http: //dx.doi.org/10.1007/978-3-319-90403-0_9

  137. [145]

    Reading Tea Leaves: How Humans Interpret Topic Models,

    J. Chang et al., “Reading Tea Leaves: How Humans Interpret Topic Models,” in Advances in Neural Information Processing Systems, Y. Bengioet al., Eds., vol. 22. Curran Associates, Inc., 2009, pp. 288–296. [Online]. Available: https://proceedings.ne urips.cc/paper/2009/file/f925...

  138. [146]

    Rotated Word Vector Representations and their Interpretability,

    S. Park, J. Bak, and A. Oh, “Rotated Word Vector Representations and their Interpretability,” inProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2017, pp. 401–411. [Online]....

  139. [147]

    Molnar,Interpretable Machine Learning

    C. Molnar,Interpretable Machine Learning. Independent, 2019. [Online]. Available: https://christophm.github.io/interpretable-ml-book/

  140. [148]

    The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations,

    A. Chatzimparmpas et al. , “The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations,” Computer Graphics Forum, vol. 39, no. 3, pp. 713–756, 6 2020. [Online]. Available: https://onlinelibrary.wiley.com/doi/10.1111/cgf.14034

  141. [149]

    Did the model understand the question?

    P. K. Mudrakartaet al., “Did the model understand the question?” inACL 2018 - 56th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers), vol. 1, 5 2018, pp. 1896–1906. [Online]. Available: https://www.aclweb.org/anthology...

  142. [150]

    Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models,

    T. Wuet al., “Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume ...

  143. [151]

    A value for N-Person Games,

    Shapley, “A value for N-Person Games,” Contributions to the Theory of Games (AM-28), Volume II , pp. 307–317, 1953. [Online]. Available: https: //apps.dtic.mil/dtic/tr/fulltext/u2/604084.pdf

  144. [152]

    Molnar,Interpreting Machine Learning Models With SHAP, 2023

    C. Molnar,Interpreting Machine Learning Models With SHAP, 2023

  145. [153]

    Quantifying Attention Flow in Transformers,

    S. Abnar and W. Zuidema, “Quantifying Attention Flow in Transformers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 4190–4197. [Online]. Available: https:/...

  146. [154]

    Techniques for interpretable machine learning,

    M. Du, N. Liu, and X. Hu, “Techniques for interpretable machine learning,” Communications of the ACM, vol. 63, no. 1, pp. 68–77, 12 2019. [Online]. Available: https://dl.acm.org/doi/10.1145/3359786 117

  147. [155]

    A causal framework for explaining the predictions of black-box sequence-to-sequence models,

    D. Alvarez-Melis and T. Jaakkola, “A causal framework for explaining the predictions of black-box sequence-to-sequence models,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Lingui...

  148. [156]

    On Identifiability in Transformers,

    G. Brunner et al. , “On Identifiability in Transformers,” in International Conference on Learning Representations (ICLR 2020), 8 2020. [Online]. Available: https://openreview.net/forum?id=BJg1f6EFDB 118

  149. [157]

    Staying True to Your Word: (How) Can Attention Become Explanation?

    M. Tutek and J. Snajder, “Staying True to Your Word: (How) Can Attention Become Explanation?” in Proceedings of the 5th Workshop on Representation Learning for NLP. Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 131–142. [Online]. Available: https:/...

  150. [158]

    Rethinking the Role of Gradient-Based Attribution Methods for Model Interpretability,

    S. Srinivas and F. Fleuret, “Rethinking the Role of Gradient-Based Attribution Methods for Model Interpretability,”ICLR 2021 - 9th International Conference on Learning Representations, 2021

  151. [159]

    Evaluating Models’ Local Decision Boundaries via Contrast Sets,

    M. Gardneret al., “Evaluating Models’ Local Decision Boundaries via Contrast Sets,” in Findings of the Association for Computational Linguistics: EMNLP 2020. Stroudsburg, PA, USA: Association for Computational Linguistics, 4 2020, pp. 1307–1323. [Online]. Available: https://ww...

  152. [160]

    Learning The Difference That Makes A Difference With Counterfactually-Augmented Data,

    D. Kaushik, E. Hovy, and Z. C. Lipton, “Learning The Difference That Makes A Difference With Counterfactually-Augmented Data,” in International Conference on Learning Representations , 2020. [Online]. Available: https: //openreview.net/forum?id=Sklgs0NFvr

  153. [161]

    T. H. Cormenet al., Introduction to Algorithms, Third Edition, 3rd ed. The MIT Press, 2009

  154. [162]

    Attention Flows are Shapley Value Explanations,

    K. Ethayarajh and D. Jurafsky, “Attention Flows are Shapley Value Explanations,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Stro...

  155. [163]

    WinoGrande: An Adversarial Winograd Schema Challenge at Scale,

    K. Sakaguchiet al., “WinoGrande: An Adversarial Winograd Schema Challenge at Scale,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 05, pp. 8732–8740, 4 2020. [Online]. Available: https://aaai.org/ojs/index.php/AAAI/articl e/view/6399

  156. [164]

    ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations,

    J. Wieting and K. Gimpel, “ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg, PA, USA: Assoc...

  157. [165]

    Rationalizing Neural Predictions,

    T. Lei, R. Barzilay, and T. Jaakkola, “Rationalizing Neural Predictions,” inProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . Stroudsburg, PA, USA: Association for Computational Linguistics, 2016, pp. 107–117. [Online]. Available: http://...

  158. [166]

    e-SNLI: Natural Language Inference with Natural Language Explanations,

    O.-M. Camburuet al., “e-SNLI: Natural Language Inference with Natural Language Explanations,” inAdvances in Neural Information Processing Systems, vol. 2018-Decem, 12 2018, pp. 9539–9549. [Online]. Available: http://arxiv.org/abs/1812.01193

  159. [167]

    Towards Explainable NLP: A Generative Explanation Framework for Text Classification,

    H. Liu, Q. Yin, and W. Y. Wang, “Towards Explainable NLP: A Generative Explanation Framework for Text Classification,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 20...

  160. [168]

    Language models are unsupervised multitask learners,

    A. Radfordet al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019. [Online]. Available: https://openai.com/blog/better-languag e-models/

  161. [169]

    PAWS: Paraphrase adversaries from word scrambling,

    Y. Zhang, J. Baldridge, and L. He, “PAWS: Paraphrase adversaries from word scrambling,” in Proceedings of the 2019 Conference of the North. Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 1298–1308. [Online]. Available: http://aclweb.org/anthology/N19-1131

  162. [170]

    RerrFact: Reduced Evidence Retrieval Representations for Scientific Claim Verification,

    A. Ranaet al., “RerrFact: Reduced Evidence Retrieval Representations for Scientific Claim Verification,” in CEUR Workshop Proceedings, vol. 3164, 2 2022, pp. 3–7. [Online]. Available: http://arxiv.org/abs/2202.02646

  163. [171]

    ERASER: A Benchmark to Evaluate Rationalized NLP Models,

    J. DeYounget al., “ERASER: A Benchmark to Evaluate Rationalized NLP Models,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 11 2020, pp. 4443–4458. [Online]. Available...

  164. [172]

    Improving Language Understanding by Generative Pre-Training,

    A. Radfordet al., “Improving Language Understanding by Generative Pre-Training,” OpenAI, 2018. [Online]. Available: https://openai.com/blog/language-unsupervised/

  165. [173]

    Recursive deep models for semantic compositionality over a sentiment treebank,

    R. Socheret al., “Recursive deep models for semantic compositionality over a sentiment treebank,” EMNLP 2013 - 2013 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, pp. 1631–1642, 2013

  166. [174]

    Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?

    P. Hase et al., “Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?” Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020, pp. 4351–4367, 2020. [Online]. Available: https://www.a...

  167. [175]

    Explaining Question Answering Models through Text Generation,

    V. Latcinnik and J. Berant, “Explaining Question Answering Models through Text Generation,” arXiv, 4 2020. [Online]. Available: http://arxiv.org/abs/2004.05569

  168. [176]

    Rationalization for explainable NLP: a survey,

    S. Gurrapuet al., “Rationalization for explainable NLP: a survey,”Frontiers in Artificial Intelligence, vol. 6, 2023

  169. [177]

    Faithfulness Tests for Natural Language Explanations,

    P. Atanasova et al., “Faithfulness Tests for Natural Language Explanations,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , vol. 2. Stroudsburg, PA, USA: Association for Computational Linguistics, 5 2023, p...

  170. [178]

    Measuring Association Between Labels and Free-Text Rationales,

    S. Wiegreffe, A. Marasović, and N. A. Smith, “Measuring Association Between Labels and Free-Text Rationales,”Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 10266–10284, 2020. [Online]. Available: http://arxiv.org/abs/2010.12762https...

  171. [179]

    Fact Checking with Insufficient Evidence,

    P. Atanasovaet al., “Fact Checking with Insufficient Evidence,”Transactions of the Association for Computational Linguistics, vol. 10, pp. 746–763, 7 2022. [Online]. Available: https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00486/112498/Fa ct-Checking-with-Insufficient...

  172. [180]

    Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting,

    M. Turpinet al., “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting,” in Thirty-seventh Conference on Neural Information Processing Systems , 5 2023, pp. 1–32. [Online]. Available: http://arxiv.org/abs/2305.04388https://ope...

  173. [181]

    Measuring Faithfulness in Chain-of-Thought Reasoning,

    T. Lanham et al., “Measuring Faithfulness in Chain-of-Thought Reasoning,”arXiv,

  174. [182]

    Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing,

    S. Wiegreffe and A. Marasović, “Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing,” in 35th Conference on Neural Information Processing Systems (NeurIPS 2021) , 2 2021. [Online]. Available: http://arxiv.org/abs/2102.12060

  175. [183]

    Translating neuralese,

    J. Andreas, A. Dragan, and D. Klein, “Translating neuralese,” inACL 2017 - 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers), 2017, pp. 232–242. [Online]. Available: http://github

  176. [184]

    e-SNLI-VE-2.0: Corrected Visual-Textual Entailment with Natural Language Explanations,

    V. Do et al., “e-SNLI-VE-2.0: Corrected Visual-Textual Entailment with Natural Language Explanations,”IEEE CVPR Workshop on Fair, Data Efficient and Trusted Computer Vision, 2020, 2020. [Online]. Available: https://github.com/

  177. [185]

    Learning to Deceive with Attention-Based Explanations,

    D. Pruthi et al., “Learning to Deceive with Attention-Based Explanations,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 4782–4793. [Online]. Available: htt...

  178. [186]

    CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations,

    L. Arras, A. Osman, and W. Samek, “CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations,”Information Fusion, vol. 81, pp. 14–40, 5 2022. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S15 66253521002335

  179. [187]

    A Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference,

    K. Zaman and Y. Belinkov, “A Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, 4 2022, pp. 1556–1576. [Online]. Available:...

  180. [188]

    Double Trouble: How to not Explain a Text Classifier’s Decisions Using Counterfactuals Synthesized by Masked Language Models?

    T. M. Pham et al., “Double Trouble: How to not Explain a Text Classifier’s Decisions Using Counterfactuals Synthesized by Masked Language Models?” in Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th Int...

  181. [189]

    Available: http://arxiv.org/abs/2307.13702

    [Online]. Available: http://arxiv.org/abs/2307.13702

  182. [190]

    On Measuring Faithfulness of Natural Language Explanations,

    L. Parcalabescu and A. Frank, “On Measuring Faithfulness of Natural Language Explanations,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2311.07466

  183. [191]

    Are Training Resources Insufficient? Predict First Then Explain!

    M. Jang and T. Lukasiewicz, “Are Training Resources Insufficient? Predict First Then Explain!” arXiv, 8 2021. [Online]. Available: http://arxiv.org/abs/2110.02056 121

  184. [192]

    Should You Mask 15% in Masked Language Modeling?

    A. Wettig et al., “Should You Mask 15% in Masked Language Modeling?”EACL 2023 - 17th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, pp. 2977–2992, 2 2023. [Online]. Available: http://arxiv.org/abs/2202.08005

  185. [193]

    An Improved Bonferroni Procedure for Multiple Tests of Significance,

    R. J. Simes, “An Improved Bonferroni Procedure for Multiple Tests of Significance,” Biometrika, vol. 73, no. 3, p. 751, 12 1986. [Online]. Available: https://www.jstor.org/stable/2336545?origin=crossref

  186. [194]

    Layer Normalization,

    J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,”Arxiv, 2016. [Online]. Available: http://arxiv.org/abs/1607.06450

  187. [195]

    Statistical Methods for Research Workers,

    R. A. Fisher, “Statistical Methods for Research Workers,” in Breakthroughs in Statistics: Methodology and Distribution , S. Kotz and N. L. Johnson, Eds. New York, NY: Springer New York, 1992, pp. 66–70. [Online]. Available: http://link.springer.com/10.1007/978-1-4612-4380-9_6

  188. [196]

    Bootstrap Methods and Their Application,

    S. T. Buckland, A. C. Davison, and D. V. Hinkley, “Bootstrap Methods and Their Application,” Biometrics, vol. 54, no. 2, p. 795, 6 1998

  189. [197]

    Annotation Artifacts in Natural Language Inference Data,

    S. Gururanganet al., “Annotation Artifacts in Natural Language Inference Data,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), vol. 2. Stroudsburg, PA, ...

  190. [198]

    The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations,

    P. Hase, H. Xie, and M. Bansal, “The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations,”Advances in Neural Information Processing Systems, vol. 5, no. NeurIPS, pp. 3650–3666, 2021

  191. [199]

    Rationales for Sequential Predictions,

    K. Vafa et al., “Rationales for Sequential Predictions,” inProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, 122 USA: Association for Computational Linguistics, 2021, pp. 10314–10332. [Online]. Available: https://aclanthol...

  192. [200]

    KS(conf): A Light-Weight Test if a Multiclass Classifier Operates Outside of Its Specifications,

    R. Sun and C. H. Lampert, “KS(conf): A Light-Weight Test if a Multiclass Classifier Operates Outside of Its Specifications,” International Journal of Computer Vision, vol. 128, no. 4, pp. 970–995, 4 2020. [Online]. Available: http://link.springer.com/10.1007/s11263-019-01232-x

  193. [201]

    $p$-DkNN: Out-of-Distribution Detection Through Statistical Testing of Deep Representations,

    A. Dziedzic et al., “$p$-DkNN: Out-of-Distribution Detection Through Statistical Testing of Deep Representations,” arXiv, 7 2022. [Online]. Available: http: //arxiv.org/abs/2207.12545 123

  194. [202]

    Backpropagated Gradient Representations for Anomaly Detection,

    G. Kwonet al., “Backpropagated Gradient Representations for Anomaly Detection,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 12366 LNCS, pp. 206–226, 7

  195. [203]

    Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey,

    B. Minet al., “Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey,”ACM Computing Surveys, vol. 56, no. 2, pp. 1–40, 2

  196. [204]

    Applications of transformer-based language models in bioinformatics: a survey,

    S. Zhanget al., “Applications of transformer-based language models in bioinformatics: a survey,” Bioinformatics Advances, vol. 3, no. 1, 1 2023. [Online]. Available: https://academic.oup.com/bioinformaticsadvances/article/doi/10.1093/bioadv/vbad0 01/6984737

  197. [205]

    Chernick and R

    Michael R. Chernick and R. A. LaBudde,An introduction to bootstrap methods with applications to R. John Wiley & Sons, 2011

  198. [206]

    SAM: The Sensitivity of Attribution Methods to Hyperparameters,

    N. Bansal, C. Agarwal, and A. Nguyen, “SAM: The Sensitivity of Attribution Methods to Hyperparameters,” in2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). IEEE, 6 2020, pp. 11–21. [Online]. Available: https://ieeexplore.ieee.org/document/9150607/

  199. [207]

    Generalized Out-of-Distribution Detection: A Survey,

    J. Yanget al., “Generalized Out-of-Distribution Detection: A Survey,”arXiv, 10 2021. [Online]. Available: http://arxiv.org/abs/2110.11334

  200. [208]

    The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only,

    G. Penedoet al., “The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2306.01116

  201. [209]

    Mistral 7B,

    A. Q. Jiang et al., “Mistral 7B,” arXiv, pp. 1–9, 2023. [Online]. Available: http://arxiv.org/abs/2310.06825

  202. [210]

    GPT-4 Technical Report,

    OpenAI, “GPT-4 Technical Report,” OpenAI, vol. 4, pp. 1–100, 3 2023. [Online]. Available: http://arxiv.org/abs/2303.08774

  203. [211]

    A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity,

    Y. Bang et al., “A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2302.04023

  204. [212]

    LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples,

    J.-Y. Yaoet al., “LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples,” arXiv, pp. 1–13, 2023. [Online]. Available: http://arxiv.org/abs/2310.01469 124

  205. [213]

    Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models,

    C. Agarwal, S. H. Tanneru, and H. Lakkaraju, “Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models,”arXiv, 2024. [Online]. Available: http://arxiv.org/abs/2402.04614

  206. [214]

    Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations,

    Y. Chen et al., “Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations,” arXiv, 2023. [Online]. Available: http: //arxiv.org/abs/2307.08678

  207. [215]

    Generative Representational Instruction Tuning,

    N. Muennighoffet al., “Generative Representational Instruction Tuning,”arXiv, 2024. [Online]. Available: http://arxiv.org/abs/2402.09906

  208. [216]

    ROUGE: A Package for Automatic Evaluation of Summaries,

    C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” inText Summarization Branches Out. Barcelona, Spain: Association for Computational Linguistics, 7 2004, pp. 74–81. [Online]. Available: https://aclanthology.org/W04-1013

  209. [217]

    Llama 2: Open Foundation and Fine-Tuned Chat Models,

    Meta, “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2307.09288

  210. [218]

    Language Models (Mostly) Know What They Know,

    Anthropic Team, “Language Models (Mostly) Know What They Know,”Anthropic, 7

  211. [219]

    Copy Suppression: Comprehensively Understanding an Attention Head,

    C. McDougallet al., “Copy Suppression: Comprehensively Understanding an Attention Head,” in NeurIPS 2023 Workshop on Attributing Model Behavior at Scale, 2023. [Online]. Available: http://arxiv.org/abs/2310.04625

  212. [220]

    Toxicity in chatgpt: Analyzing persona-assigned language models,

    A. Deshpande et al., “Toxicity in chatgpt: Analyzing persona-assigned language models,” inFindings of the Association for Computational Linguistics: EMNLP 2023. Stroudsburg, PA, USA: Association for Computational Linguistics, 2023, pp. 1236–1270. [Online]. Available: https://a...

  213. [221]

    Benchmarking and Improving Generator-Validator Consistency of Language Models,

    X. L. Li et al., “Benchmarking and Improving Generator-Validator Consistency of Language Models,” arXiv, pp. 1–15, 2023. [Online]. Available: http: //arxiv.org/abs/2310.01846

  214. [222]

    Prompt-based methods may underestimate large language models’ linguistic generalizations,

    J. Hu and R. Levy, “Prompt-based methods may underestimate large language models’ linguistic generalizations,” arXiv, 2023. [Online]. Available: http: //arxiv.org/abs/2305.13264 125

  215. [223]

    A Survey on In-context Learning,

    Q. Donget al., “A Survey on In-context Learning,”arXiv, 12 2022. [Online]. Available: http://arxiv.org/abs/2301.00234

  216. [224]

    Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing Their Input Gradients,

    A. Ross and F. Doshi-Velez, “Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing Their Input Gradients,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, pp. 1660–1669, 4 2018. [Online]. Available: http...

  217. [225]

    Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations,

    S. Huang et al. , “Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations,” arXiv, 2023. [Online]. Available: http://arxiv.org/abs/2310.11207

  218. [226]

    Rethinking Interpretability in the Era of Large Language Models,

    C. Singhet al., “Rethinking Interpretability in the Era of Large Language Models,” arXiv, 2024. [Online]. Available: http://arxiv.org/abs/2402.01761

  219. [227]

    Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment,

    Z. Li et al. , “Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment,” in Findings of the Association for Computational Linguistics: ACL 2022. Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, ...

  220. [228]

    Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?

    P. Hase and M. Bansal, “Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 20...

  221. [229]

    To what extent do human explanations of model behavior align with actual model behavior?

    G. Prasad et al., “To what extent do human explanations of model behavior align with actual model behavior?” in Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP. Stroudsburg, PA, USA: Association for Computational Linguistics...

  222. [230]

    On the Interaction of Belief Bias and Explanations,

    A. V. González, A. Rogers, and A. Søgaard, “On the Interaction of Belief Bias and Explanations,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 2930–2942. [Online]. Avail...

  223. [231]

    Human Interpretation of Saliency-based Explanation Over Text,

    H. Schuff et al., “Human Interpretation of Saliency-based Explanation Over Text,” in 2022 ACM Conference on Fairness, Accountability, and Transparency. 126 New York, NY, USA: ACM, 6 2022, pp. 611–636. [Online]. Available: https://dl.acm.org/doi/10.1145/3531146.3533127

  224. [232]

    Human-grounded Evaluations of Explanation Methods for Text Classification,

    P. Lertvittayakumjorn and F. Toni, “Human-grounded Evaluations of Explanation Methods for Text Classification,” inProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (E...

  225. [233]

    Comparing Automatic and Human Evaluation of Local Explanations for Text Classification,

    D. Nguyen, “Comparing Automatic and Human Evaluation of Local Explanations for Text Classification,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), vol. ...

  226. [234]

    Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero,

    L. Schut et al., “Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero,” arXiv, pp. 1–61, 10 2023. [Online]. Available: http://arxiv.org/abs/2310.16410

  227. [235]

    Beyond interpretability: developing a language to shape our relationships with AI,

    B. Kim, “Beyond interpretability: developing a language to shape our relationships with AI,” inThe International Conference on Learning Representations, 2022. [Online]. Available: https://iclr.cc/Conferences/2022/Schedule?showEvent=7237

  228. [236]

    Efficiently Training Low-Curvature Neural Networks,

    S. Srinivaset al., “Efficiently Training Low-Curvature Neural Networks,”Advances in Neural Information Processing Systems, vol. 35, no. NeurIPS, pp. 1–21, 6 2022. [Online]. Available: http://arxiv.org/abs/2206.07144

  229. [237]

    How Interpretable are Reasoning Explanations from Prompting Large Language Models?

    W. J. Yeoet al., “How Interpretable are Reasoning Explanations from Prompting Large Language Models?” arXiv, 2 2024. [Online]. Available: http://arxiv.org/abs/2402.11863

  230. [238]

    Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?

    C. Sen et al., “Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 20...

  231. [239]

    Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads,

    T. Caiet al., “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads,”arXiv, 2024. [Online]. Available: http://arxiv.org/abs/2401.10774

  232. [240]

    StereoSet: Measuring stereotypical bias in pre- trained language models,

    M. Nadeem, A. Bethke, and S. Reddy, “StereoSet: Measuring stereotypical bias in pre- trained language models,”ACL-IJCNLP 2021 - 59th Annual Meeting of the Association 127 for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ...

  233. [241]

    CrowS-Pairs: A challenge dataset for measuring social biases in masked language models,

    N. Nangia et al., “CrowS-Pairs: A challenge dataset for measuring social biases in masked language models,”EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, pp. 1953–1967, 2020

  234. [242]

    Fairwashing: The risk of rationalization,

    U. Aïvodjiet al., “Fairwashing: The risk of rationalization,”36th International Confer- ence on Machine Learning, ICML 2019, vol. 2019-June, pp. 240–252, 2019

  235. [243]

    Characterizing the risk of fairwashing,

    ——, “Characterizing the risk of fairwashing,”Advances in Neural Information Process- ing Systems, vol. 18, no. NeurIPS, pp. 14822–14834, 2021

  236. [244]

    Improving alignment of dialogue agents via targeted human judgements,

    A. Glaese et al., “Improving alignment of dialogue agents via targeted human judgements,” arXiv, pp. 1–77, 2022. [Online]. Available: http://arxiv.org/abs/2209.143 75

  237. [245]

    Towards a Robust Deep Neural Network against Adversarial Texts: A Survey,

    W. Wanget al., “Towards a Robust Deep Neural Network against Adversarial Texts: A Survey,” IEEE Transactions on Knowledge and Data Engineering, pp. 1–1, 2021. [Online]. Available: https://ieeexplore.ieee.org/document/9557814/

  238. [246]

    HotFlip: White-Box Adversarial Examples for Text Classification,

    J. Ebrahimiet al., “HotFlip: White-Box Adversarial Examples for Text Classification,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, vol. 2. Stroudsburg, PA, USA: Association for Computational Linguistics, 2018, pp. 31–36. [Online]....

  239. [247]

    Blockwise parallel decoding for deep autore- gressive models,

    M. Stern, N. Shazeer, and J. Uszkoreit, “Blockwise parallel decoding for deep autore- gressive models,”Advances in Neural Information Processing Systems, vol. 2018-Decem, no. Nips, pp. 10086–10095, 2018

  240. [248]

    Accelerating LLM Inference with Staged Speculative Decoding,

    B. Spector and C. Re, “Accelerating LLM Inference with Staged Speculative Decoding,” arXiv, no. Llm, 2023. [Online]. Available: http://arxiv.org/abs/2308.04623

  241. [249]

    Break the Sequential Dependency of LLM Inference Using Lookahead Decoding,

    Y. Fuet al., “Break the Sequential Dependency of LLM Inference Using Lookahead Decoding,” arXiv, 2024. [Online]. Available: http://arxiv.org/abs/2402.02057

  242. [250]

    Understanding Black-box Predictions via Influence Functions,

    P. W. Koh and P. Liang, “Understanding Black-box Predictions via Influence Functions,”34th International Conference on Machine Learning, ICML 2017, vol. 4, pp. 2976–2987, 3 2017. [Online]. Available: http://arxiv.org/abs/1703.04730

  243. [251]

    Representer Point Selection for Explaining Deep Neural Networks,

    C.-K. Yehet al., “Representer Point Selection for Explaining Deep Neural Networks,” in Advances in Neural Information Processing Systems, 11 2018, pp. 9291–9301. [Online]. Available: http://arxiv.org/abs/1811.09720

  244. [252]

    Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions,

    X. Han, B. C. Wallace, and Y. Tsvetkov, “Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions,” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational L...

  245. [253]

    FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging,

    H. Guo et al. , “FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging,” inProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 12 2021, pp. 1033...

  246. [254]

    A Generalized Representer Theorem,

    B. Schölkopf, R. Herbrich, and A. J. Smola, “A Generalized Representer Theorem,” in International Conference on Computational Learning Theory. Springer, 2001, pp. 416–426. [Online]. Available: http://link.springer.com/10.1007/3-540-44581-1_27

  247. [255]

    Estimating Training Data Influence by Tracing Gradient Descent,

    G. Pruthiet al., “Estimating Training Data Influence by Tracing Gradient Descent,” in Advances in Neural Information Processing Systems, 2 2020. [Online]. Available: http://arxiv.org/abs/2002.08484

  248. [256]

    Explaining classifiers with causal concept effect (CaCE),

    Y. Goyal, U. Shalit, and B. Kim, “Explaining classifiers with causal concept effect (CaCE),” arXiv, 7 2019. [Online]. Available: http://arxiv.org/abs/1907.07165

  249. [257]

    Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV),

    B. Kim et al., “Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV),” 35th International Conference on Machine Learning, ICML 2018, vol. 6, pp. 4186–4195, 11 2018. [Online]. Available: http://arxiv.org/abs/1711.11279

  250. [258]

    Universal Adversarial Triggers for Attacking and Analyzing NLP,

    E. Wallaceet al., “Universal Adversarial Triggers for Attacking and Analyzing NLP,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Stroudsburg, ...

  251. [259]

    Papineni et al

    K. Papineni et al. , “BLEU,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02 . Morristown, NJ, USA: Association for Computational Linguistics, 2001. [Online]. Available: http://portal.acm.org/citation.cfm?doid=1073083.1073135

  252. [260]

    Characterizations of an Empirical Influence Function for Detecting Influential Cases in Regression,

    R. D. Cook and S. Weisberg, “Characterizations of an Empirical Influence Function for Detecting Influential Cases in Regression,”Technometrics, vol. 22, no. 4, pp. 495–508, 11 1980. [Online]. Available: http://www.tandfonline.com/doi/abs/10.1080/00401706.1 980.10486199 128

  253. [261]

    Efficient estimation of word representations in vector space,

    T. Mikolovet al., “Efficient estimation of word representations in vector space,” in1st International Conference on Learning Representations, ICLR 2013 - Workshop Track Proceedings, 2013. [Online]. Available: http://ronan.collobert.com/senna/

  254. [262]

    Man is to computer programmer as woman is to homemaker? Debiasing word embeddings,

    T. Bolukbasi et al., “Man is to computer programmer as woman is to homemaker? Debiasing word embeddings,” inAdvances in Neural Information Processing Systems, 2016, pp. 4356–4364

  255. [263]

    Best practices in exploratory factor analysis: Four recommendations for getting the most from your analysis,

    A. B. Costello and J. W. Osborne, “Best practices in exploratory factor analysis: Four recommendations for getting the most from your analysis,”Practical Assessment, Research and Evaluation, vol. 10, no. 7, pp. 1–9, 2005

  256. [264]

    A general rotation criterion and its use in orthogonal rotation,

    C. B. Crawford and G. A. Ferguson, “A general rotation criterion and its use in orthogonal rotation,” Psychometrika, vol. 35, no. 3, pp. 321–332, 9 1970. [Online]. Available: http://link.springer.com/10.1007/BF02310792

  257. [265]

    Latent Dirichlet allocation,

    D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet allocation,” Journal of Machine Learning Research, vol. 3, pp. 993–1022, 2003. [Online]. Available: https://jmlr.org/papers/v3/blei03a.html

  258. [266]

    Global Explanations of Neural Networks,

    M. Ibrahimet al., “Global Explanations of Neural Networks,” inProceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. New York, NY, USA: ACM, 1 2019, pp. 279–287. [Online]. Available: https://dl.acm.org/doi/10.1145/3306618.3314230

  259. [267]

    Model Agnostic Multilevel Explanations,

    K. N. Ramamurthyet al., “Model Agnostic Multilevel Explanations,”arXiv, 3 2020. [Online]. Available: http://arxiv.org/abs/2003.06005

  260. [268]

    Are Sixteen Heads Really Better than One?

    P. Michel, O. Levy, and G. Neubig, “Are Sixteen Heads Really Better than One?” Advances in Neural Information Processing Systems, vol. 32, pp. 1–13, 5 2019. [Online]. Available: http://arxiv.org/abs/1905.10650

  261. [269]

    Compositional Explanations of Neurons,

    J. Mu and J. Andreas, “Compositional Explanations of Neurons,” in Advances in Neural Information Processing Systems , 6 2020. [Online]. Available: http: //arxiv.org/abs/2006.14032 129

  262. [270]

    Direct and Indirect Effects,

    J. Pearl, “Direct and Indirect Effects,” in Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, ser. UAI’01. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2001, p. 411–420. [Online]. Available: https://dl.acm.org/doi/10.5555/2074022.2074073

  263. [271]

    Towards automatic concept-based explanations,

    A. Ghorbaniet al., “Towards automatic concept-based explanations,” inAdvances in Neural Information Processing Systems, vol. 32, 2019

  264. [272]

    Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks,

    Y. Adi et al., “Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks,” inInternational Conference on Learning Representations (ICLR), 8 2017, pp. 1–12. [Online]. Available: http://arxiv.org/abs/1608.04207

  265. [273]

    Natural language multitasking analyzing and improving syntactic saliency of latent representations,

    G. Brunneret al., “Natural language multitasking analyzing and improving syntactic saliency of latent representations,” in31st Conference on Neural Information Processing Systems (NIPS 2017), 1 2017. [Online]. Available: http://arxiv.org/abs/1801.06024

  266. [274]

    What’s in an Embedding? Analyzing Word Embeddings through Multilingual Evaluation,

    A. Köhn, “What’s in an Embedding? Analyzing Word Embeddings through Multilingual Evaluation,” inProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2015, pp. 2067–2073. [Online...

  267. [275]

    What do you learn from context? Probing for sentence structure in contextualized word representations,

    I. Tenney et al., “What do you learn from context? Probing for sentence structure in contextualized word representations,” in7th International Conference on Learning Representations, ICLR 2019 , 2019, pp. 1–17. [Online]. Available: https://openreview.net/forum?id=SJzSgnRcKX

  268. [276]

    Anchors: High-precision model-agnostic explanations,

    M. T. Ribeiro, S. Singh, and C. Guestrin, “Anchors: High-precision model-agnostic explanations,” in32nd AAAI Conference on Artificial Intelligence, AAAI 2018, 2018, pp. 1527–1535. [Online]. Available: https://www.aaai.org/ocs/index.php/AAAI/AAAI 18/paper/view/16982 131 APPENDI...

  269. [280]

    Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies,

    T. Linzen, E. Dupoux, and Y. Goldberg, “Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies,” Transactions of the Association for 130 Computational Linguistics, vol. 4, no. 1990, pp. 521–535, 12 2016. [Online]. Available: https://direct.mit.edu/tacl/article/43378

  270. [281]

    UnNatural Language Inference,

    K. Sinhaet al., “UnNatural Language Inference,” inACL 2021 - 59th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, 2021. [Online]. Available: http://arxiv.org/abs/2101.00010

  271. [282]

    Does string-based neural MT learn source syntax?

    X. Shi, I. Padhi, and K. Knight, “Does string-based neural MT learn source syntax?” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Stroudsburg, PA, USA: Association for Computational Linguistics, 2016, pp. 1526–1534. [Online]. Availa...

  272. [288]

    The nurse said that

    first introduce an idealized version of this, which assumes optimization is done on one observation at a time (for example, SGD): TracInIdeal(˜x, x) = ∑ t∈T˜x L(y, x,θt)−L (y, x,θt+1),whereT˜x is timestep which optimized˜x (A.8) TracIn TracIn is then a relaxation of this ideal...

  273. [289]

    Analog to these papers, a few methods use cluster algorithms instead of logistic regression [273]

    have been using similar linguistic tasks and MLP probes but have extended previous analyses to multiple models and training methods. Analog to these papers, a few methods use cluster algorithms instead of logistic regression [273]. Additionally, some methods only look atword e...

  274. [290]

    In that paper, they find probes can achieve high accuracy from an untrained model unless the auxiliary dataset size is dramatically decreased

    suggest learning a probe from an untrained model as a baseline. In that paper, they find probes can achieve high accuracy from an untrained model unless the auxiliary dataset size is dramatically decreased. Similarly, Hewitt and Liang[127] use randomized datasets as a baseline...

  275. [291]

    Groundedness Because the category ofrule explanations can be very diverse,groundedness evaluation would likely depend on the specific explanation method

    be modified towardsglobal explanation, in which case it would be arule explanation. Groundedness Because the category ofrule explanations can be very diverse,groundedness evaluation would likely depend on the specific explanation method. However, generally faithfulness can be ...

  276. [292]

    No masking

    works. To compute the confidence inter- val, a dataset-aggregation is done for each seed, such that the all-observation are i.i.d.. Because some seeds do not converge for some datasets, such as bAbI-2 and bAbI-3 (as men- tioned in Section 4.2.1), those outliers and not include...

  277. [293]

    positive

    Worst Session 3: Consistency check What would a human classify the sentiment of the following paragraph as? The paragraph can contain redacted words marked with [REDACTED]. Answer only "positive", "negative", "neutral", or "unknown". Do not explain the answer. Paragraph: Ned a...

  278. [294]

    positive

    Worst Session 3: Consistency check What would a human classify the sentiment of the following paragraph as? The paragraph can contain removed words marked with [REMOVED]. Answer only "positive", "negative", "neutral", or "unknown". Do not explain the answer. Paragraph: Ned aKe...

  279. [295]

    Where is Mary?

    Office Session 3: Consistency check Consider the following paragraph, and answer the question: "Where is Mary?" The paragraph can contain redacted words marked with [REDACTED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be answered. Do not...

  280. [296]

    Where is Mary?

    Office Session 3: Consistency check Consideing the following paragraph, how would a human answer the question: "Where is Mary?" The paragraph can contain redacted words marked with [REDACTED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be ...

  281. [297]

    Where is Mary?

    Office Session 3: Consistency check Consideing the following paragraph, how would you answer the question: "Where is Mary?" The paragraph can contain redacted words marked with [REDACTED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be answ...

  282. [298]

    Where is Mary?

    Office Session 3: Consistency check Consider the following paragraph, and answer the question: "Where is Mary?" The paragraph can contain removed words marked with [REMOVED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be answered. Do not e...

  283. [299]

    Where is Mary?

    Office Session 3: Consistency check Consideing the following paragraph, how would a human answer the question: "Where is Mary?" The paragraph can contain removed words marked with [REMOVED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be an...

  284. [300]

    Where is Mary?

    Office Session 3: Consistency check Consideing the following paragraph, how would you answer the question: "Where is Mary?" The paragraph can contain removed words marked with [REMOVED]. Answer either a) "hallway", b) "office", or c) "unknown" if the question can not be answer...

  285. [2017]

    Available: https://data.quora.com/First-Quora-Dataset-Release-Quest ion-Pairs

    [Online]. Available: https://data.quora.com/First-Quora-Dataset-Release-Quest ion-Pairs

  286. [2018]

    Available: http://arxiv.org/abs/1806.03000

    [Online]. Available: http://arxiv.org/abs/1806.03000

  287. [2019]

    Available: https://openreview.net/forum?id=rJ4km2R5t7

    [Online]. Available: https://openreview.net/forum?id=rJ4km2R5t7

  288. [2020]

    Available: http://arxiv.org/abs/2007.09507

    [Online]. Available: http://arxiv.org/abs/2007.09507

  289. [2021]

    Available: https://transformer-circuits.pub/2021/framework/index.html 109

    [Online]. Available: https://transformer-circuits.pub/2021/framework/index.html 109

  290. [2022]

    Available: http://arxiv.org/abs/2207.05221

    [Online]. Available: http://arxiv.org/abs/2207.05221

  291. [2023]

    Available: https://doi.org/10.1016/j.neucom.2023.126518

    [Online]. Available: https://doi.org/10.1016/j.neucom.2023.126518

  292. [2024]

    Available: https://dl.acm.org/doi/10.1145/3605943

    [Online]. Available: https://dl.acm.org/doi/10.1145/3605943

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.