Pith. sign in

REVIEW 3 major objections 6 minor 81 references

Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods

T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper claims that post-hoc explanation methods themselves produce explanations with significant gender disparity, independent of biases in the underlying model.

desk verdict First systematic audit of gender disparity in post-hoc explanation quality for PLMs on text; the faithfulness and complexity results are credible, but the sensitivity metric looks misimplemented, so the robustness claim needs fixing. read the letter →

arxiv 2505.01198 v1 pith:GCREUDZO submitted 2025-05-02 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords genderbiasexplainabilitypost-hocexplanationsfeatureattributionfairnesslanguagemodelsfaithfulnessrobustness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to determine whether the explanations produced by post-hoc feature attribution methods are themselves fair across genders. It finds that, across three text classification tasks and five transformer-based language models, six widely used attribution methods produce explanations whose measured faithfulness, robustness, and complexity differ significantly between male and female inputs. The disparities persist even when the underlying models are trained from scratch on gender-balanced data, which the authors take as evidence that the bias is substantially contributed by the explanation methods rather than inherited from the models. A reader should care because these explanations are used to audit, debug, and justify model decisions; if the explanations are themselves skewed, transparency tools can systematically mislead for one subgroup.

What carries the argument

The argument rests on a disparity-measurement pipeline built from male-female sentence pairs and explanation-quality metrics. Six post-hoc feature attribution methods (Gradient, Integrated Gradients, Gradient×Input, IG×Input, LIME, and SHAP) produce token-importance scores; those scores are then scored by seven metrics covering faithfulness (comprehensiveness, sufficiency, and their soft counterparts), complexity (sparsity, Gini index), and robustness (sensitivity). For each metric, the male and female score distributions are compared with the Mann-Whitney U test, and the magnitude of disparity is quantified with Cohen's $d$. The key move is running the same pipeline on models trained from scratch on a gender-balanced dataset, isolating the contribution of the explanation method from the contribution of the model.

What would settle it

Check the sensitivity implementation in the released code and verify whether the perturbation search maximizes $\|\Phi(f,y)-\Phi(f,x)\|$ or the prediction error; if it maximizes prediction error, rerun the robustness comparisons with the correct objective and see whether the gender disparities in sensitivity persist.

Watch

Extended reading notes

Core claim

The central claim is that explanation quality is not gender-neutral: feature attribution methods produce significantly different faithfulness, robustness, and complexity scores for male and female inputs, with statistically significant disparity ($p \leq 0.05$) in 3,647 of 5,040 experimental combinations and considerable effect size ($|d| \geq 0.2$) in 2,761. The effect is consistent across all six methods, with IG×Input, SHAP, and LIME showing the highest rates. The authors further claim that the disparity is not merely a consequence of biased models or data: retraining BERT and GPT-2 from scratch on the gender-balanced GECO dataset still leaves over 80% of runs with significant gender disparity. The paper concludes that post-hoc explanation methods themselves can be a source of unfairness, independent of model-level bias.

Load-bearing premise

The robustness pillar depends on the sensitivity metric being computed as defined in Eq. 6, but the paper's Appendix A.6 says the PGD attack maximizes prediction error instead of explanation change, so the reported sensitivity values may not measure explanation robustness.

Editorial extensions

If this is right

  • Practitioners who audit a model's fairness by inspecting its explanations should also audit the explanations themselves, because the observed disparities can appear even when the model's predictions are not significantly biased.
  • Deploying post-hoc explanation methods in high-stakes text applications without subgroup checks can produce misleading justifications for one gender, undermining trust and potentially violating transparency obligations.
  • Training or fine-tuning on unbiased data is not a sufficient safeguard for explanation fairness; the attribution methods themselves need to be evaluated and, if needed, adjusted.
  • Because sensitivity disparity was the most frequent on the GECO datasets, explanation robustness should be included in any fairness-oriented evaluation of attribution methods.
  • Larger models such as RoBERTa reduce but do not eliminate the disparities, so model scale alone is not a remedy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the attribution methods are the source of disparity, then model-level debiasing alone will not fix explanation unfairness; explanation-evaluation suites should include per-subgroup score distributions as a standard check.
  • The appendix's discrepancy between the stated sensitivity objective and the PGD implementation used (prediction error instead of explanation change) suggests that the robustness results should be re-measured before being relied on, while the faithfulness and complexity findings do not depend on that step.
  • A testable extension is to vary the perturbation objective and radius in the sensitivity metric to see whether gender disparity in robustness changes once the explanation change is truly maximized, and to compare hard versus soft token removal in faithfulness.
  • The same disparity-measurement pipeline could be applied to non-binary gender markers or to race- and age-related cues in text, though the synthetic data generation would need to be adapted to avoid the corpus-design pitfalls the paper acknowledges.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper investigates whether post-hoc feature attribution methods (Gradient, Gradient×Input, Integrated Gradients, IG×Input, LIME, SHAP) show gender disparities in explanation faithfulness, robustness, and complexity when explaining language model predictions. Across GECO, Stereotypes, and COMPAS text datasets and five transformer models, the authors compute seven evaluation metrics, test significance with the Mann-Whitney U test, and quantify effect sizes with Cohen's d. They report that 72.4% of 5,040 dataset-model-explainer-metric-seed combinations show statistically significant disparity, that disparities persist when models are trained from scratch on GECO, and they conclude that explanation methods themselves contribute to gender bias. The paper also discusses implications for practitioners and regulators, connecting explanation fairness to frameworks such as the EU AI Act.

Significance. The paper addresses an important and underexplored question: whether post-hoc explanation methods themselves produce systematically different explanation quality across gender subgroups. The empirical scope is substantial: six explainers, five transformer models, three datasets, seven metrics, and five seeds, plus from-scratch training experiments. The authors use established metrics and statistical tests from prior literature, report effect sizes in addition to p-values, and make a clear practical case for auditing explanation fairness. If the sensitivity implementation is corrected and the causal attribution is appropriately tempered, the study would be the first systematic NLP demonstration of explanation-level gender disparity and would provide a valuable benchmark for future XAI fairness work. The current manuscript, however, does not yet support the robustness leg of the central claim and overstates the causal role of explanation methods.

major comments (3)
  1. [Appendix A.6, Eq. (6)] The sensitivity metric is defined as the worst-case relative change in the explanation under an input perturbation, max over y with ||x-y||<=r of ||Phi(f,y)-Phi(f,x)|| / ||Phi(f,x)||, but the implementation is described as a PGD attack that 'perturbs the input in the direction of the gradient maximizing the prediction error.' This attack maximizes prediction loss, not the explanation-distance term in Eq. (6). Unless the two objectives are shown to coincide, which the manuscript does not demonstrate, the reported Sensitivity values measure a different quantity: the explanation change caused by a prediction-error attack. Because sensitivity is the only robustness metric, the robustness pillar of the central claim ('faithfulness, robustness, and complexity') is unsupported as reported. Please re-run the attack with the explanation-distance objective (or provide evidence that the two objectives yield equivalent worst-case perturbations) and regenerate Tables 3-5 and the affected appendix tables.
  2. [Section 4.4 and Section 5] The disparity analysis uses the Mann-Whitney U test at p<=0.05 for each of the 5,040 dataset-model-explainer-metric-seed combinations without any correction for multiple comparisons. With the COMPAS dataset's large male subset (n=4,997), even negligible effect sizes can become statistically significant, so the headline figure of 72.4% significant disparities may be inflated. Please report multiple-testing-corrected p-values (for example, Benjamini-Hochberg within each metric/model/dataset family or across all tests) and show how many disparities remain significant after correction; alternatively, justify why uncorrected tests are appropriate for this descriptive audit.
  3. [Section 5.4 and Section 4.1] The 'training from scratch on an unbiased dataset' experiments use GECO, where the classification label is the gender expressed in the sentence. Because the model is trained to predict gender, it must rely on gender-marking tokens, and the male/female inputs differ exactly in those tokens. Persistent disparity in explanation metrics under this condition is therefore not sufficient to conclude that 'explanation methods themselves can contribute to these disparities' (Section 5.4), because token-level differences between male and female inputs are a confound. The experiment rules out biased pre-training data, but it does not rule out task-induced or token-induced disparity. Please add a control condition in which gender is task-irrelevant, match male/female inputs on token-level statistics, or rephrase the conclusion to the weaker claim that disparities persist in the absence of biased training labels.
minor comments (6)
  1. [Section 4.2 and References] The citation for FairBERTa is [31], but reference [31] is the TinyBERT distillation paper; TinyBERT is cited as [55], which is the FairBERTa/perturbation-augmentation paper. Please swap the citations.
  2. [Tables 3 and 4 captions] The captions state that cell colors indicate 'which gender has better evaluation scores,' but for sensitivity, sparsity, and sufficiency lower values are preferred; Section 5 correctly says that colors indicate higher scores. Please align the captions with the text.
  3. [Appendix E, Table 7] Values such as '0.9910.001' are missing delimiters and plus-minus signs, making the TPR/TNR/APD entries unreadable; please reformat with proper separators.
  4. [Footnote 2 and Appendix B] The paper says code and datasets are released on GitHub and made public, but it provides no repository URL or version. Please add a formal availability statement with a link.
  5. [Appendix A.6] The normalization by ||Phi(f,x)|| in Eq. (6) can be unstable when an explanation vector is near zero; this may explain the very large and highly variable Cohen's d values in Tables 8-11 (for example, FairBERTa soft comprehensiveness d = 17.86 ± 31.01 on GECO-ALL). Please consider a regularized normalization or report unnormalized changes as a robustness check.
  6. [Throughout] There are several typos, including 'thecomprehensiveness' in Section 5.2.1, 'run on Stereotypes' in Section 5.2.3, 'scoresobtained' in Figure 2's caption, and 'Table in 3' in Section 5.4. A careful proofreading pass is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the evaluation metrics, disparity tests, and controls are externally defined and empirically applied; the paper's self-citations are not load-bearing.

full rationale

This paper is an empirical audit rather than a derivation, so the circularity tests rarely apply. The metric definitions in Appendix A (comprehensiveness, sufficiency, soft variants, sparsity, Gini index, sensitivity) are taken from prior literature and from external implementations such as ferret; none is defined in terms of the male/female outcome that is later tested. Disparity is measured after the fact with the Mann-Whitney U test and Cohen's d, and no parameter is fitted to the disparity result: the pipeline in Algorithm 1 computes scores per input and then compares the two score lists, so there is no fitted-input-called-prediction step. The from-scratch GECO experiment is a designed control, described as initializing models randomly and training them on GECO-ALL or GECO-SUBJ; the claim that explanation methods themselves contribute to disparities is an inference drawn from observing disparity under that control, not a conclusion built into the metric definitions. The self-citations that appear (Inseq [59], the text summarization survey [20], and attention mechanisms [38]) are contextual examples or future-work pointers, and they are not used to justify the central empirical claim. The skeptical concern about Appendix A.6, where the PGD attack is described as perturbing the input in the direction of the gradient maximizing prediction error rather than maximizing the Eq. 6 explanation-change objective, is a measurement-validity issue about whether the reported sensitivity values reflect the defined quantity; it does not make the reported quantity equivalent to its input by construction. No circular step of any enumerated kind is present.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper is empirical and introduces no new entities. Free parameters are primarily evaluation thresholds and attack hyperparameters, several of which are unreported. The key assumptions concern metric validity and the meaningfulness of the from-scratch training setup.

free parameters (3)
  • sparsity threshold tau = 0.1
    Threshold for counting features as important in sparsity metric (Appendix A.4); chosen by hand and not varied or justified.
  • sensitivity perturbation radius r = not reported
    Radius of the ball in Eq. 6 for worst-case sensitivity; value is never specified in the paper, yet it directly controls the sensitivity scores.
  • PGD attack hyperparameters = not reported
    Number of steps, step size, and norm constraint for the PGD attack in Appendix A.6 are not given, making the sensitivity results implementation-dependent.
assumptions (4)
  • domain assumption Mann-Whitney U is appropriate for comparing male and female explanation score distributions
    Used in Section 4.4 to determine significant disparity; assumes independent samples and ordinal scores, which is reasonable for unequal group sizes.
  • domain assumption The evaluation metrics measure the intended explanation properties (faithfulness, complexity, robustness)
    The metrics are inherited from prior work; the sensitivity metric in particular is implemented in a way that may not match the definition.
  • domain assumption GECO is an unbiased dataset and training on it from scratch isolates the effect of pre-training data
    The paper defines unbiased narrowly via balanced masked pairs, but gender is still the prediction target, so gender necessarily influences the task.
  • domain assumption Randomly initialized BERT/GPT-2 can be trained on 3,220 sentences to yield meaningful classifiers for explanation evaluation
    No accuracy is reported for from-scratch models; if they memorize rather than generalize, explanation disparity may be an artifact.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods." pith.science (2026). https://pith.science/paper/GCREUDZO

@misc{pith2026250501198,
  author       = {Pith},
  title        = {Pith review of: Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GCREUDZO}},
  note         = {Machine review of arXiv:2505.01198}
}
read the original abstract

While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper, we address this gap by showing that, across three tasks and five language models, widely used post-hoc feature attribution methods exhibit significant gender disparity with respect to their faithfulness, robustness, and complexity. These disparities persist even when the models are pre-trained or fine-tuned on particularly unbiased datasets, indicating that the disparities we observe are not merely consequences of biased training data. Our results highlight the importance of addressing disparities in explanations when developing and applying explainability methods, as these can lead to biased outcomes against certain subgroups, with particularly critical implications in high-stakes contexts. Furthermore, our findings underscore the importance of incorporating the fairness of explanations, alongside overall model fairness and explainability, as a requirement in regulatory frameworks.

Figures

Figures reproduced from arXiv: 2505.01198 by the authors.

Figure 1
Figure 1. Overview of our experimental pipeline, exemplified with the GECO dataset [69]. We begin by obtaining predictions for male/female sentence pairs. We then use feature attribution methods to explain the predictions and evaluate the explanations using various metrics. We finally analyze the distributions of evaluation scores per each metric for male and female sentences and observe if the evaluations differ significantl… view at source ↗
Figure 2
Figure 2. Box-plots of evaluation scores obtained over 5 runs for each using TinyBERT on GECO, Stereotypes, and COMPAS, including the runs not resulting in statistically significant disparity. completely, prone to gender disparities in their explanations. However, this is not visible on the GECO datasets, where all runs result in statistically significant disparities in sensitivity [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Box-plots of the soft comprehensiveness and sufficiency metrics obtained over 5 runs for each using TinyBERT on GECO-ALL, Stereotypes, and COMPAS, including the runs not resulting in statistically significant disparity [PITH_FULL_IMAGE:figures/full_fig_p028_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Box-plots of evaluation scores obtained over 5 runs for each using BERT on our datasets, including the runs not resulting in statistically significant disparity [PITH_FULL_IMAGE:figures/full_fig_p029_4.png]
Figure 5
Figure 5. Figure 5: Box-plots of evaluation scores obtained over 5 runs for each using FairBERTa on our datasets, including the runs not resulting in statistically significant disparity [PITH_FULL_IMAGE:figures/full_fig_p030_5.png]
Figure 6
Figure 6. Figure 6: Box-plots of evaluation scores obtained over 5 runs for each using GPT-2 on our datasets, including the runs not resulting in statistically significant disparity [PITH_FULL_IMAGE:figures/full_fig_p031_6.png]
Figure 7
Figure 7. Figure 7: Box-plots of evaluation scores obtained over 5 runs for each using RoBERTa on our datasets, including the runs not resulting in statistically significant disparity [PITH_FULL_IMAGE:figures/full_fig_p032_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 23 canonical work pages

  1. [1]

    Definition of STAKEHOLDERS

    2025. Definition of STAKEHOLDERS. https://www.merriam-webster.com/dictionary/stakeholders

  2. [2]

    Jaakkola

    David Alvarez-Melis and Tommi S. Jaakkola. 2018. On the Robustness of Interpretability Methods. arXiv:1806.08049 [cs.LG] https://arxiv. org/abs/1806.08049

  3. [3]

    Julia Angwin, Jeff Larson, Lauren Kirchner, and Surya Mattu. 2016. Machine bias. https://www.propublica.org/article/machine- bias-risk-assessments-in-criminal-sentencing

  4. [4]

    AI Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card 1 (2024)

  5. [5]

    Leila Arras, Ahmed Osman, and Wojciech Samek. 2022. CLEVR-XAI: A benchmark dataset for the ground truth evaluation of neural network explanations. Information Fusion 81 (2022), 14–40. doi: 10.1016/j.inffus.2021.11.008

  6. [6]

    Giuseppe Attanasio, Eliana Pastor, Chiara Di Bonaventura, and Debora Nozza. 2023. ferret: a Framework for Benchmarking Explainers on Transformers. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations . Association for Computational Linguistics

  7. [7]

    Aparna Balagopalan, Haoran Zhang, Kimia Hamidieh, Thomas Hartvigsen, Frank Rudzicz, and Marzyeh Ghassemi. 2022. The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing Mach...

  8. [8]

    Esma Balkir, Svetlana Kiritchenko, Isar Nejadgholi, and Kathleen Fraser. 2022. Challenges in Applying Explainability Methods to Improve the Fairness of NLP Models. In Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022) , Apurv Verma, Yada Pruksachatkun, Kai-Wei Chang, Aram Galstyan, Jwala Dhamala, and Yang Trista Cao...

Show all 81 references
  1. [9]

    Milan Bhan, Jean-Noel Vittaut, Nicolas Chesneau, and Marie-Jeanne Lesot. 2024. Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations. arXiv preprint arXiv:2402.12038 (2024)

  2. [10]

    Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, ...

  3. [11]

    Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 1...

  4. [12]

    Stephanie Brandl, Emanuele Bugliarello, and Ilias Chalkidis. 2024. On the Interplay between Fairness and Explainability. In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024) , Anaelia Ovalle, Kai-Wei Chang, Yang Trista Cao, Ninareh Mehr...

  5. [13]

    George Chrysostomou and Nikolaos Aletras. 2022. An Empirical Study on Explanations in Out-of-Domain Settings. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavi...

  6. [14]

    Council of European Union. 2024. Council regulation (EU) no 2024/1689. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

  7. [15]

    Bach, and Himabindu Lakkaraju

    Jessica Dai, Sohini Upadhyay, Ulrich Aivodji, Stephen H. Bach, and Himabindu Lakkaraju. 2022. Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (Oxford, Un...

  8. [16]

    Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020. A Survey of the State of Explainable AI for Natural Language Processing. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Lingu...

  9. [17]

    Björn Deiseroth, Mayukh Deb, Samuel Weinbach, Manuel Brack, Patrick Schramowski, and Kristian Kersting. 2023. AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation. arXiv:2301.08110 [cs.LG]https://arxiv.org/abs/2301.08110

  10. [18]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  11. [19]

    Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020. ERASER: A Benchmark to Evaluate Rationalized NLP Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan J...

  12. [20]

    Mahdi Dhaini, Ege Erdogan, Smarth Bakshi, and Gjergji Kasneci. 2024. Explainability Meets Text Summarization: A Survey. In Proceedings of the 17th International Natural Language Generation Conference , Saad Mahamood, Nguyen Le Minh, and Daphne Ippolito (Eds.). Association for ...

  13. [21]

    Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024. Large language models on tabular data–a survey. arXiv e-prints (2024), arXiv–2402

  14. [22]

    Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed

  15. [23]

    Amirata Ghorbani, Abubakar Abid, and James Zou. 2019. Interpretation of Neural Networks Is Fragile. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 3681–3688. doi:10.1609/aaai.v33i01.33013681

  16. [24]

    Jennifer Hsia, Danish Pruthi, Aarti Singh, and Zachary Lipton. 2024. Goodhart‘s Law Applies to NLP‘s Explanation Benchmarks. In Findings of the Association for Computational Linguistics: EACL 2024 , Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguis...

  17. [25]

    Marco Huber, Meiling Fang, Fadi Boutros, and Naser Damer. 2023. Are Explainability Tools Gender Biased? A Case Study on Face Presentation Attack Detection. In 2023 31st European Signal Processing Conference (EUSIPCO) . 945–949. doi:10.23919/EUSIPCO58844.2023.10289865

  18. [26]

    Alon Jacovi. 2023. Trends in Explainable AI (XAI) Literature. arXiv:2301.05433 [cs.AI] https://arxiv.org/abs/2301.05433

  19. [27]

    Alon Jacovi and Yoav Goldberg. 2020. Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel...

  20. [28]

    Sarthak Jain and Byron C. Wallace. 2019. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy D...

  21. [29]

    Sophie Jentzsch and Cigdem Turan. 2022. Gender Bias in BERT-Measuring and Analysing Biases through Sentiment Rating in a Realistic Downstream Classification Task. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP) . 184–199

  22. [30]

    Sérgio Jesus, Catarina Belém, Vladimir Balayan, João Bento, Pedro Saleiro, Pedro Bizarro, and João Gama. 2021. How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and...

  23. [31]

    Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for Natural Language Understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (E...

  24. [32]

    Shreya Johri, Jaehwan Jeong, Benjamin A Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Leandra A Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M Van Allen, David Kim, et al. 2025. An evaluation framework for clinical use of large language models in patient interaction tasks....

  25. [33]

    Neema Kotonya and Francesca Toni. 2020. Explainable Automated Fact-Checking: A Survey. In Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Committee on Computational Linguistics, Bar...

  26. [34]

    Satyapriya Krishna, Tessa Han, Alex Gu, Steven Wu, Shahin Jabbari, and Himabindu Lakkaraju. 2024. The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective. Transactions on Machine Learning Research (2024). https://openreview.net/forum?id= jESY2WTZCe

  27. [35]

    Satyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju. 2024. Post hoc explanations of language models can improve language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New O...

  28. [36]

    How do I fool you?

    Himabindu Lakkaraju and Osbert Bastani. 2020. "How do I fool you?": Manipulating User Trust via Misleading Black Box Explanations. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society (New York, NY, USA) (AIES ’20). Association for Computing Machinery, New York,...

  29. [37]

    Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum. 2021. What do we want from Explainable Artificial Intelligence (XAI)? – A stakeholder perspective on XAI and a conceptual model guiding interdisciplinary XAI r...

  30. [38]

    Tobias Leemann, Alina Fastowski, Felix Pfeiffer, and Gjergji Kasneci. 2025. Attention Mechanisms Don’t Learn Additive Models: Rethinking Feature Importance for Transformers. arXiv:2405.13536 [cs.LG] https://arxiv.org/abs/2405.13536 Gender Bias in Explainability: Investigating ...

  31. [39]

    Xuhong Li, Mengnan Du, Jiamin Chen, Yekun Chai, Himabindu Lakkaraju, and Haoyi Xiong. 2023. M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models. InAdvances in Neural Information Processing Systems, A. Oh,...

  32. [40]

    Hui Liu, Qingyu Yin, and William Yang Wang. 2019. Towards Explainable NLP: A Generative Explanation Framework for Text Classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Ed...

  33. [41]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907. 11692

  34. [42]

    Luca Longo, Mario Brcic, Federico Cabitza, Jaesik Choi, Roberto Confalonieri, Javier Del Ser, Riccardo Guidotti, Yoichi Hayashi, Francisco Herrera, Andreas Holzinger, Richard Jiang, Hassan Khosravi, Freddy Lecue, Gianclaudio Malgieri, Andrés Páez, Wojciech Samek, Johannes Schn...

  35. [43]

    I Loshchilov. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)

  36. [44]

    Scott M Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 (2017)

  37. [45]

    Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2024. Towards Faithful Model Explanation in NLP: A Survey. Computational Linguistics 50, 2 (June 2024), 657–723. doi:10.1162/coli_a_00511

  38. [46]

    Aleksander Madry. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017)

  39. [47]

    Andreas Madsen, Himabindu Lakkaraju, Siva Reddy, and Sarath Chandar. 2024. Interpretability Needs a New Paradigm. arXiv:2405.05386 [cs.LG] https://arxiv.org/abs/2405.05386

  40. [48]

    Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022. Post-hoc Interpretability for Neural NLP: A Survey. ACM Comput. Surv. 55, 8, Article 155 (Dec. 2022), 42 pages. doi: 10.1145/3546577

  41. [49]

    Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. Proceedings of the AAAI Conference on Artificial Intelligence 35, 17 (May 2021), 14867–14875. doi:10.1...

  42. [50]

    Vishwali Mhasawade, Salman Rahman, Zoé Haskell-Craig, and Rumi Chunara. 2024. Understanding Disparities in Post Hoc Machine Learning Explanation. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT ’24). Assoc...

  43. [51]

    Katelyn Morrison, Philipp Spitzer, Violet Turri, Michelle Feng, Niklas Kühl, and Adam Perer. 2024. The Impact of Imperfect XAI on Human-AI Decision-Making. Proc. ACM Hum.-Comput. Interact. 8, CSCW1, Article 183 (April 2024), 39 pages. doi:10.1145/3641022

  44. [52]

    Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh. 2022. SHAP-Based Explanation Methods: A Review for NLP Interpretability. In Proceedings of the 29th International Conference on Computational Linguistics , Nicoletta Calzolari, Chu-Ren Huang, Ha...

  45. [53]

    Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, T...

  46. [54]

    Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Maurice van Keulen, and Christin Seifert. 2023. From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI. ACM Comput....

  47. [55]

    Rebecca Qian, Candace Ross, Jude Fernandes, Eric Michael Smith, Douwe Kiela, and Adina Williams. 2022. Perturbation Augmentation for Fairer NLP. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 9496–9521

  48. [56]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  49. [57]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144

  50. [58]

    Marko Robnik-Šikonja and Marko Bohanec. 2018. Perturbation-Based Explanations of Prediction Models . Springer International Publishing, Cham, 159–175. doi: 10.1007/978-3-319-90403-0_9

  51. [59]

    Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023. Inseq: An Interpretability Toolkit for Sequence Generation Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Danushka ...

  52. [60]

    Shlomo S Sawilowsky. 2009. New effect size rules of thumb. Journal of modern applied statistical methods 8 (2009), 597–599. 18 Dhaini et al

  53. [61]

    Jakob Schoeffer, Maria De-Arteaga, and Niklas Kühl. 2024. Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Mach...

  54. [62]

    Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013)

  55. [63]

    Samuel Sithakoul, Sara Meftah, and Clément Feutry. 2024. BEExAI: Benchmark to Evaluate Explainable AI. In Explainable Artificial Intelligence, Luca Longo, Sebastian Lapuschkin, and Christin Seifert (Eds.). Springer Nature Switzerland, Cham, 445–468

  56. [64]

    Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. In International conference on machine learning . PMLR, 3319–3328

  57. [65]

    Santosh T.y.s.s., Nina Baumgartner, Matthias Stürmer, Matthias Grabmair, and Joel Niklaus. 2024. Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset. InProceedings of the 2024 Joint International Conference on Computational...

  58. [66]

    Josef Valvoda and Ryan Cotterell. 2024. Towards Explainability in Legal Outcome Prediction Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , Kevin ...

  59. [67]

    Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating Gender Bias in Language Models Using Causal Mediation Analysis. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R....

  60. [68]

    Eric Wallace, Matt Gardner, and Sameer Singh. 2020. Interpreting Predictions of NLP Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , Aline Villavicencio and Benjamin Van Durme (Eds.). Association for Comput...

  61. [69]

    Rick Wilming, Artur Dox, Hjalmar Schulz, Marta Oliveira, Benedict Clark, and Stefan Haufe. 2024. GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations. arXiv preprint arXiv:2406.11547 (2024)

  62. [70]

    Alice Xiang and Inioluwa Deborah Raji. 2019. On the Legal Compatibility of Fairness Definitions. arXiv:1912.00761 [cs.CY] https://arxiv. org/abs/1912.00761

  63. [71]

    Wenzhuo Yang, Hung Le, Tanmay Laud, Silvio Savarese, and Steven C. H. Hoi. 2022. OmniXAI: A Library for Explainable AI. arXiv:2206.01612 [cs.LG] https://arxiv.org/abs/2206.01612

  64. [72]

    Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar. 2019. On the (In)fidelity and Sensitivity of Explanations. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d 'Alché-Buc, E. Fox, and R. Ga...

  65. [73]

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (Feb. 2024), 38 pages. doi: 10.1145/3639372

  66. [74]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  67. [75]

    Zhixue Zhao and Nikolaos Aletras. 2023. Incorporating Attribution Importance for Improving Faithfulness Metrics. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Oka...

  68. [76]

    Zhixue Zhao, George Chrysostomou, Kalina Bontcheva, and Nikolaos Aletras. 2022. On the Impact of Temporal Concept Drift on Model Explanations. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Ass...

  69. [77]

    Gandomi, Fang Chen, and Andreas Holzinger

    Jianlong Zhou, Amir H. Gandomi, Fang Chen, and Andreas Holzinger. 2021. Evaluating the Quality of Machine Learning Explanations: A Survey on Methods and Metrics. Electronics 10, 5 (2021). doi: 10.3390/electronics10050593

  70. [78]

    Julia El Zini and Mariette Awad. 2022. On the Explainability of Natural Language Processing Deep Models. ACM Comput. Surv. 55, 5, Article 103 (Dec. 2022), 31 pages. doi: 10.1145/3529755 Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods 19 A...

  71. [80]

    I was surprised to see a woman doctor articulate herself so well

    and COMPAS datasets [3], as well as the synthetic Stereotypes dataset we create and make public. We use the publicly available models from Huggingface (see Table 6) running on a single NVIDIA V100 GPU. Including fine-tuning, and generating and evaluating explanations, one mode...

  72. [81]

    significant

    with initial learning rate 0.001 and a linear learning rate schedule with 500 warm-up steps. Since our tasks are binary classification tasks, we use the binary cross-entropy loss. We observed that after one epoch of fine-tuning, the models perform hardly better than random gue...

  73. [2024]

    Computational Linguistics (2024), 1–79

    Bias and fairness in large language models: A survey. Computational Linguistics (2024), 1–79

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.