REVIEW 3 major objections 5 minor 79 references
Graph-Guided Textual Explanation Generation Framework
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read G-Tex claims that injecting a model's own attention-derived highlight explanations into its encoder as a graph reduces counterfactual unfaithfulness in generated natural language explanations by up to 12.18 percentage points.
desk verdict G-Tex is a concrete recipe with a real validity threat: the one-sided counterfactual test may reward input copying, so the headline faithfulness gain is not yet established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the highlight-explanation graph. Each input token is a node; edges are drawn according to the three highlight types—top-$k\%$ single tokens, token pairs, and span pairs—with extra edges connecting subword tokens after tokenization. This graph is encoded by a GNN layer, implemented as GraphSAGE, GAT, or GCN, which is inserted after the $3/4$-th encoder layer of T5 or BART. The GNN updates token representations by aggregating information from highlighted neighbours, and the decoder then produces the label and the natural language explanation from these augmented encoder states. The graph edges are initialized equally and are updated during fine-tuning, so the highlights act as a soft structural prior on the information flow that enters the explanation rather than as a hard constraint.
What would settle it
Re-run G-Tex with the same architecture and data, but replace the attention-derived highlight edges with equally sized edges chosen at random or from a perturbation-based attribution method such as leave-one-out. If random or perturbation-based highlights produce the same or larger faithfulness improvements, the claim that attention-based highlights are the faithful cue carrying the effect would be falsified.
Extended reading notes
Core claim
The discovery that G-Tex argues for is that a model's own high-faithfulness highlight explanations can be turned into a graph and used as a structural guide for generating its natural language explanations. The framework first trains a base encoder-decoder model for label prediction, extracts three types of highlights from its decoder attention—single important tokens, interacting token pairs across the two input parts, and interacting spans—and then builds a graph in which input tokens are nodes and edges encode the selected highlights. A graph neural network layer placed after the three-quarter point of the encoder aggregates these node representations, and the decoder generates the label and the explanation from the augmented encoder states. Across T5 and BART on e-SNLI, ComVE, and ECQA, the authors report that this lowers counterfactual unfaithfulness below both plain fine-tuning and prompt-injection baselines, by up to 12.18 percentage points, while preserving or improving label accuracy. The authors also find that interactive highlights help most for two-part inputs where relations between parts matter, while single-token highlights help most when the first input part is a fixed instruction.
Load-bearing premise
The framework assumes that the attention weights of the base model, averaged over tokens, reliably identify the input fragments the model truly uses when it makes its prediction, and that injecting those fragments as graph edges is what makes the generated explanation track the model's real reasoning.
Editorial extensions
If this is right
- Faithfulness of natural language explanations can be improved by guiding generation with the model's own high-faithfulness cues, without sacrificing answer accuracy: G-Tex reports label accuracy comparable to or better than the baselines.
- The best choice of highlight type follows the input structure: token and span interactive explanations give the largest faithfulness gains on tasks with meaningful interaction between two input parts, while token explanations work best when one input part is a fixed instruction.
- G-Tex explanations are more similar to human-written explanations in both lexical and semantic measures, and human judges rate them as less redundant and higher in overall quality.
- The method adds very few parameters (about 0.28% over T5 and 0.24% over BART) and nearly identical training time, so the faithfulness gain does not come from a large capacity increase.
Reading between the lines
- The counterfactual faithfulness test rewards explanations that mention the inserted adjective after a label change, so part of G-Tex's improvement may be a sharper ability to echo injected tokens rather than a deeper alignment with reasoning; a deletion-based or consistency-based faithfulness test would separate these.
- Because the graph controls which input tokens the decoder can see, the same wiring could steer other generation properties—factual grounding, style, or coverage—by substituting different edge sets for different criteria.
- The paper's dataset-dependent finding suggests a practical selection rule: use interactive highlights for multi-part inputs and token highlights for fixed-template inputs, which could be tested as an automatic per-instance or per-dataset choice.
- Extending the graph injection to decoder-only models would require a different way of merging graph embeddings into next-token prediction, since the decoder attends only to preceding tokens; the authors leave that adaptation open.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes G-Tex, a framework for improving the faithfulness of natural language explanations (NLEs) generated by encoder-decoder models. G-Tex first extracts post-hoc highlight explanations (highlight tokens, token interactions, span interactions) from a fine-tuned base model using attention-based attribution, then converts these highlights into a graph that is injected into the model via a GNN layer placed after the 3/4-th encoder layer. The model is fine-tuned to jointly produce the task label and the NLE, with the graph providing explicit cues about important input fragments. Experiments on e-SNLI, ComVE, and ECQA with T5-large and BART-large compare G-Tex against a fine-tuning baseline and a prompt-based highlight baseline. The paper reports reductions in counterfactual unfaithfulness (up to 12.18 percentage points in Total Unfaith), higher lexical/semantic similarity to human explanations, and favorable human evaluation scores on non-redundancy and overall quality.
Significance. If the reported effects are genuine, G-Tex would be a useful contribution: it provides an explicit architectural mechanism for steering NLE generation toward input features that the model demonstrably relies on, going beyond prompt-based or loss-based extrinsic signals. The paper is technically solid in scope: it includes three datasets, two base models, three GNN variants, a prompt baseline, five random seeds with p-values, and human evaluations. The authors also acknowledge limitations regarding decoder-only models, GNN internal mechanisms, and attribution-method choices. The main concern is not the novelty but whether the central faithfulness claim is supported by the reported numbers and whether the evaluation metric captures true alignment with model reasoning rather than a copying artifact.
major comments (3)
- [§5.1 and Table 5] The statement that "our method outperforms all baselines in counterfactual unfaithfulness and total faithfulness" is contradicted by the paper's own tables. In Table 5, BART-based Tex-GCN with span interactions on ComVE gives Total Unfaith 96.44% versus 72.82% for Fine-tuningbase, and T5-based Tex-GCN with token interactions gives 77.03% versus 73.73% for Fine-tuningbase. Table 1 similarly shows T5-based Tex-SAGE with token interactions on ComVE at 76.94% Total Unfaith versus 68.96% for Fine-tuningbase. These are not small or negligible differences, and they contradict the blanket claim in the abstract and in §5.1. The paper should either temper the claim to the specific configurations that actually improve (e.g., highlight tokens on ComVE, span interactions on e-SNLI) or report a summary statistic over all configurations that is consistent with the data.
- [§H and §3.5] The counterfactual faithfulness test is one-sided: when an inserted adjective does not change the predicted label, the generated NLE is never checked for whether it incorrectly mentions the inserted word, as the paper itself states in Appendix H ("the unchanged label provides no relevant information about the faithfulness of the NLE"). Because G-Tex injects the top-30% highlighted tokens into the encoder during fine-tuning (§3.3, §F), the model may learn to copy input spans into the NLE. Such copying would trivially pass the label-changing half of the test (the inserted word appears because it is copied) while remaining unpenalized on the label-unchanged half. The reported 12.18% reduction in Total Unfaith could therefore reflect increased lexical copying rather than better alignment with the model's reasoning. The authors should report the rate at which G-Tex and baselines include inserted adjectives when the label does not change, or use a two-sided faithfulness test, to rule out this artifact.
- [§3.2 and §3.3] The importance scores that define all three highlight-explanation types are derived exclusively from attention weights: a single head is selected and self-attention scores are averaged to produce token and interaction importances (§3.2). This choice is known to be contested (Jain and Wallace 2019; Serrano and Smith 2019), and the paper's own Limitations section acknowledges that other attribution methods (Shapley, Integrated Gradients, saliency maps) are not explored. Since the graph construction and the GNN injection both depend on these scores, the claim that G-Tex aligns NLEs with the model's "underlying reasoning" would be considerably stronger if at least one alternative attribution method were used to show that the faithfulness gains are not an artifact of the specific attention-based signal. As it stands, the central mechanism rests on a single, heavily debated attribution choice.
minor comments (5)
- [§5.1] The sentence "While G-Tex with T5 slightly underperforms the prompt baseline on ComVE with token interactive explanations" is inaccurate: the T5 Tex-SAGE token-interaction configuration in Table 1 (Total Unfaith 76.94) also underperforms Fine-tuningbase (68.96), and the gap is not slight.
- [Appendix I] The text contains a typo: "against comment sense" should be "against common sense" in the paragraph discussing ComVE dataset characteristics.
- [Appendix J] The p-values are computed only for Tex-SAGE versus Fine-tuningbase. Since the paper's claim is "outperforms all baselines," significance tests against the Prompt baseline are also needed, especially for configurations where the numerical gap is small.
- [Appendix M.3] The pairwise annotator agreement is low in several dimensions (e.g., mean agreement of 0.18 for Overall Quality on ECQA). The paper should state what level of agreement is considered acceptable and how disagreements were resolved.
- [§4.2] Checkpoint selection is performed using BLEU on the validation set. Since the paper's main outcome is faithfulness, the authors should clarify whether selecting checkpoints by BLEU could bias faithfulness results, or report faithfulness on the validation set as an alternative selection criterion.
Circularity Check
The empirical faithfulness comparison is independently computed, but the core premise that attention-based highlights are faithful rests on a load-bearing self-citation to the authors' own prior work.
-
self citation load bearing
[Abstract and §3.2 (Post Hoc Highlight Explanation and Predicted Label)]
"Previous work shows that highlight explanations extracted by attention-based methods show higher faithfulness than other explainability techniques (Sun et al., 2024). Building on this, we use attention weights as the basis for deriving importance scores for all types of highlight explanations."
The premise that attention-derived highlight explanations are faithful cues to the model's internal reasoning is not derived or independently validated in this paper; it is imported from Sun et al. (2024) and Atanasova et al. (2020a), whose author lists include the present paper's co-authors (Sun, Atanasova, Augenstein). This premise is load-bearing because the entire G-Tex pipeline begins from these highlights ('highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction'), and the abstract frames the contribution as building on this foundation. The conceptual grounding of the faithfulness-improvement claim therefore reduces to a self-citation chain rather than a first-principles derivation.
full rationale
G-Tex's headline result is an empirical comparison computed in this paper on three public datasets (e-SNLI, ComVE, ECQA) with two base models (T5-large, BART-large), and the baselines are re-run in the same experimental setup. The faithfulness evaluation uses the published counterfactual protocol of Atanasova et al. (2023); although that test is co-authored by one of the present paper's authors, it is a fixed external protocol rather than a parameter fitted by G-Tex. The method itself is a concrete architectural intervention (attention-based highlights, graph construction, GNN injection into the encoder) and is not a renamed version of the metric. The main circularity-adjacent element is the load-bearing self-citation identified above: the foundational claim that attention-based highlights are faithful comes from the authors' own prior work, not from a derivation in this paper. Because the final faithfulness scores are independently obtained and the counterfactual metric is not fitted to the method, the central claim retains independent empirical content. A separate concern that the counterfactual test is one-sided and that highlight injection may encourage input copying is a validity threat, not a definitional circularity: the inserted adjectives are not explicitly injected as graph nodes by the training procedure, the Prompt baseline also receives the same top-k highlight tokens, and no equation in the paper equates the graph structure with the tokens checked by the metric.
Assumptions & free parameters
free parameters (1)
- k (top-k% highlights) =
30
assumptions (3)
- domain assumption Attention weights identify input tokens that are most important for a model's prediction
- domain assumption The counterfactual faithfulness test measures whether an NLE reflects the model's reasoning
- ad hoc to paper Injecting a GNN layer at the 3/4 encoder depth improves information flow
Cite this review
Pith. "Pith review of Graph-Guided Textual Explanation Generation Framework." pith.science (2026). https://pith.science/paper/TIIRSUTS
@misc{pith2026241212318,
author = {Pith},
title = {Pith review of: Graph-Guided Textual Explanation Generation Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/TIIRSUTS}},
note = {Machine review of arXiv:2412.12318}
}
read the original abstract
Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions. However, recent work has questioned their faithfulness, as they may not accurately reflect the model's internal reasoning process regarding its predicted answer. In contrast, highlight explanations--input fragments critical for the model's predicted answers--exhibit measurable faithfulness. Building on this foundation, we propose G-Tex, a Graph-Guided Textual Explanation Generation framework designed to enhance the faithfulness of NLEs. Specifically, highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction. They are subsequently encoded through a graph neural network layer to guide the NLE generation, which aligns the generated explanations with the model's underlying reasoning toward the predicted answer. Experiments on T5 and BART using three reasoning datasets show that G-Tex improves NLE faithfulness by up to 12.18% compared to baseline methods. Additionally, G-Tex generates NLEs with greater semantic and lexical similarity to human-written ones. Human evaluations show that G-Tex can decrease redundant content and enhance the overall quality of NLEs. Our work presents a novel method for explicitly guiding NLE generation to enhance faithfulness, serving as a foundation for addressing broader criteria in NLE and generated text.
Figures
Reference graph
Works this paper leans on
-
[1]
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. https://doi.org/10.18653/v1/2021.acl-long.238 E xplanations for C ommonsense QA : N ew D ataset and M odels . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen...
-
[2]
David Alvarez Melis and Tommi Jaakkola. 2018. https://proceedings.neurips.cc/paper_files/paper/2018/file/3e9f0fc9b2f89e043bc6233994dfcf76-Paper.pdf Towards robust interpretability with self-explaining neural networks . In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc
work page 2018
-
[3]
Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein. 2023. https://doi.org/10.18653/v1/2023.acl-short.25 Faithfulness tests for natural language explanations . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 283--294...
-
[4]
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 a . https://doi.org/10.18653/v1/2020.emnlp-main.263 A diagnostic study of explainability techniques for text classification . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3256--3274, Online. Association for Comput...
-
[5]
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 b . https://doi.org/10.18653/v1/2020.acl-main.656 Generating fact checking explanations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7352--7364, Online. Association for Computational Linguistics
-
[6]
Milan Bhan, Jean-No \"e l Vittaut, Nicolas Chesneau, and Marie-Jeanne Lesot. 2024. https://aclanthology.org/2024.emnlp-main.615 Self- AMPLIFY : Improving small language models with self post hoc explanations . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10974--10991, Miami, Florida, USA. Association for...
work page 2024
-
[7]
Vincent D Blondel, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. 2008. Fast Unfolding of Communities in Large Networks . Journal of statistical mechanics: theory and experiment, 2008(10):P10008
work page 2008
-
[8]
Oana-Maria Camburu, Tim Rockt\" a schel, Thomas Lukasiewicz, and Phil Blunsom. 2018. https://proceedings.neurips.cc/paper_files/paper/2018/file/4c7a167bb329bd92580a99ce422d6fa6-Paper.pdf e-snli: Natural language inference with natural language explanations . In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc
2018
Show all 79 references
-
[9]
Thiago Castro Ferreira, Chris van der Lee, Emiel van Miltenburg, and Emiel Krahmer. 2019. https://doi.org/10.18653/v1/D19-1052 Neural data-to-text generation: A comparison between pipeline and end-to-end architectures . In Proceedings of the 2019 Conference on Empirical Method...
2019 doi
-
[10]
Yu-Neng Chuang, Guanchu Wang, Chia-Yuan Chang, Ruixiang Tang, Shaochen Zhong, Fan Yang, Mengnan Du, Xuanting Cai, and Xia Hu. 2024. https://arxiv.org/abs/2402.04678 FaithLM: Towards Faithful Explanations for Large Language Models . Preprint, arXiv:2402.04678
2024
-
[11]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. https://doi.org/10.18653/v1/W19-4828 What does BERT look at? an analysis of BERT ' s attention . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NL...
2019 doi
-
[12]
Zichu Fei, Qi Zhang, and Yaqian Zhou. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.201 Iterative GNN -based decoder for question generation . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2573--2582, Online and Punta Cana...
2021 doi
-
[13]
Nils Feldhus, Leonhard Hennig, Maximilian Dustin Nasert, Christopher Ebert, Robert Schwarzenberg, and Sebastian M \"o ller. 2022. Constructing natural language explanations via saliency map verbalization. arXiv preprint arXiv:2210.07222
2022 arXiv
-
[14]
Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. https://doi.org/10.18653/v1/W17-3518 The W eb NLG challenge: Generating text from RDF data . In Proceedings of the 10th International Conference on Natural Language Generation, pages 124--1...
2017 doi
-
[15]
Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/5dd9db5e033da9c6fb5ba83c7a7ebea9-Paper.pdf Inductive representation learning on large graphs . In Advances in Neural Information Processing Systems, volume 30. Curra...
2017
-
[16]
Kehang Han, Balaji Lakshminarayanan, and Jeremiah Zhe Liu. 2021. https://openreview.net/forum?id=311QRRkfrep Reliable graph neural networks for drug discovery under distributional shift . In NeurIPS 2021 Workshop on Distribution Shifts: Connecting Methods and Applications
2021
-
[17]
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2021. https://arxiv.org/abs/2005.00687 Open graph benchmark: Datasets for machine learning on graphs . Preprint, arXiv:2005.00687
2021 arXiv
-
[18]
Fan Huang, Haewoon Kwak, and Jisun An. 2023. Chain of Explanation: New Prompting Method to Generate Quality Natural Language Explanation for Implicit Hate Speech . In Companion Proceedings of the ACM Web Conference 2023, pages 90--93
2023
-
[19]
Sarthak Jain and Byron C. Wallace. 2019. https://doi.org/10.18653/v1/N19-1357 A ttention is not E xplanation . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and...
2019 doi
-
[20]
Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024. How Interpretable are Reasoning Explanations from Prompting Large Language Models? In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2148--2164
2024
-
[21]
Shailza Jolly, Pepa Atanasova, and Isabelle Augenstein. 2022. Generating Fluent Fact Checking Explanations with Unsupervised Post-editing . Information, 13(10):500
2022
-
[22]
Kipf and Max Welling
Thomas N. Kipf and Max Welling. 2017. https://openreview.net/forum?id=SJU4ayYgl Semi-supervised classification with graph convolutional networks . In International Conference on Learning Representations
2017
-
[23]
Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019. https://doi.org/10.18653/v1/N19-1238 T ext G eneration from K nowledge G raphs with G raph T ransformers . In Proceedings of the 2019 Conference of the North A merican Chapter of the ...
2019 doi
-
[24]
Satyapriya Krishna, Jiaqi Ma, Dylan Z Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju. 2023. https://openreview.net/forum?id=3H37XciUEv Post hoc explanations of language models can improve language models . In Thirty-seventh Conference on Neural Information Pr...
2023
-
[25]
Sawan Kumar and Partha Talukdar. 2020. NILE: Natural Language Inference with Faithful Natural Language Explanations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8730--8742
2020
-
[26]
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill. 2022. Can Language Models Learn from Explanations in Context? In Findings of the Association for Computational Linguistics: EMNLP 2022, ...
2022
-
[27]
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. 2023. Measuring Faithfulness in Chain-of-Thought Reasoning . arXiv preprint arXiv:2307.13702
2023 arXiv
-
[28]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...
2020 doi
-
[29]
Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2024. Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasoners . arXiv preprint arXiv:2405.18915
2024 arXiv
-
[30]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[31]
Yuxiao Lin, Yuxian Meng, Xiaofei Sun, Qinghong Han, Kun Kuang, Jiwei Li, and Fei Wu. 2021. https://doi.org/10.18653/v1/2021.findings-acl.126 B ert GCN : Transductive text classification by combining GNN and BERT . In Findings of the Association for Computational Linguistics: A...
2021 doi
-
[32]
Wei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang, Haozhao Wang, YuanKai Zhang, and Ruixuan Li. 2024 a . https://openreview.net/forum?id=eAqcVZx30k Is the MMI criterion necessary for interpretability? degenerating non-causal features to plain noise for self-rationalization . In The...
2024
-
[33]
Wei Liu, Haozhao Wang, Jun Wang, Zhiying Deng, Yuankai Zhang, Cheng Wang, and Ruixuan Li. 2024 b . Enhancing the rationale-input alignment for self-explaining rationalization. In 2024 IEEE 40th International Conference on Data Engineering (ICDE), pages 2218--2230. IEEE
2024
-
[34]
Wei Liu, Haozhao Wang, Jun Wang, Ruixuan Li, Xinyang Li, YuanKai Zhang, and Yang Qiu. 2023 a . https://doi.org/10.18653/v1/2023.acl-long.715 MGR : Multi-generator based rationalization . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023 doi
-
[35]
Wei Liu, Jun Wang, Haozhao Wang, Ruixuan Li, Zhiying Deng, YuanKai Zhang, and Yang Qiu. 2023 b . https://proceedings.neurips.cc/paper_files/paper/2023/file/87e82678c0d6e5b729398426f82e9af6-Paper-Conference.pdf D-separation for causal self-explanation . In Advances in Neural In...
2023
-
[36]
Wei Liu, Jun Wang, Haozhao Wang, Ruixuan Li, Yang Qiu, Yuankai Zhang, Jie Han, and Yixiong Zou. 2023 c . Decoupled rationalization with asymmetric learning rates: A flexible lipschitz restraint. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data M...
2023
-
[37]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[38]
Scott M Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions . Advances in neural information processing systems, 30
2017
-
[39]
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch. 2024. Towards Faithful Model Explanation in Nlp: A Survey . Computational Linguistics, pages 1--67
2024
-
[40]
Bodhisattwa Prasad Majumder, Oana-Maria Camburu, Thomas Lukasiewicz, and Julian McAuley. 2021. Knowledge-Grounded Self-Rationalization via Extractive and Natural Language Explanations . arXiv preprint arXiv:2106.13876
2021 arXiv
-
[41]
Ana Marasovic, Iz Beltagy, Doug Downey, and Matthew Peters. 2022. https://doi.org/10.18653/v1/2022.findings-naacl.31 Few-shot self-rationalization with natural language prompts . In Findings of the Association for Computational Linguistics: NAACL 2022, pages 410--424, Seattle,...
2022 doi
-
[42]
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. Wt5?! Training Text-to-Text Models to Explain Their Predictions . arXiv preprint arXiv:2004.14546
2020 arXiv
-
[43]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...
2002
-
[44]
Letitia Parcalabescu and Anette Frank. 2024. https://doi.org/10.18653/v1/2024.acl-long.329 On measuring faithfulness or self-consistency of natural language explanations . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2024 doi
-
[45]
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning . arXiv preprint arXiv:2402.13950
2024 arXiv
-
[46]
Matt Post. 2018. https://doi.org/10.18653/v1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Brussels, Belgium. Association for Computational Linguistics
2018 doi
-
[47]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the Limits of Transfer Learning with A Unified Text-to-Text Transformer . Journal of machine learning research, 21(140):1--67
2020
-
[48]
Sagnik Ray Choudhury, Pepa Atanasova, and Isabelle Augenstein. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.783 Explaining interactions between text spans . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12709--12730, Sing...
2023 doi
-
[49]
Leonardo F. R. Ribeiro, Martin Schmitt, Hinrich Sch \"u tze, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.nlp4convai-1.20 Investigating pretrained language models for graph-to-text generation . In Proceedings of the 3rd Workshop on Natural Language Processing for...
2021 doi
-
[50]
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. https://doi.org/10.1162/tacl_a_00349 A primer in BERT ology: What we know about how BERT works . Transactions of the Association for Computational Linguistics, 8:842--866
2020 doi
-
[51]
Sofia Serrano and Noah A. Smith. 2019. https://doi.org/10.18653/v1/P19-1282 Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2931--2951, Florence, Italy. Association for Computational Linguistics
2019 doi
-
[52]
Jingyi Sun, Pepa Atanasova, and Isabelle Augenstein. 2024. https://arxiv.org/abs/2406.15085 A unified framework for input feature attribution analysis . Preprint, arXiv:2406.15085
2024 arXiv
-
[53]
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic Attribution for Deep Networks . In International conference on machine learning, pages 3319--3328. PMLR
2017
-
[54]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 C ommonsense QA : A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associa...
2019 doi
-
[55]
Xuejiao Tang, Xin Huang, Wenbin Zhang, Travers B Child, Qiong Hu, Zhen Liu, and Ji Zhang. 2021. Cognitive Visual Commonsense Reasoning Using Dynamic Working Memory . In Big Data Analytics and Knowledge Discovery: 23rd International Conference, DaWaK 2021, Virtual Event, Septem...
2021
-
[56]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting . Advances in Neural Information Processing Systems, 36
2024
-
[57]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems
2017
-
[58]
Petar Veli c kovi \'c , Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li \`o , and Yoshua Bengio. 2018. Graph Attention Networks . In International Conference on Learning Representations
2018
-
[59]
Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020. https://doi.org/10.18653/v1/2020.semeval-1.39 S em E val-2020 task 4: Commonsense validation and explanation . In Proceedings of the Fourteenth Workshop on Semantic Evaluation, pages 307--...
2020 doi
-
[60]
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023 a . https://doi.org/10.18653/v1/2023.emnlp-main.609 Label words are anchors: An information flow perspective for understanding in-context learning . In Proceedings of the 2023 Conferenc...
2023 doi
-
[61]
PINTO: Faithful Language Reasoning Using Prompt-Generated Rationales
PeiFeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen, and Xiang Ren. PINTO: Faithful Language Reasoning Using Prompt-Generated Rationales . In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022
2022
-
[62]
Peifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao, Bing Yin, and Xiang Ren. 2023 b . SCOTT: Self-Consistent Chain-of-Thought Distillation . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5546--5558
2023
-
[63]
Ronald L Wasserstein and Nicole A Lazar. 2016. The asa statement on p-values: context, process, and purpose
2016
-
[64]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models . Advances in neural information processing systems, 35:24824--24837
2022
-
[65]
Sarah Wiegreffe, Ana Marasovi \'c , and Noah A Smith. 2021. Measuring Association Between Labels and Free-Text Rationales . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10266--10284
2021
-
[66]
Neemesh Yadav, Sarah Masud, Vikram Goyal, Md Shad Akhtar, and Tanmoy Chakraborty. 2024. Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech . arXiv preprint arXiv:2406.03953
2024
-
[67]
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. https://doi.org/10.18653/v1/2021.naacl-main.45 QA - GNN : Reasoning with language models and knowledge graphs for question answering . In Proceedings of the 2021 Conference of the North Ame...
2021 doi
-
[68]
Shuzhou Yuan and Michael Faerber. 2023. https://aclanthology.org/2023.ranlp-1.133 Evaluating generative models for graph-to-text generation . In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, pages 1256--1264, Varna, Bulgari...
2023
-
[69]
Shuzhou Yuan and Michael F \"a rber. 2024. https://aclanthology.org/2024.findings-naacl.58 G ra SAME : Injecting token-level structural information to pretrained language models via graph-guided self-attention mechanism . In Findings of the Association for Computational Lingui...
2024
-
[70]
Shuzhou Yuan, Ercong Nie, Michael F \"a rber, Helmut Schmid, and Hinrich Schuetze. 2024. https://doi.org/10.18653/v1/2024.findings-acl.237 GNN avi: Navigating the information flow in large language models by graph neural network . In Findings of the Association for Computation...
2024 doi
-
[71]
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. https://proceedings.neurips.cc/paper/2021/file/e4d2b6e6fdeca3e60e0f1a62fee3d9dd-Paper.pdf Bartscore: Evaluating generated text as text generation . In Advances in Neural Information Processing Systems, volume 34, pages 27263--...
2021
-
[72]
Jiang Zhang, Qiong Wu, Yiming Xu, Cheng Cao, Zheng Du, and Konstantinos Psounis. 2024 a . Efficient Toxic Content Detection by Bootstrapping and Distilling Large Language Models . In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21779--21787
2024
-
[73]
Qingru Zhang, Xiaodong Yu, Chandan Singh, Xiaodong Liu, Liyuan Liu, Jianfeng Gao, Tuo Zhao, Dan Roth, and Hao Cheng. 2024 b . Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering . arXiv preprint arXiv:2409.10790
2024 arXiv
-
[74]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with bert . In International Conference on Learning Representations
2020
-
[75]
X Zhang, A Bosselut, M Yasunaga, H Ren, P Liang, C Manning, and J Leskovec. 2022 a . GreaseLM: Graph REASoning Enhanced Language Models for Question Answering . In International Conference on Representation Learning (ICLR)
2022
-
[76]
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022 b . Automatic Chain of Thought Prompting in Large Language Models . arXiv preprint arXiv:2210.03493
2022 arXiv
-
[77]
Meyer, and Steffen Eger
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. https://doi.org/10.18653/v1/D19-1053 M over S core: Text generation evaluating with contextualized embeddings and earth mover distance . In Proceedings of the 2019 Conference on Empirical ...
2019 doi
-
[78]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[79]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.