REVIEW 4 major objections 4 minor 55 references
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Single-pass decoding boosts context-faithful QA answers by 17.7%
desk verdict A solid, well-scoped decoding paper whose probing result is the strongest part; referee should push on per-step detector validity and alpha consistency. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the attention ratio, defined for a context token $j$ in head $h$ of layer $l$ as $r^j_{l,h} = a^j_{l,h} / \sum_{j \in C} a^j_{l,h}$, the share of attention that token receives among all context tokens, which normalizes away attention-sink noise and cross-head magnitude differences. Aggregating these ratios across heads gives a feature vector per token; a logistic-regression classifier trained on those vectors, with ground-truth 'utilized' meaning the token matches the gold answer, becomes the Context Utilization Detector. At inference the classifier's learned coefficients weight the ratios to form utilization scores $s_j$, normalized into a utilization distribution $U$; a top-rank constraint keeps only tokens already ranked in the top-$R$ of the generation distribution, and the final adjustment $P' = P + \alpha H_{\mathrm{norm}}(P) \cdot U_{\mathrm{top}}$ amplifies those tokens in proportion to the current token-level normalized entropy. This mechanism carries the argument: it converts the claimed attention-ratio signal into a probability-mass shift within one decoding pass, using only top-10 attention heads and a 100-sample training set.
What would settle it
Retrain the probing detector using labels taken from attention maps at the second or later generated tokens, or labeling non-answer support tokens as utilized. The paper's claim predicts AUC comparable to the reported >0.99; if the classifier instead falls toward chance, the signal is specific to first-token answer-word recognition and DAGCD's amplification has no demonstrated mechanism beyond the first generated token.
Extended reading notes
Core claim
The paper's central claim is that context faithfulness hallucinations in retrieval-augmented generation are not caused by the model ignoring the supplied context, but by its failing to prioritize context tokens it has already identified as relevant. The supporting observations are that wrong answers show higher uncertainty (average normalized entropy 0.36 vs 0.29 for correct answers, average maximum softmax probability 0.25 vs 0.41), and that in 66% of wrong cases the gold-answer token is ranked within the top 10 of the token-level distribution. A logistic-regression probe trained on per-head attention ratios classifies 'utilized' context tokens with AUC above 0.99 across six domains, even with only 100 training samples. DAGCD operationalizes this signal: at each decoding step it detects utilized context tokens from attention ratios, builds a utilization distribution over them, and adjusts the generation distribution as $P' = P + \alpha H_{\mathrm{norm}}(P) \cdot U_{\mathrm{top}}$, where $H_{\mathrm{norm}}(P)$ is the normalized entropy and $\alpha$ is a per-model constant. The paper reports that this adjustment outperforms greedy decoding, CAD, and COIECD on seven open-book QA datasets, with the largest gains on multi-hop reasoning and adversarial-swap settings.
Load-bearing premise
The load-bearing premise, which the paper's limitations section flags, is that a detector trained to label 'utilized' context tokens as those matching the gold answer at the first generated token continues to identify the context tokens that should be amplified at every later decoding step, including multi-token answers.
Editorial extensions
If this is right
- Because DAGCD reuses attention already computed during greedy decoding and needs no second forward pass, it maintains the theoretical time complexity of greedy decoding while reporting average EM gains of 17.67% on pretrained models and 2.25% on instruction-tuned models.
- The largest reported gains are in settings where retrieved evidence is hardest to prioritize: multi-hop HotpotQA (18.80% EM on Mistral-7B) and adversarial NQ-swap (74.52% EM on Mistral-7B), suggesting the method corrects failures of evidence prioritization rather than generic decoding artifacts.
- The detector transfers across domains with AUC above 0.99 when trained on 100 samples from HotpotQA, so applying DAGCD to a new open-book QA setup requires only a tiny calibration set rather than dataset-specific detector training.
- On instruction-tuned models the absolute gains shrink but DAGCD still leads all baselines, implying fine-tuning already compresses much of the uncertainty signal the method exploits.
Reading between the lines
- If the attention-ratio signal is genuinely a utilization signal, the same classifier output should provide a per-token interpretability map of which retrieved sentences influenced a generation; that map could be tested as a diagnostic for RAG failures in long-context and tool-use settings.
- The detector is trained on gold-answer-matching tokens at the first generated position; the claim that it detects utilization at every later step is an extrapolation that the paper does not test. Relabeling training data with tokens from later decoding positions, or with non-answer support tokens, would settle whether the signal is utilization or answer-word recognition.
- The paper's own limitations flag that the scaling factor $\alpha$ needs per-model calibration; a natural extension is to make $\alpha$ self-tuning by deriving it from a running estimate of normalized entropy, which would remove the manual tuning step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Dynamic Attention-Guided Context Decoding (DAGCD), a single-pass decoding method aimed at reducing context faithfulness hallucinations in open-book QA. The authors first present analyses linking token-level uncertainty to unfaithful answers, then train a logistic-regression probing classifier on 'attention ratio' features to identify context tokens that are 'utilized,' where positive labels are context tokens matching the gold answer. At inference, DAGCD applies this detector at every decoding step to build a utilization distribution, restricts it to top-ranked tokens, and adds an entropy-scaled version of this distribution to the next-token probabilities (Eq. 6). Experiments across seven QA datasets and six LLAMA/Mistral models report consistent EM/F1 gains over greedy decoding and over CAD and COIECD baselines, with headline improvements of 17.67% EM on pretrained models and 2.25% on instruction-tuned models.
Significance. If the stated mechanism were established, DAGCD would be a useful lightweight contribution: it requires no second decoding pass, uses a small logistic-regression detector that transfers across models and datasets with reported AUC above 0.99, and is accompanied by ablations over training size, top-rank constraint, and scaling factor. The reproducible recipe (attention-ratio features, top-K head selection, entropy-scaled additive adjustment) is concrete and falsifiable. However, the significance is currently limited by a gap between the probing setup and the inference-time use of the detector: the classifier is trained only on first-token gold-answer-matching labels, yet it is applied at every decoding step to all context tokens. Until per-step and non-answer-token behavior is validated, the reported gains cannot be confidently attributed to the 'context utilization signal' that the paper claims to exploit. The paper is honest about alpha calibration difficulty and classifier robustness in its Limitations section, but those admitted limitations interact with the core evaluation and need to be addressed by additional experiments.
major comments (4)
- [§3.2, §4.1, Eq. (6)] The detector is trained under conditions that differ from its inference-time use. Section 3.2 constructs positive labels as context tokens that match the gold answer, and the feature vectors are extracted at the first generated token after context concatenation. Section 4.1 then applies the same detector at every decoding step to all context tokens and builds the utilization distribution used in Eq. (6). Nothing in Sections 2–5 or the appendices measures detector accuracy or agreement as a function of decoding step, nor for tokens that support multi-hop reasoning without being the answer string. For multi-token answers, which are common in SQuAD, NQ, and HotpotQA, later decoding steps are out of distribution for the detector. The central claim—that DAGCD amplifies actually utilized context tokens—requires per-step validation, for example by reporting detector precision/recall at decoding positions 1, 2, 3, ... and on non-answer context tokens that appear in correct multi-token generations.
- [§4.2, §5.3] The paper never ablates the detector itself. DAGCD's adjustment is P' = P + α·H_norm(P)·U_top, where U_top is built from detector-filtered utilization scores. A natural control is to replace the detector output with a no-classifier baseline, e.g., U computed directly from raw attention ratios without thresholding, or with a random labeling of context tokens. The ablations in Figure 5 vary detector training data size, top-rank constraint, and α, but none removes the detector. Given that the positive training labels are gold-answer tokens, part of the EM gain could arise simply from amplifying any high-attention, top-ranked context token; without the no-detector control, the load-bearing claim that the learned classifier is necessary is unsupported.
- [§5.1, Appendix E.2] The reported hyperparameter configuration is internally inconsistent. Section 5.1 states that α is set to 2 for pretrained models and 4 for instruction-tuned models. Appendix E.2 and Figure 13 report that on HotpotQA the optimal α for Mistral-7B is 5, while Mistral-7B-Instruct stabilizes only at α = 13. The abstract and Section 1 summarize gains as aggregate EM improvements of 17.67% and 2.25%, but it is unclear whether Table 1 uses the fixed values from Section 5.1 or per-model optimal values from the appendix. This distinction matters because selecting α per model from the test set would affect the fairness of the comparison to baselines. The authors should state explicitly which α values produced Table 1, and should present sensitivity results for all examined models, not only LLaMA2 and Mistral on HotpotQA.
- [§3.2, §4.1] The interpretation of the probing result as evidence of a 'context utilization signal' is stronger than the labels support. Positive examples are defined as context tokens that match the gold answer, so the classifier learns to recognize answer-word-like tokens under open-book conditions, not to detect all tokens that contribute to reasoning or to faithful multi-token generation. The high AUC (>0.99) and cross-dataset generalization are consistent with this narrower reading. The paper should either soften the 'fundamental mechanism' claim (Section 3.3 and Conclusion) or add a probing evaluation with richer utilization labels, for example tokens that are copied into a correct answer span at the step they are generated, or tokens whose removal changes the model's prediction.
minor comments (4)
- [§5.2, Table 1] The row label 'OURs' is a typo; it should read 'DAGCD' or 'Ours' consistently.
- [References] References for Meng et al. 2022a and 2022b are identical in content, as are the two Olsson et al. 2022 entries; these duplicate entries should be merged or disambiguated.
- [§4.3, Eq. (6)] The adjusted distribution P' is not renormalized. Since the paper reports greedy decoding, argmax is unaffected by the missing normalization, but the method description and any sampling-based use would require renormalization; this should be stated explicitly.
- [Appendix A.2] The prompt contains a space before the colon in 'information: {context}', which is likely intentional but should be checked for consistency with the prompt templates in Appendix F.
Circularity Check
The 'context utilization' signal is defined by gold-answer labels, and DAGCD amplifies exactly those tokens; part of the EM gain is built into the supervision.
-
fitted input called prediction
[Section 3.2 (Data Construction), Section 4.3 (Eq. 6), Section 5.1 (Metrics)]
"Context tokens were labeled as positive (utilized) if they corresponded to the gold answer, and negative (non-utilized) otherwise. ... The adjusted generation distribution P ′ is computed as: P ′ = P + αHnorm(P ) · Utop. ... we use EM and F1 score metrics to evaluate the performance."
The detector's target concept ('utilized') is defined as 'corresponded to the gold answer.' The utilization scores s_j (Eq. 4) and the distribution U_top injected in Eq. 6 are built from this detector's outputs. The paper's headline result is EM/F1 against the same gold-answer labels. So the decoding adjustment is, by construction, amplifying tokens the detector was trained to recognize as gold-answer tokens; part of the reported EM gain is the direct effect of the training supervision rather than evidence of an independent 'context utilization signal.' Held-out transfer and cross-domain AUC provide independent content, but the 'fundamental mechanism' interpretation is confounded with answer-span detection.
full rationale
The paper is not relying on self-citation or an imported uniqueness theorem. The method is an empirical, externally benchmarked decoding strategy, and its comparisons against greedy decoding, CAD, and COIECD are self-contained. The partial circularity lies in the definitional loop: the probing classifier that motivates the method is trained to label tokens matching the gold answer as 'utilized,' and the decoding update (Eq. 6) then boosts tokens classified as utilized, while the evaluation metric is EM/F1 against those same gold answers. This makes some of the improvement a supervised answer-spotting effect rather than a demonstration of a general, task-independent attention signal. The lack of per-step validation of the detector (trained on first-token features but applied at every decoding step) is a correctness risk, not a circularity, and is noted separately.
Assumptions & free parameters
free parameters (4)
- Scaling factor alpha =
2 for pretrained, 4 for instruction-tuned; appendix reports 5 for Mistral-7B and above 13 for Mistral-7B-Instruct
- Top-K attention heads K =
10
- Top-R rank constraint R =
10
- Detector training set size =
100 samples (50 positive, 50 negative) from HotpotQA
assumptions (4)
- domain assumption The attention ratio defined in Eq. 1 measures how much a context token is utilized.
- ad hoc to paper A classifier trained on gold-answer tokens as 'utilized' generalizes to all context tokens and all decoding steps.
- ad hoc to paper The additive adjustment P + alpha H_norm(P) U_top improves faithfulness without renormalization.
- domain assumption Normalized entropy is a reliable per-token uncertainty signal for scaling context amplification.
invented entities (1)
-
Context utilization signal
Cite this review
Pith. "Pith review of Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models." pith.science (2026). https://pith.science/paper/O7ZKTARQ
@misc{pith2026250101059,
author = {Pith},
title = {Pith review of: Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/O7ZKTARQ}},
note = {Machine review of arXiv:2501.01059}
}
read the original abstract
Large language models (LLMs) often exhibit Context Faithfulness Hallucinations, where outputs deviate from retrieved information due to incomplete context integration. Our analysis reveals a strong correlation between token-level uncertainty and hallucinations. We hypothesize that attention mechanisms inherently encode context utilization signals, supported by probing analysis. Based on these insights, we propose Dynamic Attention-Guided Context Decoding (DAGCD), a lightweight framework that leverages attention distributions and uncertainty signals in a single-pass decoding. Experiments on open-book QA datasets demonstrate DAGCD's effectiveness, yielding significant improvements in faithfulness and robustness while preserving computational efficiency.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. https://api.semanticscholar.org/CorpusID:257532815 Gpt-4 technical report . ArXiv
work page 2023
-
[2]
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.627 Understanding and overcoming the challenges of efficient transformer quantization . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7947--7969, Online and Punta Cana, Dominican Republic. Associati...
-
[3]
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877--1901
work page 2020
-
[4]
Sehyun Choi, Tianqing Fang, Zhaowei Wang, and Yangqiu Song. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.867 KCTS : Knowledge-constrained tree search decoding with token-level hallucination detection . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14035--14053, Singapore. Association for Computationa...
-
[5]
Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, Yoon Kim, and James R. Glass. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.84 Lookback lens: Detecting and mitigating contextual hallucinations in large language models using only attention maps . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, ...
-
[6]
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024 b . https://openreview.net/forum?id=Th6NyL07na Dola: Decoding by contrasting layers improves factuality in large language models . In The Twelfth International Conference on Learning Representations
work page 2024
-
[7]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. https://doi.org/10.18653/v1/W19-4828 What does BERT look at? an analysis of BERT ' s attention . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276--286, Florence, Italy. Association for Computational Linguistics
-
[8]
Souvik Das, Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, and Dong Yu. 2025. Entropy guided extrapolative decoding to improve factuality in large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 6589--6600
work page 2025
Show all 55 references
-
[9]
Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017. Searchqa: A new q&a dataset augmented with context from a search engine. arXiv preprint arXiv:1704.05179
2017 arXiv
-
[10]
Wenqi Fan, Yujuan Ding, Liang bo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. https://api.semanticscholar.org/CorpusID:269740933 A survey on rag meeting llms: Towards retrieval-augmented large language models . In Knowledge Discovery and Data Mining
2024
-
[11]
Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.59 F act KB : Generalizable factuality evaluation using language models enhanced with factual knowledge . In Proceedings of the 2023 Conference on Empirical ...
2023 doi
-
[12]
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019. https://doi.org/10.18653/v1/D19-5801 MRQA 2019 shared task: Evaluating generalization in reading comprehension . In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pa...
2019 doi
-
[13]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. https://api.semanticscholar.org/CorpusID:266359151 Retrieval-augmented generation for large language models: A survey . ArXiv, abs/2312.10997
2023 arXiv
-
[14]
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.751 Dissecting recall of factual associations in auto-regressive language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language ...
2023 doi
-
[15]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. https://proceedings.mlr.press/v119/guu20a.html Retrieval augmented language model pre-training . In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of ...
2020
-
[16]
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021. Self-attention attribution: Interpreting information interactions inside transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 12963--12971
2021
-
[17]
Dan Hendrycks and Kevin Gimpel. 2017. https://openreview.net/forum?id=Hkg4TI9xl A baseline for detecting misclassified and out-of-distribution examples in neural networks . In International Conference on Learning Representations
2017
-
[18]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023 a . A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on I...
2023
-
[19]
Yuheng Huang, Jiayang Song, Zhijie Wang, Huaming Chen, and Lei Ma. 2023 b . https://api.semanticscholar.org/CorpusID:259991714 Look before you leap: An exploratory study of uncertainty measurement for large language models . ArXiv, abs/2307.10236
2023 arXiv
-
[20]
Sarthak Jain and Byron C. Wallace. 2019. https://doi.org/10.18653/v1/N19-1357 Attention is not explanation . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and S...
2019 doi
-
[21]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. https://doi.org/10.1145/3571730 Survey of hallucination in natural language generation . ACM Comput. Surv., 55(12)
2023 doi
-
[22]
Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Xiaojian Jiang, Jiexin Xu, Li Qiuxia, and Jun Zhao. 2024 a . https://aclanthology.org/2024.lrec-main.1466 Tug-of-war between knowledge: Exploring and resolving knowledge conflicts in retrieval-augmented language models . In Procee...
2024
-
[23]
Zhuoran Jin, Pengfei Cao, Hongbang Yuan, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao. 2024 b . https://doi.org/10.18653/v1/2024.findings-acl.70 Cutting off the head ends the conflict: A mechanism for interpreting and mitigating knowledge conflicts ...
2024 doi
-
[24]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. https://doi.org/10.18653/v1/P17-1147 T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension . In Proceedings of the 55th Annual Meeting of the Association for Computational...
2017 doi
-
[25]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019 doi
-
[26]
Deren Lei, Yaxi Li, Mengya Hu, Mingyu Wang, and Xi Yun. 2023. https://openreview.net/forum?id=byNSabn9cA Chain of natural language inference for reducing large language model hallucinations . In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following
2023
-
[27]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\" u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\" a schel, Sebastian Riedel, and Douwe Kiela. 2020. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f78...
2020
-
[28]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. https://doi.org/10.18653/v1/2023.acl-long.687 Contrastive decoding: Open-ended text generation as optimization . In Proceedings of the 61st Annual...
2023 doi
-
[29]
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Yi Wang, Zhonghao Wang, Feiyu Xiong, et al. 2024. Internal consistency and self-feedback in large language models: A survey. arXiv preprint arXiv:2407.14507
2024 arXiv
-
[30]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[31]
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.565 Entity-based knowledge conflicts in question answering . In Proceedings of the 2021 Conference on Empirical Methods in Natural L...
2021 doi
-
[32]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 a . Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359--17372
2022
-
[33]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 b . Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359--17372
2022
-
[35]
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 b . In-context learning and induction heads. arXiv preprint arXiv:2209.11895
2022 arXiv
-
[36]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. https://doi.org/10.18653/v1/D16-1264 SQ u AD : 100,000+ questions for machine comprehension of text . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383-...
2016 doi
-
[37]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. https://doi.org/10.18653/v1/P17-1099 Get to the point: Summarization with pointer-generator networks . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...
2017 doi
-
[38]
Dan Shi, Renren Jin, Tianhao Shen, Weilong Dong, Xinwei Wu, and Deyi Xiong. 2024 a . https://openreview.net/forum?id=ZfXRAqbBKX IRCAN : Mitigating knowledge conflicts in LLM generation via identifying and reweighting context-aware neurons . In The Thirty-eighth Annual Conferen...
2024
-
[39]
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024 b . https://doi.org/10.18653/v1/2024.naacl-short.69 Trusting your evidence: Hallucinate less with context-aware decoding . In Proceedings of the 2024 Conference of the North America...
2024 doi
-
[40]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. https://api.semanticscholar.org/CorpusID:259950998 Llama 2: Open foundation and fine-tuned chat models . ArX...
2023 arXiv
-
[41]
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017. https://doi.org/10.18653/v1/W17-2623 N ews QA : A machine comprehension dataset . In Proceedings of the 2nd Workshop on Representation Learning for NLP , pages ...
2017 doi
-
[42]
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019. Attention interpretability across nlp tasks. arXiv preprint arXiv:1909.11218
2019 arXiv
-
[43]
Jesse Vig and Yonatan Belinkov. 2019. https://doi.org/10.18653/v1/W19-4808 Analyzing the structure of attention in a transformer language model . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63--76, Florence, It...
2019 doi
-
[44]
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2023. https://api.semanticscholar.org/CorpusID:263672149 Freshllms: Refreshing large language models with search engine augmentation . In Annu...
2023
-
[45]
Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. 2024. Adacad: Adaptively decoding to balance conflicts between contextual and parametric knowledge. arXiv preprint arXiv:2409.07394
2024 arXiv
-
[46]
Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu. 2024. Retrieval head mechanistically explains long-context factuality. arXiv preprint arXiv:2404.15574
2024 arXiv
-
[47]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. https://openreview.net/forum?id=NG7sS51zVF Efficient streaming language models with attention sinks . In The Twelfth International Conference on Learning Representations
2024
-
[48]
Song Xu, Haoran Li, Peng Yuan, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020. https://doi.org/10.18653/v1/2020.acl-main.125 Self-attention guided copy mechanism for abstractive summarization . In Proceedings of the 58th Annual Meeting of the Association for Computational Ling...
2020 doi
-
[49]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://doi.org/10.18653/v1/D18-1259 H otpot QA : A dataset for diverse, explainable multi-hop question answering . In Proceedings of the 2018 Conference...
2018 doi
-
[50]
Xiaowei Yuan, Zhao Yang, Yequan Wang, Shengping Liu, Jun Zhao, and Kang Liu. 2024. https://doi.org/10.18653/v1/2024.findings-acl.234 Discerning and resolving knowledge conflicts through adaptive decoding with contextual information-entropy constraint . In Findings of the Assoc...
2024 doi
-
[51]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with bert . In International Conference on Learning Representations
2020
-
[52]
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology, 15(2):1--38
2024
-
[53]
Zifan Zheng, Yezhaohui Wang, Yuxin Huang, Shichao Song, Bo Tang, Feiyu Xiong, and Zhiyu Li. 2024. Attention heads of large language models: A survey. arXiv preprint arXiv:2409.03752
2024 arXiv
-
[54]
Wenxuan Zhou, Sheng Zhang, Hoifung Poon, and Muhao Chen. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.968 Context-faithful prompting for large language models . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14544--14556, Singapore. As...
2023 doi
-
[55]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[56]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.