REVIEW 16 cited by
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermediate reasoning steps for explaining their behavior. Self-explanations have seen widespread adoption owing to their conversational and plausible nature. However, there is little to no understanding of their faithfulness. In this work, we discuss the dichotomy between faithfulness and plausibility in SEs generated by LLMs. We argue that while LLMs are adept at generating plausible explanations -- seemingly logical and coherent to human users -- these explanations do not necessarily align with the reasoning processes of the LLMs, raising concerns about their faithfulness. We highlight that the current trend towards increasing the plausibility of explanations, primarily driven by the demand for user-friendly interfaces, may come at the cost of diminishing their faithfulness. We assert that the faithfulness of explanations is critical in LLMs employed for high-stakes decision-making. Moreover, we emphasize the need for a systematic characterization of faithfulness-plausibility requirements of different real-world applications and ensure explanations meet those needs. While there are several approaches to improving plausibility, improving faithfulness is an open challenge. We call upon the community to develop novel methods to enhance the faithfulness of self explanations thereby enabling transparent deployment of LLMs in diverse high-stakes settings.
Forward citations
Cited by 16 Pith papers
-
Mechanistic Attention Guidance for Agent Memory Refinement
Attention patterns reveal how an AI agent uses memory, and using these patterns to rewrite memory improves task performance and memory efficiency.
-
PEDANTIC: A Dataset for the Automatic Examination of Definiteness in Patent Claims
PEDANTIC provides the first public dataset of 14k patent claims labeled with examiner-cited reasons for indefiniteness, along with baselines showing LLMs still lag logistic regression on binary prediction.
-
Training Large Language Models for Self-Explanation Faithfulness
RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions
Human experts are cautious about intervening in online discussions while six open-source LLMs are eager to step in, and a fine-tuned ModernBert classifier predicts real facilitator interventions more reliably than any...
-
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
A skeleton-first reasoning generation method reduces answer anchoring in reverse chain-of-thought traces, while semantic suppression increases latent anchoring.
-
The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models
Topological features of reasoning-trace embeddings correlate with Smith-Waterman alignment to expert AIME solutions more than graph metrics do, but the paper does not validate this out of sample.
-
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
An LLM-augmented synthetic data pipeline produces the largest public eviction-focused SDoH dataset (14 categories) and fine-tuned open LLMs that outperform prompt-optimized GPT-4o on the authors' test sets.
-
Energy-Based Transformers are Scalable Learners and Thinkers
Energy-Based Transformers learn to predict by gradient-descent minimization of a learned energy function, and the paper reports faster pretraining scaling and inference-time thinking gains over Transformer++ and Diffu...
-
TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning
TRUE checks whether LLM reasoning traces are self-sufficient by executing them blind, maps neighboring reasoning paths into a DAG, and ranks recurring failure modes by Shapley values.
-
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.
-
From Plausible to Actionable: A Position on LLM Self-Explanations
Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.
-
Governing Generative AI Across Financial Institutions: A Framework for Generative AI Risk Control
GAICF maps SR 26-2 model-risk principles into approved-use gates, risk tiers, evidence checks, and output monitoring for generative AI outside the formal model boundary.
-
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.
-
Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
KG-SMILE applies perturbation and linear regression to a knowledge graph to attribute which entities and relations drive a GraphRAG system's answers.
-
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
An argument, grounded in social-science measurement theory, that LLM-as-judge adoption has outpaced validity and reliability testing, with an analysis of four underlying assumptions.
Discussion (0). Continue with ORCID to comment.