Pith. sign in

REVIEW 16 cited by

Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04614 v3 pith:P6GXSDZY submitted 2024-02-07 cs.CL

classification cs.CL
keywords faithfulnessexplanationsllmsplausibilitylanguageapplicationshigh-stakesimproving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermediate reasoning steps for explaining their behavior. Self-explanations have seen widespread adoption owing to their conversational and plausible nature. However, there is little to no understanding of their faithfulness. In this work, we discuss the dichotomy between faithfulness and plausibility in SEs generated by LLMs. We argue that while LLMs are adept at generating plausible explanations -- seemingly logical and coherent to human users -- these explanations do not necessarily align with the reasoning processes of the LLMs, raising concerns about their faithfulness. We highlight that the current trend towards increasing the plausibility of explanations, primarily driven by the demand for user-friendly interfaces, may come at the cost of diminishing their faithfulness. We assert that the faithfulness of explanations is critical in LLMs employed for high-stakes decision-making. Moreover, we emphasize the need for a systematic characterization of faithfulness-plausibility requirements of different real-world applications and ensure explanations meet those needs. While there are several approaches to improving plausibility, improving faithfulness is an open challenge. We call upon the community to develop novel methods to enhance the faithfulness of self explanations thereby enabling transparent deployment of LLMs in diverse high-stakes settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mechanistic Attention Guidance for Agent Memory Refinement

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Attention patterns reveal how an AI agent uses memory, and using these patterns to rewrite memory improves task performance and memory efficiency.

  2. PEDANTIC: A Dataset for the Automatic Examination of Definiteness in Patent Claims

    cs.CL 2025-05 conditional novelty 7.0 of 10

    PEDANTIC provides the first public dataset of 14k patent claims labeled with examiner-cited reasons for indefiniteness, along with baselines showing LLMs still lag logistic regression on binary prediction.

  3. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  4. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  5. To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions

    cs.HC 2026-05 conditional novelty 6.0 of 10

    Human experts are cautious about intervening in online discussions while six open-source LLMs are eager to step in, and a fine-tuned ModernBert classifier predicts real facilitator interventions more reliably than any...

  6. Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

    cs.CL 2026-02 conditional novelty 6.0 of 10

    A skeleton-first reasoning generation method reduces answer anchoring in reverse chain-of-thought traces, while semantic suppression increases latent anchoring.

  7. The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models

    cs.AI 2025-10 unverdicted novelty 6.0 of 10

    Topological features of reasoning-trace embeddings correlate with Smith-Waterman alignment to expert AIME solutions more than graph metrics do, but the paper does not validate this out of sample.

  8. SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An LLM-augmented synthetic data pipeline produces the largest public eviction-focused SDoH dataset (14 categories) and fine-tuned open LLMs that outperform prompt-optimized GPT-4o on the authors' test sets.

  9. Energy-Based Transformers are Scalable Learners and Thinkers

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Energy-Based Transformers learn to predict by gradient-descent minimization of a learned energy function, and the paper reports faster pretraining scaling and inference-time thinking gains over Transformer++ and Diffu...

  10. TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning

    cs.LG 2026-02 reject novelty 5.0 of 10

    TRUE checks whether LLM reasoning traces are self-sufficient by executing them blind, maps neighboring reasoning paths into a DAG, and ranks recurring failure modes by Shapley values.

  11. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

    cs.HC 2025-02 conditional novelty 5.0 of 10

    Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.

  12. From Plausible to Actionable: A Position on LLM Self-Explanations

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.

  13. Governing Generative AI Across Financial Institutions: A Framework for Generative AI Risk Control

    q-fin.RM 2026-07 unverdicted novelty 4.0 of 10

    GAICF maps SR 26-2 model-risk principles into approved-use gates, risk tiers, evidence checks, and output monitoring for generative AI outside the formal model boundary.

  14. A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

    cs.CL 2025-09 conditional novelty 4.0 of 10

    A structured literature survey concluding that reasoning capabilities do not automatically make LLMs more trustworthy and can introduce new vulnerabilities in safety, robustness, and privacy.

  15. Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

    cs.AI 2025-09 reject novelty 4.0 of 10

    KG-SMILE applies perturbation and linear regression to a knowledge graph to attribute which entities and relations drive a GraphRAG system's answers.

  16. Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

    cs.CL 2025-08 conditional novelty 4.0 of 10

    An argument, grounded in social-science measurement theory, that LLM-as-judge adoption has outpaced validity and reliability testing, with an analysis of four underlying assumptions.

Pith tools