Pith. sign in

REVIEW 2 cited by

Linearly-Interpretable Concept Embedding Models for Text Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14335 v2 pith:MW5Q4YRS submitted 2024-06-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsconceptinterpretableanalysisbeencbmsconceptsembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite their success, Large-Language Models (LLMs) still face criticism due to their lack of interpretability. Traditional post-hoc interpretation methods, based on attention and gradient-based analysis, offer limited insights as they only approximate the model's decision-making processes and have been proved to be unreliable. For this reason, Concept-Bottleneck Models (CBMs) have been lately proposed in the textual field to provide interpretable predictions based on human-understandable concepts. However, CBMs still exhibit several limitations due to their architectural constraints limiting their expressivity, to the absence of task-interpretability when employing non-linear task predictors and for requiring extensive annotations that are impractical for real-world text data. In this paper, we address these challenges by proposing a novel Linearly Interpretable Concept Embedding Model (LICEM) going beyond the current accuracy-interpretability trade-off. LICEMs classification accuracy is better than existing interpretable models and matches black-box ones. We show that the explanations provided by our models are more interveneable and causally consistent with respect to existing solutions. Finally, we show that LICEMs can be trained without requiring any concept supervision, as concepts can be automatically predicted when using an LLM backbone.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ECO-Concept uses slot attention plus LLM-based comprehensibility feedback to learn explainable text concepts without concept annotations.

  2. A Self-Explainable Deep Architecture for Security Applications

    cs.CR 2026-08 conditional novelty 4.0 of 10

    XSec is a prototype-based deep architecture that jointly classifies security data and produces deterministic feature-importance explanations from learned masks and class-specific prototypes.

Pith tools