REVIEW 9 cited by
Post-hoc Concept Bottleneck Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Concept Bottleneck Models (CBMs) map the inputs onto a set of interpretable concepts (``the bottleneck'') and use the concepts to make predictions. A concept bottleneck enhances interpretability since it can be investigated to understand what concepts the model "sees" in an input and which of these concepts are deemed important. However, CBMs are restrictive in practice as they require dense concept annotations in the training data to learn the bottleneck. Moreover, CBMs often do not match the accuracy of an unrestricted neural network, reducing the incentive to deploy them in practice. In this work, we address these limitations of CBMs by introducing Post-hoc Concept Bottleneck models (PCBMs). We show that we can turn any neural network into a PCBM without sacrificing model performance while still retaining the interpretability benefits. When concept annotations are not available on the training data, we show that PCBM can transfer concepts from other datasets or from natural language descriptions of concepts via multimodal models. A key benefit of PCBM is that it enables users to quickly debug and update the model to reduce spurious correlations and improve generalization to new distributions. PCBM allows for global model edits, which can be more efficient than previous works on local interventions that fix a specific prediction. Through a model-editing user study, we show that editing PCBMs via concept-level feedback can provide significant performance gains without using data from the target domain or model retraining.
Forward citations
Cited by 9 Pith papers
-
Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis
UniCon unifies dermoscopic and clinical concept vocabularies in a shared text-embedding codebook, achieving state-of-the-art interpretable diagnosis and test-time intervention on skin lesion benchmarks.
-
Explainable Novel Category Discovery in Semantic Concept Space
xNCD routes novel category discovery through a CLIP-aligned concept bottleneck, matching strong NCD baselines while producing intrinsic cluster- and instance-level concept explanations.
-
Understanding and evaluating computer vision models through the lens of counterfactuals
Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.
-
MVP-CBM:Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image Classification
By modeling which visual layers best explain each diagnostic concept and sparsely fusing multi-layer concept activations, MVP-CBM improves accuracy and interpretability over prior concept bottleneck models on seven me...
-
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.
-
Towards Interpretable PolSAR Image Classification: Polarimetric Scattering Mechanism Informed Concept Bottleneck and Kolmogorov-Arnold Network
A concept bottleneck model built from polarimetric target decomposition plus a Kolmogorov-Arnold Network gives PolSAR classification with human-auditable concept predictions and symbolic decision formulas.
-
DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models
DeCoDe combines concept bottleneck models with learning to defer to select per-instance among AI-only, human-only, and AI+human strategies, reporting accuracy gains over binary deferral baselines on three image datasets.
-
Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey
A broad modality-aware survey of explainable AI methods for biomedical imaging, covering heatmap, concept, text, and latent-space approaches plus tools, metrics, and vision-language models.
-
Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification
SATs aggregate local saliency maps over semantic segments to produce semi-global rankings of feature influence, exposing shortcut reliance even when out-of-distribution accuracy barely changes.
Discussion (0). Continue with ORCID to comment.