Pith. sign in

REVIEW 9 cited by

Post-hoc Concept Bottleneck Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15480 v2 pith:G5HBA3ZA submitted 2022-05-31 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords bottleneckconceptconceptsmodelcbmsmodelspcbmdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Concept Bottleneck Models (CBMs) map the inputs onto a set of interpretable concepts (``the bottleneck'') and use the concepts to make predictions. A concept bottleneck enhances interpretability since it can be investigated to understand what concepts the model "sees" in an input and which of these concepts are deemed important. However, CBMs are restrictive in practice as they require dense concept annotations in the training data to learn the bottleneck. Moreover, CBMs often do not match the accuracy of an unrestricted neural network, reducing the incentive to deploy them in practice. In this work, we address these limitations of CBMs by introducing Post-hoc Concept Bottleneck models (PCBMs). We show that we can turn any neural network into a PCBM without sacrificing model performance while still retaining the interpretability benefits. When concept annotations are not available on the training data, we show that PCBM can transfer concepts from other datasets or from natural language descriptions of concepts via multimodal models. A key benefit of PCBM is that it enables users to quickly debug and update the model to reduce spurious correlations and improve generalization to new distributions. PCBM allows for global model edits, which can be more efficient than previous works on local interventions that fix a specific prediction. Through a model-editing user study, we show that editing PCBMs via concept-level feedback can provide significant performance gains without using data from the target domain or model retraining.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 37 citations worldwide. Full citation record

  1. Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniCon unifies dermoscopic and clinical concept vocabularies in a shared text-embedding codebook, achieving state-of-the-art interpretable diagnosis and test-time intervention on skin lesion benchmarks.

  2. Explainable Novel Category Discovery in Semantic Concept Space

    cs.CV 2026-07 conditional novelty 6.0 of 10

    xNCD routes novel category discovery through a CLIP-aligned concept bottleneck, matching strong NCD baselines while producing intrinsic cluster- and instance-level concept explanations.

  3. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  4. MVP-CBM:Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image Classification

    cs.CV 2025-06 conditional novelty 6.0 of 10

    By modeling which visual layers best explain each diagnostic concept and sparsely fusing multi-layer concept activations, MVP-CBM improves accuracy and interpretability over prior concept bottleneck models on seven me...

  5. Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.

  6. Towards Interpretable PolSAR Image Classification: Polarimetric Scattering Mechanism Informed Concept Bottleneck and Kolmogorov-Arnold Network

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A concept bottleneck model built from polarimetric target decomposition plus a Kolmogorov-Arnold Network gives PolSAR classification with human-auditable concept predictions and symbolic decision formulas.

  7. DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models

    cs.AI 2025-05 conditional novelty 5.0 of 10

    DeCoDe combines concept bottleneck models with learning to defer to select per-instance among AI-only, human-only, and AI+human strategies, reporting accuracy gains over binary deferral baselines on three image datasets.

  8. Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A broad modality-aware survey of explainable AI methods for biomedical imaging, covering heatmap, concept, text, and latent-space approaches plus tools, metrics, and vision-language models.

  9. Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SATs aggregate local saliency maps over semantic segments to produce semi-global rankings of feature influence, exposing shortcut reliance even when out-of-distribution accuracy barely changes.

Pith tools