Pith. sign in

REVIEW 5 cited by

Do Concept Bottleneck Models Learn as Intended?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.04289 v1 pith:LNX4ZISJ submitted 2021-05-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsbottleneckconceptconceptsinterpretabilitymeetanythingbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Concept bottleneck models map from raw inputs to concepts, and then from concepts to targets. Such models aim to incorporate pre-specified, high-level concepts into the learning procedure, and have been motivated to meet three desiderata: interpretability, predictability, and intervenability. However, we find that concept bottleneck models struggle to meet these goals. Using post hoc interpretability methods, we demonstrate that concepts do not correspond to anything semantically meaningful in input space, thus calling into question the usefulness of concept bottleneck models in their current form.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.

  2. Locality-aware Concept Bottleneck Model

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A label-free concept bottleneck model using per-concept prototypes aligned by CLIP to localize concept predictions to the correct image regions.

  3. Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.

  4. ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

    cs.AI 2026-07 conditional novelty 4.0 of 10

    A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.

  5. A Geometric Unification of Concept Learning with Concept Cones

    cs.AI 2025-12 conditional novelty 4.0 of 10

    CBMs and SAEs both learn nonnegative linear concept cones; a new containment metric suite scores SAE dictionaries against CBM concepts.

Pith tools