REVIEW 5 cited by
Do Concept Bottleneck Models Learn as Intended?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Concept bottleneck models map from raw inputs to concepts, and then from concepts to targets. Such models aim to incorporate pre-specified, high-level concepts into the learning procedure, and have been motivated to meet three desiderata: interpretability, predictability, and intervenability. However, we find that concept bottleneck models struggle to meet these goals. Using post hoc interpretability methods, we demonstrate that concepts do not correspond to anything semantically meaningful in input space, thus calling into question the usefulness of concept bottleneck models in their current form.
Forward citations
Cited by 5 Pith papers
-
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.
-
Locality-aware Concept Bottleneck Model
A label-free concept bottleneck model using per-concept prototypes aligned by CLIP to localize concept predictions to the correct image regions.
-
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.
-
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.
-
A Geometric Unification of Concept Learning with Concept Cones
CBMs and SAEs both learn nonnegative linear concept cones; a new containment metric suite scores SAE dictionaries against CBM concepts.
Discussion (0). Sign in to comment.