REVIEW 13 cited by
Label-Free Concept Bottleneck Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Concept bottleneck models (CBM) are a popular way of creating more interpretable neural networks by having hidden layer neurons correspond to human-understandable concepts. However, existing CBMs and their variants have two crucial limitations: first, they need to collect labeled data for each of the predefined concepts, which is time consuming and labor intensive; second, the accuracy of a CBM is often significantly lower than that of a standard neural network, especially on more complex datasets. This poor performance creates a barrier for adopting CBMs in practical real world applications. Motivated by these challenges, we propose Label-free CBM which is a novel framework to transform any neural network into an interpretable CBM without labeled concept data, while retaining a high accuracy. Our Label-free CBM has many advantages, it is: scalable - we present the first CBM scaled to ImageNet, efficient - creating a CBM takes only a few hours even for very large datasets, and automated - training it for a new dataset requires minimal human effort. Our code is available at https://github.com/Trustworthy-ML-Lab/Label-free-CBM. Finally, in Appendix B we conduct a large scale user evaluation of the interpretability of our method.
Forward citations
Cited by 13 Pith papers
-
When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate
MAGNETS learns unsupervised, mask-based concepts to make time-series regression predictions additively interpretable, recovering ground-truth temporal rules on synthetic tasks and beating interpretable baselines on mo...
-
Explainable Novel Category Discovery in Semantic Concept Space
xNCD routes novel category discovery through a CLIP-aligned concept bottleneck, matching strong NCD baselines while producing intrinsic cluster- and instance-level concept explanations.
-
TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
TimeSAE trains a sparse autoencoder with counterfactual and consistency losses to explain black-box time series predictions, claiming better faithfulness and out-of-distribution robustness than eight baselines.
-
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
A 'tool bottleneck' framework—VLM tool selection plus learned spatial fusion—matches or beats black-box classifiers, especially on scarce data.
-
Exploring the Rashomon Set for Concept-Based Models
A shared frozen backbone plus per-model LoRA adapters and a concept-diversity loss trains a set of accurate CBMs that reason through different concepts.
-
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
Neural Concept Verifier trains image classifiers so predictions must rely on small, verifiable subsets of extracted concepts rather than raw pixel masks.
-
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
DiSciPLE uses LLM-guided evolution to discover interpretable Python programs that predict geospatial quantities, outperforming black-box deep nets on population density and on out-of-distribution generalization.
-
Enhancing Performance of Explainable AI Models with Constrained Concept Refinement
Constrained Concept Refinement slightly adjusts concept embeddings under a small-radius constraint, improving accuracy of explainable classifiers and cutting training time by about 10x on large image benchmarks.
-
Survival Concept-Based Learning Models
SurvCBM and SurvRCM combine concept bottleneck learning with Cox and Beran survival models, and SurvCBM achieves the best C-index and concept F1 on synthetic MNIST and CIFAR experiments.
-
Interpretable Failure Detection with Human-Level Concepts
ORCA ranks concept activations from CLIP and uses the rank-weighted agreement of the top-K concepts with the predicted category as its confidence score, improving failure-detection FPR on several benchmarks.
-
Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
Class-name-free accuracy of LLM descriptors is 15.5% versus 59.5% with class names; selecting attributes from target images raises attribute-only accuracy to 23.8% (45.5% uncapped).
-
A Concept-based approach to Voice Disorder Detection
Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.
-
Towards Interpretable PolSAR Image Classification: Polarimetric Scattering Mechanism Informed Concept Bottleneck and Kolmogorov-Arnold Network
A concept bottleneck model built from polarimetric target decomposition plus a Kolmogorov-Arnold Network gives PolSAR classification with human-auditable concept predictions and symbolic decision formulas.
Discussion (0). Continue with ORCID to comment.