Pith. sign in

REVIEW 8 cited by

Promises and Pitfalls of Black-Box Concept Learning Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.13314 v1 pith:SVPP5BZ4 submitted 2021-06-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsconceptlearningblack-boxinformationabilitybeyondconcepts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning models that incorporate concept learning as an intermediate step in their decision making process can match the performance of black-box predictive models while retaining the ability to explain outcomes in human understandable terms. However, we demonstrate that the concept representations learned by these models encode information beyond the pre-defined concepts, and that natural mitigation strategies do not fully work, rendering the interpretation of the downstream prediction misleading. We describe the mechanism underlying the information leakage and suggest recourse for mitigating its effects.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 20 citations worldwide. Full citation record

  1. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.

  2. Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention

    cs.CV 2026-06 conditional novelty 7.0 of 10

    Part-factorized CBM with Gaussian spatial prior matches supervised 88.85% top-1 accuracy on CUB-200-2011 while raising pointing accuracy to 52.6% and works with 0.5% keypoint data or PCA foreground only.

  3. The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.

  4. Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Neural Concept Verifier trains image classifiers so predictions must rely on small, verifiable subsets of extracted concepts rather than raw pixel masks.

  5. Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.

  6. Interpretable Hierarchical Concept Reasoning through Attention-Guided Graph Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    H-CMR is a concept-based classifier whose concept and task predictions are made by attention-selected logic rules over a learned acyclic concept graph.

  7. ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

    cs.AI 2026-07 conditional novelty 4.0 of 10

    A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.

  8. Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs

    cs.LG 2025-07 reject novelty 4.0 of 10

    BAGEL trains per-layer logistic-regression probes on CLIP-defined concepts and compares per-class concept probabilities with dataset-level concept frequencies, visualizing the alignment in a knowledge graph.

Pith tools