Pith. sign in

REVIEW 35 references

Training-free Test-time Improvement for Explainable Medical Image Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.18070 v1 pith:WLCCZ7N4 submitted 2025-06-22 cs.CV

classification cs.CV
keywords conceptmedicalclassificationconceptsimagemodelsexplainableimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning-based medical image classification techniques are rapidly advancing in medical image analysis, making it crucial to develop accurate and trustworthy models that can be efficiently deployed across diverse clinical scenarios. Concept Bottleneck Models (CBMs), which first predict a set of explainable concepts from images and then perform classification based on these concepts, are increasingly being adopted for explainable medical image classification. However, the inherent explainability of CBMs introduces new challenges when deploying trained models to new environments. Variations in imaging protocols and staining methods may induce concept-level shifts, such as alterations in color distribution and scale. Furthermore, since CBM training requires explicit concept annotations, fine-tuning models solely with image-level labels could compromise concept prediction accuracy and faithfulness - a critical limitation given the high cost of acquiring expert-annotated concept labels in medical domains. To address these challenges, we propose a training-free confusion concept identification strategy. By leveraging minimal new data (e.g., 4 images per class) with only image-level labels, our approach enhances out-of-domain performance without sacrificing source domain accuracy through two key operations: masking misactivated confounding concepts and amplifying under-activated discriminative concepts. The efficacy of our method is validated on both skin and white blood cell images. Our code is available at: https://github.com/riverback/TF-TTI-XMed.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 12 canonical work pages

  1. [1]

    Data in brief30, 105474 (2020)

    Acevedo, A., Merino, A., Alférez, S., Molina, Á., Boldú, L., Rodellar, J.: A dataset of microscopic peripheral blood cell images for development of automatic recogni- tion systems. Data in brief30, 105474 (2020)

  2. [2]

    Nature Medicine28(9), 1773–1784 (2022)

    Acosta, J.N., Falcone, G.J., Rajpurkar, P., Topol, E.J.: Multimodal biomedical ai. Nature Medicine28(9), 1773–1784 (2022)

  3. [3]

    npj Digital Medicine8(1), 38 (2025)

    Balendran,A.,Beji,C.,Bouvier,F.,Khalifa,O.,Evgeniou,T.,Ravaud,P.,Porcher, R.: A scoping review of robustness concepts for machine learning in healthcare. npj Digital Medicine8(1), 38 (2025)

  4. [4]

    Nature biomedical engineering7(6), 719–742 (2023)

    Chen, R.J., Wang, J.J., Williamson, D.F., Chen, T.Y., Lipkova, J., Lu, M.Y., Sahai, S., Mahmood, F.: Algorithmic fairness in artificial intelligence for medicine and healthcare. Nature biomedical engineering7(6), 719–742 (2023)

  5. [5]

    In: The Thirteenth International Conference on Learning Representations (2025)

    Choi, J., Raghuram, J., Li, Y., Jha, S.: Adaptive concept bottleneck for foundation models under distribution shifts. In: The Thirteenth International Conference on Learning Representations (2025)

  6. [6]

    In: MICCAI (10)

    Chowdhury, T.F., Phan, V.M.H., Liao, K., To, M., Xie, Y., van den Hengel, A., Verjans, J.W., Liao, Z.: Adacbm: An adaptive concept bottleneck model for ex- plainable and accurate diagnosis. In: MICCAI (10). Lecture Notes in Computer Science, vol. 15010, pp. 35–45. Springer (2024)

  7. [7]

    Science advances 8(31), eabq6147 (2022)

    Daneshjou, R., Vodrahalli, K., Novoa, R.A., Jenkins, M., Liang, W., Rotemberg, V., Ko, J., Swetter, S.M., Bailey, E.E., Gevaert, O., et al.: Disparities in derma- tology ai performance on a diverse, curated clinical image set. Science advances 8(31), eabq6147 (2022)

  8. [8]

    In: NeurIPS (2022)

    Daneshjou,R.,Yüksekgönül,M.,Cai,Z.R.,Novoa,R.A.,Zou,J.Y.:Skincon:Askin disease dataset densely annotated by domain experts for fine-grained debugging and analysis. In: NeurIPS (2022)

Show all 35 references
  1. [9]

    He et al

    Fajtl, J., Welikala, R.A., Barman, S., Chambers, R., Bolter, L., Anderson, J., Olvera-Barrios, A., Shakespeare, R., Egan, C., Owen, C.G., et al.: Trustworthy evaluationofclinicalaiforanalysisofmedicalimagesindiversepopulations.NEJM AI1(9), AIoa2400353 (2024) 10 H. He et al

  2. [10]

    In: CVPR Workshops

    Groh, M., Harris, C., Soenksen, L., Lau, F., Han, R., Kim, A., Koochek, A., Badri, O.: Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset. In: CVPR Workshops. pp. 1820–1828. Computer Vision Foundation / IEEE (2021)

  3. [11]

    AI magazine40(2), 44–58 (2019)

    Gunning, D., Aha, D.: Darpa’s explainable artificial intelligence (xai) program. AI magazine40(2), 44–58 (2019)

  4. [12]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    He, H., Zhu, L., Zhang, X., Zeng, S., Chen, Q., Lu, Y.: V2c-cbm: Building con- cept bottlenecks with vision-to-concept tokenizer. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39(3), pp. 3401–3409 (2025)

  5. [13]

    In: CVPR

    He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. pp. 770–778. IEEE Computer Society (2016)

  6. [14]

    In: AAAI

    Huang, Q., Song, J., Hu, J., Zhang, H., Wang, Y., Song, M.: On the concept trustworthiness in concept bottleneck models. In: AAAI. pp. 21161–21168. AAAI Press (2024)

  7. [15]

    In: ISIC/Care- AI/MedAGI/DeCaF@MICCAI

    Kim, I., Kim, J., Choi, J., Kim, H.J.: Concept bottleneck with visual con- cept filtering for explainable medical image classification. In: ISIC/Care- AI/MedAGI/DeCaF@MICCAI. Lecture Notes in Computer Science, vol. 14393, pp. 225–233. Springer (2023)

  8. [16]

    arXiv preprint arXiv:1412.6980 (2014)

    Kingma, D.P.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)

  9. [17]

    In: ICML

    Koh, P.W., Nguyen, T., Tang, Y.S., Mussmann, S., Pierson, E., Kim, B., Liang, P.: Concept bottleneck models. In: ICML. Proceedings of Machine Learning Research, vol. 119, pp. 5338–5348. PMLR (2020)

  10. [18]

    Scientific reports12(1), 1123 (2022)

    Kouzehkanan, Z.M., Saghari, S., Tavakoli, S., Rostami, P., Abaszadeh, M., Mirzadeh, F., Satlsar, E.S., Gheidishahran, M., Gorgi, F., Mohammadi, S., et al.: A large dataset of white blood cells containing cell locations and types, along with segmented nuclei and cytoplasm. Scie...

  11. [19]

    Scientific Reports13(1), 2160 (2023)

    Li, M., Lin, C., Ge, P., Li, L., Song, S., Zhang, H., Lu, L., Liu, X., Zheng, F., Zhang, S., et al.: A deep learning model for detection of leukocytes under various interference factors. Scientific Reports13(1), 2160 (2023)

  12. [20]

    Liang, J., He, R., Tan, T.: A comprehensive survey on test-time adaptation under distribution shifts. Int. J. Comput. Vis.133(1), 31–64 (2025)

  13. [21]

    In: NeurIPS (2024)

    Mai, Z., Chowdhury, A., Zhang, P., Tu, C., Chen, H., Pahuja, V., Berger-Wolf, T.Y., Gao, S., Stewart, C.V., Su, Y., Chao, W.: Fine-tuning is fine, if calibrated. In: NeurIPS (2024)

  14. [22]

    McCloskey, M., Cohen, N.J.: Catastrophic interference in connectionist networks: Thesequentiallearningproblem.In:Psychologyoflearningandmotivation,vol.24, pp. 109–165. Elsevier (1989)

  15. [23]

    Na- ture616(7956), 259–265 (2023)

    Moor, M., Banerjee, O., Abad, Z.S.H., Krumholz, H.M., Leskovec, J., Topol, E.J., Rajpurkar, P.: Foundation models for generalist medical artificial intelligence. Na- ture616(7956), 259–265 (2023)

  16. [24]

    In: ICLR

    Oikarinen, T.P., Das, S., Nguyen, L.M., Weng, T.: Label-free concept bottleneck models. In: ICLR. OpenReview.net (2023)

  17. [25]

    Pang, W., Ke, X., Tsutsui, S., Wen, B.: Integrating clinical knowledge into concept bottleneckmodels.In:MICCAI(4).LectureNotesinComputerScience,vol.15004, pp. 243–253. Springer (2024)

  18. [26]

    NEJM AI2(1), AIp2400672 (2025) Training-free Test-time Improvement for Medical CBMs 11

    Sabuncu, M.R., Wang, A.Q., Nguyen, M.: Ethical use of artificial intelligence in medical diagnostics demands a focus on accuracy, not fairness. NEJM AI2(1), AIp2400672 (2025) Training-free Test-time Improvement for Medical CBMs 11

  19. [27]

    Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad- cam: Visual explanations from deep networks via gradient-based localization. Int. J. Comput. Vis.128(2), 336–359 (2020)

  20. [28]

    In: ICML

    Shin, S., Jo, Y., Ahn, S., Lee, N.: A closer look at the intervention procedure of concept bottleneck models. In: ICML. Proceedings of Machine Learning Research, vol. 202, pp. 31504–31520. PMLR (2023)

  21. [29]

    arXiv preprint arXiv:1409.1556 (2014)

    Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)

  22. [30]

    In: ICML

    Steinmann, D., Stammer, W., Friedrich, F., Kersting, K.: Learning to intervene on concept bottlenecks. In: ICML. OpenReview.net (2024)

  23. [31]

    In: ICASSP

    Tsutsui, S., Su, Z., Wen, B.: Benchmarking white blood cell classification under domain shift. In: ICASSP. pp. 1–5. IEEE (2023)

  24. [32]

    Nature medicine30(1), 1257–1268 (2024)

    Varghese, C., Harrison, E.M., O’Grady, G., Topol, E.J.: Artificial intelligence in surgery. Nature medicine30(1), 1257–1268 (2024)

  25. [33]

    In: NeurIPS (2024)

    Yang, Y., Gandhi, M., Wang, Y., Wu, Y., Yao, M.S., Callison-Burch, C., Gee, J.C., Yatskar, M.: A textbook remedy for domain shifts: Knowledge priors for medical image analysis. In: NeurIPS (2024)

  26. [34]

    In: CVPR

    Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., Yatskar, M.: Language in a bottle: Language model guided concept bottlenecks for interpretable image classification. In: CVPR. pp. 19187–19197. IEEE (2023)

  27. [35]

    Nature medicine31(1), 45–59 (2025)

    Zhang, K., Yang, X., Wang, Y., Yu, Y., Huang, N., Li, G., Li, X., Wu, J.C., Yang, S.: Artificial intelligence in drug development. Nature medicine31(1), 45–59 (2025)

Pith tools