Pith. sign in

REVIEW 3 major objections 4 minor 57 references

MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A lung nodule AI needs only 1% of labels to match fully supervised accuracy.

desk verdict Solid within-protocol annotation-efficiency result; the exceeding-SOTA headline rests on cross-protocol baselines that were never re-run on MERA's own split. read the letter →

arxiv 2504.19357 v1 pith:ZYSAHUVJ submitted 2025-04-27 cs.CV

classification cs.CV
keywords explainableartificialintelligencelungnodulediagnosisself-supervisedlearningVisionTransformersemi-supervisedactiveLIDCdatasetreducedannotation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that a lung nodule malignancy classifier can be built almost without labels: using only 1% of annotated training samples on the LIDC protocol, MERA reports malignancy accuracy of 86.22 ± 2.51%, close to the 87.56 ± 0.61% it reaches with full annotation and competitive with fully supervised methods. The same model predicts clinically defined nodule attributes at 91–96% accuracy and emits four kinds of explanations—attention maps, nearest-case retrieval, latent-space clustering, and attribute-based concept explanations—most of which need no labels at all. The motivation is practical: expert annotation of CT nodules is scarce and costly, and trust in diagnostic AI depends on explanations that align with radiological reasoning.

What carries the argument

The load-bearing component is the two-stage training schedule. Stage 1 uses DINO self-supervised contrastive learning on a ViT-Small encoder (starting from ImageNet-pretrained DINO weights) to map 32 × 32 axial CT patches into a semantically organised latent space. Stage 2 trains one linear predictor per nodule attribute plus a malignancy predictor on the concatenation of image features and predicted attributes; the annotation exploitation mechanism selects seeds by k-means clustering, requests labels for low-confidence samples while pseudo-labelling high-confidence ones, and periodically reinitialises the predictors ('quenching') to curb confirmation bias. The argument is that Stage 1 supplies most of the representation, so Stage 2 needs only a handful of labels.

What would settle it

Train the Stage 1 encoder from scratch on unlabelled LIDC CT patches with no ImageNet weights and measure malignancy accuracy at 1% annotation; if accuracy falls toward the 79.19% level reported in Table 3 rather than staying near 86%, the 1%-annotation result is an artefact of pretrained representation transfer, not of self-supervised learning on the target domain. Re-running the literature baselines on MERA's exact 730-nodule split would similarly test whether the performance advantage survives equal footing.

Watch

Extended reading notes

Core claim

MERA's central claim is that a self-supervised Vision Transformer, trained with DINO-style contrastive learning on unlabelled nodule patches, creates a latent space with enough semantic separability that a linear predictor trained on a tiny labelled seed can classify malignancy and nodule attributes nearly as well as a fully supervised model. The predictor is trained with sparse seeding via clustering, then refined by semi-supervised active learning with dynamic pseudo-labels and periodic reinitialisation ('quenching'). On the paper's 730-nodule LIDC split, 1% annotation yields 86.22 ± 2.51% malignancy accuracy and 91–96% per-attribute accuracy, versus 87.56 ± 0.61% with full labels. Explanations are generated intrinsically: t-SNE clustering shows malignancy-correlated groupings, k-nearest neighbours supply case-based rationale, averaged self-attention maps localise diagnostically relevant features, and predicted nodule attributes feed the malignancy decision.

Load-bearing premise

The central claim that 1% annotations suffice rests on initializing the encoder with ImageNet self-supervised weights, since removing that initialization drops two-stage malignancy accuracy from 87.56% to 79.19% in Table 3.

Editorial extensions

If this is right

  • If the finding holds, a diagnostic AI for lung nodules can be trained with about five labelled nodules on a 518-nodule training set, cutting annotation cost by roughly two orders of magnitude.
  • The four explanation channels (attention maps, nearest cases, cluster structure, attribute concepts) are generated without additional supervision, except for attribute prediction, so explainability does not require a large labelled dataset.
  • Simultaneous high accuracy on all nodule attributes means the concept explanations are not just rhetorical; they can serve as a checkable intermediate output before the malignancy verdict.
  • The reported robustness under annotation reduction—comparable accuracy at 1%, 10%, and 100% labels—suggests the model keeps working as expert labels become scarce.
  • The method's reliance on hundreds of unlabelled CT patches indicates that unlabelled data curation, not annotation volume, becomes the main practical requirement for deploying it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: because removing the ImageNet DINO initialisation drops two-stage ViT malignancy accuracy from 87.56% to 79.19% (Table 3), the paper supports 'a good latent representation plus a few labels', not yet 'unlabelled lung CT alone suffices'. A definitive test would pretrain on unlabelled thoracic CT without ImageNet and re-measure the 1% accuracy.
  • Editorial inference: the 'exceeding state-of-the-art' comparisons use literature numbers from different data protocols (1,149–4,252 nodules, 3D volumes, extra supervision). Re-running those baselines on MERA's exact split would settle whether the advantage is real or protocol-driven.
  • Editorial inference: the same two-stage recipe—self-supervised encoder, clustered seed selection, and dynamic pseudo-labelling with quenching—could transfer to other lesion-classification tasks with scarce labels, provided a similar semantic attribute set exists; that transfer is not tested in the paper.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes MERA, a two-stage pipeline for lung nodule malignancy diagnosis that combines self-supervised contrastive pretraining of a Vision Transformer with a weakly supervised hierarchical predictor over radiologist-defined nodule attributes. The authors claim that with only 1% of annotated samples MERA reaches diagnostic accuracy comparable to or exceeding that of state-of-the-art methods that use full annotation, while also providing model-level, instance-level, local visual, and concept-level explanations. Experiments on the LIDC dataset report 86.22±2.51% malignancy accuracy with 1% annotations versus 87.56±0.61% with full annotations under the same protocol, together with per-attribute accuracies above 90% and a series of qualitative explanation case studies.

Significance. The core same-protocol result, 1% versus 100% annotation (86.22±2.51% versus 87.56±0.61%, Table 2), is internally consistent and represents a genuinely valuable empirical contribution: it shows that sparse seeding followed by dynamic pseudo-labelling with quenching can stabilize training in a very-low-label regime. The paper also ships its code, reports standard deviations, and includes ablations of the seeding, acquisition, pseudo-labelling, and quenching components. If the central comparison to prior work were placed on equal footing, the proposed pipeline would be a meaningful step toward transparent, low-annotation medical image diagnosis. The positive assessment is conditional, however, because the headline 'exceeding state-of-the-art' claim currently rests on published numbers from substantially different data protocols that are never re-run on the authors' split.

major comments (3)
  1. [Abstract, Sec. 4.1.1, Tables 1-2] The central claim that MERA with 1% annotations 'exceeds state-of-the-art methods requiring full annotation' is not supported by the reported experiments. All five cited baselines (HSCNN, X-Caps, MSN-JCN, MTMR, WeakSup) are taken from published numbers obtained on different protocols (1149-4252 nodules, 3D volumes or multiple 2D slices, and in some cases additional supervision such as segmentation masks or diameter information), a limitation the manuscript itself acknowledges in Sec. 4.1.1. The only full-annotation method re-run on MERA's own 730-nodule, 70/30 nodule-level split is an end-to-end ResNet-50, which reaches 88.08% (Table 3), above both MERA's full-annotation 87.56% and its 1%-annotation 86.22±2.51%. The 'exceeding' part of the headline is therefore not testable from the reported data; the authors should either re-run the baselines on their split or explicitly downgrade the claim to 'comparable' with a clear protocol caveat.
  2. [Sec. 4.1.2, Table 3] The 1%-annotation result is conditional on initializing the ViT with ImageNet-pretrained self-supervised DINO weights. Table 3 shows that removing this initialization drops ViT two-stage malignancy accuracy from 87.56% to 79.19%, so the reported absolute performance largely inherits ImageNet-learned representations rather than demonstrating that unlabelled lung CT alone plus 1% labels suffices. This does not invalidate the annotation-efficiency comparison between 1% and 100% labels, because both use the same initialization, but the abstract and Sec. 5 should state this dependency explicitly when describing the method as 'primarily unsupervised.'
  3. [Sec. 3.1.2, Table 4 footnote] The manuscript presents uncertainty-sampling active learning as a core component of the annotation exploitation mechanism, but the footnote to Table 4 states that the 1% column 'Does not contain requested annotations.' This means the headline 1%-annotation result is obtained without active learning, using only sparse seeding, pseudo-labelling, and quenching. Please state this explicitly in the method description and abstract, or clarify how the annotation budget is counted if actively requested labels are in fact used at the 1% setting.
minor comments (4)
  1. [Table 4] In the 1% column, the rows 'sparse integrated entropy' and 'sparse malignancy confidence' report exactly the same mean and standard deviation (86.22±2.51). Please clarify whether this is a typo or a genuine coincidence under the no-requested-annotations setting.
  2. [Sec. 4.3, Table 1] For the k-NN case-based explanation results under partial annotation, the paper should specify how unlabeled training samples are treated when selecting the k nearest neighbors and when assigning the majority label, since with only 10% of training labels the nearest neighbors may often lack ground-truth annotations.
  3. [Fig. 11] The annotations 'Annotaion 10x' and 'Accuracy 0.43%' are ambiguous, and the figure would benefit from stating explicitly which baseline is used and how its annotation reduction is performed.
  4. [Throughout] There are numerous rendering artifacts and typos, such as 'di fferent' and 'o ffer' instead of 'different' and 'offer', and 'inference phrase' in Sec. 3 instead of 'inference phase'; these should be corrected during copy-editing.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the headline 1%-annotation accuracy is measured on a held-out nodule-level test split, and the attribute-to-malignancy composition in Eq. (5) is a genuine intermediate prediction rather than a restatement of the inputs.

full rationale

I find no equation-level circularity. Stage 1 learns features with self-supervised DINO on unlabeled data, and Stage 2 trains linear attribute and malignancy predictors on sparse seeds plus pseudo-labels; all reported accuracies in Tables 1-4 and Figures 9-11 are evaluated on the held-out 30% test split defined in Sec. 4.1.1. Equation (5) composes predicted nodule attributes into the malignancy predictor, so the concept explanations are genuine intermediate outputs, not a re-encoding of the malignancy label. The only near-circular component is the self-training loop in Eq. (4), where the model's own confident predictions on unlabeled training samples become pseudo-labels; however, the paper discloses this and introduces quenching to counter confirmation bias, and the pseudo-labeled set is the training set, not the test set, so the reported test accuracy is not forced by construction. The self-citations [56,57] set an evaluation tolerance for attribute accuracy and are not load-bearing for the central claim. The cross-protocol comparison to HSCNN, X-Caps, MSN-JCN, MTMR, and WeakSup is a benchmark-comparability risk rather than a circularity, and the dependence on ImageNet-pretrained DINO weights is disclosed in Sec. 4.1.2 and ablated in Table 3.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on an ImageNet-pretrained DINO ViT encoder (external, non-medical knowledge), a published but restrictive LIDC preprocessing protocol with single 2D-slice patches, and a self-training loop whose pseudo-label quality is never measured. Stage-2 hyperparameters (seed count, quenching interval) are hand-chosen. No invented entities are introduced.

free parameters (3)
  • Number of seed clusters n (1% annotation budget) = about 5 (1% of 518 training nodules)
    Stage-2 k-means seeding (Sec. 3.1.2): these about 5 labeled nodules carry the entire supervised signal at the headline setting; Tab. 4 shows random seeding collapses 1% accuracy to 80.02 +/- 8.56, so this choice is performance-critical.
  • Quenching schedule = first quench at 100 seed epochs, then every 10 epochs for 50 more epochs
    Hand-chosen training schedule for the annotation exploitation mechanism (Sec. 4.1.2); no sensitivity analysis is reported, and it interacts with the pseudo-label loop.
  • DINO hyperparameters = tau_pri = 0.04, tau_aux = 0.1, momentum m from 0.996 to 1, K = 65536
    Adopted without tuning from DINO [14]; listed because Stage-1 feature quality determines the 1%-annotation results.
assumptions (4)
  • domain assumption ImageNet self-supervised DINO weights are used to initialize the encoder, and their representation is the substrate for the 1%-annotation predictors.
    Sec. 4.1.2: 'starting from the weights pretrained unsupervisedly on ImageNet'. Tab. 3 shows removing this init drops ViT two-stage accuracy from 87.56% to 79.19%, so 'unsupervised' inherits ImageNet knowledge.
  • domain assumption The LIDC preprocessing protocol of [51] is valid for comparing with prior work.
    Sec. 4.1.1: nodules with fewer than 3 radiologist ratings or slice thickness over 2.5 mm are discarded, median aggregation, malignancy binarized at 3, one 2D central-slice patch per nodule. This yields 730 nodules, a protocol different from all methods MERA compares against.
  • domain assumption A single 2D central axial slice carries enough information to predict malignancy and the eight nodule attributes.
    The model never sees full 3D nodule volumes; methods using 3D volumes (HSCNN, MTMR, WeakSup) report the highest malignancy numbers (89 to 93.5).
  • ad hoc to paper The model's own confident predictions can be used as ground-truth labels for continued training (pseudo-labeling).
    Eq. (4) trains on pseudo-labels with periodic resets; pseudo-label accuracy is never measured, yet the 91 to 96% per-attribute accuracies at 1% annotation flow through this loop.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis." pith.science (2026). https://pith.science/paper/ZYSAHUVJ

@misc{pith2026250419357,
  author       = {Pith},
  title        = {Pith review of: MERA: Multimodal and Multiscale Self-Explanatory Model with Considerably Reduced Annotation for Lung Nodule Diagnosis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZYSAHUVJ}},
  note         = {Machine review of arXiv:2504.19357}
}
read the original abstract

Lung cancer, a leading cause of cancer-related deaths globally, emphasises the importance of early detection for better patient outcomes. Pulmonary nodules, often early indicators of lung cancer, necessitate accurate, timely diagnosis. Despite Explainable Artificial Intelligence (XAI) advances, many existing systems struggle providing clear, comprehensive explanations, especially with limited labelled data. This study introduces MERA, a Multimodal and Multiscale self-Explanatory model designed for lung nodule diagnosis with considerably Reduced Annotation requirements. MERA integrates unsupervised and weakly supervised learning strategies (self-supervised learning techniques and Vision Transformer architecture for unsupervised feature extraction) and a hierarchical prediction mechanism leveraging sparse annotations via semi-supervised active learning in the learned latent space. MERA explains its decisions on multiple levels: model-level global explanations via semantic latent space clustering, instance-level case-based explanations showing similar instances, local visual explanations via attention maps, and concept explanations using critical nodule attributes. Evaluations on the public LIDC dataset show MERA's superior diagnostic accuracy and self-explainability. With only 1% annotated samples, MERA achieves diagnostic accuracy comparable to or exceeding state-of-the-art methods requiring full annotation. The model's inherent design delivers comprehensive, robust, multilevel explanations aligned closely with clinical practice, enhancing trustworthiness and transparency. Demonstrated viability of unsupervised and weakly supervised learning lowers the barrier to deploying diagnostic AI in broader medical domains. Our complete code is open-source available: https://github.com/diku-dk/credanno.

Figures

Figures reproduced from arXiv: 2504.19357 by the authors.

Figure 1
Figure 1. Multimodal and multiscale explanations as an intrinsic driving force of the decision: local visual explanations via attention maps, model-level global explanations through semantic latent space clustering, instance-level case-based explanations providing similar instances, and concept explanations based on critical nodule attributes. existing XAI systems struggle to provide clear, compre￾hensive multi-level explanat… view at source ↗
Figure 2
Figure 2. Method overview of the data-/annotation-efficient train￾ing. In Stage 1, an encoder is trained using self-supervised contrastive learning to map the input nodule images to a semantically meaningful latent space. In Stage 2, the proposed annotation exploitation mech￾anism conducts semi-supervised active learning with sparse seeding and training quenching in the learned space, to jointly exploit the ex￾tracted feature… view at source ↗
Figure 3
Figure 3. t-SNE visualisation of features extracted from testing images. Data points are coloured using ground truth annotations. Malignancy shows highly separable in the learned space, and semantically correlates with the clustering in each nodule attribute. ably separable in both malignancy and nodule attributes. This provides the possibility to train the initial predic￾tors using only a very small number of seed annotation… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Case-based explanation of a malignant nodule. Nearest training and testing samples illustrating similarities in poorly defined margins, lobulation, and spiculation, which are typical morphological characteristics suggesting malignancy. (a) CT image and visual ex￾planat…
Figure 5
Figure 5. Figure 5: Case-based explanation of a benign nodule.. Nearest training and testing samples display the same solid texture, smooth and sharp margins, and lack of lobulations or spiculations, mirroring the features of the target nodule. 9 [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Local visual explanation of benign nodules. From left to right: original lung nodule image patch on an axial chest CT, visual explanation of 4 competitor methods, visual explanation of our proposed method, our predicted nodule attributes and malignancy (green font indi…
Figure 7
Figure 7. Figure 7: Local visual explanation of malignant nodules. From left to right: original lung nodule image patch on an axial chest CT, visual explanation of 4 competitor methods, visual explanation of our proposed method, our predicted nodule attributes and malignancy (green font i…
Figure 8
Figure 8. Figure 8: Local visual explanation of incorrect predictions. From left to right: original lung nodule image patch on an axial chest CT, visual explanation of 4 competitor methods, visual explanation of our proposed method, our predicted nodule attributes and malignancy (green fo…
Figure 9
Figure 9. Figure 9: Performance comparison, in terms of prediction accu￾racy (%) of nodule attributes and malignancy. Observe that MERA achieves simultaneously high accuracy in predicting malignancy and all nodule attributes, regardless of using either full or partial annota￾tions. 0 1 2 …
Figure 10
Figure 10. Figure 10: Probabilities of correctly predicting a certain number of attributes for a given nodule sample. Observe that MERA shows a more prominent probability of simultaneously predicting all 8 nod￾ule attributes correctly. show that when using only 518 among the 730 nod￾ule sa…
Figure 11
Figure 11. Figure 11: Influence of annotation reduction on MERA, MERA without annotation exploitation mechanism (AEM) and a representa￾tive CNN baseline , in terms of nodule malignancy prediction accu￾racy. With the proposed AEM, MERA achieves comparable or even higher accuracy with 10x fe…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 33 canonical work pages

  1. [1]

    del Ciello, P

    A. del Ciello, P. Franchi, A. Contegiacomo, G. Cicchetti, L. Bonomo, A. R. Larici, Missed lung cancer: When, where, and why?, Diagnostic and Interventional Radiology 23 (2) (2017) 118–126. doi:10.5152/dir.2016.16187

  2. [2]

    Vlahos, K

    I. Vlahos, K. Stefanidis, S. Sheard, A. Nair, C. Sayer, J. Moser, Lung cancer screening: Nodule identification and characteriza- tion, Translational Lung Cancer Research 7 (3) (2018) 288–303. doi:10.21037/tlcr.2018.05.02

  3. [3]

    P. J. Mazzone, L. Lam, Evaluating the Patient With a Pulmonary Nodule: A Review, JAMA 327 (3) (2022) 264. doi:10.1001/ jama.2021.24287

  4. [4]

    S. Shen, S. X. Han, D. R. Aberle, A. A. Bui, W. Hsu, An in- terpretable deep hierarchical semantic convolutional neural net- work for lung nodule malignancy classification, Expert Systems with Applications 128 (2019) 84–95. doi:10.1016/j.eswa. 2019.01.048

  5. [5]

    LaLonde, D

    R. LaLonde, D. Torigian, U. Bagci, Encoding Visual Attributes in Capsules for Explainable Medical Diagnoses, in: Medical Image Computing and Computer Assisted Intervention – MIC- CAI 2020, Lecture Notes in Computer Science, Springer Inter- national Publishing, Cham, 2020, pp. 294–304. doi:10.1007/ 978-3-030-59710-8_29

  6. [6]

    W. Chen, Q. Wang, D. Yang, X. Zhang, C. Liu, Y . Li, End-to- End Multi-Task Learning for Lung Nodule Segmentation and Diagnosis, in: 2020 25th International Conference on Pattern Recognition (ICPR), IEEE, Milan, Italy, 2021, pp. 6710–6717. doi:10.1109/ICPR48806.2021.9412218

  7. [7]

    L. Liu, Q. Dou, H. Chen, J. Qin, P.-A. Heng, Multi-Task Deep Model With Margin Ranking Loss for Lung Nodule Analysis, IEEE Transactions on Medical Imaging 39 (3) (2020) 718–728. doi:10.1109/TMI.2019.2934577

  8. [8]

    Joshi, J

    A. Joshi, J. Sivaswamy, G. D. Joshi, Lung nodule malignancy classification with weakly supervised explanation generation, Journal of Medical Imaging 8 (04) (Aug. 2021).doi:10.1117/ 1.JMI.8.4.044502

Show all 57 references
  1. [9]

    Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature Machine Intelligence 1 (5) (2019) 206–215

    C. Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature Machine Intelligence 1 (5) (2019) 206–215. doi:10. 1038/s42256-019-0048-x

  2. [10]

    B. H. van der Velden, H. J. Kuijf, K. G. Gilhuijs, M. A. Viergever, Explainable artificial intelligence (XAI) in deep learning-based medical image analysis, Medical Image Analysis 79 (2022) 102470. doi:10.1016/j.media.2022.102470

  3. [11]

    Salahuddin, H

    Z. Salahuddin, H. C. Woodru ff, A. Chatterjee, P. Lambin, Trans- parency of deep neural networks for medical image analysis: A review of interpretability methods, Computers in Biology and Medicine 140 (2022) 105111. doi:10.1016/j.compbiomed. 2021.105111

  4. [12]

    MacMahon, D

    H. MacMahon, D. P. Naidich, J. M. Goo, K. S. Lee, A. N. C. Leung, J. R. Mayo, A. C. Mehta, Y . Ohno, C. A. Powell, M. Prokop, G. D. Rubin, C. M. Schaefer-Prokop, W. D. Travis, P. E. Van Schil, A. A. Bankier, Guidelines for Management of Incidental Pulmonary Nodules Detected on...

  5. [13]

    C. Chen, O. Li, D. Tao, A. Barnett, C. Rudin, J. K. Su, This Looks Like That: Deep Learning for Interpretable Image Recog- nition, in: Advances in Neural Information Processing Systems, V ol. 32, Curran Associates, Inc., 2019

  6. [14]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bo- janowski, A. Joulin, Emerging Properties in Self-Supervised Vi- sion Transformers, in: Proceedings of the IEEE /CVF Interna- tional Conference on Computer Vision, 2021, pp. 9650–9660

  7. [15]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, in: Inter- national Conference on Learning Re...

  8. [16]

    K. Wang, D. Zhang, Y . Li, R. Zhang, L. Lin, Cost-Effective Ac- tive Learning for Deep Image Classification, IEEE Transactions on Circuits and Systems for Video Technology 27 (12) (2017) 2591–2600. doi:10.1109/TCSVT.2016.2589879

  9. [17]

    Liang, M

    H. Liang, M. Hu, Y . Ma, L. Yang, J. Chen, L. Lou, C. Chen, Y . Xiao, Performance of Deep-Learning Solutions on Lung Nodule Malignancy Classification: A Systematic Review, Life 13 (9) (2023) 1911. doi:10.3390/life13091911

  10. [18]

    M. A. Balcı, L. M. Batrancea, ¨O. Akg ¨uller, A. Nichita, A Series-Based Deep Learning Approach to Lung Nodule Im- age Classification, Cancers 15 (3) (2023) 843. doi:10.3390/ cancers15030843

  11. [19]

    S. G. Armato, G. McLennan, L. Bidaut, M. F. McNitt-Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Hen- schke, E. A. Hoffman, E. A. Kazerooni, H. MacMahon, E. J. R. van Beek, D. Yankelevitz, A. M. Biancardi, P. H. Bland, M. S. Brown, R. M. Engelmann, G. E. Ladera...

  12. [20]

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, A. Torralba, Learn- ing Deep Features for Discriminative Localization, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Las Vegas, NV , USA, 2016, pp. 2921–2929. doi:10.1109/CVPR.2016.319

  13. [21]

    H. Wang, Z. Wang, M. Du, F. Yang, Z. Zhang, S. Ding, P. Mardziel, X. Hu, Score-CAM: Score-Weighted Visual Expla- nations for Convolutional Neural Networks, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Work- shops (CVPRW), IEEE, Seattle, W A, USA, 202...

  14. [22]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization, International Journal of Com- puter Vision (Oct. 2019).arXiv:1610.02391, doi:10.1007/ s11263-019-01228-7

  15. [23]

    Chattopadhay, A

    A. Chattopadhay, A. Sarkar, P. Howlader, V . N. Balasubra- manian, Grad-CAM ++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks, in: 2018 IEEE Winter Conference on Applications of Computer Vision 19 (W ACV), 2018, pp. 839–847. doi:10.1109/WACV.20...

  16. [24]

    Omeiza, S

    D. Omeiza, S. Speakman, C. Cintas, K. Weldermariam, Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models (Aug. 2019). arXiv:1908.01224, doi:10.48550/arXiv. 1908.01224

  17. [25]

    R. Fu, Q. Hu, X. Dong, Y . Guo, Y . Gao, B. Li, Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs (Aug. 2020). arXiv:2008.02312, doi:10.48550/ arXiv.2008.02312

  18. [26]

    S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M ¨uller, W. Samek, On Pixel-Wise Explanations for Non-Linear Clas- sifier Decisions by Layer-Wise Relevance Propagation, PLOS ONE 10 (7) (2015) e0130140.doi:10.1371/journal.pone. 0130140

  19. [27]

    B ¨ohle, F

    M. B ¨ohle, F. Eitel, M. Weygandt, K. Ritter, Layer-Wise Rele- vance Propagation for Explaining Deep Neural Network Deci- sions in MRI-Based Alzheimer’s Disease Classification, Fron- tiers in Aging Neuroscience 11 (2019). doi:10.3389/fnagi. 2019.00194

  20. [28]

    P. W. Koh, T. Nguyen, Y . S. Tang, S. Mussmann, E. Pierson, B. Kim, P. Liang, Concept Bottleneck Models, in: Proceed- ings of the 37th International Conference on Machine Learning, PMLR, 2020, pp. 5338–5348

  21. [30]

    Mohammadjafari, M

    S. Mohammadjafari, M. Cevik, M. Thanabalasingam, A. Basar, A. D. N. Initiative, Using ProtoPNet for Interpretable Alzheimer’s Disease Classification, Proceedings of the Cana- dian Conference on Artificial Intelligence (Jun. 2021). doi: 10.21428/594757db.fb59ce6c

  22. [31]

    E. Kim, S. Kim, M. Seo, S. Yoon, XProtoNet: Diagnosis in Chest Radiography With Global and Local Explanations, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15719–15728

  23. [32]

    Singh, K.-C

    G. Singh, K.-C. Yow, These do not Look Like Those: An In- terpretable Deep Learning Model for Image Recognition, IEEE Access 9 (2021) 41482–41493. doi:10.1109/ACCESS.2021. 3064838

  24. [33]

    Ho ffmann, C

    A. Ho ffmann, C. Fanconi, R. Rade, J. Kohler, This Looks Like That... Does it? Shortcomings of Latent Space Prototype Inter- pretability in Deep Networks (Jun. 2021).arXiv:2105.02968, doi:10.48550/arXiv.2105.02968

  25. [34]

    K. He, H. Fan, Y . Wu, S. Xie, R. Girshick, Momentum Con- trast for Unsupervised Visual Representation Learning, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), IEEE, Seattle, W A, USA, 2020, pp. 9726–9735. doi:10.1109/CVPR42600.2020.00975

  26. [35]

    X. Chen, K. He, Exploring Simple Siamese Representation Learning, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Nashville, TN, USA, 2021, pp. 15745–15753. doi:10.1109/CVPR46437.2021. 01549

  27. [36]

    Grill, F

    J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Ghesh- laghi Azar, B. Piot, k. kavukcuoglu, R. Munos, M. Valko, Boot- strap Your Own Latent - A New Approach to Self-Supervised Learning, in: Advances in Neural ...

  28. [37]

    T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A Simple Frame- work for Contrastive Learning of Visual Representations, in: Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, pp. 1597–1607

  29. [38]

    Sowrirajan, J

    H. Sowrirajan, J. Yang, A. Y . Ng, P. Rajpurkar, MoCo Pretrain- ing Improves Representation and Transferability of Chest X-ray Models, in: Medical Imaging with Deep Learning, 2021

  30. [39]

    Y . N. T. Vu, R. Wang, N. Balachandar, C. Liu, A. Y . Ng, P. Ra- jpurkar, MedAug: Contrastive learning leveraging patient meta- data improves representations for chest X-ray interpretation, in: Proceedings of the 6th Machine Learning for Healthcare Con- ference, PMLR, 2021, pp...

  31. [40]

    Russakovsky, J

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, L. Fei-Fei, ImageNet Large Scale Visual Recognition Chal- lenge, International Journal of Computer Vision 115 (3) (2015) 211–252. doi:10.1007/s11263-015-0816-y

  32. [41]

    X. Chen, S. Xie, K. He, An Empirical Study of Training Self- Supervised Vision Transformers, in: 2021 IEEE /CVF Interna- tional Conference on Computer Vision (ICCV), IEEE, Mon- treal, QC, Canada, 2021, pp. 9620–9629. doi:10.1109/ ICCV48922.2021.00950

  33. [42]

    Caron, I

    M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, A. Joulin, Unsupervised learning of visual features by contrasting cluster assignments, in: Advances in Neural Information Processing Systems, V ol. 33, Curran Associates, Inc., 2020, pp. 9912–9924

  34. [43]

    Touvron, M

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, H. Jegou, Training data-efficient image transformers & distilla- tion through attention, in: Proceedings of the 38th International Conference on Machine Learning, PMLR, 2021, pp. 10347– 10357

  35. [44]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is All you Need, in: Advances in Neural Information Processing Systems, V ol. 30, Curran Associates, Inc., 2017

  36. [45]

    Salimans, D

    T. Salimans, D. P. Kingma, Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Net- works, in: Advances in Neural Information Processing Systems, V ol. 29, Curran Associates, Inc., 2016

  37. [46]

    R. Hu, B. M. Namee, S. J. Delany, O ff to a Good Start: Using Clustering to Select the Initial Training Set in Active Learning, in: Twenty-Third International FLAIRS Conference, 2010

  38. [47]

    Settles, Uncertainty Sampling, Springer International Publishing, Cham, 2012, pp

    B. Settles, Uncertainty Sampling, Springer International Publishing, Cham, 2012, pp. 11–20. doi:10.1007/ 978-3-031-01560-1_2

  39. [48]

    Cascante-Bonilla, F

    P. Cascante-Bonilla, F. Tan, Y . Qi, V . Ordonez, Curriculum La- beling: Revisiting Pseudo-Labeling for Semi-Supervised Learn- ing, Proceedings of the AAAI Conference on Artificial Intelli- gence 35 (8) (2021) 6912–6920. doi:10.1609/aaai.v35i8. 16852

  40. [49]

    Zhang, Y

    B. Zhang, Y . Wang, W. Hou, HAO. WU, J. Wang, M. Okumura, T. Shinozaki, FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling, in: Advances in Neural In- formation Processing Systems, V ol. 34, Curran Associates, Inc., 2021, pp. 18408–18419

  41. [50]

    Arazo, D

    E. Arazo, D. Ortego, P. Albert, N. E. O’Connor, K. McGuin- ness, Pseudo-Labeling and Confirmation Bias in Deep Semi- Supervised Learning, in: 2020 International Joint Conference on Neural Networks (IJCNN), IEEE, Glasgow, United Kingdom, 2020, pp. 1–8. doi:10.1109/IJCNN48605.20...

  42. [51]

    Baltatzis, K.-M

    V . Baltatzis, K.-M. Bintsi, L. L. Folgoc, O. E. Martinez Man- zanera, S. Ellis, A. Nair, S. Desai, B. Glocker, J. A. Schnabel, The Pitfalls of Sample Selection: A Case Study on Lung Nod- ule Classification, in: Predictive Intelligence in Medicine, V ol. 12928, Springer Intern...

  43. [52]

    E. A. Kazerooni, J. H. Austin, W. C. Black, D. S. Dyer, T. R. Hazelton, A. N. Leung, M. F. McNitt-Gray, R. F. Munden, S. Pipavath, ACR–STR Practice Parameter for the Perfor- mance and Reporting of Lung Cancer Screening Thoracic Com- puted Tomography (CT): 2014 (Resolution 4)*,...

  44. [53]

    Loshchilov, F

    I. Loshchilov, F. Hutter, Fixing Weight Decay Regularization in Adam (Feb. 2018)

  45. [54]

    Loshchilov, F

    I. Loshchilov, F. Hutter, SGDR: Stochastic Gradient Descent with Warm Restarts (Nov. 2016)

  46. [55]

    Al-Shabi, B

    M. Al-Shabi, B. L. Lan, W. Y . Chan, K.-H. Ng, M. Tan, Lung nodule classification using deep Local–Global net- works, International Journal of Computer Assisted Radiol- ogy and Surgery 14 (10) (2019) 1815–1819. doi:10.1007/ s11548-019-01981-7

  47. [56]

    J. Lu, C. Yin, O. Krause, K. Erleben, M. B. Nielsen, S. Dark- ner, Reducing Annotation Need in Self-explanatory Models for Lung Nodule Diagnosis, in: Interpretability of Machine In- telligence in Medical Image Computing, V ol. 13611, Springer Nature Switzerland, Cham, 2022, pp...

  48. [57]

    J. Lu, C. Yin, K. Erleben, M. B. Nielsen, S. Darkner, cRedAnno+: Annotation Exploitation In Self-Explanatory Lung Nodule Diagnosis, in: 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), IEEE, Cartagena, Colombia, 2023, pp. 1–5. doi:10.1109/ISBI53787.2023. 10...

  49. [211]

    doi:10.1007/978-3-030-87602-9_19 . 20

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.