Pith. sign in

REVIEW 4 major objections 5 minor 111 references

MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Masked inpainting with diffusion models weakens shortcut cues and lifts accuracy under domain shift

desk verdict A clean, incremental diffusion-inpainting method for debiasing medical classifiers; the tables support the narrow claim, but the paper never verifies its key assumption that spurious features stay outside the ROI mask. read the letter →

arxiv 2411.10686 v1 pith:UVUSOMJP submitted 2024-11-16 cs.CV cs.LG

classification cs.CVcs.LG
keywords spuriouscorrelationsmedicalimagingdiffusionmodelsimageinpaintingdataaugmentationdomaingeneralizationDreamboothdermoscopy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MaskMedPaint is a data-augmentation procedure that uses a text-to-image diffusion model to repaint the non-informative parts of medical images so they resemble a target hospital's images, while leaving the diagnostic region untouched. The paper argues that this breaks spurious correlations—such as the ruler markings that commonly accompany melanoma in dermoscopy—and improves performance when the model is evaluated on a new domain. The central claim is that with roughly 100 unlabeled target images, MaskMedPaint raises target-domain accuracy across two medical datasets (ISIC 2018 dermoscopy and chest X-rays) and two natural datasets (Waterbirds and iWildCam), without requiring clinicians to describe the spurious feature in words. If correct, this offers a validation-friendly, data-driven route to debiasing classifiers for real-world hospital shifts.

What carries the argument

The central object is a three-stage diffusion pipeline. Stage one finetunes Stable Diffusion on labeled source images using class-name prompts. Stage two removes the region of interest from target images with LaMa, then finetunes the stage-one model with Dreambooth on the remaining backgrounds paired with a dummy token such as 'target'. Stage three protects the region of interest with a segmentation mask and uses the finetuned diffusion model to inpaint everything outside it, conditioning on both the class name and the dummy token. The mask is the load-bearing piece: it confines style transfer to the background, so class-relevant anatomy is preserved while the spurious shortcut is repainted into target-domain appearance.

What would settle it

Run MaskMedPaint on a dataset where the spurious cue overlaps the lesion, for example a surgical marker inside the tumor boundary, using the same segmentation masks; if target-domain accuracy gains vanish or reverse relative to the baseline, the spatial-separability assumption is violated.

Watch

Extended reading notes

Core claim

The paper's central claim is that masked inpainting with a diffusion model can transfer source images into the target domain's visual style while preserving the region of interest, and that this transfer is enough to substantially reduce the classifier's dependence on spurious features. Concretely, MaskMedPaint first finetunes a text-to-image diffusion model on labeled source images, then uses Dreambooth to teach it the look of about 100 unlabeled target images whose regions of interest have been removed, and finally inpaints the source images' backgrounds under prompts that combine the class name and a dummy target token. The augmented images are added to the training set. On ISIC 2018 with a ruler/melanoma spurious correlation, target accuracy rises from 0.146 (baseline) to 0.344; on a MIMIC-CXR to NIH ChestXray14 shift, target AUROC rises from 0.490 to 0.546; on Waterbirds, target accuracy rises from 0.264 to 0.571. The study frames this as evidence that generative data augmentation can mitigate shortcuts that are hard to describe in natural language.

Load-bearing premise

The spurious features targeted for removal must lie entirely outside the region of interest that the segmentation mask protects in both source and target images.

Editorial extensions

If this is right

  • If correct, medical imaging teams could mitigate hospital-to-hospital shifts using a small unlabeled sample from the target site, without needing clinicians to articulate which visual cues are spurious.
  • The method suggests that the key bottleneck for debiasing is not generating counterfactuals but accurately segmenting the region of interest; better masks should make the same pipeline transferable to other spurious features.
  • The ablation on Waterbirds indicates that roughly 2500 generated images approximate the benefit of adding 100 real target images, so generated data could serve as a stopgap when real target data are scarce.
  • The ISIC result shows that even simply masking the background (the Masked baseline) helps substantially, implying that forcing the classifier to ignore the background accounts for much of the gain, and MaskMedPaint recovers some source accuracy that pure masking loses.
  • Because gains appear for both localized artifact shortcuts (rulers) and global shifts (hospital acquisition, grayscale-to-color, night-to-day), the mechanism appears general across different types of spurious correlations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The method's reliance on segmentation quality suggests a natural stress test: apply MaskMedPaint with progressively coarser or noisier ROI masks and measure target accuracy; the advantage should degrade gracefully as mask precision declines.
  • The pipeline could be inverted as a feature-discovery tool: by comparing classifier predictions on original versus inpainted images, one could localize which background changes drive predictions, offering a data-driven way to audit spurious cues without manual framing.
  • Because the authors report that 10–20 target images cause Dreambooth memorization and reduced diversity, a practical extension would be to regularize generation (e.g., with textual inversion or multiple concepts) to lower the target-sample requirement below 100.
  • The success on global shifts such as hospital-to-hospital chest X-rays hints that masked inpainting might also align multi-site imaging protocols, though the paper does not test that directly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. MaskMedPaint proposes a diffusion-based data augmentation method for reducing spurious correlations in medical and natural image classification. The pipeline first fine-tunes a text-to-image diffusion model on labeled source images, then personalizes it with unlabeled target-domain background images via Dreambooth, and finally generates augmented source images by inpainting everything outside a protected region-of-interest mask while conditioning on class and target-domain tokens. The method is evaluated on ISIC 2018 dermoscopy with a ruler artifact, a MIMIC-to-NIH chest X-ray transfer, and the Waterbirds and iWildCam natural-image benchmarks. The central empirical claim is that MaskMedPaint improves target-domain accuracy over a base classifier with no augmentation, given roughly 100 unlabeled target images.

Significance. If the central claim holds, MaskMedPaint would be a useful tool for medical imaging, where spurious features are often difficult to describe linguistically and target data are scarce. The paper has notable strengths: it evaluates on both medical and natural datasets, reports confidence intervals over multiple seeds, includes a deliberate oracle and a ground-truth classifier as sanity checks on the generated counterfactuals, and candidly states a load-bearing limitation in Appendix A.6. The method is also reproducible in principle because the code and dataset links are promised. However, the significance is currently bounded by the fact that the reported gains over the Base baseline are not yet tied to the proposed mechanism, and the experimental protocol leaves key sources of selection bias unaddressed.

major comments (4)
  1. [Section 3 / Appendix A.6] The central mechanism is never verified on the actual benchmark images. MaskMedPaint removes a spurious feature only if that feature lies entirely outside the preserved ROI mask, a condition that Appendix A.6 explicitly acknowledges ('we assume that the spurious feature can be separated from the region of interest identified by the segmentation model'). Yet the paper does not specify which segmentation model is used for ISIC, CXR, Waterbirds, or iWildCam, and it reports no statistics on the spatial relationship between the spurious features (e.g., rulers in ISIC) and the ROI masks. The ISIC split filters out images with patches but does not quantify how many images contain rulers that overlap the lesion boundary. Without this information, the reported ISIC target accuracy gain of 0.344 vs. 0.146 for Base cannot be attributed to counterfactual removal of the ruler rather than to background restyling and regularization. Please specify the segmentation method and report per-dataset overlap statistics, or otherwise demonstrate that the assumed separation holds.
  2. [Section 4.3] The selection of diffusion hyperparameters is underspecified and could compromise the target-domain comparisons. Section 4.3 states that the authors vary strength over {0.5, 0.7, 0.9, 1.0} and guidance scale over {7.5, 15, 20}, but it does not say how the final per-dataset values were chosen. Since the target split is unlabeled and the validation splits in Appendix A.1 appear to contain only source groups, it is not clear whether the sweep was evaluated on the target test set, which would amount to tuning on the test distribution. Please describe the selection protocol explicitly, including what criterion was used and which data were accessed, so that the reported improvements are not the result of selection over hyperparameters.
  3. [Section 5.1 / Table 1] In the ISIC experiment, the Masked baseline achieves the highest target accuracy (0.385) and MaskMedPaint (0.344) has overlapping confidence intervals with it, as the text itself notes. This means the comparison to Base does not establish that masked inpainting is what drives the improvement; simply erasing everything outside the ROI yields statistically indistinguishable performance. Because ISIC is the primary medical demonstration of artifact removal, this weakens the paper's central claim that MaskMedPaint's specific generation mechanism is responsible for the target-domain gains. Please report a formal significance test between MaskMedPaint and Masked on ISIC, and discuss whether the improvement over Base is attributable to masking alone rather than to the diffusion-based background transfer.
  4. [Appendix A.1 / A.5] The number of unlabeled target images used in the main experiments is not consistently reported or varied for most datasets. Appendix A.6 states that 'approximately 100' target images are assumed and that 10-20 yields suboptimal results, but the CXR split in Table 6 uses 100 extra images while the exact numbers for ISIC and iWildCam are not stated in the main text. Since the method's applicability depends on this sample-size regime, please state the exact number of target images per dataset and, if possible, include an ablation on this number for a medical dataset as is already done for Waterbirds in Figure 6.
minor comments (5)
  1. [Section 2.1] There is a typo: 'automately construct' should be 'automatically construct'.
  2. [Section 5.2 / Figure 3] The method is referred to as 'MaskedMedPaint' in the text of Section 5.2 and in Figure 3, whereas the rest of the paper uses 'MaskMedPaint'. Please unify the name.
  3. [Appendix A.1] The Waterbirds split table is difficult to parse because the four-group structure is not labeled with column headings that distinguish species from background. Please reformat the table so the group definitions (landbird on land, landbird on water, waterbird on land, waterbird on water) are explicit.
  4. [Section 5.1] The 'ground-truth' classifier that detects rulers in generated images is described only by its accuracy (0.985). Please state the architecture, training set size, and whether it was trained on generated images only or on a mix of real and generated images, so the sanity check is interpretable.
  5. [Section 4.2] The Masked baseline is defined as training on 'only the ROI, and the remaining area masked out,' but the paper does not say how the ROI mask is obtained for each dataset. Since the method's comparison with Masked is critical, the mask source should be specified here or in the appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: MaskMedPaint's target-domain gains are empirical and not reducible to a fitted input or self-citation.

full rationale

MaskMedPaint's claimed contribution is an empirical augmentation method. The derivation chain is: (1) fine-tune a diffusion model on labeled source images with class prompts; (2) fine-tune with Dreambooth on unlabeled target backgrounds obtained by removing the ROI with LaMa; (3) inpaint source images outside the protected ROI conditioned on class plus target token; (4) train a classifier on original plus augmented images; (5) evaluate on a held-out balanced target test set. The target test labels are never used to fit the diffusion model, the Dreambooth concept, the inpainting strength or guidance, or the classifier. The unlabeled target images are used only to define the target background style, which is the intended mechanism rather than a hidden label leak. No parameter in the pipeline is defined in terms of the target accuracy, and no equation reduces the reported target accuracy to an input. The 'ground-truth' ruler and background classifiers are sanity checks, not inputs to the method. The stated limitation that spurious features must be separable from the ROI (Appendix A.6) is an assumption about the data, not a circular definition; it concerns whether the mechanism will work, not whether the result is equivalent to the input. The fact that the simple Masked baseline sometimes outperforms MaskMedPaint (ISIC target 0.385 vs 0.344) further shows the reported gains are not forced by construction. The only mild concern is that hyperparameter selection over strength and guidance values is not fully specified; even if tuned on target performance, that would be selection bias rather than circularity. No self-citations are load-bearing.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on ML hyperparameters that are chosen by hand (strength, guidance scale, augmentation count) and on domain assumptions about ROI/spurious-feature separability and the sufficiency of ~100 target images. There are no invented physical entities. The most fragile premise is that spurious features live outside the segmentation mask.

free parameters (3)
  • diffusion generation strength = searched over {0.5, 0.7, 0.9, 1.0}
    Chosen by hand in Section 4.3; the paper does not state whether the selection used a validation set or the target test set, so it may be fit to the benchmark.
  • guidance scale = searched over {7.5, 15, 20}
    Same as strength; a generation hyperparameter that changes how strongly the prompt controls the inpainted background and therefore affects target accuracy.
  • number of generated augmented images = 1000/2500/5000 in the Waterbirds ablation; not specified per dataset
    The amount of augmented data added to the training set is a free choice; the paper only ablates it for Waterbirds.
assumptions (4)
  • domain assumption Spurious features can be separated from the region of interest; the segmentation mask correctly identifies the ROI in source and target images.
    Explicitly stated as a limitation in Appendix A.6: if spurious features overlap the protected ROI, the inpainted counterfactuals cannot break the correlation. The whole pipeline depends on this spatial separation.
  • domain assumption Approximately 100 unlabeled target images are sufficient to finetune Dreambooth so it learns target background style without memorizing the target images.
    The method assumes this scale of target data (Section 3, Appendix A.6); the authors report that 10-20 images lead to suboptimal results, so the method only works in this data regime.
  • domain assumption A Stable Diffusion model finetuned on source class names preserves class-discriminative features inside the ROI while inpainting the background.
    Step 3 combines class label and target dummy token in the prompt; if inpainting degrades or alters the ROI, the augmentation would corrupt labels. The paper does not provide a quantitative check of ROI preservation, only visual examples and a ruler-presence classifier.
  • domain assumption For global shifts such as MIMIC to NIH CXR, the domain difference can be captured by repainting the non-ROI background with a single dummy token.
    The CXR experiment treats the whole dataset shift as style/background shift; if the shift also affects the diagnostic region (e.g., pathology appearance), the method may not transfer.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations." pith.science (2026). https://pith.science/paper/UVUSOMJP

@misc{pith2026241110686,
  author       = {Pith},
  title        = {Pith review of: MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UVUSOMJP}},
  note         = {Machine review of arXiv:2411.10686}
}
read the original abstract

Spurious features associated with class labels can lead image classifiers to rely on shortcuts that don't generalize well to new domains. This is especially problematic in medical settings, where biased models fail when applied to different hospitals or systems. In such cases, data-driven methods to reduce spurious correlations are preferred, as clinicians can directly validate the modified images. While Denoising Diffusion Probabilistic Models (Diffusion Models) show promise for natural images, they are impractical for medical use due to the difficulty of describing spurious medical features. To address this, we propose Masked Medical Image Inpainting (MaskMedPaint), which uses text-to-image diffusion models to augment training images by inpainting areas outside key classification regions to match the target domain. We demonstrate that MaskMedPaint enhances generalization to target domains across both natural (Waterbirds, iWildCam) and medical (ISIC 2018, Chest X-ray) datasets, given limited unlabeled target images.

Figures

Figures reproduced from arXiv: 2411.10686 by the authors.

Figure 1
Figure 1. MaskMedPaint pipeline. In step 1, the text-to-image diffusion model is finetuned on the source dataset Dsource, with the class names formatted directly in the text prompts. The diffusion model learns class-conditional features (e.g. benign versus malignant). In step 2, we preprocess the target images by segmenting the ROI and removing the ROI with an inpainting model. We obtain background images of the target domain… view at source ↗
Figure 2
Figure 2. MaskMedPaint Image Generation for ISIC 2018 Dermoscopic Shift. All pos￾sible combinations of skin lesion conditions and rulers in the generated augmentations. In the ISIC 2018 dataset, the artifact of surgical ruler marking is spuriously correlated with the pres￾ence of melanoma. From [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. MaskMedPaint Image Generation for CXR dataset shift. Example of the source MIMIC-CXR image (left) aug￾mented to NIH style with MaskMedPaint (middle). For reference, a CXR from NIH (right). We evaluate global shifts in a medical setting with a transfer between Chest X-ray datasets. With a trans￾fer from MIMIC-CXR to NIH ChestXray14, we see from [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: MaskMedPaint Image Generation for Waterbirds Shift. All possible combina￾tions of bird species and backgrounds in the generated augmentations. Landbirds on water and waterbirds on land are cat￾egories not present in the original source dataset and are counterfactuals g…
Figure 5
Figure 5. Figure 5: MaskMedPaint Image Generation for iWildcam Shifts. Example of the origi￾nal source image (left) adapted to the style of the target domain through MaskMed￾Paint augmentation (middle). For refer￾ence, we provide the real target domain im￾ages from the same class (right).…
Figure 6
Figure 6. Figure 6: Number of Real versus Generated Images Added (Waterbirds). The overall test accuracy (left), source accu￾racy (middle), and target accuracy (right) of adding 10, 20, 50, 100, and 200 real images from the target distribution. As dashed lines, we have the mean accuracy o…
Figure 7
Figure 7. Figure 7: Different Image Generation Meth￾ods (ISIC). The overall test accuracy (left), source accuracy (middle), and tar￾get accuracy (right) of our MaskMedPaint method and the baselines of Text2Img and Img2Img generation. CIs over 5 seeds. to the text prompt. In the ISIC datas…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

111 extracted references · 62 canonical work pages

  1. [1]

    Debiasing skin lesion datasets and models? not so fast

    Alceu Bissoto, Eduardo Valle, and Sandra Avila. Debiasing skin lesion datasets and models? not so fast. In ISIC Skin Image Anaylsis Workshop, 2020 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2020 a

  2. [2]

    Debiasing skin lesion datasets and models? not so fast

    Alceu Bissoto, Eduardo Valle, and Sandra Avila. Debiasing skin lesion datasets and models? not so fast. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 740--741, 2020 b

  3. [3]

    Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups

    Billie Bonevski, Madeleine Randell, Chris Paul, Kathy Chapman, Laura Twyman, Jamie Bryant, Irena Brozek, and Clare Hughes. Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups. BMC medical research methodology, 14: 0 1--29, 2014

  4. [4]

    Instructpix2pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392--18402, 2023

  5. [6]

    Ai for radiographic covid-19 detection selects shortcuts over signal

    Alex J DeGrave, Joseph D Janizek, and Su-In Lee. Ai for radiographic covid-19 detection selects shortcuts over signal. Nature Machine Intelligence, 3 0 (7): 0 610--619, 2021

  6. [7]

    Diffusion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34: 0 8780--8794, 2021

  7. [8]

    Using language to extend to unseen domains

    Lisa Dunlap, Clara Mohri, Devin Guillory, Han Zhang, Trevor Darrell, Joseph E Gonzalez, Aditi Raghunathan, and Anna Rohrbach. Using language to extend to unseen domains. In The Eleventh International Conference on Learning Representations, 2022

  8. [9]

    Diversify your vision datasets with automatic diffusion-based augmentation

    Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. Diversify your vision datasets with automatic diffusion-based augmentation. Advances in Neural Information Processing Systems, 36, 2024

Show all 111 references
  1. [11]

    Unsupervised domain adaptation by backpropagation

    Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180--1189. PMLR, 2015

  2. [12]

    Domain-adversarial training of neural networks

    Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Fran c ois Laviolette, Mario March, and Victor Lempitsky. Domain-adversarial training of neural networks. Journal of machine learning research, 17 0 (59): 0 1--35, 2016

  3. [14]

    Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup

    Zongbo Han, Zhipeng Liang, Fan Yang, Liu Liu, Lanqing Li, Yatao Bian, Peilin Zhao, Bingzhe Wu, Changqing Zhang, and Jianhua Yao. Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup. Advances in Neural Information Processing Systems, 35: 0 3...

  4. [15]

    Densely connected convolutional networks

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700--4708, 2017

  5. [17]

    Key challenges for delivering clinical impact with artificial intelligence

    Christopher J Kelly, Alan Karthikesalingam, Mustafa Suleyman, Greg Corrado, and Dominic King. Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17: 0 1--9, 2019

  6. [20]

    Wilds: A benchmark of in-the-wild distribution shifts

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. In International conference on machine learning, p...

  7. [21]

    Deep domain generalization via conditional invariant adversarial networks

    Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep domain generalization via conditional invariant adversarial networks. In Proceedings of the European conference on computer vision (ECCV), pages 624--639, 2018

  8. [23]

    Torchvision: Pytorch's computer vision library

    TorchVision maintainers and contributors. Torchvision: Pytorch's computer vision library. https://github.com/pytorch/vision, 2016

  9. [24]

    Spuriosity rankings: Sorting data to measure and mitigate biases

    Mazda Moayeri, Wenxiao Wang, Sahil Singla, and Soheil Feizi. Spuriosity rankings: Sorting data to measure and mitigate biases. Advances in Neural Information Processing Systems, 36, 2024

  10. [25]

    Learning from failure: De-biasing classifier from biased classifier

    Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. Learning from failure: De-biasing classifier from biased classifier. Advances in Neural Information Processing Systems, 33: 0 20673--20684, 2020

  11. [26]

    Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis

    Meike Nauta, Ricky Walsh, Adam Dubowski, and Christin Seifert. Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis. Diagnostics, 12 0 (1): 0 40, 2021

  12. [27]

    Hidden stratification causes clinically meaningful failures in machine learning for medical imaging

    Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher R \'e . Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In Proceedings of the ACM conference on health, inference, and learning, pages 151--159, 2020

  13. [28]

    Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification

    Eduardo HP Pooch, Pedro Ballester, and Rodrigo C Barros. Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification. In Thoracic Image Analysis: Second International Workshop, TIA 2020, Held in Conjunction with MICCAI 2020, Lima, ...

  14. [29]

    Explainable, trustworthy, and ethical machine learning for healthcare: A survey

    Khansa Rasheed, Adnan Qayyum, Mohammed Ghaly, Ala Al-Fuqaha, Adeel Razi, and Junaid Qadir. Explainable, trustworthy, and ethical machine learning for healthcare: A survey. Computers in Biology and Medicine, 149: 0 106043, 2022

  15. [30]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684--10695, 2022

  16. [32]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500...

  17. [34]

    Transparency of deep neural networks for medical image analysis: A review of interpretability methods

    Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee, and Philippe Lambin. Transparency of deep neural networks for medical image analysis: A review of interpretability methods. Computers in biology and medicine, 140: 0 105111, 2022

  18. [35]

    Resolution-robust large mask inpainting with fourier convolutions

    Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the IEEE/CVF winte...

  19. [36]

    The importance of interpretability and visualization in machine learning for applications in medicine and health care

    Alfredo Vellido. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural computing and applications, 32 0 (24): 0 18069--18083, 2020

  20. [37]

    Diffusers: State-of-the-art diffusion models

    Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022

  21. [38]

    Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases

    Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and R Summers. Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In IEEE CVPR, volume 7, page 46. sn, 2017

  22. [39]

    Preparing medical imaging data for machine learning

    Martin J Willemink, Wojciech A Koszek, Cailin Hardell, Jie Wu, Dominik Fleischmann, Hugh Harvey, Les R Folio, Ronald M Summers, Daniel L Rubin, and Matthew P Lungren. Preparing medical imaging data for machine learning. Radiology, 295 0 (1): 0 4--15, 2020

  23. [40]

    Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition

    Julia K Winkler, Christine Fink, Ferdinand Toberer, Alexander Enk, Teresa Deinlein, Rainer Hofmann-Wellenhof, Luc Thomas, Aimilios Lallas, Andreas Blum, Wilhelm Stolz, et al. Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep ...

  24. [41]

    Controllable invariance through adversarial feature learning

    Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig. Controllable invariance through adversarial feature learning. Advances in neural information processing systems, 30, 2017

  25. [42]

    Paint by example: Exemplar-based image editing with diffusion models

    Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18381--18391, 2023

  26. [43]

    Improving out-of-distribution robustness via selective augmentation

    Huaxiu Yao, Yu Wang, Sai Li, Linjun Zhang, Weixin Liang, James Zou, and Chelsea Finn. Improving out-of-distribution robustness via selective augmentation. In International Conference on Machine Learning, pages 25407--25437. PMLR, 2022

  27. [44]

    Cutmix: Regularization strategy to train strong classifiers with localizable features

    Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023--6032, 2019

  28. [45]

    Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study

    John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric Karl Oermann. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS medicine, 15 0 (11): 0 e1002683, 2018

  29. [47]

    Model agnostic sample reweighting for out-of-distribution learning

    Xiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang, Renzhe Xu, Peng Cui, and Tong Zhang. Model agnostic sample reweighting for out-of-distribution learning. In International Conference on Machine Learning, pages 27203--27221. PMLR, 2022

  30. [48]

    & Hein, M

    Neuhaus, Y., Augustin, M., Boreiko, V. & Hein, M. Spurious features everywhere-large-scale detection of harmful spurious features in imagenet. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 20235-20246 (2023)

  31. [49]

    & Lopez-Paz, D

    Arjovsky, M., Bottou, L., Gulrajani, I. & Lopez-Paz, D. Invariant risk minimization. ArXiv Preprint ArXiv:1907.02893 . (2019)

  32. [50]

    & Tsotsos, J

    Rosenfeld, A., Zemel, R. & Tsotsos, J. The elephant in the room. ArXiv Preprint ArXiv:1808.03305 . (2018)

  33. [51]

    & Lopez-Paz, D

    Gulrajani, I. & Lopez-Paz, D. In search of lost domain generalization. ArXiv Preprint ArXiv:2007.01434 . (2020)

  34. [52]

    & Saria, S

    Subbaswamy, A. & Saria, S. From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics . 21, 345-352 (2020)

  35. [53]

    & Qadir, J

    Rasheed, K., Qayyum, A., Ghaly, M., Al-Fuqaha, A., Razi, A. & Qadir, J. Explainable, trustworthy, and ethical machine learning for healthcare: A survey. Computers In Biology And Medicine . 149 pp. 106043 (2022)

  36. [54]

    & King, D

    Kelly, C., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine . 17 pp. 1-9 (2019)

  37. [55]

    & Wichmann, F

    Geirhos, R., Jacobsen, J., Michaelis, C., Zemel, R., Brendel, W., Bethge, M. & Wichmann, F. Shortcut learning in deep neural networks. Nature Machine Intelligence . 2, 665-673 (2020)

  38. [56]

    & Feizi, S

    Singla, S. & Feizi, S. Salient imagenet: How to discover spurious features in deep learning?. ArXiv Preprint ArXiv:2110.04301 . (2021)

  39. [57]

    & Dahdouh, S

    Boland, C., Goatman, K., Tsaftaris, S. & Dahdouh, S. There Are No Shortcuts To Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis. Medical Imaging With Deep Learning . (2024)

  40. [58]

    & Tschannen, M

    Minderer, M., Bachem, O., Houlsby, N. & Tschannen, M. Automatic shortcut removal for self-supervised representation learning. International Conference On Machine Learning . pp. 6927-6937 (2020)

  41. [59]

    & Perona, P

    Beery, S., Van Horn, G. & Perona, P. Recognition in terra incognita. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 456-473 (2018)

  42. [60]

    & Oermann, E

    Zech, J., Badgeley, M., Liu, M., Costa, A., Titano, J. & Oermann, E. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Medicine . 15, e1002683 (2018)

  43. [61]

    & Lee, S

    DeGrave, A., Janizek, J. & Lee, S. AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence . 3, 610-619 (2021)

  44. [62]

    & Barros, R

    Pooch, E., Ballester, P. & Barros, R. Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification. Thoracic Image Analysis: Second International Workshop, TIA 2020, Held In Conjunction With MICCAI 2020, Lima, Peru, October 8, 2020,...

  45. [63]

    & Ranganath, R

    Puli, A., Zhang, L., Oermann, E. & Ranganath, R. Out-of-distribution generalization in the presence of nuisance-induced spurious correlations. ArXiv Preprint ArXiv:2107.00520 . (2021)

  46. [64]

    Oakden-Rayner, L., Dunnmon, J., Carneiro, G. & Ré, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proceedings Of The ACM Conference On Health, Inference, And Learning . pp. 151-159 (2020)

  47. [65]

    & Lee, S

    Janizek, J., Erion, G., DeGrave, A. & Lee, S. An adversarial approach for the robust classification of pneumonia from chest radiographs. Proceedings Of The ACM Conference On Health, Inference, And Learning . pp. 69-79 (2020)

  48. [66]

    & Others Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition

    Winkler, J., Fink, C., Toberer, F., Enk, A., Deinlein, T., Hofmann-Wellenhof, R., Thomas, L., Lallas, A., Blum, A., Stolz, W. & Others Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network fo...

  49. [67]

    & Seifert, C

    Nauta, M., Walsh, R., Dubowski, A. & Seifert, C. Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis. Diagnostics . 12, 40 (2021)

  50. [68]

    & Lempitsky, V

    Ganin, Y. & Lempitsky, V. Unsupervised domain adaptation by backpropagation. International Conference On Machine Learning . pp. 1180-1189 (2015)

  51. [69]

    & Neubig, G

    Xie, Q., Dai, Z., Du, Y., Hovy, E. & Neubig, G. Controllable invariance through adversarial feature learning. Advances In Neural Information Processing Systems . 30 (2017)

  52. [70]

    & Tao, D

    Li, Y., Tian, X., Gong, M., Liu, Y., Liu, T., Zhang, K. & Tao, D. Deep domain generalization via conditional invariant adversarial networks. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 624-639 (2018)

  53. [71]

    & Lempitsky, V

    Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., March, M. & Lempitsky, V. Domain-adversarial training of neural networks. Journal Of Machine Learning Research . 17, 1-35 (2016)

  54. [72]

    & Wilson, A

    Kirichenko, P., Izmailov, P. & Wilson, A. Last layer re-training is sufficient for robustness to spurious correlations. ArXiv Preprint ArXiv:2204.02937 . (2022)

  55. [73]

    & Feizi, S

    Moayeri, M., Wang, W., Singla, S. & Feizi, S. Spuriosity rankings: Sorting data to measure and mitigate biases. Advances In Neural Information Processing Systems . 36 (2024)

  56. [74]

    & Wilson, A

    Yang, W., Kirichenko, P., Goldblum, M. & Wilson, A. Chroma-vae: Mitigating shortcut learning with generative classifiers. Advances In Neural Information Processing Systems . 35 pp. 20351-20365 (2022)

  57. [75]

    & Sra, S

    Robinson, J., Sun, L., Yu, K., Batmanghelich, K., Jegelka, S. & Sra, S. Can contrastive learning avoid shortcut solutions?. Advances In Neural Information Processing Systems . 34 pp. 4974-4986 (2021)

  58. [76]

    & Liang, P

    Sagawa, S., Koh, P., Hashimoto, T. & Liang, P. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. ArXiv Preprint ArXiv:1911.08731 . (2019)

  59. [77]

    & Liang, P

    Sagawa, S., Raghunathan, A., Koh, P. & Liang, P. An investigation of why overparameterization exacerbates spurious correlations. International Conference On Machine Learning . pp. 8346-8356 (2020)

  60. [78]

    & Finn, C

    Liu, E., Haghgoo, B., Chen, A., Raghunathan, A., Koh, P., Sagawa, S., Liang, P. & Finn, C. Just train twice: Improving group robustness without training group information. International Conference On Machine Learning . pp. 6781-6792 (2021)

  61. [79]

    & Shin, J

    Nam, J., Cha, H., Ahn, S., Lee, J. & Shin, J. Learning from failure: De-biasing classifier from biased classifier. Advances In Neural Information Processing Systems . 33 pp. 20673-20684 (2020)

  62. [80]

    & Zhang, T

    Zhou, X., Lin, Y., Pi, R., Zhang, W., Xu, R., Cui, P. & Zhang, T. Model agnostic sample reweighting for out-of-distribution learning. International Conference On Machine Learning . pp. 27203-27221 (2022)

  63. [81]

    & Yao, J

    Han, Z., Liang, Z., Yang, F., Liu, L., Li, L., Bian, Y., Zhao, P., Wu, B., Zhang, C. & Yao, J. Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup. Advances In Neural Information Processing Systems . 35 pp. 37704-37718 (2022)

  64. [82]

    & Lopez-Paz, D

    Zhang, H., Cisse, M., Dauphin, Y. & Lopez-Paz, D. mixup: Beyond empirical risk minimization. ArXiv Preprint ArXiv:1710.09412 . (2017)

  65. [83]

    & Yoo, Y

    Yun, S., Han, D., Oh, S., Chun, S., Choe, J. & Yoo, Y. Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 6023-6032 (2019)

  66. [84]

    & Finn, C

    Yao, H., Wang, Y., Li, S., Zhang, L., Liang, W., Zou, J. & Finn, C. Improving out-of-distribution robustness via selective augmentation. International Conference On Machine Learning . pp. 25407-25437 (2022)

  67. [85]

    & Darrell, T

    Dunlap, L., Umino, A., Zhang, H., Yang, J., Gonzalez, J. & Darrell, T. Diversify your vision datasets with automatic diffusion-based augmentation. Advances In Neural Information Processing Systems . 36 (2024)

  68. [86]

    & Zou, J

    Wu, S., Yuksekgonul, M., Zhang, L. & Zou, J. Discover and cure: Concept-aware mitigation of spurious correlation. International Conference On Machine Learning . pp. 37765-37786 (2023)

  69. [87]

    & Kim, J

    Kim, B., Kim, H., Kim, K., Kim, S. & Kim, J. Learning not to learn: Training deep neural networks with biased data. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 9012-9020 (2019)

  70. [88]

    & Rohrbach, A

    Dunlap, L., Mohri, C., Guillory, D., Zhang, H., Darrell, T., Gonzalez, J., Raghunathan, A. & Rohrbach, A. Using language to extend to unseen domains. The Eleventh International Conference On Learning Representations . (2022)

  71. [89]

    & Raff, E

    Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L. & Raff, E. Vqgan-clip: Open domain image generation and editing with natural language guidance. European Conference On Computer Vision . pp. 88-105 (2022)

  72. [90]

    & Feizi, S

    Kattakinda, P., Levine, A. & Feizi, S. Invariant learning via diffusion dreamed distribution shifts. ArXiv Preprint ArXiv:2211.10370 . (2022)

  73. [91]

    & Böttinger, K

    Müller, N., Jacobs, J., Williams, J. & Böttinger, K. Localized Shortcut Removal. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 3720-3724 (2023)

  74. [92]

    & Bigdeli, S

    Weng, N., Pegios, P., Feragen, A., Petersen, E. & Bigdeli, S. Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation. ArXiv Preprint ArXiv:2312.14223 . (2023)

  75. [93]

    & Others Wilds: A benchmark of in-the-wild distribution shifts

    Koh, P., Sagawa, S., Marklund, H., Xie, S., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R., Gao, I. & Others Wilds: A benchmark of in-the-wild distribution shifts. International Conference On Machine Learning . pp. 5637-5664 (2021)

  76. [94]

    & Others Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic)

    Codella, N., Rotemberg, V., Tschandl, P., Celebi, M., Dusza, S., Gutman, D., Helba, B., Kalloo, A., Liopyris, K., Marchetti, M. & Others Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). ArXiv Prepri...

  77. [95]

    & Avila, S

    Bissoto, A., Valle, E. & Avila, S. Debiasing skin lesion datasets and models? not so fast. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition Workshops . pp. 740-741 (2020)

  78. [96]

    & Horng, S

    Johnson, A., Pollard, T., Greenbaum, N., Lungren, M., Deng, C., Peng, Y., Lu, Z., Mark, R., Berkowitz, S. & Horng, S. MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs. ArXiv Preprint ArXiv:1901.07042 . (2019)

  79. [97]

    & Others Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

    Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K. & Others Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. Proceedings Of The AAAI Conference On Artificial Intell...

  80. [98]

    & De La Iglesia-Vaya, M

    Bustos, A., Pertusa, A., Salinas, J. & De La Iglesia-Vaya, M. Padchest: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis . 66 pp. 101797 (2020)

  81. [99]

    & Summers, R

    Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M. & Summers, R. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recog...

  82. [100]

    & Others Learning transferable visual models from natural language supervision

    Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. & Others Learning transferable visual models from natural language supervision. International Conference On Machine Learning . pp. 8748-8763 (2021)

  83. [101]

    & Sun, J

    Wang, Z., Wu, Z., Agarwal, D. & Sun, J. Medclip: Contrastive learning from unpaired medical images and text. ArXiv Preprint ArXiv:2210.10163 . (2022)

  84. [102]

    & Lempitsky, V

    Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K. & Lempitsky, V. Resolution-robust large mask inpainting with fourier convolutions. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Visio...

  85. [103]

    & Ommer, B

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. High-resolution image synthesis with latent diffusion models. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 10684-10695 (2022)

  86. [104]

    & Wolf, T

    Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S. & Wolf, T. Diffusers: State-of-the-art diffusion models. GitHub Repository . (2022), https://github.com/huggingface/diffusers

  87. [105]

    Kingma, D. & Ba, J. Adam: A method for stochastic optimization. ArXiv Preprint ArXiv:1412.6980 . (2014)

  88. [106]

    & Sun, J

    He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 770-778 (2016)

  89. [107]

    & Weinberger, K

    Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Densely connected convolutional networks. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 4700-4708 (2017)

  90. [108]

    & Contributors TorchVision: PyTorch's Computer Vision library

    Maintainers, T. & Contributors TorchVision: PyTorch's Computer Vision library. GitHub Repository . (2016), https://github.com/pytorch/vision

  91. [109]

    & Nichol, A

    Dhariwal, P. & Nichol, A. Diffusion models beat gans on image synthesis. Advances In Neural Information Processing Systems . 34 pp. 8780-8794 (2021)

  92. [110]

    & Keutzer, K

    Liao, P., Li, X., Liu, X. & Keutzer, K. The artbench dataset: Benchmarking generative models with artworks. ArXiv Preprint ArXiv:2206.11404 . (2022)

  93. [111]

    & Wen, F

    Yang, B., Gu, S., Zhang, B., Zhang, T., Chen, X., Sun, X., Chen, D. & Wen, F. Paint by example: Exemplar-based image editing with diffusion models. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 18381-18391 (2023)

  94. [112]

    & Efros, A

    Brooks, T., Holynski, A. & Efros, A. Instructpix2pix: Learning to follow image editing instructions. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 18392-18402 (2023)

  95. [113]

    & Cohen-Or, D

    Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A., Chechik, G. & Cohen-Or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. ArXiv Preprint ArXiv:2208.01618 . (2022)

  96. [114]

    & Aberman, K

    Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M. & Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 22500-22510 (2023)

  97. [115]

    The importance of interpretability and visualization in machine learning for applications in medicine and health care

    Vellido, A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Computing And Applications . 32, 18069-18083 (2020)

  98. [116]

    & Lambin, P

    Salahuddin, Z., Woodruff, H., Chatterjee, A. & Lambin, P. Transparency of deep neural networks for medical image analysis: A review of interpretability methods. Computers In Biology And Medicine . 140 pp. 105111 (2022)

  99. [117]

    & Lungren, M

    Willemink, M., Koszek, W., Hardell, C., Wu, J., Fleischmann, D., Harvey, H., Folio, L., Summers, R., Rubin, D. & Lungren, M. Preparing medical imaging data for machine learning. Radiology . 295, 4-15 (2020)

  100. [118]

    & Hughes, C

    Bonevski, B., Randell, M., Paul, C., Chapman, K., Twyman, L., Bryant, J., Brozek, I. & Hughes, C. Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups. BMC Medical Research Methodology . 14 p...

  101. [119]

    & Selbst, A

    Raji, I., Kumar, I., Horowitz, A. & Selbst, A. The fallacy of AI functionality. Proceedings Of The 2022 ACM Conference On Fairness, Accountability, And Transparency . pp. 959-972 (2022)

  102. [120]

    & Avila, S

    Bissoto, A., Valle, E. & Avila, S. Debiasing Skin Lesion Datasets and Models? Not So Fast. ISIC Skin Image Anaylsis Workshop, 2020 IEEE Conference On Computer Vision And Pattern Recognition Workshops (CVPRW) . (2020)

  103. [121]

    & Summers, R

    Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M. & Summers, R. Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. IEEE CVPR . 7 pp. 46 (2017)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.