REVIEW 4 major objections 5 minor 111 references
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Masked inpainting with diffusion models weakens shortcut cues and lifts accuracy under domain shift
desk verdict A clean, incremental diffusion-inpainting method for debiasing medical classifiers; the tables support the narrow claim, but the paper never verifies its key assumption that spurious features stay outside the ROI mask. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a three-stage diffusion pipeline. Stage one finetunes Stable Diffusion on labeled source images using class-name prompts. Stage two removes the region of interest from target images with LaMa, then finetunes the stage-one model with Dreambooth on the remaining backgrounds paired with a dummy token such as 'target'. Stage three protects the region of interest with a segmentation mask and uses the finetuned diffusion model to inpaint everything outside it, conditioning on both the class name and the dummy token. The mask is the load-bearing piece: it confines style transfer to the background, so class-relevant anatomy is preserved while the spurious shortcut is repainted into target-domain appearance.
What would settle it
Run MaskMedPaint on a dataset where the spurious cue overlaps the lesion, for example a surgical marker inside the tumor boundary, using the same segmentation masks; if target-domain accuracy gains vanish or reverse relative to the baseline, the spatial-separability assumption is violated.
Extended reading notes
Core claim
The paper's central claim is that masked inpainting with a diffusion model can transfer source images into the target domain's visual style while preserving the region of interest, and that this transfer is enough to substantially reduce the classifier's dependence on spurious features. Concretely, MaskMedPaint first finetunes a text-to-image diffusion model on labeled source images, then uses Dreambooth to teach it the look of about 100 unlabeled target images whose regions of interest have been removed, and finally inpaints the source images' backgrounds under prompts that combine the class name and a dummy target token. The augmented images are added to the training set. On ISIC 2018 with a ruler/melanoma spurious correlation, target accuracy rises from 0.146 (baseline) to 0.344; on a MIMIC-CXR to NIH ChestXray14 shift, target AUROC rises from 0.490 to 0.546; on Waterbirds, target accuracy rises from 0.264 to 0.571. The study frames this as evidence that generative data augmentation can mitigate shortcuts that are hard to describe in natural language.
Load-bearing premise
The spurious features targeted for removal must lie entirely outside the region of interest that the segmentation mask protects in both source and target images.
Editorial extensions
If this is right
- If correct, medical imaging teams could mitigate hospital-to-hospital shifts using a small unlabeled sample from the target site, without needing clinicians to articulate which visual cues are spurious.
- The method suggests that the key bottleneck for debiasing is not generating counterfactuals but accurately segmenting the region of interest; better masks should make the same pipeline transferable to other spurious features.
- The ablation on Waterbirds indicates that roughly 2500 generated images approximate the benefit of adding 100 real target images, so generated data could serve as a stopgap when real target data are scarce.
- The ISIC result shows that even simply masking the background (the Masked baseline) helps substantially, implying that forcing the classifier to ignore the background accounts for much of the gain, and MaskMedPaint recovers some source accuracy that pure masking loses.
- Because gains appear for both localized artifact shortcuts (rulers) and global shifts (hospital acquisition, grayscale-to-color, night-to-day), the mechanism appears general across different types of spurious correlations.
Reading between the lines
- The method's reliance on segmentation quality suggests a natural stress test: apply MaskMedPaint with progressively coarser or noisier ROI masks and measure target accuracy; the advantage should degrade gracefully as mask precision declines.
- The pipeline could be inverted as a feature-discovery tool: by comparing classifier predictions on original versus inpainted images, one could localize which background changes drive predictions, offering a data-driven way to audit spurious cues without manual framing.
- Because the authors report that 10–20 target images cause Dreambooth memorization and reduced diversity, a practical extension would be to regularize generation (e.g., with textual inversion or multiple concepts) to lower the target-sample requirement below 100.
- The success on global shifts such as hospital-to-hospital chest X-rays hints that masked inpainting might also align multi-site imaging protocols, though the paper does not test that directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MaskMedPaint proposes a diffusion-based data augmentation method for reducing spurious correlations in medical and natural image classification. The pipeline first fine-tunes a text-to-image diffusion model on labeled source images, then personalizes it with unlabeled target-domain background images via Dreambooth, and finally generates augmented source images by inpainting everything outside a protected region-of-interest mask while conditioning on class and target-domain tokens. The method is evaluated on ISIC 2018 dermoscopy with a ruler artifact, a MIMIC-to-NIH chest X-ray transfer, and the Waterbirds and iWildCam natural-image benchmarks. The central empirical claim is that MaskMedPaint improves target-domain accuracy over a base classifier with no augmentation, given roughly 100 unlabeled target images.
Significance. If the central claim holds, MaskMedPaint would be a useful tool for medical imaging, where spurious features are often difficult to describe linguistically and target data are scarce. The paper has notable strengths: it evaluates on both medical and natural datasets, reports confidence intervals over multiple seeds, includes a deliberate oracle and a ground-truth classifier as sanity checks on the generated counterfactuals, and candidly states a load-bearing limitation in Appendix A.6. The method is also reproducible in principle because the code and dataset links are promised. However, the significance is currently bounded by the fact that the reported gains over the Base baseline are not yet tied to the proposed mechanism, and the experimental protocol leaves key sources of selection bias unaddressed.
major comments (4)
- [Section 3 / Appendix A.6] The central mechanism is never verified on the actual benchmark images. MaskMedPaint removes a spurious feature only if that feature lies entirely outside the preserved ROI mask, a condition that Appendix A.6 explicitly acknowledges ('we assume that the spurious feature can be separated from the region of interest identified by the segmentation model'). Yet the paper does not specify which segmentation model is used for ISIC, CXR, Waterbirds, or iWildCam, and it reports no statistics on the spatial relationship between the spurious features (e.g., rulers in ISIC) and the ROI masks. The ISIC split filters out images with patches but does not quantify how many images contain rulers that overlap the lesion boundary. Without this information, the reported ISIC target accuracy gain of 0.344 vs. 0.146 for Base cannot be attributed to counterfactual removal of the ruler rather than to background restyling and regularization. Please specify the segmentation method and report per-dataset overlap statistics, or otherwise demonstrate that the assumed separation holds.
- [Section 4.3] The selection of diffusion hyperparameters is underspecified and could compromise the target-domain comparisons. Section 4.3 states that the authors vary strength over {0.5, 0.7, 0.9, 1.0} and guidance scale over {7.5, 15, 20}, but it does not say how the final per-dataset values were chosen. Since the target split is unlabeled and the validation splits in Appendix A.1 appear to contain only source groups, it is not clear whether the sweep was evaluated on the target test set, which would amount to tuning on the test distribution. Please describe the selection protocol explicitly, including what criterion was used and which data were accessed, so that the reported improvements are not the result of selection over hyperparameters.
- [Section 5.1 / Table 1] In the ISIC experiment, the Masked baseline achieves the highest target accuracy (0.385) and MaskMedPaint (0.344) has overlapping confidence intervals with it, as the text itself notes. This means the comparison to Base does not establish that masked inpainting is what drives the improvement; simply erasing everything outside the ROI yields statistically indistinguishable performance. Because ISIC is the primary medical demonstration of artifact removal, this weakens the paper's central claim that MaskMedPaint's specific generation mechanism is responsible for the target-domain gains. Please report a formal significance test between MaskMedPaint and Masked on ISIC, and discuss whether the improvement over Base is attributable to masking alone rather than to the diffusion-based background transfer.
- [Appendix A.1 / A.5] The number of unlabeled target images used in the main experiments is not consistently reported or varied for most datasets. Appendix A.6 states that 'approximately 100' target images are assumed and that 10-20 yields suboptimal results, but the CXR split in Table 6 uses 100 extra images while the exact numbers for ISIC and iWildCam are not stated in the main text. Since the method's applicability depends on this sample-size regime, please state the exact number of target images per dataset and, if possible, include an ablation on this number for a medical dataset as is already done for Waterbirds in Figure 6.
minor comments (5)
- [Section 2.1] There is a typo: 'automately construct' should be 'automatically construct'.
- [Section 5.2 / Figure 3] The method is referred to as 'MaskedMedPaint' in the text of Section 5.2 and in Figure 3, whereas the rest of the paper uses 'MaskMedPaint'. Please unify the name.
- [Appendix A.1] The Waterbirds split table is difficult to parse because the four-group structure is not labeled with column headings that distinguish species from background. Please reformat the table so the group definitions (landbird on land, landbird on water, waterbird on land, waterbird on water) are explicit.
- [Section 5.1] The 'ground-truth' classifier that detects rulers in generated images is described only by its accuracy (0.985). Please state the architecture, training set size, and whether it was trained on generated images only or on a mix of real and generated images, so the sanity check is interpretable.
- [Section 4.2] The Masked baseline is defined as training on 'only the ROI, and the remaining area masked out,' but the paper does not say how the ROI mask is obtained for each dataset. Since the method's comparison with Masked is critical, the mask source should be specified here or in the appendix.
Circularity Check
No significant circularity: MaskMedPaint's target-domain gains are empirical and not reducible to a fitted input or self-citation.
full rationale
MaskMedPaint's claimed contribution is an empirical augmentation method. The derivation chain is: (1) fine-tune a diffusion model on labeled source images with class prompts; (2) fine-tune with Dreambooth on unlabeled target backgrounds obtained by removing the ROI with LaMa; (3) inpaint source images outside the protected ROI conditioned on class plus target token; (4) train a classifier on original plus augmented images; (5) evaluate on a held-out balanced target test set. The target test labels are never used to fit the diffusion model, the Dreambooth concept, the inpainting strength or guidance, or the classifier. The unlabeled target images are used only to define the target background style, which is the intended mechanism rather than a hidden label leak. No parameter in the pipeline is defined in terms of the target accuracy, and no equation reduces the reported target accuracy to an input. The 'ground-truth' ruler and background classifiers are sanity checks, not inputs to the method. The stated limitation that spurious features must be separable from the ROI (Appendix A.6) is an assumption about the data, not a circular definition; it concerns whether the mechanism will work, not whether the result is equivalent to the input. The fact that the simple Masked baseline sometimes outperforms MaskMedPaint (ISIC target 0.385 vs 0.344) further shows the reported gains are not forced by construction. The only mild concern is that hyperparameter selection over strength and guidance values is not fully specified; even if tuned on target performance, that would be selection bias rather than circularity. No self-citations are load-bearing.
Assumptions & free parameters
free parameters (3)
- diffusion generation strength =
searched over {0.5, 0.7, 0.9, 1.0}
- guidance scale =
searched over {7.5, 15, 20}
- number of generated augmented images =
1000/2500/5000 in the Waterbirds ablation; not specified per dataset
assumptions (4)
- domain assumption Spurious features can be separated from the region of interest; the segmentation mask correctly identifies the ROI in source and target images.
- domain assumption Approximately 100 unlabeled target images are sufficient to finetune Dreambooth so it learns target background style without memorizing the target images.
- domain assumption A Stable Diffusion model finetuned on source class names preserves class-discriminative features inside the ROI while inpainting the background.
- domain assumption For global shifts such as MIMIC to NIH CXR, the domain difference can be captured by repainting the non-ROI background with a single dummy token.
Cite this review
Pith. "Pith review of MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations." pith.science (2026). https://pith.science/paper/UVUSOMJP
@misc{pith2026241110686,
author = {Pith},
title = {Pith review of: MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations},
year = {2026},
howpublished = {\url{https://pith.science/paper/UVUSOMJP}},
note = {Machine review of arXiv:2411.10686}
}
read the original abstract
Spurious features associated with class labels can lead image classifiers to rely on shortcuts that don't generalize well to new domains. This is especially problematic in medical settings, where biased models fail when applied to different hospitals or systems. In such cases, data-driven methods to reduce spurious correlations are preferred, as clinicians can directly validate the modified images. While Denoising Diffusion Probabilistic Models (Diffusion Models) show promise for natural images, they are impractical for medical use due to the difficulty of describing spurious medical features. To address this, we propose Masked Medical Image Inpainting (MaskMedPaint), which uses text-to-image diffusion models to augment training images by inpainting areas outside key classification regions to match the target domain. We demonstrate that MaskMedPaint enhances generalization to target domains across both natural (Waterbirds, iWildCam) and medical (ISIC 2018, Chest X-ray) datasets, given limited unlabeled target images.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Debiasing skin lesion datasets and models? not so fast
Alceu Bissoto, Eduardo Valle, and Sandra Avila. Debiasing skin lesion datasets and models? not so fast. In ISIC Skin Image Anaylsis Workshop, 2020 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2020 a
2020
-
[2]
Debiasing skin lesion datasets and models? not so fast
Alceu Bissoto, Eduardo Valle, and Sandra Avila. Debiasing skin lesion datasets and models? not so fast. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 740--741, 2020 b
2020
-
[3]
Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups
Billie Bonevski, Madeleine Randell, Chris Paul, Kathy Chapman, Laura Twyman, Jamie Bryant, Irena Brozek, and Clare Hughes. Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups. BMC medical research methodology, 14: 0 1--29, 2014
2014
-
[4]
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392--18402, 2023
2023
-
[6]
Ai for radiographic covid-19 detection selects shortcuts over signal
Alex J DeGrave, Joseph D Janizek, and Su-In Lee. Ai for radiographic covid-19 detection selects shortcuts over signal. Nature Machine Intelligence, 3 0 (7): 0 610--619, 2021
2021
-
[7]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34: 0 8780--8794, 2021
2021
-
[8]
Using language to extend to unseen domains
Lisa Dunlap, Clara Mohri, Devin Guillory, Han Zhang, Trevor Darrell, Joseph E Gonzalez, Aditi Raghunathan, and Anna Rohrbach. Using language to extend to unseen domains. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[9]
Diversify your vision datasets with automatic diffusion-based augmentation
Lisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang, Joseph E Gonzalez, and Trevor Darrell. Diversify your vision datasets with automatic diffusion-based augmentation. Advances in Neural Information Processing Systems, 36, 2024
2024
Show all 111 references
-
[11]
Unsupervised domain adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180--1189. PMLR, 2015
2015
-
[12]
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Fran c ois Laviolette, Mario March, and Victor Lempitsky. Domain-adversarial training of neural networks. Journal of machine learning research, 17 0 (59): 0 1--35, 2016
2016
-
[14]
Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup
Zongbo Han, Zhipeng Liang, Fan Yang, Liu Liu, Lanqing Li, Yatao Bian, Peilin Zhao, Bingzhe Wu, Changqing Zhang, and Jianhua Yao. Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup. Advances in Neural Information Processing Systems, 35: 0 3...
2022
-
[15]
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700--4708, 2017
2017
-
[17]
Key challenges for delivering clinical impact with artificial intelligence
Christopher J Kelly, Alan Karthikesalingam, Mustafa Suleyman, Greg Corrado, and Dominic King. Key challenges for delivering clinical impact with artificial intelligence. BMC medicine, 17: 0 1--9, 2019
2019
-
[20]
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. In International conference on machine learning, p...
2021
-
[21]
Deep domain generalization via conditional invariant adversarial networks
Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep domain generalization via conditional invariant adversarial networks. In Proceedings of the European conference on computer vision (ECCV), pages 624--639, 2018
2018
-
[23]
Torchvision: Pytorch's computer vision library
TorchVision maintainers and contributors. Torchvision: Pytorch's computer vision library. https://github.com/pytorch/vision, 2016
2016
-
[24]
Spuriosity rankings: Sorting data to measure and mitigate biases
Mazda Moayeri, Wenxiao Wang, Sahil Singla, and Soheil Feizi. Spuriosity rankings: Sorting data to measure and mitigate biases. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[25]
Learning from failure: De-biasing classifier from biased classifier
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. Learning from failure: De-biasing classifier from biased classifier. Advances in Neural Information Processing Systems, 33: 0 20673--20684, 2020
2020
-
[26]
Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis
Meike Nauta, Ricky Walsh, Adam Dubowski, and Christin Seifert. Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis. Diagnostics, 12 0 (1): 0 40, 2021
2021
-
[27]
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher R \'e . Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In Proceedings of the ACM conference on health, inference, and learning, pages 151--159, 2020
2020
-
[28]
Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification
Eduardo HP Pooch, Pedro Ballester, and Rodrigo C Barros. Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification. In Thoracic Image Analysis: Second International Workshop, TIA 2020, Held in Conjunction with MICCAI 2020, Lima, ...
2020
-
[29]
Explainable, trustworthy, and ethical machine learning for healthcare: A survey
Khansa Rasheed, Adnan Qayyum, Mohammed Ghaly, Ala Al-Fuqaha, Adeel Razi, and Junaid Qadir. Explainable, trustworthy, and ethical machine learning for healthcare: A survey. Computers in Biology and Medicine, 149: 0 106043, 2022
2022
-
[30]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684--10695, 2022
2022
-
[32]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500...
2023
-
[34]
Transparency of deep neural networks for medical image analysis: A review of interpretability methods
Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee, and Philippe Lambin. Transparency of deep neural networks for medical image analysis: A review of interpretability methods. Computers in biology and medicine, 140: 0 105111, 2022
2022
-
[35]
Resolution-robust large mask inpainting with fourier convolutions
Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the IEEE/CVF winte...
2022
-
[36]
The importance of interpretability and visualization in machine learning for applications in medicine and health care
Alfredo Vellido. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural computing and applications, 32 0 (24): 0 18069--18083, 2020
2020
-
[37]
Diffusers: State-of-the-art diffusion models
Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022
2022
-
[38]
Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and R Summers. Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In IEEE CVPR, volume 7, page 46. sn, 2017
2017
-
[39]
Preparing medical imaging data for machine learning
Martin J Willemink, Wojciech A Koszek, Cailin Hardell, Jie Wu, Dominik Fleischmann, Hugh Harvey, Les R Folio, Ronald M Summers, Daniel L Rubin, and Matthew P Lungren. Preparing medical imaging data for machine learning. Radiology, 295 0 (1): 0 4--15, 2020
2020
-
[40]
Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition
Julia K Winkler, Christine Fink, Ferdinand Toberer, Alexander Enk, Teresa Deinlein, Rainer Hofmann-Wellenhof, Luc Thomas, Aimilios Lallas, Andreas Blum, Wilhelm Stolz, et al. Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep ...
2019
-
[41]
Controllable invariance through adversarial feature learning
Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig. Controllable invariance through adversarial feature learning. Advances in neural information processing systems, 30, 2017
2017
-
[42]
Paint by example: Exemplar-based image editing with diffusion models
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18381--18391, 2023
2023
-
[43]
Improving out-of-distribution robustness via selective augmentation
Huaxiu Yao, Yu Wang, Sai Li, Linjun Zhang, Weixin Liang, James Zou, and Chelsea Finn. Improving out-of-distribution robustness via selective augmentation. In International Conference on Machine Learning, pages 25407--25437. PMLR, 2022
2022
-
[44]
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023--6032, 2019
2019
-
[45]
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study
John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric Karl Oermann. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS medicine, 15 0 (11): 0 e1002683, 2018
2018
-
[47]
Model agnostic sample reweighting for out-of-distribution learning
Xiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang, Renzhe Xu, Peng Cui, and Tong Zhang. Model agnostic sample reweighting for out-of-distribution learning. In International Conference on Machine Learning, pages 27203--27221. PMLR, 2022
2022
-
[48]
& Hein, M
Neuhaus, Y., Augustin, M., Boreiko, V. & Hein, M. Spurious features everywhere-large-scale detection of harmful spurious features in imagenet. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 20235-20246 (2023)
2023
-
[49]
& Lopez-Paz, D
Arjovsky, M., Bottou, L., Gulrajani, I. & Lopez-Paz, D. Invariant risk minimization. ArXiv Preprint ArXiv:1907.02893 . (2019)
2019 arXiv
-
[50]
& Tsotsos, J
Rosenfeld, A., Zemel, R. & Tsotsos, J. The elephant in the room. ArXiv Preprint ArXiv:1808.03305 . (2018)
2018 arXiv
-
[51]
& Lopez-Paz, D
Gulrajani, I. & Lopez-Paz, D. In search of lost domain generalization. ArXiv Preprint ArXiv:2007.01434 . (2020)
2020 arXiv
-
[52]
& Saria, S
Subbaswamy, A. & Saria, S. From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics . 21, 345-352 (2020)
2020
-
[53]
& Qadir, J
Rasheed, K., Qayyum, A., Ghaly, M., Al-Fuqaha, A., Razi, A. & Qadir, J. Explainable, trustworthy, and ethical machine learning for healthcare: A survey. Computers In Biology And Medicine . 149 pp. 106043 (2022)
2022
-
[54]
& King, D
Kelly, C., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine . 17 pp. 1-9 (2019)
2019
-
[55]
& Wichmann, F
Geirhos, R., Jacobsen, J., Michaelis, C., Zemel, R., Brendel, W., Bethge, M. & Wichmann, F. Shortcut learning in deep neural networks. Nature Machine Intelligence . 2, 665-673 (2020)
2020
-
[56]
& Feizi, S
Singla, S. & Feizi, S. Salient imagenet: How to discover spurious features in deep learning?. ArXiv Preprint ArXiv:2110.04301 . (2021)
2021 arXiv
-
[57]
& Dahdouh, S
Boland, C., Goatman, K., Tsaftaris, S. & Dahdouh, S. There Are No Shortcuts To Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis. Medical Imaging With Deep Learning . (2024)
2024
-
[58]
& Tschannen, M
Minderer, M., Bachem, O., Houlsby, N. & Tschannen, M. Automatic shortcut removal for self-supervised representation learning. International Conference On Machine Learning . pp. 6927-6937 (2020)
2020
-
[59]
& Perona, P
Beery, S., Van Horn, G. & Perona, P. Recognition in terra incognita. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 456-473 (2018)
2018
-
[60]
& Oermann, E
Zech, J., Badgeley, M., Liu, M., Costa, A., Titano, J. & Oermann, E. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Medicine . 15, e1002683 (2018)
2018
-
[61]
& Lee, S
DeGrave, A., Janizek, J. & Lee, S. AI for radiographic COVID-19 detection selects shortcuts over signal. Nature Machine Intelligence . 3, 610-619 (2021)
2021
-
[62]
& Barros, R
Pooch, E., Ballester, P. & Barros, R. Can we trust deep learning based diagnosis? the impact of domain shift in chest radiograph classification. Thoracic Image Analysis: Second International Workshop, TIA 2020, Held In Conjunction With MICCAI 2020, Lima, Peru, October 8, 2020,...
2020
-
[63]
& Ranganath, R
Puli, A., Zhang, L., Oermann, E. & Ranganath, R. Out-of-distribution generalization in the presence of nuisance-induced spurious correlations. ArXiv Preprint ArXiv:2107.00520 . (2021)
2021 arXiv
-
[64]
Oakden-Rayner, L., Dunnmon, J., Carneiro, G. & Ré, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proceedings Of The ACM Conference On Health, Inference, And Learning . pp. 151-159 (2020)
2020
-
[65]
& Lee, S
Janizek, J., Erion, G., DeGrave, A. & Lee, S. An adversarial approach for the robust classification of pneumonia from chest radiographs. Proceedings Of The ACM Conference On Health, Inference, And Learning . pp. 69-79 (2020)
2020
-
[66]
& Others Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition
Winkler, J., Fink, C., Toberer, F., Enk, A., Deinlein, T., Hofmann-Wellenhof, R., Thomas, L., Lallas, A., Blum, A., Stolz, W. & Others Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network fo...
2019
-
[67]
& Seifert, C
Nauta, M., Walsh, R., Dubowski, A. & Seifert, C. Uncovering and correcting shortcut learning in machine learning models for skin cancer diagnosis. Diagnostics . 12, 40 (2021)
2021
-
[68]
& Lempitsky, V
Ganin, Y. & Lempitsky, V. Unsupervised domain adaptation by backpropagation. International Conference On Machine Learning . pp. 1180-1189 (2015)
2015
-
[69]
& Neubig, G
Xie, Q., Dai, Z., Du, Y., Hovy, E. & Neubig, G. Controllable invariance through adversarial feature learning. Advances In Neural Information Processing Systems . 30 (2017)
2017
-
[70]
& Tao, D
Li, Y., Tian, X., Gong, M., Liu, Y., Liu, T., Zhang, K. & Tao, D. Deep domain generalization via conditional invariant adversarial networks. Proceedings Of The European Conference On Computer Vision (ECCV) . pp. 624-639 (2018)
2018
-
[71]
& Lempitsky, V
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., March, M. & Lempitsky, V. Domain-adversarial training of neural networks. Journal Of Machine Learning Research . 17, 1-35 (2016)
2016
-
[72]
& Wilson, A
Kirichenko, P., Izmailov, P. & Wilson, A. Last layer re-training is sufficient for robustness to spurious correlations. ArXiv Preprint ArXiv:2204.02937 . (2022)
2022 arXiv
-
[73]
& Feizi, S
Moayeri, M., Wang, W., Singla, S. & Feizi, S. Spuriosity rankings: Sorting data to measure and mitigate biases. Advances In Neural Information Processing Systems . 36 (2024)
2024
-
[74]
& Wilson, A
Yang, W., Kirichenko, P., Goldblum, M. & Wilson, A. Chroma-vae: Mitigating shortcut learning with generative classifiers. Advances In Neural Information Processing Systems . 35 pp. 20351-20365 (2022)
2022
-
[75]
& Sra, S
Robinson, J., Sun, L., Yu, K., Batmanghelich, K., Jegelka, S. & Sra, S. Can contrastive learning avoid shortcut solutions?. Advances In Neural Information Processing Systems . 34 pp. 4974-4986 (2021)
2021
-
[76]
& Liang, P
Sagawa, S., Koh, P., Hashimoto, T. & Liang, P. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. ArXiv Preprint ArXiv:1911.08731 . (2019)
2019 arXiv
-
[77]
& Liang, P
Sagawa, S., Raghunathan, A., Koh, P. & Liang, P. An investigation of why overparameterization exacerbates spurious correlations. International Conference On Machine Learning . pp. 8346-8356 (2020)
2020
-
[78]
& Finn, C
Liu, E., Haghgoo, B., Chen, A., Raghunathan, A., Koh, P., Sagawa, S., Liang, P. & Finn, C. Just train twice: Improving group robustness without training group information. International Conference On Machine Learning . pp. 6781-6792 (2021)
2021
-
[79]
& Shin, J
Nam, J., Cha, H., Ahn, S., Lee, J. & Shin, J. Learning from failure: De-biasing classifier from biased classifier. Advances In Neural Information Processing Systems . 33 pp. 20673-20684 (2020)
2020
-
[80]
& Zhang, T
Zhou, X., Lin, Y., Pi, R., Zhang, W., Xu, R., Cui, P. & Zhang, T. Model agnostic sample reweighting for out-of-distribution learning. International Conference On Machine Learning . pp. 27203-27221 (2022)
2022
-
[81]
& Yao, J
Han, Z., Liang, Z., Yang, F., Liu, L., Li, L., Bian, Y., Zhao, P., Wu, B., Zhang, C. & Yao, J. Umix: Improving importance weighting for subpopulation shift via uncertainty-aware mixup. Advances In Neural Information Processing Systems . 35 pp. 37704-37718 (2022)
2022
-
[82]
& Lopez-Paz, D
Zhang, H., Cisse, M., Dauphin, Y. & Lopez-Paz, D. mixup: Beyond empirical risk minimization. ArXiv Preprint ArXiv:1710.09412 . (2017)
2017 arXiv
-
[83]
& Yoo, Y
Yun, S., Han, D., Oh, S., Chun, S., Choe, J. & Yoo, Y. Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceedings Of The IEEE/CVF International Conference On Computer Vision . pp. 6023-6032 (2019)
2019
-
[84]
& Finn, C
Yao, H., Wang, Y., Li, S., Zhang, L., Liang, W., Zou, J. & Finn, C. Improving out-of-distribution robustness via selective augmentation. International Conference On Machine Learning . pp. 25407-25437 (2022)
2022
-
[85]
& Darrell, T
Dunlap, L., Umino, A., Zhang, H., Yang, J., Gonzalez, J. & Darrell, T. Diversify your vision datasets with automatic diffusion-based augmentation. Advances In Neural Information Processing Systems . 36 (2024)
2024
-
[86]
& Zou, J
Wu, S., Yuksekgonul, M., Zhang, L. & Zou, J. Discover and cure: Concept-aware mitigation of spurious correlation. International Conference On Machine Learning . pp. 37765-37786 (2023)
2023
-
[87]
& Kim, J
Kim, B., Kim, H., Kim, K., Kim, S. & Kim, J. Learning not to learn: Training deep neural networks with biased data. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 9012-9020 (2019)
2019
-
[88]
& Rohrbach, A
Dunlap, L., Mohri, C., Guillory, D., Zhang, H., Darrell, T., Gonzalez, J., Raghunathan, A. & Rohrbach, A. Using language to extend to unseen domains. The Eleventh International Conference On Learning Representations . (2022)
2022
-
[89]
& Raff, E
Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L. & Raff, E. Vqgan-clip: Open domain image generation and editing with natural language guidance. European Conference On Computer Vision . pp. 88-105 (2022)
2022
-
[90]
& Feizi, S
Kattakinda, P., Levine, A. & Feizi, S. Invariant learning via diffusion dreamed distribution shifts. ArXiv Preprint ArXiv:2211.10370 . (2022)
2022 arXiv
-
[91]
& Böttinger, K
Müller, N., Jacobs, J., Williams, J. & Böttinger, K. Localized Shortcut Removal. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 3720-3724 (2023)
2023
-
[92]
& Bigdeli, S
Weng, N., Pegios, P., Feragen, A., Petersen, E. & Bigdeli, S. Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation. ArXiv Preprint ArXiv:2312.14223 . (2023)
2023 arXiv
-
[93]
& Others Wilds: A benchmark of in-the-wild distribution shifts
Koh, P., Sagawa, S., Marklund, H., Xie, S., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R., Gao, I. & Others Wilds: A benchmark of in-the-wild distribution shifts. International Conference On Machine Learning . pp. 5637-5664 (2021)
2021
-
[94]
& Others Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic)
Codella, N., Rotemberg, V., Tschandl, P., Celebi, M., Dusza, S., Gutman, D., Helba, B., Kalloo, A., Liopyris, K., Marchetti, M. & Others Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the international skin imaging collaboration (isic). ArXiv Prepri...
2019 arXiv
-
[95]
& Avila, S
Bissoto, A., Valle, E. & Avila, S. Debiasing skin lesion datasets and models? not so fast. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition Workshops . pp. 740-741 (2020)
2020
-
[96]
& Horng, S
Johnson, A., Pollard, T., Greenbaum, N., Lungren, M., Deng, C., Peng, Y., Lu, Z., Mark, R., Berkowitz, S. & Horng, S. MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs. ArXiv Preprint ArXiv:1901.07042 . (2019)
2019 arXiv
-
[97]
& Others Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K. & Others Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. Proceedings Of The AAAI Conference On Artificial Intell...
2019
-
[98]
& De La Iglesia-Vaya, M
Bustos, A., Pertusa, A., Salinas, J. & De La Iglesia-Vaya, M. Padchest: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis . 66 pp. 101797 (2020)
2020
-
[99]
& Summers, R
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M. & Summers, R. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recog...
2017
-
[100]
& Others Learning transferable visual models from natural language supervision
Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. & Others Learning transferable visual models from natural language supervision. International Conference On Machine Learning . pp. 8748-8763 (2021)
2021
-
[101]
& Sun, J
Wang, Z., Wu, Z., Agarwal, D. & Sun, J. Medclip: Contrastive learning from unpaired medical images and text. ArXiv Preprint ArXiv:2210.10163 . (2022)
2022 arXiv
-
[102]
& Lempitsky, V
Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K. & Lempitsky, V. Resolution-robust large mask inpainting with fourier convolutions. Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Visio...
2022
-
[103]
& Ommer, B
Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. High-resolution image synthesis with latent diffusion models. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 10684-10695 (2022)
2022
-
[104]
& Wolf, T
Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S. & Wolf, T. Diffusers: State-of-the-art diffusion models. GitHub Repository . (2022), https://github.com/huggingface/diffusers
2022
-
[105]
Kingma, D. & Ba, J. Adam: A method for stochastic optimization. ArXiv Preprint ArXiv:1412.6980 . (2014)
2014 arXiv
-
[106]
& Sun, J
He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 770-778 (2016)
2016
-
[107]
& Weinberger, K
Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, K. Densely connected convolutional networks. Proceedings Of The IEEE Conference On Computer Vision And Pattern Recognition . pp. 4700-4708 (2017)
2017
-
[108]
& Contributors TorchVision: PyTorch's Computer Vision library
Maintainers, T. & Contributors TorchVision: PyTorch's Computer Vision library. GitHub Repository . (2016), https://github.com/pytorch/vision
2016
-
[109]
& Nichol, A
Dhariwal, P. & Nichol, A. Diffusion models beat gans on image synthesis. Advances In Neural Information Processing Systems . 34 pp. 8780-8794 (2021)
2021
-
[110]
& Keutzer, K
Liao, P., Li, X., Liu, X. & Keutzer, K. The artbench dataset: Benchmarking generative models with artworks. ArXiv Preprint ArXiv:2206.11404 . (2022)
2022 arXiv
-
[111]
& Wen, F
Yang, B., Gu, S., Zhang, B., Zhang, T., Chen, X., Sun, X., Chen, D. & Wen, F. Paint by example: Exemplar-based image editing with diffusion models. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 18381-18391 (2023)
2023
-
[112]
& Efros, A
Brooks, T., Holynski, A. & Efros, A. Instructpix2pix: Learning to follow image editing instructions. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 18392-18402 (2023)
2023
-
[113]
& Cohen-Or, D
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A., Chechik, G. & Cohen-Or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. ArXiv Preprint ArXiv:2208.01618 . (2022)
2022 arXiv
-
[114]
& Aberman, K
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M. & Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition . pp. 22500-22510 (2023)
2023
-
[115]
The importance of interpretability and visualization in machine learning for applications in medicine and health care
Vellido, A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Computing And Applications . 32, 18069-18083 (2020)
2020
-
[116]
& Lambin, P
Salahuddin, Z., Woodruff, H., Chatterjee, A. & Lambin, P. Transparency of deep neural networks for medical image analysis: A review of interpretability methods. Computers In Biology And Medicine . 140 pp. 105111 (2022)
2022
-
[117]
& Lungren, M
Willemink, M., Koszek, W., Hardell, C., Wu, J., Fleischmann, D., Harvey, H., Folio, L., Summers, R., Rubin, D. & Lungren, M. Preparing medical imaging data for machine learning. Radiology . 295, 4-15 (2020)
2020
-
[118]
& Hughes, C
Bonevski, B., Randell, M., Paul, C., Chapman, K., Twyman, L., Bryant, J., Brozek, I. & Hughes, C. Reaching the hard-to-reach: a systematic review of strategies for improving health and medical research with socially disadvantaged groups. BMC Medical Research Methodology . 14 p...
2014
-
[119]
& Selbst, A
Raji, I., Kumar, I., Horowitz, A. & Selbst, A. The fallacy of AI functionality. Proceedings Of The 2022 ACM Conference On Fairness, Accountability, And Transparency . pp. 959-972 (2022)
2022
-
[120]
& Avila, S
Bissoto, A., Valle, E. & Avila, S. Debiasing Skin Lesion Datasets and Models? Not So Fast. ISIC Skin Image Anaylsis Workshop, 2020 IEEE Conference On Computer Vision And Pattern Recognition Workshops (CVPRW) . (2020)
2020
-
[121]
& Summers, R
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M. & Summers, R. Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. IEEE CVPR . 7 pp. 46 (2017)
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.