REVIEW 3 major objections 5 minor 49 references
NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read NOVA is a new evaluation-only benchmark of rare brain MRI pathologies where GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B all exhibit large performance drops across anomaly localization, image captioning, and diagnostic reasoning.
desk verdict NOVA is a genuinely useful new benchmark that is undermined by a false 'double-blinded' claim in its annotation protocol, yet the dataset and baseline results are worth serious consideration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The benchmark asks AI models to do three things: point to the abnormal region in an image (localization), write a short description of what they see (captioning), and make a diagnosis when given the image plus the clinical history (reasoning). The authors tested three leading vision-language models: GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B. Performance was low across the board. For localization, the best model found only about 38% of abnormalities at a lenient overlap threshold and produced hundreds of false positives. For captioning, clinical term F1 scores were below 20%, and for diagnosis, top-1 accuracy was around 24%.
The results suggest that these large, general-purpose models do not transfer well to rare medical conditions. However, the benchmark has some caveats. The radiologists who drew the bounding boxes had access to the original clinical description and diagnosis, while the models only see the image in the localization task. Also, each case is a single 2D slice, not a full 3D volume, and the dataset comes from a single European teaching repository. These factors could make the localization task harder or less fair than intended.
Extended reading notes
Core claim
The paper claims NOVA is the first benchmark to jointly evaluate anomaly localization, visual captioning, and diagnostic reasoning in brain MRI with 281 rare pathologies, and that leading VLMs show 'substantial performance drops across all tasks', establishing it as a rigorous testbed. If true, current vision-language models are unreliable for rare brain MRI analysis in zero-shot settings.
Load-bearing premise
The ground-truth bounding boxes in Task 1 were drawn by radiologists who had access to the full Eurorad clinical description and diagnosis (Section 3.2), whereas models receive only the MRI image. If the anomalies are visually identifiable only with clinical context, the localization scores do not purely measure visual anomaly detection, undermining the benchmark's validity for Task 1.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. NOVA is an evaluation-only benchmark compiled from 906 Eurorad brain MRI cases covering 281 rare diagnoses. The paper defines three tasks: anomaly localization from a single released MRI slice via bounding boxes, image captioning, and diagnostic reasoning from clinical history plus image caption. The authors benchmark GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B and report substantial performance drops on all three tasks relative to natural-image benchmarks, arguing that current vision-language models are unreliable for rare brain MRI analysis in zero-shot settings. The dataset and annotations are released under a CC BY-NC-SA license.
Significance. If the validity concerns are resolved, NOVA addresses a genuine gap: no existing benchmark jointly evaluates anomaly localization, clinical captioning, and diagnostic reasoning on rare brain MRI. The breadth of diagnoses (281), the real-world acquisition heterogeneity, the expert multi-reader annotation protocol, and the inference-only design are strengths, and the authors explicitly acknowledge possible pretraining contamination. The paper also avoids fitted parameters, so the benchmark is not self-fulfilling in the usual sense. However, the two central validity issues—unblinded ground-truth boxes for Task 1 and GPT-4o serving as both judge and contestant in Task 3—must be addressed before the benchmark can support the headline claim that VLMs are unreliable in open-world clinical brain MRI.
major comments (3)
- [§3.2, Abstract] The abstract states that NOVA provides 'double-blinded expert bounding-box annotations,' but Section 3.2 reports that 'Each case was independently labeled by two readers, who reviewed the full original Eurorad clinical description and associated metadata to inform their annotations.' The readers therefore knew the clinical history and likely the diagnosis before drawing boxes, and the senior adjudicator also had this context. Since Task 1 (Section 4.1) gives models only the image, low localization scores conflate visual anomaly detection with the information asymmetry between annotators and models. This is an internal inconsistency: the stated protocol contradicts the 'double-blinded' claim. The authors should either re-annotate a representative subset with image-only readers to quantify the effect, or rescope Task 1 as localization with clinical priors and soften the corresponding claims about VLM visual unreliability.
- [§4.3, Table 3] Task 3 uses GPT-4o 'to perform semantic matching between predictions and ground truth labels,' while GPT-4o is itself one of the models evaluated on the same task. This introduces a self-scoring circularity: GPT-4o's own predictions may be judged more leniently than those of other models if the judge favors its own output distribution. The authors should use a different judge model (e.g., a strong open-source model) or human raters, and should report inter-annotator agreement on a sample of diagnostic predictions to confirm that the semantic matching is not biased.
- [§3.3, §6] The release format is described as 'uniformly sized 480×480 grayscale PNG slices,' and Section 6 acknowledges that only 2D slices are provided, yet the manuscript repeatedly refers to 'scans' and claims to evaluate anomaly localization in brain MRI. A single axial, sagittal, or coronal slice may not contain the full lesion burden or the most representative plane, and the paper does not state how the slice was selected for each case. This limits the external validity of the localization and captioning results. The authors should document the slice-selection rule and report the proportion of cases in which the selected slice actually contains the annotated abnormality.
minor comments (5)
- [Throughout] The spelling 'NOVA' and 'NOV A' is used inconsistently; please unify the dataset name.
- [§5.1, Table 1] The sentence 'over 600 false-positive boxes were recorded' is ambiguous because Table 1 reports FP30 values of 899, 1163, and 672 per model; the text should explicitly state whether this is per-model or total.
- [§3.3] Please clarify whether each 'case' is represented by exactly one slice or by multiple slices, and explain how multiple ground-truth bounding boxes per case map onto the released slice(s).
- [Table 3] The column header 'Cov.' is not defined in the caption; please state explicitly that it denotes the fraction of ground-truth labels covered by the model's prediction vocabulary.
- [§4.2, Table 2] The binary normal-versus-abnormal F1 values in Table 2 are extremely low (2.4–11.3%); since all Eurorad cases are pathological, the paper should explain how the binary label is derived and whether the metric measures any meaningful signal for these models.
Assumptions & free parameters
free parameters (2)
- Consensus IoU threshold =
0.3
- Single 2D slice selection =
1 per case
assumptions (5)
- domain assumption Eurorad cases are representative of rare brain MRI pathologies in clinical practice.
- domain assumption A single 2D slice per case is sufficient for the three benchmark tasks.
- ad hoc to paper Radiologist bounding boxes drawn with knowledge of the clinical history are a valid visual ground truth for anomaly localization.
- domain assumption The 281 diagnosis labels are mutually exclusive and correctly assigned.
- standard math Standard evaluation metrics (mAP, BLEU, METEOR, exact keyword matching) are appropriate for the three tasks.
Cite this review
Pith. "Pith review of NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI." pith.science (2026). https://pith.science/paper/ADFRWIEA
@misc{pith2026250514064,
author = {Pith},
title = {Pith review of: NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI},
year = {2026},
howpublished = {\url{https://pith.science/paper/ADFRWIEA}},
note = {Machine review of arXiv:2505.14064}
}
abstract
In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Out-of-distribution detection identifies whether an input stems from an unseen distribution, while open-world recognition flags such inputs to ensure the system remains robust as ever-emerging, previously $unknown$ categories appear and must be addressed without retraining. Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging. However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use. We therefore present $NOVA$, a challenging, real-life $evaluation-only$ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an $extreme$ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B) reveal substantial performance drops across all tasks, establishing NOVA as a rigorous testbed for advancing models that can detect, localize, and reason about truly unknown anomalies.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
https://brain-development.org/ixi-dataset/
Ixi dataset. https://brain-development.org/ixi-dataset/ . Accessed: 2023-02-15
work page 2023
-
[2]
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare V oss, editors,Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan,...
2005
-
[3]
Bercea, Benedikt Wiestler, Daniel Rueckert, and Julia A Schnabel
Cosmin I. Bercea, Benedikt Wiestler, Daniel Rueckert, and Julia A Schnabel. Generalizing unsupervised anomaly detection: Towards unbiased pathology screening. In Medical Imaging with Deep Learning, 2023
work page 2023
-
[4]
Diffusion models with implicit guidance for medical anomaly detection
Cosmin I Bercea, Benedikt Wiestler, Daniel Rueckert, and Julia A Schnabel. Diffusion models with implicit guidance for medical anomaly detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 211–220. Springer, 2024
work page 2024
-
[5]
Cosmin I Bercea, Benedikt Wiestler, Daniel Rueckert, and Julia A Schnabel. Evaluating normative representation learning in generative ai for robust anomaly detection in brain imaging. Nature Communications, 16(1):1624, 2025
work page 2025
-
[6]
Mvtec ad — a compre- hensive real-world dataset for unsupervised anomaly detection
Paul Bergmann, Michael Fauser, David Sattlegger, and Carsten Steger. Mvtec ad — a compre- hensive real-world dataset for unsupervised anomaly detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9584–9592, 2019
work page 2019
-
[7]
Salinas, and María de la Iglesia-Vayá
Agustín Bustos, Antonio Pertusa, Jose M. Salinas, and María de la Iglesia-Vayá. Padchest: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis, 66:101797, 2020
work page 2020
-
[8]
Padchest-gr: A bilingual chest x-ray dataset for grounded radiology report generation
Daniel C Castro, Aurelia Bustos, Shruthi Bannur, Stephanie L Hyland, Kenza Bouzid, Maria Teodora Wetscherek, Maria Dolores Sánchez-Valverde, Lara Jaques-Pérez, Lourdes Pérez-Rodríguez, Kenji Takeda, et al. Padchest-gr: A bilingual chest x-ray dataset for grounded radiology report generation. arXiv preprint arXiv:2411.05085, 2024
arXiv 2024
Show all 49 references
-
[9]
Unsupervised lesion detection via image restoration with a normative prior
Xiaoran Chen, Suhang You, Kerem Can Tezcan, and Ender Konukoglu. Unsupervised lesion detection via image restoration with a normative prior. Medical Image Analysis, 64:101713, 2020
2020
-
[10]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv...
2010 arXiv
-
[11]
Promptdet: Towards open-vocabulary detection using uncurated images
Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma. Promptdet: Towards open-vocabulary detection using uncurated images. In European conference on computer vision, pages 701–717. Springer, 2022
2022
-
[12]
Recontrast: Domain-specific anomaly detection via contrastive reconstruction
Jia Guo, Shuai Lu, Lize Jia, Weihang Zhang, and Huiqi Li. Recontrast: Domain-specific anomaly detection via contrastive reconstruction. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36...
2023
-
[13]
Developing generalist foun- dation models from a multimodal dataset for 3d computed tomography
Ibrahim Ethem Hamamci, Sezgin Er, Furkan Almas, et al. Developing generalist foun- dation models from a multimodal dataset for 3d computed tomography. arXiv preprint arXiv:2403.17834, 2024
2024
-
[14]
Deep anomaly detection with outlier exposure
Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. arXiv preprint arXiv:1812.04606, 2018
2018 arXiv
-
[15]
Isles 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset
Moritz R Hernandez Petzsche, Ezequiel de la Rosa, Uta Hanning, Roland Wiest, Waldo Valenzuela, Mauricio Reyes, Maria Meyer, Sook-Lei Liew, Florian Kofler, Ivan Ezhov, et al. Isles 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset. Scientific Da...
2022
-
[16]
Out-of-distribution detection in medical image analysis: A survey
Zesheng Hong, Yubiao Yue, Yubin Chen, Lele Cong, Huanjie Lin, Yuanmei Luo, Mini Han Wang, Weidong Wang, Jialong Xu, Xiaoqi Yang, et al. Out-of-distribution detection in medical image analysis: A survey. arXiv preprint arXiv:2404.18279, 2024
2024 arXiv
-
[17]
Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 6:317, 2019
2019
-
[18]
Y . W. Kim and L. T. Mansfield. Fool me twice: Delayed diagnoses in radiology with emphasis on perpetuated errors. AJR. American Journal of Roentgenology, 202(3):465–470, 2014
2014
-
[19]
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M Beery, Jure Leskovec, Anshu...
2021
-
[20]
Emergency triage of brain computed tomography via anomaly detection with a deep generative model
Seungjun Lee, Boryeong Jeong, Minjee Kim, Ryoungwoo Jang, Wooyul Paik, Jiseon Kang, Won Chung, Gil-Sun Hong, and Namkug Kim. Emergency triage of brain computed tomography via anomaly detection with a deep generative model. Nature Communications, 13:4251, 07 2022
2022
-
[21]
A novel public MR image dataset of multiple sclerosis patients with lesion segmentations based on multi-rater consensus
Žiga Lesjak, Alina Galimzianova, Andrej Koren, et al. A novel public MR image dataset of multiple sclerosis patients with lesion segmentations based on multi-rater consensus. Neuroin- formatics, 16(1):51–63, 2018
2018
-
[22]
Cutpaste: Self-supervised learning for anomaly detection and localization
Chun-Liang Li, Kihyuk Sohn, Jinsung Yoon, and Tomas Pfister. Cutpaste: Self-supervised learning for anomaly detection and localization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9664–9674, 2021
2021
-
[23]
Lo, ., and et al
Sook-Lei Liew, Bethany P. Lo, ., and et al. Miarnda R. Donnelly. A large, curated, open-source stroke neuroimaging dataset to improve lesion segmentation algorithms. Scientific Data, 9, 2022
2022
-
[24]
A benchmarking crisis in biomedical machine learning
Faisal Mahmood. A benchmarking crisis in biomedical machine learning. Nature Medicine, 31(4):1060–1060, 2025
2025
-
[25]
Marcus, Aditya F
Daniel S. Marcus, Aditya F. Fotenos, John G. Csernansky, John C. Morris, and Randy L. Buckner. Open access series of imaging studies: longitudinal MRI data in nondemented and demented older adults. Journal of Cognitive Neuroscience, 22(12):2677–2684, 2010
2010
-
[26]
Bjoern H. Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth Gerstner, Marc-André Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buend...
1993
-
[27]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , ...
2002
-
[28]
Petersen, Paul S
Ronald C. Petersen, Paul S. Aisen, Laurel A. Beckett, Michael C. Donohue, Anthony C. Gamst, Danielle J. Harvey, Clifford R. Jr. Jack, William J. Jagust, Leslie M. Shaw, Arthur W. Toga, John Q. Trojanowski, and Michael W. Weiner. Alzheimer’s disease neuroimaging initiative (adn...
2010
-
[29]
Generative probabilistic novelty detection with adversarial autoencoders
Stanislav Pidhorskyi, Ranya Almohsen, and Gianfranco Doretto. Generative probabilistic novelty detection with adversarial autoencoders. Advances in neural information processing systems, 31, 2018
2018
-
[30]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pa...
2021
-
[31]
Do imagenet classifiers generalize to imagenet? In Proceedings of the International Conference on Machine Learning, pages 5389–5400
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In Proceedings of the International Conference on Machine Learning, pages 5389–5400. PMLR, 2019
2019
-
[32]
Buizer, Antonio Federico, Thomas Gasser, Samuel Groeschel, Sanja Hermanns, Thomas Klockgether, Ingeborg Krägeloh-Mann, G
Carola Reinhard, Anne-Catherine Bachoud-Lévi, Tobias Bäumer, Enrico Bertini, Alicia Brunelle, Annemieke I. Buizer, Antonio Federico, Thomas Gasser, Samuel Groeschel, Sanja Hermanns, Thomas Klockgether, Ingeborg Krägeloh-Mann, G. Bernhard Landwehrmeyer, Is- abelle Leber, Alfons...
2020
-
[33]
Towards total recall in industrial anomaly detection
Karsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf, Thomas Brox, and Peter Gehler. Towards total recall in industrial anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14318–14328, 2022
2022
-
[34]
A unifying review of deep and shallow anomaly detection
Lukas Ruff, Jacob R Kauffmann, Robert A Vandermeulen, Grégoire Montavon, Wojciech Samek, Marius Kloft, Thomas G Dietterich, and Klaus-Robert Müller. A unifying review of deep and shallow anomaly detection. Proceedings of the IEEE, 109(5):756–795, 2021
2021
-
[35]
f-anogan: Fast unsupervised anomaly detection with generative adversarial networks
Thomas Schlegl, Philipp Seeböck, Sebastian M Waldstein, Georg Langs, and Ursula Schmidt- Erfurth. f-anogan: Fast unsupervised anomaly detection with generative adversarial networks. Medical Image Analysis, 54:30–44, 2019
2019
-
[36]
Natural synthetic anomalies for self-supervised anomaly detection and localization
Hannah M Schlüter, Jeremy Tan, Benjamin Hou, and Bernhard Kainz. Natural synthetic anomalies for self-supervised anomaly detection and localization. In European Conference on Computer Vision, pages 474–489. Springer, 2022
2022
-
[37]
Uk biobank: An open access resource for identifying the causes of a wide range of complex diseases of middle and old age
Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, et al. Uk biobank: An open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Medicine, 12(3):e1001779, 2015
2015
-
[38]
Taylor, Nitin Williams, Rhodri Cusack, Tibor Auer, Meredith A
Jason R. Taylor, Nitin Williams, Rhodri Cusack, Tibor Auer, Meredith A. Shafto, Marie Dixon, Lorraine K. Tyler, Cam-CAN, and Richard N. Henson. The cambridge centre for ageing and neuroscience (cam-can) data repository: Structural and functional mri, meg, and cognitive data fr...
2017
-
[39]
Bridging ood detection and generalization: A graph-theoretic view
Han Wang and Yixuan Li. Bridging ood detection and generalization: A graph-theoretic view. Advances in Neural Information Processing Systems, 2024
2024
-
[40]
Diffusion models for medical anomaly detection
Julia Wolleb, Florentin Bieder, Robin Sandkühler, and Philippe C Cattin. Diffusion models for medical anomaly detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 35–45. Springer, 2022
2022
-
[41]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115, 2024. 12
2024 arXiv
-
[42]
Openood: Benchmarking generalized out-of- distribution detection
Jingkang Yang, Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, WENXUAN PENG, Haoqi Wang, Guangyao Chen, Bo Li, Yiyou Sun, Xuefeng Du, Kaiyang Zhou, Wayne Zhang, Dan Hendrycks, Yixuan Li, and Ziwei Liu. Openood: Benchmarking generalized out-of- distribution detection. In S...
2022
-
[43]
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In European Conference on Computer Vision, pages 69–85. Springer, 2016
2016
-
[44]
fastmri: An open dataset and benchmarks for accelerated mri
Jure Zbontar, Florian Knoll, Anuroop Sriram, Tullie Murrell, Zhengnan Huang, Matthew J Muckley, Aaron Defazio, Ruben Stern, Patricia Johnson, Mary Bruno, et al. fastmri: An open dataset and benchmarks for accelerated mri. arXiv preprint arXiv:1811.08839, 2018
2018 arXiv
-
[45]
fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data
Ruiyang Zhao, Burhaneddin Yaman, Yuxin Zhang, Russell Stewart, Austin Dixon, Florian Knoll, Zhengnan Huang, Yvonne W Lui, Michael S Hansen, and Matthew P Lungren. fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data. Scienti...
2022
-
[46]
Out-of-distribution detection learning with unreliable out-of-distribution sources
Haotian Zheng, Qizhou Wang, Zhen Fang, Xiaobo Xia, Feng Liu, Tongliang Liu, and Bo Han. Out-of-distribution detection learning with unreliable out-of-distribution sources. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Inform...
2023
-
[47]
A survey on open-vocabulary detection and segmentation: Past, present, and future
Chaoyang Zhu and Long Chen. A survey on open-vocabulary detection and segmentation: Past, present, and future. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
2024
-
[48]
Mood 2020: A public benchmark for out-of-distribution detection and localization on medical images
David Zimmerer, Peter M Full, Fabian Isensee, Paul Jäger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Ross, Annika Reinke, Antanas Kascenas, et al. Mood 2020: A public benchmark for out-of-distribution detection and localization on medical images. IEEE Transactions on Medi...
2020
-
[49]
Unsu- pervised anomaly localization using variational auto-encoders
David Zimmerer, Fabian Isensee, Jens Petersen, Simon Kohl, and Klaus Maier-Hein. Unsu- pervised anomaly localization using variational auto-encoders. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, Octo...
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.