Pith. sign in

REVIEW 2 major objections 100 references

MozzaVID: Mozzarella Volumetric Image Dataset

T0 review · 2 major / 0 minor · reviewed 2026-05-23 · grok-4.3

Pith's one-line read MozzaVID supplies X-ray CT volumes of 25 mozzarella types to train and benchmark volumetric deep-learning models.

desk verdict MozzaVID is a straightforward release of a new multi-resolution CT dataset for 25-class mozzarella classification aimed at volumetric benchmarking. read the letter →

arxiv 2412.04880 v3 submitted 2024-12-06 cs.CV eess.IV

classification cs.CVeess.IV
keywords volumetricdatasetCTimagingmozzarellamicrostructurecheeseclassification3DdeeplearningfoodstructureanalysisX-raytomography
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces MozzaVID, a dataset of computed tomography scans that capture the internal microstructure of mozzarella cheese samples. It covers 149 samples from 25 distinct cheese types and supplies the volumes at three different resolutions, yielding between 591 and 37,824 images per version. The work responds to the scarcity of clean, labeled volumetric datasets that researchers can use to compare and improve three-dimensional deep-learning architectures. Beyond model development, the scans also let users examine how cheese microstructure varies across types and imaging scales.

What carries the argument

The MozzaVID dataset of labeled 3D CT volumes, provided at multiple resolutions, that serves as both a classification benchmark and a microstructural reference.

What would settle it

Any classifier trained on the training split fails to exceed random-guess accuracy on a held-out test set of the same cheese types.

Watch

Extended reading notes

Core claim

The authors acquire and label X-ray CT volumes of mozzarella, organize them into a classification task over 25 cheese types and 149 samples, and release the data at three resolutions so that volumetric networks can be trained and evaluated on food microstructures that are both complex and disordered.

Load-bearing premise

The CT volumes and their cheese-type labels are clean enough and representative enough to support reliable distinction among the 25 varieties.

Editorial extensions

If this is right

  • Volumetric deep-learning models gain a standardized benchmark for comparing architectures on three-dimensional data.
  • Researchers can measure how spatial resolution affects classification performance on disordered food structures.
  • The dataset supports studies that link visible microstructure features to cheese type without destructive sampling.
  • Algorithms developed on MozzaVID can be tested for robustness on other complex, non-periodic materials.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same acquisition and labeling approach could be repeated for other food products to create comparable volumetric benchmarks.
  • Success on MozzaVID may indicate whether a model will handle similar tasks in medical or materials CT imaging.
  • The multi-resolution design allows direct experiments on the trade-off between detail and computational cost in 3D classification.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper introduces MozzaVID, a volumetric classification dataset of X-ray CT images of mozzarella cheese microstructures. It enables classification of 25 cheese types across 149 samples and supplies the data in three resolutions, yielding dataset instances with 591 to 37,824 images. The work targets benchmarking of volumetric deep-learning models while also supporting investigation of food microstructure properties.

Significance. If the released volumes and labels prove clean and representative, the dataset would address the noted shortage of volumetric imaging benchmarks and provide a testbed for algorithms handling complex, disordered structures. The multi-resolution releases and public exploration link constitute concrete strengths for reproducibility and downstream use in both general volumetric methods and food-science applications.

major comments (2)
  1. [Abstract] Abstract: the assertion that the dataset is 'large, clean, and versatile' supplies no quantitative evidence (label accuracy, inter-sample consistency, exclusion criteria, or diversity statistics) to support the 25-class classification claim.
  2. [Abstract] Abstract, paragraph 2: the central assumption that the CT volumes and labels are 'sufficiently clean and representative' for reliable classification is load-bearing yet unsupported by any description of imaging parameters, sample preparation, labeling protocol, or quality-control steps.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the detailed feedback on the abstract. The comments correctly identify that the abstract makes claims without direct supporting evidence or references. We address each point below and will revise the abstract in the next version.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the assertion that the dataset is 'large, clean, and versatile' supplies no quantitative evidence (label accuracy, inter-sample consistency, exclusion criteria, or diversity statistics) to support the 25-class classification claim.

    Authors: We agree that the abstract does not supply quantitative evidence for these descriptors. The main text provides sample counts, resolution variants, and class distribution, but the abstract itself does not reference specific metrics such as label accuracy or exclusion criteria. We will revise the abstract to remove or qualify the phrase 'large, clean, and versatile' and add a brief pointer to the relevant statistics and methods sections. revision: yes

  2. Referee: [Abstract] Abstract, paragraph 2: the central assumption that the CT volumes and labels are 'sufficiently clean and representative' for reliable classification is load-bearing yet unsupported by any description of imaging parameters, sample preparation, labeling protocol, or quality-control steps.

    Authors: The manuscript body contains dedicated sections on imaging parameters, sample preparation, and labeling. However, the abstract does not cite or summarize these elements. We will revise the abstract to include a short reference to the methods and quality-control procedures so that the assumption is no longer unsupported within the abstract itself. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No derivations, predictions or fitted quantities; dataset release paper

full rationale

The manuscript is a pure data-release contribution. It describes the acquisition, labeling and formatting of CT volumes of mozzarella samples for 25-class classification but contains no equations, no model derivations, no parameter fitting, no predictions, and no load-bearing self-citations. The central claim is simply the existence and utility of the released dataset files, which are evaluated externally rather than by internal logical reduction. No step reduces to its own inputs by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Dataset release paper; contains no mathematical model, fitted parameters, axioms, or postulated entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MozzaVID: Mozzarella Volumetric Image Dataset." pith.science (2026). https://pith.science/paper/2412.04880

@misc{pith2026241204880,
  author       = {Pith},
  title        = {Pith review of: MozzaVID: Mozzarella Volumetric Image Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2412.04880}},
  note         = {Machine review of arXiv:2412.04880}
}
read the original abstract

Influenced by the complexity of volumetric imaging, there is a shortage of established datasets useful for benchmarking volumetric deep-learning models. As a consequence, new and existing models are not easily comparable, limiting the development of architectures optimized specifically for volumetric data. To counteract this trend, we introduce MozzaVID -- a large, clean, and versatile volumetric classification dataset. Our dataset contains X-ray computed tomography (CT) images of mozzarella microstructure and enables the classification of 25 cheese types and 149 cheese samples. We provide data in three different resolutions, resulting in three dataset instances containing from 591 to 37,824 images. While targeted for developing general-purpose volumetric algorithms, the dataset also facilitates investigating the properties of mozzarella microstructure. The complex and disordered nature of food structures brings a unique challenge, where a choice of appropriate imaging method, scale, and sample size is not trivial. With this dataset, we aim to address these complexities, contributing to more robust structural analysis models and a deeper understanding of food structure. The dataset can be explored through: https://papieta.github.io/MozzaVID/

Figures

Figures reproduced from arXiv: 2412.04880 by the authors.

Figure 1
Figure 1. Comparison of typical volumetric and 2D dataset sizes. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Mozzarella samples wrapped in parafilm and mounted [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Sketch of the three proposed dataset configurations. The [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: UMAP generated from second-to-last layer feature representations of the best-performing model in the coarse-grained classifi [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Overview of the variation in the normalized experimental design parameters in the first 24 cheese types (coarse-grained classes). [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: PCA of the experimental design parameters used to de [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: UMAP generated from second-to-last layer feature representations of the best-performing model in the fine-grained classification [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Overview of slices from each cheese type, forming the 25 coarse-grained classes. [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: Example slices from the fine-grained classes. Each row represents a set of six samples from one cheese type (coarse-grained [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

100 extracted references · 100 canonical work pages

  1. [1]

    Mariam Andersson, Hans Martin Kjer, Jonathan Rafael- Patino, Alexandra Pacureanu, Bente Pakkenberg, Jean- Philippe Thiran, Maurice Ptito, Martin Bech, Anders Bjorholm Dahl, Vedrana Andersen Dahl, and Tim B. Dyrby. Axon morphology is modulated by the local environment and impacts the noninvasive investigation of its struc- ture–function relationship. Proce...

  2. [2]

    Armato, III, Geoffrey McLennan, Luc Bidaut, Michael F

    Samuel G. Armato, III, Geoffrey McLennan, Luc Bidaut, Michael F. McNitt-Gray, Charles R. Meyer, Anthony P. Reeves, Binsheng Zhao, Denise R. Aberle, Claudia I. Hen- schke, Eric A. Hoffman, Ella A. Kazerooni, Heber MacMa- hon, Edwin J. R. Van Beek, David Yankelevitz, Alberto M. Biancardi, Peyton H. Bland, Matthew S. Brown, Roger M Engelmann, Gary E. Laderac...

  3. [3]

    Freymann, Justin S

    Ujjwal Baid, Satyam Ghodasara, Suyash Mohan, Michel Bilello, Evan Calabrese, Errol Colak, Keyvan Farahani, Jayashree Kalpathy-Cramer, Felipe Campos Kitamura, Sarthak Pati, Luciano Prevedello, Jeffrey Rudie, Chiharu Sako, Russell Shinohara, Timothy Bergquist, Rong Chai, James Eddy, Julia Elliott, Walter Reade, Thomas Schaffter, Thomas Yu, Jiaxin Zheng, Chr...

  4. [4]

    Kirby, John B

    Spyridon Bakas, Hamed Akbari, Aristeidis Sotiras, Michel Bilello, Martin Rozycki, Justin S. Kirby, John B. Freymann, Keyvan Farahani, and Christos Davatzikos. Advancing the cancer genome atlas glioma MRI collections with expert seg- mentation labels and radiomic features.Scientific Data, 4(1): 170117, 2017. 3

  5. [5]

    Ramona Bast, Prateek Sharma, Hannah K. B. Easton, Tzvetelin T. Dessev, Mita Lad, and Peter A. Munro. Ten- sile testing to quantitate the anisotropy and strain hardening of mozzarella cheese. International Dairy Journal, 44:6–14,

  6. [6]

    Prospectus of cultured meat - advancing meat alternatives

    Zuhaib Fayaz Bhat and Hina Fayaz. Prospectus of cultured meat - advancing meat alternatives. Journal of Food Science and Technology, 48(2):125–140, 2011. 2

  7. [7]

    Patrick Bilic, Patrick Christ, Hongwei Bran Li, Eugene V orontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Loh ¨ofer, Julian Walter Holch, Wieland Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev- Cohain, Michal Drozdzal, Michal Marianne Amitai, Re- fael Vivanti, Jacob Sosn...

  8. [8]

    ViT-V-Net: Vision transformer for unsupervised volumetric medical image registration

    Junyu Chen, Yufan He, Eric Frey, Ye Li, and Yong Du. ViT-V-Net: Vision transformer for unsupervised volumetric medical image registration. In Medical Imaging with Deep Learning, 2021. 8

Show all 100 references
  1. [9]

    Lungren, Shaoting Zhang, Lei Xing, Le Lu, Alan Yuille, and Yuyin Zhou

    Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie, Ehsan Adeli, Yan Wang, Matthew P. Lungren, Shaoting Zhang, Lei Xing, Le Lu, Alan Yuille, and Yuyin Zhou. TransUNet: rethinking the u-net architecture design for medical image segm...

  2. [10]

    Schwing, Alexan- der Kirillov, and Rohit Girdhar

    Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. Proceedings of the Ieee Computer Society Conference on Computer Vi- sion and Pattern Recognition, 2022-:1280–1289, 2022. 8

  3. [11]

    Determination of hip-joint loading patterns of living and extinct mammals using an inverse Wolff’s law approach

    Patrik Christen, Keita Ito, Frietson Galis, and Bert van Riet- bergen. Determination of hip-joint loading patterns of living and extinct mammals using an inverse Wolff’s law approach. Biomechanics and Modeling in Mechanobiology, 14(2):427– 432, 2015. 2

  4. [12]

    Lienkamp, Thomas Brox, and Olaf Ronneberger

    ¨Ozg¨un C ¸ ic ¸ek, Ahmed Abdulkadir, Soeren S. Lienkamp, Thomas Brox, and Olaf Ronneberger. 3D U-Net: Learn- ing dense volumetric segmentation from sparse annotation. In Medical Image Computing and Computer-Assisted Inter- 9 vention (MICCAI 2016) , pages 424–432. Lecture Note...

  5. [13]

    Crippa, E

    M. Crippa, E. Solazzo, D. Guizzardi, F. Monforti-Ferrario, F. N. Tubiello, and A. Leip. Food systems are responsible for a third of global anthropogenic GHG emissions. Nature Food, 2(3):198–209, 2021. 2

  6. [14]

    Cunningham, Imran A

    John A. Cunningham, Imran A. Rahman, Stephan Lauten- schlager, Emily J. Rayfield, and Philip C. J. Donoghue. A virtual world of paleontology. Trends in Ecology and Evolu- tion, 29(6):347–357, 2014. 1

  7. [15]

    Ching, K

    Francesco De Carlo, Do ˘ga G¨ursoy, Daniel J. Ching, K. Joost Batenburg, Wolfgang Ludwig, Lucia Mancini, Federica Marone, Rajmund Mokso, Dani ¨el M. Pelt, Jan Sijbers, and Mark Rivers. TomoBank: A tomographic data repository for computational x-ray science. Measurement Science...

  8. [16]

    ImageNet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proceedings of Computer Vision and Pattern Recognition (CVPR). IEEE, 2009. 3

  9. [17]

    The MNIST database of handwritten digit images for machine learning research [best of the web].IEEE Signal Processing Magazine, 29(6):141–142, 2012

    Li Deng. The MNIST database of handwritten digit images for machine learning research [best of the web].IEEE Signal Processing Magazine, 29(6):141–142, 2012. 1, 3

  10. [18]

    Kevin Zhou

    Yang Deng, Ce Wang, Yuan Hui, Qian Li, Jun Li, Shiwei Luo, Mengke Sun, Quan Quan, Shuxin Yang, You Hao, Pengbo Liu, Honghu Xiao, Chunpeng Zhao, Xinbao Wu, and S. Kevin Zhou. CTSpine1K: A large-scale dataset for spinal vertebrae segmentation in computed tomography. arXiv:2105.1...

  11. [19]

    Dobson and A

    S. Dobson and A. G. Marangoni. Methodology and develop- ment of a high-protein plant-based cheese alternative. Cur- rent Research in Food Science, 7:100632, 2023. 2

  12. [20]

    Computer-aided diagnosis in medical imaging: Historical review, current status and future potential

    Kunio Doi. Computer-aided diagnosis in medical imaging: Historical review, current status and future potential. Com- puterized Medical Imaging and Graphics, 31(4-5):198–211,

  13. [21]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...

  14. [22]

    X-ray computed tomography for quality inspection of agricultural products: A review

    Zhe Du, Yongguang Hu, Noman Ali Buttar, and Ashraf Mah- mood. X-ray computed tomography for quality inspection of agricultural products: A review. Food Science & Nutrition, 7(10):3146–3160, 2019. 1

  15. [23]

    Anton du Plessis and William P. Boshoff. A review of X-ray computed tomography of concrete and asphalt construction materials. Construction and Building Materials , 199:637– 651, 2019. 1

  16. [24]

    Ran Feng, Sylvain Barjon, Frans W. J. van den Berg, Søren Kristian Lillevang, and Lilia Ahrn ´e. Effect of res- idence time in the cooker-stretcher on mozzarella cheese composition, structure and functionality. Journal of Food Engineering, 309:110690, 2021. 4

  17. [25]

    van der Berg, Rajmund Mokso, Søren Kristian Lillevang, and Lilia Ahrn ´e

    Ran Feng, Franciscus Winfried J. van der Berg, Rajmund Mokso, Søren Kristian Lillevang, and Lilia Ahrn ´e. Struc- tural, rheological and functional properties of extruded moz- zarella cheese influenced by the properties of the renneted casein gels. Food Hydrocolloids, 137:1083...

  18. [26]

    Frisullo, J

    P. Frisullo, J. Laverse, R. Marino, and M.A. Del Nobile. X-ray computed tomography to study processed meat mi- crostructure. Journal of Food Engineering , 94(3–4):283– 289, 2009. 2

  19. [27]

    A re- view of applications of CT imaging on fiber reinforced com- posites

    Yantao Gao, Wenfeng Hu, Sanfa Xin, and Lijuan Sun. A re- view of applications of CT imaging on fiber reinforced com- posites. Journal of Composite Materials , 56(1):133–164,

  20. [28]

    S. C. Garcea, Y . Wang, and P. J. Withers. X-ray computed to- mography of polymer composites. Composites Science and Technology, 156:305–319, 2018. 2

  21. [29]

    Godoi, Sangeeta Prakash, and Bhesh R

    Fernanda C. Godoi, Sangeeta Prakash, and Bhesh R. Bhan- dari. 3D printing technologies applied for food design: Sta- tus and prospects. Journal of Food Engineering, 179:44–54,

  22. [30]

    3D semantic segmentation with submanifold sparse convolutional networks

    Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3D semantic segmentation with submanifold sparse convolutional networks. Proceedings of Computer Vision and Pattern Recognition (CVPR) , pages 9224–9232,

  23. [31]

    Granlund

    Goesta H. Granlund. In search of a general picture process- ing operator. Computer Graphics and Image Processing , 8 (2):155–173, 1978. 8

  24. [32]

    Haralick, Karthikeyan Shanmugam, and Its’Hak Dinstein

    Robert M. Haralick, Karthikeyan Shanmugam, and Its’Hak Dinstein. Textural features for image classification. IEEE Transactions on Systems, Man, and Cybernetics , SMC3(6): 610–621, 1973. 8

  25. [33]

    Roth, and Daguang Xu

    Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang, Dong Yang, Holger R. Roth, and Daguang Xu. Swin UNETR: Swin transformers for semantic segmentation of brain tu- mors in MRI images. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries (BrainLes 2021), pa...

  26. [34]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceedings of Computer Vision and Pattern Recognition (CVPR), pages 770–778. IEEE, 2016. 2, 5

  27. [35]

    Statistical shape models for 3D medical image segmentation: A review.Med- ical Image Analysis, 13(4):543–563, 2009

    Tobias Heimann and Hans Peter Meinzer. Statistical shape models for 3D medical image segmentation: A review.Med- ical Image Analysis, 13(4):543–563, 2009. 1

  28. [36]

    The KiTS19 challenge data: 300 kidney tumor cases with clinical context, CT semantic segmenta- tions, and surgical outcomes

    Nicholas Heller, Niranjan Sathianathen, Arveen Kalapara, Edward Walczak, Keenan Moore, Heather Kaluzniak, Joel Rosenberg, Paul Blake, Zachary Rengel, Makinna Oestre- ich, Joshua Dean, Michael Tradewell, Aneri Shah, Resha Tejpaul, Zachary Edgerton, Matthew Peterson, Shaneab- ba...

  29. [37]

    Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xi- aoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan, Guangrui 10 Mu, Zhiyong Lin, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang, Yixin Wang, Feng Hou, Jiawei Yang, Guang- wei Xiong, Jiang Tian, Cheng Zhong, Jun Ma, Jack Rick-...

  30. [38]

    BugNIST a large volumetric dataset for object detection under domain shift

    Patrick Møller Jensen, Vedrana Andersen Dahl, Rebecca Engberg, Carsten Gundlach, Hans Marin Kjer, and An- ders Bjorholm Dahl. BugNIST a large volumetric dataset for object detection under domain shift. Proceedings of the 18th European Conference on Computer Vision – Eccv 2024,...

  31. [39]

    Christensen, Vedrana A

    Niels Jeppesen, Anders N. Christensen, Vedrana A. Dahl, and Anders B. Dahl. Sparse layered graphs for multi-object segmentation. In Proceedings of Computer Vision and Pat- tern Recognition (CVPR), pages 12774–12782. IEEE, 2020. 2

  32. [40]

    Jeppesen, L

    N. Jeppesen, L. P. Mikkelsen, A. B. Dahl, A. N. Christensen, and V . A. Dahl. Quantifying effects of manufacturing meth- ods on fiber orientation in unidirectional composites using structure tensor analysis. Composites Part A: Applied Sci- ence and Manufacturing, 149:106541, 2021. 8

  33. [41]

    Martulli, Martin Kerschbaum, Ivan Sergeichev, Yentl Swolfs, and Stepan V

    Radmir Karamov, Luca M. Martulli, Martin Kerschbaum, Ivan Sergeichev, Yentl Swolfs, and Stepan V . Lomov. Micro- CT based structure tensor analysis of fibre orientation in ran- dom fibre composites versus high-fidelity fibre identification methods. Composite Structures, 235:11...

  34. [42]

    Anderson, Jimmy Huynh, Jeff Gelb, Jouni Freund, and Alp Karakoc ¸

    ¨Ozg¨ur Keles ¸, Eric H. Anderson, Jimmy Huynh, Jeff Gelb, Jouni Freund, and Alp Karakoc ¸. Stochastic fracture of addi- tively manufactured porous composites. Scientific Reports, 8 (1):15437, 2018. 2

  35. [43]

    NeuralVDB: High-resolution sparse volume representation using hierar- chical neural networks

    Doyub Kim, Minjae Lee, and Ken Museth. NeuralVDB: High-resolution sparse volume representation using hierar- chical neural networks. ACM Transactions on Graphics, 43 (2):20, 2024. 2

  36. [44]

    Muscle struc- ture assessment using synchrotron radiation x-ray micro- computed tomography in murine with cerebral ischemia.Sci- entific Reports, 14(1):26825, 2024

    Subok Kim, Sanghun Jang, and Onseok Lee. Muscle struc- ture assessment using synchrotron radiation x-ray micro- computed tomography in murine with cerebral ischemia.Sci- entific Reports, 14(1):26825, 2024. 2

  37. [45]

    The open images dataset V4: Unified image classification, object detection, and visual relationship detection at scale

    Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Ui- jlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset V4: Unified image classification, object detection, and vi...

  38. [46]

    Vision transformer for small-size datasets, 2021

    Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song. Vision transformer for small-size datasets, 2021. 8

  39. [47]

    Lo, Miranda R

    Sook-Lei Liew, Bethany P. Lo, Miranda R. Donnelly, Artemis Zavaliangos-Petropulu, Jessica N. Jeong, Giuseppe Barisano, Alexandre Hutton, Julia P. Simon, Julia M. Ju- liano, Anisha Suri, Zhizhuo Wang, Aisha Abdullah, Jun Kim, Tyler Ard, Nerisa Banaj, Michael R. Borich, Lara A. ...

  40. [48]

    Lawrence Zitnick

    Tsung Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. Lec- ture Notes in Computer Science (including Subseries Lecture Notes in Artificial Intelligence and Lectur...

  41. [49]

    Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Ar- naud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen A. W. M. van der Laak, Bram van Gin- neken, and Clara I. S ´anchez. A survey on deep learning in medical image analysis. Medical Image Analysis, 42:60–88,

  42. [50]

    Kevin Zhou

    Pengbo Liu, Hu Han, Yuanqi Du, Heqin Zhu, Yinhao Li, Feng Gu, Honghu Xiao, Jun Li, Chunpeng Zhao, Li Xiao, Xinbao Wu, and S. Kevin Zhou. Deep Learning to Segment Pelvic Bones: Large-scale CT Datasets and Baseline Mod- els, 2021. 3

  43. [51]

    Deep learning face attributes in the wild

    Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), pages 3730–3738. IEEE, 2015. 1, 3

  44. [52]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. Pro- ceedings of International Conference on Computer Vision (ICCV), pages 9992–10002, 2021. 2, 5, 8

  45. [53]

    A ConvNet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A ConvNet for the 2020s. In Proceedings of Computer Vision and Pattern Recognition (CVPR), pages 11966–11976. IEEE, 2022. 2, 5

  46. [54]

    Decoupled weight de- cay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight de- cay regularization. 7th International Conference on Learn- ing Representations, Iclr 2019, 2019. 5

  47. [55]

    3D-CNN- PyTorch: PyTorch implementation for 3dCNNs for medical images

    3D-CNN-PyTorch maintainers and contributors. 3D-CNN- PyTorch: PyTorch implementation for 3dCNNs for medical images. https://github.com/xmuyzz/3D-CNN- PyTorch, 2022. 5 11

  48. [56]

    TorchVision: Py- Torch’s computer vision library

    TorchVision maintainers and contributors. TorchVision: Py- Torch’s computer vision library. https://github. com/pytorch/vision, 2016. 5

  49. [57]

    Maire and P

    E. Maire and P. J. Withers. Quantitative X-ray tomography. International Materials Reviews, 59(1):1–43, 2014. 1

  50. [58]

    Marcus, Tracy H

    Daniel S. Marcus, Tracy H. Wang, Jamie Parker, John G. Csernansky, John C. Morris, and Randy L. Buckner. Open access series of imaging studies (OASIS): Cross-sectional MRI data in young, middle aged, nondemented, and de- mented older adults. Journal of Cognitive Neuroscience, ...

  51. [59]

    Recipe1m+: A dataset for learning cross-modal embeddings for cooking recipes and food images

    Javier Marin, Aritro Biswas, Ferda Ofli, Nicholas Hynes, Amaia Salvador, Yusuf Aytar, Ingmar Weber, and Antonio Torralba. Recipe1m+: A dataset for learning cross-modal embeddings for cooking recipes and food images. IEEE Transactions on Pattern Analysis and Machine Intelligenc...

  52. [60]

    Comparing vision transformers and convolutional neural net- works for image classification: A literature review

    Jos ´e Maur ´ıcio, In ˆes Domingues, and Jorge Bernardino. Comparing vision transformers and convolutional neural net- works for image classification: A literature review. Applied Sciences (switzerland), 13(9):5521, 2023. 8

  53. [61]

    UMAP: Uniform manifold approximation and projection

    Leland McInnes, John Healy, Nathaniel Saul, and Lukas Großberger. UMAP: Uniform manifold approximation and projection. Journal of Open Source Software , 3(29):861,

  54. [62]

    SANet: A slice-aware network for pulmonary nodule detection

    Jie Mei, Ming-Ming Cheng, Gang Xu, Lan-Ruo Wan, and Huan Zhang. SANet: A slice-aware network for pulmonary nodule detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(8):4374–4387, 2021. 1, 3

  55. [63]

    Bjoern H. Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth Gerstner, Marc-Andre Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buend...

  56. [64]

    Morozov, Anna E

    Sergey P. Morozov, Anna E. Andreychenko, Ivan A. Blokhin, Pavel B. Gelezhe, Anna P. Gonchar, Alexander E. Nikolaev, Nikolay A. Pavlov, Valeria Yu. Chernina, and Vic- tor A. Gombolevskiy. MosMedData: data set of 1110 chest CT scans performed during the covid-19 epidemic. Digita...

  57. [65]

    Rita Verdelho, Diogo J

    Tiago Mota, M. Rita Verdelho, Diogo J. Ara ´ujo, Alceu Bis- soto, Carlos Santiago, and Catarina Barata. MMIST-ccRCC: A real world medical dataset for the development of multi- modal systems. In Proceedings of Computer Vision and Pat- tern Recognition (CVPR), pages 2395–2403. I...

  58. [66]

    M. F. Mridha, Akibur Rahman Prodeep, A. S.M.Morshedul Hoque, Md Rashedul Islam, Aklima Akter Lima, Muham- mad Mohsin Kabir, Md Abdul Hamid, and Yutaka Watanobe. A comprehensive survey on the progress, process, and chal- lenges of lung cancer detection and classification.Journa...

  59. [67]

    Rusch, and R

    National Academies of Sciences, Engineering, and Medicine, Health and Medicine Division, Board on Population Health, Public Health Practice, Roundtable on Environmental Health Sciences, Research, and Medicine, E. Rusch, and R. Pool. Principles and Obstacles for Sharing Data fr...

  60. [68]

    Applications of x-ray micro-computed tomography and small-angle x-ray scattering techniques in food systems: A concise review

    Sunday Olakanmi, Chithra Karunakaran, and Digvir Jayas. Applications of x-ray micro-computed tomography and small-angle x-ray scattering techniques in food systems: A concise review. Journal of Food Engineering, 342:111355,

  61. [69]

    Feature-centered first order structure tensor scale-space in 2d and 3d

    Pawel Tomasz Pieta, Anders Bjorholm Dahl, Jeppe Revall Frisvad, Siavash Arjomand Bigdeli, and Anders Nymark Christensen. Feature-centered first order structure tensor scale-space in 2d and 3d. Ieee Access, 13:9766–9779, 2025. 8

  62. [70]

    3D imaging in material science: Application of X-ray tomography

    Luc Salvo, Michel Su ´ery, Ariane Marmottant, Nathalie Limodin, and Dominique Bernard. 3D imaging in material science: Application of X-ray tomography. Comptes Rendus Physique, 11(9-10):641–649, 2010. 1

  63. [71]

    MobileNetV2: Inverted residuals and linear bottlenecks

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang Chieh Chen. MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of Computer Vision and Pattern Recognition (CVPR) , pages 4510–4520,

  64. [72]

    Correlative imaging of the murine hind limb vasculature and muscle tissue by microCT and light mi- croscopy

    Laura Schaad, Ruslan Hlushchuk, S ´ebastien Barr´e, Roberto Gianni-Barrera, David Haberth ¨ur, Andrea Banfi, and Valentin Djonov. Correlative imaging of the murine hind limb vasculature and muscle tissue by microCT and light mi- croscopy. Scientific Reports, 7(1):41842, 2017. 2

  65. [73]

    X-ray micro-computed tomography ( µCT) for non-destructive characterisation of food microstructure

    Letitia Schoeman, Paul Williams, Anton du Plessis, and Marena Manley. X-ray micro-computed tomography ( µCT) for non-destructive characterisation of food microstructure. Trends in Food Science & Technology, 47:10–24, 2016. 1

  66. [74]

    Arnaud Arindra Adiyoso Setio, Alberto Traverso, Thomas de Bel, Moira S. N. Berens, Cas van den Bogaard, Pier- giorgio Cerello, Hao Chen, Qi Dou, Maria Evelina Fan- tacci, Bram Geurts, Robbert van der Gugten, Pheng Ann Heng, Bart Jansen, Michael M. J. de Kaste, Valentin Ko- tov...

  67. [75]

    Deep learning in medical image analysis

    Dinggang Shen, Guorong Wu, and Heung-Il Suk. Deep learning in medical image analysis. Annual Review of Biomedical Engineering, 19(1):221–248, 2017. 1

  68. [76]

    Direct ob- servation and measurement of fiber architecture in short fiber-polymer composite foam through micro-CT imaging

    Hongbin Shen, Steven Nutt, and David Hull. Direct ob- servation and measurement of fiber architecture in short fiber-polymer composite foam through micro-CT imaging. Composites Science and Technology, 64(13-14):2113–2120,

  69. [77]

    Singh, Lipo Wang, Sukrit Gupta, Haveesh Goli, Parasuraman Padmanabhan, and Bal ´azs Guly ´as

    Satya P. Singh, Lipo Wang, Sukrit Gupta, Haveesh Goli, Parasuraman Padmanabhan, and Bal ´azs Guly ´as. 3D deep learning on medical images: A review. Sensors, 20(18):1– 24, 2020. 1

  70. [78]

    Mark D. Sutton. Tomographic techniques for the study of exceptionally preserved fossils. Proceedings of the Royal Society B: Biological Sciences, 275(1643):1587–1593, 2008. 1

  71. [79]

    EfficientNet: Rethinking model scaling for convolutional neural networks

    Mingxing Tan and Quoc Le. EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning , pages 6105–6114. PMLR, 2019. 2

  72. [80]

    Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh

    Yucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth, Bennett Landman, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh. Self-supervised pre-training of swin trans- formers for 3D medical image analysis. In Proceedings of Computer Vision and Pattern Recognition (CVPR) , pages 20698...

  73. [81]

    Timmins, Irene C

    Kimberley M. Timmins, Irene C. van der Schaaf, Edwin Bennink, Ynte M. Ruigrok, Xingle An, Michael Baumgart- ner, Pascal Bourdon, Riccardo De Feo, Tommaso Di Noto, Florian Dubost, Augusto Fava-Sanches, Xue Feng, Corentin Giroud, Inteneural Group, Minghui Hu, Paul F. Jaeger, Juh...

  74. [82]

    vit-pytorch

    vit-pytorch maintainers and contributors. vit-pytorch. https : / / github . com / lucidrains / vit - pytorch, 2020. 5

  75. [83]

    C. L. Walsh, P. Tafforeau, W. L. Wagner, D. J. Jafree, A. Bel- lier, C. Werlein, M. P. K¨uhnel, E. Boller, S. Walker-Samuel, J. L. Robertus, D. A. Long, J. Jacob, S. Marussi, E. Brown, N. Holroyd, D. D. Jonigk, M. Ackermann, and P. D. Lee. Imaging intact human organs with loca...

  76. [84]

    Visualizing 3D food mi- crostructure using tomographic methods: Advantages and disadvantages

    Zi Wang, Els Herremans, Siem Janssen, Dennis Cantre, Pieter Verboven, and Bart Nicola¨ı. Visualizing 3D food mi- crostructure using tomographic methods: Advantages and disadvantages. Annual Review of Food Science and Tech- nology, 9(1):323–343, 2018. 2

  77. [85]

    Mapillary street-level sequences: A dataset for lifelong place recognition

    Frederik Warburg, Soren Hauberg, Manuel Lopez- Antequera, Pau Gargallo, Yubin Kuang, and Javier Civera. Mapillary street-level sequences: A dataset for lifelong place recognition. In Proceedings of Computer Vision and Pattern Recognition (CVPR) , pages 2626–2635. IEEE, 2020. 3

  78. [86]

    Meyer, Maurice Pradella, Daniel Hinck, Alexander W

    Jakob Wasserthal, Hanns Christian Breit, Manfred T. Meyer, Maurice Pradella, Daniel Hinck, Alexander W. Sauter, To- bias Heye, Daniel T. Boll, Joshy Cyriac, Shan Yang, Michael Bach, and Martin Segeroth. TotalSegmentator: Robust seg- mentation of 104 anatomic structures in CT i...

  79. [87]

    Hoi, and Qianru Sun

    Xiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim, Steven C.H. Hoi, and Qianru Sun. A large-scale benchmark for food im- age segmentation. In Proceedings of the 29th ACM Inter- national Conference on Multimedia , pages 506–515. ACM,

  80. [88]

    Fashion- MNIST: a novel image dataset for benchmarking machine learning algorithms

    Han Xiao, Kashif Rasul, and Roland V ollgraf. Fashion- MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv:1708.07747 [cs.LG], 2017. 3

  81. [89]

    Alvarez, and Ping Luo

    Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. SegFormer: Simple and efficient design for semantic segmentation with transform- ers. Advances in Neural Information Processing Systems 34 (neurips 2021), 34, 2021. 8

  82. [90]

    Aggregated residual transformations for deep neural networks

    Saining Xie, Ross Girshick, Piotr Doll ´ar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proceedings of Computer Vision and Pattern Recognition (CVPR), pages 5987–5995. IEEE, 2017. 2

  83. [91]

    CoTr: Efficiently bridging CNN and transformer for 3D medical image segmentation

    Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Yong Xia. CoTr: Efficiently bridging CNN and transformer for 3D medical image segmentation. In Medical Image Computing and Computer Assisted Intervention (MICCAI 2021), pages 171–180. Lecture Notes in Computer Science, V ol. 12903....

  84. [92]

    MedM- NIST v2 - a large-scale lightweight benchmark for 2d and 3d biomedical image classification

    Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. MedM- NIST v2 - a large-scale lightweight benchmark for 2d and 3d biomedical image classification. Scientific Data, 10(1),

  85. [93]

    3d deep learning for efficient and ro- bust landmark detection in volumetric data

    Yefeng Zheng, David Liu, Bogdan Georgescu, Hien Nguyen, and Dorin Comaniciu. 3d deep learning for efficient and ro- bust landmark detection in volumetric data. In Medical Im- age Computing and Computer-Assisted Intervention (MIC- CAI), pages 565–572. Springer, 2015. 2

  86. [94]

    nn- Former: V olumetric medical image segmentation via a 3D transformer

    Hong Yu Zhou, Jiansen Guo, Yinghao Zhang, Xiaoguang Han, Lequan Yu, Liansheng Wang, and Yizhou Yu. nn- Former: V olumetric medical image segmentation via a 3D transformer. IEEE Transactions on Image Processing , 32: 4036–4045, 2023. 2

  87. [95]

    Kevin Zhou, Hayit Greenspan, Christos Davatzikos, James S

    S. Kevin Zhou, Hayit Greenspan, Christos Davatzikos, James S. Duncan, Bram Van Ginneken, Anant Madabhushi, Jerry L. Prince, Daniel Rueckert, and Ronald M. Summers. A review of deep learning in medical imaging: Imaging 13 traits, technology trends, case studies with progress hi...

  88. [96]

    Scans In Fig

    Data visualization 8.1. Scans In Fig. 2, we introduce a set of example scan slices that il- lustrate the structural variability across the cheese samples. Subsequently, in Fig. 4c, we use slices from all scans to ex- plore the representation learned by one of the models. To pr...

  89. [97]

    Learning rate fine tuning In Sec

    Experiments 9.1. Learning rate fine tuning In Sec. 4, we outline the experimental design, including the investigated models and ablation studies. The training setup across all models is standardized to ensure a fair comparison of the models. However, certain hyperparameters sh...

  90. [98]

    Random 90° rotation ( p = 0.5). 1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 Coarse-grained class Temperature Rotor speed Additive type -1.17 -1.17 -1.17 1.44 1.44 1.07 0.14 -1.17 0.14 0.14 1.44 -1.17 1.44 0.14 -0.80 0.14 1.44 -1.17 0.56 0.14 0.56 -0.80 -0....

  91. [99]

    Random flipping in X and Y axis ( p = 0.5)

  92. [100]

    Elements 1 and 3 were retained from the original pipeline

    Random rotation in a −30° to 30° range ( p = 0.5). Elements 1 and 3 were retained from the original pipeline. The combined set of transforms covers most of the possible structure orientations while minimizing information loss at the most extreme angles. The results of the stud...

Pith tools

Reviewed May 23, 2026 · model on record in the stance chip above.