REVIEW 55 references
SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce the Shape-Image Multimodal Network (SIM-Net), a novel 2D image classification architecture that integrates 3D point cloud representations inferred directly from RGB images. Our key contribution lies in a pixel-to-point transformation that converts 2D object masks into 3D point clouds, enabling the fusion of texture-based and geometric features for enhanced classification performance. SIM-Net is particularly well-suited for the classification of digitized herbarium specimens (a task made challenging by heterogeneous backgrounds), non-plant elements, and occlusions that compromise conventional image-based models. To address these issues, SIM-Net employs a segmentation-based preprocessing step to extract object masks prior to 3D point cloud generation. The architecture comprises a CNN encoder for 2D image features and a PointNet-based encoder for geometric features, which are fused into a unified latent space. Experimental evaluations on herbarium datasets demonstrate that SIM-Net consistently outperforms ResNet101, achieving gains of up to 9.9% in accuracy and 12.3% in F-score. It also surpasses several transformer-based state-of-the-art architectures, highlighting the benefits of incorporating 3D structural reasoning into 2D image classification tasks.
Reference graph
Works this paper leans on
-
[1]
Complete Image Point cloud Dataset (CIPD): Combines the full set of original unsegmented images with the point clouds generated from their corresponding segmented images
-
[2]
AlexNet, [2] introduced key innovations such as the ReLU activation function, dropout layers to mitigate overfitting, and overlapping pooling
Related work 2D image classification: 2D image classification has advanced considerably over the last decade. AlexNet, [2] introduced key innovations such as the ReLU activation function, dropout layers to mitigate overfitting, and overlapping pooling. VGGNet [3], emphasized the depth, utilizi ng 16 convolutional layers. GoogLeNet/Inception [4] introduced...
-
[3]
The core of our approach involves transforming 2D images into 3D point clouds that represent the targeted objects
Approach Our goal is to improve image classification by fusing 2D images and 3D point cloud data in a multimodal architecture to identify the characteristics of objects of the same nature, such as herbarium images. The core of our approach involves transforming 2D images into 3D point clouds that represent the targeted objects. To achieve this, we first u...
-
[4]
The first se t of experiments ( cf
Experiments In this section, we present a set of experiments designed to evaluate the effectiveness of the PointNet, PointNet++ and SIM -Net models using the point clouds derived from herbarium 2D images, with ResNet serving as the baseline for comparison. The first se t of experiments ( cf. Table 2 and Table 3) assesses the performance of PointNet [7] an...
-
[5]
Complete Segmented Image Point cloud Dataset (CSIPD): Pairs the entire collection of segmented images with the point clouds derived from the same segmented images
-
[6]
Selected Segmented Image Point cloud Dataset (SSIPD) : Merges the selected segmented images with the point clouds generated from these accurately segmented images
-
[7]
Additionally, we generated two enriched point cloud datasets for each trait–one within the selected series and one within the complete series
Selected Image Point cloud Dataset (SIPD) : Consists of the selected set of original unsegmented images that correspond to the well -segmented images, along with their associated point clouds from the well-segmented counterparts. Additionally, we generated two enriched point cloud datasets for each trait–one within the selected series and one within the c...
-
[8]
Conclusion The foundation of our approach lies in transforming the same dataset into multiple representations to leverage the strengths of different architectures for feature extraction. In this study, we convert 2D herbarium images into point clouds, enabling the use of a multimodal architecture that combines a convolutional neural network for detailed p...
2023
Show all 55 references
-
[9]
W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural 24 Language Supervision. arXiv preprint arXiv:2103.00020. Retrieved fro...
2021 arXiv
-
[10]
LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient -based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324. https://doi.org/10.1109/5.726791
1998 doi
-
[11]
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2017). ImageNet classification with deep convolutional neural networks. Communications of the ACM, 60(6), 84–90. https://doi.org/10.1145/3065386
2017 doi
-
[12]
Simonyan, K., & Zisserman, A. (2015). Very Deep Convolutional Networks for Large -Scale Image Recognition. arXiv preprint arXiv:1409.1556. Retrieved from https://arxiv.org/abs/1409.1556
2015 arXiv
-
[13]
Going deeper with convolutions
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. Going deeper with convolutions. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Los Alamitos, CA, USA: IEEE Computer Society; 2015. p. 1–9. https://doi.ieeecomputersociety.org/10.1109/CV...
2015
-
[14]
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770 –778). Las Vegas, NV, USA. https://doi.org/10.1109/CVPR.2016.90
2016 doi
-
[15]
N., Kaiser, L., & Polosukhin, I
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2023). Attention Is All You Need. arXiv preprint arXiv:1706.03762. Retrieved from https://arxiv.org/abs/1706.03762
2023 arXiv
-
[16]
Q., Su, H., Kaichun, M., & Guibas, L
Charles, R. Q., Su, H., Kaichun, M., & Guibas, L. J. (2017). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 77–85). Honolulu, HI, USA. https://doi.org/10...
2017 doi
-
[17]
-W., Lee, K., & Toutanova, K
Devlin, J., Chang, M. -W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805. Retrieved from https://arxiv.org/abs/1810.04805
2019 arXiv
-
[18]
Moayeri, M., Pope, P., Balaji, Y., & Feizi, S. (2022). A comprehensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) (pp. 1906...
2022
-
[19]
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:201...
2021 arXiv
-
[20]
D., Ellwood, E
Lorieul, T., Pearson, K. D., Ellwood, E. R., Goëau, H., Molino, J.-F., Sweeney, P. W., Yost, J. M., Sachs, J., Mata-Montero, E., Nelson, G., Soltis, P. S., Bonnet, P., & Joly, A. (2019). Toward a large -scale and deep phenological stage annotation of herbarium specimens: Case ...
2019 doi
-
[21]
Biodiversity Data Journal 8: e57090
Younis S, Schmidt M, Weiland C, Dressler S, Seeger B, Hickler T (2020) Detection and annotation of plant organs from digitised herbarium scans using deep learning. Biodiversity Data Journal 8: e57090. https://doi.org/10.3897/BDJ.8.e57090
2020 doi
-
[22]
Abdelaziz, A., Bassem, B., & Walid, W. (2022). A deep learning -based approach for detecting plant organs from digitized herbarium specimen images. Ecological Informatics, 69, 101590. https://doi.org/10.1016/j.ecoinf.2022.101590
2022
-
[23]
Biodiversity Information Science and Standards 8: e135629
Sklab Y, Ariouat H, Boujydah Y, Qacami Y, Prifti E, Zucker J-daniel, Vignes Lebbe R, Chenin E (2024) Towards a Deep Learning -Powered Herbarium Image Analysis Platform. Biodiversity Information Science and Standards 8: e135629. https://doi.org/10.3897/biss.8.135629
2024 doi
-
[24]
Sklab, Y., Ariouat, H., Prifti, E., Zucker, J.-D., & Chenin, E. (2025). Identification of non-plant elements in herbarium images using YOLO. In 17th African Conference on Research in Computer Science and Applied Mathematics, CARI 2024, Bejaïa, Algeria, November 24–26, 2024. ht...
2025 doi
-
[25]
Sklab, E
Ariouat, H., Y. Sklab, E. Prifti, J. -D. Zucker, and E. Chenin. 2025. Enhancing plant morphological trait identification in herbarium collections through deep learning –based segmentation. Applications in Plant Sciences 13(2): e70000. https://doi.org/10.1002/aps3.70000
2025 doi
-
[26]
(2024) Enhancing YOLOv7 for plant organs detection using attention-gate mechanism
Ariouat H, Sklab Y, Pignal M, Jabbour F, Lebbe RV, Prifti E, et al. (2024) Enhancing YOLOv7 for plant organs detection using attention-gate mechanism. In: Pacific-Asia Conference on Knowledge Discovery and Data Mining Curran Associates. https://doi.org/10.1007/978-981-97-2253-2_18
2024 doi
-
[27]
Guo, R., Li, D., & Han, Y. (2021). Deep multi -scale and multi -modal fusion for 3D object detection. Pattern Recognition Letters, 151, 236–242. https://doi.org/10.1016/j.patrec.2021.08.028 25
2021 doi
-
[28]
T., Zhao, J., & Itti, L
Leksut, J. T., Zhao, J., & Itti, L. (2020). Learning visual variation for object recognition. Image and Vision Computing, 98, 103912. https://doi.org/10.1016/j.imavis.2020.103912
2020
-
[29]
Sahraoui, M., Sklab, Y., Pignal, M., Vignes Lebbe, R., & Guigue, V. (2023). Leveraging multimodality for biodiversity data: Exploring joint representations of species descriptions and specimen images using CLIP. Biodiversity Information Science and Standards, 7, e112666. https...
2023 doi
-
[30]
Zhao, J., Wang, Y., Cao, Y., Guo, M., Huang, X., Zhang, R., Dou, X., Niu, X., Cui, Y., & Wang, J. (2021). The fusion strategy of 2D and 3D information based on deep learning: A review. Remote Sensing, 13(20),
2021
-
[31]
S., Do, N.-T., Kim, S.-H., Yang, H.-J., & Lee, G.-S
[31] Ly, T. S., Do, N.-T., Kim, S.-H., Yang, H.-J., & Lee, G.-S. (2019). A novel 2D and 3D multimodal approach for in -the-wild facial expression recognition. Image and Vision Computing, 92, 103817. https://doi.org/10.1016/j.imavis.2019.10.003
2019 doi
-
[32]
Huang, G., Liu, Z., Van Der Maaten , L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017) (pp. 4700–4708). IEEE. https://doi.org/10.1109/CVPR.2017.243
2017 doi
-
[33]
R., Yi, L., Su, H., & Guibas, L
Qi, C. R., Yi, L., Su, H., & Guibas, L. J. (2017). PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS 2017) (pp. 5105–5114)
2017
-
[34]
Engel, N., Belagiannis, V., & Dietmayer, K. (2021). Point Transformer. IEEE Access, 9, 134826–134840. https://doi.org/10.1109/ACCESS.2021.3116304
2021
-
[35]
Ma, X., Qin, C., You, H., Ran, H., & Fu, Y. (2022). Rethinking network design and local geometry in point cloud: A simple residual MLP framework. In International Conference on Learning Representations (ICLR 2022). https://openreview.net/forum?id=3Pbra-_u76D
2022
-
[36]
A., Elhoseiny, M., & Ghanem, B
Qian, G., Li, Y., Peng, H., Mai, J., Al Kader Hammoud, H. A., Elhoseiny, M., & Ghanem, B. (2022). PointNeXt: Revisiting PointNet++ with improved training and scaling strategies. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS 2022)
2022
-
[37]
Chen, Y., Yang, B., Liang, M., & Urtasun, R. (2019). Learning joint 2D -3D representations for depth completion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 10022–10031). IEEE. https://doi.org/10.1109/ICCV.2019.01012
2019
-
[38]
L., Qin, J., & Wei, M
Gu, L., Yan, X., Cui, P., Gong, L., Xie, H., Wang, F. L., Qin, J., & Wei, M. (2024). PointSee: Image enhances point cloud. IEEE Transactions on Visualization and Computer Graphics, 30(6), 6291 –6308. https://doi.org/10.1109/TVCG.2023.3331779
2024
-
[39]
Poliyapram, V., Wang, W., & Nakamura, R. (2019). A point -wise LiDAR and image multimodal fusion network (PMNet) for aerial point cloud 3D semantic segmentation. Remote Sensing, 11(24), 2961. https://doi.org/10.3390/rs11242961
2019 doi
-
[40]
Pradelle, O., Chaine, R., Wendland, D., & Digne, J. (2023). Lightweight integration of 3D features to improve 2D image segmentation. Computers & Graphics, 114, 326 –336. https://doi.org/10.1016/j.cag.2023.06.004
2023 doi
- [41]
- [42]
-
[43]
-D., Prifti, E., & Chenin, E
[33] Ariouat, H., Sklab, Y., Pignal, M., Vignes Lebbe, R., Zucker, J. -D., Prifti, E., & Chenin, E. (2023). Extracting masks from herbarium specimen images based on object detection and image segmentation techniques. Biodiversity Information Science an d Standards, 7, e112161....
2023 doi
-
[44]
Yang, C.-K., Chen, M.-H., Chuang, Y.-Y., & Lin, Y.-Y. (2023). 2D-3D interlaced transformer for point cloud segmentation with scene -level supervision. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 977 –987). IEEE. https://doi.org/1...
2023
-
[45]
Zhao, L., Lu, J., & Zhou, J. (2021). Similarity-aware fusion network for 3D semantic segmentation. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 1585–1592). IEEE. https://doi.org/10.1109/IROS51168.2021.9636494
2021
-
[46]
Palladin, E., Dietze, R., Narayanan, P., Bijelic, M., & Heide, F. (2024). SAMFusion: Sensor-adaptive multimodal fusion for 3D object detection in adverse weather. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings,...
2024 doi
-
[48]
Ge, Z., Cao, G., Li, X., & Fu, P. (2020). Hyperspectral image classification method based on 2D –3D CNN and multibranch feature fusion. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13, 5776–5788. https://doi.org/10.1109/JSTARS.2020.3024841
2020
- [49]
- [50]
- [52]
-
[53]
L., Morcos, A
d’Ascoli, S., Touvron, H., Leavitt, M. L., Morcos, A. S., Biroli, G., & Sagun, L. (2022). ConViT: Improving vision transformers with soft convolutional inductive biases. Journal of Statistical Mechanics: Theory and Experiment, 2022(11), 114005. https://doi.org/10.1088/1742-5468/ac9830
2022 doi
- [54]
- [55]
-
[128]
Color -coded coordinate masks
Each model underwent training for 100 epochs, employing the Adam optimizer with an initial learning rate of 0.001 and a weight decay set at 1e -4. A StepLR scheduler wa s applied, systematically reducing the learning rate by 30% every 20 epochs. The cross -entropy loss functio...
-
[4029]
https://doi.org/10.3390/rs13204029
Discussion (0). Continue with ORCID to comment.