Pith. sign in

REVIEW 55 references

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.18683 v1 pith:D7IUU6QL submitted 2025-06-23 cs.CV cs.AI

classification cs.CVcs.AI
keywords classificationsim-netpointfeaturesimageobjectarchitecturecloud
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Shape-Image Multimodal Network (SIM-Net), a novel 2D image classification architecture that integrates 3D point cloud representations inferred directly from RGB images. Our key contribution lies in a pixel-to-point transformation that converts 2D object masks into 3D point clouds, enabling the fusion of texture-based and geometric features for enhanced classification performance. SIM-Net is particularly well-suited for the classification of digitized herbarium specimens (a task made challenging by heterogeneous backgrounds), non-plant elements, and occlusions that compromise conventional image-based models. To address these issues, SIM-Net employs a segmentation-based preprocessing step to extract object masks prior to 3D point cloud generation. The architecture comprises a CNN encoder for 2D image features and a PointNet-based encoder for geometric features, which are fused into a unified latent space. Experimental evaluations on herbarium datasets demonstrate that SIM-Net consistently outperforms ResNet101, achieving gains of up to 9.9% in accuracy and 12.3% in F-score. It also surpasses several transformer-based state-of-the-art architectures, highlighting the benefits of incorporating 3D structural reasoning into 2D image classification tasks.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 19 canonical work pages

  1. [1]

    Complete Image Point cloud Dataset (CIPD): Combines the full set of original unsegmented images with the point clouds generated from their corresponding segmented images

  2. [2]

    AlexNet, [2] introduced key innovations such as the ReLU activation function, dropout layers to mitigate overfitting, and overlapping pooling

    Related work 2D image classification: 2D image classification has advanced considerably over the last decade. AlexNet, [2] introduced key innovations such as the ReLU activation function, dropout layers to mitigate overfitting, and overlapping pooling. VGGNet [3], emphasized the depth, utilizi ng 16 convolutional layers. GoogLeNet/Inception [4] introduced...

  3. [3]

    The core of our approach involves transforming 2D images into 3D point clouds that represent the targeted objects

    Approach Our goal is to improve image classification by fusing 2D images and 3D point cloud data in a multimodal architecture to identify the characteristics of objects of the same nature, such as herbarium images. The core of our approach involves transforming 2D images into 3D point clouds that represent the targeted objects. To achieve this, we first u...

  4. [4]

    The first se t of experiments ( cf

    Experiments In this section, we present a set of experiments designed to evaluate the effectiveness of the PointNet, PointNet++ and SIM -Net models using the point clouds derived from herbarium 2D images, with ResNet serving as the baseline for comparison. The first se t of experiments ( cf. Table 2 and Table 3) assesses the performance of PointNet [7] an...

  5. [5]

    Complete Segmented Image Point cloud Dataset (CSIPD): Pairs the entire collection of segmented images with the point clouds derived from the same segmented images

  6. [6]

    Selected Segmented Image Point cloud Dataset (SSIPD) : Merges the selected segmented images with the point clouds generated from these accurately segmented images

  7. [7]

    Additionally, we generated two enriched point cloud datasets for each trait–one within the selected series and one within the complete series

    Selected Image Point cloud Dataset (SIPD) : Consists of the selected set of original unsegmented images that correspond to the well -segmented images, along with their associated point clouds from the well-segmented counterparts. Additionally, we generated two enriched point cloud datasets for each trait–one within the selected series and one within the c...

  8. [8]

    Conclusion The foundation of our approach lies in transforming the same dataset into multiple representations to leverage the strengths of different architectures for feature extraction. In this study, we convert 2D herbarium images into point clouds, enabling the use of a multimodal architecture that combines a convolutional neural network for detailed p...

Show all 55 references
  1. [9]

    W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I

    Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning Transferable Visual Models From Natural 24 Language Supervision. arXiv preprint arXiv:2103.00020. Retrieved fro...

  2. [10]

    LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient -based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324. https://doi.org/10.1109/5.726791

  3. [11]

    Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2017). ImageNet classification with deep convolutional neural networks. Communications of the ACM, 60(6), 84–90. https://doi.org/10.1145/3065386

  4. [12]

    Simonyan, K., & Zisserman, A. (2015). Very Deep Convolutional Networks for Large -Scale Image Recognition. arXiv preprint arXiv:1409.1556. Retrieved from https://arxiv.org/abs/1409.1556

  5. [13]

    Going deeper with convolutions

    Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. Going deeper with convolutions. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Los Alamitos, CA, USA: IEEE Computer Society; 2015. p. 1–9. https://doi.ieeecomputersociety.org/10.1109/CV...

  6. [14]

    He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770 –778). Las Vegas, NV, USA. https://doi.org/10.1109/CVPR.2016.90

  7. [15]

    N., Kaiser, L., & Polosukhin, I

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2023). Attention Is All You Need. arXiv preprint arXiv:1706.03762. Retrieved from https://arxiv.org/abs/1706.03762

  8. [16]

    Q., Su, H., Kaichun, M., & Guibas, L

    Charles, R. Q., Su, H., Kaichun, M., & Guibas, L. J. (2017). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 77–85). Honolulu, HI, USA. https://doi.org/10...

  9. [17]

    -W., Lee, K., & Toutanova, K

    Devlin, J., Chang, M. -W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805. Retrieved from https://arxiv.org/abs/1810.04805

  10. [18]

    Moayeri, M., Pope, P., Balaji, Y., & Feizi, S. (2022). A comprehensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022) (pp. 1906...

  11. [19]

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:201...

  12. [20]

    D., Ellwood, E

    Lorieul, T., Pearson, K. D., Ellwood, E. R., Goëau, H., Molino, J.-F., Sweeney, P. W., Yost, J. M., Sachs, J., Mata-Montero, E., Nelson, G., Soltis, P. S., Bonnet, P., & Joly, A. (2019). Toward a large -scale and deep phenological stage annotation of herbarium specimens: Case ...

  13. [21]

    Biodiversity Data Journal 8: e57090

    Younis S, Schmidt M, Weiland C, Dressler S, Seeger B, Hickler T (2020) Detection and annotation of plant organs from digitised herbarium scans using deep learning. Biodiversity Data Journal 8: e57090. https://doi.org/10.3897/BDJ.8.e57090

  14. [22]

    Abdelaziz, A., Bassem, B., & Walid, W. (2022). A deep learning -based approach for detecting plant organs from digitized herbarium specimen images. Ecological Informatics, 69, 101590. https://doi.org/10.1016/j.ecoinf.2022.101590

  15. [23]

    Biodiversity Information Science and Standards 8: e135629

    Sklab Y, Ariouat H, Boujydah Y, Qacami Y, Prifti E, Zucker J-daniel, Vignes Lebbe R, Chenin E (2024) Towards a Deep Learning -Powered Herbarium Image Analysis Platform. Biodiversity Information Science and Standards 8: e135629. https://doi.org/10.3897/biss.8.135629

  16. [24]

    Sklab, Y., Ariouat, H., Prifti, E., Zucker, J.-D., & Chenin, E. (2025). Identification of non-plant elements in herbarium images using YOLO. In 17th African Conference on Research in Computer Science and Applied Mathematics, CARI 2024, Bejaïa, Algeria, November 24–26, 2024. ht...

  17. [25]

    Sklab, E

    Ariouat, H., Y. Sklab, E. Prifti, J. -D. Zucker, and E. Chenin. 2025. Enhancing plant morphological trait identification in herbarium collections through deep learning –based segmentation. Applications in Plant Sciences 13(2): e70000. https://doi.org/10.1002/aps3.70000

  18. [26]

    (2024) Enhancing YOLOv7 for plant organs detection using attention-gate mechanism

    Ariouat H, Sklab Y, Pignal M, Jabbour F, Lebbe RV, Prifti E, et al. (2024) Enhancing YOLOv7 for plant organs detection using attention-gate mechanism. In: Pacific-Asia Conference on Knowledge Discovery and Data Mining Curran Associates. https://doi.org/10.1007/978-981-97-2253-2_18

  19. [27]

    Guo, R., Li, D., & Han, Y. (2021). Deep multi -scale and multi -modal fusion for 3D object detection. Pattern Recognition Letters, 151, 236–242. https://doi.org/10.1016/j.patrec.2021.08.028 25

  20. [28]

    T., Zhao, J., & Itti, L

    Leksut, J. T., Zhao, J., & Itti, L. (2020). Learning visual variation for object recognition. Image and Vision Computing, 98, 103912. https://doi.org/10.1016/j.imavis.2020.103912

  21. [29]

    Sahraoui, M., Sklab, Y., Pignal, M., Vignes Lebbe, R., & Guigue, V. (2023). Leveraging multimodality for biodiversity data: Exploring joint representations of species descriptions and specimen images using CLIP. Biodiversity Information Science and Standards, 7, e112666. https...

  22. [30]

    Zhao, J., Wang, Y., Cao, Y., Guo, M., Huang, X., Zhang, R., Dou, X., Niu, X., Cui, Y., & Wang, J. (2021). The fusion strategy of 2D and 3D information based on deep learning: A review. Remote Sensing, 13(20),

  23. [31]

    S., Do, N.-T., Kim, S.-H., Yang, H.-J., & Lee, G.-S

    [31] Ly, T. S., Do, N.-T., Kim, S.-H., Yang, H.-J., & Lee, G.-S. (2019). A novel 2D and 3D multimodal approach for in -the-wild facial expression recognition. Image and Vision Computing, 92, 103817. https://doi.org/10.1016/j.imavis.2019.10.003

  24. [32]

    Huang, G., Liu, Z., Van Der Maaten , L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017) (pp. 4700–4708). IEEE. https://doi.org/10.1109/CVPR.2017.243

  25. [33]

    R., Yi, L., Su, H., & Guibas, L

    Qi, C. R., Yi, L., Su, H., & Guibas, L. J. (2017). PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS 2017) (pp. 5105–5114)

  26. [34]

    Engel, N., Belagiannis, V., & Dietmayer, K. (2021). Point Transformer. IEEE Access, 9, 134826–134840. https://doi.org/10.1109/ACCESS.2021.3116304

  27. [35]

    Ma, X., Qin, C., You, H., Ran, H., & Fu, Y. (2022). Rethinking network design and local geometry in point cloud: A simple residual MLP framework. In International Conference on Learning Representations (ICLR 2022). https://openreview.net/forum?id=3Pbra-_u76D

  28. [36]

    A., Elhoseiny, M., & Ghanem, B

    Qian, G., Li, Y., Peng, H., Mai, J., Al Kader Hammoud, H. A., Elhoseiny, M., & Ghanem, B. (2022). PointNeXt: Revisiting PointNet++ with improved training and scaling strategies. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS 2022)

  29. [37]

    Chen, Y., Yang, B., Liang, M., & Urtasun, R. (2019). Learning joint 2D -3D representations for depth completion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 10022–10031). IEEE. https://doi.org/10.1109/ICCV.2019.01012

  30. [38]

    L., Qin, J., & Wei, M

    Gu, L., Yan, X., Cui, P., Gong, L., Xie, H., Wang, F. L., Qin, J., & Wei, M. (2024). PointSee: Image enhances point cloud. IEEE Transactions on Visualization and Computer Graphics, 30(6), 6291 –6308. https://doi.org/10.1109/TVCG.2023.3331779

  31. [39]

    Poliyapram, V., Wang, W., & Nakamura, R. (2019). A point -wise LiDAR and image multimodal fusion network (PMNet) for aerial point cloud 3D semantic segmentation. Remote Sensing, 11(24), 2961. https://doi.org/10.3390/rs11242961

  32. [40]

    Pradelle, O., Chaine, R., Wendland, D., & Digne, J. (2023). Lightweight integration of 3D features to improve 2D image segmentation. Computers & Graphics, 114, 326 –336. https://doi.org/10.1016/j.cag.2023.06.004

  33. [41]

    -Y., Feichtenhofer, C., Darrell, T., & Xie, S

    Liu, Z., Mao, H., Wu, C. -Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A ConvNet for the 2020s. arXiv preprint arXiv:2201.03545. https://doi.org/10.48550/arXiv.2201.03545

  34. [42]

    [32] Ronneberger, O., Fischer, P., & Brox, T. (2015). U -Net: Convolutional networks for biomedical image segmentation. arXiv preprint arXiv:1505.04597. https://doi.org/10.48550/arXiv.1505.04597

  35. [43]

    -D., Prifti, E., & Chenin, E

    [33] Ariouat, H., Sklab, Y., Pignal, M., Vignes Lebbe, R., Zucker, J. -D., Prifti, E., & Chenin, E. (2023). Extracting masks from herbarium specimen images based on object detection and image segmentation techniques. Biodiversity Information Science an d Standards, 7, e112161....

  36. [44]

    Yang, C.-K., Chen, M.-H., Chuang, Y.-Y., & Lin, Y.-Y. (2023). 2D-3D interlaced transformer for point cloud segmentation with scene -level supervision. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 977 –987). IEEE. https://doi.org/1...

  37. [45]

    Zhao, L., Lu, J., & Zhou, J. (2021). Similarity-aware fusion network for 3D semantic segmentation. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 1585–1592). IEEE. https://doi.org/10.1109/IROS51168.2021.9636494

  38. [46]

    Palladin, E., Dietze, R., Narayanan, P., Bijelic, M., & Heide, F. (2024). SAMFusion: Sensor-adaptive multimodal fusion for 3D object detection in adverse weather. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings,...

  39. [48]

    Ge, Z., Cao, G., Li, X., & Fu, P. (2020). Hyperspectral image classification method based on 2D –3D CNN and multibranch feature fusion. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13, 5776–5788. https://doi.org/10.1109/JSTARS.2020.3024841

  40. [49]

    Bao, H., Dong, L., Piao, S., & Wei, F. (2022). BEiT: BERT pre -training of image transformers. arXiv preprint arXiv:2106.08254. https://doi.org/10.48550/arXiv.2106.08254

  41. [50]

    V., & Tan, M

    Dai, Z., Liu, H., Le, Q. V., & Tan, M. (2021). CoAtNet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803. https://doi.org/10.48550/arXiv.2106.04803

  42. [52]

    S., & Xie, S

    Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., & Xie, S. (2023). ConvNeXt V2: Co - designing and scaling ConvNets with masked autoencoders. arXiv preprint arXiv:2301.00808. https://doi.org/10.48550/arXiv.2301.00808

  43. [53]

    L., Morcos, A

    d’Ascoli, S., Touvron, H., Leavitt, M. L., Morcos, A. S., Biroli, G., & Sagun, L. (2022). ConViT: Improving vision transformers with soft convolutional inductive biases. Journal of Statistical Mechanics: Theory and Experiment, 2022(11), 114005. https://doi.org/10.1088/1742-5468/ac9830

  44. [54]

    Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030. https://doi.org/10.48550/arXiv.2103.14030

  45. [55]

    Ding, M., Xiao, B., Codella, N., Luo, P., Wang, J., & Yuan, L. (2022). DaViT: Dual attention vision transformers. arXiv preprint arXiv:2204.03645. https://doi.org/10.48550/arXiv.2204.03645

  46. [128]

    Color -coded coordinate masks

    Each model underwent training for 100 epochs, employing the Adam optimizer with an initial learning rate of 0.001 and a weight decay set at 1e -4. A StepLR scheduler wa s applied, systematically reducing the learning rate by 30% every 20 epochs. The cross -entropy loss functio...

  47. [4029]

    https://doi.org/10.3390/rs13204029

Pith tools