REVIEW 3 major objections 5 minor 67 references
The paper claims that a single lightweight architecture, CURE, fuses arbitrarily many heterogeneous medical modalities—images, multi-omics, clinical records, and wearable time series—into one shared representation, and that this representat
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 13:28 UTC pith:4FWWOFQL
load-bearing objection A lightweight fusion architecture with real paired-regime value, but the headline unpaired results are multi-task gains, not evidence of cross-modal fusion. the 3 major comments →
Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
CURE is presented as a state-of-the-art multimodal fusion framework built on a sequential fusion loop. Each HyFuse layer takes the current shared representation and the next modality, refines both with an Efficient Multimodal Residual Convolution (EMRC) block, passes them through a Hybrid-Space Aware Attention Mixer (HySAM) that computes attention in Poincaré and Lorentz hyperbolic spaces and a quantum-inspired space simultaneously, then fuses them with learnable gates (LLF) and refines them again (SIR). Because modalities are fused one at a time rather than all-pairs, the cost scales linearly with the number of modalities. The paper reports that CURE variants outperform sixteen baselines on
What carries the argument
HyFuse (Hybrid Geometry Aware Fusion) layer, the modular fusion cell that carries the argument. It contains four components: EMRC (multi-scale depthwise/pointwise convolutions for cheap feature extraction), HySAM (a hybrid-space attention mixer combining Multimodal Hyperbolic Dual-Geometry Attention over Poincaré and Lorentz models with Multimodal Quantum-Inspired Attention, fused by learnable gating), LLF (learnable late fusion that masks missing modalities), and SIR (shared information refinement). The key structural claim is that cascading these layers sequentially—rather than fusing all modalities in parallel—produces modality-order-invariant shared representations at linear cost.
Load-bearing premise
The empirical case assumes that fusing datasets with no subject-level alignment (e.g., skin images from one cohort, multi-omics from another, EHR from a third) is a meaningful test of multimodal fusion; if the unpaired gains mainly reflect multi-task regularization, the central generalization claim would not transfer to genuinely paired medical data.
What would settle it
Train CURE on the unpaired groups after randomly permuting the sample indices within each modality stream. If accuracy and AUC do not drop materially, cross-modal alignment is not doing the work, and the reported fusion gains can be attributed to task regularization. Conversely, run the same architecture on a paired cohort (same patients contributing WSI and omics) and compare against a late-fusion baseline with identical backbones; the cross-modal attention should add a measurable margin.
If this is right
- CURE's claim implies that a single network can be extended to new medical modalities by appending one HyFuse layer, without retraining the fusion topology from scratch.
- The reported missing-modality results suggest that the same trained model can serve deployments where some data sources are unavailable, since the learnable gates zero out absent streams.
- The efficiency numbers (0.29–2.82 GFLOPs) imply that multimodal fusion becomes feasible on edge or resource-constrained clinical hardware, not just large GPU clusters.
- If the modality-order invariance is real, system builders do not need to enforce a canonical ordering of clinical inputs before training.
Where Pith is reading between the lines
- The unpaired evaluation protocol treats separate, unaligned datasets as 'modalities,' so the reported cross-modal gains may actually come from multi-task regularization across datasets; a paired subject-level benchmark would be needed to verify genuine multimodal fusion.
- The paper does not isolate whether the non-Euclidean (hyperbolic/quantum) machinery, rather than the cascaded fusion topology, drives the improvement; an equally parameterized Euclidean attention mixer would be a natural control.
- A testable extension: shuffle the sample indices across the unpaired streams; if CURE's performance is unchanged, its 'fusion' is not exploiting cross-modal correspondences, and the framework is better described as a unified multi-task learner.
- If the sequential design holds up, it points toward streaming and federated settings where modalities arrive over time, since the shared representation is built incrementally.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CURE, a lightweight cascaded multimodal fusion framework whose core is the HyFuse layer, comprising residual multi-scale convolutions (EMRC), a hybrid hyperbolic/quantum attention mixer (HySAM), learnable late fusion (LLF), and shared-information refinement (SIR). Modalities are processed sequentially and fused through HyFuse layers, yielding a shared representation x_C used for downstream classification, mortality, and survival tasks. The paper claims state-of-the-art performance across 16 public medical datasets, with up to ≈3.97% accuracy gains and ≈87.8% FLOPs reductions relative to prior multimodal fusion baselines. Evaluation includes unpaired settings, where each dataset is treated as an independent modality stream, and paired WSI+omics survival benchmarks on BLCA and KIRP, plus ablations, missing-modality tests, and resolution scaling.
Significance. The proposed modular architecture and the availability of code are strengths, as are the paired survival experiments and the detailed ablations in Tables 2(c) and 3. If the central claim of efficient and generalizable multimodal fusion were established, the work could be a practical contribution for resource-constrained medical AI. However, the primary evidence for cross-modal fusion rests on the unpaired protocol of Table 1, in which HAM10000, SIPaKMeD, TCGA-BRCA, and MIMIC-III are treated as four fused modality streams despite having no subject-level alignment. Under that protocol, the HyFuse cross-modal attention cannot operate on samples with actual cross-modal correspondence, so the reported gains are more plausibly attributable to multi-task regularization and dataset-specific heads than to multimodal fusion. The paired BLCA/KIRP results are a proper fusion test, but they involve only two datasets and the margins over MMLego are modest (≈2.5 and ≈1.5 C-index points). Thus the headline generalization claims are not yet supported.
major comments (3)
- [Sec. 3.1.2, Eqs. (8)–(10)] The unpaired protocol defines each dataset as an independent modality stream and then runs a single CURE pass over, e.g., HAM10000 images, SIPaKMeD images, TCGA-BRCA multi-omics, and MIMIC-III EHR. Because these datasets come from different cohorts/institutions and have no subject-level alignment, the HyFuse cross-modal attention (Eqs. 1–13) operates on randomly batched samples that have no semantic correspondence. The resulting shared representation x_C is therefore a multi-task shared trunk with dataset-specific heads, and the large margins in Table 1 (e.g., +1.42 ACC on HAM10000, +5.8 C-index on BRCA, +5.55 ACC on MORT) could reflect multi-task regularization, dataset-specific heads, or implicit augmentation rather than cross-modal fusion. The abstract's headline '≈3.97%' claim relies heavily on this table. Please re-frame the unpaired experiments as multi-task learning over independe
- [Sec. 4; Table 1] The HySAM module is under-specified. In Eq. (8), δ_i is described as a channel-wise bias obtained from MQIA, but Eq. (9) defines the MQIA output as a Softmax attention map; no equation or text maps A^Q_i to δ_i. Additionally, the index i is overloaded: the outer modality index and the inner summation index in Eq. (8) both use i, and the notation δ_{i,l} is not introduced. Eq. (9) also contains an ambiguous expression — the parentheses/operators around |q_i|^2, MLE(ψ_i), and h^{-2} are incomplete — and the hyperparameter h is never given a value. Because HySAM is the key novel component, these omissions prevent reproduction and make it difficult to assess whether the reported gains follow from the stated mechanism.
- [Sec. 4; Table 1] Table 1 reports no standard deviations despite the statement that all experiments use five random seeds. Many entries are at saturation (AUC 99.99 and 99.95; ACC 99.81), and the claimed margins over the strongest baseline are often small in absolute terms. Without variance estimates or significance testing, the reader cannot judge whether the reported improvements are reliable. In addition, App. E states that baselines originally designed for other tasks were adapted by removing decoders/heads and attaching task-specific heads; this adaptation protocol should be reported per baseline, since it can materially affect relative rankings.
minor comments (5)
- [Algorithm 1, line 10] In the MSIL loop, the else branch calls HyFuse(x^{S'}_i, x_{i+1}) when i=1, but x^{S'}_1 is not yet defined. The special case for i==2 suggests the intended first call uses (x_1, x_2); please fix the indexing in the pseudocode.
- [Sec. 4 dataset list vs. App. F] The paper says '16 public datasets', but Appendix F, Table 7 introduces KVASIR and MIT-BIH, which are not listed among D1–D16 in Sec. 4. Please reconcile the dataset count or clarify that these are additional validation datasets.
- [Abstract vs. Sec. 4.1] The abstract states a performance improvement of up to ≈3.97%, while Sec. 4.1 reports gains of up to ≈5.8% over the strongest competing method on Table 1. Please make the reported numbers consistent and define the exact aggregation used for the abstract's headline figure.
- [Eq. (10) block text] The MAFG block is referred to as 'MFAG' in the sentence preceding Eq. (10). Please correct the typo.
- [Sec. 3.1.2, Eq. (9)] The expression for the quantum attention weights is hard to parse. Please define each term explicitly, including whether the norm is Euclidean or Lorentzian, and what the argument of the Softmax is.
Circularity Check
No significant circularity: the paper's claims are empirical benchmark comparisons, not derivations that reduce to their own inputs.
full rationale
CURE is an architecture paper; its headline claims (Tables 1, 2) are comparisons against external baselines trained under a common protocol. I found no fitted parameter renamed as a prediction, no theorem whose conclusion is assumed in its premises, and no equation that reduces to its own input by construction. The HyFuse equations (Eqs. 1-15) are standard attention, gating, and residual-convolution modules; they define the architecture but do not derive any result from the data being evaluated. The unpaired MSIL protocol (Footnote 5; App. E) explicitly states that datasets have no subject-level alignment and are treated as independent modality streams; this is a legitimate validity concern about whether the unpaired setup tests cross-modal fusion versus multi-task regularization, but it is not circularity because the reported numbers are empirical outcomes rather than consequences of the protocol's definition. Self-citations such as DRIFA-Net [19] appear as baselines and related work, not as load-bearing justification for CURE's correctness or uniqueness. No self-citation chain forces the central claim, and no ansatz is smuggled in solely via the authors' prior work. Accordingly, the paper's derivation chain is self-contained in the sense required here; the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- learnable curvature c~ = c × (1/e) Σ σ(f_j), c = clip(e^k, 0.1, 10.0) =
trained k and f ∈ R^e
- quantum interaction hyperparameter h =
not reported
- LLF masking constant o =
not reported
- task-modality loss weights λ_t^M =
not reported
- modality-specific learnable scalars L, P, c_i, α_i, β_l, η_real, η_imag =
learned
axioms (6)
- ad hoc to paper Unrelated datasets from different cohorts can be treated as independent "modality streams" and fused by HyFuse without subject-level alignment
- domain assumption Non-imaging omics/EHR vectors can be reshaped into pseudo-image tensors R^{H×W×C} without loss of structure
- domain assumption Poincaré ball and Lorentz hyperboloid embedding formulas (Eqs. 4, 7) preserve the needed frequency structure for attention
- standard math The Born rule |q_i|² = q_i q*_i is applicable to the complex projection of DCT features for attention weighting
- domain assumption At inference at least one modality must be present
- domain assumption Four-fold label-preserving augmentation (rotation, translation, blur) does not distort downstream task labels
invented entities (2)
-
Quantum-inspired complex state q_i = ψ_i·(η_real + jη_imag)
no independent evidence
-
Learnable curvature c~ with fractal scaling weights
no independent evidence
read the original abstract
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively, which in turn limits performance improvements. Second, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI applications. Finally, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. To address these challenges, we propose a novel MFL framework - Cascaded Unified Representation Learning for Efficient Fusion Network (CURE) - a lightweight and scalable framework that progressively integrates various modalities through a novel efficient Hybrid Geometry Aware Fusion layer (HyFuse), where each HyFuse layer is sequentially learned for each modality, making the framework adaptable and generalizable. Within HyFuse, an efficient residual convolution module captures rich multi-scale features to ensure cost-effective learning, while a hybrid-space aware attention mixer learns coarse-to-fine structural cues to better preserve cross-modal relationships. Complementary learnable late-fusion and shared information refinement modules are then employed to learn robust modality-order-invariant shared representations, which in turn yields consistent performance improvements. Extensive evaluations on 16 public datasets show that CURE outperforms leading multimodal fusion methods, boosting performance by up to 3.97% and lowering computational costs by up to 87.8%, ensuring more effective and reliable predictions.
Figures
Reference graph
Works this paper leans on
-
[1]
I. S. A. Abdelhalim, M. F. Mohamed, and Y. B. Mahdy. 2021. Data augmentation for skin lesion using self-attention based progressive generative adversarial network. Expert Systems with Applications165 (2021), 113922
2021
-
[2]
Yajun An, Jiale Chen, Huan Lin, Zhenbing Liu, Siyang Feng, Hualong Zhang, Rushi Lan, Zaiyi Liu, and Xipeng Pan. 2025. CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Joy Dhar et al. Prediction. InProceedings of the AAAI Conference on Artificial I...
2025
-
[3]
Plamen Angelov and Eduardo Soares. 2020. Towards explainable deep neural networks (xDNN).Neural Networks130 (2020), 185–194
2020
-
[4]
Anguita, A
D. Anguita, A. Ghio, L. Oneto, X. Parra, and J.L. Reyes-Ortiz. 2013. A Public Domain Dataset for Human Activity Recognition Using Smartphones. In21st European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN)
2013
-
[5]
Ujjwal Baid, Satyam Ghodasara, Suyash Mohan, Michel Bilello, Evan Calabrese, Errol Colak, Keyvan Farahani, Jayashree Kalpathy-Cramer, Felipe C Kitamura, Sarthak Pati, et al . 2021. The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification.arXiv preprint arXiv:2107.02314(2021)
Pith/arXiv arXiv 2021
-
[6]
Oresti Banos, Rafael Garcia, Juan A Holgado-Terriza, Miguel Damas, Hector Pomares, Ignacio Rojas, Alejandro Saez, and Claudia Villalonga. 2014. mHealth- Droid: a novel framework for agile development of mobile health applications. In Ambient Assisted Living and Daily Activities: 6th International Work-Conference, IW AAL 2014, Belfast, UK, December 2-5, 20...
2014
-
[7]
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross- attention multi-scale vision transformer for image classification. InProceedings of the IEEE/CVF international conference on computer vision. 357–366
2021
-
[8]
Richard J Chen, Ming Y Lu, Jingwen Wang, Drew FK Williamson, Scott J Rodig, Neal I Lindeman, and Faisal Mahmood. 2020. Pathomic fusion: an integrated framework for fusing histopathology and genomic features for cancer diagnosis and prognosis.IEEE Transactions on Medical Imaging41, 4 (2020), 757–770
2020
-
[9]
Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2021. Fully Hyperbolic Neural Networks. arXiv:2105.14686 [cs.CL] doi:10.48550/arXiv.2105.14686 Submitted on 31 May 2021; last revised 16 Mar 2022 (this version, v3)
- [10]
-
[11]
Iris Cong, Soonwon Choi, and Mikhail D. Lukin. 2018. Quantum Convolutional Neural Networks. arXiv:1810.03787 [quant-ph] Revised version posted May 2, 2019
Pith/arXiv arXiv 2018
-
[12]
Can Cui, Haichun Yang, Yaohong Wang, Shilin Zhao, Zuhayr Asad, Lori A Coburn, Keith T Wilson, Bennett A Landman, and Yuankai Huo. 2023. Deep multimodal fusion of image and non-image data in disease diagnosis and prognosis: a review. Progress in Biomedical Engineering5, 2 (2023), 022001
2023
-
[13]
Y. Cui, Y. Tao, W. Ren, and A. Knoll. 2023. Dual-domain attention for image deblurring. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 479–487
2023
-
[14]
Joy Dhar, Puneet Goyal, Maryam Haghighat, Nayyar Zaidi, Ferdous Sohel, Bao Q Vo, and KC Santosh. 2025. Towards building robust models for unimodal and multimodal medical imaging data.Information fusion(2025), 103822
2025
-
[15]
Joy Dhar, Manish Kumar Pandey, Debashis Das Chakladar, Maryam Haghighat, Azadeh Alavi, Sajib Mistry, and Nayyar Zaidi. 2026. HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1831–1840
2026
-
[16]
Joy Dhar, Kapil Rana, and Puneet Goyal. 2024. Uncertainty-RIFA-Net: Uncer- tainty Aware Robust Information Fusion Attention Network for Brain Tumors Classification in MRI Images. InInternational Conference on Pattern Recognition. Springer, 311–327
2024
-
[17]
Joy Dhar, Song Xia, Manish Kumar Pandey, Maryam Haghighat, Azadeh Alavi, Ferdous Sohel, Wenyu Zhang, and Nayyar Zaidi. 2026. Certified vs. Empirical Adversarial Robust-ness via Hybrid Convolutions with Attention Stochasticity. arXiv preprint arXiv:2605.01519(2026)
Pith/arXiv arXiv 2026
-
[18]
Joy Dhar, Nayyar Zaidi, and Maryam Haghighat. 2026. Effective and Robust Multimodal Medical Image Analysis. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 188–199
2026
-
[19]
Joy Dhar, Nayyar Zaidi, Maryam Haghighat, Sudipta Roy, Puneet Goyal, Azadeh Alavi, and Vikas Kumar. 2025. Multimodal Fusion Learning with Dual Attention for Medical Imaging. In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 4362–4371. doi:10.1109/WACV61041.2025.00428
arXiv 2025
-
[20]
Octavian-Eugen Ganea, Gary Becigneul, and Thomas Hofmann. 2018. Hyper- bolic Attention Networks. InAdvances in Neural Information Processing Systems, Vol. 31
2018
-
[21]
Yash Goyal, Akrit Mohapatra, Devi Parikh, and Dhruv Batra. 2016. Towards transparent ai systems: Interpreting visual question answering models.arXiv preprint arXiv:1608.08974(2016)
Pith/arXiv arXiv 2016
-
[22]
Brian C. Hall. 2013.Quantum Theory for Mathematicians. Graduate Texts in Mathematics, Vol. 267. Springer New York, New York, NY. 14–15, 58 pages. doi:10.1007/978-1-4614-7116-5
-
[23]
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi. 2023. Neigh- borhood attention transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6185–6194
2023
-
[24]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition. 770–778
2016
-
[25]
X. He, Y. Wang, S. Zhao, and X. Chen. 2023. Co-attention fusion network for multimodal skin cancer diagnosis.Pattern Recognition133 (2023), 108990
2023
-
[26]
Konstantin Hemker, Nikola Simidjievski, and Mateja Jamnik. 2024. HEALNet: Multimodal fusion for heterogeneous biomedical data.Advances in Neural Infor- mation Processing Systems37 (2024), 64479–64498
2024
-
[27]
Konstantin Hemker, Nikola Simidjievski, and Mateja Jamnik. 2025. Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in Biomedicine. arXiv:2405.19950 [cs.LG] https://arxiv.org/abs/2405.19950
Pith/arXiv arXiv 2025
-
[28]
Peizhong Hou, Haiyang Wang, Tianming Li, and Junchi Yan. 2024. HSA: Hy- perbolic Self-Attention for Sequential Recommendation. InWeb and Big Data (APWeb-W AIM 2023). Lecture Notes in Computer Science, Vol. 14333. 250–264. doi:10.1007/978-981-97-2387-4_17
-
[29]
S. C. Huang, L. Shen, M. P. Lungren, and S. Yeung. 2021. Gloria: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image Recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3942–3951
2021
-
[30]
Md Mofijul Islam and Tariq Iqbal. 2020. Hamlet: A hierarchical multimodal attention-based human activity recognition algorithm. In2020 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS). IEEE, 10285–10292
2020
-
[31]
Md Mofijul Islam and Tariq Iqbal. 2022. Mumu: Cooperative multitask learning- based guided multimodal fusion. InProceedings of the AAAI conference on artificial intelligence, Vol. 36. 1043–1051
2022
-
[32]
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. InInternational conference on machine learning. PMLR, 4651–4664
2021
-
[33]
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database.Scientific data3, 1 (2016), 1–9
2016
-
[34]
Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L Iuzzolino, and Kazuhito Koishida. 2020. MMTM: Multimodal transfer module for CNN fusion. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13289–13299
2020
-
[35]
Daniel Kermany. 2018. Labeled optical coherence tomography (oct) and chest x-ray images for classification.Mendeley data(2018)
2018
-
[36]
Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang, Zhiyang He, Xi- aodong Tao, and S Kevin Zhou. 2024. Carzero: Cross-attention alignment for radiology zero-shot classification. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11137–11146
2024
-
[37]
Guangxi Li, Xuanqiang Zhao, and Xin Wang. 2022. Quantum Self-Attention Neural Networks for Text Classification.arXiv preprint arXiv:2205.05625(2022). doi:10.48550/arXiv.2205.05625
-
[38]
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2022. Foundations and trends in multimodal machine learning: Principles, challenges, and open questions.arXiv preprint arXiv:2209.03430(2022)
Pith/arXiv arXiv 2022
-
[39]
M Linehan, R Gautam, S Kirk, Y Lee, C Roche, E Bonaccio, and R Jarosz. 2016. Radiology data from the cancer genome atlas cervical kidney renal papillary cell carcinoma [KIRP] collection.Cancer Imaging Arch10 (2016), K9
2016
-
[40]
W Lingle, BJ Erickson, ML Zuley, R Jarosz, E Bonaccio, J Filippini, and N Gruszauskas. 2016. Radiology data from the cancer genome atlas breast in- vasive carcinoma [TCGA-BRCA] collection.The Cancer Imaging Archive10, K9 (2016), 5
2016
-
[41]
Chang Liu, Henghui Ding, Yulun Zhang, and Xudong Jiang. 2023. Multi-modal mutual attention and iterative interaction for referring image segmentation.IEEE Transactions on Image Processing32 (2023), 3054–3065
2023
-
[42]
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng
-
[43]
S Mourya, S Kant, P Kumar, A Gupta, and R Gupta. 2019. ALL Challenge Dataset of ISBI. 2019.The Cancer Imaging Archive(2019)
2019
-
[44]
Ju-Hyeon Nam, Nur Suriza Syazwany, Su Jung Kim, and Sang-Chul Lee. 2024. Modality-agnostic Domain Generalizable Medical Image Segmentation by Multi- Frequency in Multi-Scale Attention. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11480–11491
2024
-
[45]
Cancer Genome Atlas Research Network et al. 2014. Comprehensive molecular characterization of urothelial bladder carcinoma.Nature507, 7492 (2014), 315
2014
-
[46]
Maximillian Nickel and Douwe Kiela. 2017. Poincaré Embeddings for Learn- ing Hierarchical Representations. InAdvances in Neural Information Processing Systems 30. 6341–6350. doi:10.5555/3295222.3295381
arXiv 2017
-
[47]
Nielsen and Isaac L
Michael A. Nielsen and Isaac L. Chuang. 2010.Quantum Computation and Quantum Information(10th anniversary edition ed.). Cambridge University Press, Cambridge, UK
2010
-
[48]
Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey.IEEE Transactions on pattern analysis and machine intelligence44, 12 (2021), 10023–10044. Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
2021
-
[49]
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8238–8247
2022
-
[50]
Maria E Plissiti, Panagiotis Dimitrakopoulos, Giorgos Sfikas, Christophoros Nikou, Orestis Krikoni, and Avraam Charchanti. 2018. Sipakmed: A new dataset for feature and image based classification of normal and pathological cervical cells in pap smear images. In2018 25th IEEE International Conference on Image Processing (ICIP). IEEE, 3144–3148
2018
-
[51]
Md Mostafijur Rahman, Mustafa Munir, and Radu Marculescu. 2024. Emcad: Efficient multi-scale convolutional attention decoding for medical image segmen- tation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11769–11779
2024
-
[52]
Jinjing Shi, Ren-Xin Zhao, Wenxuan Wang, Shichao Zhang, and Xuelong Li. 2022. QSAN: A Near-term Achievable Quantum Self-Attention Network.arXiv preprint arXiv:2207.07563(2022). doi:10.48550/arXiv.2207.07563
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2207.07563 2022
-
[53]
Jinjing Shi, Ren-Xin Zhao, Wenxuan Wang, Shichao Zhang, and Xuelong Li
-
[54]
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. 2021. How to train your vit? data, augmentation, and regularization in vision transformers.arXiv preprint arXiv:2106.10270(2021)
Pith/arXiv arXiv 2021
-
[55]
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2818–2826
2016
-
[56]
Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions.Scientific data5, 1 (2018), 1–9
2018
-
[57]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 30
2017
-
[58]
Yikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu, Yu Rong, and Junzhou Huang. 2020. Deep multimodal fusion by channel exchanging.Advances in neural information processing systems33 (2020), 4835–4845
2020
-
[59]
Yikai Wang, Fuchun Sun, Ming Lu, and Anbang Yao. 2020. Learning deep multi- modal feature representation with asymmetric multi-layer fusion. InProceedings of the 28th ACM International Conference on Multimedia. 3902–3910
2020
-
[60]
Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. CBAM: Convolutional Block Attention Module. InProceedings of the European Conference on Computer Vision (ECCV). 3–19
2018
-
[61]
Yingxue Xu and Hao Chen. 2023. Multimodal optimal transport-based co- attention transformer with global structure consistency for survival prediction. InProceedings of the IEEE/CVF international conference on computer vision. 21241– 21251
2023
-
[62]
Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. 2023. Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification.Scientific Data10, 1 (2023), 41
2023
-
[63]
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2018. Shufflenet: An ex- tremely efficient convolutional neural network for mobile devices. InProceedings of the IEEE conference on computer vision and pattern recognition. 6848–6856
2018
-
[64]
Ce Zheng, Xianpeng Liu, Guo-Jun Qi, and Chen Chen. 2023. Potter: Pooling attention transformer for efficient human mesh recovery. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1611–1620
2023
-
[65]
Wujie Zhou, Shaohua Dong, Meixin Fang, and Lu Yu. 2023. CACFNet: Cross- modal attention cascaded fusion network for RGB-T urban scene parsing.IEEE Transactions on Intelligent Vehicles(2023). A Appendix – Additional Related Work We summarize the research gaps—highlighting how the CURE framework achieves an optimal performance–efficiency trade-off relative ...
2023
-
[2021]
InProceedings of the AAAI Conference on Artificial Intelligence, Vol
Smil: Multimodal learning with severely missing modality. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 2302–2310
-
[2024]
QSAN: A near-term achievable quantum self-attention network.IEEE Transactions on Neural Networks and Learning Systems(2024)
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.