Pith. sign in

REVIEW 3 major objections 5 minor 67 references

The paper claims that a single lightweight architecture, CURE, fuses arbitrarily many heterogeneous medical modalities—images, multi-omics, clinical records, and wearable time series—into one shared representation, and that this representat

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 13:28 UTC pith:4FWWOFQL

load-bearing objection A lightweight fusion architecture with real paired-regime value, but the headline unpaired results are multi-task gains, not evidence of cross-modal fusion. the 3 major comments →

arxiv 2607.19086 v1 pith:4FWWOFQL submitted 2026-07-21 cs.CV

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention

classification cs.CV
keywords multimodal fusionmedical image analysishyperbolic attentionquantum-inspired attentionheterogeneous medical dataefficient deep learningsurvival predictioncascaded fusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that a single lightweight architecture, CURE, can fuse arbitrarily many heterogeneous medical modalities—images, multi-omics, clinical records, and wearable time series—into one shared representation, and that this shared representation beats specialized state-of-the-art fusion methods on sixteen public datasets while using up to about 88% less computation. The architectural bet is a cascaded, order-invariant pipeline: modalities enter one at a time through a HyFuse layer that combines multi-scale convolution, a hybrid hyperbolic-plus-quantum attention mixer, learnable gating, and shared-information refinement. If the claim holds, multimodal medical AI no longer needs heavy paired-data fusion stacks; the same model can scale to new modalities by appending a HyFuse layer, and it keeps working when some modalities are missing at test time. The paper also claims that these gains are not bought with parameters—the CURE variants range from 3.1M to 14.4M parameters and 0.29 to 2.82 GFLOPs.

Core claim

CURE is presented as a state-of-the-art multimodal fusion framework built on a sequential fusion loop. Each HyFuse layer takes the current shared representation and the next modality, refines both with an Efficient Multimodal Residual Convolution (EMRC) block, passes them through a Hybrid-Space Aware Attention Mixer (HySAM) that computes attention in Poincaré and Lorentz hyperbolic spaces and a quantum-inspired space simultaneously, then fuses them with learnable gates (LLF) and refines them again (SIR). Because modalities are fused one at a time rather than all-pairs, the cost scales linearly with the number of modalities. The paper reports that CURE variants outperform sixteen baselines on

What carries the argument

HyFuse (Hybrid Geometry Aware Fusion) layer, the modular fusion cell that carries the argument. It contains four components: EMRC (multi-scale depthwise/pointwise convolutions for cheap feature extraction), HySAM (a hybrid-space attention mixer combining Multimodal Hyperbolic Dual-Geometry Attention over Poincaré and Lorentz models with Multimodal Quantum-Inspired Attention, fused by learnable gating), LLF (learnable late fusion that masks missing modalities), and SIR (shared information refinement). The key structural claim is that cascading these layers sequentially—rather than fusing all modalities in parallel—produces modality-order-invariant shared representations at linear cost.

Load-bearing premise

The empirical case assumes that fusing datasets with no subject-level alignment (e.g., skin images from one cohort, multi-omics from another, EHR from a third) is a meaningful test of multimodal fusion; if the unpaired gains mainly reflect multi-task regularization, the central generalization claim would not transfer to genuinely paired medical data.

What would settle it

Train CURE on the unpaired groups after randomly permuting the sample indices within each modality stream. If accuracy and AUC do not drop materially, cross-modal alignment is not doing the work, and the reported fusion gains can be attributed to task regularization. Conversely, run the same architecture on a paired cohort (same patients contributing WSI and omics) and compare against a late-fusion baseline with identical backbones; the cross-modal attention should add a measurable margin.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • CURE's claim implies that a single network can be extended to new medical modalities by appending one HyFuse layer, without retraining the fusion topology from scratch.
  • The reported missing-modality results suggest that the same trained model can serve deployments where some data sources are unavailable, since the learnable gates zero out absent streams.
  • The efficiency numbers (0.29–2.82 GFLOPs) imply that multimodal fusion becomes feasible on edge or resource-constrained clinical hardware, not just large GPU clusters.
  • If the modality-order invariance is real, system builders do not need to enforce a canonical ordering of clinical inputs before training.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The unpaired evaluation protocol treats separate, unaligned datasets as 'modalities,' so the reported cross-modal gains may actually come from multi-task regularization across datasets; a paired subject-level benchmark would be needed to verify genuine multimodal fusion.
  • The paper does not isolate whether the non-Euclidean (hyperbolic/quantum) machinery, rather than the cascaded fusion topology, drives the improvement; an equally parameterized Euclidean attention mixer would be a natural control.
  • A testable extension: shuffle the sample indices across the unpaired streams; if CURE's performance is unchanged, its 'fusion' is not exploiting cross-modal correspondences, and the framework is better described as a unified multi-task learner.
  • If the sequential design holds up, it points toward streaming and federated settings where modalities arrive over time, since the shared representation is built incrementally.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes CURE, a lightweight cascaded multimodal fusion framework whose core is the HyFuse layer, comprising residual multi-scale convolutions (EMRC), a hybrid hyperbolic/quantum attention mixer (HySAM), learnable late fusion (LLF), and shared-information refinement (SIR). Modalities are processed sequentially and fused through HyFuse layers, yielding a shared representation x_C used for downstream classification, mortality, and survival tasks. The paper claims state-of-the-art performance across 16 public medical datasets, with up to ≈3.97% accuracy gains and ≈87.8% FLOPs reductions relative to prior multimodal fusion baselines. Evaluation includes unpaired settings, where each dataset is treated as an independent modality stream, and paired WSI+omics survival benchmarks on BLCA and KIRP, plus ablations, missing-modality tests, and resolution scaling.

Significance. The proposed modular architecture and the availability of code are strengths, as are the paired survival experiments and the detailed ablations in Tables 2(c) and 3. If the central claim of efficient and generalizable multimodal fusion were established, the work could be a practical contribution for resource-constrained medical AI. However, the primary evidence for cross-modal fusion rests on the unpaired protocol of Table 1, in which HAM10000, SIPaKMeD, TCGA-BRCA, and MIMIC-III are treated as four fused modality streams despite having no subject-level alignment. Under that protocol, the HyFuse cross-modal attention cannot operate on samples with actual cross-modal correspondence, so the reported gains are more plausibly attributable to multi-task regularization and dataset-specific heads than to multimodal fusion. The paired BLCA/KIRP results are a proper fusion test, but they involve only two datasets and the margins over MMLego are modest (≈2.5 and ≈1.5 C-index points). Thus the headline generalization claims are not yet supported.

major comments (3)
  1. [Sec. 3.1.2, Eqs. (8)–(10)] The unpaired protocol defines each dataset as an independent modality stream and then runs a single CURE pass over, e.g., HAM10000 images, SIPaKMeD images, TCGA-BRCA multi-omics, and MIMIC-III EHR. Because these datasets come from different cohorts/institutions and have no subject-level alignment, the HyFuse cross-modal attention (Eqs. 1–13) operates on randomly batched samples that have no semantic correspondence. The resulting shared representation x_C is therefore a multi-task shared trunk with dataset-specific heads, and the large margins in Table 1 (e.g., +1.42 ACC on HAM10000, +5.8 C-index on BRCA, +5.55 ACC on MORT) could reflect multi-task regularization, dataset-specific heads, or implicit augmentation rather than cross-modal fusion. The abstract's headline '≈3.97%' claim relies heavily on this table. Please re-frame the unpaired experiments as multi-task learning over independe
  2. [Sec. 4; Table 1] The HySAM module is under-specified. In Eq. (8), δ_i is described as a channel-wise bias obtained from MQIA, but Eq. (9) defines the MQIA output as a Softmax attention map; no equation or text maps A^Q_i to δ_i. Additionally, the index i is overloaded: the outer modality index and the inner summation index in Eq. (8) both use i, and the notation δ_{i,l} is not introduced. Eq. (9) also contains an ambiguous expression — the parentheses/operators around |q_i|^2, MLE(ψ_i), and h^{-2} are incomplete — and the hyperparameter h is never given a value. Because HySAM is the key novel component, these omissions prevent reproduction and make it difficult to assess whether the reported gains follow from the stated mechanism.
  3. [Sec. 4; Table 1] Table 1 reports no standard deviations despite the statement that all experiments use five random seeds. Many entries are at saturation (AUC 99.99 and 99.95; ACC 99.81), and the claimed margins over the strongest baseline are often small in absolute terms. Without variance estimates or significance testing, the reader cannot judge whether the reported improvements are reliable. In addition, App. E states that baselines originally designed for other tasks were adapted by removing decoders/heads and attaching task-specific heads; this adaptation protocol should be reported per baseline, since it can materially affect relative rankings.
minor comments (5)
  1. [Algorithm 1, line 10] In the MSIL loop, the else branch calls HyFuse(x^{S'}_i, x_{i+1}) when i=1, but x^{S'}_1 is not yet defined. The special case for i==2 suggests the intended first call uses (x_1, x_2); please fix the indexing in the pseudocode.
  2. [Sec. 4 dataset list vs. App. F] The paper says '16 public datasets', but Appendix F, Table 7 introduces KVASIR and MIT-BIH, which are not listed among D1–D16 in Sec. 4. Please reconcile the dataset count or clarify that these are additional validation datasets.
  3. [Abstract vs. Sec. 4.1] The abstract states a performance improvement of up to ≈3.97%, while Sec. 4.1 reports gains of up to ≈5.8% over the strongest competing method on Table 1. Please make the reported numbers consistent and define the exact aggregation used for the abstract's headline figure.
  4. [Eq. (10) block text] The MAFG block is referred to as 'MFAG' in the sentence preceding Eq. (10). Please correct the typo.
  5. [Sec. 3.1.2, Eq. (9)] The expression for the quantum attention weights is hard to parse. Please define each term explicitly, including whether the norm is Euclidean or Lorentzian, and what the argument of the Softmax is.

Circularity Check

0 steps flagged

No significant circularity: the paper's claims are empirical benchmark comparisons, not derivations that reduce to their own inputs.

full rationale

CURE is an architecture paper; its headline claims (Tables 1, 2) are comparisons against external baselines trained under a common protocol. I found no fitted parameter renamed as a prediction, no theorem whose conclusion is assumed in its premises, and no equation that reduces to its own input by construction. The HyFuse equations (Eqs. 1-15) are standard attention, gating, and residual-convolution modules; they define the architecture but do not derive any result from the data being evaluated. The unpaired MSIL protocol (Footnote 5; App. E) explicitly states that datasets have no subject-level alignment and are treated as independent modality streams; this is a legitimate validity concern about whether the unpaired setup tests cross-modal fusion versus multi-task regularization, but it is not circularity because the reported numbers are empirical outcomes rather than consequences of the protocol's definition. Self-citations such as DRIFA-Net [19] appear as baselines and related work, not as load-bearing justification for CURE's correctness or uniqueness. No self-citation chain forces the central claim, and no ansatz is smuggled in solely via the authors' prior work. Accordingly, the paper's derivation chain is self-contained in the sense required here; the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 2 invented entities

The central empirical claims rest on the unpaired grouping protocol, the pseudo-image reshaping of tabular data, and a set of hand-chosen geometric/quantum-inspired modules. None of these are derived from first principles or externally validated; they are design choices. The paired WSI+omics experiments provide the cleanest test of actual cross-modal fusion.

free parameters (5)
  • learnable curvature c~ = c × (1/e) Σ σ(f_j), c = clip(e^k, 0.1, 10.0) = trained k and f ∈ R^e
    Introduced in Eq. 4 to control Poincaré/Lorentz geometry. Bounded but otherwise unconstrained; no independent evidence for the fractal-scaling parameterization.
  • quantum interaction hyperparameter h = not reported
    Appears in Eq. 9 for MQIA attention. Chosen by hand; no sensitivity analysis is provided.
  • LLF masking constant o = not reported
    Eq. 12 uses a constant o with (1-mask_i) to make missing-modality weights zero. Exact zero only holds as o → -∞; the value is unspecified.
  • task-modality loss weights λ_t^M = not reported
    Eq. 16 weights each task-modality loss. Values are not reported or ablated, so the multitask objective is not fully specified.
  • modality-specific learnable scalars L, P, c_i, α_i, β_l, η_real, η_imag = learned
    Introduced ad hoc in Eqs. 3, 8–10 for gating and complex projection. They are trained network parameters rather than values derived from first principles.
axioms (6)
  • ad hoc to paper Unrelated datasets from different cohorts can be treated as independent "modality streams" and fused by HyFuse without subject-level alignment
    Footnote 5 and App. E define the unpaired setting; this is load-bearing for the 16-dataset generalization claim.
  • domain assumption Non-imaging omics/EHR vectors can be reshaped into pseudo-image tensors R^{H×W×C} without loss of structure
    Sec. 3 problem formulation; all fusion operators assume spatial tensor inputs.
  • domain assumption Poincaré ball and Lorentz hyperboloid embedding formulas (Eqs. 4, 7) preserve the needed frequency structure for attention
    Used in PIL/LIL/MQIA; no theoretical guarantee is provided for this specific DCT-derived input.
  • standard math The Born rule |q_i|² = q_i q*_i is applicable to the complex projection of DCT features for attention weighting
    Eq. 9 invokes a quantum amplitude interpretation. The rule itself is standard, but its use here is an analogy, not quantum physics.
  • domain assumption At inference at least one modality must be present
    Sec. 3.1.3 states this; LLF cannot handle the case where all modalities are missing.
  • domain assumption Four-fold label-preserving augmentation (rotation, translation, blur) does not distort downstream task labels
    App. D generates a 4× training set; the method assumes these geometric/blur variants preserve the diagnostic label.
invented entities (2)
  • Quantum-inspired complex state q_i = ψ_i·(η_real + jη_imag) no independent evidence
    purpose: Captures long-range dependencies and produces the MQIA attention/bias in HySAM
    An architectural device in Eq. 9; it makes no external falsifiable prediction and is not a physically implemented quantum state.
  • Learnable curvature c~ with fractal scaling weights no independent evidence
    purpose: Controls hyperbolic embedding geometry in PIL/LIL
    A hand-designed trainable parameterization; no independent evidence or theoretical constraint beyond the clip bounds.

pith-pipeline@v1.3.0-alltime-deepseek · 27228 in / 15234 out tokens · 144141 ms · 2026-08-01T13:28:43.802954+00:00 · methodology

0 comments
read the original abstract

Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively, which in turn limits performance improvements. Second, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI applications. Finally, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. To address these challenges, we propose a novel MFL framework - Cascaded Unified Representation Learning for Efficient Fusion Network (CURE) - a lightweight and scalable framework that progressively integrates various modalities through a novel efficient Hybrid Geometry Aware Fusion layer (HyFuse), where each HyFuse layer is sequentially learned for each modality, making the framework adaptable and generalizable. Within HyFuse, an efficient residual convolution module captures rich multi-scale features to ensure cost-effective learning, while a hybrid-space aware attention mixer learns coarse-to-fine structural cues to better preserve cross-modal relationships. Complementary learnable late-fusion and shared information refinement modules are then employed to learn robust modality-order-invariant shared representations, which in turn yields consistent performance improvements. Extensive evaluations on 16 public datasets show that CURE outperforms leading multimodal fusion methods, boosting performance by up to 3.97% and lowering computational costs by up to 87.8%, ensuring more effective and reliable predictions.

Figures

Figures reproduced from arXiv: 2607.19086 by Chen Chen, Ferdous Sohel, Joy Dhar, Manish Kumar Pandey, Maryam Haghighat, Nayyar Zaidi, Puneet Goyal.

Figure 1
Figure 1. Figure 1: Comparative high-level overview of various MFL methods – (Left) architecture of intermediate or late fusion methods (e.g., DRIFA-Net [19], MuMu [31], MOTCAT [61]) and (Right) architecture of our proposed CURE framework. To address this question, classical multimodal fusion paradigms— early fusion1 [32], late fusion2 [10, 29–31], and intermediate fu￾sion3 [19, 61] (see [PITH_FULL_IMAGE:figures/full_fig_p00… view at source ↗
Figure 3
Figure 3. Figure 3: Overview of CURE framework, comprising the MSIL and HMML phases. (a) Overview of MSIL phase utilizing HyFuse layer. (b) Overview of HMML phase. (c) Layout of HyFuse layer, composed of EMRC, HySAM, LLF, and SIR modules, to learn robust shared features (𝑥 𝑆 ′ 𝑗 ); this, in turn, yields modality-order-invariant shared representations 𝑥 𝐶 via the intermediate fusion in (a); See Sec. 3.1.3–3.1.4). frequency-dom… view at source ↗
Figure 4
Figure 4. Figure 4: (a) Illustration of the HySAM module’s role in learning more expressive shared representations 𝑥 𝑆 𝑖 for each 𝑖 ∈ {1, . . . ,𝑚}. (b) Overview of the HySAM module, which employs the HQMGA block – comprising the MHDGA and MQIA mechanisms and a MAFG block that fuses the attention maps from MHDGA and MQIA to learn hybrid-space aware shared attention maps 𝐴𝑖 . fusion. Our aim is to learn a function – F (·) that… view at source ↗
Figure 5
Figure 5. Figure 5: (a) Overview of the MHDGA module, which incorporates PIL and LIL to learn dual-geometry-aware attention maps. (b) Detailed view of the PIL branch for computing Poincaré attention weights. (c–d) Overview of the LIL–MQIA mutual-guidance loop (red arrows), which couples Lorentzian hyperbolic embeddings with quantum-inspired interactions to yield richer, co-evolving attention weights. Concretely, LIL embeds fe… view at source ↗
Figure 6
Figure 6. Figure 6: Visual comparison of discriminative regions high￾lighted by our proposed CURE variant (e.g., CURE-50) and seven top￾performing state-of-the-art methods using the Grad-CAM technique on two benchmark datasets: HAM10000 (top row) and SIPaKMeD (bot￾tom row) This adaptive fusion ensures robust integration of hyperbolic and quantum structural priors, enhancing cross-modal interactions and enabling the learning o… view at source ↗
Figure 7
Figure 7. Figure 7: Architecture of the Efficient Multimodal Residual Convolution (EMRC) and Shared Information Refinement (SIR) modules. (A) The EMRC module integrates the Modality-specific Heterogeneous Convolutions Fusion (MHCF) block (shown in B) to progressively refine multimodal representations 𝑥 ′ 𝑖 , where 𝑖 ∈ [1 : 𝑚]. (C) The SIR module takes the shared representations 𝑥 𝑆 𝑖 produced by the HySAM module and refines t… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    I. S. A. Abdelhalim, M. F. Mohamed, and Y. B. Mahdy. 2021. Data augmentation for skin lesion using self-attention based progressive generative adversarial network. Expert Systems with Applications165 (2021), 113922

  2. [2]

    Yajun An, Jiale Chen, Huan Lin, Zhenbing Liu, Siyang Feng, Hualong Zhang, Rushi Lan, Zaiyi Liu, and Xipeng Pan. 2025. CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Joy Dhar et al. Prediction. InProceedings of the AAAI Conference on Artificial I...

  3. [3]

    Plamen Angelov and Eduardo Soares. 2020. Towards explainable deep neural networks (xDNN).Neural Networks130 (2020), 185–194

  4. [4]

    Anguita, A

    D. Anguita, A. Ghio, L. Oneto, X. Parra, and J.L. Reyes-Ortiz. 2013. A Public Domain Dataset for Human Activity Recognition Using Smartphones. In21st European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN)

  5. [5]

    Ujjwal Baid, Satyam Ghodasara, Suyash Mohan, Michel Bilello, Evan Calabrese, Errol Colak, Keyvan Farahani, Jayashree Kalpathy-Cramer, Felipe C Kitamura, Sarthak Pati, et al . 2021. The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification.arXiv preprint arXiv:2107.02314(2021)

  6. [6]

    Oresti Banos, Rafael Garcia, Juan A Holgado-Terriza, Miguel Damas, Hector Pomares, Ignacio Rojas, Alejandro Saez, and Claudia Villalonga. 2014. mHealth- Droid: a novel framework for agile development of mobile health applications. In Ambient Assisted Living and Daily Activities: 6th International Work-Conference, IW AAL 2014, Belfast, UK, December 2-5, 20...

  7. [7]

    Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross- attention multi-scale vision transformer for image classification. InProceedings of the IEEE/CVF international conference on computer vision. 357–366

  8. [8]

    Richard J Chen, Ming Y Lu, Jingwen Wang, Drew FK Williamson, Scott J Rodig, Neal I Lindeman, and Faisal Mahmood. 2020. Pathomic fusion: an integrated framework for fusing histopathology and genomic features for cancer diagnosis and prognosis.IEEE Transactions on Medical Imaging41, 4 (2020), 757–770

  9. [9]

    Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2021. Fully Hyperbolic Neural Networks. arXiv:2105.14686 [cs.CL] doi:10.48550/arXiv.2105.14686 Submitted on 31 May 2021; last revised 16 Mar 2022 (this version, v3)

  10. [10]

    Cheng, J

    J. Cheng, J. Liu, H. Kuang, and J. Wang. 2022. A Fully Automated Multimodal MRI-Based Multi-Task Learning for Glioma Segmentation and IDH Genotyping. IEEE Transactions on Medical Imaging41, 6 (June 2022), 1520–1532. doi:10.1109/ tmi.2022.3142321

  11. [11]

    Iris Cong, Soonwon Choi, and Mikhail D. Lukin. 2018. Quantum Convolutional Neural Networks. arXiv:1810.03787 [quant-ph] Revised version posted May 2, 2019

  12. [12]

    Can Cui, Haichun Yang, Yaohong Wang, Shilin Zhao, Zuhayr Asad, Lori A Coburn, Keith T Wilson, Bennett A Landman, and Yuankai Huo. 2023. Deep multimodal fusion of image and non-image data in disease diagnosis and prognosis: a review. Progress in Biomedical Engineering5, 2 (2023), 022001

  13. [13]

    Y. Cui, Y. Tao, W. Ren, and A. Knoll. 2023. Dual-domain attention for image deblurring. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 479–487

  14. [14]

    Joy Dhar, Puneet Goyal, Maryam Haghighat, Nayyar Zaidi, Ferdous Sohel, Bao Q Vo, and KC Santosh. 2025. Towards building robust models for unimodal and multimodal medical imaging data.Information fusion(2025), 103822

  15. [15]

    Joy Dhar, Manish Kumar Pandey, Debashis Das Chakladar, Maryam Haghighat, Azadeh Alavi, Sajib Mistry, and Nayyar Zaidi. 2026. HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1831–1840

  16. [16]

    Joy Dhar, Kapil Rana, and Puneet Goyal. 2024. Uncertainty-RIFA-Net: Uncer- tainty Aware Robust Information Fusion Attention Network for Brain Tumors Classification in MRI Images. InInternational Conference on Pattern Recognition. Springer, 311–327

  17. [17]

    Joy Dhar, Song Xia, Manish Kumar Pandey, Maryam Haghighat, Azadeh Alavi, Ferdous Sohel, Wenyu Zhang, and Nayyar Zaidi. 2026. Certified vs. Empirical Adversarial Robust-ness via Hybrid Convolutions with Attention Stochasticity. arXiv preprint arXiv:2605.01519(2026)

  18. [18]

    Joy Dhar, Nayyar Zaidi, and Maryam Haghighat. 2026. Effective and Robust Multimodal Medical Image Analysis. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 188–199

  19. [19]

    Joy Dhar, Nayyar Zaidi, Maryam Haghighat, Sudipta Roy, Puneet Goyal, Azadeh Alavi, and Vikas Kumar. 2025. Multimodal Fusion Learning with Dual Attention for Medical Imaging. In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 4362–4371. doi:10.1109/WACV61041.2025.00428

  20. [20]

    Octavian-Eugen Ganea, Gary Becigneul, and Thomas Hofmann. 2018. Hyper- bolic Attention Networks. InAdvances in Neural Information Processing Systems, Vol. 31

  21. [21]

    Yash Goyal, Akrit Mohapatra, Devi Parikh, and Dhruv Batra. 2016. Towards transparent ai systems: Interpreting visual question answering models.arXiv preprint arXiv:1608.08974(2016)

  22. [22]

    Brian C. Hall. 2013.Quantum Theory for Mathematicians. Graduate Texts in Mathematics, Vol. 267. Springer New York, New York, NY. 14–15, 58 pages. doi:10.1007/978-1-4614-7116-5

  23. [23]

    Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi. 2023. Neigh- borhood attention transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6185–6194

  24. [24]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition. 770–778

  25. [25]

    X. He, Y. Wang, S. Zhao, and X. Chen. 2023. Co-attention fusion network for multimodal skin cancer diagnosis.Pattern Recognition133 (2023), 108990

  26. [26]

    Konstantin Hemker, Nikola Simidjievski, and Mateja Jamnik. 2024. HEALNet: Multimodal fusion for heterogeneous biomedical data.Advances in Neural Infor- mation Processing Systems37 (2024), 64479–64498

  27. [27]

    Konstantin Hemker, Nikola Simidjievski, and Mateja Jamnik. 2025. Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in Biomedicine. arXiv:2405.19950 [cs.LG] https://arxiv.org/abs/2405.19950

  28. [28]

    Peizhong Hou, Haiyang Wang, Tianming Li, and Junchi Yan. 2024. HSA: Hy- perbolic Self-Attention for Sequential Recommendation. InWeb and Big Data (APWeb-W AIM 2023). Lecture Notes in Computer Science, Vol. 14333. 250–264. doi:10.1007/978-981-97-2387-4_17

  29. [29]

    S. C. Huang, L. Shen, M. P. Lungren, and S. Yeung. 2021. Gloria: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image Recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3942–3951

  30. [30]

    Md Mofijul Islam and Tariq Iqbal. 2020. Hamlet: A hierarchical multimodal attention-based human activity recognition algorithm. In2020 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS). IEEE, 10285–10292

  31. [31]

    Md Mofijul Islam and Tariq Iqbal. 2022. Mumu: Cooperative multitask learning- based guided multimodal fusion. InProceedings of the AAAI conference on artificial intelligence, Vol. 36. 1043–1051

  32. [32]

    Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. InInternational conference on machine learning. PMLR, 4651–4664

  33. [33]

    Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database.Scientific data3, 1 (2016), 1–9

  34. [34]

    Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L Iuzzolino, and Kazuhito Koishida. 2020. MMTM: Multimodal transfer module for CNN fusion. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13289–13299

  35. [35]

    Daniel Kermany. 2018. Labeled optical coherence tomography (oct) and chest x-ray images for classification.Mendeley data(2018)

  36. [36]

    Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang, Zhiyang He, Xi- aodong Tao, and S Kevin Zhou. 2024. Carzero: Cross-attention alignment for radiology zero-shot classification. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11137–11146

  37. [37]

    Guangxi Li, Xuanqiang Zhao, and Xin Wang. 2022. Quantum Self-Attention Neural Networks for Text Classification.arXiv preprint arXiv:2205.05625(2022). doi:10.48550/arXiv.2205.05625

  38. [38]

    Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2022. Foundations and trends in multimodal machine learning: Principles, challenges, and open questions.arXiv preprint arXiv:2209.03430(2022)

  39. [39]

    M Linehan, R Gautam, S Kirk, Y Lee, C Roche, E Bonaccio, and R Jarosz. 2016. Radiology data from the cancer genome atlas cervical kidney renal papillary cell carcinoma [KIRP] collection.Cancer Imaging Arch10 (2016), K9

  40. [40]

    W Lingle, BJ Erickson, ML Zuley, R Jarosz, E Bonaccio, J Filippini, and N Gruszauskas. 2016. Radiology data from the cancer genome atlas breast in- vasive carcinoma [TCGA-BRCA] collection.The Cancer Imaging Archive10, K9 (2016), 5

  41. [41]

    Chang Liu, Henghui Ding, Yulun Zhang, and Xudong Jiang. 2023. Multi-modal mutual attention and iterative interaction for referring image segmentation.IEEE Transactions on Image Processing32 (2023), 3054–3065

  42. [42]

    Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng

  43. [43]

    S Mourya, S Kant, P Kumar, A Gupta, and R Gupta. 2019. ALL Challenge Dataset of ISBI. 2019.The Cancer Imaging Archive(2019)

  44. [44]

    Ju-Hyeon Nam, Nur Suriza Syazwany, Su Jung Kim, and Sang-Chul Lee. 2024. Modality-agnostic Domain Generalizable Medical Image Segmentation by Multi- Frequency in Multi-Scale Attention. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11480–11491

  45. [45]

    Cancer Genome Atlas Research Network et al. 2014. Comprehensive molecular characterization of urothelial bladder carcinoma.Nature507, 7492 (2014), 315

  46. [46]

    Maximillian Nickel and Douwe Kiela. 2017. Poincaré Embeddings for Learn- ing Hierarchical Representations. InAdvances in Neural Information Processing Systems 30. 6341–6350. doi:10.5555/3295222.3295381

  47. [47]

    Nielsen and Isaac L

    Michael A. Nielsen and Isaac L. Chuang. 2010.Quantum Computation and Quantum Information(10th anniversary edition ed.). Cambridge University Press, Cambridge, UK

  48. [48]

    Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey.IEEE Transactions on pattern analysis and machine intelligence44, 12 (2021), 10023–10044. Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

  49. [49]

    Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8238–8247

  50. [50]

    Maria E Plissiti, Panagiotis Dimitrakopoulos, Giorgos Sfikas, Christophoros Nikou, Orestis Krikoni, and Avraam Charchanti. 2018. Sipakmed: A new dataset for feature and image based classification of normal and pathological cervical cells in pap smear images. In2018 25th IEEE International Conference on Image Processing (ICIP). IEEE, 3144–3148

  51. [51]

    Md Mostafijur Rahman, Mustafa Munir, and Radu Marculescu. 2024. Emcad: Efficient multi-scale convolutional attention decoding for medical image segmen- tation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11769–11779

  52. [52]

    Jinjing Shi, Ren-Xin Zhao, Wenxuan Wang, Shichao Zhang, and Xuelong Li. 2022. QSAN: A Near-term Achievable Quantum Self-Attention Network.arXiv preprint arXiv:2207.07563(2022). doi:10.48550/arXiv.2207.07563

  53. [53]

    Jinjing Shi, Ren-Xin Zhao, Wenxuan Wang, Shichao Zhang, and Xuelong Li

  54. [54]

    Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. 2021. How to train your vit? data, augmentation, and regularization in vision transformers.arXiv preprint arXiv:2106.10270(2021)

  55. [55]

    Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2818–2826

  56. [56]

    Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions.Scientific data5, 1 (2018), 1–9

  57. [57]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems (NeurIPS), Vol. 30

  58. [58]

    Yikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu, Yu Rong, and Junzhou Huang. 2020. Deep multimodal fusion by channel exchanging.Advances in neural information processing systems33 (2020), 4835–4845

  59. [59]

    Yikai Wang, Fuchun Sun, Ming Lu, and Anbang Yao. 2020. Learning deep multi- modal feature representation with asymmetric multi-layer fusion. InProceedings of the 28th ACM International Conference on Multimedia. 3902–3910

  60. [60]

    Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. CBAM: Convolutional Block Attention Module. InProceedings of the European Conference on Computer Vision (ECCV). 3–19

  61. [61]

    Yingxue Xu and Hao Chen. 2023. Multimodal optimal transport-based co- attention transformer with global structure consistency for survival prediction. InProceedings of the IEEE/CVF international conference on computer vision. 21241– 21251

  62. [62]

    Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. 2023. Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification.Scientific Data10, 1 (2023), 41

  63. [63]

    Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2018. Shufflenet: An ex- tremely efficient convolutional neural network for mobile devices. InProceedings of the IEEE conference on computer vision and pattern recognition. 6848–6856

  64. [64]

    Ce Zheng, Xianpeng Liu, Guo-Jun Qi, and Chen Chen. 2023. Potter: Pooling attention transformer for efficient human mesh recovery. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1611–1620

  65. [65]

    Wujie Zhou, Shaohua Dong, Meixin Fang, and Lu Yu. 2023. CACFNet: Cross- modal attention cascaded fusion network for RGB-T urban scene parsing.IEEE Transactions on Intelligent Vehicles(2023). A Appendix – Additional Related Work We summarize the research gaps—highlighting how the CURE framework achieves an optimal performance–efficiency trade-off relative ...

  66. [2021]

    InProceedings of the AAAI Conference on Artificial Intelligence, Vol

    Smil: Multimodal learning with severely missing modality. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 2302–2310

  67. [2024]

    QSAN: A near-term achievable quantum self-attention network.IEEE Transactions on Neural Networks and Learning Systems(2024)