Pith. sign in

REVIEW 5 cited by

Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.18947 v3 pith:GXOYTINI submitted 2024-04-27 cs.LG cs.AI

classification cs.LGcs.AI
keywords multimodaldatafusiondifferentlow-qualitymodalitieschallengescomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal fusion focuses on integrating information from multiple modalities with the goal of more accurate prediction, which has achieved remarkable progress in a wide range of scenarios, including autonomous driving and medical diagnosis. However, the reliability of multimodal fusion remains largely unexplored especially under low-quality data settings. This paper surveys the common challenges and recent advances of multimodal fusion in the wild and presents them in a comprehensive taxonomy. From a data-centric view, we identify four main challenges that are faced by multimodal fusion on low-quality data, namely (1) noisy multimodal data that are contaminated with heterogeneous noises, (2) incomplete multimodal data that some modalities are missing, (3) imbalanced multimodal data that the qualities or properties of different modalities are significantly different and (4) quality-varying multimodal data that the quality of each modality dynamically changes with respect to different samples. This new taxonomy will enable researchers to understand the state of the field and identify several potential directions. We also provide discussion for the open problems in this field together with interesting future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Multimodal Learning via Imbalanced Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Asymmetric Representation Learning reweights each modality's gradient by the inverse of its prediction variance, improving multimodal accuracy on CREMA-D, Kinetics-Sounds, AVE, MOSI, and UCF101.

  2. Improving Multimodal Learning Balance and Sufficiency through Data Remixing

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Data Remixing improves multimodal learning by decoupling samples into per-modality subsets and training each batch on a single modality, yielding accuracy gains on CREMAD and Kinetic-Sounds.

  3. RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

    cs.LG 2025-06 conditional novelty 5.0 of 10

    RollingQ rotates the classification query in a multimodal Transformer toward a rebalanced direction so attention stops over-favoring a single modality, restoring dynamic fusion and improving accuracy.

  4. Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion

    cs.AI 2025-07 reject novelty 4.0 of 10

    MoCME combines expert-network fusion weighted by estimated mutual information and entropy-based negative sampling, and reports state-of-the-art multi-modal knowledge graph completion on five benchmarks.

  5. Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A structured review of multimodal emotion recognition in conversations, covering datasets, feature processing, methods, and open challenges, with emphasis on recent LLM-based approaches.

Pith tools