REVIEW 8 cited by
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be added or removed at any time. Thus, it is crucial to enable the machine to tackle the mismatch and unequal-scale problems of modality combinations between training and inference. In this paper, we tackle these problems from a new perspective: "Modalities Help Modalities". Intuitively, we present OmniBind, a novel two-stage learning framework that can achieve any modality combinations and interaction. It involves teaching data-constrained, a.k.a, student, modalities to be aligned with the well-trained data-abundant, a.k.a, teacher, modalities. This subtly enables the adaptive fusion of any modalities to build a unified representation space for any combinations. Specifically, we propose Cross-modal Alignment Distillation (CAD) to address the unequal-scale problem between student and teacher modalities and effectively align student modalities into the teacher modalities' representation space in stage one. We then propose an Adaptive Fusion (AF) module to fuse any modality combinations and learn a unified representation space in stage two. To address the mismatch problem, we aggregate existing datasets and combine samples from different modalities by the same semantics. This way, we build the first dataset for training and evaluation that consists of teacher (image, text) and student (touch, thermal, event, point cloud, audio) modalities and enables omni-bind for any of them. Extensive experiments on the recognition task show performance gains over prior arts by an average of 4.05 % on the arbitrary modality combination setting. It also achieves state-of-the-art performance for a single modality, e.g., touch, with a 4.34 % gain.
Forward citations
Cited by 8 Pith papers
-
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
A multi-modal semantic segmentation framework that processes RGB and non-RGB sensors separately, matches labels in two stages, and aligns cross-modal queries with a VAE refiner.
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.
-
Gramian Multimodal Representation Learning and Alignment
GRAM replaces cosine similarity with the Gramian volume of the parallelotope formed by multiple modality embeddings, and a volume-based contrastive loss improves multimodal retrieval and classification.
-
Omnidirectional Spatial Modeling from Correlated Panoramas
The authors create a cross-frame panoramic VQA benchmark from 3D scene data and show that GRPO fine-tuning of Qwen2.5-VL raises its score on that benchmark.
-
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
MAGIC++ trains a semantic segmentation backbone with all available sensors and then uses the plain backbone at test time, reporting strong average results on arbitrary sensor combinations.
-
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
AnySeg trains a segmentor to handle arbitrary combinations of visual modalities through unimodal and cross-modal distillation, improving mean mIoU by +6.37% on MUSES and +6.15% on DELIVER over prior state-of-the-art.
-
MLLMs are Deeply Affected by Modality Bias
A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.
-
EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation
EGFormer dynamically scores and drops the least useful sensor modality at each processing stage, cutting parameters by up to 91 percent and GFLOPs by half while keeping segmentation accuracy competitive.
Discussion (0). Continue with ORCID to comment.