Pith. sign in

REVIEW 1 cited by

U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15365 v1 pith:TT6IM3PZ submitted 2024-05-24 cs.CV

classification cs.CV
keywords multimodalfusionsegmentationsemanticmodelunbiasedacrossbiases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal semantic segmentation is a pivotal component of computer vision and typically surpasses unimodal methods by utilizing rich information set from various sources.Current models frequently adopt modality-specific frameworks that inherently biases toward certain modalities. Although these biases might be advantageous in specific situations, they generally limit the adaptability of the models across different multimodal contexts, thereby potentially impairing performance. To address this issue, we leverage the inherent capabilities of the model itself to discover the optimal equilibrium in multimodal fusion and introduce U3M: An Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation. Specifically, this method involves an unbiased integration of multimodal visual data. Additionally, we employ feature fusion at multiple scales to ensure the effective extraction and integration of both global and local features. Experimental results demonstrate that our approach achieves superior performance across multiple datasets, verifing its efficacy in enhancing the robustness and versatility of semantic segmentation in diverse settings. Our code is available at U3M-multimodal-semantic-segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    FGAseg combines a pixel-text alignment transformer, a text-pixel alignment loss, and similarity-based pseudo-masks to achieve state-of-the-art open-vocabulary segmentation on multiple benchmarks.

Pith tools