Pith. sign in

REVIEW 4 cited by

Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04220 v1 pith:QCUPCLNT submitted 2024-12-05 cs.CV cs.AI

classification cs.CVcs.AI
keywords segmentationmulti-modalmodalitiesperformanceacrossanythingdownstreamexperts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent Segment Anything Model (SAM) represents a significant breakthrough in scaling segmentation models, delivering strong performance across various downstream applications in the RGB modality. However, directly applying SAM to emerging visual modalities, such as depth and event data results in suboptimal performance in multi-modal segmentation tasks. In this paper, we make the first attempt to adapt SAM for multi-modal semantic segmentation by proposing a Mixture of Low-Rank Adaptation Experts (MoE-LoRA) tailored for different input visual modalities. By training only the MoE-LoRA layers while keeping SAM's weights frozen, SAM's strong generalization and segmentation capabilities can be preserved for downstream tasks. Specifically, to address cross-modal inconsistencies, we propose a novel MoE routing strategy that adaptively generates weighted features across modalities, enhancing multi-modal feature integration. Additionally, we incorporate multi-scale feature extraction and fusion by adapting SAM's segmentation head and introducing an auxiliary segmentation head to combine multi-scale features for improved segmentation performance effectively. Extensive experiments were conducted on three multi-modal benchmarks: DELIVER, MUSES, and MCubeS. The results consistently demonstrate that the proposed method significantly outperforms state-of-the-art approaches across diverse scenarios. Notably, under the particularly challenging condition of missing modalities, our approach exhibits a substantial performance gain, achieving an improvement of 32.15% compared to existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multi-modal semantic segmentation framework that processes RGB and non-RGB sensors separately, matches labels in two stages, and aligns cross-modal queries with a VAE refiner.

  2. Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A partial, frozen CLIP block mounted on a segmentation backbone, plus selective distillation to CLIP's CLS token, improves zero-shot semantic segmentation by about 1 hIoU point on two datasets.

  3. MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection

    cs.CV 2024-12 reject novelty 5.0 of 10

    MAGIC++ trains a semantic segmentation backbone with all available sensors and then uses the plain backbone at test time, reporting strong average results on arbitrary sensor combinations.

  4. EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    EGFormer dynamically scores and drops the least useful sensor modality at each processing stage, cutting parameters by up to 91 percent and GFLOPs by half while keeping segmentation accuracy competitive.

Pith tools