Pith. sign in

REVIEW 4 cited by

Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02581 v2 pith:DBGXMKBM submitted 2025-03-04 cs.CV cs.ROeess.IV

classification cs.CVcs.ROeess.IV
keywords perceptionsemanticpotentialsam2segmentationshifnettasksanything
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The perception capability of robotic systems relies on the richness of the dataset. Although Segment Anything Model 2 (SAM2), trained on large datasets, demonstrates strong perception potential in perception tasks, its inherent training paradigm prevents it from being suitable for RGB-T tasks. To address these challenges, we propose SHIFNet, a novel SAM2-driven Hybrid Interaction Paradigm that unlocks the potential of SAM2 with linguistic guidance for efficient RGB-Thermal perception. Our framework consists of two key components: (1) Semantic-Aware Cross-modal Fusion (SACF) module that dynamically balances modality contributions through text-guided affinity learning, overcoming SAM2's inherent RGB bias; (2) Heterogeneous Prompting Decoder (HPD) that enhances global semantic information through a semantic enhancement module and then combined with category embeddings to amplify cross-modal semantic consistency. With 32.27M trainable parameters, SHIFNet achieves state-of-the-art segmentation performance on public benchmarks, reaching 89.8% on PST900 and 67.8% on FMB, respectively. The framework facilitates the adaptation of pre-trained large models to RGB-T segmentation tasks, effectively mitigating the high costs associated with data collection while endowing robotic systems with comprehensive perception capabilities. The source code will be made publicly available at https://github.com/iAsakiT3T/SHIFNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multi-modal semantic segmentation framework that processes RGB and non-RGB sensors separately, matches labels in two stages, and aligns cross-modal queries with a VAE refiner.

  2. Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A partial, frozen CLIP block mounted on a segmentation backbone, plus selective distillation to CLIP's CLS token, improves zero-shot semantic segmentation by about 1 hIoU point on two datasets.

  3. Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    AnySeg trains a segmentor to handle arbitrary combinations of visual modalities through unimodal and cross-modal distillation, improving mean mIoU by +6.37% on MUSES and +6.15% on DELIVER over prior state-of-the-art.

  4. EGFormer: Towards Efficient and Generalizable Multimodal Semantic Segmentation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    EGFormer dynamically scores and drops the least useful sensor modality at each processing stage, cutting parameters by up to 91 percent and GFLOPs by half while keeping segmentation accuracy competitive.

Pith tools