Pith. sign in

REVIEW 12 cited by

Cross-Modality Fusion Transformer for Multispectral Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.00273 v4 pith:SSVVFNM6 submitted 2021-10-30 eess.IV cs.CV

classification eess.IVcs.CV
keywords detectionfusiontransformercross-modalitymultispectralobjectapproacheffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effective cross-modality feature fusion approach, named Cross-Modality Fusion Transformer (CFT) in this paper. Unlike prior CNNs-based works, guided by the transformer scheme, our network learns long-range dependencies and integrates global contextual information in the feature extraction stage. More importantly, by leveraging the self attention of the transformer, the network can naturally carry out simultaneous intra-modality and inter-modality fusion, and robustly capture the latent interactions between RGB and Thermal domains, thereby significantly improving the performance of multispectral object detection. Extensive experiments and ablation studies on multiple datasets demonstrate that our approach is effective and achieves state-of-the-art detection performance. Our code and models are available at https://github.com/DocF/multispectral-object-detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A frozen-backbone framework that uses pretrained DINOv3 register tokens as a bidirectional cross-modal bottleneck reports the highest mAP50-95 on LLVIP, M3FD, DroneVehicle, and FLIR-Aligned among the compared methods.

  2. Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-inference decoupled fusion framework (ADCR) with semantic mask exchange and learnable channel competition improves RGB-infrared object detection, with the largest gains on aerial drone imagery.

  3. Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

    cs.CV 2026-07 conditional novelty 6.0 of 10

    DHNet with patch alignment and dual hypergraph fusion reaches SOTA RGBT video object detection on VT-VOD50 and the new large-scale DVT-VOD1000 benchmark.

  4. InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    QualGate-regulated RGB guidance during training produces efficient IR-only and dual-modal detectors that match or beat equal-fusion baselines under low light and adverse weather.

  5. WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    WaveMamba fuses RGB and infrared features in the wavelet domain and reports an average mAP gain of about 4.5 points over prior methods on four public benchmarks.

  6. M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision

    cs.CV 2025-07 conditional novelty 6.0 of 10

    M-SpecGene is a Siamese masked-autoencoder foundation model for RGB-thermal vision, trained on the RGBT550K dataset with a GMM-CMSS progressive masking strategy, and evaluated on four downstream tasks.

  7. CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection

    cs.CV 2026-08 reject novelty 5.0 of 10

    CFGPNet combines RepViT-based backbones, cross-modal attention gating, and attention-based fusion to achieve SOTA mAP on FLIR, M3FD, LLVIP, VEDAI, and MFAD.

  8. From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety

    cs.CV 2025-09 accept novelty 5.0 of 10

    A survey organizing recent camera-based AI methods for vulnerable road user safety into four interlocking visual tasks and four open deployment challenges.

  9. Multispectral Detection Transformer with Infrared-Centric Feature Fusion

    cs.CV 2025-05 conditional novelty 5.0 of 10

    IC-Fusion, an infrared-centric transformer detector with a lightweight RGB backbone and gated fusion modules, achieves state-of-the-art mAP on LLVIP and competitive mAP on FLIR.

  10. M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

    cs.CV 2026-01 conditional novelty 4.0 of 10

    M2I2HA adds intra-modal and cross-modal hypergraph attention modules to a YOLO-style detector and reports the best average precision on DroneVehicle and FLIR, while on LLVIP and VEDAI prior methods score higher on the...

  11. HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A Mamba-based fusion network with a channel-aware decoder reports competitive or state-of-the-art results on RGB-thermal and RGB-depth hidden-object detection benchmarks.

  12. LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    LASFNet fuses RGB and infrared features in one lightweight stage with attention modules, reporting similar or better detection accuracy than heavier multimodal detectors.

Pith tools