REVIEW 12 cited by
Cross-Modality Fusion Transformer for Multispectral Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effective cross-modality feature fusion approach, named Cross-Modality Fusion Transformer (CFT) in this paper. Unlike prior CNNs-based works, guided by the transformer scheme, our network learns long-range dependencies and integrates global contextual information in the feature extraction stage. More importantly, by leveraging the self attention of the transformer, the network can naturally carry out simultaneous intra-modality and inter-modality fusion, and robustly capture the latent interactions between RGB and Thermal domains, thereby significantly improving the performance of multispectral object detection. Extensive experiments and ablation studies on multiple datasets demonstrate that our approach is effective and achieves state-of-the-art detection performance. Our code and models are available at https://github.com/DocF/multispectral-object-detection.
Forward citations
Cited by 12 Pith papers
-
RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection
A frozen-backbone framework that uses pretrained DINOv3 register tokens as a bidirectional cross-modal bottleneck reports the highest mAP50-95 on LLVIP, M3FD, DroneVehicle, and FLIR-Aligned among the compared methods.
-
Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion
A training-inference decoupled fusion framework (ADCR) with semantic mask exchange and learnable channel competition improves RGB-infrared object detection, with the largest gains on aerial drone imagery.
-
Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark
DHNet with patch alignment and dual hypergraph fusion reaches SOTA RGBT video object detection on VT-VOD50 and the new large-scale DVT-VOD1000 benchmark.
-
InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection
QualGate-regulated RGB guidance during training produces efficient IR-only and dual-modal detectors that match or beat equal-fusion baselines under low light and adverse weather.
-
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
WaveMamba fuses RGB and infrared features in the wavelet domain and reports an average mAP gain of about 4.5 points over prior methods on four public benchmarks.
-
M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
M-SpecGene is a Siamese masked-autoencoder foundation model for RGB-thermal vision, trained on the RGBT550K dataset with a GMM-CMSS progressive masking strategy, and evaluated on four downstream tasks.
-
CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection
CFGPNet combines RepViT-based backbones, cross-modal attention gating, and attention-based fusion to achieve SOTA mAP on FLIR, M3FD, LLVIP, VEDAI, and MFAD.
-
From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
A survey organizing recent camera-based AI methods for vulnerable road user safety into four interlocking visual tasks and four open deployment challenges.
-
Multispectral Detection Transformer with Infrared-Centric Feature Fusion
IC-Fusion, an infrared-centric transformer detector with a lightweight RGB backbone and gated fusion modules, achieves state-of-the-art mAP on LLVIP and competitive mAP on FLIR.
-
M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention
M2I2HA adds intra-modal and cross-modal hypergraph attention modules to a YOLO-style detector and reports the best average precision on DroneVehicle and FLIR, while on LLVIP and VEDAI prior methods score higher on the...
-
HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
A Mamba-based fusion network with a channel-aware decoder reports competitive or state-of-the-art results on RGB-thermal and RGB-depth hidden-object detection benchmarks.
-
LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection
LASFNet fuses RGB and infrared features in one lightweight stage with attention modules, reporting similar or better detection accuracy than heavier multimodal detectors.
Discussion (0). Continue with ORCID to comment.