REVIEW 3 major objections 4 minor 145 references
CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read CFGPNet claims that visible–thermal object detection reaches leading results on five benchmarks by having each modality re-weight the other through exchanged spatial attention maps, with a gradient-programmed auxiliary branch carrying the…
desk verdict A detailed, transparent architecture paper whose headline numbers are undermined by an internal contradiction between the specified final loss and the ablation table; the reported SOTA is not reproducible until that is resolved. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is CrossCEA, a cross-attention module that swaps compact spatial reliability maps between the visible and thermal branches. Inside each CEA unit, the input feature map is split into channel groups; axis-pooled descriptors pass through depthwise 1D convolutions, group normalization, and sigmoid to form height- and width-wise gates, while a parallel global branch builds channel queries; a temperature-scaled interaction between the local and global paths produces a single 2D attention map $\Omega$ per modality. CrossCEA then gates each stream with the other stream's map, $\tilde{F}^I = F^I \odot \Omega^V$ and $\tilde{F}^V = F^V \odot \Omega^I$, so that only bounded attention weights, not raw feature values, cross between modalities. The second mechanism is ASAF, which concatenates the enhanced streams and runs two parallel paths — a DenseNet-style dense aggregation refined by CBAM, and an ELA-guided multi-branch selector (MBatt) whose element-wise max keeps the strongest attention candidate at each location — before a 1×1 projection yields the compact fused tensor. The third mechanism is optimization-level: the PGI auxiliary branch instantiates a second detection path whose loss (weighted 0.25) sends programmed gradients back through tapped intermediate features of the main branch, and the branch stays active at inference.
What would settle it
Retrain the strongest published RGB–T baselines under CFGPNet's exact protocol — 600 epochs from scratch at 640×640, the same batch-size-adjusted schedules, the same dataset-specific NMS IoU thresholds chosen on 10% held-out validation subsets, and best-checkpoint selection — and compare on the five benchmarks. If a baseline matches or exceeds CFGPNet under that protocol, the architecture-attributed gains are refuted; a complementary check replaces CrossCEA's swapped attention maps with concatenation or early feature mixing at matched computation and asks whether the mAP50:95 margin survives.
Extended reading notes
Core claim
The paper's central claim is that a multispectral detector built from three interacting mechanisms — a RepViT-block GELAN backbone with separate visible and infrared weights, a Cross Computation Efficient Attention (CrossCEA) stage that exchanges compact spatial attention maps between the two streams, and an Attention Selection and Aggregation Fusion (ASAF) network that condenses the enhanced streams into one tensor — outperforms recent RGB–T detectors on every benchmark it evaluates. The best reported configurations reach 80.7% mAP50 / 45.0% mAP50:95 on FLIR, 89.9% / 63.4% on M3FD, 97.8% / 68.9% on LLVIP, 83.3% / 56.9% on VEDAI, and 83.4% / 61.8% on MFAD. The author would add that the framework is trained from scratch in three scales (21.0M, 71.7M, and 180.9M parameters), that the programmable-gradient auxiliary branch remains active at inference and contributes candidate detections merged by non-maximum suppression, and that the MFAD ablations attribute the gains to the combination of the RepViT backbone, CrossCEA, ASAF, and the PGI branch rather than to any single module.
Load-bearing premise
The reported superiority over prior detectors assumes that CFGPNet's longer training schedule (600 epochs from scratch), per-dataset NMS IoU threshold tuning, and best-checkpoint selection are not what produce the gains, since no published baseline is retrained under the same protocol.
Editorial extensions
If this is right
- If the reported results hold, the 'exchange attention maps, not features' principle offers a lightweight alternative to dense Transformer cross-attention for RGB–T fusion, since CrossCEA adds no token-wise attention and the three scales report 18.6–52.2 FPS at 640×640 input.
- The compact 21.0M-parameter variant claims to match or beat much larger published fusion detectors on FLIR, M3FD, and LLVIP, suggesting strong parameter-efficient RGB–T detection is possible without ImageNet-pretrained initialization.
- The MFAD ablation quantifies the optimization contribution: removing the PGI branch drops mAP50 from 79.6 to 74.1 and mAP50:95 from 56.7 to 53.3, so the gradient-programming path is a substantial part of the reported gain.
- The paper's VEDAI evaluation uses Split 1 for all variants and includes only single-split reports from prior work, so its VEDAI comparisons are conditional on that split being representative of the benchmark.
Reading between the lines
- The paper never runs the clean attribution experiment: retraining strong published baselines under its exact 600-epoch, from-scratch protocol with identical NMS tuning and checkpoint selection; without that control, the size of the architecture-level margin over prior work is an open question.
- The attention-map-swap principle is a transferable hypothesis for any aligned sensor pair with location-dependent reliability — RGB-depth, RGB-event, or multi-spectral satellite imagery — and could be tested at matched compute against concatenation, learned-gate, or Transformer fusion.
- The ablations show that pointwise and structural alignment losses (MSE, SSIM) actively degrade accuracy, which suggests forcing modalities toward one another is counterproductive; a useful follow-up is to test whether any cross-modal alignment term is needed once the backbone and fusion are strong.
- The abstract's LLVIP headline pair mixes scales: 97.8% mAP50 comes from CFGPNet-m and 68.9% mAP50:95 from CFGPNet-e, and no single variant achieves both; with accuracy reported as single-run point estimates, small margins on saturated benchmarks may fall within run-to-run variance, which the conclusion itself flags as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CFGPNet, a multispectral (RGB-T) object detection framework combining a RepViT-based GELAN backbone, a CrossCEA cross-modal attention module, an ASAF fusion module, and a programmable-gradient auxiliary branch (PGI). Three model scales (m/c/e) are designed and evaluated on FLIR, M3FD, LLVIP, VEDAI, and MFAD, where the paper reports state-of-the-art or near-state-of-the-art results. The authors also provide detailed layer-wise configuration tables, release code/data/models, and conduct ablation studies on MFAD to support the contribution of individual components.
Significance. If the reported results are reproducible, CFGPNet would represent a practical advance in multispectral detection, combining competitive accuracy with a favorable efficiency trade-off. The paper is commendable for its detailed architectural descriptions, the release of code and fine-tuned models, and the broad evaluation across five benchmarks with per-class breakdowns. The ablations are extensive and mostly controlled, and the disclosure of dataset-specific NMS thresholds and the use of a single VEDAI split is transparent. However, the current manuscript contains a load-bearing internal contradiction in the training-loss specification that prevents verification of which configuration produced the headline results; this issue, together with a lack of same-protocol baselines, makes the central claim not yet defensible as written.
major comments (3)
- [Section 3.6 vs. Table 12] The training objective is specified inconsistently. In Section 3.6, the total loss is defined as the YOLOv9 detection losses plus all three proposed alignment terms, with λ_CE−KL=1.0, λ_SSIM=3.0, and λ_MSE=300.0, and the text states that these coefficients are kept fixed across datasets. However, Table 12 of the ablation study reports that this exact combination (MSE+SSIM+CE-KL) yields only 74.3 mAP50 / 51.5 mAP50:95 on MFAD, while CFGPNet-m is reported at 79.6 / 56.7. The same table identifies CE-KL alone as the best loss, and the text in Section 4.5 explicitly states that "CE-KL is selected for the final model." These two statements are mutually contradictory. If the final models were trained with the Section 3.6 loss, the headline results are contradicted by the paper's own ablation; if the final models used CE-KL only, then Section 3.6 and its hyperparameters do not describe the actual training objective. Since the loss function is central to the optimization, the reported state-of-the-art numbers cannot be attributed to the proposed architecture, and independent reproduction is blocked. The authors must correct this inconsistency and specify exactly which loss was used for each reported result, ideally by providing training logs or rerunning the affected experiments.
- [Sections 4.3 and 4.4] The comparison with prior methods is not conducted under equal conditions. All CFGPNet variants are trained from scratch for 600 epochs, use dataset-specific NMS IoU thresholds selected on a 10% validation subset (Table 2), and are evaluated with the best checkpoint over the training schedule (as stated for VEDAI). In contrast, baseline results are copied from their original publications, which may use different training budgets, data splits, input resolutions, and post-processing settings. Consequently, the reported gains cannot be isolated to the architectural contributions (CrossCEA, ASAF, PGI). To support the claim of architecture-level superiority, the authors should provide at least one strong baseline (e.g., a YOLOv9-based dual-stream detector, or the best performing published method) trained under the same 600-epoch schedule, same data augmentation, same NMS calibration, and same checkpoint-selection protocol. Without such a control, the state-of-the-art claim is not convincingly established.
- [Table 12 and Section 3.6] Even if the loss-contradiction is resolved, the ablation provides only weak support for the proposed feature-alignment loss. In Table 12, the "None" configuration achieves 79.7 mAP50 and 56.6 mAP50:95, while CE-KL alone achieves 79.6 and 56.7, i.e., a 0.1-point improvement on mAP50:95 and a 0.1-point drop on mAP50. The text acknowledges this but nevertheless selects CE-KL. Given that one of the paper's stated contributions is the auxiliary feature-alignment supervision of CrossCEA, the negligible and sometimes negative effect of the alignment loss should be discussed more candidly. The authors should either provide additional ablations (e.g., different λ values, other datasets) showing a consistent benefit, or moderate the claim that this loss is a key enabler of the reported performance.
minor comments (4)
- [Table 14 / Section 4.5] The MBatt branch ablation reports results for 2 to 6 branches, but the text refers to the case of a single branch ("When only one branch is used..."); either add the N_b=1 row to the table or adjust the discussion to avoid referencing a configuration that is not tabulated.
- [Tables 3 vs 10–16] The FPS values differ by an order of magnitude between the main comparison (e.g., CFGPNet-m: 52.2 FPS) and the ablation tables (e.g., 1463.4 FPS for the same model). Although the batched timing protocol is explained, the presentation is confusing. Please clearly label ablation FPS as "relative batched timing" in every table caption and in the text wherever such values appear.
- [Section 4.4.4 (VEDAI)] The sentence "each variant is evaluated using its best-performing checkpoint over the same 600-epoch training schedule" should specify whether the best checkpoint is chosen using the validation subset or the test split. If the test set was used for checkpoint selection, the reported numbers would be optimistically biased.
- [Abstract and Section 4.4] The claim of "strong and consistent performance" is weaker on FLIR, where the mAP50 gain over the best published method (ERFF, 80.6) is only 0.1 point for CFGPNet-e. Please present a more nuanced interpretation of the margin of improvement on each dataset.
Circularity Check
No significant circularity: the reported gains are empirical benchmark measurements, and the loss-function inconsistency is a reproducibility flaw rather than a derivation that reduces to its inputs.
full rationale
CFGPNet's central claims are empirical detection scores on five public benchmarks produced by training the proposed architecture. There is no derivation chain in which a predicted quantity is constructed from the same quantity, and no fitted parameter is renamed as a prediction. The NMS IoU thresholds are tuned on a held-out 10% validation subset and then fixed before evaluation; this is standard post-processing calibration and does not make the reported mAP values equivalent to the tuning objective. The learnable aggregation factors are explicitly ablated and discarded, so they are not fitted inputs masquerading as predictions. The paper does not rely on load-bearing self-citations: references to YOLOv9, RepViT, EMA, ELA, DenseNet, and CBAM are external published methods, and no uniqueness theorem from the authors is invoked. The internal contradiction between §3.6, which specifies the final objective as YOLOv9 losses plus MSE, SSIM, and CE-KL with the stated coefficients, and Table 12, which shows that MSE+SSIM+CE-KL degrades CFGPNet-m to 74.3/51.5 and that CE-KL alone is selected, is a serious reproducibility and correctness flaw that prevents unambiguous attribution of the reported gains to the proposed modules. However, this is not circularity in the derivation sense: the reported numbers are still measurements on external benchmarks, and no equation reduces an output to an input by construction. Single-run point estimates and differences in training budgets between CFGPNet and published baselines are statistical fairness concerns, not circularity. Overall, the paper is self-contained as an empirical architecture study, so the circularity score is 0.
Assumptions & free parameters
free parameters (9)
- NMS IoU threshold (FLIR) =
0.6
- NMS IoU threshold (M3FD) =
0.5
- NMS IoU threshold (LLVIP) =
0.65
- NMS IoU threshold (VEDAI) =
0.45
- NMS IoU threshold (MFAD) =
0.55
- Loss weights lambda_CE-KL, lambda_SSIM, lambda_MSE =
1.0, 3.0, 300.0
- rho_PGI (auxiliary branch weight) =
0.25
- Detection loss gains (box, cls, DFL) =
7.5, 0.5, 1.5
- Number of MBatt branches Nb =
3
assumptions (4)
- standard math Softmax and scaled dot-product attention operations are computed as defined (Eq. 14).
- domain assumption The registered RGB-T pairs share the same field of view and are synchronized, so cross-modal attention map swapping is meaningful.
- domain assumption The reported results of compared methods were obtained under comparable protocols and are accurate as cited.
- ad hoc to paper The alignment losses (CE-KL, SSIM, MSE) improve fusion; however, the ablation shows these losses do not clearly help, making this an ad hoc assumption of the paper.
Cite this review
Pith. "Pith review of CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection." pith.science (2026). https://pith.science/paper/BD36R2GG
@misc{pith2026260806205,
author = {Pith},
title = {Pith review of: CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/BD36R2GG}},
note = {Machine review of arXiv:2608.06205}
}
read the original abstract
RGB--T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low-light, adverse-weather, and complex multi-scale environments. However, existing methods still suffer from insufficient cross-modal interaction, unstable fusion from modality distribution gaps, and the high computational cost of heavy attention-based architectures. To address these issues, CFGPNet is proposed, a Cross-Attention-Based Fused Gradient Programmed Network framework for multispectral object detection. CFGPNet uses an improved GELAN backbone with RepViT-style re-parameterized blocks to strengthen feature representation while preserving computational efficiency. A Cross Computation Efficient Attention (CrossCEA) module is introduced to enhance cross-modal feature interaction and reduce redundant information transfer between visible and thermal branches. To generate compact and discriminative fused representations, an Attention Selection and Aggregation Fusion (ASAF) network combines dense feature aggregation with selective attention-based emphasis. Moreover, a programmable-gradient auxiliary branch is integrated into each CFGPNet variant to improve gradient delivery and optimization quality. Experiments on five public multispectral benchmarks, FLIR, M3FD, LLVIP, VEDAI, and MFAD, demonstrate that CFGPNet achieves strong and consistent performance across diverse scenes, object scales, and modality balances. In particular, the framework attains 80.7% mAP50 / 45.0% mAP50:95 on FLIR, 89.9% / 63.4% on M3FD, and 97.8% / 68.9% on LLVIP. It also reaches 83.3% / 56.9% on VEDAI and 83.4% / 61.8% on MFAD. These results show that CFGPNet is an effective, practical solution offering useful accuracy--efficiency trade-offs across three model scales. The code, data, and fine-tuned models are available at https://github.com/NimaHatami99/CFGPNet.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
C. Bao, J. Cao, Q. Hao, Y . Cheng, Y . Ning, T. Zhao, Dual-yolo architecture from infrared and visible images for object detection, Sensors 23 (6) (2023) 2934
2023
-
[2]
K. Hu, Y . He, Y . Li, J. Zhao, S. Chen, Y . Kang, Ei 2 det: Edge-guided illumination-aware interactive learning for visible-infrared object detection, IEEE Transactions on Circuits and Systems for Video Technology (2025)
2025
-
[3]
Zhang, X
J. Zhang, X. Song, Y . Li, D. Liang, Z. Zhang, J. Cai, Adaptive dual cross-attention network for multispectral object detection in autonomous driving, Expert Systems with Applications (2026) 132012
2026
-
[4]
Z. Tang, Z. Wu, M. Li, J. Wen, B. Zhang, Y . Xu, J. Li, Adaptive fine-grained fusion network for multimodal uav object detection, IEEE Transactions on Image Pro- cessing (2026)
2026
-
[5]
G. Li, G. Ren, J. Wang, M. Zhi, Z. Yu, B. Jiang, H. Guan, Q. Guo, Cross-modal edge-enhanced detector for uav- based multispectral object detection, Scientific Reports (2025)
2025
- [6]
-
[7]
S. Wang, G. Sun, L. Dong, B. Zheng, Carnet: Cross- attention guided feature reconstruction for rgbt object de- tection, Available at SSRN 5292837
-
[8]
H. Li, L. Xiao, L. Cao, D. Wu, Y . Liu, Y . Li, Y . Zhang, H. Bao, Crossmodalnet: A dual-modal object detec- tion network based on cross-modal fusion and channel interaction, Expert Systems with Applications (2025) 129677
2025
Show all 145 references
-
[9]
J. Ma, P. Hu, Dmfusion-yolov8: A difference-aware modality fusion framework for infrared–visible object detection, Signal, Image and Video Processing 19 (18) (2025) 1448. 37
2025
-
[10]
Z. Chen, Y . Qian, X. Yang, C. Wang, M. Yang, Amfd: Distillation via adaptive multimodal fusion for multi- spectral pedestrian detection, IEEE Transactions on Mul- timedia (2025)
2025
-
[11]
Zhang, K
L. Zhang, K. Dong, Y . Song, G. Zhang, G. Yan, T. Liu, Y . Wang, Y . Li, X. Li, Multimodal object detection method based on bidirectional dynamic sampling and adaptive cross-modal fusion, Optics & Laser Technology 192 (2025) 113996
2025
-
[12]
M. Cui, J. Nie, H. Sun, J. Xie, J. Cao, Y . Pang, X. Li, Multispectral remote sensing object detection via se- lective cross-modal interaction and aggregation, Neural Networks (2025) 108533
2025
-
[13]
C. Liu, X. Ma, X. Yang, Y . Zhang, Y . Dong, Como: Cross-mamba interaction and offset-guided fusion for multimodal object detection, Information Fusion 125 (2026) 103414
2026
-
[14]
H. Xu, X. Yuan, J. Wang, Y . Wang, Scvi: A semi- coupled visible-infrared small object detection method based on multimodal proposal-level probability fusion strategy, Neurocomputing (2026) 132688
2026
-
[15]
J. Hu, L. Ni, B. Peng, T. Li, Wtcafnet: A wavelet trans- form and cross-attention modality-adaptive fusion net- work for multispectral object detection, Signal Process- ing (2025) 110446
2025
-
[16]
E. Wang, J. Li, T. Yu, S. Xu, P. Qu, L. Na, Y . Chen, Y . Cheng, A uav detection in complex environments method based on cross-modal fusion of infrared and vis- ible images, Signal, Image and Video Processing 20 (2) (2026) 69
2026
-
[17]
W. Wu, X. Zhang, H. Yin, H. Zeng, C. Wei, L. Yu, Y . Zhang, Cdfnet: Cross-dimension fusion network with dual feature enhancement for multimodal object detec- tion, Expert Systems with Applications (2026) 132380
2026
-
[18]
C. Zhao, B. Mo, J. Zhao, Y . Tao, D. Zhao, Cmifdf: A lightweight cross-modal image fusion and weight- sharing object detection network framework, Infrared Physics & Technology 145 (2025) 105631
2025
-
[19]
X. Duan, X. Liu, Z. Li, J. Lei, S. Li, J. Zhang, Dfas- cma: A decoupled feature adaptive sharing and cross- modulation approach for uav multi-source object detec- tion, Information Fusion (2026) 104320
2026
-
[20]
X. Chen, S. Xu, S. Hu, X. Ma, Dgfd: A dual-graph con- volutional network for image fusion and low-light object detection, Information Fusion 119 (2025) 103025
2025
-
[21]
Cheng, H
M. Cheng, H. Huang, X. Liu, H. Mo, X. Zhao, S. Wu, Lefuse: Joint low-light enhancement and image fusion for nighttime infrared and visible images, Neurocomput- ing 626 (2025) 129592
2025
-
[22]
Y . Shao, Q. Huang, et al., Mod-yolo: Multispectral ob- ject detection based on transformer dual-stream yolo, Pattern Recognition Letters 183 (2024) 26–34
2024
-
[23]
Cheng, B
B. Cheng, B. Xu, W. Gan, Q. Wang, Multimodal collab- orative interactive soft fusion network for rgb-infrared aerial image object detection, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sens- ing 19 (2025) 3206–3218
2025
-
[24]
K. Bai, L. He, S. Ma, J. Dang, Osfgnet: Object saliency- driven feature aggregation network for multi-modal re- mote sensing object detection, IEEE Geoscience and Re- mote Sensing Letters (2026)
2026
-
[25]
X. Wang, J. Xi, F. Yang, Y . Yang, M. Li, Pfi-net: A par- allel feature interaction network for infrared and visible target detection, Pattern Recognition (2025) 113003
2025
-
[26]
M. Yuan, X. Shi, N. Wang, Y . Wang, X. Wei, Improving rgb-infrared object detection with cascade alignment- guided transformer, Information Fusion 105 (2024) 102246
2024
-
[27]
F. Yang, W. Li, L. Li, M. Yang, J. Zhang, Dwsf-net: A dynamic wavelet-based spatial-frequency fusion net- work for multispectral object detection, IEEE Transac- tions on Multimedia (2026)
2026
-
[28]
Z. Wang, T. Tian, Regional defeats global: An effi- cient regional feature fusion via convolutional architec- ture for multispectral object detection, Information Fu- sion (2026) 104110
2026
-
[29]
X. Wu, L. Wang, J. Guan, H. Ji, L. Xu, Y . Hou, A. Fei, Dhanet: Dual-stream hierarchical interaction networks for multimodal drone object detection, IEEE Transac- tions on Geoscience and Remote Sensing (2025)
2025
-
[30]
Zhang, M
J. Zhang, M. Gao, Y . Wang, C. Li, D. Fang, Z. Wei, Hffn: Hierarchical feature fusion network for rgb-infrared ob- ject detection, Infrared Physics & Technology (2026) 106429
2026
-
[31]
H. Wu, P. Yuan, W. Liu, A. Wang, L. Yu, G. Molnar, Ivd- net: An adaptive dual-branch network for small object detection in multi-modal remote sensing images, IEEE Journal of Selected Topics in Applied Earth Observa- tions and Remote Sensing (2026)
2026
-
[32]
Z. Tang, Y . Xie, T. Xu, X.-J. Wu, J. Kittler, Learning bi-directional fusion and deformation-sensitive loss for rgb-t tiny object detection, Information Fusion (2025) 103985
2025
-
[33]
W. Wu, H. Zhang, X. Zhang, H. Yin, Y . Zhang, Lightweight modal-guided cross-attention fusion net- work for visible-infrared object detection, Pattern Recognition (2026) 113350. 38
2026
-
[34]
J. Liu, J. Ni, Z. Zhang, Y . Gu, S. X. Yang, Maftnet: Mul- timodal adaptive fusion-based transformer network for infrared and visible image uav object detection, IEEE Sensors Journal (2026)
2026
-
[35]
Z. Hou, J. Zhao, X. Li, S. Ma, X. Yang, L. Pu, Mul- timodal detection transformer with multiscale cross- modal feature fusion and selective query recollection, In- frared Physics & Technology (2025) 106232
2025
-
[36]
Wang, I.-H
C.-Y . Wang, I.-H. Yeh, H.-Y . Mark Liao, Yolov9: Learn- ing what you want to learn using programmable gradient information, in: European conference on computer vi- sion, Springer, 2024, pp. 1–21
2024
-
[37]
Wang, H.-Y
C.-Y . Wang, H.-Y . M. Liao, Y .-H. Wu, P.-Y . Chen, J.-W. Hsieh, I.-H. Yeh, Cspnet: A new backbone that can en- hance learning capability of cnn, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2020, pp. 390–391
2020
-
[38]
C.-Y . Wang, A. Bochkovskiy, H.-Y . M. Liao, Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 7464–7475
2023
-
[39]
X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, J. Sun, Repvgg: Making vgg-style convnets great again, in: Pro- ceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, 2021, pp. 13733–13742
2021
-
[40]
A. Wang, H. Chen, Z. Lin, J. Han, G. Ding, Repvit: Re- visiting mobile cnn from vit perspective, in: Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 15909–15920
2024
-
[41]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255
2009
-
[42]
J. Hu, L. Shen, G. Sun, Squeeze-and-excitation net- works, in: Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2018, pp. 7132– 7141
2018
-
[43]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, At- tention is all you need, Advances in neural information processing systems 30 (2017)
2017
-
[44]
W. Luo, Y . Li, R. Urtasun, R. Zemel, Understanding the effective receptive field in deep convolutional neural net- works, Advances in neural information processing sys- tems 29 (2016)
2016
-
[45]
F. Yu, V . Koltun, Multi-scale context aggregation by dilated convolutions, arXiv preprint arXiv:1511.07122 (2015)
2015 arXiv
-
[46]
Zhang, X
X. Zhang, X. Zhou, M. Lin, J. Sun, Shufflenet: An ex- tremely efficient convolutional neural network for mo- bile devices, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6848–6856
2018
-
[47]
Touvron, M
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, H. Jégou, Going deeper with image transformers, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 32–42
2021
-
[48]
Y . Sun, B. Cao, P. Zhu, Q. Hu, Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning, IEEE Transactions on Circuits and Systems for Video Technology 32 (10) (2022) 6700–6713
2022
-
[49]
Ouyang, S
D. Ouyang, S. He, G. Zhang, M. Luo, H. Guo, J. Zhan, Z. Huang, Efficient multi-scale attention module with cross-spatial learning, in: ICASSP 2023-2023 IEEE in- ternational conference on acoustics, speech and signal processing (ICASSP), IEEE, 2023, pp. 1–5
2023
-
[50]
W. Xu, Y . Wan, W. Zhao, Ela: efficient location attention for deep convolution neural networks, Journal of Real- Time Image Processing 22 (4) (2025) 1–14
2025
-
[51]
Q. Hou, D. Zhou, J. Feng, Coordinate attention for ef- ficient mobile network design, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 13713–13722
2021
-
[52]
Abdelfattah, A
A. Abdelfattah, A. Haidar, S. Tomov, J. Dongarra, Per- formance, design, and autotuning of batched gemm for gpus, in: International Conference on High Performance Computing, Springer, 2016, pp. 21–38
2016
-
[53]
Huang, Z
G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely connected convolutional networks, in: Proceed- ings of the IEEE conference on computer vision and pat- tern recognition, 2017, pp. 4700–4708
2017
-
[54]
S. Woo, J. Park, J.-Y . Lee, I. S. Kweon, Cbam: Convo- lutional block attention module, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 3–19
2018
-
[55]
C. E. Shannon, A mathematical theory of communica- tion, The Bell system technical journal 27 (3) (1948) 379–423
1948
-
[56]
Kullback, Information theory and statistics, Courier Corporation, 1997
S. Kullback, Information theory and statistics, Courier Corporation, 1997
1997
-
[57]
Z. Wang, A. C. Bovik, H. R. Sheikh, E. P. Simoncelli, Image quality assessment: from error visibility to struc- tural similarity, IEEE transactions on image processing 13 (4) (2004) 600–612
2004
-
[58]
M. A. K. Raiaan, S. Sakib, N. M. Fahad, A. Al Mamun, M. A. Rahman, S. Shatabda, M. S. H. Mukta, A system- atic review of hyperparameter optimization techniques in 39 convolutional neural networks, Decision Analytics Jour- nal 11 (2024) 100470
2024
-
[59]
Zheng, P
Z. Zheng, P. Wang, D. Ren, W. Liu, R. Ye, Q. Hu, W. Zuo, Enhancing geometric factors in model learning and inference for object detection and instance segmen- tation, IEEE transactions on cybernetics 52 (8) (2021) 8574–8586
2021
-
[60]
X. Li, W. Wang, L. Wu, S. Chen, X. Hu, J. Li, J. Tang, J. Yang, Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection, Advances in neural information processing systems 33 (2020) 21002–21012
2020
-
[61]
Teledyne FLIR, FREE Teledyne FLIR Thermal Dataset for Algorithm Training, Available online: https://oem.flir.com/solutions/automotive/ adas-dataset-form/, (Accessed on 25 December
-
[62]
J. Liu, X. Fan, Z. Huang, G. Wu, R. Liu, W. Zhong, Z. Luo, Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse in- frared and visible for object detection, in: Proceedings of the IEEE/CVF conference on computer vision and pat- tern reco...
2022
-
[63]
X. Jia, C. Zhu, M. Li, W. Tang, W. Zhou, Llvip: A visible-infrared paired dataset for low-light vision, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 3496–3504
2021
-
[64]
Razakarivony, F
S. Razakarivony, F. Jurie, Vehicle detection in aerial imagery: A small target detection benchmark, Journal of Visual Communication and Image Representation 34 (2016) 187–203
2016
-
[65]
Zhang, E
H. Zhang, E. Fromont, S. Lefevre, B. Avignon, Multi- spectral fusion for object detection with cyclic fuse-and- refine blocks, in: 2020 IEEE International conference on image processing (ICIP), IEEE, 2020, pp. 276–280
2020
-
[66]
Everingham, L
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, A. Zisserman, The pascal visual object classes (voc) challenge, International journal of computer vision 88 (2) (2010) 303–338
2010
-
[67]
Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp
R. Girshick, Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 1440–1448
2015
-
[68]
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.- Y . Fu, A. C. Berg, Ssd: Single shot multibox detector, in: European conference on computer vision, Springer, 2016, pp. 21–37
2016
-
[69]
T.-Y . Lin, P. Goyal, R. Girshick, K. He, P. Dollár, Focal loss for dense object detection, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980–2988
2017
-
[70]
Jocher, A
G. Jocher, A. Chaurasia, A. Stoken, J. Borovec, Y . Kwon, K. Michael, J. Fang, C. Wong, Z. Yifu, D. Montes, et al., ultralytics/yolov5: v6. 2-yolov5 clas- sification models, apple m1, reproducibility, clearml and deci. ai integrations, Zenodo (2022)
2022
-
[71]
Jocher, A
G. Jocher, A. Chaurasia, J. Qiu, Ultralytics yolo,https: //github.com/ultralytics/ultralytics(2023)
2023
-
[72]
Khanam, M
R. Khanam, M. Hussain, Yolov11: An overview of the key architectural enhancements, arXiv preprint arXiv:2410.17725 (2024)
2024 arXiv
-
[73]
Z. Ge, S. Liu, F. Wang, Z. Li, J. Sun, Yolox: Exceed- ing yolo series in 2021, arXiv preprint arXiv:2107.08430 (2021)
2021 arXiv
-
[74]
Zhang, F
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, H.-Y . Shum, Dino: Detr with improved denoising anchor boxes for end-to-end object detection, arXiv preprint arXiv:2203.03605 (2022)
2022 arXiv
-
[75]
Y . Chen, X. Dai, D. Chen, M. Liu, X. Dong, L. Yuan, Z. Liu, Mobile-former: Bridging mobilenet and trans- former, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5270–5279
2022
-
[76]
X. Liu, H. Peng, N. Zheng, Y . Yang, H. Hu, Y . Yuan, Ef- ficientvit: Memory efficient vision transformer with cas- caded group attention, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14420–14430
2023
-
[77]
Zhang, J
J. Zhang, J. Lei, W. Xie, Z. Fang, Y . Li, Q. Du, Supery- olo: Super resolution assisted object detection in mul- timodal remote sensing imagery, IEEE Transactions on Geoscience and Remote Sensing 61 (2023) 1–15
2023
-
[78]
Y . Zeng, T. Liang, Y . Jin, Y . Li, Mmi-det: Exploring multi-modal integration for visible and infrared object detection, IEEE Transactions on Circuits and Systems for Video Technology 34 (11) (2024) 11198–11213
2024
-
[79]
Qingyun, H
F. Qingyun, H. Dapeng, W. Zhaokui, Cross-modality fu- sion transformer for multispectral object detection, arXiv preprint arXiv:2111.00273 (2021)
2021 arXiv
-
[80]
J. Shen, Y . Chen, Y . Liu, X. Zuo, H. Fan, W. Yang, Ica- fusion: Iterative cross-attention guided feature fusion for multispectral object detection, Pattern Recognition 145 (2024) 109913
2024
-
[81]
Y . Cao, J. Bin, J. Hamari, E. Blasch, Z. Liu, Multi- modal object detection by channel switching and spatial attention, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 403–411. 40
2023
-
[82]
H. Fu, S. Wang, P. Duan, C. Xiao, R. Dian, S. Li, Z. Li, Lraf-net: Long-range attention fusion network for visible–infrared object detection, IEEE Transactions on Neural Networks and Learning Systems 35 (10) (2023) 13232–13245
2023
-
[83]
H. Fu, H. Liu, J. Yuan, X. He, J. Lin, Z. Li, Yolo- adaptor: a fast adaptive one-stage detector for non- aligned visible-infrared object detection, IEEE Transac- tions on Intelligent Vehicles (2024)
2024
-
[84]
S. Lee, J. Park, J. Park, Crossformer: Cross-guided at- tention for multi-modal object detection, Pattern Recog- nition Letters 179 (2024) 144–150
2024
-
[85]
J. Jang, J. Lee, J. Paik, Camdet: Condition-adaptive mul- tispectral object detection using a visible-thermal trans- lation model, in: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2025, pp. 1–5
2025
-
[86]
C. Sun, Y . Chen, X. Qiu, R. Li, L. You, Mrd-yolo: A multispectral object detection algorithm for complex road scenes, Sensors 24 (10) (2024) 3222
2024
-
[87]
Y . Zhao, Y . Gao, X. Yang, L. Yang, Multispectral target detection based on deep feature fusion of visible and in- frared modalities, Applied Sciences 15 (11) (2025) 5857
2025
-
[88]
X. He, T. Yang, T. Yan, H. Li, Y . Ge, Z. Ren, Z. Liu, J. Jiang, C. Tang, Efficient layer-wise cross-view calibra- tion and aggregation for multispectral object detection, Electronics 15 (3) (2026) 498
2026
-
[89]
X. Yang, Y . Liu, W. Pan, G. Chu, J. Zhang, J. Zhao, Z. Man, X. Cao, M2i2ha: A multi-modal object detec- tion method based on intra-and inter-modal hypergraph attention, arXiv preprint arXiv:2601.14776 (2026)
2026 arXiv
-
[90]
W. Xu, Y . Yang, Jfdet: Joint fusion and detection for multimodal remote sensing imagery, Remote Sensing 18 (1) (2026) 176
2026
-
[91]
J. Jang, C. Park, H. Kim, J. Lee, J. Paik, Multispec- tral object detection enhanced by cross-modal informa- tion complementary and cosine similarity channel re- sampling modules, in: 2025 IEEE/CVF Winter Con- ference on Applications of Computer Vision (W ACV), IEEE, 2025, pp....
2025
-
[92]
Zhang, L
Q. Zhang, L. Tabaro, A. M. Abdelmoniem, J. An, Dlrmamba: Distilling low-rank mamba for edge multispectral fusion object detection, arXiv preprint arXiv:2603.06920 (2026)
2026
-
[93]
K. Li, Z. Zhong, Z. Luo, H. Tian, K. Wang, H. Jiang, D. Xiang, W. Tang, Fcat: Frequency-domain cross- attention for all-weather multispectral object detection in low-altitude uav security inspection of urban and indus- trial areas, Remote Sensing 18 (5) (2026) 826
2026
-
[94]
J. Li, C. Sui, J. Wang, J. Zhou, Pmdet: Patch-aware en- hancement and fusion for multispectral object detection, Remote Sensing 18 (7) (2026) 1068
2026
-
[95]
Y . Chen, J. Ye, X. Wan, Tf-yolo: a transformer–fusion- based yolo detector for multimodal pedestrian detection in autonomous driving scenes, World Electric Vehicle Journal 14 (12) (2023) 352
2023
-
[96]
Zhang, H
Y . Zhang, H. Yu, Y . He, X. Wang, W. Yang, Illumination-guided rgbt object detection with inter-and intra-modality fusion, IEEE Transactions on Instrumen- tation and Measurement 72 (2023) 1–13
2023
-
[97]
F. Yang, B. Liang, W. Li, J. Zhang, Multidimensional fusion network for multispectral object detection, IEEE Transactions on Circuits and Systems for Video Technol- ogy (2024)
2024
-
[98]
H. Zhu, Y . Chen, X. Pan, Y . He, J. Wu, Modality-guided feature alignment and complementary enhancement for infrared-visible object detection, IEEE Transactions on Circuits and Systems for Video Technology (2026)
2026
-
[99]
F. Liu, C. Gao, F. Chen, P. Li, J. Guo, D. Meng, A fusion- enhanced network for infrared and visible high-level vi- sion tasks, IEEE Transactions on Multimedia (2025)
2025
-
[100]
K. Li, D. Wang, Z. Hu, S. Li, W. Ni, L. Zhao, Q. Wang, Fd2-net: Frequency-driven feature decomposition net- work for infrared-visible object detection, in: Proceed- ings of the AAAI Conference on Artificial Intelligence, V ol. 39, 2025, pp. 4797–4805
2025
-
[101]
W. Wu, X. Zhang, H. Yin, S. Dai, H. Zhang, Y . Zhang, Fredft: Frequency domain fusion transformer for visible-infrared object detection, arXiv preprint arXiv:2511.10046 (2025)
2025
-
[102]
H. Yang, J. Fang, Y . Zhu, X. Zhao, Y . Guo, X. Zhang, X. Hu, X. Yang, Q. Ming, Crossweaver: Towards efficient cross-modal interweaving and decoupling for weakly-aligned multispectral object detection, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vi- sion and Patte...
2026
-
[103]
S. Peng, R. Xue, Y . Tong, Z. Wang, H. Yang, Multispec- tral object detection via edge-enhanced and frequency- aware fusion network, IEEE Journal of Selected Top- ics in Applied Earth Observations and Remote Sensing (2025)
2025
-
[104]
Y . Shao, T. Shi, Representation space constrained learn- ing with modality decoupling for multimodal object de- tection, arXiv preprint arXiv:2511.15433 (2025)
2025
-
[105]
H. Xiao, J. Zhuang, B. Yang, J. Hu, J. Zhu, Z. Lian, S. Zhao, Y . Zhou, Saff: A spatially-aware fusion frame- work for effective and efficient aerial object detection, IEEE Transactions on Geoscience and Remote Sensing (2026). 41
2026
-
[106]
Huang, Z
K. Huang, Z. Zhang, C. Shi, L. Luo, J. Shi, Y . Liu, De- tection drives an end-to-end fusion of infrared and visible images based on diffusion models, IEEE Transactions on Image Processing (2026)
2026
-
[107]
L. Lu, S. Zhang, Y . Gu, B. Liu, B. Liu, A cross-modal hierarchical enhanced fusion method for object detection in intelligent transportation systems, Digital Signal Pro- cessing (2026) 106094
2026
-
[108]
C. Yuan, W. Li, Q. Zhou, Q. Li, C. Li, T. Gong, A fusion-enhanced infrared-visible object detection net- work with target difference sensitivity, Neurocomputing (2026) 133640
2026
-
[109]
Z. Zhu, X. Song, G. Zhou, A. Gong, Q. Zhu, Acse: Ad- vantage complementation and salient feature-enhanced fusion for visible-infrared object detection, Journal of King Saud University Computer and Information Sci- ences (2026)
2026
-
[110]
Y . Hu, H. Jin, Modality-aware fusion and selection for robust multispectral pedestrian detection (2026)
2026
-
[111]
H. Wang, X. Liu, Mutual distillation attribute fusion network for multimodal vehicle object detection, IEEE Transactions on Vehicular Technology (2026)
2026
-
[112]
X. Wang, X. Chen, H. Fan, W. Ren, S. Wang, Y . Tang, L. Liu, Z. Han, Seeing only the focus: Rgb-t object- aware region enhancement for object detection in harsh environments, IEEE Transactions on Multimedia (2026)
2026
-
[113]
Y . Liu, G. Sun, Q. Guo, Y . Wu, Df-net: A dual-modal de- tection network based on decoupled deformable dilated convolution and progressive coordinated fusion, IEEE Sensors Journal (2026)
2026
-
[114]
C. Yang, X. Zhang, Y . Xiao, F. Meng, Wd-fqdet: Mul- tispectral detection transformer via wavelet decomposi- tion and frequency-aware query learning, arXiv preprint arXiv:2605.13621 (2026)
2026 arXiv
-
[115]
P. Lu, S. Yu, Z. Li, Alignfree-net: A registration-free hybrid framework for infrared-visible object detection via semantic feature matching, in: 2026 11th Interna- tional Conference on Intelligent Computing and Signal Processing (ICSP), IEEE, 2026, pp. 2107–2111
2026
-
[116]
H. R. Medeiros, F. A. G. Pena, M. Aminbeidokhti, T. Dubail, E. Granger, M. Pedersoli, Hallucidet: hallu- cinating rgb modality for person detection through priv- ileged information, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp....
2024
-
[117]
Z. Wang, F. Colonnier, J. Zheng, J. Acharya, W. Jiang, K. Huang, Tirdet: Mono-modality thermal infrared ob- ject detection based on prior thermal-to-visible transla- tion, in: Proceedings of the 31st ACM International Con- ference on Multimedia, 2023, pp. 2663–2672
2023
-
[118]
Z. Wang, H. Shen, W. Jiang, K. Huang, A fourier- transform-based framework with asymptotic attention for mobile thermal infrared object detection, IEEE Sen- sors Journal 24 (13) (2024) 21012–21024
2024
-
[119]
H. Liu, F. Jin, H. Zeng, H. Pu, B. Fan, Image en- hancement guided object detection in visually degraded scenes, IEEE transactions on neural networks and learn- ing systems 35 (10) (2023) 14164–14177
2023
-
[120]
L. Tang, Z. Chen, J. Huang, J. Ma, Camf: An inter- pretable infrared and visible image fusion network based on class activation mapping, IEEE Transactions on Mul- timedia 26 (2023) 4776–4791
2023
-
[121]
W. Zhao, S. Xie, F. Zhao, Y . He, H. Lu, Metafusion: Infrared and visible image fusion via meta-feature em- bedding from object detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13955–13965
2023
-
[122]
Z. Zhao, H. Bai, Y . Zhu, J. Zhang, S. Xu, Y . Zhang, K. Zhang, D. Meng, R. Timofte, L. Van Gool, Ddfm: De- noising diffusion model for multi-modality image fusion, in: Proceedings of the IEEE/CVF international confer- ence on computer vision, 2023, pp. 8082–8093
2023
-
[123]
W. Dong, H. Zhu, S. Lin, X. Luo, Y . Shen, G. Guo, B. Zhang, Fusion-mamba for cross-modality object de- tection, IEEE Transactions on Multimedia (2025)
2025
-
[124]
Huang, W
Z. Huang, W. Li, Y . Zhang, J. Guo, J. Zheng, G. Ji, Y . Tao, Cdfit: A transformer using cross-modal dual- stream feature interaction for multispectral pedestrian detection, IEEE Transactions on Intelligent Transporta- tion Systems (2026)
2026
-
[125]
F. Meng, A. Hong, H. Tang, G. Tong, Fqdnet: A fusion- enhanced quad-head network for rgb-infrared object de- tection, Remote Sensing 17 (6) (2025) 1095
2025
-
[126]
H. Wang, L. Jin, G. Wang, W. Liu, Q. Shi, Y . Hou, J. Liu, Rgb-fir multimodal pedestrian detection with cross- modality context attentional model, Sensors 25 (13) (2025) 3854
2025
-
[127]
J. Gong, Z. Yuan, W. Li, W. Li, Y . Guo, B. Guo, A lightweight upsampling and cross-modal feature fusion- based algorithm for small-object detection in uav im- agery, Electronics 15 (2) (2026) 298
2026
-
[128]
C. Chai, Y . Wang, X. Hu, S. Han, Efficient multi-spectral pedestrian detection network for occlusion and far dis- tance targets, Available at SSRN 5312893
-
[129]
R. Lan, Y . Zhang, Z. Wang, W. Liu, R. Yang, Kdet-hpfl: A personalized federated learning framework for multi- modal pedestrian detection with adaptive feature selec- tion, IEEE Internet of Things Journal (2025). 42
2025
-
[130]
T. Weng, X. Niu, Lmdenet: A lightweight rgb-ir object detection network for low-light remote sensing images, Sensors 26 (4) (2026) 1130
2026
-
[131]
X. Zhai, F. Xu, Y . Pan, Mcff-det: Multispectral coarse- to-fine fusion for object detection, in: 2025 IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP), IEEE, 2025, pp. 1–6
2025
-
[132]
R. Chen, H. Sun, H. Fan, P. Wu, Vif-yolo: Visible- infrared fusion yolo for remote sensing small object de- tection, in: 2025 IEEE 31th International Conference on Parallel and Distributed Systems (ICPADS), IEEE, 2025, pp. 1–8
2025
-
[133]
X. Guo, F. Yang, L. Ji, Def-net: A dual-modal feature enhancement and fusion network for infrared and visible object detection, PloS one 21 (4) (2026) e0345815
2026
-
[134]
H. Li, J. Ma, J. Zhang, H. Luo, H. Zuo, Y . Wu, T. Tan, Hyperdet: Cross-modal hypergraph fusion for enhanced object detection in uav imagery, Journal of King Saud University Computer and Information Sciences (2026)
2026
-
[135]
Q. Hu, H. Yu, Z. Zhou, S. Li, Iaf-rtdetr: Illumina- tion evaluation-driven multimodal object detection net- work for infrared–visible dual-source fusion, Electronics 15 (6) (2026) 1332
2026
-
[136]
Zhang, G
Z. Zhang, G. Lv, Y . Gao, G. Ma, Mdsf-det: Modality decoupling and synergistic fusion detector, in: ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2026, pp. 10377–10381
2026
-
[137]
C. Guo, T. Zhao, Z. Zhang, Yolo-msff: An improved yolo11-based multispectral feature fusion method for object detection, in: 2026 International Conference on Generative Artificial Intelligence and Information Secu- rity (GAIIS), IEEE, 2026, pp. 574–583
2026
-
[138]
H. Wu, J. Rong, Q. Zhou, D. Yang, W. Wang, A guided fusion network based on cross-scale semantic alignment for multi-spectral object detection, Pattern Recognition (2026) 114024
2026
-
[139]
X. Xue, H. Chen, K. Song, Y . Yan, B. Li, Q. Meng, Transfer graph reasoning network for misaligned visible- thermal object detection, Knowledge-Based Systems (2026) 116271
2026
-
[140]
L. Yu, Z. Han, Edge-guided progressive feature comple- mentary network for visible–infrared remote sensing ob- ject detection, International Journal of Remote Sensing (2026) 1–28
2026
-
[141]
H. Guo, C. Sun, J. Zhang, W. Zhang, N. Zhang, Mmyfnet: Multi-modality yolo fusion network for ob- ject detection in remote sensing images, Remote Sensing 16 (23) (2024) 4451
2024
-
[142]
Z. Wu, F. Zhang, B. Yang, Multimodal uav object de- tection based on progressive bi-directional attention and relational bayesian constraints (2026)
2026
-
[143]
Cheng, Y
Q. Cheng, Y . Jiang, Y . Gao, Y . Qiu, Y . Tang, X. Tu, Yolo- ch: A cross-modal feature interaction and screening- based dual-stream network for uav small object detec- tion, Drones 10 (5) (2026) 350
2026
-
[144]
Jocher, A
G. Jocher, A. Chaurasia, A. Stoken, J. Borovec, Y . Kwon, K. Michael, J. Fang, Z. Yifu, C. Wong, D. Montes, et al., ultralytics/yolov5: v7. 0-yolov5 sota realtime instance segmentation, Zenodo (2022)
2022
-
[145]
A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, G. Ding, Yolov10: Real-time end-to-end object detec- tion, Advances in neural information processing systems 37 (2024) 107984–108011. Nima Hatamireceived the B.Sc. de- gree in Biomedical Engineering and the M.Sc. degree in Ele...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.