Pith. sign in

REVIEW 4 major objections 6 minor 154 references

A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper's benchmark experiments claim that YOLOv11m achieves the highest vehicle detection accuracy on two aerial datasets, while the nano-scale YOLO11n consumes the least energy per frame on an embedded drone platform.

desk verdict Useful survey, but its central experimental claim is contradicted by its own tables: an undefined YOLOv26(m) outranks YOLOv11m. read the letter →

arxiv 2608.07571 v1 pith:G5LWZBRX submitted 2026-08-04 cs.CV

classification cs.CV
keywords UAVtrafficmonitoringvehicledetectionYOLOdeeplearningenergyefficiencyaerialbenchmarkVisDroneAMVD
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey reviews deep-learning vehicle detection for drone-based traffic monitoring and adds its own head-to-head experiments. On the VisDrone benchmark, YOLOv11m reaches the highest mAP@0.5 at 46.3%, and on the high-altitude Cyprus dataset AMVD it reaches 95.1% mAP@0.5 with 20.0 million parameters, beating all tested variants including YOLOv12m. On an embedded edge platform, the nano-scale YOLO11n sustains 43.7 frames per second at 0.229 joules per frame under a 10-watt power limit, making it the paper's recommended choice for energy-constrained, long-endurance flights. The paper concludes that medium-sized YOLO detectors currently offer the best accuracy-efficiency trade-off for UAV traffic monitoring, while nano models are preferable when battery life is the binding constraint.

What carries the argument

The machinery is the standardized comparative benchmark: both aerial datasets resized to 640 by 640 pixels, training on four Tesla V100 GPUs under a fixed PyTorch environment, evaluation with mAP@0.5 and mAP@0.5:0.95, and edge inference on the Jetson Orin NX across four power modes using the energy-per-frame metric $E_f = P_{\text{limit}} / \text{FPS}$. The central objects are the YOLO detector family spanning nano to medium scales, together with the two complementary benchmarks, VisDrone and AMVD, chosen for their contrasting altitudes and geographies.

What would settle it

Re-running the stated benchmark protocol on VisDrone2019-val with an official YOLOv26(m) checkpoint, or confirming that no such official checkpoint exists, would settle whether the reported 47.8 mAP@0.5 and the resulting ranking are reproducible.

Watch

Extended reading notes

Core claim

The paper's central empirical claim is a quantitative map of the accuracy-latency-energy trade-off for current YOLO detectors in UAV traffic monitoring. Trained identically at 640 by 640 pixels on VisDrone2019-val and AMVD, YOLOv11m outperforms all compared detectors, including YOLOv12m and a model listed as YOLOv26(m), while YOLO11n provides the lowest energy per frame on the Jetson Orin NX edge platform. The authors interpret this as evidence that model scale, not mere recency, drives deployable performance, and that deployment planning should be guided by an energy-per-frame metric defined as the power limit divided by the inference speed.

Load-bearing premise

The benchmark rankings rest on a model called YOLOv26(m) that appears only in the results tables with no description, citation, or training details anywhere in the text, so the paper gives no way to verify its existence or that it was trained under the same conditions as the other detectors.

Editorial extensions

If this is right

  • For precision-critical traffic tasks, a medium-scale detector such as YOLOv11m appears to be the current best choice; for battery-limited long-endurance flights, a nano-scale model such as YOLO11n is preferable.
  • Future evaluations of drone-deployed detectors should report energy-per-frame alongside accuracy, since frames per second alone overstates viability on power-constrained hardware.
  • Performance gains in recent YOLO versions concentrate on small and occluded object categories, indicating that progress in aerial detection is driven by improvements relevant to high-altitude, small-object scenarios.
  • The same standardized benchmark protocol can be applied to newer detector families beyond the YOLO lineage, providing a direct comparison path for future architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The results tables include a model called YOLOv26(m) with specific accuracy, parameter, and speed numbers but no description or citation anywhere in the text; this should be verified before relying on the ranking, since the paper's model-evolution discussion rests on it.
  • Combining the recommended YOLOv11m with slice-aided inference, which the paper cites as improving small-object average precision by up to 14.5%, would likely yield larger gains than model swaps alone.
  • The energy-per-frame metric could be extended to whole-system budgets that include data transmission and multi-UAV coordination, factors the paper lists as open challenges but does not quantify.
  • Because the two datasets are geographically narrow and mostly captured in clear daytime conditions, the rankings may not transfer to fog, rain, or night operations that the paper itself identifies as underexplored.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript is a survey and experimental benchmarking study of vision-based vehicle detection for UAV-based traffic monitoring. It reviews architectural components of detectors and YOLO variants, describes a range of aerial datasets, and reports new comparative experiments on the VisDrone and Aerial Multi-Vehicle Detection (AMVD) datasets, together with an edge-deployment energy benchmark on an NVIDIA Jetson Orin NX. The claimed main findings are that medium YOLO variants, especially YOLOv11m, offer the best accuracy-efficiency trade-off for UAV deployment, while the nano variant YOLOv11n is the most energy-efficient choice for power-limited operations.

Significance. If fully supported, the paper would provide a practically useful model-selection reference for UAV traffic monitoring, combining accuracy, parameter count, latency, and energy consumption on both a public benchmark and a real-world aerial dataset. The survey component is broad and reasonably current, and the edge benchmarking across multiple power modes addresses a genuine deployment need. However, the central benchmark conclusions are not currently supportable because an undocumented model appears in the main tables and several textual claims are contradicted by those same tables; the experimental protocol also lacks key reproducibility information. The paper's significance can therefore only be assessed after the experiments are corrected, documented, and made internally consistent.

major comments (4)
  1. [Section VI-B and Section VI-C, Tables III-V] The headline claims that 'YOLOv11m achieved the highest mAP@0.5 at 46.3%' on VisDrone and that 'the YOLOv11m architecture emerged as the superior candidate' on AMVD are contradicted by the paper's own tables. Table III and Table IV report YOLOv26(m) at 47.8% mAP@0.5 on VisDrone, and Table V reports YOLOv26(m) at 95.7% on AMVD, both above the corresponding YOLOv11m values. YOLOv26(m) is never cited, described, referenced, or given a training configuration anywhere outside these tables. Please either remove the YOLOv26 entries or provide a proper citation, architecture description, and training setup, and then revise all 'highest' and 'superior' statements so that they are consistent with the final tables.
  2. [Section VI-A] The reproducibility of all benchmark results is not established. The text lists hardware and software versions but gives no learning rate, optimizer, batch size, number of epochs, augmentation pipeline, data split, or random seeds, and no table reports error bars or repeated-run statistics. This matters because several headline comparisons are close, for example YOLOv11m versus YOLOv12m on VisDrone (46.3 vs. 45.7) and on AMVD (95.1 vs. 94.6). Please report the full training protocol, and either provide variance over multiple runs or explicitly state that the numbers are single-run point estimates.
  3. [Section VI-D, E_f definition] The energy-per-frame metric is defined as E_f = P_limit/FPS, which assumes the platform draws exactly the configured power limit. On an embedded GPU with dynamic voltage and frequency scaling, actual power draw is typically below the power cap and varies with utilization, so this metric can systematically misstate J/Frame and can change efficiency rankings. Please measure or report actual power consumption, or relabel the quantity as an upper-bound proxy and discuss the limitation in the deployment analysis.
  4. [Table V, YOLOv9(c) row] The YOLOv9(c) row in Table V lists mAP@0.5 = 90.2 while the class APs are Car 97.0, Bus 92.8, and Truck 90.2; the mean of these class APs is approximately 93.3, which is inconsistent with the reported mAP. Since this table directly supports the benchmark conclusion, the entry must be corrected or the computation explained, for example if mAP is computed over a different class set or with different rounding.
minor comments (6)
  1. [Section VII-A-7] The last two paragraphs of this subsection are duplicated: 'Real-Time Processing Constraints' is immediately followed by a second paragraph beginning 'Real-time vehicle detection is crucial for UAV applications...' with nearly identical content. One copy should be removed.
  2. [Section V-A and Table II] The VisDrone statistics are inconsistent between the text, which reports 288 video clips, 261,908 frames, 10,209 static images, and more than 2.6 million marked boxes, and Table II, which reports '10k images, 263 videos' and '2.5M' instances. Please use consistent and clearly defined rounded values.
  3. [Tables IV-VI and text] Notation is inconsistent: Table VI uses 'YOLO11n/s/m' while the text and other tables use 'YOLOv11...'; similarly, 'FLOPS' and 'FLOPs' are used interchangeably. Please standardize terminology throughout.
  4. [Section V-B-1] The pNeuma description states that trajectories 'are captured along with the same frequency of 0.04-seconds computed as the highest frame rate of the video'; this sentence is unclear and should be rewritten to state the actual sampling rate.
  5. [Section V and Section VI] The AMVD dataset was created by co-authors of this paper (reference [76]), and the evaluation protocol on it is not specified. This is not improper, but the paper should explicitly disclose the relationship and state the exact train/validation split and evaluation subset used for Table V.
  6. [Section VI-B] The sentence 'Our experiments results align with previous work on [98]–[100]' is vague and grammatically awkward; please specify which results are being compared and how the cited works support that comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation found: the benchmark results are externally measured on public datasets, and the paper's self-citations are not load-bearing.

full rationale

The paper's load-bearing empirical content is the comparative benchmark in Section VI. The claimed accuracies (e.g., YOLOv11m at 46.3% mAP@0.5 on VisDrone and 95.1% on AMVD) are presented as measured outcomes of training and evaluation runs on external benchmarks, not as quantities derived from the claims they are meant to support. The experimental setup in Section VI-A fixes hardware, image size, and training stack, and the tables report raw mAP, parameters, FLOPs, FPS, and J/Frame values; none of these is a fitted parameter renamed as a prediction. The energy-per-frame metric E_f = P_limit / FPS is a definitional transformation of two measured quantities, not a circular prediction. Self-citations to EdgeNet, DroNet, STVD, and AMVD describe the authors' prior detectors and datasets, but they do not supply the asserted rankings; even the AMVD dataset, co-created by one of the authors, is an independent public benchmark whose annotations were not constructed to force any particular YOLO variant to win. No uniqueness theorem or prior result is imported to forbid alternative model choices. Separately, the appearance of an undescribed YOLOv26(m) in Tables III-VI with scores above the claimed 'highest' YOLOv11m is a verifiability and consistency defect, and the unsupported 'highest mAP' statements are questionable, but this is not circularity: the table entries are not defined in terms of the conclusions. Hence no circular step is present.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The central claims depend on standard training and evaluation assumptions for object detection benchmarks. The only unusual entity is the referenced model YOLOv26, which is treated as a known object but is given no citation, architecture description, or external source; this is better classified as an unverifiable artifact than as an invented entity. The AMVD dataset is a co-author-created resource and is publicly available via DOI, so it is not an invented entity.

free parameters (1)
  • Training hyperparameters (learning rate, epochs, batch size, augmentation) = not reported
    The mAP and FPS numbers depend on these choices, but the paper does not state them. Without this information, the benchmark results cannot be reproduced or independently verified.
assumptions (2)
  • domain assumption VisDrone and AMVD annotations are accurate ground truth for vehicle detection.
    All mAP calculations treat the provided bounding boxes as perfect references; any label noise would bias the reported accuracies.
  • domain assumption Resizing all images to 640x640 preserves enough detail for fair cross-model comparison.
    The experiments standardize input resolution at 640x640, but the paper does not analyze whether this resolution penalizes small-object detection differently across models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions." pith.science (2026). https://pith.science/paper/G5LWZBRX

@misc{pith2026260807571,
  author       = {Pith},
  title        = {Pith review of: A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G5LWZBRX}},
  note         = {Machine review of arXiv:2608.07571}
}
read the original abstract

In Intelligent Transportation System (ITS), unmanned aerial vehicle (UAV)-based surveillance offers an innovative solution to traffic surveillance with wide coverage and real-time data collection capabilities. In comparison to fixed ground-based infrastructure, UAVs are able to respond to dynamic traffic but present challenges such as vehicle detection at varying altitudes, compensation for motion-induced image variations and efficient processing of high-resolution images. Deep learning has been largely beneficial on improving the detection accuracy; however, for practical deployment, a critical assessment of the accuracy, latency, and harmonization with current transportation systems needs to be carefully considered. This survey reviews recent advancements in the UAV-based traffic monitoring, with a primary focus being deep neural network models for traffic analytics in various urban settings. Three main challenges identified in the literature are ensuring compatibility with traffic control systems, achieving real-time processing to optimize traffic flow, and maintaining robust detection in different environmental conditions. Existing solutions often lack comprehensive frameworks for utilizing UAV captured data to respond to incidents and manage traffic effectively. Future research should focus on optimal detection models, edge processing, and adaptive control integration to improve the responsiveness of urban traffic management.

Figures

Figures reproduced from arXiv: 2608.07571 by the authors.

Figure 1
Figure 1. Overall workflow of the UAV-based intelligent traffic monitoring system. Multiple UAV nodes acquire real-time images of road intersections, perform [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A generic architectural workflow for vision-based UAV traffic monitoring. The process begins with preprocessing to handle high-resolution inputs. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. (a) Examples images from the VisDrone Dataset [75], this collection [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Illustration of multi-camera multi-vehicle tracking in UAV-based traffic [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Model selection of vehicle detection based on different illumination [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Demonstration of Vision Language Models (VLMs) in traffic monitor [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

154 extracted references · 63 canonical work pages

  1. [1]

    Data- driven intelligent transportation systems: A survey,

    J. Zhang, F.-Y . Wang, K. Wang, W.-H. Lin, X. Xu, and C. Chen, “Data- driven intelligent transportation systems: A survey,”IEEE Transactions on Intelligent Transportation Systems, vol. 12, no. 4, pp. 1624–1639, 2011

  2. [2]

    Explainable ai for safe and trustworthy autonomous driving: A system- atic review,

    A. Kuznietsov, B. Gyevnar, C. Wang, S. Peters, and S. V . Albrecht, “Explainable ai for safe and trustworthy autonomous driving: A system- atic review,”IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 12, pp. 19 342–19 364, 2024

  3. [3]

    Navigation of a uav network for optimal surveillance of a group of ground targets moving along a road,

    A. V . Savkin and H. Huang, “Navigation of a uav network for optimal surveillance of a group of ground targets moving along a road,”IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 7, pp. 9281–9285, 2021

  4. [4]

    UA V-enabled intelligent transportation systems for the smart city: Applications and challenges,

    H. Menouar, I. Guvenc, K. Akkaya, A. S. Uluagac, A. Kadri, and A. Tuncer, “UA V-enabled intelligent transportation systems for the smart city: Applications and challenges,”IEEE Communications Mag- azine, vol. 55, no. 3, pp. 22–28, 2017. 17

  5. [5]

    Applications of unmanned aerial vehicle (uav) in road safety, traffic and highway infrastructure management: Recent advances and challenges,

    F. Outay, H. A. Mengash, and M. Adnan, “Applications of unmanned aerial vehicle (uav) in road safety, traffic and highway infrastructure management: Recent advances and challenges,”Transportation Re- search Part A: Policy and Practice, vol. 141, pp. 116–129, 2020

  6. [6]

    A study on vehicle detection through aerial images: Various challenges, issues and applications,

    S. Kumar, E. Rajan, and S. Rani, “A study on vehicle detection through aerial images: Various challenges, issues and applications,” in2021 International Conference on Computing, Communication, and Intelligent Systems (ICCCIS). IEEE, 2021, pp. 504–509

  7. [7]

    Advances of uavs toward future transportation: The state-of-the-art, challenges, and opportuni- ties,

    A. Gupta, T. Afrin, E. Scully, and N. Yodo, “Advances of uavs toward future transportation: The state-of-the-art, challenges, and opportuni- ties,”Future Transportation, vol. 1, no. 2, pp. 326–350, 2021

  8. [8]

    Vehicle detection from uav imagery with deep learning: A review,

    A. Bouguettaya, H. Zarzour, A. Kechida, and A. M. Taberkit, “Vehicle detection from uav imagery with deep learning: A review,”IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 11, pp. 6047–6067, 2021

Show all 154 references
  1. [9]

    Machine learning for UA V- aided ITS: A review with comparative study,

    A. Telikani, A. Sarkar, B. Du, and J. Shen, “Machine learning for UA V- aided ITS: A review with comparative study,”IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 11, pp. 15 388–15 406, 2024

  2. [10]

    UA V video vehicle detection: Benchmark and baseline,

    Y . Xiao, J. Wang, Z. Zhao, B. Jiang, C. Li, and J. Tang, “UA V video vehicle detection: Benchmark and baseline,”IEEE Transactions on Geoscience and Remote Sensing, vol. 63, p. 5609814, 2025

  3. [11]

    A systematic review of drone based road traffic monitoring system,

    I. Bisio, C. Garibotto, H. Haleem, F. Lavagetto, and A. Sciarrone, “A systematic review of drone based road traffic monitoring system,”IEEE Access, vol. 10, pp. 101 537–101 555, 2022

  4. [12]

    Unmanned aerial vehicle-aided intelligent transportation systems: Vision, challenges, and opportunities,

    A. Telikani, A. Sarkar, B. Du, F. Santoso, J. Shen, J. Yan, J. Yong, and E. Yap, “Unmanned aerial vehicle-aided intelligent transportation systems: Vision, challenges, and opportunities,”IEEE Communications Surveys & Tutorials, vol. PP, pp. 1–1, 2025, early Access

  5. [13]

    Developing a more reliable framework for extracting traffic data from a uav video,

    X. Li and J. Wu, “Developing a more reliable framework for extracting traffic data from a uav video,”IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 11, pp. 12 272–12 283, 2023

  6. [14]

    Efficient joint deployment of multi-UA Vs for target tracking in traffic big data,

    L. Sun, J. Wang, J. Wang, L. Lin, and M. Gen, “Efficient joint deployment of multi-UA Vs for target tracking in traffic big data,”IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 7, pp. 7780–7791, 2024

  7. [15]

    On-demand routing for urban V ANETs using cooperating UA Vs,

    O. S. Oubbati, N. Chaib, A. Lakas, and S. Bitam, “On-demand routing for urban V ANETs using cooperating UA Vs,” in2018 Interna- tional Conference on Smart Communications in Network Technologies (SaCoNeT). IEEE, 2018, pp. 108–113

  8. [16]

    You only look once: Unified, real-time object detection,

    J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 779–788

  9. [17]

    Combining monocular and stereo- vision for real-time vehicle ranging and tracking on multilane high- ways,

    S. Sivaraman and M. M. Trivedi, “Combining monocular and stereo- vision for real-time vehicle ranging and tracking on multilane high- ways,” in2011 14th International IEEE Conference on Intelligent Transportation Systems (ITSC). IEEE, 2011, pp. 1249–1254

  10. [18]

    Urban traffic monitoring and analysis using unmanned aerial vehicles (uavs): A systematic literature review,

    E. V . Butil ˘a and R. G. Boboc, “Urban traffic monitoring and analysis using unmanned aerial vehicles (uavs): A systematic literature review,” Remote Sensing, vol. 14, no. 3, p. 620, 2022

  11. [19]

    A survey of deep learning techniques for vehicle detection from uav images,

    S. Srivastava, S. Narayan, and S. Mittal, “A survey of deep learning techniques for vehicle detection from uav images,”Journal of Systems Architecture, vol. 117, p. 102152, 2021

  12. [20]

    A review on deep learning in uav remote sensing,

    L. P. Osco, J. M. Junior, A. P. M. Ramos, L. A. de Castro Jorge, S. N. Fatholahi, J. de Andrade Silva, E. T. Matsubara, H. Pistori, W. N. Gonc ¸alves, and J. Li, “A review on deep learning in uav remote sensing,”International Journal of Applied Earth Observation and Geoinforma...

  13. [21]

    Uav assistance paradigm: State-of-the-art in applications and challenges,

    B. Alzahrani, O. S. Oubbati, A. Barnawi, M. Atiquzzaman, and D. Al- ghazzawi, “Uav assistance paradigm: State-of-the-art in applications and challenges,”Journal of Network and Computer Applications, vol. 166, p. 102706, 2020

  14. [22]

    A survey of unmanned aerial vehicles (uavs) for traffic monitoring,

    K. Kanistras, G. Martins, M. J. Rutherford, and K. P. Valavanis, “A survey of unmanned aerial vehicles (uavs) for traffic monitoring,” in2013 International Conference on Unmanned Aircraft Systems (ICUAS). IEEE, 2013, pp. 221–234

  15. [23]

    Application of deep learning method for real-time traffic analysis using uav,

    H. Park, S. Byun, and H. Lee, “Application of deep learning method for real-time traffic analysis using uav,”Journal of the Korean Society of Surveying, Geodesy, Photogrammetry and Cartography, vol. 38, no. 4, pp. 353–361, 2020

  16. [24]

    Analysis of the occlusion interference problem in target tracking,

    S. Zhang, K. Zheng, and S. Huaiyuan, “Analysis of the occlusion interference problem in target tracking,”Mathematical Problems in Engineering, vol. 2022, no. 1, p. 4605111, 2022

  17. [25]

    Performance analysis of vehicle detection algorithm in aerial traffic videos,

    S. Liu, H. Liu, W. Shi, S. Wang, M. Shi, L. Wang, and T. Mao, “Performance analysis of vehicle detection algorithm in aerial traffic videos,” in2019 International Conference on Virtual Reality and Visualization (ICVRV). IEEE, 2019, pp. 59–64

  18. [26]

    Target detection and recognition for traffic congestion in smart cities using deep learning-enabled uavs: A review and analysis,

    S. Iftikhar, M. Asim, Z. Zhang, A. Muthanna, J. Chen, M. El-Affendi, A. Sedik, and A. A. Abd El-Latif, “Target detection and recognition for traffic congestion in smart cities using deep learning-enabled uavs: A review and analysis,”Applied Sciences, vol. 13, no. 6, p. 3995, 2023

  19. [27]

    A survey of object detection for uavs based on deep learning,

    G. Tang, J. Ni, Y . Zhao, Y . Gu, and W. Cao, “A survey of object detection for uavs based on deep learning,”Remote Sensing, vol. 16, no. 1, p. 149, 2023

  20. [28]

    Deep learning for un- manned aerial vehicle-based object detection and tracking: A survey,

    X. Wu, W. Li, D. Hong, R. Tao, and Q. Du, “Deep learning for un- manned aerial vehicle-based object detection and tracking: A survey,” IEEE Geoscience and Remote Sensing Magazine, vol. 10, no. 1, pp. 91–124, 2021

  21. [29]

    A comparison of deep learning-based object detection for unmanned aerial vehicle,

    T. Li, “A comparison of deep learning-based object detection for unmanned aerial vehicle,”Applied and Computational Engineering, vol. 47, pp. 186–192, 2024

  22. [30]

    Small object detection: A comprehensive survey on challenges, techniques and real-world applications,

    M. Nikouei, B. Baroutian, S. Nabavi, F. Taraghi, A. Aghaei, A. Sajedi, and M. E. Moghaddam, “Small object detection: A comprehensive survey on challenges, techniques and real-world applications,”arXiv preprint arXiv:2503.20516, 2025

  23. [31]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  24. [32]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” inInternational Conference on Learning Representations (ICLR), 2015

  25. [33]

    Mobilenets: Efficient convolutional neural networks for mobile vision applications,

    A. G. Howard, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,”arXiv preprint arXiv:1704.04861, 2017

  26. [34]

    Edgenet: Balancing accuracy and performance for edge-based convolutional neural network object detectors,

    G. Plastiras, C. Kyrkou, and T. Theocharides, “Edgenet: Balancing accuracy and performance for edge-based convolutional neural network object detectors,” inProceedings of the 13th International Conference on Distributed Smart Cameras, 2019, pp. 1–6

  27. [35]

    Dronet: Efficient convolutional neural network detector for real-time uav applications,

    C. Kyrkou, G. Plastiras, T. Theocharides, S. I. Venieris, and C.-S. Bouganis, “Dronet: Efficient convolutional neural network detector for real-time uav applications,” in2018 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2018, pp. 967–972

  28. [36]

    Feature pyramid networks for object detection,

    T.-Y . Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2117–2125

  29. [37]

    Path aggregation network for instance segmentation,

    S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8759–8768

  30. [38]

    Efficientdet: Scalable and efficient object detection,

    M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 10 781–10 790

  31. [39]

    Ssd: Single shot multibox detector,

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y . Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” inComputer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14. Springer, 2016, pp. 21– 37

  32. [40]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137– 1149, 2016

  33. [41]

    Yolox: Exceeding yolo series in 2021,

    Z. Ge, “Yolox: Exceeding yolo series in 2021,”arXiv preprint arXiv:2107.08430, 2021

  34. [42]

    Fcos: Fully convolutional one- stage object detection,

    Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional one- stage object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 9627–9636

  35. [43]

    Efficient convnet- based object detection for unmanned aerial vehicles by selective tile processing,

    G. Plastiras, C. Kyrkou, and T. Theocharides, “Efficient convnet- based object detection for unmanned aerial vehicles by selective tile processing,” inProceedings of the 12th International Conference on Distributed Smart Cameras, 2018, pp. 1–6

  36. [44]

    Slicing aided hyper inference and fine-tuning for small object detection,

    F. C. Akyon, S. O. Altinuc, and A. Temizel, “Slicing aided hyper inference and fine-tuning for small object detection,” in2022 IEEE International Conference on Image Processing (ICIP). IEEE, 2022, pp. 966–970

  37. [45]

    Pc-yolo11s: a lightweight and effective feature extraction method for small target image detection,

    Z. Wang, Y . Su, F. Kang, L. Wang, Y . Lin, Q. Wu, H. Li, and Z. Cai, “Pc-yolo11s: a lightweight and effective feature extraction method for small target image detection,”Sensors, vol. 25, no. 2, p. 348, 2025

  38. [46]

    Lud-yolo: A novel lightweight object detection network for unmanned aerial vehicle,

    Q. Fan, Y . Li, M. Deveci, K. Zhong, and S. Kadry, “Lud-yolo: A novel lightweight object detection network for unmanned aerial vehicle,” Information Sciences, vol. 686, p. 121366, 2025

  39. [47]

    End-to-end object detection with transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in 18 European Conference on Computer Vision. Springer, 2020, pp. 213– 229

  40. [48]

    Deformable detr: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,”arXiv preprint arXiv:2010.04159, 2020

  41. [49]

    Detrs with collaborative hybrid assignments training,

    Z. Zong, G. Song, and Y . Liu, “Detrs with collaborative hybrid assignments training,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 6748–6758

  42. [50]

    Detrs beat yolos on real-time object detection,

    Y . Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y . Liu, and J. Chen, “Detrs beat yolos on real-time object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 16 965–16 974

  43. [51]

    Simple open-vocabulary object detection,

    M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shenet al., “Simple open-vocabulary object detection,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 728–755

  44. [52]

    Dino: Detr with improved denoising anchor boxes for end-to- end object detection,

    H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y . Shum, “Dino: Detr with improved denoising anchor boxes for end-to- end object detection,”arXiv preprint arXiv:2203.03605, 2022

  45. [53]

    Transformers in vision: A survey,

    S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,”ACM Computing Surveys (CSUR), vol. 54, no. 10s, pp. 1–41, 2022

  46. [54]

    Swin transformer: Hierarchical vision transformer using shifted win- dows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted win- dows,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10 012–10 022

  47. [55]

    Rf- detr: Neural architecture search for real-time detection transformers,

    I. Robinson, P. Robicheaux, M. Popov, D. Ramanan, and N. Peri, “Rf- detr: Neural architecture search for real-time detection transformers,” arXiv preprint arXiv:2511.09554, 2025

  48. [56]

    Efficientformer: Vision transformers at mobilenet speed,

    Y . Li, G. Yuan, Y . Wen, J. Hu, G. Evangelidis, S. Tulyakov, Y . Wang, and J. Ren, “Efficientformer: Vision transformers at mobilenet speed,” Advances in neural information processing systems, vol. 35, pp. 12 934–12 949, 2022

  49. [57]

    Benchmark analysis of representative deep neural network architectures,

    S. Bianco, R. Cadene, L. Celona, and P. Napoletano, “Benchmark analysis of representative deep neural network architectures,”IEEE Access, vol. 6, pp. 64 270–64 277, 2018

  50. [58]

    Simple online and realtime tracking with a deep association metric,

    N. Wojke, A. Bewley, and D. Paulus, “Simple online and realtime tracking with a deep association metric,” in2017 IEEE International Conference on Image Processing (ICIP). IEEE, 2017, pp. 3645–3649

  51. [59]

    Strongsort: Make deepsort great again,

    Y . Du, Z. Zhao, Y . Song, Y . Zhao, F. Su, T. Gong, and H. Meng, “Strongsort: Make deepsort great again,”IEEE Transactions on Multi- media, vol. 25, pp. 8725–8737, 2023

  52. [60]

    Bytetrack: Multi-object tracking by associating every detection box,

    Y . Zhang, P. Sun, Y . Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu, and X. Wang, “Bytetrack: Multi-object tracking by associating every detection box,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 1–21

  53. [61]

    Efficientnet: Rethinking model scaling for con- volutional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for con- volutional neural networks,” inInternational Conference on Machine Learning. PMLR, 2019, pp. 6105–6114

  54. [62]

    Comparison of cspdarknet53, cspresnext- 50, and efficientnet-b0 backbones on yolo v4 as object detector,

    M. Mahasin and I. A. Dewi, “Comparison of cspdarknet53, cspresnext- 50, and efficientnet-b0 backbones on yolo v4 as object detector,”Inter- national Journal of Engineering, Science and Information Technology, vol. 2, no. 3, pp. 64–72, 2022

  55. [63]

    Mobilenet-ca-yolo: An improved yolov7 based on the mobilenetv3 and attention mechanism for rice pests and diseases detection,

    L. Jia, T. Wang, Y . Chen, Y . Zang, X. Li, H. Shi, and L. Gao, “Mobilenet-ca-yolo: An improved yolov7 based on the mobilenetv3 and attention mechanism for rice pests and diseases detection,”Agri- culture, vol. 13, no. 7, p. 1285, 2023

  56. [64]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141

  57. [65]

    Cbam: Convolutional block attention module,

    S. Woo, J. Park, J.-Y . Lee, and I. S. Kweon, “Cbam: Convolutional block attention module,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 3–19

  58. [66]

    Fully convolutional one-stage 3d object detection on lidar range images,

    Z. Tian, X. Chu, X. Wang, X. Wei, and C. Shen, “Fully convolutional one-stage 3d object detection on lidar range images,”Advances in Neural Information Processing Systems, vol. 35, pp. 34 899–34 911, 2022

  59. [67]

    Focal loss for dense object detection,

    T. Lin, “Focal loss for dense object detection,”Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 2980–2988, 2017

  60. [68]

    Generalized intersection over union: A metric and a loss for bounding box regression,

    H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 658–666

  61. [69]

    Vit-yolo: Transformer-based yolo for object detection,

    Z. Zhang, X. Lu, G. Cao, Y . Yang, L. Jiao, and F. Liu, “Vit-yolo: Transformer-based yolo for object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 2799–2808

  62. [70]

    You only look one-level feature,

    Q. Chen, Y . Wang, T. Yang, X. Zhang, J. Cheng, and J. Sun, “You only look one-level feature,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 13 039–13 048

  63. [71]

    Yolov12: Attention-centric real- time object detectors,

    Y . Tian, Q. Ye, and D. Doermann, “Yolov12: Attention-centric real- time object detectors,”arXiv preprint arXiv:2502.12524, 2025

  64. [72]

    Drone-yolo: An efficient neural network method for target detection in drone images,

    Z. Zhang, “Drone-yolo: An efficient neural network method for target detection in drone images,”Drones, vol. 7, no. 8, p. 526, 2023

  65. [73]

    Repvgg: Mak- ing vgg-style convnets great again,

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, “Repvgg: Mak- ing vgg-style convnets great again,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 13 733–13 742

  66. [74]

    Dau-yolo: A lightweight and effective method for small object detection in uav images,

    W. Zeyu, L. Yizhou, X. Zhuodong, S. Ke, and Z. Feizhou, “Dau-yolo: A lightweight and effective method for small object detection in uav images,”Remote Sensing, vol. 17, no. 10, p. 1768, 2025

  67. [75]

    Visdrone- det2019: The vision meets drone object detection in image challenge results,

    D. Du, P. Zhu, L. Wen, X. Bian, H. Lin, Q. Huet al., “Visdrone- det2019: The vision meets drone object detection in image challenge results,” inProceedings of the IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). Seoul, South Korea: IEEE, 2019, pp. 213–226

  68. [76]

    Aerial multi-vehicle detection dataset,

    R. Makrigiorgis, P. Kolios, and C. Kyrkou, “Aerial multi-vehicle detection dataset,” 2022, version 1.0, Available at: https://doi.org/10. 5281/zenodo.7053442

  69. [77]

    On the new era of urban traffic monitoring with massive drone data: The pNEUMA large-scale field experiment,

    E. Barmpounakis and N. Geroliminis, “On the new era of urban traffic monitoring with massive drone data: The pNEUMA large-scale field experiment,”Transportation Research Part C: Emerging Technologies, vol. 111, pp. 50–71, 2020

  70. [78]

    A critical evaluation of the next generation simulation (ngsim) vehicle trajectory dataset,

    B. Coifman and L. Li, “A critical evaluation of the next generation simulation (ngsim) vehicle trajectory dataset,”Transportation Research Part B: Methodological, vol. 105, pp. 362–377, 2017

  71. [79]

    Learning social etiquette: Human trajectory understanding in crowded scenes,

    A. Robicquet, A. Sadeghian, A. Alahi, and S. Savarese, “Learning social etiquette: Human trajectory understanding in crowded scenes,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14. Springer, 2...

  72. [80]

    Vsai: A multi- view dataset for vehicle detection in complex scenarios using aerial images,

    J. Wang, X. Teng, Z. Li, Q. Yu, Y . Bian, and J. Wei, “Vsai: A multi- view dataset for vehicle detection in complex scenarios using aerial images,”Drones, vol. 6, no. 7, p. 161, 2022

  73. [81]

    The monet dataset: Multimodal drone thermal dataset recorded in rural scenarios,

    L. Riz, A. Caraffa, M. Bortolon, M. L. Mekhalfi, D. Boscaini, A. Moura, J. Antunes, A. Dias, H. Silva, A. Leonidouet al., “The monet dataset: Multimodal drone thermal dataset recorded in rural scenarios,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  74. [82]

    The unmanned aerial vehicle benchmark: Object detection and tracking,

    D. Du, Y . Qi, H. Yu, Y . Yang, K. Duan, G. Li, W. Zhang, Q. Huang, and Q. Tian, “The unmanned aerial vehicle benchmark: Object detection and tracking,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 370–386

  75. [83]

    Au-air: A multi-modal unmanned aerial vehicle dataset for low altitude traffic surveillance,

    I. Bozcan and E. Kayacan, “Au-air: A multi-modal unmanned aerial vehicle dataset for low altitude traffic surveillance,” in2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 8504–8510

  76. [84]

    The ad4che dataset and its application in typical congestion scenarios of traffic jam pilot systems,

    Y . Zhang, C. Wang, R. Yu, L. Wang, W. Quan, Y . Gao, and P. Li, “The ad4che dataset and its application in typical congestion scenarios of traffic jam pilot systems,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 5, pp. 3312–3323, 2023

  77. [85]

    Automatum data: Drone-based highway dataset for the development and validation of automated driving software for research and commercial applications,

    P. Spannaus, P. Zechel, and K. Lenz, “Automatum data: Drone-based highway dataset for the development and validation of automated driving software for research and commercial applications,” in2021 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2021, pp. 1372– 1377

  78. [86]

    Sind: A drone dataset at signalized intersection in china,

    Y . Xu, W. Shao, J. Li, K. Yang, W. Wang, H. Huang, C. Lv, and H. Wang, “Sind: A drone dataset at signalized intersection in china,” in 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2022, pp. 2471–2478

  79. [87]

    Interaction dataset: An international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps,

    W. Zhan, L. Sun, D. Wang, H. Shi, A. Clausse, M. Naumann, J. Kum- merle, H. Konigshof, C. Stiller, A. de La Fortelleet al., “Interaction dataset: An international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps,”arXiv preprint ar...

  80. [88]

    Spatiotemporal object detection for improved aerial vehicle detection in traffic monitoring,

    K. Telegraph and C. Kyrkou, “Spatiotemporal object detection for improved aerial vehicle detection in traffic monitoring,”IEEE Trans- actions on Artificial Intelligence, 2024

  81. [89]

    Dota: A large-scale dataset for object detection in aerial images,

    G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang, “Dota: A large-scale dataset for object detection in aerial images,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2018, pp. 3974–3983. 19

  82. [90]

    Ts4net: Two-stage sample selective strategy for rotating object detection,

    J. Zhou, K. Feng, W. Li, J. Han, and F. Pan, “Ts4net: Two-stage sample selective strategy for rotating object detection,”Neurocomputing, vol. 501, pp. 753–764, 2022

  83. [91]

    Anchor-free oriented proposal generator for object detection,

    G. Cheng, J. Wang, K. Li, X. Xie, C. Lang, Y . Yao, and J. Han, “Anchor-free oriented proposal generator for object detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–11, 2022

  84. [92]

    Towards large-scale small object detection: Survey and benchmarks,

    G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, X. Xie, and J. Han, “Towards large-scale small object detection: Survey and benchmarks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 11, pp. 13 467–13 488, 2023

  85. [93]

    Mor-uav: A benchmark dataset and baselines for moving object recognition in uav videos,

    M. Mandal, L. K. Kumar, and S. K. Vipparthi, “Mor-uav: A benchmark dataset and baselines for moving object recognition in uav videos,” in Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 2626–2635

  86. [94]

    Cascade r-cnn: Delving into high quality object detection,

    Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6154–6162

  87. [95]

    Dmnet: A network archi- tecture using dilated convolution and multiscale mechanisms for spa- tiotemporal fusion of remote sensing images,

    W. Li, X. Zhang, Y . Peng, and M. Dong, “Dmnet: A network archi- tecture using dilated convolution and multiscale mechanisms for spa- tiotemporal fusion of remote sensing images,”IEEE Sensors Journal, vol. 20, no. 20, pp. 12 190–12 202, 2020

  88. [96]

    Cdnet 2014: An expanded change detection benchmark dataset,

    Y . Wang, P.-M. Jodoin, F. Porikli, J. Konrad, Y . Benezeth, and P. Ishwar, “Cdnet 2014: An expanded change detection benchmark dataset,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2014, pp. 387–394

  89. [97]

    Msa-yolo: a remote sensing object detection model based on multi-scale strip attention,

    Z. Su, J. Yu, H. Tan, X. Wan, and K. Qi, “Msa-yolo: a remote sensing object detection model based on multi-scale strip attention,”Sensors, vol. 23, no. 15, p. 6811, 2023

  90. [98]

    Cf-yolo for small target detection in drone imagery based on yolov11 algorithm,

    C. Wang, Y . Han, C. Yang, M. Wu, Z. Chen, L. Yun, and X. Jin, “Cf-yolo for small target detection in drone imagery based on yolov11 algorithm,”Scientific Reports, vol. 15, no. 1, p. 16741, 2025

  91. [99]

    Sl-yolo: A stronger and lighter drone target detection model,

    D. Chen and L. Zhang, “Sl-yolo: A stronger and lighter drone target detection model,”arXiv preprint arXiv:2411.11477, 2024

  92. [100]

    Enhancing uav aerial image analysis: Integrating advanced sahi techniques with real-time detection models on the visdrone dataset,

    M. Muzammul, A. Algarni, Y . Y . Ghadi, and M. Assam, “Enhancing uav aerial image analysis: Integrating advanced sahi techniques with real-time detection models on the visdrone dataset,”IEEE Access, vol. 12, pp. 21 621–21 633, 2024

  93. [101]

    Ensembling object detection models for robust and reliable malaria parasite detection in thin blood smear microscopic images,

    E. ¨Ozbilge, E. G ¨uler, and E. Ozbilge, “Ensembling object detection models for robust and reliable malaria parasite detection in thin blood smear microscopic images,”IEEE Access, vol. 12, pp. 60 747–60 764, 2024

  94. [102]

    Ufpmp-det: Toward accurate and efficient object detection on drone imagery,

    Y . Huang, J. Chen, and D. Huang, “Ufpmp-det: Toward accurate and efficient object detection on drone imagery,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 1, 2022, pp. 1026– 1033

  95. [103]

    Orientation-and scale- invariant multi-vehicle detection and tracking from unmanned aerial videos,

    J. Wang, S. Simeonova, and M. Shahbazi, “Orientation-and scale- invariant multi-vehicle detection and tracking from unmanned aerial videos,”Remote Sensing, vol. 11, no. 18, p. 2155, 2019

  96. [104]

    Car detection from low-altitude uav imagery with the faster r-cnn,

    Y . Xu, G. Yu, Y . Wang, X. Wu, and Y . Ma, “Car detection from low-altitude uav imagery with the faster r-cnn,”Journal of Advanced Transportation, vol. 2017, no. 1, p. 2823617, 2017

  97. [105]

    A UA V-UGV cooperative system: Patrolling and energy management for urban monitoring,

    O. S. Oubbati, J. Alotaibi, F. Alromithy, M. Atiquzzaman, and M. R. Altimania, “A UA V-UGV cooperative system: Patrolling and energy management for urban monitoring,”IEEE Transactions on Vehicular Technology, 2025

  98. [106]

    Fast r-cnn,

    R. Girshick, “Fast r-cnn,”Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 1440–1448, 2015

  99. [107]

    Shufflenet: An extremely effi- cient convolutional neural network for mobile devices,

    X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely effi- cient convolutional neural network for mobile devices,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6848–6856

  100. [108]

    Mo- bilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mo- bilenetv2: Inverted residuals and linear bottlenecks,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4510–4520

  101. [109]

    Histograms of oriented gradients for human detection,

    N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), vol. 1. IEEE, 2005, pp. 886–893

  102. [110]

    Rich feature hierarchies for accurate object detection and semantic segmentation,

    R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 580–587

  103. [111]

    Spatial pyramid pooling in deep convolutional networks for visual recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 9, pp. 1904– 1916, 2015

  104. [112]

    R-fcn: Object detection via region- based fully convolutional networks,

    J. Dai, Y . Li, K. He, and J. Sun, “R-fcn: Object detection via region- based fully convolutional networks,”Advances in Neural Information Processing Systems, vol. 29, 2016

  105. [113]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 2961–2969

  106. [114]

    Selective search for object recognition,

    J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,”International Journal of Computer Vision, vol. 104, pp. 154–171, 2013

  107. [115]

    Generative adversarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 27, 2014, pp. 2672–2680

  108. [116]

    Unsupervised learning of video representations using lstms,

    N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” inInternational Con- ference on Machine Learning. PMLR, 2015, pp. 843–852

  109. [117]

    Deep high-resolution representation learning for visual recognition,

    J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wanget al., “Deep high-resolution representation learning for visual recognition,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 10, pp. 3349–3364, 2020

  110. [118]

    Scale-aware trident networks for object detection,

    Y . Li, Y . Chen, N. Wang, and Z. Zhang, “Scale-aware trident networks for object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6054–6063

  111. [119]

    Cornernet: Detecting objects as paired key- points,

    H. Law and J. Deng, “Cornernet: Detecting objects as paired key- points,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 734–750

  112. [120]

    Deformable model-based vehicle tracking and recognition using 3-d constrained multiple-kernels and kalman filter,

    T. Liu and Y . Liu, “Deformable model-based vehicle tracking and recognition using 3-d constrained multiple-kernels and kalman filter,” IEEE Access, vol. 9, pp. 90 346–90 357, 2021

  113. [121]

    Development of uav-based target tracking and recognition systems,

    S. Wang, F. Jiang, B. Zhang, R. Ma, and Q. Hao, “Development of uav-based target tracking and recognition systems,”IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 8, pp. 3409–3422, 2019

  114. [122]

    Optimizing disaster response with UA V-mounted RIS and HAP-enabled edge computing in 6G networks,

    J. Alotaibi, O. S. Oubbati, M. Atiquzzaman, F. Alromithy, and M. R. Altimania, “Optimizing disaster response with UA V-mounted RIS and HAP-enabled edge computing in 6G networks,”Journal of Network and Computer Applications, p. 104213, 2025

  115. [123]

    Aware channel-wise attentive network for vehicle re-identification,

    T.-S. Chen, M.-Y . Lee, C.-T. Liu, and S.-Y . Chien, “Aware channel-wise attentive network for vehicle re-identification,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 574–575

  116. [124]

    Deep multi-view spatial-temporal network for taxi demand prediction,

    H. Yao, F. Wu, J. Ke, X. Tang, Y . Jia, S. Lu, P. Gong, J. Ye, and Z. Li, “Deep multi-view spatial-temporal network for taxi demand prediction,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018

  117. [125]

    Uav swarm intelligence: Recent advances and future trends,

    Y . Zhou, B. Rao, and W. Wang, “Uav swarm intelligence: Recent advances and future trends,”IEEE Access, vol. 8, pp. 183 856–183 878, 2020

  118. [126]

    Multi-sensor optimal data fusion based on the adaptive fading unscented kalman filter,

    B. Gao, G. Hu, S. Gao, Y . Zhong, and C. Gu, “Multi-sensor optimal data fusion based on the adaptive fading unscented kalman filter,” Sensors, vol. 18, no. 2, p. 488, 2018

  119. [127]

    Robust stereo visual inertial odometry for fast autonomous flight,

    K. Sun, K. Mohta, B. Pfrommer, M. Watterson, S. Liu, Y . Mulgaonkar, C. J. Taylor, and V . Kumar, “Robust stereo visual inertial odometry for fast autonomous flight,”IEEE Robotics and Automation Letters, vol. 3, no. 2, pp. 965–972, 2018

  120. [128]

    Flight time minimization of uav for data collection over wireless sensor networks,

    J. Gong, T.-H. Chang, C. Shen, and X. Chen, “Flight time minimization of uav for data collection over wireless sensor networks,”IEEE Journal on Selected Areas in Communications, vol. 36, no. 9, pp. 1942–1954, 2018

  121. [129]

    Advanced computer vision for extracting georeferenced vehicle trajectories from drone imagery,

    R. Fonod, H. Cho, H. Yeo, and N. Geroliminis, “Advanced computer vision for extracting georeferenced vehicle trajectories from drone imagery,”Transportation Research Part C: Emerging Technologies, vol. 178, p. 105205, 2025

  122. [130]

    Multisource fusion uav cluster cooperative positioning using information geometry,

    C. Tang, Y . Wang, L. Zhang, Y . Zhang, and H. Song, “Multisource fusion uav cluster cooperative positioning using information geometry,” Remote Sensing, vol. 14, no. 21, p. 5491, 2022

  123. [131]

    A lightweight multidimensional feature network for small object detection on uavs,

    W. Yang, Q. He, and Z. Li, “A lightweight multidimensional feature network for small object detection on uavs,”Pattern Analysis and Applications, vol. 28, no. 1, pp. 1–24, 2025

  124. [132]

    Uav-yolov5: A swin-transformer- enabled small object detection model for long-range uav images,

    J. Li, C. Xie, S. Wu, and Y . Ren, “Uav-yolov5: A swin-transformer- enabled small object detection model for long-range uav images,” Annals of Data Science, pp. 1–30, 2024

  125. [133]

    Atbhc-yolo: aggregate trans- former and bidirectional hybrid convolution for small object detection,

    D. Liao, J. Zhang, Y . Tao, and X. Jin, “Atbhc-yolo: aggregate trans- former and bidirectional hybrid convolution for small object detection,” Complex & Intelligent Systems, vol. 11, no. 1, p. 38, 2025. 20

  126. [134]

    Improved multi-scale small target detection by uav,

    K. Sun, D. Li, and Y . Song, “Improved multi-scale small target detection by uav,”Multimedia Tools and Applications, pp. 1–15, 2024

  127. [135]

    Simultaneous localization and mapping (slam) and data fusion in unmanned aerial vehicles: Recent advances and challenges,

    A. Gupta and X. Fernando, “Simultaneous localization and mapping (slam) and data fusion in unmanned aerial vehicles: Recent advances and challenges,”Drones, vol. 6, no. 4, p. 85, 2022

  128. [136]

    Simsf: A scale insensitive multi-sensor fusion framework for unmanned aerial vehicles based on graph optimization,

    B. Dai, Y . He, L. Yang, Y . Su, Y . Yue, and W. Xu, “Simsf: A scale insensitive multi-sensor fusion framework for unmanned aerial vehicles based on graph optimization,”IEEE Access, vol. 8, pp. 118 273– 118 284, 2020

  129. [137]

    Robust ins/gps sensor fusion for uav local- ization using sdre nonlinear filtering,

    A. Nemra and N. Aouf, “Robust ins/gps sensor fusion for uav local- ization using sdre nonlinear filtering,”IEEE Sensors Journal, vol. 10, no. 4, pp. 789–798, 2010

  130. [138]

    Sensor fusion for attitude estimation and pid control of quadrotor uav,

    A. Noordin, M. A. M. Basri, and Z. Mohamed, “Sensor fusion for attitude estimation and pid control of quadrotor uav,”International Journal of Electrical and Electronic Engineering and Telecommuni- cations, vol. 7, no. 4, pp. 183–189, 2018

  131. [139]

    Stnn: A spatio-temporal neural network for traffic predictions,

    Z. He, C.-Y . Chow, and J.-D. Zhang, “Stnn: A spatio-temporal neural network for traffic predictions,”IEEE Transactions on Intelligent Transportation Systems, vol. 22, no. 12, pp. 7642–7651, 2020

  132. [140]

    Lightweight multi-frame integration for robust yolo object detection in videos,

    Y . Quan, B. Kiefer, M. Messmer, and A. Zell, “Lightweight multi-frame integration for robust yolo object detection in videos,”arXiv preprint arXiv:2506.20550, 2025

  133. [141]

    Learning spatiotemporal features with 3d convolutional networks,

    D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” inProceed- ings of the IEEE International Conference on Computer Vision, 2015, pp. 4489–4497

  134. [142]

    Teinet: Towards an efficient architecture for video recognition,

    Z. Liu, D. Luo, Y . Wang, L. Wang, Y . Tai, C. Wang, J. Li, F. Huang, and T. Lu, “Teinet: Towards an efficient architecture for video recognition,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 11 669–11 676

  135. [143]

    Yolo with adaptive frame control for real- time object detection applications,

    J. Lee and K.-i. Hwang, “Yolo with adaptive frame control for real- time object detection applications,”Multimedia Tools and Applications, vol. 81, no. 25, pp. 36 375–36 396, 2022

  136. [144]

    Energy-efficient data collection in uav enabled wireless sensor network,

    C. Zhan, Y . Zeng, and R. Zhang, “Energy-efficient data collection in uav enabled wireless sensor network,”IEEE Wireless Communications Letters, vol. 7, no. 3, pp. 328–331, 2017

  137. [145]

    Deep reinforcement learning for uav navigation through massive mimo technique,

    H. Huang, Y . Yang, H. Wang, Z. Ding, H. Sari, and F. Adachi, “Deep reinforcement learning for uav navigation through massive mimo technique,”IEEE Transactions on Vehicular Technology, vol. 69, no. 1, pp. 1117–1121, 2019

  138. [146]

    Video-based vehicle counting framework,

    Z. Dai, H. Song, X. Wang, Y . Fang, X. Yun, Z. Zhang, and H. Li, “Video-based vehicle counting framework,”IEEE Access, vol. 7, pp. 64 460–64 470, 2019

  139. [147]

    Real- time unmanned aerial vehicle-based traffic state estimation for multi- regional traffic networks,

    K. Theocharides, C. Menelaou, Y . Englezou, and S. Timotheou, “Real- time unmanned aerial vehicle-based traffic state estimation for multi- regional traffic networks,”Transportation Research Record, vol. 2678, no. 8, pp. 1–12, 2024

  140. [148]

    Adversarial examples: Attacks and defenses for deep learning,

    X. Yuan, P. He, Q. Zhu, and X. Li, “Adversarial examples: Attacks and defenses for deep learning,”IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 9, pp. 2805–2824, 2019

  141. [149]

    Explaining and harnessing adversarial examples,

    I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,”arXiv preprint arXiv:1412.6572, 2014

  142. [150]

    Practical black-box attacks against machine learning,

    N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami, “Practical black-box attacks against machine learning,” in Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, 2017, pp. 506–519

  143. [151]

    Robust physical-world attacks on deep learning visual classification,

    K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, and D. Song, “Robust physical-world attacks on deep learning visual classification,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 1625–1634

  144. [152]

    Model agnostic defense against adversarial patch attacks on object detection in unmanned aerial vehicles,

    S. Pathak, S. Shrestha, and A. AlMahmoud, “Model agnostic defense against adversarial patch attacks on object detection in unmanned aerial vehicles,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024, pp. 2586–2593

  145. [153]

    Robust adversarial attacks detection based on explainable deep reinforcement learning for uav guidance and planning,

    T. Hickling, N. Aouf, and P. Spencer, “Robust adversarial attacks detection based on explainable deep reinforcement learning for uav guidance and planning,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 10, pp. 4381–4394, 2023

  146. [154]

    On the robustness of semantic segmentation models to adversarial attacks,

    A. Arnab, O. Miksik, and P. H. Torr, “On the robustness of semantic segmentation models to adversarial attacks,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 888–897

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.