Pith. sign in

REVIEW 4 major objections 6 minor 37 references

MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read MID, a new shore-based dataset with 135,884 oriented-box ship instances, gives busy-port detection a realistic occlusion-heavy benchmark that satellite and SAR datasets lack.

desk verdict Potentially valuable shore-based OBB ship dataset, but the unexplained 15,050-to-5,673 frame selection gap must be addressed before the diversity claims are credible. read the letter →

arxiv 2412.05871 v3 pith:WONACGWJ submitted 2024-12-08 cs.CV

classification cs.CV
keywords orientedboundingboxshipdetectionopticalshore-baseddatasetdenseocclusionsmallobjectmaritimesituationalawarenessvideo-derivedYOLObaselines
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces MID, a shore-based optical dataset of 5,673 images annotated with oriented bounding boxes, containing 135,884 ship instances extracted from 43 real navigation videos. It argues that existing ship datasets—chiefly satellite or SAR-based, horizontal-box, or single-task datasets such as HRSID, SSDD, NWPU-10, HRSC2016, ShipRSImageNet, and SeaShip—do not cover dense occlusion, small-target clustering, and complex ship interactions in busy ports. MID is designed to fill that gap by encoding eight weather conditions, wide scale and aspect-ratio variation, multiple viewpoints, graded occlusion levels, and crossing, overtaking, and head-on encounters. The paper evaluates ten YOLO-family detectors with oriented-box heads as baselines, intended to serve as a reference for future work on maritime situational awareness, tracking, and trajectory prediction.

What carries the argument

The central object is the oriented bounding box (OBB): a rotated rectangle specified by four corner coordinates, chosen over axis-aligned boxes because ships appear at arbitrary headings and dense pixel overlap makes horizontal boxes inaccurate. Equally load-bearing is the dataset's annotation and organization scheme—each image is tied to a video ID and frame ID extracted at one frame every 176 frames (roughly every 6 seconds) from 1920×1080 shore-mounted cameras, giving a temporal ordering that detection alone would not provide. Around this scheme, the paper builds a difficulty taxonomy (weather, scale, aspect ratio, occlusion degree, background, and collision type) that lets the dataset be sliced into focused test conditions, and it runs ten YOLO-family detectors with OBB heads under fixed training settings to supply reference numbers.

What would settle it

Annotate all 15,050 extracted frames, or a random sample of the 9,377 frames that MID leaves out, and compare their instance counts, occlusion rates, weather conditions, and video coverage with MID's 5,673 images; a systematic mismatch would show that MID's diversity statistics describe the chosen subset, not the captured navigation scenes.

Watch

Extended reading notes

Core claim

MID is a video-derived optical dataset whose images carry a time dimension (video ID and frame ID) and point-based oriented bounding box annotations in the form of four corner coordinates. The dataset's defining claim is that dense occlusion and interaction-rich scenes, not just clean single-ship views, are the norm in real port monitoring: 16% of instances are at least slightly occluded, 837 instances are almost fully or fully occluded, and roughly half of all instances are tiny (at most 16 pixels) while the other half are extra-large (above 256 pixels). By including these cases alongside rain, fog, lens water droplets, overexposure, and multiple camera viewpoints, the authors aim to provide a harder and more realistic training and evaluation ground than existing datasets, one that supports both supervised and semi-supervised learning and downstream tasks such as tracking, trajectory extraction and prediction, and traffic information analysis. The baseline runs of ten YOLO variants with oriented-box heads are presented as first reference results on this benchmark.

Load-bearing premise

The load-bearing premise is that the 5,673 annotated images fairly represent the 15,050 frames extracted from the 43 videos, yet the paper gives no selection or filtering procedure between the two sets.

Editorial extensions

If this is right

  • Detectors trained on MID should transfer better to crowded port and narrow-channel monitoring than models trained on satellite, SAR, or single-target datasets, because the training distribution includes occluded, tiny, and overlapping ships.
  • The video ID and frame ID naming makes MID usable for tracking and trajectory extraction without extra alignment, directly supporting speed estimation and ship counting.
  • The graded occlusion annotations let researchers measure how detection performance degrades as occlusion increases, and provide a test set for occlusion-aware detectors.
  • The extreme scale distribution—about half tiny and half extra-large instances—stresses multiscale detectors and makes the dataset a demanding benchmark for small-target detection.
  • The fixed training settings and ten baseline configurations provide a reproducible comparison point for future oriented-box detectors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper does not say how the 5,673 annotated images were selected from the 15,050 extracted frames, the dataset's diversity statistics implicitly assume that this subset represents the full video corpus; releasing the selection rule or all frames would let users test that assumption.
  • The occlusion labels include fully occluded instances, which are not visible; this opens a route to evaluating track-based re-identification or weakly supervised detection that the paper does not develop.
  • All data come from one port area during 10 days in March, so generalizations to other seasons, regions, or port layouts are plausible but untested; a straightforward check is training on MID and testing on a second port's footage.
  • The time dimension plus OBB annotations could support a unified detection-and-tracking benchmark with occlusion-conditioned metrics, a construction the paper leaves for future versions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript introduces MID, a shore-based optical maritime dataset with 5,673 images and 135,884 oriented-bounding-box (OBB) annotations derived from 43 video segments recorded by port surveillance cameras. The paper documents dataset organization, analyzes diversity in terms of weather, scale, aspect ratio, background, occlusion, and collision scenarios, and reports detection baselines for ten YOLO-family variants with OBB heads. The central claim is that MID fills a gap in existing ship datasets by providing dense, occluded, multi-scale real-world maritime interactions and therefore supports detection, tracking, and trajectory-prediction research.

Significance. If the dataset is released as described, MID is a potentially valuable community resource: shore-based OBB video-derived data with temporal ordering is scarce, and the internal statistics are arithmetically consistent (the instance counts in Tables II and III sum to 135,884, and 135,884/5,673 = 23.95). The paper also provides reproducible-looking OBB conversion tooling and a public release plan. The significance is conditional, however, on resolving the undocumented gap between the 15,050 extracted frames and the 5,673 released images, and on clarifying the definitions behind the scale and occlusion statistics.

major comments (4)
  1. [III.A and IV.B] The paper states in Section III.A that 43 video segments, each sampled at one frame per 176 frames, yield 350 images per video and 15,050 original images, yet the released dataset and all subsequent statistics use 5,673 images. No filtering, exclusion, or subsampling procedure is described anywhere in the manuscript. This is load-bearing because the dataset's claims to reflect real-world shore-based maritime distributions rest on the representativeness of the final image set; the phrase in Section III.B.3, 'we collected as much occluded data as possible,' and the abstract's mention of 'manually supplemented annotations' suggest curation rather than random sampling. Please specify exactly how 5,673 images were obtained from 15,050 frames, report any exclusion criteria (e.g., empty frames, blur, annotation difficulty), and state whether selection was randomized and, if so, with what seed or protocol.
  2. [Table V] The reported recall of 0.917 for YOLOv10s-obb head is a striking outlier: every other model in the table has recall between 0.688 and 0.722, including YOLOv11s-obb, which has the same mAP50 of 84.9. No explanation or experimental note accompanies this value, and it is unlikely to be correct as reported. Because the baseline comparison is one of the paper's central evaluation claims, please verify the YOLOv10s result, report corrected numbers, and, ideally, include variance over multiple runs or seeds.
  3. [Table II and Section IV.B] The scale categories in Table II are labeled only as pixel thresholds (e.g., 'Tiny Instances ≤ 16 pixels'), without specifying whether the threshold refers to bounding-box area, long side, short side, or some other quantity. The resulting distribution—65,748 tiny instances and 66,984 extra-large instances, with almost no instances in the intermediate bins—is surprising for a dataset advertised as multi-scale and needs explanation. If the threshold is on area, 16 square pixels is far below a plausible annotatable ship size; if it is on side length, the units are unspecified. Please define the scale measure and discuss the apparent bimodality, which may also indicate that the 'tiny' and 'extra-large' bins are not measuring what the text implies.
  4. [III.B.3 and IV.F] The paper's central occlusion contribution lacks a formal definition. Section III.B.3 states that the authors 'annotate both the visible parts of the hull and the obscured sections at different visibility ratios,' but each object has a single OBB; it is unclear how a single box can encode both visible and occluded portions or how the occlusion percentage in Table III is computed. Section IV.F says 16% of the dataset contains occlusion, but Table III reports instance counts, not image counts. Please define the occlusion ratio, describe the annotation protocol for occluded targets (e.g., is the box drawn around the full extent or only the visible part?), and report occlusion statistics at both image and instance levels.
minor comments (6)
  1. [IV.A] The weather categories 'silty,' 'fuzzy,' and 'color-distorted' are non-standard and are not defined; please replace them with standard meteorological or visual categories or provide quantitative criteria.
  2. [IV.F] The sentence '16% of the dataset contains varying levels of occlusion' should say '16% of instances' if it refers to Table III, or should be recomputed at image level.
  3. [Table VI] The column 'Time Dimension Year' is confusing: the checkmark for Ours appears to refer to the video/frame ID naming convention rather than to temporal annotations; please clarify what is being compared.
  4. [V] The baseline experiments evaluate models trained and tested on MID only; a cross-dataset evaluation (e.g., fine-tune on MID and test on HRSID, HRSC2016, or SeaShip) would substantiate the claim that MID improves generalization to real-world complex scenes.
  5. [III.B] Annotation quality is stated to be ensured by four experienced annotators over three months, but no inter-annotator agreement, quality-control, or re-check procedure is described; a brief protocol statement would be valuable for dataset reliability.
  6. [References] Reference [30] for YOLO is incomplete (missing co-authors and publication details), and there are occasional formatting inconsistencies in the reference list (e.g., incomplete venue names).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper constructs a dataset and reports benchmark evaluations; no result is derived from or equivalent to its inputs.

full rationale

This is a dataset-and-benchmark paper, not a derivation. The central contribution is the MID image set with OBB annotations; its construction is described in Section III (frame extraction at one frame every 176 frames, annotation by professionals) and its properties are then measured in Section IV (weather, scale, aspect ratio, background, occlusion, collision statistics). These statistics are descriptive of the released images, not predictions obtained from a fitted model, so the paper does not claim to predict any quantity from first principles. The utility claim is supported by evaluation of ten YOLO-family detectors on a train/validation/test split of the same dataset (Section V), which is standard benchmark practice rather than circularity: the baseline results are measured performance, not quantities constructed from the models' own assumptions. There are no fitted parameters masquerading as predictions, and no load-bearing self-citations; the few citations to prior datasets and detectors are external. The unexplained reduction from 15,050 extracted frames to 5,673 released images (Section III.A vs. the abstract) is a genuine reporting gap about sample selection and representativeness, but it is a completeness and validity concern, not a circularity: nothing in the paper's argument reduces by definition to its own inputs. Therefore the circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The dataset claim rests on hand-chosen category thresholds (scale, occlusion), a sampling design that is not fully described, and gap assumptions about prior datasets; no fitted mathematical parameters are used in a derivation.

free parameters (3)
  • Occlusion degree thresholds = 10%, 20%, 50%, 90%
    Hand-chosen cutoffs in Table III define unobstructed, slight, partial, severe, and complete occlusion; these thresholds determine the 16% occlusion statistic used to support the dataset's occlusion-challenge claim.
  • Scale category thresholds = 16, 32, 96, 256 pixels
    Hand-chosen pixel cutoffs in Table II define tiny, small, medium, large, and extra-large categories; they affect the reported scale distribution and the claim that MID emphasizes small targets.
  • Frame extraction interval = 1 frame per 176 frames (approx. 6 seconds)
    Design choice in Section III-A determines the maximum possible frame count of 15,050, from which the final 5,673 images were selected without a described filter.
assumptions (3)
  • domain assumption Oriented bounding boxes are more accurate than horizontal bounding boxes for ship detection in complex scenes.
    Invoked in Section III-A (Figs. 3-4) to justify OBB annotation; not proven within the paper, but standard in rotated object detection literature.
  • domain assumption Existing ship detection datasets (HRSID, SSDD, NWPU-10) do not adequately cover dense occlusion and interaction scenarios.
    Stated in the Introduction and Section II to motivate MID; no quantitative comparison of occlusion or density statistics with prior datasets is provided.
  • domain assumption The camera installations and selected water areas (41 square km, 43 video segments) are representative of busy port and narrow-channel navigation.
    Assumed in Section III; no external validation is given that these scenes represent the general distribution of real-world maritime traffic.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios." pith.science (2026). https://pith.science/paper/WONACGWJ

@misc{pith2026241205871,
  author       = {Pith},
  title        = {Pith review of: MID: A Comprehensive Shore-Based Dataset for Multi-Scale Dense Ship Occlusion and Interaction Scenarios},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WONACGWJ}},
  note         = {Machine review of arXiv:2412.05871}
}
read the original abstract

This paper introduces the Maritime Ship Navigation Behavior Dataset (MID), designed to address challenges in ship detection within complex maritime environments using Oriented Bounding Boxes (OBB). MID contains 5,673 images with 135,884 finely annotated target instances, supporting both supervised and semi-supervised learning. It features diverse maritime scenarios such as ship encounters under varying weather, docking maneuvers, small target clustering, and partial occlusions, filling critical gaps in datasets like HRSID, SSDD, and NWPU-10. MID's images are sourced from high-definition video clips of real-world navigation across 43 water areas, with varied weather and lighting conditions (e.g., rain, fog). Manually curated annotations enhance the dataset's variety, ensuring its applicability to real-world demands in busy ports and dense maritime regions. This diversity equips models trained on MID to better handle complex, dynamic environments, supporting advancements in maritime situational awareness. To validate MID's utility, we evaluated 10 detection algorithms, providing an in-depth analysis of the dataset, detection results from various models, and a comparative study of baseline algorithms, with a focus on handling occlusions and dense target clusters. The results highlight MID's potential to drive innovation in intelligent maritime traffic monitoring and autonomous navigation systems. The dataset will be made publicly available at https://github.com/VirtualNew/MID_DataSet.

Figures

Figures reproduced from arXiv: 2412.05871 by the authors.

Figure 1
Figure 1. Video capture device. autonomous unmanned ships and maritime regulatory systems, where data from these complex scenes can effectively enhance the robustness and generalization ability of algorithms in real￾world applications. The core of our work is to establish a large dataset similar to existing visual tasks, but unlike current datasets, our target domain focuses on the perception of maritime ship navigation behav… view at source ↗
Figure 3
Figure 3. HBB annotation method. frames are extracted. The panoramic fisheye camera is used to capture a wide range of video data. The visible light zoom camera and the pan-tilt-zoom (PTZ) camera not only capture high-definition video images in a single direction but can also rotate at any angle to obtain video from different perspectives, while allowing for zoom adjustments to adapt to various scales. We divide each recorded… view at source ↗
Figure 4
Figure 4. OBB annotation method [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: Overall file structure of MID. professionals with experience in ship navigation management and labeling, ensuring the quality and consistency of the annotations. Using Horizontal Bounding Boxes (HBB) can lead to inaccuracies in enclosing complex-shaped objects, resulti…
Figure 6
Figure 6. Figure 6: Diversity of MID [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Examples images of different weather variations in [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 10
Figure 10. Figure 10: Proportion chart of different scale instances in MID. [PITH_FULL_IMAGE:figures/full_fig_p006_10.png]
Figure 13
Figure 13. Figure 13: AR distribution in MID [PITH_FULL_IMAGE:figures/full_fig_p007_13.png]
Figure 12
Figure 12. Figure 12: Distribution of ship instances of each image in MID. [PITH_FULL_IMAGE:figures/full_fig_p007_12.png]
Figure 15
Figure 15. Figure 15: Five main different backgrounds in MID [PITH_FULL_IMAGE:figures/full_fig_p008_15.png]
Figure 16
Figure 16. Figure 16: Sixteen different backgrounds in MID based on the [PITH_FULL_IMAGE:figures/full_fig_p008_16.png]
Figure 17
Figure 17. Figure 17: Examples images of different occlusion degrees. The [PITH_FULL_IMAGE:figures/full_fig_p008_17.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 32 canonical work pages

  1. [1]

    An outlook on the future marine traffic management system for autonomous ships,

    M. Martelli, A. Virdis, A. Gotta, P. Cassar `a, and M. Di Summa, “An outlook on the future marine traffic management system for autonomous ships,” IEEE Access , vol. 9, pp. 157 316–157 328, 2021

  2. [2]

    Developments in global seatrade and container shipping markets: their effects on the port industry and private sector involve- ment,

    H. J. Peters, “Developments in global seatrade and container shipping markets: their effects on the port industry and private sector involve- ment,” Int. J. Marit. Econ. , vol. 3, no. 1, pp. 3–26, 2001

  3. [3]

    Internet of things for smart ports: Technologies and challenges,

    Y . Yang, M. Zhong, H. Yao, F. Yu, X. Fu, and O. Postolache, “Internet of things for smart ports: Technologies and challenges,” IEEE Instrum. Meas. Mag. , vol. 21, no. 1, pp. 34–43, 2018

  4. [4]

    Multi-stage and multi- topology analysis of ship traffic complexity for probabilistic collision detection,

    X. Xin, Z. Yang, K. Liu, J. Zhang, and X. Wu, “Multi-stage and multi- topology analysis of ship traffic complexity for probabilistic collision detection,” Expert Syst. Appl. , vol. 213, p. 118890, 2023

  5. [5]

    A sidelobe-aware small ship detection network for synthetic aperture radar imagery,

    Y . Zhou, H. Liu, F. Ma, Z. Pan, and F. Zhang, “A sidelobe-aware small ship detection network for synthetic aperture radar imagery,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–16, 2023

  6. [6]

    Sensors and ai techniques for situational awareness in au- tonomous ships: A review,

    S. Thombre, Z. Zhao, H. Ramm-Schmidt, J. M. V . Garc´ıa, T. Malkam¨aki, S. Nikolskiy, T. Hammarberg, H. Nuortie, M. Z. H. Bhuiyan, S. S ¨arkk¨a et al. , “Sensors and ai techniques for situational awareness in au- tonomous ships: A review,” IEEE trans. Intell. Transp. Syst. , vol. 23, no. 1, pp. 64–83, 2020

  7. [7]

    The ocean-going autonomous ship—challenges and threats,

    A. Felski and K. Zwolak, “The ocean-going autonomous ship—challenges and threats,” J. Mar . Sci. Eng. , vol. 8, no. 1, p. 41, 2020

  8. [8]

    Weather-aware object detection method for maritime surveillance systems,

    M. Chen, J. Sun, K. Aida, and A. Takefusa, “Weather-aware object detection method for maritime surveillance systems,” Future Gener . Comp. Sy., vol. 151, pp. 111–123, 2024

Show all 37 references
  1. [9]

    Object detection in a maritime environment: Performance evaluation of background subtraction methods,

    D. K. Prasad, C. K. Prasath, D. Rajan, L. Rachmawati, E. Rajabally, and C. Quek, “Object detection in a maritime environment: Performance evaluation of background subtraction methods,” IEEE trans. Intell. Transp. Syst., vol. 20, no. 5, pp. 1787–1802, 2018

  2. [10]

    Ship detection in high-resolution optical imagery based on anomaly detector and local shape feature,

    Z. Shi, X. Yu, Z. Jiang, and B. Li, “Ship detection in high-resolution optical imagery based on anomaly detector and local shape feature,” IEEE Trans. Geosci. Remote Sens. , vol. 52, no. 8, pp. 4511–4523, 2013

  3. [11]

    How big data enriches maritime research–a critical review of automatic identification system (ais) data applications,

    D. Yang, L. Wu, S. Wang, H. Jia, and K. X. Li, “How big data enriches maritime research–a critical review of automatic identification system (ais) data applications,” Transp. Rev., vol. 39, no. 6, pp. 755–773, 2019

  4. [12]

    Ship detection with high-resolution hf skywave radar,

    J. Barnum, “Ship detection with high-resolution hf skywave radar,” IEEE J. OCEANIC ENG. , vol. 11, no. 2, pp. 196–209, 1986

  5. [13]

    Conservation science and policy applications of the marine vessel automatic identification system (ais)—a review,

    M. Robards, G. Silber, J. Adams, J. Arroyo, D. Lorenzini, K. Schwehr, and J. Amos, “Conservation science and policy applications of the marine vessel automatic identification system (ais)—a review,” B. MAR. SCI., vol. 92, no. 1, pp. 75–103, 2016

  6. [14]

    Ship de- tection with spectral analysis of synthetic aperture radar: A comparison of new and well-known algorithms,

    A. Marino, M. J. Sanjuan-Ferrer, I. Hajnsek, and K. Ouchi, “Ship de- tection with spectral analysis of synthetic aperture radar: A comparison of new and well-known algorithms,” Remote Sens. , vol. 7, no. 5, pp. 5416–5439, 2015

  7. [15]

    Ship detection for visual maritime surveillance from non-stationary platforms,

    Y . Zhang, Q.-Z. Li, and F.-N. Zang, “Ship detection for visual maritime surveillance from non-stationary platforms,” OCEAN ENG. , vol. 141, pp. 53–63, 2017

  8. [16]

    Research of target detection and classification techniques using millimeter-wave radar and vision sensors,

    Z. Wang, X. Miao, Z. Huang, and H. Luo, “Research of target detection and classification techniques using millimeter-wave radar and vision sensors,” Remote Sens. , vol. 13, no. 6, p. 1064, 2021

  9. [17]

    Dataset and benchmark for ship detection in complex optical remote sensing image,

    J. Hu, X. Zhi, T. Shi, J. Wang, Y . Li, and X. Sun, “Dataset and benchmark for ship detection in complex optical remote sensing image,” IEEE Trans. Geosci. Remote Sens. , 2024

  10. [18]

    Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,

    S. Wei, X. Zeng, Q. Qu, M. Wang, H. Su, and J. Shi, “Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,” IEEE Access , vol. 8, pp. 120 234–120 254, 2020

  11. [19]

    Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,

    T. Zhang, X. Zhang, J. Li, X. Xu, B. Wang, X. Zhan, Y . Xu, X. Ke, T. Zeng, H. Su et al., “Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,” Remote Sens., vol. 13, no. 18, p. 3690, 2021

  12. [20]

    Object detection and instance segmentation in remote sensing imagery based on precise mask r-cnn,

    H. Su, S. Wei, M. Yan, C. Wang, J. Shi, and X. Zhang, “Object detection and instance segmentation in remote sensing imagery based on precise mask r-cnn,” in Proc. IEEE Int. Geosci. Remote Sens. Symp. IEEE, 2019, pp. 1454–1457

  13. [21]

    Hq- isnet: High-quality instance segmentation for remote sensing imagery,

    H. Su, S. Wei, S. Liu, J. Liang, C. Wang, J. Shi, and X. Zhang, “Hq- isnet: High-quality instance segmentation for remote sensing imagery,” Remote Sens. , vol. 12, no. 6, p. 989, 2020

  14. [22]

    A high resolution optical satellite image dataset for ship recognition and some new baselines,

    Z. Liu, L. Yuan, L. Weng, and Y . Yang, “A high resolution optical satellite image dataset for ship recognition and some new baselines,” in Proc. Int. Conf. Pattern Recognit. Appl. Methods , vol. 2. SciTePress, 2017, pp. 324–331

  15. [23]

    Shiprsimagenet: A large-scale fine-grained dataset for ship detection in high-resolution optical remote sensing images,

    Z. Zhang, L. Zhang, Y . Wang, P. Feng, and R. He, “Shiprsimagenet: A large-scale fine-grained dataset for ship detection in high-resolution optical remote sensing images,” EEE J. Sel. Top Appl. Earth Obs. Remote Sens., vol. 14, pp. 8458–8472, 2021

  16. [24]

    Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,

    Y . Han, X. Yang, T. Pu, and Z. Peng, “Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,” IEEE Trans. Geosci. Remote Sens. , vol. 60, pp. 1–18, 2021

  17. [25]

    Seaships: A large-scale precisely annotated dataset for ship detection,

    Z. Shao, W. Wu, Z. Wang, W. Du, and C. Li, “Seaships: A large-scale precisely annotated dataset for ship detection,” IEEE Trans. Multimedia., vol. 20, no. 10, pp. 2593–2604, 2018

  18. [26]

    A discriminatively trained, multiscale, deformable part model,

    P. Felzenszwalb, D. McAllester, and D. Ramanan, “A discriminatively trained, multiscale, deformable part model,” in Proc. IEEE Conf. Com- put. Vision Pattern Recognit. Ieee, 2008, pp. 1–8

  19. [27]

    Rich feature hierarchies for accurate object detection and semantic segmentation,

    R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2014, pp. 580–587

  20. [28]

    Fast r-cnn,

    R. Girshick, “Fast r-cnn,” arXiv preprint arXiv:1504.08083 , 2015

  21. [29]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 6, pp. 1137–1149, 2016

  22. [30]

    You only look once: Unified, real-time object detection,

    J. Redmon, “You only look once: Unified, real-time object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2016

  23. [31]

    Ssd: Single shot multibox detector,

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y . Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in Proc. European Conf. Comput. Vision . Springer, 2016, pp. 21–37

  24. [32]

    Focal loss for dense object detection,

    T.-Y . Ross and G. Doll ´ar, “Focal loss for dense object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2017, pp. 2980– 2988

  25. [33]

    End-to-end object detection with transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. European Conf. Comput. Vision . Springer, 2020, pp. 213–229

  26. [34]

    Efficientdet: Scalable and efficient object detection,

    M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” in Proc. IEEE Conf. Comput. Vision Pattern Recognit. , 2020, pp. 10 781–10 790

  27. [35]

    Yolov6: A single-stage object detection framework for industrial applications,

    C. Li, L. Li, H. Jiang, K. Weng, Y . Geng, L. Li, Z. Ke, Q. Li, M. Cheng, W. Nie et al. , “Yolov6: A single-stage object detection framework for industrial applications,” arXiv preprint arXiv:2209.02976 , 2022

  28. [36]

    Yolov9: Learning what you want to learn using programmable gradient information,

    C.-Y . Wang, I.-H. Yeh, and H.-Y . Mark Liao, “Yolov9: Learning what you want to learn using programmable gradient information,” in Proc. European Conf. Comput. Vision . Springer, 2025, pp. 1–21

  29. [37]

    Yolov10: Real-time end-to-end object detection,

    A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding, “Yolov10: Real-time end-to-end object detection,” arXiv preprint arXiv:2405.14458, 2024

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.