Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

Recent Advances in Deep Learning for Object Detection

T0 review · 3 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This survey claims that deep-learning object detection is best organized as a two-stage/one-stage dichotomy, with components, learning strategies, and benchmarks making up the rest of the map.

desk verdict A useful but imperfect map of deep learning object detection as of 2019; the narrative is solid, but the benchmark tables need fixing before you trust them. read the letter →

arxiv 1908.03673 v1 pith:WIBCY2D7 submitted 2019-08-10 cs.CV cs.LGcs.MM

classification cs.CVcs.LGcs.MM
keywords objectdetectiondeeplearningconvolutionalneuralnetworkstwo-stagedetectorsone-stagefeaturepyramidanchor-freeinstancesegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that recent deep-learning object detection is best understood through a single split: two-stage detectors that first propose regions and then classify them, versus one-stage detectors that predict classes and boxes directly from feature maps. It claims that performance is determined by a small set of reusable components, namely backbone architecture, proposal generation, and feature representation, together with learning strategies such as data augmentation, imbalance sampling, localization refinement, and cascade learning. If the map is right, it gives researchers and practitioners a reliable way to locate any detector, compare the accuracy-versus-speed tradeoff, and choose where to improve next. The survey also assembles benchmark tables and reviews face and pedestrian detection as the main applied settings.

What carries the argument

The machinery that carries the survey is its taxonomy, built on the two-stage/one-stage dichotomy as the root. The taxonomy organizes every reviewed method into detection components, learning strategies, and applications and benchmarks, so that a detector is described by choices within each category rather than by a single headline number. Within that structure, the review treats feature pyramids and multi-scale feature learning as the key mechanism for handling scale variation, and class-imbalance handling through hard negative mining and focal loss as the key mechanism for training one-stage detectors. The IoU-based evaluation metrics and the benchmark tables are the instruments that make the taxonomy's comparisons concrete.

What would settle it

Retrain two detectors from the same benchmark table under an identical protocol, using the same training data, backbone, augmentation, and test settings, and compare mAP; if the ranking inverts or the gap collapses, the survey's cross-method comparisons are not reliable.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that every important deep object detector falls into one of two paradigms: two-stage detectors, which use a proposal generator to produce a sparse set of candidate regions and then classify and refine each one, and one-stage detectors, which skip proposal generation and make dense predictions at every location. The survey further claims that the factors controlling detection quality decompose into detection components, including backbone networks, proposal generation, multi-scale and deformable feature learning, and region encoding, plus learning strategies such as data augmentation, imbalance sampling, localization refinement, cascade learning, and test-time processing such as non-maximum suppression. It presents benchmark tables on Pascal VOC and MS COCO that track the accuracy and speed of representative detectors, and it identifies anchor-free, keypoint-based detectors and AutoML-designed architectures as the most active directions going forward.

Load-bearing premise

The survey assumes that the benchmark scores reported by different papers can be compared directly in Tables 2 and 3, even though those papers differ in training data, test-time augmentation, backbone, and implementation details.

Editorial extensions

If this is right

  • If the taxonomy is right, a new detector can be located and understood by its place in the two-stage/one-stage split and by its component choices, giving the field a stable reference map as of mid-2019.
  • The accuracy-versus-speed tradeoff between the two families becomes a design expectation rather than an accident: two-stage detectors set accuracy records while one-stage detectors set real-time records.
  • Class imbalance and scale variation emerge as first-order training problems, so progress in sampling strategies, loss design, and feature pyramids should keep improving accuracy across both families.
  • The benchmark tables imply that anchor-free keypoint detectors and AutoML-based architectures are the directions most likely to push state-of-the-art results next.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The two-family split will probably blur as components migrate, since two-stage detectors adopt anchor-free heads and one-stage detectors adopt cascaded refinement, so a later survey may need a finer-grained axis.
  • Because the benchmark tables mix training data, backbones, and test-time augmentation, the reported rankings should be read as approximate; identical-protocol re-runs could reorder methods.
  • A testable extension of the survey's decomposition is to ablate components one at a time within a fixed backbone, and the largest mAP swings would identify which part of the taxonomy carries the field's progress.
  • The review's list of open problems suggests that low-shot detection and detection-specific backbones may matter as much as anchor-free design in the next stage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This manuscript is a survey of deep learning methods for visual object detection, organized into detection components, learning strategies, and applications and benchmarks. It reviews two-stage and one-stage detectors, backbone architectures, proposal generation, feature learning, training and testing strategies, and specialized tasks such as face and pedestrian detection, along with public benchmarks and future directions. The paper also provides two large comparison tables (Pascal VOC and MS COCO) intended to summarize the state of the art.

Significance. If the survey's content is accurate, it offers a useful structured overview of a fast-moving field, with broad coverage and a substantial reference list. The taxonomy (two-stage vs. one-stage, components, learning strategies, applications) is reasonable and the milestone timeline is helpful for orientation. However, the central value of such a survey depends on the reliability of its compiled benchmark results, and the concrete misattributions identified below materially reduce that reliability. The contribution is therefore of moderate significance and needs a careful revision of the benchmark tables and related citations before it can serve as a dependable map of the field.

major comments (3)
  1. [Table 2 (Section 7)] The R-CNN row reports 66.0 mAP on VOC2007 with an asterisk claiming that the model is trained only on VOC2007 trainval. This is inconsistent with the original R-CNN paper, which reports 58.5 mAP for a VGG-16 model trained on VOC2007 trainval; the 66.0 figure corresponds to a different configuration (with additional VOC2012 data and/or ensembled bounding-box regression). As written, the entry misleads readers about the training protocol and inflates the comparison for that row. Please re-verify the number and either correct the value or clarify the exact configuration in the footnote.
  2. [Table 3 (Section 7)] The DeepRegionlets row lists ResNet-101 as the backbone, but the original ECCV 2018 paper reports its COCO test-dev result (39.3 AP, 59.8 AP50) with a VGG-16 backbone and a multi-scale testing variant. The combination of a wrong backbone and the missing '++' marker means the row is not comparable to other rows in the same table that use standard inference. Please correct the backbone and indicate the test-time protocol, or remove the row if the original protocol cannot be cleanly accommodated.
  3. [Table 3 (Section 7)] The table mixes base models, entries marked '++' (multi-scale testing, horizontal flip, etc.), and entries such as DeepRegionlets that use multi-scale testing without the marker. The caption defines '++' but does not state that rows with '++' and rows without it are not directly comparable. Given that the survey's organizing claim is to provide a reliable comparative map of detector performance, the table should include per-row protocol indications (or a separate column) so that readers can make valid comparisons; otherwise the benchmark tables may mislead rather than inform.
minor comments (7)
  1. [Table 2] The SPP-net row cites reference [2], but the correct reference for SPP-net is [47].
  2. [Table 2 footnote] The footnote contains duplicated words ('the the model is trained') and a subject-verb agreement error ('the model are trained'); please correct both.
  3. [Section 3.2.2] The name 'Single-Shot Mulibox Detector' should be 'Single-Shot Multibox Detector'.
  4. [Section 3.5.2] The phrase 'Precise ROI Pooing' should be 'Precise ROI Pooling'.
  5. [Section 4.2.1] The name 'Hosong et al.' should be 'Hosang et al.' (the authors of 'Learning non-maximum suppression').
  6. [Section 8] The word 'sveral' should be 'several'.
  7. [Section 7] The opening sentence lists 'Pascal VOC2007, VOC2007 and MSCOCO'; the second occurrence should be 'VOC2012'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey organizes prior published results and makes no derivation or prediction claim that reduces to its inputs.

full rationale

This paper is a literature survey of deep-learning object detection. Its organizing claim is taxonomic (two-stage vs one-stage detectors, components, learning strategies, applications, benchmarks), and its factual content consists of descriptions of prior work plus benchmark tables reproducing numbers from the cited papers. There is no derivation chain in which an output is constructed from an input in a way that makes the output equivalent to the input by definition. The authors' self-citations (e.g., refs [13], [183], [188]) appear only as examples of logo detection, face detection, and related applications; they are not invoked as load-bearing evidence for any taxonomic or comparative claim, and the survey's structure does not depend on accepting those specific results. The benchmark tables could be criticized for inconsistent evaluation protocols across entries, but that is a correctness/comparability concern, not circularity: the table entries are reported external results, not predictions generated from a fitted parameter in this paper. No fitted input is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via self-citation. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

No free parameters or invented entities appear in this review. The only background assumption is the reliability of the cited literature and the representativeness of the selected papers.

assumptions (1)
  • domain assumption The cited papers are accurately summarized and their reported results are reliable as published.
    The survey does not re-run any experiments, so its entire content rests on trusting the primary sources it cites.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Recent Advances in Deep Learning for Object Detection." pith.science (2026). https://pith.science/paper/WIBCY2D7

@misc{pith2026190803673,
  author       = {Pith},
  title        = {Pith review of: Recent Advances in Deep Learning for Object Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WIBCY2D7}},
  note         = {Machine review of arXiv:1908.03673}
}
read the original abstract

Object detection is a fundamental visual recognition problem in computer vision and has been widely studied in the past decades. Visual object detection aims to find objects of certain target classes with precise localization in a given image and assign each object instance a corresponding class label. Due to the tremendous successes of deep learning based image classification, object detection techniques using deep learning have been actively studied in recent years. In this paper, we give a comprehensive survey of recent advances in visual object detection with deep learning. By reviewing a large body of recent related work in literature, we systematically analyze the existing object detection frameworks and organize the survey into three major parts: (i) detection components, (ii) learning strategies, and (iii) applications & benchmarks. In the survey, we cover a variety of factors affecting the detection performance in detail, such as detector architectures, feature learning, proposal generation, sampling strategies, etc. Finally, we discuss several future directions to facilitate and spur future research for visual object detection with deep learning. Keywords: Object Detection, Deep Learning, Deep Convolutional Neural Networks

Figures

Figures reproduced from arXiv: 1908.03673 by the authors.

Figure 1
Figure 1. Comparison of different visual recognition tasks in computer vision. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Major milestone in object detection research based on deep convolution neural networks since 2012. The trend in the last year has been designing object [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy of key methodologies in this survey. We categorize various contributions for deep learning based object detection into three major categories: [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Overview of different two-stage detection frameworks for generic object detection. Red dotted rectangles denote the outputs that define the loss functions. [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Overview of different one-stage detection frameworks for generic object detection. Red rectangles denotes the outputs that define the objective functions. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Diagram of RPN [34]. Each position of the feature map connects [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Four paradigms for multi-scale feature learning. Top Left: [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: General framework for feature combination. Top-down features are 2 [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Example of failure case of detection in high IoU threshold. Purple box [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: Duplicate predictions are eliminated by NMS operation. The most [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Some examples of Pascal VOC, MSCOCO, Open Images and LVIS. [PITH_FULL_IMAGE:figures/full_fig_p030_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of Simultaneous Localization and Mapping with an Envision in 6G Wireless Networks

    cs.RO 2019-08 conditional novelty 1.0 of 10

    A broad review of Lidar, visual, and fused SLAM systems, with an unquantified vision for SLAM using future 6G terahertz wireless networks.

Reference graph

Works this paper leans on

256 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [2]

    Girshick, J

    R. Girshick, J. Donahue, T. Darrell, J. Malik, Rich feature hierarchies for accurate object detection and semantic segmentation, in: CVPR, 2014

  2. [47]

    K. He, X. Zhang, S. Ren, J. Sun, Spatial pyramid pooling in deep con- 35 volutional networks for visual recognition, in: ECCV , 2014

  3. [1]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recog- nition, in: CVPR, 2016

  4. [3]

    K. He, G. Gkioxari, P. Doll ´ar, R. Girshick, Mask r-cnn, in: ICCV , 2017

  5. [4]

    L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A. L. Yuille, Se- mantic image segmentation with deep convolutional nets and fully con- nected crfs, in: arXiv preprint arXiv:1412.7062, 2014. 34

  6. [5]

    Y . Sun, D. Liang, X. Wang, X. Tang, Deepid3: Face recognition with very deep neural networks, in: arXiv preprint arXiv:1502.00873, 2015

  7. [6]

    Y . Sun, Y . Chen, X. Wang, X. Tang, Deep learning face representation by joint identification-verification, in: NeurIPS, 2014

  8. [7]

    W. Liu, Y . Wen, Z. Yu, M. Li, B. Raj, L. Song, Sphereface: Deep hyper- sphere embedding for face recognition, in: CVPR, 2017

Show all 256 references
  1. [8]

    J. Li, X. Liang, S. Shen, T. Xu, J. Feng, S. Yan, Scale-aware fast r-cnn for pedestrian detection, in: IEEE Transactions on Multimedia, 2018

  2. [9]

    Hosang, M

    J. Hosang, M. Omran, R. Benenson, B. Schiele, Taking a deeper look at pedestrians, in: CVPR, 2015

  3. [10]

    Angelova, A

    A. Angelova, A. Krizhevsky, V . Vanhoucke, A. S. Ogale, D. Ferguson, Real-time pedestrian detection with deep network cascades., in: BMVC, 2015

  4. [11]

    Karpathy, G

    A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, L. Fei-Fei, Large-scale video classification with convolutional neural networks, in: CVPR, 2014

  5. [12]

    Mobahi, R

    H. Mobahi, R. Collobert, J. Weston, Deep learning from temporal coher- ence in video, in: Annual International Conference on Machine Learn- ing, 2009

  6. [13]

    S. C. Hoi, X. Wu, H. Liu, Y . Wu, H. Wang, H. Xue, Q. Wu, Logo-net: Large-scale deep logo detection and brand recognition with deep region- based convolutional networks, in: arXiv preprint arXiv:1511.02462, 2015

  7. [14]

    H. Su, X. Zhu, S. Gong, Deep learning logo detection with data ex- pansion by synthesising context, in: 2017 IEEE Winter Conference on Applications of Computer Vision (W ACV), 2017

  8. [15]

    H. Su, S. Gong, X. Zhu, Scalable deep learning logo detection, in: arXiv preprint arXiv:1803.11417, 2018

  9. [16]

    Vedaldi, V

    A. Vedaldi, V . Gulshan, M. Varma, A. Zisserman, Multiple kernels for object detection, in: ICCV , 2009

  10. [17]

    Viola, M

    P. Viola, M. Jones, Rapid object detection using a boosted cascade of simple features, in: CVPR, 2001

  11. [18]

    Harzallah, F

    H. Harzallah, F. Jurie, C. Schmid, Combining efficient object localiza- tion and image classification, in: ICCV , 2009

  12. [19]

    Dalal, B

    N. Dalal, B. Triggs, Histograms of oriented gradients for human detec- tion, in: CVPR, 2005

  13. [20]

    Viola, M

    P. Viola, M. J. Jones, Robust real-time face detection, in: IJCV , 2004

  14. [21]

    D. G. Lowe, Object recognition from local scale-invariant features, in: ICCV , 1999

  15. [22]

    Lienhart, J

    R. Lienhart, J. Maydt, An extended set of haar-like features for rapid ob- ject detection, in: International Conference on Image Processing, 2002

  16. [23]

    H. Bay, T. Tuytelaars, L. Van Gool, Surf: Speeded up robust features, in: ECCV , 2006

  17. [24]

    M. A. Hearst, S. T. Dumais, E. Osuna, J. Platt, B. Scholkopf, Support vector machines, in: IEEE Intelligent Systems and their applications, 1998

  18. [25]

    Opitz, R

    D. Opitz, R. Maclin, Popular ensemble methods: An empirical study, in: Journal of artificial intelligence research, 1999

  19. [26]

    Freund, R

    Y . Freund, R. E. Schapire, et al., Experiments with a new boosting algo- rithm, in: ICML, 1996

  20. [27]

    Y . Yu, J. Zhang, Y . Huang, S. Zheng, W. Ren, C. Wang, K. Huang, T. Tan, Object detection by context and boosted hog-lbp, in: PASCAL VOC Challenge, 2010

  21. [28]

    Felzenszwalb, R

    P. Felzenszwalb, R. Girshick, D. McAllester, D. Ramanan, Discrimina- tively trained mixtures of deformable part models, in: PASCAL VOC Challenge, 2008

  22. [29]

    Everingham, L

    M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, A. Zisserman, The pascal visual object classes (voc) challenge, in: IJCV , 2010

  23. [30]

    P. F. Felzenszwalb, R. B. Girshick, D. McAllester, D. Ramanan, Object detection with discriminatively trained part-based models, in: TPAMI, 2010

  24. [31]

    D. G. Lowe, Distinctive image features from scale-invariant keypoints, in: IJCV , 2004

  25. [32]

    Ojala, M

    T. Ojala, M. Pietikainen, T. Maenpaa, Multiresolution gray-scale and rotation invariant texture classification with local binary patterns, in: TPAMI, 2002

  26. [33]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, in: NeurIPS, 2012

  27. [34]

    S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time ob- ject detection with region proposal networks, in: NeurIPS, 2015

  28. [35]

    Fukushima, S

    K. Fukushima, S. Miyake, Neocognitron: A self-organizing neural net- work model for a mechanism of visual pattern recognition, in: Compe- tition and cooperation in neural nets, 1982

  29. [36]

    LeCun, L

    Y . LeCun, L. Bottou, Y . Bengio, P. Haffner, Gradient-based learning applied to document recognition, in: Proceedings of the IEEE, 1998

  30. [37]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: CVPR, 2009

  31. [38]

    Girshick, Fast r-cnn, in: ICCV , 2015

    R. Girshick, Fast r-cnn, in: ICCV , 2015

  32. [39]

    T.-Y . Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, S. Belongie, Fea- ture pyramid networks for object detection, in: CVPR, 2017

  33. [40]

    Redmon, S

    J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: Unified, real-time object detection, in: CVPR, 2016

  34. [41]

    Redmon, A

    J. Redmon, A. Farhadi, Yolo9000: better, faster, stronger, in: CVPR, 2017

  35. [42]

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y . Fu, A. C. Berg, SSD: Single shot multibox detector, in: ECCV , 2016

  36. [43]

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, P. Doll ´ar, Focal loss for dense object detection, in: ICCV , 2017

  37. [44]

    Fidler, R

    S. Fidler, R. Mottaghi, A. Yuille, R. Urtasun, Bottom-up segmentation for top-down detection, in: CVPR, 2013

  38. [45]

    J. R. Uijlings, K. E. Van De Sande, T. Gevers, A. W. Smeulders, Selec- tive search for object recognition, in: IJCV , 2013

  39. [46]

    Kleban, X

    J. Kleban, X. Xie, W.-Y . Ma, Spatial pyramid mining for logo detection in natural scenes, in: Multimedia and Expo, 2008 IEEE International Conference on, 2008

  40. [48]

    C. L. Zitnick, P. Doll ´ar, Edge boxes: Locating object proposals from edges, in: ECCV , 2014

  41. [49]

    Z. Cai, N. Vasconcelos, Cascade r-cnn: Delving into high quality object detection, in: CVPR, 2018

  42. [50]

    T. Kong, A. Yao, Y . Chen, F. Sun, Hypernet: Towards accurate region proposal generation and joint object detection, in: CVPR, 2016

  43. [51]

    S. Bell, C. Lawrence Zitnick, K. Bala, R. Girshick, Inside-outside net: Detecting objects in context with skip pooling and recurrent neural net- works, in: CVPR, 2016

  44. [52]

    J. Dai, Y . Li, K. He, J. Sun, R-fcn: Object detection via region-based fully convolutional networks, in: NeurIPS, 2016

  45. [53]

    K. Kang, W. Ouyang, H. Li, X. Wang, Object detection from video tubelets with convolutional neural networks, in: CVPR, 2016

  46. [54]

    W. Han, P. Khorrami, T. L. Paine, P. Ramachandran, M. Babaeizadeh, H. Shi, J. Li, S. Yan, T. S. Huang, Seq-nms for video object detection, in: arXiv preprint arXiv:1602.08465, 2016

  47. [55]

    Rayat Imtiaz Hossain, J

    M. Rayat Imtiaz Hossain, J. Little, Exploiting temporal information for 3d human pose estimation, in: ECCV , 2018

  48. [56]

    Pavlakos, X

    G. Pavlakos, X. Zhou, K. G. Derpanis, K. Daniilidis, Coarse-to-fine vol- umetric prediction for single-image 3d human pose, in: CVPR, 2017

  49. [57]

    P. O. Pinheiro, T.-Y . Lin, R. Collobert, P. Doll´ar, Learning to refine ob- ject segments, in: ECCV , 2016

  50. [58]

    P. O. Pinheiro, R. Collobert, P. Doll ´ar, Learning to segment object can- didates, in: NeurIPS, 2015

  51. [59]

    J. Dai, K. He, J. Sun, Instance-aware semantic segmentation via multi- task network cascades, in: CVPR, 2016

  52. [60]

    Huang, L

    Z. Huang, L. Huang, Y . Gong, C. Huang, X. Wang, Mask scoring r-cnn, in: CVPR, 2019

  53. [61]

    Sermanet, D

    P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, Y . LeCun, Overfeat: Integrated recognition, localization and detection using con- volutional networks, in: arXiv preprint arXiv:1312.6229, 2013

  54. [62]

    Ioffe, C

    S. Ioffe, C. Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, in: ICML, 2015

  55. [63]

    H. Law, J. Deng, Cornernet: Detecting objects as paired keypoints, in: ECCV , 2018

  56. [64]

    X. Zhou, D. Wang, P. Kr ¨ahenb¨uhl, Objects as points, in: arXiv preprint arXiv:1904.07850, 2019

  57. [65]

    K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, Q. Tian, Centernet: Keypoint triplets for object detection, in: arXiv preprint arXiv:1904.08189, 2019

  58. [66]

    Robbins, S

    H. Robbins, S. Monro, A stochastic approximation method, in: The an- nals of mathematical statistics, 1951

  59. [67]

    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: arXiv preprint arXiv:1412.6980, 2014

  60. [68]

    V . Nair, G. E. Hinton, Rectified linear units improve restricted boltzmann machines, in: ICML, 2010

  61. [69]

    Simonyan, A

    K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: arXiv preprint arXiv:1409.1556, 2014

  62. [70]

    K. He, X. Zhang, S. Ren, J. Sun, Identity mappings in deep residual networks, in: ECCV , Springer, 2016

  63. [71]

    Huang, Z

    G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely con- nected convolutional networks., in: CVPR, 2017

  64. [72]

    Y . Chen, J. Li, H. Xiao, X. Jin, S. Yan, J. Feng, Dual path networks, in: NeurIPS, 2017, pp. 4467–4475

  65. [73]

    S. Xie, R. Girshick, P. Doll ´ar, Z. Tu, K. He, Aggregated residual trans- formations for deep neural networks, in: CVPR, 2017

  66. [74]

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, H. Adam, Mobilenets: Efficient convolu- tional neural networks for mobile vision applications, in: arXiv preprint arXiv:1704.04861, 2017

  67. [75]

    Szegedy, W

    C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. Reed, D. Anguelov, D. Er- han, V . Vanhoucke, A. Rabinovich, Going deeper with convolutions, in: CVPR, 2015

  68. [76]

    Szegedy, V

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna, Rethinking the inception architecture for computer vision, in: CVPR, 2016

  69. [77]

    Szegedy, S

    C. Szegedy, S. Ioffe, V . Vanhoucke, A. A. Alemi, Inception-v4, inception-resnet and the impact of residual connections on learning., in: AAAI, 2017

  70. [78]

    Z. Li, C. Peng, G. Yu, X. Zhang, Y . Deng, J. Sun, Detnet: A backbone network for object detection, in: ECCV , 2018

  71. [79]

    Newell, K

    A. Newell, K. Yang, J. Deng, Stacked hourglass networks for human pose estimation, in: ECCV , 2016

  72. [80]

    Alexe, T

    B. Alexe, T. Deselaers, V . Ferrari, Measuring the objectness of image windows, in: TPAMI, 2012

  73. [81]

    Rahtu, J

    E. Rahtu, J. Kannala, M. Blaschko, Learning a category independent object detection cascade, in: ICCV , 2011

  74. [82]

    P. F. Felzenszwalb, D. P. Huttenlocher, Efficient graph-based image seg- mentation, in: IJCV , 2004

  75. [83]

    Manen, M

    S. Manen, M. Guillaumin, L. Van Gool, Prime object proposals with randomized prim’s algorithm, in: CVPR, 2013

  76. [84]

    Carreira, C

    J. Carreira, C. Sminchisescu, Cpmc: Automatic object segmentation us- ing constrained parametric min-cuts, in: TPAMI, 2011

  77. [85]

    Endres, D

    I. Endres, D. Hoiem, Category-independent object proposals with di- verse ranking, in: TPAMI, 2014

  78. [86]

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, C. L. Zitnick, Microsoft coco: Common objects in context, in: ECCV , 2014

  79. [87]

    Zhang, X

    S. Zhang, X. Zhu, Z. Lei, H. Shi, X. Wang, S. Z. Li, S3fd: Single shot scale-invariant face detector, in: ICCV , 2017

  80. [88]

    W. Luo, Y . Li, R. Urtasun, R. Zemel, Understanding the effective recep- tive field in deep convolutional neural networks, in: NeurIPS, 2016

  81. [89]

    C. Zhu, R. Tao, K. Luu, M. Savvides, Seeing small faces from robust anchors perspective, in: CVPR, 2018

  82. [90]

    L. J. Z. X. Lele Xie, Yuliang Liu, Derpn: Taking a further step toward more general object detection, in: AAAI, 2019. 36

  83. [91]

    Ghodrati, A

    A. Ghodrati, A. Diba, M. Pedersoli, T. Tuytelaars, L. Van Gool, Deep- proposal: Hunting objects by cascading deep convolutional layers, in: ICCV , 2015

  84. [92]

    Zhang, L

    S. Zhang, L. Wen, X. Bian, Z. Lei, S. Z. Li, Single-shot refinement neural network for object detection, in: CVPR, 2018

  85. [93]

    T. Yang, X. Zhang, Z. Li, W. Zhang, J. Sun, Metaanchor: Learning to detect objects with customized anchors, in: NeurIPS, 2018

  86. [94]

    Tychsen-Smith, L

    L. Tychsen-Smith, L. Petersson, Denet: Scalable real-time object detec- tion with directed sparse sampling, in: ICCV , 2017

  87. [95]

    C. Zhu, Y . He, M. Savvides, Feature selective anchor-free module for single-shot object detection, in: CVPR, 2019

  88. [96]

    Y . Lu, T. Javidi, S. Lazebnik, Adaptive object detection using adjacency and zoom prediction, in: CVPR, 2016

  89. [97]

    J. Dai, H. Qi, Y . Xiong, Y . Li, G. Zhang, H. Hu, Y . Wei, Deformable convolutional networks, in: ICCV , 2017

  90. [98]

    Singh, L

    B. Singh, L. S. Davis, An analysis of scale invariance in object detection–snip, in: CVPR, 2018

  91. [99]

    P. Hu, D. Ramanan, Finding tiny faces, in: CVPR, 2017

  92. [100]

    F. Yang, W. Choi, Y . Lin, Exploit all the layers: Fast and accurate cnn ob- ject detector with scale dependent pooling and cascaded rejection clas- sifiers, in: CVPR, 2016

  93. [101]

    Y . Liu, H. Li, J. Yan, F. Wei, X. Wang, X. Tang, Recurrent scale approx- imation for object detection in cnn, in: ICCV , 2017

  94. [102]

    Shrivastava, R

    A. Shrivastava, R. Sukthankar, J. Malik, A. Gupta, Beyond skip con- nections: Top-down modulation for object detection, in: arXiv preprint arXiv:1612.06851, 2016

  95. [103]

    H. Wang, Q. Wang, M. Gao, P. Li, W. Zuo, Multi-scale location-aware kernel representation for object detection, in: CVPR, 2018

  96. [104]

    K.-H. Kim, S. Hong, B. Roh, Y . Cheon, M. Park, Pvanet: deep but lightweight neural networks for real-time object detection, in: arXiv preprint arXiv:1608.08021, 2016

  97. [105]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: International Conference on Medi- cal image computing and computer-assisted intervention, 2015

  98. [106]

    Z. Cai, Q. Fan, R. S. Feris, N. Vasconcelos, A unified multi-scale deep convolutional neural network for fast object detection, in: ECCV , 2016

  99. [107]

    Z. Shen, Z. Liu, J. Li, Y .-G. Jiang, Y . Chen, X. Xue, Dsod: Learning deeply supervised object detectors from scratch, in: ICCV , 2017

  100. [108]

    S. Liu, D. Huang, Y . Wang, Receptive field block net for accurate and fast object detection, in: ECCV , 2018

  101. [109]

    J. Ren, X. Chen, J. Liu, W. Sun, J. Pang, Q. Yan, Y .-W. Tai, L. Xu, Accurate single stage detector using recurrent rolling convolution, in: CVPR, 2017

  102. [110]

    Jeong, H

    J. Jeong, H. Park, N. Kwak, Enhancement of ssd by concatenating fea- ture maps for object detection, in: arXiv preprint arXiv:1705.09587, 2017

  103. [111]

    P. Zhou, B. Ni, C. Geng, J. Hu, Y . Xu, Scale-transferrable object detec- tion, in: CVPR, 2018

  104. [112]

    C.-Y . Fu, W. Liu, A. Ranga, A. Tyagi, A. C. Berg, Dssd: Deconvolu- tional single shot detector, in: arXiv preprint arXiv:1701.06659, 2017

  105. [113]

    S. Woo, S. Hwang, I. S. Kweon, Stairnet: Top-down semantic aggrega- tion for accurate one shot detection, in: 2018 IEEE Winter Conference on Applications of Computer Vision (W ACV), 2018

  106. [114]

    H. Li, Y . Liu, W. Ouyang, X. Wang, Zoom out-and-in network with recursive training for object proposal, in: arXiv preprint arXiv:1702.05711, 2017

  107. [115]

    T. Kong, F. Sun, W. Huang, H. Liu, Deep feature pyramid reconfigura- tion for object detection, in: ECCV , 2018

  108. [116]

    Q. Zhao, T. Sheng, Y . Wang, Z. Tang, Y . Chen, L. Cai, H. Ling, M2det: A single-shot object detector based on multi-level feature pyramid net- work, in: AAAI, 2019

  109. [117]

    Z. Li, F. Zhou, Fssd: Feature fusion single shot multibox detector, in: arXiv preprint arXiv:1712.00960, 2017

  110. [118]

    K. Lee, J. Choi, J. Jeong, N. Kwak, Residual features and uni- fied prediction network for single stage detection, in: arXiv preprint arXiv:1707.05031, 2017

  111. [119]

    Cui, Mdssd: Multi-scale deconvolutional single shot detector for small objects, in: arXiv preprint arXiv:1805.07009, 2018

    L. Cui, Mdssd: Multi-scale deconvolutional single shot detector for small objects, in: arXiv preprint arXiv:1805.07009, 2018

  112. [120]

    T. Kong, F. Sun, A. Yao, H. Liu, M. Lu, Y . Chen, Ron: Reverse con- nection with objectness prior networks for object detection, in: CVPR, 2017

  113. [121]

    B. Lim, S. Son, H. Kim, S. Nah, K. Mu Lee, Enhanced deep residual networks for single image super-resolution, in: CVPR workshops, 2017

  114. [122]

    W. Shi, J. Caballero, F. Husz ´ar, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, Z. Wang, Real-time single image and video super- resolution using an efficient sub-pixel convolutional neural network, in: CVPR, 2016

  115. [123]

    Jiang, R

    B. Jiang, R. Luo, J. Mao, T. Xiao, Y . Jiang, Acquisition of localization confidence for accurate object detection, in: ECCV , 2018

  116. [124]

    Y . Zhai, J. Fu, Y . Lu, H. Li, Feature selective networks for object detec- tion, in: CVPR, 2018

  117. [125]

    Y . Zhu, C. Zhao, J. Wang, X. Zhao, Y . Wu, H. Lu, Couplenet: Coupling global structure with local parts for object detection, in: ICCV , 2017

  118. [126]

    Galleguillos, S

    C. Galleguillos, S. Belongie, Context based object categorization: A critical survey, in: Computer vision and image understanding, 2010

  119. [127]

    Ouyang, X

    W. Ouyang, X. Wang, X. Zeng, S. Qiu, P. Luo, Y . Tian, H. Li, S. Yang, Z. Wang, C.-C. Loy, et al., Deepid-net: Deformable deep convolutional neural networks for object detection, in: CVPR, 2015

  120. [128]

    W. Chu, D. Cai, Deep feature based contextual model for object detec- tion, in: Neurocomputing, 2018

  121. [129]

    Y . Zhu, R. Urtasun, R. Salakhutdinov, S. Fidler, segdeepm: Exploiting segmentation and context in deep neural networks for object detection, in: CVPR, 2015

  122. [130]

    X. Chen, A. Gupta, Spatial memory for context reasoning in object de- tection, in: ICCV , 2017

  123. [131]

    Gidaris, N

    S. Gidaris, N. Komodakis, Object detection via a multi-region and se- 37 mantic segmentation-aware cnn model, in: ICCV , 2015

  124. [132]

    Cheng, Y

    B. Cheng, Y . Wei, H. Shi, R. Feris, J. Xiong, T. Huang, Revisiting rcnn: On awakening the classification power of faster rcnn, in: ECCV , 2018

  125. [133]

    X. Zhao, S. Liang, Y . Wei, Pseudo mask augmented object detection, in: CVPR, 2018

  126. [134]

    Zhang, S

    Z. Zhang, S. Qiao, C. Xie, W. Shen, B. Wang, A. L. Yuille, Single-shot object detection with enriched semantics, Tech. rep. (2018)

  127. [135]

    Shrivastava, A

    A. Shrivastava, A. Gupta, Contextual priming and feedback for faster r-cnn, in: ECCV , 2016

  128. [136]

    B. Li, T. Wu, L. Zhang, R. Chu, Auto-context r-cnn, in: arXiv preprint arXiv:1807.02842, 2018

  129. [137]

    Y . Liu, R. Wang, S. Shan, X. Chen, Structure inference net: Object detection using scene-level context and instance-level relationships, in: CVPR, 2018

  130. [138]

    H. Hu, J. Gu, Z. Zhang, J. Dai, Y . Wei, Relation networks for object detection, in: CVPR, 2018

  131. [139]

    J. Gu, H. Hu, L. Wang, Y . Wei, J. Dai, Learning region features for object detection, in: ECCV , 2018

  132. [140]

    H. Xu, X. Lv, X. Wang, Z. Ren, R. Chellappa, Deep regionlets for object detection, in: ECCV , 2018

  133. [141]

    Z. Chen, S. Huang, D. Tao, Context refinement for object detection, in: ECCV , 2018

  134. [142]

    X. Zeng, W. Ouyang, B. Yang, J. Yan, X. Wang, Gated bi-directional cnn for object detection, in: ECCV , 2016

  135. [143]

    J. Li, Y . Wei, X. Liang, J. Dong, T. Xu, J. Feng, S. Yan, Attentive con- texts for object detection, in: IEEE Transactions on Multimedia, 2017

  136. [144]

    S. L. Xizhou Zhu, Han Hu, J. Dai, Deformable convnets v2: More de- formable, better results, in: CVPR, 2019

  137. [145]

    Girshick, F

    R. Girshick, F. Iandola, T. Darrell, J. Malik, Deformable part models are convolutional neural networks, in: CVPR, 2015

  138. [146]

    Singh, M

    B. Singh, M. Najibi, L. S. Davis, Sniper: Efficient multi-scale training, in: NeurIPS, 2018

  139. [147]

    Y . L. Buyu Li, X. Wang, Gradient harmonized single-stage detector, in: AAAI, 2019

  140. [148]

    Shrivastava, A

    A. Shrivastava, A. Gupta, R. Girshick, Training region-based object de- tectors with online hard example mining, in: CVPR, 2016

  141. [149]

    Gidaris, N

    S. Gidaris, N. Komodakis, Locnet: Improving localization accuracy for object detection, in: CVPR, 2016

  142. [150]

    Zagoruyko, A

    S. Zagoruyko, A. Lerer, T.-Y . Lin, P. O. Pinheiro, S. Gross, S. Chintala, P. Doll´ar, A multipath network for object detection, in: BMVC, 2016

  143. [151]

    X. Lu, B. Li, Y . Yue, Q. Li, J. Yan, Grid r-cnn, in: CVPR, 2019

  144. [152]

    Tychsen-Smith, L

    L. Tychsen-Smith, L. Petersson, Improving object localization with fit- ness nms and bounded iou loss, in: arXiv preprint arXiv:1711.00164, 2017

  145. [153]

    B. Yang, J. Yan, Z. Lei, S. Z. Li, Craft objects from images, in: CVPR, 2016

  146. [154]

    Goodfellow, J

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, Y . Bengio, Generative adversarial nets, in: NeurIPS, 2014

  147. [155]

    J.-Y . Zhu, T. Park, P. Isola, A. A. Efros, Unpaired image-to-image trans- lation using cycle-consistent adversarial networkss, in: ICCV , 2017

  148. [156]

    Radford, L

    A. Radford, L. Metz, S. Chintala, Unsupervised representation learn- ing with deep convolutional generative adversarial networks, in: arXiv preprint arXiv:1511.06434, 2015

  149. [157]

    Brock, J

    A. Brock, J. Donahue, K. Simonyan, Large scale gan training for high fidelity natural image synthesis, in: arXiv preprint arXiv:1809.11096, 2018

  150. [158]

    J. Li, X. Liang, Y . Wei, T. Xu, J. Feng, S. Yan, Perceptual generative adversarial networks for small object detection, in: CVPR, 2017

  151. [159]

    X. Wang, A. Shrivastava, A. Gupta, A-fast-rcnn: Hard positive genera- tion via adversary for object detection, in: CVPR, 2017

  152. [160]

    R. G. Kaiming He, P. Dollro, Rethinking imagenet pre-training, in: arXiv preprint arXiv:1811.08883, 2018

  153. [161]

    R. Zhu, S. Zhang, X. Wang, L. Wen, H. Shi, L. Bo, T. Mei, Scratchdet: Exploring to train single-shot object detectors from scratch, in: CVPR, 2019

  154. [162]

    Z. Shen, H. Shi, R. Feris, L. Cao, S. Yan, D. Liu, X. Wang, X. Xue, T. S. Huang, Learning object detectors from scratch with gated recurrent feature pyramids, in: arXiv preprint arXiv:1712.00886, 2017

  155. [163]

    Hinton, O

    G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network, in: arXiv preprint arXiv:1503.02531, 2015

  156. [164]

    Q. Li, S. Jin, J. Yan, Mimicking very efficient network for object detec- tion, in: CVPR, 2017

  157. [165]

    Bodla, B

    N. Bodla, B. Singh, R. Chellappa, L. S. Davis, Soft-nms – improving object detection with one line of code, in: ICCV , 2017

  158. [166]

    Hosang, R

    J. Hosang, R. Benenson, B. Schiele, Learning non-maximum suppres- sion, in: CVPR, 2017

  159. [167]

    Huang, V

    J. Huang, V . Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y . Song, S. Guadarrama, et al., Speed/accuracy trade-offs for modern convolutional object detectors, in: CVPR, 2017

  160. [168]

    Z. Li, C. Peng, G. Yu, X. Zhang, Y . Deng, J. Sun, Light-head r- cnn: In defense of two-stage object detector, in: arXiv preprint arXiv:1711.07264, 2017

  161. [169]

    Sandler, A

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L.-C. Chen, Inverted residuals and linear bottlenecks: Mobile networks for classification, de- tection and segmentation, in: arXiv preprint arXiv:1801.04381, 2018

  162. [170]

    A. Wong, M. J. Shafiee, F. Li, B. Chwyl, Tiny ssd: A tiny single-shot detection deep convolutional neural network for real-time embedded ob- ject detection, in: arXiv preprint arXiv:1802.06488, 2018

  163. [171]

    Y . Li, J. Li, W. Lin, J. Li, Tiny-dsod: Lightweight object detection for resource-restricted usages, in: arXiv preprint arXiv:1807.11013, 2018

  164. [172]

    K. S. D. A. Shang, Wenling, H. Lee., Understanding and improving convolutional neural networks via concatenated rectified linear units, in: ICML, 2016

  165. [173]

    Y . D. Kim, E. Park, S. Yoo, T. Choi, L. Yang, D. Shin, Compression of deep convolutional neural networks for fast and low power mobile 38 applications, in: Computer Science, 2015

  166. [174]

    Y . He, X. Zhang, J. Sun, Channel pruning for accelerating very deep neural networks, in: ICCV , 2017

  167. [175]

    Y . Gong, L. Liu, M. Yang, L. Bourdev, Compressing deep convolutional networks using vector quantization, in: Computer Science, 2014

  168. [176]

    Y . Lin, S. Han, H. Mao, Y . Wang, W. J. Dally, Deep gradient compres- sion: Reducing the communication bandwidth for distributed training, in: arXiv preprint arXiv:1712.01887, 2017

  169. [177]

    J. Wu, L. Cong, Y . Wang, Q. Hu, J. Cheng, Quantized convolutional neural networks for mobile devices, in: CVPR, 2016

  170. [178]

    S. Han, H. Mao, W. J. Dally, Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, in: Fiber, 2015

  171. [179]

    S. Han, J. Pool, J. Tran, W. Dally, Learning both weights and connec- tions for efficient neural network, in: NeurIPS, 2015

  172. [180]

    Osuna, R

    E. Osuna, R. Freund, F. Girosit, Training support vector machines: an application to face detection, in: CVPR, 1997

  173. [181]

    R ¨atsch, S

    M. R ¨atsch, S. Romdhani, T. Vetter, Efficient face detection by a cas- caded support vector machine using haar-like features, in: Joint Pattern Recognition Symposium, 2004

  174. [182]

    Romdhani, P

    S. Romdhani, P. Torr, B. Scholkopf, A. Blake, Computationally efficient face detection, in: ICCV , 2001

  175. [183]

    X. Sun, P. Wu, S. C. Hoi, Face detection using deep learning: An im- proved faster rcnn approach, in: Neurocomputing, 2018

  176. [184]

    Y . Liu, M. D. Levine, Multi-path region-based convolutional neural net- work for accurate detection of unconstrained” hard faces”, in: Computer and Robot Vision (CRV), 2017 14th Conference on, 2017

  177. [185]

    X. Tang, D. K. Du, Z. He, J. Liu, Pyramidbox: A context-assisted single shot face detector, in: ECCV , 2018

  178. [186]

    C. Chi, S. Zhang, J. Xing, Z. Lei, S. Z. Li, X. Zou, Selective refine- ment network for high performance face detection, in: arXiv preprint arXiv:1809.02693, 2018

  179. [187]

    J. Li, Y . Wang, C. Wang, Y . Tai, J. Qian, J. Yang, C. Wang, J. Li, F. Huang, Dsfd: Dual shot face detector, in: CVPR, 2019

  180. [188]

    Zhang, X

    J. Zhang, X. Wu, J. Zhu, S. C. Hoi, Feature agglomeration networks for single stage face detection, in: arXiv preprint arXiv:1712.00721, 2017

  181. [189]

    Najibi, P

    M. Najibi, P. Samangouei, R. Chellappa, L. Davis, Ssh: Single stage headless face detector, in: ICCV , 2017

  182. [190]

    Z. Hao, Y . Liu, H. Qin, J. Yan, X. Li, X. Hu, Scale-aware face detection, in: CVPR, 2017

  183. [191]

    H. Wang, Z. Li, X. Ji, Y . Wang, Face r-cnn, in: arXiv preprint arXiv:1706.01061, 2017

  184. [192]

    Zhang, Z

    K. Zhang, Z. Zhang, Z. Li, Y . Qiao, Jjoint face detection and alignment using multi-task cascaded convolutional networks, in: IEEE Signal Pro- cessing Letters, 2016

  185. [193]

    Samangouei, M

    P. Samangouei, M. Najibi, L. Davis, R. Chellappa, Face-magnet: Magnifying feature maps to detect small faces, in: arXiv preprint arXiv:1803.05258, 2018

  186. [194]

    Zhang, X

    C. Zhang, X. Xu, D. Tu, Face detection using improved faster rcnn, in: arXiv preprint arXiv:1802.02142, 2018

  187. [195]

    C. Zhu, Y . Zheng, K. Luu, M. Savvides, Cms-rcnn: Contextual multi- scale region-based cnn for unconstrained face detection, in: Deep Learn- ing for Biometrics, 2017

  188. [196]

    B. Yu, D. Tao, Anchor cascade for efficient face detection, in: arXiv preprint arXiv:1805.03363, 2018

  189. [197]

    Zhang, Z

    K. Zhang, Z. Zhang, H. Wang, Z. Li, Y . Qiao, W. Liu, Detecting faces using inside cascaded contextual cnn, in: ICCV , 2017

  190. [198]

    Y . Wang, X. Ji, Z. Zhou, H. Wang, Z. Li, Detecting faces us- ing region-based fully convolutional networks, in: arXiv preprint arXiv:1709.05256, 2017

  191. [199]

    V . Jain, E. Learned-Miller, Fddb: A benchmark for face detection in un- constrained settings, Tech. Rep. UM-CS-2010-009, University of Mas- sachusetts, Amherst (2010)

  192. [200]

    S. Yang, P. Luo, C.-C. Loy, X. Tang, Wwider face: A face detection benchmark, in: CVPR, 2016

  193. [201]

    J. Han, W. Nam, P. Dollar, Local decorrelation for improved detection, in: NeurIPS, 2014

  194. [202]

    Doll ´ar, Z

    P. Doll ´ar, Z. Tu, P. Perona, S. Belongie, Integral channel features, in: BMVC, 2009

  195. [203]

    Doll ´ar, R

    P. Doll ´ar, R. Appel, S. Belongie, P. Perona, Fast feature pyramids for object detection, in: TPAMI, 2014

  196. [204]

    C. P. Papageorgiou, M. Oren, T. Poggio, A general framework for object detection, in: ICCV , 1998

  197. [205]

    Brazil, X

    G. Brazil, X. Yin, X. Liu, Illuminating pedestrians via simultaneous de- tection & segmentation, in: arXiv preprint arXiv:1706.08564, 2017

  198. [206]

    X. Du, M. El-Khamy, J. Lee, L. Davis, Fused dnn: A deep neural net- work fusion approach to fast and robust pedestrian detection, in: IEEE Winter Conference on Applications of Computer Vision (W ACV), 2017

  199. [207]

    S. Wang, J. Cheng, H. Liu, M. Tang, Pcn: Part and context information for pedestrian detection with cnns, in: arXiv preprint arXiv:1804.04483, 2018

  200. [208]

    D. Xu, W. Ouyang, E. Ricci, X. Wang, N. Sebe, Learning cross-modal deep representations for robust pedestrian detection, in: CVPR, 2017

  201. [209]

    Benenson, M

    R. Benenson, M. Omran, J. Hosang, B. Schiele, Ten years of pedestrian detection, what have we learned?, in: ECCV , 2014

  202. [210]

    Z. Cai, M. Saberian, N. Vasconcelos, Learning complexity-aware cas- cades for deep pedestrian detection, in: ICCV , 2015

  203. [211]

    Sermanet, K

    P. Sermanet, K. Kavukcuoglu, S. Chintala, Y . LeCun, Pedestrian detec- tion with unsupervised multi-stage feature learning, in: CVPR, 2013

  204. [212]

    Zhang, L

    L. Zhang, L. Lin, X. Liang, K. He, Is faster r-cnn doing well for pedes- trian detection?, in: ECCV , 2016

  205. [213]

    X. Wang, T. Xiao, Y . Jiang, S. Shao, J. Sun, C. Shen, Repulsion loss: Detecting pedestrians in a crowd, in: CVPR, 2018

  206. [214]

    Zhang, L

    S. Zhang, L. Wen, X. Bian, Z. Lei, S. Z. Li, Occlusion-aware r-cnn: Detecting pedestrians in a crowd, in: ECCV , 2018

  207. [215]

    J. Mao, T. Xiao, Y . Jiang, Z. Cao, What can help pedestrian detection?, 39 in: CVPR, 2017

  208. [216]

    Y . Tian, P. Luo, X. Wang, X. Tang, Deep learning strong parts for pedes- trian detection, in: CVPR, 2015

  209. [217]

    Ouyang, X

    W. Ouyang, X. Wang, Joint deep learning for pedestrian detection, in: ICCV , 2013

  210. [218]

    Mathias, R

    M. Mathias, R. Benenson, R. Timofte, L. Van Gool, Handling occlusions with franken-classifiers, in: ICCV , 2013

  211. [219]

    Ouyang, X

    W. Ouyang, X. Zeng, X. Wang, Modeling mutual visibility relationship in pedestrian detection, in: CVPR, 2013

  212. [220]

    G. Duan, H. Ai, S. Lao, A structural filter approach to human detection, in: ECCV , 2010

  213. [221]

    Enzweiler, A

    M. Enzweiler, A. Eigenstetter, B. Schiele, D. M. Gavrila, Multi-cue pedestrian classification with partial occlusion handling, in: CVPR, 2010

  214. [222]

    C. Zhou, J. Yuan, Bi-box regression for pedestrian detection and occlu- sion estimation, in: ECCV , 2018

  215. [223]

    Ouyang, X

    W. Ouyang, X. Wang, A discriminative deep model for pedestrian de- tection with occlusion handling, in: CVPR, 2012

  216. [224]

    S. Tang, M. Andriluka, B. Schiele, Detection and tracking of occluded people, in: IJCV , 2014

  217. [225]

    Ouyang, X

    W. Ouyang, X. Wang, Single-pedestrian detection aided by multi- pedestrian detection, in: CVPR, 2013

  218. [226]

    V . D. Shet, J. Neumann, V . Ramesh, L. S. Davis, Bilattice-based logical reasoning for human detection, in: CVPR, 2007

  219. [227]

    Y . Zhou, L. Liu, L. Shao, M. Mellor, Dave: A unified framework for fast vehicle detection and annotation, in: ECCV , 2016

  220. [228]

    Gebru, J

    T. Gebru, J. Krause, Y . Wang, D. Chen, J. Deng, L. Fei-Fei, Fine-grained car detection for visual census estimation, in: AAAI, 2017

  221. [229]

    Majid Azimi, Shuffledet: Real-time vehicle detection network in on- board embedded uav imagery, in: ECCV , 2018

    S. Majid Azimi, Shuffledet: Real-time vehicle detection network in on- board embedded uav imagery, in: ECCV , 2018

  222. [230]

    Z. Zhu, D. Liang, S. Zhang, X. Huang, B. Li, S. Hu, Traffic-sign detec- tion and classification in the wild, in: CVPR, 2016

  223. [231]

    A. Pon, O. Adrienko, A. Harakeh, S. L. Waslander, A hierarchical deep architecture and mini-batch selection method for joint traffic sign and light detection, in: Conference on Computer and Robot Vision (CRV), 2018

  224. [232]

    W. Ke, J. Chen, J. Jiao, G. Zhao, Q. Ye, Srn: side-output residual net- work for object symmetry detection in the wild, in: CVPR, 2017

  225. [233]

    W. Shen, K. Zhao, Y . Jiang, Y . Wang, Z. Zhang, X. Bai, Object skeleton extraction in natural images by fusing scale-associated deep side out- puts, in: CVPR, 2016

  226. [234]

    Kuznetsova, H

    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, T. Duerig, et al., The open images dataset v4: Unified image classification, object detection, and visual re- lationship detection at scale, in: arXiv preprint arXiv:1811.0...

  227. [235]

    Gupta, P

    A. Gupta, P. Dollar, R. Girshick, Lvis: A dataset for large vocabulary instance segmentation, in: CVPR, 2019

  228. [236]

    Bodla, B

    N. Bodla, B. Singh, R. Chellappa, L. S. Davis, Soft-nms–improving ob- ject detection with one line of code, in: ICCV , 2017

  229. [237]

    Y . Wu, K. He, Group normalization, in: ECCV , 2018

  230. [238]

    S. Liu, L. Qi, H. Qin, J. Shi, J. Jia, Path aggregation network for instance segmentation, in: CVPR, 2018

  231. [239]

    Y . Li, Y . Chen, N. Wang, Z. Zhang, Scale-aware trident networks for object detection, in: arXiv preprint arXiv:1901.01892, 2019

  232. [240]

    X. Zhou, J. Zhuo, P. Krahenbuhl, Bottom-up object detection by group- ing extreme and center points, in: CVPR, 2019

  233. [241]

    Z. Tian, C. Shen, H. Chen, T. He, Fcos: Fully convolutional one-stage object detection, in: arXiv preprint arXiv:1904.01355, 2019

  234. [242]

    Zhang, R

    S. Zhang, R. Benenson, B. Schiele, Citypersons: A diverse dataset for pedestrian detection, in: CVPR, 2017

  235. [243]

    Cordts, M

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benen- son, U. Franke, S. Roth, B. Schiele, The cityscapes dataset for semantic urban scene understanding, in: CVPR, 2016

  236. [244]

    Dollar, C

    P. Dollar, C. Wojek, B. Schiele, P. Perona, Pedestrian detection: An evaluation of the state of the art, in: TPAMI, 2012

  237. [245]

    A. Ess, B. Leibe, L. Van Gool, Depth and appearance for mobile scene analysis, in: ICCV , 2007

  238. [246]

    Geiger, P

    A. Geiger, P. Lenz, C. Stiller, R. Urtasun, Vision meets robotics: The kitti dataset, 2013

  239. [247]

    B. Zoph, V . Vasudevan, J. Shlens, Q. V . Le, Learning transferable archi- tectures for scalable image recognition, in: CVPR, 2018

  240. [248]

    M. Tan, Q. V . Le, Efficientnet: Rethinking model scaling for convolu- tional neural networks, in: arXiv preprint arXiv:1905.11946, 2019

  241. [249]

    Y . Chen, T. Yang, X. Zhang, G. Meng, C. Pan, J. Sun, Detnas: Neural architecture search on object detection, in: arXiv preprint arXiv:1903.10979, 2019

  242. [250]

    Ghiasi, T.-Y

    G. Ghiasi, T.-Y . Lin, Q. V . Le, Nas-fpn: Learning scalable feature pyra- mid architecture for object detection, in: CVPR, 2019

  243. [251]

    B. Zoph, E. D. Cubuk, G. Ghiasi, T.-Y . Lin, J. Shlens, Q. V . Le, Learn- ing data augmentation strategies for object detection, in: arXiv preprint arXiv:1906.11172, 2019

  244. [252]

    X. Dong, L. Zheng, F. Ma, Y . Yang, D. Meng, Few-example object de- tection with model communication, in: TPAMI, 2018

  245. [253]

    Schwartz, L

    E. Schwartz, L. Karlinsky, J. Shtok, S. Harary, M. Marder, S. Pankanti, R. Feris, A. Kumar, R. Giries, A. M. Bronstein, Repmet: Representative- based metric learning for classification and one-shot object detection, in: CVPR, 2019

  246. [254]

    H. Chen, Y . Wang, G. Wang, Y . Qiao, Lstd: A low-shot transfer detector for object detection, in: AAAI, 2018

  247. [255]

    C. Peng, T. Xiao, Z. Li, Y . Jiang, X. Zhang, K. Jia, G. Yu, J. Sun, Megdet: A large mini-batch object detector, in: CVPR, 2018

  248. [256]

    Shmelkov, C

    K. Shmelkov, C. Schmid, K. Alahari, Incremental learning of object de- tectors without catastrophic forgetting, in: ICCV , 2017. 40

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.