Pith. sign in

REVIEW 11 cited by

YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.02696 v1 pith:KFTUO2YQ submitted 2022-07-06 cs.CV

classification cs.CV
keywords accuracyobjectyolov7detectorsspeeddetectora100cascade-mask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

YOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 160 FPS and has the highest accuracy 56.8% AP among all known real-time object detectors with 30 FPS or higher on GPU V100. YOLOv7-E6 object detector (56 FPS V100, 55.9% AP) outperforms both transformer-based detector SWIN-L Cascade-Mask R-CNN (9.2 FPS A100, 53.9% AP) by 509% in speed and 2% in accuracy, and convolutional-based detector ConvNeXt-XL Cascade-Mask R-CNN (8.6 FPS A100, 55.2% AP) by 551% in speed and 0.7% AP in accuracy, as well as YOLOv7 outperforms: YOLOR, YOLOX, Scaled-YOLOv4, YOLOv5, DETR, Deformable DETR, DINO-5scale-R50, ViT-Adapter-B and many other object detectors in speed and accuracy. Moreover, we train YOLOv7 only on MS COCO dataset from scratch without using any other datasets or pre-trained weights. Source code is released in https://github.com/WongKinYiu/yolov7.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 845 citations worldwide. Full citation record

  1. Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Supervising a pre-localization priority map, whether from boxes, real gaze, or transferred pseudo-gaze, improves infrared weak-small UAV detection over DINO-DETR.

  2. Resonance-enhanced integrated acousto-optic beam steering

    physics.app-ph 2026-03 unverdicted novelty 6.0 of 10

    A TFLN ring-resonator-enhanced acousto-optic beam steerer reaches 26% efficiency and 18° FOV and supports FMCW LiDAR via electro-optic resonance locking.

  3. SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models

    cs.RO 2025-11 conditional novelty 6.0 of 10

    SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.

  4. Integrated Detection and Tracking Based on Radar Range-Doppler Feature

    eess.SP 2025-09 conditional novelty 6.0 of 10

    A deep learning detector and Kalman tracker are integrated with a three-channel Range-Doppler input, confidence-adaptive measurement noise, and feature-based data association, improving low-SNR radar detection and tracking.

  5. Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice

    cs.CV 2025-09 reject novelty 5.0 of 10

    ARTSS, a deep learning pipeline for automated Sharp/van der Heijde rheumatoid arthritis scoring from hand X-rays, reports MAE 0.95 and 99% joint detection, but its key results table contains a mathematically impossibl...

  6. Synthesizing Reality: Leveraging the Generative AI-Powered Platform Midjourney for Construction Worker Detection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Training YOLOv7 on 11,992 Midjourney-synthesized construction worker images transfers to real construction photos, with AP0.5 0.937 and AP0.5:0.95 0.642.

  7. FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection

    cs.CV 2026-07 conditional novelty 4.0 of 10

    FedTR (public pre-train + FedAvg fine-tune) reaches 95.5%/94.2% end-to-end word accuracy for industrial label text recognition under homogeneous/heterogeneous plant data, matching centralized performance while keeping...

  8. Single Domain Generalization in Diabetic Retinopathy: A Neuro-Symbolic Learning Approach

    cs.CV 2025-09 reject novelty 4.0 of 10

    KG-DG fuses YOLO-derived lesion features with a frozen ViT via confidence-based fusion and claims gains in diabetic retinopathy domain generalization, but the central KL-divergence mechanism and the MDG headline are c...

  9. Data-Efficient Challenges in Visual Inductive Priors: A Retrospective

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A retrospective of four data-limited computer vision challenges finds that ensembles and heavy augmentation, not novel inductive priors, drove winning performance.

  10. A smart fridge with AI-enabled food computing

    eess.SY 2025-09 reject novelty 2.0 of 10

    A smart fridge system uses YOLO and compares BCE, focal, and adaptive focal losses, finding BCE best calibrated despite the abstract's claim that focal loss fixes calibration.

  11. YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges

    cs.CV 2025-08 unverdicted

    A survey of YOLO object detectors from version 1 to version 11 that compiles architectures, benchmarks, and applications, with several factual inconsistencies.

Pith tools