REVIEW 11 cited by
YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
YOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 160 FPS and has the highest accuracy 56.8% AP among all known real-time object detectors with 30 FPS or higher on GPU V100. YOLOv7-E6 object detector (56 FPS V100, 55.9% AP) outperforms both transformer-based detector SWIN-L Cascade-Mask R-CNN (9.2 FPS A100, 53.9% AP) by 509% in speed and 2% in accuracy, and convolutional-based detector ConvNeXt-XL Cascade-Mask R-CNN (8.6 FPS A100, 55.2% AP) by 551% in speed and 0.7% AP in accuracy, as well as YOLOv7 outperforms: YOLOR, YOLOX, Scaled-YOLOv4, YOLOv5, DETR, Deformable DETR, DINO-5scale-R50, ViT-Adapter-B and many other object detectors in speed and accuracy. Moreover, we train YOLOv7 only on MS COCO dataset from scratch without using any other datasets or pre-trained weights. Source code is released in https://github.com/WongKinYiu/yolov7.
Forward citations
Cited by 11 Pith papers
-
Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR
Supervising a pre-localization priority map, whether from boxes, real gaze, or transferred pseudo-gaze, improves infrared weak-small UAV detection over DINO-DETR.
-
Resonance-enhanced integrated acousto-optic beam steering
A TFLN ring-resonator-enhanced acousto-optic beam steerer reaches 26% efficiency and 18° FOV and supports FMCW LiDAR via electro-optic resonance locking.
-
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.
-
Integrated Detection and Tracking Based on Radar Range-Doppler Feature
A deep learning detector and Kalman tracker are integrated with a three-channel Range-Doppler input, confidence-adaptive measurement noise, and feature-based data association, improving low-SNR radar detection and tracking.
-
Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
ARTSS, a deep learning pipeline for automated Sharp/van der Heijde rheumatoid arthritis scoring from hand X-rays, reports MAE 0.95 and 99% joint detection, but its key results table contains a mathematically impossibl...
-
Synthesizing Reality: Leveraging the Generative AI-Powered Platform Midjourney for Construction Worker Detection
Training YOLOv7 on 11,992 Midjourney-synthesized construction worker images transfers to real construction photos, with AP0.5 0.937 and AP0.5:0.95 0.642.
-
FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection
FedTR (public pre-train + FedAvg fine-tune) reaches 95.5%/94.2% end-to-end word accuracy for industrial label text recognition under homogeneous/heterogeneous plant data, matching centralized performance while keeping...
-
Single Domain Generalization in Diabetic Retinopathy: A Neuro-Symbolic Learning Approach
KG-DG fuses YOLO-derived lesion features with a frozen ViT via confidence-based fusion and claims gains in diabetic retinopathy domain generalization, but the central KL-divergence mechanism and the MDG headline are c...
-
Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
A retrospective of four data-limited computer vision challenges finds that ensembles and heavy augmentation, not novel inductive priors, drove winning performance.
-
A smart fridge with AI-enabled food computing
A smart fridge system uses YOLO and compares BCE, focal, and adaptive focal losses, finding BCE best calibrated despite the abstract's claim that focal loss fixes calibration.
-
YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
A survey of YOLO object detectors from version 1 to version 11 that compiles architectures, benchmarks, and applications, with several factual inconsistencies.
Discussion (0). Sign in to comment.