Pith. sign in

REVIEW 18 cited by

D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13842 v1 pith:3LUDM6QH submitted 2024-10-17 cs.CV

classification cs.CV
keywords d-finelocalizationfine-grainedmodelsregressionaccuracyachievesd-fine-l
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and pretrained models: https://github.com/Peterande/D-FINE.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ContextShift: A Controlled Benchmark for Context Dependence in Object Detection

    cs.CV 2026-06 conditional novelty 7.0 of 10

    ContextShift benchmark on COCO reveals up to 227% more false negatives and 44% fewer predictions under controlled context changes, non-monotonic NPMI response, and gains from context-aware augmentation.

  2. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Ev-DTAD improves event-based object detection accuracy and speed by using hierarchical temporal aggregation at the representation level and frequency-aware hypergraph fusion at the model level.

  3. Training-Free Semantic Multi-Object Tracking with Vision-Language Models

    cs.CV 2026-04 conditional novelty 7.0 of 10

    TF-SMOT composes pretrained vision-language models into a training-free pipeline that reaches state-of-the-art tracking and improved summary quality on the BenSMOT benchmark.

  4. WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    WUTDet is a 100K-image ship detection dataset with benchmarks indicating Transformer models outperform CNN and Mamba architectures in accuracy and small-object detection for complex maritime environments.

  5. SARES-DEIM: Sparse Mixture-of-Experts Meets DETR for Robust SAR Ship Detection

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SARES-DEIM achieves 76.4% mAP50:95 and 93.8% mAP50 on HRSID by routing SAR features through sparse frequency and wavelet experts plus a high-resolution preservation neck, outperforming prior YOLO and SAR detectors.

  6. MORE: A Multilingual Document Parsing Benchmark and Evaluation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MORE provides a 149-language, structure-aware document parsing benchmark from real PDFs and reports baselines showing specialized OCR models still fail on tables and rare scripts.

  7. From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Proposes DERNet with Decompose-Enhance-Reconstruct operator and three plug-and-play modules to shift small object detection from spatial to spectral feature processing, claiming better performance than YOLOv11 with 1/...

  8. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Ev-DTAD introduces hierarchical temporal aggregation for event representation and frequency-aware hypergraph fusion for feature reasoning, delivering accuracy and speed gains on Gen1, 1Mpx/Gen4, and eTraM event detect...

  9. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Ev-DTAD combines hierarchical temporal aggregation into a pseudo-RGB representation with frequency-aware hypergraph fusion to improve accuracy and speed in event-based object detection on Gen1, Gen4, and eTraM datasets.

  10. ZoomSpec: A Physics-Guided Coarse-to-Fine Framework for Wideband Spectrum Sensing

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    ZoomSpec achieves 78.1 mAP@0.5:0.95 on the SpaceNet dataset by combining log-space STFT, a coarse proposal net, adaptive heterodyne filtering, and dual-domain fine recognition to improve narrowband visibility in wideb...

  11. Hyper-FEOD: Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Hyper-FEOD fuses RGB and event data via sparse hypergraph cross-modal fusion and region-specialized MoE experts to improve accuracy-efficiency in object detection.

  12. RiO-DETR: DETR for Real-time Oriented Object Detection

    cs.CV 2026-03 conditional novelty 6.0 of 10

    RiO-DETR gives the first real-time oriented DETR, matching or beating CNN real-time detectors on DOTA-1.0, DIOR-R, and FAIR-1M-2.0 with a new speed-accuracy trade-off.

  13. PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    PS-Track sets a new state-of-the-art for point-supervised multi-object tracking by converting point seeds into temporally consistent pseudo-labels via Temporal-Feedback Prompting, Point-Excited Wavelet Attention, and ...

  14. M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    M^2C-EvDet proposes Adaptive Frequency-Decoupled Feature Distillation (AF^2D^2) and Multi-Order Relational Distillation (MORD) modules to reduce the performance gap between event-based and frame-based object detection.

  15. When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    Forensic fine-tuning of vision foundation models leaves semantic structure intact ("semantic fallback"); suppressing CLIP-estimated semantic subspaces via SVD is claimed to yield more generalizable AI-image detectors.

  16. FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection

    cs.CV 2025-09 unverdicted novelty 5.0 of 10

    FMC-DETR proposes a frequency-decoupled fusion framework with WeKat backbone, MDFC coordination, and CPF fusion modules that claims state-of-the-art results on remote sensing object detection benchmarks.

  17. RT-SDGOD: Real-Time Single-Domain Generalized Object Detection

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    RT-SDGDet applies one-to-many supervision, Discriminative Evidence Diversity Learning, and Dual-view Evidence Consistency Learning during training to reduce missed detections in real-time object detectors under unseen...

  18. Hyper-FEOD: Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

    cs.CV 2026-04 reject novelty 4.0 of 10

    Reported frame-event detection SOTA from a claimed sparse-hypergraph + MoE design, but the paper supplies neither the MoE nor the sparse selection and instead provides a masked distillation loss.

Pith tools