Pith. sign in

REVIEW 3 cited by

detrex: Benchmarking Detection Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07265 v2 pith:YELUQDHC submitted 2023-06-12 cs.CV

classification cs.CV
keywords detrexdetectiondetr-basedmodelsunifiedalgorithmsbenchmarkcodebase
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The DEtection TRansformer (DETR) algorithm has received considerable attention in the research community and is gradually emerging as a mainstream approach for object detection and other perception tasks. However, the current field lacks a unified and comprehensive benchmark specifically tailored for DETR-based models. To address this issue, we develop a unified, highly modular, and lightweight codebase called detrex, which supports a majority of the mainstream DETR-based instance recognition algorithms, covering various fundamental tasks, including object detection, segmentation, and pose estimation. We conduct extensive experiments under detrex and perform a comprehensive benchmark for DETR-based models. Moreover, we enhance the performance of detection transformers through the refinement of training hyper-parameters, providing strong baselines for supported algorithms.We hope that detrex could offer research communities a standardized and unified platform to evaluate and compare different DETR-based models while fostering a deeper understanding and driving advancements in DETR-based instance recognition. Our code is available at https://github.com/IDEA-Research/detrex. The project is currently being actively developed. We encourage the community to use detrex codebase for further development and contributions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A flexible-query transformer detector that separates cross-attention localization from self-attention deduplication reports consistent accuracy gains over DINO across five backbones.

  2. Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Cross-DINO improves small-object detection in DETR-like detectors by mixing MLP backbone features, a cross-coding fusion module, and a category-size soft-label loss, achieving 36.4% APs on COCO.

  3. What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A framework that generates image-level object concepts with a vision-language model before region segmentation improves open-vocabulary segmentation on multiple benchmarks.

Pith tools