Pith. sign in

REVIEW 9 cited by

Simple Open-Vocabulary Object Detection with Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.06230 v2 pith:IGY5WSB3 submitted 2022-05-12 cs.CV

classification cs.CV
keywords detectionobjectpre-trainingopen-vocabularyimage-textimprovementsmodelsscaling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the long-tailed and open-vocabulary setting, where training data is relatively scarce. In this paper, we propose a strong recipe for transferring image-text models to open-vocabulary object detection. We use a standard Vision Transformer architecture with minimal modifications, contrastive image-text pre-training, and end-to-end detection fine-tuning. Our analysis of the scaling properties of this setup shows that increasing image-level pre-training and model size yield consistent improvements on the downstream detection task. We provide the adaptation strategies and regularizations needed to attain very strong performance on zero-shot text-conditioned and one-shot image-conditioned object detection. Code and models are available on GitHub.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A shared-weight stereo ViT with cross-attention fusion segments the catheter in two views and regresses 3D tip forces, claiming state-of-the-art results on synthetic X-ray datasets.

  2. On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking

    cs.IR 2025-08 unverdicted novelty 6.0 of 10

    Multimodal entity linking models are vulnerable to visual adversarial perturbations, and the proposed retrieval-augmented LLM method (LLM-RetLink) reportedly improves accuracy by 0.4% to 35.7%.

  3. Text-guided Generation of Efficient Personalized Inspection Plans

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A training-free pipeline uses a vision-language model and segmentation to convert text instructions into smooth, order-respecting drone inspection trajectories in known 3D maps.

  4. Instance Segmentation of Scene Sketches Using Natural Image Priors

    cs.CV 2025-02 conditional novelty 6.0 of 10

    InkLayer combines class-agnostic fine-tuning of Grounding DINO, SAM masks, and depth-based foreground refinement to achieve state-of-the-art instance segmentation on scene sketches across styles.

  5. A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

    cs.RO 2025-02 conditional novelty 6.0 of 10

    IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.

  6. Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Textual Inversion is applied to open-vocabulary object detectors to learn a few new tokens while keeping the VLM frozen and preserving zero-shot abilities.

  7. (Almost) Free Modality Stitching of Foundation Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A hypernetwork that generates connector weights for all image-text model pairs can rank pairs like grid search at about 10x lower training cost, but the best connector lags grid search by a few points.

  8. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  9. Towards Wearable Interfaces for Robotic Caregiving

    cs.RO 2025-02 conditional novelty 4.0 of 10

    This paper reports lessons from prior HAT teleoperation studies, preliminary shared-control gains, and a passive-control concept for robot-assisted feeding that is not yet evaluated.

Pith tools