REVIEW 9 cited by
Simple Open-Vocabulary Object Detection with Vision Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the long-tailed and open-vocabulary setting, where training data is relatively scarce. In this paper, we propose a strong recipe for transferring image-text models to open-vocabulary object detection. We use a standard Vision Transformer architecture with minimal modifications, contrastive image-text pre-training, and end-to-end detection fine-tuning. Our analysis of the scaling properties of this setup shows that increasing image-level pre-training and model size yield consistent improvements on the downstream detection task. We provide the adaptation strategies and regularizations needed to attain very strong performance on zero-shot text-conditioned and one-shot image-conditioned object detection. Code and models are available on GitHub.
Forward citations
Cited by 9 Pith papers
-
TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
A shared-weight stereo ViT with cross-attention fusion segments the catheter in two views and regresses 3D tip forces, claiming state-of-the-art results on synthetic X-ray datasets.
-
On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking
Multimodal entity linking models are vulnerable to visual adversarial perturbations, and the proposed retrieval-augmented LLM method (LLM-RetLink) reportedly improves accuracy by 0.4% to 35.7%.
-
Text-guided Generation of Efficient Personalized Inspection Plans
A training-free pipeline uses a vision-language model and segmentation to convert text instructions into smooth, order-respecting drone inspection trajectories in known 3D maps.
-
Instance Segmentation of Scene Sketches Using Natural Image Priors
InkLayer combines class-agnostic fine-tuning of Grounding DINO, SAM masks, and depth-based foreground refinement to achieve state-of-the-art instance segmentation on scene sketches across styles.
-
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.
-
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
Textual Inversion is applied to open-vocabulary object detectors to learn a few new tokens while keeping the VLM frozen and preserving zero-shot abilities.
-
(Almost) Free Modality Stitching of Foundation Models
A hypernetwork that generates connector weights for all image-text model pairs can rank pairs like grid search at about 10x lower training cost, but the best connector lags grid search by a few points.
-
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.
-
Towards Wearable Interfaces for Robotic Caregiving
This paper reports lessons from prior HAT teleoperation studies, preliminary shared-control gains, and a passive-control concept for robot-assisted feeding that is not yet evaluated.
Discussion (0). Continue with ORCID to comment.