REVIEW 6 cited by
F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present F-VLM, a simple open-vocabulary object detection method built upon Frozen Vision and Language Models. F-VLM simplifies the current multi-stage training pipeline by eliminating the need for knowledge distillation or detection-tailored pretraining. Surprisingly, we observe that a frozen VLM: 1) retains the locality-sensitive features necessary for detection, and 2) is a strong region classifier. We finetune only the detector head and combine the detector and VLM outputs for each region at inference time. F-VLM shows compelling scaling behavior and achieves +6.5 mask AP improvement over the previous state of the art on novel categories of LVIS open-vocabulary detection benchmark. In addition, we demonstrate very competitive results on COCO open-vocabulary detection benchmark and cross-dataset transfer detection, in addition to significant training speed-up and compute savings. Code will be released at the https://sites.google.com/view/f-vlm/home
Forward citations
Cited by 6 Pith papers
-
Visual Textualization for Image Prompted Object Detection
Visual textualization projects support images into the text feature space and prompts an unmodified OVLM, achieving strong few-shot and open-set detection results.
-
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
Detection Prompt Optimization (DetPO) improves few-shot object detection with black-box MLLMs by iteratively refining text prompts from TP/FP/FN errors on few-shot examples, gaining up to 9.7 mAP over prior black-box methods.
-
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
EP-VLM uses event-camera motion data to sparsify image patches before a vision-language model processes them, cutting FLOPs by about half with a small accuracy drop.
-
RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
A dual-loop VLM plus VLA system uses visual prompts and closed-loop monitoring to perform chemistry lab manipulations with reported gains in success and safety-compliance over baseline robot policies.
-
Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
An online EM algorithm fits class-conditional Gaussians to the test stream from CLIP text-embedding initializations, improving test-time adaptation accuracy over prior methods on 15 benchmarks.
-
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.
Discussion (0). Sign in to comment.