Pith. sign in

REVIEW 6 cited by

F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15639 v2 pith:KONG6QI2 submitted 2022-09-30 cs.CV

classification cs.CV
keywords detectionf-vlmopen-vocabularyfrozenadditionbenchmarkdetectorlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present F-VLM, a simple open-vocabulary object detection method built upon Frozen Vision and Language Models. F-VLM simplifies the current multi-stage training pipeline by eliminating the need for knowledge distillation or detection-tailored pretraining. Surprisingly, we observe that a frozen VLM: 1) retains the locality-sensitive features necessary for detection, and 2) is a strong region classifier. We finetune only the detector head and combine the detector and VLM outputs for each region at inference time. F-VLM shows compelling scaling behavior and achieves +6.5 mask AP improvement over the previous state of the art on novel categories of LVIS open-vocabulary detection benchmark. In addition, we demonstrate very competitive results on COCO open-vocabulary detection benchmark and cross-dataset transfer detection, in addition to significant training speed-up and compute savings. Code will be released at the https://sites.google.com/view/f-vlm/home

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Visual Textualization for Image Prompted Object Detection

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Visual textualization projects support images into the text feature space and prompts an unmodified OVLM, achieving strong few-shot and open-set detection results.

  2. DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Detection Prompt Optimization (DetPO) improves few-shot object detection with black-box MLLMs by iteratively refining text prompts from TP/FP/FN errors on few-shot examples, gaining up to 9.7 mAP over prior black-box methods.

  3. Event-Priori-Based Vision-Language Model for Efficient Visual Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EP-VLM uses event-camera motion data to sparsify image patches before a vision-language model processes them, cutting FLOPs by about half with a small accuracy drop.

  4. RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A dual-loop VLM plus VLA system uses visual prompts and closed-loop monitoring to perform chemistry lab manipulations with reported gains in success and safety-compliance over baseline robot policies.

  5. Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An online EM algorithm fits class-conditional Gaussians to the test stream from CLIP text-embedding initializations, improving test-time adaptation accuracy over prior methods on 15 benchmarks.

  6. MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.

Pith tools