Pith. sign in

REVIEW 5 cited by

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.18954 v1 pith:PRH3LIQU submitted 2025-01-31 cs.CV

classification cs.CV
keywords largemodelopen-vocabularylanguagellmdetcaptionsdatasetdetector
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https://github.com/iSEE-Laboratory/LLMDet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fusing open-set and open-ended query sets raises LVIS detection accuracy, with the largest gains on rare categories, and lets one model run in either mode.

  2. MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Ten leading VLMs mostly fail to report removed essential object parts as missing, and simulated detector evidence, image tools, longer reasoning, and an easier fine-tune barely improve accuracy.

  3. Reinforced Visual Perception with Tools

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.

  4. RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution

    cs.CV 2025-08 conditional novelty 5.0 of 10

    RAGSR combines region-level vision-language captions with regional attention masks to improve fine-grained detail generation in diffusion-based super-resolution.

  5. Kitchen Robotic Manipulation utilizing Foundation Models

    cs.RO 2026-08 conditional novelty 4.0 of 10

    A modular perception pipeline using off-the-shelf foundation models achieves 89.12% ADI on a custom kitchen dishware dataset and performs real robot sink-to-dishwasher and cup-stacking tasks without retraining.

Pith tools