REVIEW 5 cited by
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https://github.com/iSEE-Laboratory/LLMDet.
Forward citations
Cited by 5 Pith papers
-
VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
Fusing open-set and open-ended query sets raises LVIS detection accuracy, with the largest gains on rare categories, and lets one model run in either mode.
-
MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts
Ten leading VLMs mostly fail to report removed essential object parts as missing, and simulated detector evidence, image tools, longer reasoning, and an easier fine-tune barely improve accuracy.
-
Reinforced Visual Perception with Tools
ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.
-
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
RAGSR combines region-level vision-language captions with regional attention masks to improve fine-grained detail generation in diffusion-based super-resolution.
-
Kitchen Robotic Manipulation utilizing Foundation Models
A modular perception pipeline using off-the-shelf foundation models achieves 89.12% ADI on a custom kitchen dishware dataset and performs real robot sink-to-dishwasher and cup-stacking tasks without retraining.
Discussion (0). Continue with ORCID to comment.