REVIEW 3 cited by
GLIPv2: Unifying Localization and Vision-Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies localization pre-training and Vision-Language Pre-training (VLP) with three pre-training tasks: phrase grounding as a VL reformulation of the detection task, region-word contrastive learning as a novel region-word level contrastive learning task, and the masked language modeling. This unification not only simplifies the previous multi-stage VLP procedure but also achieves mutual benefits between localization and understanding tasks. Experimental results show that a single GLIPv2 model (all model weights are shared) achieves near SoTA performance on various localization and understanding tasks. The model also shows (1) strong zero-shot and few-shot adaption performance on open-vocabulary object detection tasks and (2) superior grounding capability on VL understanding tasks. Code will be released at https://github.com/microsoft/GLIP.
Forward citations
Cited by 3 Pith papers
-
Hi-TOPS: Hierarchical Topology-aware Scoring Prior for 3D Part Decomposition
Hi-TOPS, a training-free meso-scale Flow-Freeze prior with TSDF-guided superquadric fitting, achieves competitive 3D part decomposition and the best mIoU on PartNet (55.87).
-
Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
Textual Inversion is applied to open-vocabulary object detectors to learn a few new tokens while keeping the VLM frozen and preserving zero-shot abilities.
-
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
AnySynth is a single synthetic-data pipeline that produces layouts, images, and annotations for multiple vision tasks, and its data improves benchmark scores by small but consistent margins.
Discussion (0). Continue with ORCID to comment.