Pith. sign in

REVIEW 3 cited by

GLIPv2: Unifying Localization and Vision-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.05836 v2 pith:GXNI7QLD submitted 2022-06-12 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords tasksunderstandinglocalizationglipv2modeldetectionpre-trainingvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies localization pre-training and Vision-Language Pre-training (VLP) with three pre-training tasks: phrase grounding as a VL reformulation of the detection task, region-word contrastive learning as a novel region-word level contrastive learning task, and the masked language modeling. This unification not only simplifies the previous multi-stage VLP procedure but also achieves mutual benefits between localization and understanding tasks. Experimental results show that a single GLIPv2 model (all model weights are shared) achieves near SoTA performance on various localization and understanding tasks. The model also shows (1) strong zero-shot and few-shot adaption performance on open-vocabulary object detection tasks and (2) superior grounding capability on VL understanding tasks. Code will be released at https://github.com/microsoft/GLIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-TOPS: Hierarchical Topology-aware Scoring Prior for 3D Part Decomposition

    cs.GR 2026-08 conditional novelty 6.0 of 10

    Hi-TOPS, a training-free meso-scale Flow-Freeze prior with TSDF-guided superquadric fitting, achieves competitive 3D part decomposition and the best mIoU on PartNet (55.87).

  2. Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Textual Inversion is applied to open-vocabulary object detectors to learn a few new tokens while keeping the VLM frozen and preserving zero-shot abilities.

  3. AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

    cs.CV 2024-11 conditional novelty 5.0 of 10

    AnySynth is a single synthetic-data pipeline that produces layouts, images, and annotations for multiple vision tasks, and its data improves benchmark scores by small but consistent margins.

Pith tools