Pith. sign in

REVIEW 5 cited by

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.16334 v1 pith:IRDBZJGH submitted 2024-12-20 cs.CV

classification cs.CV
keywords tasksself-supervisedtextvisualalignmentclipdensedinov2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual features are not readily aligned with language, hindering their adoption in open-vocabulary tasks. Our method, named dino.txt, unlocks this new ability for DINOv2, a widely used self-supervised visual encoder. We build upon the LiT training strategy, which trains a text encoder to align with a frozen vision model but leads to unsatisfactory results on dense tasks. We propose several key ingredients to improve performance on both global and dense tasks, such as concatenating the [CLS] token with the patch average to train the alignment and curating data using both text and image modalities. With these, we successfully train a CLIP-like model with only a fraction of the computational cost compared to CLIP while achieving state-of-the-art results in zero-shot classification and open-vocabulary semantic segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    DinoSplat-OV does training-free open-vocabulary segmentation on remote sensing images by combining DINOv3 features with Laplacian smoothing and Gaussian-splatting upsampling, reaching 42.9 mIoU on UDD5 and 37.5 averag...

  2. Twins: Learn to Predict Unified Representations with Focal Loss

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.

  3. Generalizable Object Re-Identification via Visual In-Context Prompting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VICP uses an LLM to generate per-category visual prompts for a frozen DINOv2, enabling few-shot generalization to unseen object categories in re-identification without parameter updates.

  4. Automated data curation for self-supervised learning in underwater acoustic analysis

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A data-curation pipeline combining AIS ship data with hierarchical k-means clustering improves self-supervised learning for underwater acoustic ship classification.

  5. Exploring Image-Text Alignment for Radio Galaxy Morphologies

    astro-ph.IM 2026-07 conditional novelty 5.0 of 10

    Text captions of radio galaxy images can classify FR-I vs FR-II morphologies comparably to image embeddings, but LoRA fine-tuning improves local class coherence without improving global image-text alignment.

Pith tools