Pith. sign in

REVIEW 4 cited by

HyperCLIP: Adapting Vision-Language models with Hypernetworks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.16777 v1 pith:65OEXJD7 submitted 2024-12-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords encoderimagemodelshyperclipvision-languagehypernetworktexttrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised vision-language models trained with contrastive objectives form the basis of current state-of-the-art methods in AI vision tasks. The success of these models is a direct consequence of the huge web-scale datasets used to train them, but they require correspondingly large vision components to properly learn powerful and general representations from such a broad data domain. This poses a challenge for deploying large vision-language models, especially in resource-constrained environments. To address this, we propose an alternate vision-language architecture, called HyperCLIP, that uses a small image encoder along with a hypernetwork that dynamically adapts image encoder weights to each new set of text inputs. All three components of the model (hypernetwork, image encoder, and text encoder) are pre-trained jointly end-to-end, and with a trained HyperCLIP model, we can generate new zero-shot deployment-friendly image classifiers for any task with a single forward pass through the text encoder and hypernetwork. HyperCLIP increases the zero-shot accuracy of SigLIP trained models with small image encoders by up to 3% on ImageNet and 5% on CIFAR-100 with minimal training throughput overhead.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuARI: Query Adaptive Retrieval Improvement

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A hypernetwork predicts a query-specific low-rank linear projection that reshapes frozen VLM embeddings, improving retrieval on ILIAS and INQUIRE.

  2. Parameter-Dynamic Adaptive Fusion and Calibration Network for RGBT Tracking

    cs.CV 2026-08 conditional novelty 6.0 of 10

    PAFCNet dynamically generates target-conditioned parameters for multimodal fusion and spatio-temporal calibration in RGBT tracking, achieving competitive benchmark results.

  3. WeightCLIP: Aligning Datasets and Models for Weight Space Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Contrastive dataset–weight alignment reshapes weight-space latents so dataset prompts retrieve, generate, and refine neural nets better than prior weight-space methods.

  4. (Almost) Free Modality Stitching of Foundation Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A hypernetwork that generates connector weights for all image-text model pairs can rank pairs like grid search at about 10x lower training cost, but the best connector lags grid search by a few points.

Pith tools