Pith. sign in

Paper Citation Record · LEDGER

Adapting Vision-Language Models Without Labels: A Comprehensive Survey

As of 8 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 4 inbound Pith citation observations for arXiv:2508.05547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05547 v1

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:17:06.729087Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T23:10:08.033785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T12:54:40.377437Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7b6bc77-500a-472d-865f-2f872bda4b73 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Learning transferable visual models from natural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.411463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.411463Z digest=sha256:fa7867090172e5dd6f5e01afc1bb1dc5694f4b4a369c7af49e1997060e520f3e

Observation b31b2cf6-d9fc-4828-b62b-048077f931a8 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.440245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.440245Z digest=sha256:4566093317a682f81b50bf38cf1abb90c5b7d691150d5a1db6db98d307a50d99

Observation c25732c9-0c30-4640-bae1-4ef40b5f49f0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.506315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.506315Z digest=sha256:b9fff3f0f22015b5f4f133e83f841e8b271a325958543b06fa861f32f92d5d42

Observation a6d4c706-f83b-4876-9841-aa92642aeb6c · outbound

This paper cites Visual instruction tuning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Visual instruction tuning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.573874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.573874Z digest=sha256:878dbb7af7188bbf3465f3f703a424ce6ac3a08dfbcec9cb76125b651ccd334d

Observation 871f2069-6a30-41ea-8e04-b01f027af56a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.616036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.616036Z digest=sha256:c6483838be8ea865771fbd3aa3263ae9861bfea973c9ef7671bf2f69716c97d7

Observation f656fd03-df30-4533-8938-90be9d1ffa37 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.674684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.674684Z digest=sha256:2c74f80f4dc36e229b562f723c25756d0cc43bff67b5c17c3be8004602d11e6f

Observation 3beeec1f-bdac-4ce2-877c-6bee9c981e19 · outbound

This paper cites CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.779796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.779796Z digest=sha256:365b64cf8393f53ba8f1c753c20e7985d8fcaaf2265a1d14a03a2173d3c387ee

Observation a2af32f6-8260-4f30-8822-38d51e270b87 · outbound

This paper cites Unseen visual anomaly generation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Unseen visual anomaly generation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.831119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.831119Z digest=sha256:d22263b214a1ee7647c03eca94bf37754ee1cfefe60184e589e9ab07dc1a319d

Observation fd09b1c1-6566-4185-b79b-c82440621154 · outbound

This paper cites Probabilistic embeddings for cross-modal retrieval,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Probabilistic embeddings for cross-modal retrieval,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.910821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.910821Z digest=sha256:180e971bbd86d90b73997bc8188fc5cfdd5328cf7e605553367aa9683a02f34d

Observation ba6eecd3-805d-46a9-87f9-170a1bbb4d0b · outbound

This paper cites Learning to prompt for vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Learning to prompt for vision-language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:03.951281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:03.951281Z digest=sha256:43b218dbb1512abdbd6bf92292d09f04bc9583352db36a157e7b93a0182eb08c

Observation a204632c-839e-4b13-b0f0-d6ad72687abb · outbound

This paper cites Conditional prompt learning for vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Conditional prompt learning for vision-language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.008476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.008476Z digest=sha256:67bca8abb4d8bc3c10b579ac322e576112746a722ee9d070948fb9d9b75bdad4

Observation dc17dc4d-4f79-4b3b-be30-068c5e9da5a6 · outbound

This paper cites MaPLe: Multi-modal Prompt Learning.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey MaPLe: Multi-modal Prompt Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.069590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.069590Z digest=sha256:4a6b4a7afb355ee586a587f205bd14996aeb1b5fea6a39c8dc7c2c71b0cfa5b9

Observation 37b296ab-d061-46b0-af81-05101b73d4aa · outbound

This paper cites A hard-to- beat baseline for training-free clip-based adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A hard-to- beat baseline for training-free clip-based adaptation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.125513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.125513Z digest=sha256:ed7e26e3901bcfbb588fd687d79ff289cffdcb5dedbe8e15ddfb2771dd47f546

Observation c7d5e9f4-b8b5-48a1-b6f1-03b734620888 · outbound

This paper cites Adapting visual category models to new domains,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Adapting visual category models to new domains,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.180514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.180514Z digest=sha256:528b7ee788ca5b917e406556b962ed5705ce7dc8108e1078456b9b3ff919ea17

Observation ae56fbcf-d5fb-4397-b0b8-c5b88ba57c63 · outbound

This paper cites Visual classification via description from large language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Visual classification via description from large language models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.240023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.240023Z digest=sha256:b05317bf971ead5ee777b3dad3b513d1f59d1b5fff2ba1c324fd6a772697983c

Observation 43569026-8929-4db7-a916-50f67bb0c716 · outbound

This paper cites Extract free dense labels from clip,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Extract free dense labels from clip,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.287319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.287319Z digest=sha256:4f069d128ad72389744676ac32dec67507a158adafaef5a8d7aa543be7ad0e29

Observation c3942aff-a4f6-4c99-945b-09a255845dd3 · outbound

This paper cites Unsupervised Prompt Learning for Vision-Language Models.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Unsupervised Prompt Learning for Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.343681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.343681Z digest=sha256:5a9233799f97f2e0061e56cad2ceb09d7a67ff467161504c9cdb48e54cb76a42

Observation 68b5ba6b-c2c9-496f-b81f-0d203c35a7ca · outbound

This paper cites Test-time prompt tuning for zero-shot generalization in vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Test-time prompt tuning for zero-shot generalization in vision-language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.396702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.396702Z digest=sha256:6a6bb0e595a306c6cf1ad1412bac1ade280aa1372fe30f8c2b73e168886496ff

Observation 788dd14e-fd13-4fbc-ad9d-ba36c89c061c · outbound

This paper cites Swapprompt: Test-time prompt adaptation for vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Swapprompt: Test-time prompt adaptation for vision-language models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.466211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.466211Z digest=sha256:1cbf9a97979e7b7c78ce49b0f427cb80b77f178a36f0d1d8b4f011b77eac6700

Observation c3c3c35f-78cd-46df-b823-32edca0210c0 · outbound

This paper cites The illusion of progress? a critical look at test-time adaptation for vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey The illusion of progress? a critical look at test-time adaptation for vision-language models,

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-05T23:17:15.006412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:17:04.506791Z digest=sha256:97d36a4af3870c7bbc24ed9b0b36bcb3cbb25d8a36e5fb901d07233e9e49a2bd

Observation e124257a-b3a7-41f6-8f27-cb4a0b3476b3 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classi- fication,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey What does a platypus look like? generating customized prompts for zero-shot image classi- fication,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.542044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.542044Z digest=sha256:3c1065a43f45a80ff0ded93d94599bf917b28ebd661689f692286a9ab5f1491c

Observation 7fca3fa5-e0cd-4f0b-a89b-0fdc766993b2 · outbound

This paper cites Improving Zero-Shot Models with Label Distribution Priors.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Improving Zero-Shot Models with Label Distribution Priors

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:17:14.928712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:17:04.589997Z digest=sha256:3b2198e42aa3ac256c6f45ec3de5f987727011229b78ad36bc8532cff446ab34

Observation 447f5e3e-7f2c-4d92-aa47-da94c1dec0c3 · outbound

This paper cites Online zero-shot classification with clip,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Online zero-shot classification with clip,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.649384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.649384Z digest=sha256:177a8fb391c68ed87ebadcbe20b5b6d548eec3a0a735e4fbe25da93c3555f023

Observation 0d32d8bf-997f-4b63-b1ed-f92733283f9a · outbound

This paper cites Align your prompts: Test- time prompting with distribution alignment for zero-shot generaliza- tion,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Align your prompts: Test- time prompting with distribution alignment for zero-shot generaliza- tion,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.682395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.682395Z digest=sha256:f708d375cb2d3c98ae57e59e7ce2f23ca370ab13369adab9f1110df4d3b940cd

Observation 40ceada1-b4e4-47a8-9dcc-76ad2c947cd3 · outbound

This paper cites Efficient test-time adaptation of vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Efficient test-time adaptation of vision-language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.745293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.745293Z digest=sha256:e6e9a037943cff54de23d1ba0882a083fe784ac2fbeb4956177c34f26e6b5205

Observation d26b76ca-92e6-460b-a75d-33fff16c2494 · outbound

This paper cites Masked unsupervised self-training for label-free image classification,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Masked unsupervised self-training for label-free image classification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.829357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.829357Z digest=sha256:c0ac59d034e29e599c7958ddeaed7530d89184f667cb3a586161bea2e7d32927

Observation 2588556e-e9e0-440e-b2f1-903b9e85aba7 · outbound

This paper cites Realistic unsupervised clip fine-tuning with universal entropy optimization,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Realistic unsupervised clip fine-tuning with universal entropy optimization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.889601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.889601Z digest=sha256:b6bc4a2cf4c2591f7735df475427a17b05213e4c65a4f8cc2700cf46254f3d37

Observation e42f5036-091d-4e86-9a67-2fb5ac166d2d · outbound

This paper cites Reco: Retrieve and co-segment for zero-shot transfer,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Reco: Retrieve and co-segment for zero-shot transfer,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.931657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.931657Z digest=sha256:34bbc04051963bf0cf7cf9681da221438d485751f4466f07b5508a011c051b68

Observation ea382039-b3dc-40c6-96ab-0c3079b25544 · outbound

This paper cites Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:04.977067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:04.977067Z digest=sha256:4f6f21a0105261e79d50f48358b1ef53c69022749312ed2da26281e7d1d50ee9

Observation 9386ebcb-8a65-4608-9c31-ec0d63031710 · outbound

This paper cites Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Can language-guided unsupervised adaptation improve medical image classification using unpaired images and texts?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.046092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.046092Z digest=sha256:76d07a007558177023d724b53a8ac301b755464dc44d711f2db60b92f89c4ab0

Observation 0b96e069-5ce9-46fb-ba23-618196242cbd · outbound

This paper cites A chatgpt aided explainable framework for zero-shot medical image diagnosis,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A chatgpt aided explainable framework for zero-shot medical image diagnosis,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.070471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.070471Z digest=sha256:e87248729e92c37712302d100aeb4e5be6241b9a8a91163283ba2b7d830327da

Observation e432cc70-4a70-4a78-834c-ee7a7e2fb61c · outbound

This paper cites Text-enhanced zero-shot action recognition: A training-free approach,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Text-enhanced zero-shot action recognition: A training-free approach,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.138131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.138131Z digest=sha256:a080d691f332ac8f4a69df21e12f9236dd71604dd7872a13c14ef06c554af615

Observation efb7bc96-b047-4ae8-ab73-df044ce405da · outbound

This paper cites Dts-tpt: dual temporal-sync test-time prompt tuning for zero-shot activity recogni- tion,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Dts-tpt: dual temporal-sync test-time prompt tuning for zero-shot activity recogni- tion,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.189370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.189370Z digest=sha256:ee14c3ec299272f27a4054e71ff983304bf04c8a0c75be7098c9f518eec4adc2

Observation 2d5e7e2c-9e54-4fd9-9876-a1e3ec263850 · outbound

This paper cites Pouf: Prompt- oriented unsupervised fine-tuning for large pre-trained models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Pouf: Prompt- oriented unsupervised fine-tuning for large pre-trained models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.253496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.253496Z digest=sha256:9061546bf2ecd3ec330e08e659a2970bd996f50a54094c4033249d011d7745cc

Observation a4d550c2-8289-4794-a8c9-5bea88691405 · outbound

This paper cites Label propagation for zero-shot classification with vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Label propagation for zero-shot classification with vision-language models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.308737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.308737Z digest=sha256:dbc54f3d93b1d99e24f244d36426d18b6288ef83e7bc595f6d5f7e37fb69d571

Observation 2bc59056-7b73-454a-8dd2-bed00cc3b9be · outbound

This paper cites On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.371056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.371056Z digest=sha256:97ec96af1821c8e54db9247785a5375b29dddfaca025fac99a58ae6095145b5c

Observation ea13eed2-730c-4441-91de-ef40aadca1f7 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Vision-language models for vision tasks: A survey,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.461471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.461471Z digest=sha256:2856b2f1d6f91ea76cfefb6a47416b8c0cc1810ec8e494f0a4fd896e5bd32886

Observation e5790b02-7f2d-4dad-a3f2-33261cc94dfd · outbound

This paper cites Advances in multimodal adaptation and generalization: From traditional approaches to foundation models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Advances in multimodal adaptation and generalization: From traditional approaches to foundation models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.538275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.538275Z digest=sha256:d571ff42764339500355aa51a398e107adbf3dd700723147ec498920ffe89b21

Observation f8531e7b-6bdf-4510-b747-0e5a4a7f87ed · outbound

This paper cites Generalizing vision-language models to novel domains: A comprehensive survey.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Generalizing vision-language models to novel domains: A comprehensive survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.585412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.585412Z digest=sha256:4ad658a8141a3f9318de45c091dd6a9921df7b64c6733c3ca9299e0a638b7ecc

Observation bf69ec72-773f-48b1-8cca-2da2df8b104c · outbound

This paper cites A comprehensive survey on test-time adaptation under distribution shifts,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A comprehensive survey on test-time adaptation under distribution shifts,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.671452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.671452Z digest=sha256:baf1386bc9307f4a8909d26f312f2db78620f6dafd67c18e7b15bf60e5e717a5

Observation 842cf4bc-c21d-4531-8f59-f062c69b0885 · outbound

This paper cites In search of lost online test-time adaptation: A survey,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey In search of lost online test-time adaptation: A survey,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.713046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.713046Z digest=sha256:339466dce6d145d1a0160ca6171ff769632978a76dcc0d66bef37d79896decfa

Observation d0d3fa61-1766-4e2b-9cae-d689705ac2fb · outbound

This paper cites Beyond Model Adaptation at Test Time: A Survey.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Beyond Model Adaptation at Test Time: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.771665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.771665Z digest=sha256:db8f1707e4936dc4c7edf34b85fa38786ee4c9ee831138ce194c92e7a6d93f2c

Observation e716f882-24dc-4ea4-b2f4-4fc56b7832f5 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.859329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.859329Z digest=sha256:0f8d2df3ec031375a59c832bed87915b1d95c004438cbc648bbe36df0d4b221e

Observation c30dc607-dab7-4d22-bae1-125ce99ae1eb · outbound

This paper cites Attention is all you need,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Attention is all you need,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.910520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.910520Z digest=sha256:baa1863000ebcf1bfcbf1db4002eb0bcc6c87d4cab7b6326810a2f62d886f69b

Observation 200d380b-0fc0-4d23-91a2-4c94b932b669 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.951097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.951097Z digest=sha256:27baec9205a0754743c9ce626aeb87fcd82a9e3059dcfd88d18840341ad8ed73

Observation aaf5043a-1208-4a05-9244-460064c6cb40 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.035332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.035332Z digest=sha256:41c9799bd97fb243a19dd0f1999709320eb52ee03c38922f361f41fb073966e2

Observation 57e4bcf7-a40f-4541-86cc-45e50f583cc9 · outbound

This paper cites Image captioning with semantic attention,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Image captioning with semantic attention,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.095913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.095913Z digest=sha256:dd64326b6d41b9950fca20691a4d39f53b76555ba711c2a4539548f05f512446

Observation ee405704-3202-4be4-a1ab-e3073d8f9d58 · outbound

This paper cites Visual question answering: A survey of methods and datasets,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Visual question answering: A survey of methods and datasets,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.198250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.198250Z digest=sha256:94c908ec38d642bad0c8b5ec0fd90eb7c01904d4bd707dad9944b8b9fc8ff1ab

Observation b11b603b-f8e3-497e-be26-5e4aa4056122 · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey High- resolution image synthesis with latent diffusion models,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.276256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.276256Z digest=sha256:d33a4136e8b1126ee9b05f3d09d3b4f0492e0daa4141ade918c18c225f518c8f

Observation 02359158-71fa-4a62-9dd2-f003c7892612 · outbound

This paper cites Deep supervised cross-modal retrieval,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Deep supervised cross-modal retrieval,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.317978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.317978Z digest=sha256:60800566c9ffb4dd065cafd8a3784f8434121ce5fbbeeaa0fc2859c120d779f9

Observation 83f1a346-24ac-4557-93a1-f3d22b65a286 · outbound

This paper cites A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.384111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.384111Z digest=sha256:7f5486a38c07f21c9a3bb2ff392cd42f1bc54172a6be7ce569b7df3ffbc1c2b1

Observation 61ef6ba4-54e1-479b-b3c1-e5a66aef93b7 · outbound

This paper cites Learning to detect unseen object classes by between-class attribute transfer,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Learning to detect unseen object classes by between-class attribute transfer,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.445225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.445225Z digest=sha256:7a51018d6565038b8ca606d08aeb6fe47816ff5dd8e70f3b18377c343617f820

Observation c38ba418-de6d-4f35-9acc-7332b1c1c721 · outbound

This paper cites Describing objects by their attributes,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Describing objects by their attributes,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.511582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.511582Z digest=sha256:2be1b1335ae28139423e45baaf4a171e3d1cdbbc73ea403f15d80de0a984ed92

Observation 7df03e6e-1b56-43ad-8a7d-7a52637ed111 · outbound

This paper cites Devise: A deep visual-semantic embedding model,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Devise: A deep visual-semantic embedding model,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.519380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.519380Z digest=sha256:dec2fa5fb1feda127370241e4e9e90f42c43592c2acdd6044c89b529a3aa2307

Observation 3536aa89-4f5e-4bed-9844-69467245b3d1 · outbound

This paper cites An embarrassingly simple approach to zero-shot learning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey An embarrassingly simple approach to zero-shot learning,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.524000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.524000Z digest=sha256:5574b79a32a61e10569d640e699eb605e41de72ebc2e986b3d34a48c07e7b6af

Observation bd00ec66-9ca8-4ca5-bea0-5d1ad47cabce · outbound

This paper cites Generalized zero-shot learning via synthesized examples,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Generalized zero-shot learning via synthesized examples,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.528338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.528338Z digest=sha256:5340a0a6fb9cf22ac29b4db20137bef9152fb289f2b73e6802a83f4f9ff5b633

Observation c293621c-f38d-4d18-af71-7364ada6ef61 · outbound

This paper cites Multi-modal cycle-consistent generalized zero-shot learning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Multi-modal cycle-consistent generalized zero-shot learning,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.532571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.532571Z digest=sha256:ef9137071827843eea39c337ec5f8712c83069a01ae247464071f3e00b1ae2da

Observation fd4c9846-139c-474a-beec-29711a63aee4 · outbound

This paper cites f-vaegan-d2: A feature generating framework for any-shot learning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey f-vaegan-d2: A feature generating framework for any-shot learning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.537074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.537074Z digest=sha256:908c77feddc6c1be48674aea65fac67e01c8c61beb6d04417424000446154935

Observation c72e388f-fee6-4710-90e6-e6bbd9715452 · outbound

This paper cites An empirical study and analysis of generalized zero-shot learning for object recognition in the wild,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey An empirical study and analysis of generalized zero-shot learning for object recognition in the wild,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.541736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.541736Z digest=sha256:93bf3d9c807f1a4d7344b03f0780de3954cf20baa16717c14cfa76d7034ece83

Observation e9ef69b8-6ff8-43b3-8a42-d537f92582ce · outbound

This paper cites Zero-shot learning-the good, the bad and the ugly,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Zero-shot learning-the good, the bad and the ugly,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.546148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.546148Z digest=sha256:2a05fc9d07e314b68e16a61f54f789146821e6a2fbd2a85858925eabcc8f37f0

Observation 0e9509dd-9864-47a8-aba1-19cbb38e52f9 · outbound

This paper cites A review of generalized zero-shot learning meth- ods,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A review of generalized zero-shot learning meth- ods,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.550354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.550354Z digest=sha256:cc69bb2ed9b4b36c158f5ae05586b2c63f133266afe6e1f2d5be9a88c93d7e4f

Observation 9f0d1a5f-f402-4753-9031-61e150c492d7 · outbound

This paper cites A survey of zero-shot learning: Settings, methods, and applications,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A survey of zero-shot learning: Settings, methods, and applications,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.554621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.554621Z digest=sha256:8a20a333db3a6e05c5f2e496b5c2ab95ee5fe2b8cfdb9719a1308a30d914a874

Observation cedb5ca5-4f96-49a2-800f-3bbf385f578d · outbound

This paper cites Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:17:14.764669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:17:06.559119Z digest=sha256:e8fa460160c4d25d1141ff5b379e68b2080ef697687b3cc6b78560842db14c17

Observation 295c011c-078c-4ccc-a1bf-3ebd20648b9e · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Clip-adapter: Better vision-language models with feature adapters,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.563682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.563682Z digest=sha256:5e7f424a5dc03a8cdce06f502da0b2f77eaccc021473ebe27d85bf895e258be2

Observation 6bdd2d81-edc0-4840-8603-264b06a8fbad · outbound

This paper cites Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.568562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.568562Z digest=sha256:bce99415dfcd22afc0edb8f38a34c6ad28fee01ccb55a08a9f3dcfef1d04108c

Observation 508ca37d-ec6c-41f9-9955-d8a8c376ed91 · outbound

This paper cites Low-rank few-shot adaptation of vision- language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Low-rank few-shot adaptation of vision- language models,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.572790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.572790Z digest=sha256:a41acc1262fff8bcdd65dcb34798ab89a9112c5dcfa8cca8aec01d4c6cbc6362

Observation be2961b0-4a51-4b47-995f-fc2953daa493 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Visual Classification via Description from Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.577174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.577174Z digest=sha256:d6b00ace7424f0f8c1d954370e5e5a71c1d362e2adf210ebca50b9ba13621881

Observation 647103fe-dd31-47ff-ae31-feed2951415b · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Denseclip: Language-guided dense prediction with context- aware prompting,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.581490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.581490Z digest=sha256:a1212289b20a6395cb5168aec637d60674b2b3ce5b404d48745fb3600c81ce89

Observation a0d3e291-6774-4c4b-b33c-5f3ddc703e08 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.585528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.585528Z digest=sha256:66fcfb68a1ebda7e64fa377d342bf09c8191120f871f090b1b8af15b11bf7921

Observation fb8af6a8-7d5e-4f5e-a6b1-9d43331e4572 · outbound

This paper cites Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:17:14.695358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:17:06.589956Z digest=sha256:25f2904317f5d6c9c2188d959c5479dd6ffcc56d6f8a9e512a861fa2423a402b

Observation 19322e0b-ce33-4da2-b424-77e88498a2e5 · outbound

This paper cites Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.594580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.594580Z digest=sha256:3b088af6728d1c66f507a332a0c0b63fffbef81c824a60f886aaaa4ed018406b

Observation 8e7bf716-5549-4aad-9d98-634e953e60d6 · outbound

This paper cites Exploiting local feature patterns for unsupervised domain adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Exploiting local feature patterns for unsupervised domain adaptation,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.598787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.598787Z digest=sha256:2b2167fab66761e035bb725dbd1a475a629a80f138928f62c5185392442e58a5

Observation ce03de49-3ecc-4add-90dd-634f432ef179 · outbound

This paper cites Contrastive test-time adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Contrastive test-time adaptation,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.602801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.602801Z digest=sha256:8a2a357cc32ed63272f2519d3f0db74c259ab67f8072b4ad995c2fa469e2e52f

Observation d8d98dbb-c695-44bf-afd4-c9a410ea73b1 · outbound

This paper cites Universal source-free domain adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Universal source-free domain adaptation,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.607476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.607476Z digest=sha256:d53f7bdd92da5845103769a2e79c40285b7aca61076c9191c8ded01c2289ca5d

Observation 338b0a1c-c833-4f03-aea6-808e47f1a14f · outbound

This paper cites Model adaptation: Historical contrastive learning for unsupervised domain adaptation without source data,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Model adaptation: Historical contrastive learning for unsupervised domain adaptation without source data,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.611750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.611750Z digest=sha256:73f46046bdb872fee066cf9b3c14a692f026992dfe07c37358c877f881bc0219

Observation 1861f6fd-2daa-4cf2-9940-09ceecbc5bc9 · outbound

This paper cites Domain adaptation without source data,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Domain adaptation without source data,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.615871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.615871Z digest=sha256:ee6aa0d9fa3972d298725803906f5e58d873693715af4e5fb23d5390d8e7baa1

Observation 7fa2327c-4e39-4d95-85d2-6db255df0849 · outbound

This paper cites A comprehensive survey on source-free domain adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A comprehensive survey on source-free domain adaptation,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.620242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.620242Z digest=sha256:872c13ceb860363b518874b1403d98fcbfc274812c2e3ce45168b1e394db1c02

Observation f4ce1167-0449-48b6-9528-9dba3aff35d9 · outbound

This paper cites Source-free unsu- pervised domain adaptation: A survey,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Source-free unsu- pervised domain adaptation: A survey,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.624666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.624666Z digest=sha256:1d4b5f2a62c5105f1cbc37ff729a4e6f09fbfef9689610491ec3df9b4660e53c

Observation 5b0de065-0440-4bfc-9d68-f2fc31c4619d · outbound

This paper cites Tent: Fully test-time adaptation by entropy minimization,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Tent: Fully test-time adaptation by entropy minimization,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.628996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.628996Z digest=sha256:cceadff728ef28557c016b4e835bf1b0de6d9a63c15e49ddeeb3d31b1f914a5f

Observation df703f37-0672-45f7-9e34-c28f0e8d0c7a · outbound

This paper cites Towards robust multimodal open-set test-time adaptation via adaptive entropy-aware optimization,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Towards robust multimodal open-set test-time adaptation via adaptive entropy-aware optimization,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.633652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.633652Z digest=sha256:27abf748e59fadbc09c759a6a5987ae404783207f40740a993cfc5a76cb4ed83

Observation 76497343-3cc5-4a9a-be0f-0ba739fcdaae · outbound

This paper cites Efficient test-time model adaptation without forgetting,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Efficient test-time model adaptation without forgetting,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.638132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.638132Z digest=sha256:23b2c6dab23df48b48f152418bb163df6a509c0fb832f55456c997c2164c4c5e

Observation dc774672-c694-45f1-a379-0a41fb77b784 · outbound

This paper cites Sotta: Robust test-time adaptation on noisy data streams,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Sotta: Robust test-time adaptation on noisy data streams,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.642480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.642480Z digest=sha256:52a17d8862836763d34198c519e95e993d26367561b703ff0b1e95ea87160bf1

Observation 62cbaaf8-0c64-44be-8665-0a97165148f2 · outbound

This paper cites Continual test-time domain adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Continual test-time domain adaptation,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.648686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.648686Z digest=sha256:4f2563563fa894a69af82d64fe156f3922828498fd77244f8e4cae9c4c02c71e

Observation dac2b887-aefc-4861-b6b5-6dda7765fbf6 · outbound

This paper cites Ecotta: Memory- efficient continual test-time adaptation via self-distilled regularization,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Ecotta: Memory- efficient continual test-time adaptation via self-distilled regularization,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.653156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.653156Z digest=sha256:6b54ebf499589f7fc503d05ae8a8d53ec41475ff9db238099bf1b31ecf76533d

Observation e4377699-4fc5-4b85-8662-e509ee2fe5ae · outbound

This paper cites Diverse data augmentation with diffusions for effective test-time prompt tuning,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Diverse data augmentation with diffusions for effective test-time prompt tuning,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.657503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.657503Z digest=sha256:ad71eb04c54d409339e22b5dabdd49d3c6779b583c2cdd601daa4955f4ef18b4

Observation 4b059712-fd22-4477-b17a-1ad35fee2618 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.662323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.662323Z digest=sha256:e338d6b48dce5f460f302bf0cf9799c9fb63db24a4cfa8fee6a20e5e4e6ab031

Observation fde89f92-ca02-403b-80ba-a5ea5b6555bb · outbound

This paper cites Sigmoid loss for language image pre-training,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Sigmoid loss for language image pre-training,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.667461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.667461Z digest=sha256:f7260604d9c98f9f7b12a6dcaed26d9c5364e986d0b592a7fd84a78f425a52d1

Observation 1763b2da-8de3-4af7-83ce-b084a89c1661 · outbound

This paper cites Chils: Zero-shot image classification with hierarchical label sets,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Chils: Zero-shot image classification with hierarchical label sets,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.671557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.671557Z digest=sha256:f1fbbd460d297670bc7db0ff304f77d3aa5016973dc4386b801618402b41bf59

Observation fdef4187-3409-4253-91db-20dcf330a566 · outbound

This paper cites Sus-x: Training-free name- only transfer of vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Sus-x: Training-free name- only transfer of vision-language models,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.680502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.680502Z digest=sha256:66de40d90ba859be1f6061ee19b7424e587dd9a3076c8f60c43e372e0c11e0d5

Observation fe24334b-0686-42d4-b91d-db55f481f4fc · outbound

This paper cites Neural priming for sample- efficient adaptation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Neural priming for sample- efficient adaptation,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.685106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.685106Z digest=sha256:54881217dd6b24be68ef30f2f5a2b55e39b209f8d6dffa8032bc7f30fb850c4b

Observation d0e8afd2-32b2-41d1-9647-5c5193546abc · outbound

This paper cites Just say the name: Online continual learning with category names only via data generation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Just say the name: Online continual learning with category names only via data generation,

Reference 92

Resolution
verified exact
raw_fallback, observed 2026-08-05T23:17:14.657902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:17:06.689789Z digest=sha256:3fbbb3ed3e3c3da69eace1f2d34e89ffc364d3e78672ce8a1b0af0c66b1c4abc

Observation 794fb8eb-5dba-4078-a90a-b01a70774e3d · outbound

This paper cites Calip: Zero-shot enhancement of clip with parameter-free attention,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Calip: Zero-shot enhancement of clip with parameter-free attention,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.693916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.693916Z digest=sha256:6af359c0d0dda4f47280762598cc98c3483be20f1ee3f516d0641f4c54b2daf4

Observation 520db2c9-ed3c-4e35-baec-e294df52c37f · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Sclip: Rethinking self-attention for dense vision-language inference,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.698630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.698630Z digest=sha256:281aa59b40541d8a5cf94c2a7a2846555a987bd5d0463678cc2aab2b7bfe76c6

Observation dd531c91-a9a6-473a-8f67-3bc93551ea42 · outbound

This paper cites Proxyclip: Proxy attention improves clip for open-vocabulary segmentation,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Proxyclip: Proxy attention improves clip for open-vocabulary segmentation,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.702964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.702964Z digest=sha256:1b6e0af4ec23ce40cc7e3425ff0d2dc3031ecdd12b80b6c2550b9f8664ba5280

Observation 6909462e-7405-4d0d-a95f-a8a0c478f0fd · outbound

This paper cites Language models are few-shot learners,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Language models are few-shot learners,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.707303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.707303Z digest=sha256:ff274b5f07b86c3d65e4ba170f2e3ded003e6d40e25cfc0c59700693373926ad

Observation 58118cb6-2a68-4a9c-985c-e96e0734f571 · outbound

This paper cites Meta-prompting for automating zero- shot visual recognition with llms,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Meta-prompting for automating zero- shot visual recognition with llms,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.711448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.711448Z digest=sha256:f10b3312b73d7d7b19331d457239359356dd954b488417915184e20688661db9

Observation 0f584590-04b1-4ba6-841b-cf602dbaa2ac · outbound

This paper cites The neglected tails in vision-language models,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey The neglected tails in vision-language models,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.715573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.715573Z digest=sha256:836a52427bfc9884a013fb63fa0bfea92f0cd8f14ca3f48250b7c00abcfa7cd3

Observation 0a162a8a-de35-43ad-9d0c-ab0017399326 · outbound

This paper cites Introducing chatgpt,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Introducing chatgpt,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.720336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.720336Z digest=sha256:a1e1768558bf6f6a9ddbb6e2c127eeb67e005e355e93d347a89aaf1f2713e454

Observation 3dacc8ed-c96d-498c-8816-f21d7e5c368f · outbound

This paper cites Prompting scientific names for zero-shot species recognition,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Prompting scientific names for zero-shot species recognition,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.724722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.724722Z digest=sha256:59a86d3b3e430a11938577c88c67846b7d22b898308e113f01db4b6bdbce6c7c

Observation 1470bece-a105-4470-b068-35d2358f5e17 · outbound

This paper cites Waffling around for performance: Visual classification with random words and broad concepts,.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Waffling around for performance: Visual classification with random words and broad concepts,

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.729087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.729087Z digest=sha256:c4088b2fa23dba4d0f73ea2012ee629057f01f517931525be891b4e6e2e11bc9

Pith citing papers

Observation a924611e-c64a-4e38-af94-4b9e9e490014 · inbound

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference cites this paper.

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference Adapting Vision-Language Models Without Labels: A Comprehensive Survey

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.164716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:27:02.544831Z digest=sha256:183bc78f541f0a79f520a49e77a21ed5c298a3638eff46aa55a029e6c3e5acec

Observation 63954e70-ee0f-45f8-a0c9-8e86a17425f9 · inbound

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference cites this paper.

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference Adapting Vision-Language Models Without Labels: A Comprehensive Survey

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.188003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:42:36.088959Z digest=sha256:7284d109e019eab504a6ea3fa9c828c90dc0a7f28d5a634e97984e25d4ed19e8

Observation 21374d3e-ec90-4c91-86c3-b6e740d7c3d3 · inbound

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models cites this paper.

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models Adapting Vision-Language Models Without Labels: A Comprehensive Survey

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.378855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T09:57:02.377866Z digest=sha256:0d0cb26d8695dfde94d733f34d4a75fd51388155a6107bd8068944fadfa405df

Observation 7646928a-0959-4ed0-9ea9-14e3b74238db · inbound

USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning cites this paper.

USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning Adapting Vision-Language Models Without Labels: A Comprehensive Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T23:10:08.033785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T23:10:08.033785Z digest=sha256:15b0832cf89b2cc2231fe37e3c602e2d74eafe2b28db243c4d38a2775a75ce5e