Pith. sign in

Paper Citation Record · LEDGER

CLIP-Adapter: Better Vision-Language Models with Feature Adapters

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2110.04544.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.04544 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:52.768260Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T23:23:36.310592Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b72a58df-b73e-4dff-bf74-7be0b4222d5e · inbound

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion cites this paper.

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:08:55.450615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T18:08:55.311069Z digest=sha256:fe1a0abae66617ed639a39fe8e2585ec328f8fc725dfc68bb451c6920480aee8

Observation bbf81b39-5674-4d5a-898e-052eaf597d6d · inbound

Adding Conditional Control to Text-to-Image Diffusion Models cites this paper.

Adding Conditional Control to Text-to-Image Diffusion Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:10.987817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:43:10.880338Z digest=sha256:6112e362562862c1d168cc467c60e3fa541fd428d05380cc2639739b72413f7c

Observation c097a2d0-d901-4d7a-83f4-fb2e73ffe158 · inbound

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection cites this paper.

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:36.979529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T11:11:36.847364Z digest=sha256:50d4ade82449975dc57e7d60529829b94bc980c6e33646558ea1938690304992

Observation 28e64406-67e1-4019-ba21-738b2169d270 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:07:42.951425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:a3bec339281320b1f16a9898b060438bd0400f1d71d5bf11bb20ef5014129899

Observation 61a35d66-312c-4a3f-be7a-6e7bfc6afb7b · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:41:04.876978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:c5e64c6694d6d58ce35ff673eaf6da45cf87f9e4f531a8b9b7740615713710ea

Observation 6424d0d0-15a5-4132-a410-f31e6449ce9d · inbound

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels cites this paper.

Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 195

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:35:48.111287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T16:35:47.826165Z digest=sha256:38ea0421e41d3bfdd6249764654709a9a5654c41ba0e4748a1d93a83d98a75c2

Observation a2026741-05e0-45bf-b428-d6a4a0b0eb3f · inbound

Robust Adaptation of Foundation Models with Black-Box Visual Prompting cites this paper.

Robust Adaptation of Foundation Models with Black-Box Visual Prompting CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:23:36.314337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T23:23:03.562550Z digest=sha256:ba65148597ba39d00ca40d982e9eb6d29a9fcf4b3632399ea8f36c9d76955670

Observation 816bed74-d1f1-4642-bb44-4ad8b8361707 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:52.768260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:52.768260Z digest=sha256:f3c516d0ca371f9792da4eb49614df40b681e79508d194702ad46b928a17189e

Observation e0bfc597-ddae-4ac7-9bf9-2bf7e9f83ac8 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:46.278075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:46.278075Z digest=sha256:76934ca470376c6b6618fe2678914fb1c06a42e72f16461bfab65bc27472fad4

Observation 8d07c900-70c3-42d7-ae87-1fd1be7b6793 · inbound

Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting cites this paper.

Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:47.428858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:47.428858Z digest=sha256:325643ec4227a42028305dee43ff0e3fdfb71194307d16060ac5db0bb3708529

Observation 7b745bcb-5c88-475d-b6fc-c15536b35469 · inbound

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization cites this paper.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:52.021631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:52.021631Z digest=sha256:6f83933192001e1dc18d7af0b8d7d5a0fb674da701cafc2b1f0aff85fb334271

Observation 0041a258-e634-4362-89d4-a2981ace20f3 · inbound

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning cites this paper.

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:01.932831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:01.932831Z digest=sha256:ee23fd376e9d8b1440b9e817441946960a1d10368e36dc07814852042acb2af6

Observation 062cd90d-afa2-40bf-b4ed-a5ecf1e1fef6 · inbound

GLAD: Generalizable Tuning for Vision-Language Models cites this paper.

GLAD: Generalizable Tuning for Vision-Language Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.530702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.530702Z digest=sha256:d5cc8cff0dab72745238f62c66dcb48b6b724a6dc186d4d9748b89773918d0b5

Observation 5b1d91b1-9982-4e84-9702-900d907448db · inbound

TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking cites this paper.

TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:31.568388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:31.568388Z digest=sha256:40d7a6540980f11696c52b0058530f413ee2a42e5f63a226a1213434c691efc6

Observation 050025fe-d975-4125-8a71-5a79d424d946 · inbound

Regularizing Subspace Redundancy of Low-Rank Adaptation cites this paper.

Regularizing Subspace Redundancy of Low-Rank Adaptation CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.803336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:23:34.803336Z digest=sha256:20ea66dac4686287b2890df01c8773a1fd863a65607241a109f91046f6da82f9

Observation 2a6c4d74-0d82-42b3-95e9-63219634dfb2 · inbound

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model cites this paper.

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:32.816317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:32.816317Z digest=sha256:b53adf2ed31a4f5bec01c47d1a44e43aa04c8a4997ba60d18e3a7f848816009a

Observation 79b66bf9-aea1-449c-a6a8-283d765ac7da · inbound

Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts cites this paper.

Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:23:01.918374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:22:35.204623Z digest=sha256:583986a356b5e9737906ab64df116780dace4bc7c1a1866ed2b5f687bc677580

Observation afbdcc36-e4aa-4436-8e10-50a70d4dc60e · inbound

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes cites this paper.

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.709803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:19:15.019645Z digest=sha256:c19d9545d32265307e8a604b5ab87d1ce597e2261945d5b7952f14219f3206e0

Observation bb36a05e-73be-4bb2-ab22-3274e1625924 · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.912472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:52f4e1b79dcf8aed0a4351cbf6d7a353dee6108c21dc6566e02066983196a547

Observation a6d60b2f-8e99-4e12-bb6e-c15450028fd8 · inbound

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models cites this paper.

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:22.563299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:56:31.648428Z digest=sha256:ecdadca1c13254c3c65f194435ca29dc27bcbb7425685bd557ce9814b0e02141