Pith. sign in

Paper Citation Record · LEDGER

Learning to Prompt for Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2109.01134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.01134 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:20:08.334206Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:07.768237Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2607
arxiv_reference, observed 2026-07-04T20:00:07.768237Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68a3ef13-421f-44f3-a3a6-9243823bc9d8 · inbound

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion cites this paper.

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion Learning to Prompt for Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:08:55.541350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T18:08:55.311069Z digest=sha256:607664b0c456d9af19dce609a4e16804d4c31be35af7513e8b8d28c3c1bf7b7a

Observation d7d90f01-0bf6-4483-be44-a1e77b96d2e4 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Learning to Prompt for Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.452379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:9076648ca0c2266c3483e2db4253012d8b9fff47ebfb6212e83ac29a6678b481

Observation 57e163cd-281f-4af1-b3db-a4b1ea8f8487 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey Learning to Prompt for Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.334206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.334206Z digest=sha256:0d3ef975ee8033bc0030a9eff32bd2e8a50d1292bfabc616bc81c603be4072cd

Observation 2b551534-5774-43ca-bc29-26542cf35cca · inbound

StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection cites this paper.

StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection Learning to Prompt for Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:29.866668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:43:29.866668Z digest=sha256:a53646cff2cd9604f126572e2716b0a05b57909c53cbeb6a21bff894362f4462

Observation 6495820b-8c45-4a35-966d-7ba4266044cf · inbound

SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction cites this paper.

SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction Learning to Prompt for Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.412320Z digest=sha256:6b1ead483183b300fb19e84fa254d2df180442beb09a533c4303a80a978d9b3a

Observation 52bc2d64-fc76-4e3f-8c65-74ef080be272 · inbound

$\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization cites this paper.

$\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization Learning to Prompt for Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:14:08.549417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:14:08.549417Z digest=sha256:33e430b33a14ec4ea40aa51ad0bde77bae0bea6b5b3d246cd87f9354152c8090

Observation ef9fd95f-1da4-4d7d-af66-76f365e15dc0 · inbound

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion cites this paper.

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion Learning to Prompt for Vision-Language Models

Reference 12

Resolution
verified exact
doi, observed 2026-05-16T20:18:23.191256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:15:50.210484Z digest=sha256:0aa2970281c9d436b82697eb4786826e680009938121d3fb8a7ac12593b49e58

Observation 3d304dad-0597-4e3e-a9a5-8a963d34d842 · inbound

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification cites this paper.

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification Learning to Prompt for Vision-Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T04:45:20.976440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:41:38.183119Z digest=sha256:ecb044251f0931eb72f6ef7a9f1c46c7e7a4e27e2229cea2d38e52c0da39df1a

Observation a5c6457d-c08f-41be-a024-fc9916c5d2bf · inbound

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction cites this paper.

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction Learning to Prompt for Vision-Language Models

Reference 81

Resolution
metadata mismatch
doi, observed 2026-05-09T23:14:36.510352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T23:13:34.852488Z digest=sha256:f303bbe05742db109d0419be538ae4633686d7f3faf254e731e48f58a975ba55

Observation 2aa58e67-2cf1-4014-8423-da1565bb5cfa · inbound

Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model cites this paper.

Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model Learning to Prompt for Vision-Language Models

Reference 16

Resolution
verified exact
doi, observed 2026-05-08T18:28:58.211029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:26:24.377120Z digest=sha256:5d9e5e40cae910243a017f7328df340edf72a129dc9e78fd7fa0dede986abdaf

Observation fefcfcbb-8348-4a23-98db-fd3fecc40e7d · inbound

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models cites this paper.

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models Learning to Prompt for Vision-Language Models

Reference 6

Resolution
verified exact
doi, observed 2026-05-10T19:05:44.275069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T19:04:20.674374Z digest=sha256:14b1ee9bd91e6d8bd9de49e81692f27579b618de511fbb89ab45d10a9a1581a1

Observation 42ba4757-df8c-4608-bdb5-e337fe221c0c · inbound

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection cites this paper.

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection Learning to Prompt for Vision-Language Models

Reference 38

Resolution
verified exact
doi, observed 2026-05-09T01:24:39.013742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T01:26:11.811914Z digest=sha256:0badae53d9254ec49be5ef6d61ac82e3fddb7f4d69a41894c397a45df6a71c3d

Observation 142f6e9b-f692-49b2-a0b8-b0229d78405e · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs Learning to Prompt for Vision-Language Models

Reference 14

Resolution
metadata mismatch
doi, observed 2026-05-08T21:29:15.071225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:f156e001db6b76269594dffbed3545e0a418eb045aab5a59df569793e4c3e4d9

Observation dd73e5e7-bfd3-4153-83c4-51e12cce0f41 · inbound

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning cites this paper.

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning Learning to Prompt for Vision-Language Models

Reference 33

Resolution
verified exact
doi, observed 2026-05-13T01:42:02.392171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:41:07.685217Z digest=sha256:9261cc8465407968f22bfc3c4f3793c7b5b8e96a04728ce745414f6bd803faa3

Observation e15eada8-be81-4a30-9434-4e2fa8cb232e · inbound

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning cites this paper.

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning Learning to Prompt for Vision-Language Models

Reference 33

Resolution
verified exact
doi, observed 2026-05-14T21:17:58.410163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:14:38.092918Z digest=sha256:f25c8033db119d4971fa34e75f63f23fa25620c316e98ad627148345a4c84400

Observation 6de3b287-6771-48c9-859f-749beb818b69 · inbound

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation cites this paper.

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation Learning to Prompt for Vision-Language Models

Reference 43

Resolution
verified exact
doi, observed 2026-05-14T19:57:52.550230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T19:55:33.741521Z digest=sha256:760abe35faee6fb1bcecee31c706d8ba93cd78309c504424739154a63dde1bad

Observation 65b83ab6-516d-4082-b17d-0a6e127439a7 · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Learning to Prompt for Vision-Language Models

Reference 178

Resolution
verified exact
doi, observed 2026-05-20T14:08:20.777711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:6280b734f8c6d08574b5c581ee1bd2fb3ed7dad59bd7fa7557b7459614c89013

Observation 276e70c1-fb5b-4276-883f-82c75cdbbf38 · inbound

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency cites this paper.

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency Learning to Prompt for Vision-Language Models

Reference 73

Resolution
verified exact
doi, observed 2026-05-20T11:38:14.383828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:33:39.121508Z digest=sha256:518e7cbb3f98d4f36ad5abbcd97fa4af26b27d05c76b7ec8a38e1b7c9d325d55

Observation 0aec0bb8-3d96-445c-b9aa-b45811663fdd · inbound

PERL: Parameter Efficient Reasoning in CLIP Latent Space cites this paper.

PERL: Parameter Efficient Reasoning in CLIP Latent Space Learning to Prompt for Vision-Language Models

Reference 38

Resolution
verified exact
doi, observed 2026-05-20T11:48:14.619516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:45:15.339048Z digest=sha256:371ba1e63f9ca79cc03b402159028bf86fd303513a64cd965b72fa3d7ed1a359

Observation e4e39745-277a-47bb-ac9a-de4988b1e285 · inbound

Steering Vision-Language Models with Joint Sparse Autoencoders cites this paper.

Steering Vision-Language Models with Joint Sparse Autoencoders Learning to Prompt for Vision-Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:07.769905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:56:57.716246Z digest=sha256:98d1f8e3fcd0ca38c6f068b34aa98b4a2a10876e21128ec40f9abd0563be87b0