Pith. sign in

Paper Citation Record · LEDGER

What does CLIP know about a red circle? Visual prompt engineering for VLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2304.06712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.06712 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:56.874054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T03:33:50.786236Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ac6869bc-71bd-42a0-afd8-8f19df449366 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:06.303819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:1eec7101d16343f330b73c891314a1805415cdae613e94a612198414ea5c9025

Observation d66c9429-a5b5-4afc-ba0d-b70f8ad9d70b · inbound

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition cites this paper.

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:33:50.788761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T03:31:58.848180Z digest=sha256:69e618a448b4b298ed6445bc604d612c9193f58ba0f3def7836ffadc2b24675b

Observation dcb66970-d321-4245-98b2-a09e0c80b1e4 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.551709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d355d683170fe92ef3d3cd2d6593f62b18fc5df03c18fc74b1f569e9cb449cb7

Observation cef1a705-6e1d-4fe4-b967-8f7369bcafec · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:56.874054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:56.874054Z digest=sha256:f4f42dd7c78a42fb609f9dc603a119c41d49cf86706527cde9e4a8b8d405a568

Observation 28371d84-9f29-4461-852d-8375f006c0d9 · inbound

ConText: Driving In-context Learning for Text Removal and Segmentation cites this paper.

ConText: Driving In-context Learning for Text Removal and Segmentation What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:11.648041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:01:11.648041Z digest=sha256:8fb8a009ecf07326442f38c6fad89cfd951d618430a2faca9ef0f795acc9daa7

Observation b7290177-dbcc-48bf-afe5-da6924556643 · inbound

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning cites this paper.

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:04.350951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:17:04.350951Z digest=sha256:3fdd56a70f03aefbd99a3acbfb375dd83fea9be92085bf11b5bf486f50ec5fb4

Observation 15dba7cd-b3ed-4867-80aa-c3808a26b4e3 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:22.821052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:22.821052Z digest=sha256:c4a6f900be81deaa605e9d9dd4df0ebb0d30416feb647cb3d3f20acd72acd0e9