Pith. sign in

Paper Citation Record · LEDGER

What does CLIP know about a red circle? Visual prompt engineering for VLMs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2304.06712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.06712 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:56.874054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T03:33:50.786236Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ac6869bc-71bd-42a0-afd8-8f19df449366 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:06.303819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:52fa29b1653a93effd677128aa4b52f87347cc83c14e5e53f2042f12611e336c

Observation d66c9429-a5b5-4afc-ba0d-b70f8ad9d70b · inbound

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition cites this paper.

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:33:50.788761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T03:31:58.848180Z digest=sha256:7619bc55c17747af003e18d0b539ae43f201f86281fd01ac7ec12bb56aa43c63

Observation dcb66970-d321-4245-98b2-a09e0c80b1e4 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.551709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f48514613957ce0c66576f322addc956dacfd588230301a4e9f18eafbb2d3219

Observation cef1a705-6e1d-4fe4-b967-8f7369bcafec · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:56.874054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:56.874054Z digest=sha256:e5fd08fce815b08a729bbbbe49b6cc10a575563425988a07bd5e760fda758724

Observation 28371d84-9f29-4461-852d-8375f006c0d9 · inbound

ConText: Driving In-context Learning for Text Removal and Segmentation cites this paper.

ConText: Driving In-context Learning for Text Removal and Segmentation What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:11.648041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:01:11.648041Z digest=sha256:8fb8a009ecf07326442f38c6fad89cfd951d618430a2faca9ef0f795acc9daa7

Observation b7290177-dbcc-48bf-afe5-da6924556643 · inbound

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning cites this paper.

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:04.350951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:17:04.350951Z digest=sha256:3bd4fa33986a92092c72d746fbfa8b4d2ea87145aaad92506bcfe5642f836a2b

Observation 15dba7cd-b3ed-4867-80aa-c3808a26b4e3 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:22.821052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:22.821052Z digest=sha256:c4a6f900be81deaa605e9d9dd4df0ebb0d30416feb647cb3d3f20acd72acd0e9