Pith. sign in

Paper Citation Record · LEDGER

CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2203.07190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.07190 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:05:28.318675Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:50:00.732703Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1dc487ff-e6c1-4cc3-99e6-44a31205782b · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.734828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:1856e6c6ca41b03abfd69f4754a48abeefb43f391b6e5ed1cc0dc3635c235c15

Observation adad59bb-9f9f-406c-9c95-8386147827c9 · inbound

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion cites this paper.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.318675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.318675Z digest=sha256:32abee9f6e33951279b571e84062635e8b8ff5129a28f690d1542f2d8e6dfb8f

Observation 6cacea6f-70fa-4dcc-b2fe-fbd51df36e67 · inbound

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search cites this paper.

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:25.803771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:25.803771Z digest=sha256:b5f5d28c775e30418f08fec69aa9ce53bab366ac5ff0224152714940f2829953

Observation f4bf2ebf-9f7f-49f2-8a34-66c2ce042822 · inbound

(Almost) Free Modality Stitching of Foundation Models cites this paper.

(Almost) Free Modality Stitching of Foundation Models CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:23.709123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:47:23.709123Z digest=sha256:912dbe6105c4743104d91820975e0314f7d80604e81128876276e0e01f6cdf67

Observation b20cc5d5-8999-419f-b50d-de5067b26b04 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.024416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.024416Z digest=sha256:24a6c122844901eea17c1e29bf265c79b66193f818fcbf22d5ac9df618bb6f26

Observation 56c330f5-d58b-42c0-a4f4-b76187d16700 · inbound

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation cites this paper.

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:43.175206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:43.175206Z digest=sha256:a651c072712cc8d9c554c8d758fa5010ec1ed98110902776358c11ddf2fbb4b6

Observation 849ce55d-f81e-411d-a613-d97ed986faaf · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.089156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:55dc618c3fff66796e25a3a7661ab3d8ee53c0d6c30e534d365a481ba8e5697c