Pith. sign in

Paper Citation Record · LEDGER

Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2110.05208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05208 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:07.932304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:47:23.261860Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3088681-764d-47e2-bfed-00f40ccd0737 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.475201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:da3dc348c6c3e2e05494726878d8356700d8190574c51ce94fc59fcfa6641d56

Observation 43973210-4fda-49f0-b1e1-5407148d264f · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:20.309226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:ca51301c6c1965ec2bb3f93e18a0124c03a3e88bd4184224e943369000f50d15

Observation 2b4ca0b7-0483-459c-8a3a-09de53af94b3 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.972218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:b1c6a1d6e3bf0e2ef9987295e7b5b1468a7a3446e57ff052ec2dd0ddee4a096e

Observation 9066309a-9e9b-43bb-8d46-653d2303d4e6 · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.755350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.755350Z digest=sha256:169fae67b52f78d2ba5fe64fa28c615b84a8b90b6bc014d3490dfc1724671c05

Observation e3248fd1-ec62-4470-8ee1-387d7052c612 · inbound

HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives cites this paper.

HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:55:23.799740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:55:23.799740Z digest=sha256:36962274b936acfe8d908425e73542d35ab0765e4fc1802c5878268a50023c66

Observation 9e9d7ef6-4019-4fdb-ad6a-63b163c52140 · inbound

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers cites this paper.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.135248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.135248Z digest=sha256:71cdecc01aa9cc868487299f89eb2efbff711f246c06de93b8e3bb6a2c45dd3e

Observation 5a182fec-8b24-4af0-aba6-5b67af6419af · inbound

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining cites this paper.

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:55.429777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:55.429777Z digest=sha256:2ad563a11dc3ee9c287c8c94dab201188f93b79012932760906ae862bf9e5dd0

Observation 7827851f-bd7d-4a94-8440-0bda523518d5 · inbound

FLAIR: VLM with Fine-grained Language-informed Image Representations cites this paper.

FLAIR: VLM with Fine-grained Language-informed Image Representations Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:06.735220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:06.735220Z digest=sha256:ebec2f9ce9a493a7b95bd76908670cf794eed261f7ad5d4f2b5ec8e646db1558

Observation eaed1dd4-93c9-428d-b2f0-f0d9497deb5b · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.418704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.418704Z digest=sha256:5708bc9062d0cb8d428c50a6c16b09d7f61139c1ad73df0ae76c2f29b6aba7d8

Observation 55cfc67b-9615-45a6-b1c6-22fd8b29633a · inbound

DiffCLIP: Few-shot Language-driven Multimodal Classifier cites this paper.

DiffCLIP: Few-shot Language-driven Multimodal Classifier Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:02.387573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:02.387573Z digest=sha256:4e5a0b5e1769c110b81ea68d661f91890e0484d10beed57242197bd97cc30947

Observation 36871028-b91e-4da9-a0f0-6ea3e3c58058 · inbound

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples cites this paper.

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:24.360452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:32:24.360452Z digest=sha256:25d41ca7031c5bc0152994ee7313194e5a98d25b02684e9e2bda295aac653e2c

Observation c8d1c715-6f9e-4bdd-a4f7-801b17483251 · inbound

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework cites this paper.

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:10:44.688935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:10:44.688935Z digest=sha256:429992c93b7945c75f9166fb961fa69b8f8ba69091d1f1cdf47e855d66ceb9d2

Observation 4ff8a9fe-0978-4103-95c2-21f22b981205 · inbound

Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting cites this paper.

Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:10:19.572614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:10:19.572614Z digest=sha256:adc40f77a0645a72fff69a96ee19a3eec1c076dbb9f3369a16353dd4d0c4a86a

Observation e7fbefa1-7d61-472e-9f7b-ca5749a063c4 · inbound

ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models cites this paper.

ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:40:27.171479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:40:27.171479Z digest=sha256:5ecfef2d2c993ba792b1d734a115b15463832718a51c5749b8a72192c732a3f1

Observation 1af0c70c-95e4-4bc5-ae4f-476ff02b4339 · inbound

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis cites this paper.

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:32:36.832435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:32:36.832435Z digest=sha256:728c69a56ffa87953ba081d05c9363c502eabf2e5c8b6c2a09ecd32be051ca5d

Observation 7eaf461c-ec4e-4ac5-bf04-695ec715f0a5 · inbound

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion cites this paper.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.238464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.238464Z digest=sha256:c0681e9a3288353d785fb20f073edaf9f06922916b8e06fd81e487e4fd8e769c

Observation d77a6c48-9565-4f39-98f8-80179d8c550b · inbound

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis cites this paper.

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:11.100535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:11.100535Z digest=sha256:158d7ea1cf4624e596850f85946d2eb14f4a0b4d573249104e156cfbabd00270

Observation 016af8a9-8f2a-410a-80f5-dbd356d88a0f · inbound

Aligning Proteins and Language: A Foundation Model for Protein Retrieval cites this paper.

Aligning Proteins and Language: A Foundation Model for Protein Retrieval Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:49.294185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:50:49.294185Z digest=sha256:4e077acd3e6866e7479a5ae56676aa25db50472e8d06547edfdcd2eaa996e7df

Observation 34035c79-8394-499e-ab5e-55eef93e6e98 · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:20.470954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:20.470954Z digest=sha256:00e16b873f3137529632f1a95ba72ee761d8fbd2c92239b3595029aa1f47b91e

Observation a9d6ab1f-8d29-4de2-b802-510eeafda76c · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:31.213119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:31.213119Z digest=sha256:1a24c1b0360e9addb916f665390f4735d84f60dea3873f856c8daa5fcec80395

Observation 7568b478-a665-4c51-b333-02e35e594616 · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.571959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.571959Z digest=sha256:6bac12d0d3804268071daef56eaa188a5600f8827964e8b2e564ec031e2a90ea

Observation 31375c32-c378-4279-8aa1-deb280288a65 · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.369563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:0c5d355575587116b0c48146bea2a05010c32e0289fc0de38f35eeeec5687f83

Observation b90db992-73a4-4d65-b59c-e8ee759f25d4 · inbound

Neutral-Reference Prompting for Vision-Language Models cites this paper.

Neutral-Reference Prompting for Vision-Language Models Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.591922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T19:15:54.152498Z digest=sha256:9ca2849ddd46ed2f0419c926fec132328b656fa95eb76ac20e241c64d6fab3b1

Observation 40426464-81bd-4cec-aac9-4c6a226316bf · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.752092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:d223749575e01ad6b9a87d3deb0c84b43339984a4b5a1d54a713aaf57bc51076

Observation e444a6ec-ce9f-4c29-9e81-52cb1ed6fd91 · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:47:23.263021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:c30e4a9a226d4d7d98183122a779cca9bb6330a640dc5cb63d07243927ae9ad0

Observation 89e3af63-ede3-4e24-89ff-40271e8def59 · inbound

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition cites this paper.

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:07.932304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:37:07.932304Z digest=sha256:6496451e81e5e2f3e4a612ef0c4a1b0f9ab4437c71816e96b4d4c6fb997b7c2f