Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2203.02053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.02053 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:19.855755Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

98
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad9eeb77-3851-4f76-8702-3702171fdb4a · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.094675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:112380640404711ade69384e69d027918b36eb5f1d8203efab34402a61ee22b6

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:a900f4d256ebadc8ae0686541584487f97c98cbe8ab0b63d994dc30f2b257e70

Observation 4ea027f2-e3fe-4a8d-8cd9-a931b167e2b1 · inbound

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps cites this paper.

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:37:23.778509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:37:23.778509Z digest=sha256:c3570c156633d8bec5bab7af9920f84ecb3657550b2092a8d46f1919722b8383

Observation fcc47bb4-a808-48f3-84ad-483bb1702b32 · inbound

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models cites this paper.

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:50.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:50.461183Z digest=sha256:85f4147653fbeeb2a9272860bc8000a17c932ad069a69254e80de9aac53d6bdf

Observation eb9a4f0d-d496-4c51-98a7-aed0c5d3577c · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.947358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:5db707e9401e3d6ad8eb8b7e08df85414576033388e53b678ae30159cd640b74

Observation 31fade60-4c4a-48ce-9548-16a5c26c91dd · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:e90c533de368df820dcab1dcc189cd842713436c658d42ad1be5427acfcca71d

Observation 65e85d52-f2ec-425b-92e6-29a8222780e3 · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.934858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:9ff4178f1ecc0ecfd2749ddbbc1a0a039972493ebec607b4cfe33a151efa736c

Observation 90dd2a2a-699c-4669-a914-e563b100455d · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.955356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:dfb786d7eac82506c500e4aafbb836907ea3975e11f8c8346c9d51ea0b33148d

Observation 672f9bdc-1ebe-4665-a0c0-58a83e959e86 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.740739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:f7e208a9212466dd168417676ea5da7f4af72cef9617c39ef361c08514c25c7f

Observation 4ed090c5-521d-49d9-8a7c-e078e60cf360 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.736252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:8c4157d24433e30faf63c81be14a4ee93bbd7e85026cc54a6326fc0a1aa93a90

Observation 65b23aaa-f7d2-4ab2-a270-9cc6ba0fc6c0 · inbound

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift cites this paper.

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:05.858809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T02:12:54.796350Z digest=sha256:6f18ef2281195b935526a8f1b744f54cf9e05469aa8a93346a2023d6e6456c95

Observation 77e2541c-9271-478f-9107-975d68981210 · inbound

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning cites this paper.

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.665073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:32:00.361991Z digest=sha256:a051430b95bb8b2bbc8fa4354897e002b3c7e5b609b151ecd26ff9fd93af3bda

Observation c5ef43e4-90d6-4834-a917-68e5c3113c5b · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.050488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:651f178ceb910162639ea7ee3e3e91343f1e47a30d1000c0ef347ffb177ac650

Observation 40b2cdd8-71d1-4a78-b8e7-790c42dd38ce · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.796470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:e034ce5e31dd67eff4c2a384c106a3e3cb543622b7573d74e4a126bdfcff5cbd

Observation d65a7949-bf34-45f5-aecf-fbe5eee2f8a7 · inbound

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms cites this paper.

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:34:13.357363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T03:27:31.177489Z digest=sha256:862f99d9d0c29a8215df3eabbab777b3032aa832ec3ff0001a1114ce870cc1a1

Observation 58328478-9e2b-4265-9cf4-83b2d7999dc3 · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.287239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.287239Z digest=sha256:64350eeec0f58e92afc5f3b4c73764840b28b9a485a3b0b02960657ff901f987

Observation fdc3b59e-ffdc-4230-b0ab-23375b4f97d9 · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.890280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.890280Z digest=sha256:8ec6e408001f8b5a9c83be234fe0204431f3cefc0c7ca1159e72037d03ca4a39