Pith. sign in

Paper Citation Record · LEDGER

On the modality gap and the contrastive loss in multi-modal representation learning

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2607.10698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10698 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T09:55:43.412471Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41f32ed8-c815-499c-ae3f-455d7144bad6 · outbound

This paper cites an unresolved cited work.

On the modality gap and the contrastive loss in multi-modal representation learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:211bd3db607a9ba4877a9b9908031eddad2fc79498e22a1bd04d2d590b723a0e

Observation 5ec96b19-09be-461c-a827-81bf234957f5 · outbound

This paper cites International conference on machine learning , pages=.

On the modality gap and the contrastive loss in multi-modal representation learning International conference on machine learning , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:53bf0d311fbab5bd6353e571a74ab7d727a73262c94bc5718321a935764f24c2

Observation 7d2dd094-c7e7-4bd3-b877-d184fb62bccc · outbound

This paper cites Welle and M.

On the modality gap and the contrastive loss in multi-modal representation learning Welle and M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:337286602e184146b81d95299888142076ceaefd7a8400ee8cc77b41a0351731

Observation 41717dd9-21de-4123-91e6-c58ec63bbf7a · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision , issn =.

On the modality gap and the contrastive loss in multi-modal representation learning Learning Transferable Visual Models From Natural Language Supervision , issn =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:6d1d00644ca379897e5c3b243aebc2b152575205608e584171721b13b2f233e9

Observation dd5339ae-ce24-4d18-ab0c-59ccd891267b · outbound

This paper cites Mind the Gap: Preserving and Compensating for the Modality Gap in.

On the modality gap and the contrastive loss in multi-modal representation learning Mind the Gap: Preserving and Compensating for the Modality Gap in

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:5828271a58fa8f8e98684e12d523222fa9ac1a12d0458d94fba4c4fba99f8c44

Observation aaa20988-4bfe-48ae-b506-7093ef8b35f1 · outbound

This paper cites Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning.

On the modality gap and the contrastive loss in multi-modal representation learning Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:ec5af4019bca747672cf5a7b558b804666b6f503266ea2bdf07eb9b40f88018d

Observation facefa64-37b2-4c5e-b531-fab1f1d52c42 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

On the modality gap and the contrastive loss in multi-modal representation learning It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:ad49755d13ec2c9642608572f66808266ba82f65bd766cd0c62f807f311abaac

Observation c343c166-52ce-4d12-824e-6580531435fc · outbound

This paper cites Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP.

On the modality gap and the contrastive loss in multi-modal representation learning Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:8631a63504a703b14b3c82090d2301a1f0f03744c4aab5a6d347805a536feadd

Observation fd5d053f-626a-41a5-b30a-92475cfc6c00 · outbound

This paper cites The Double-Ellipsoid Geometry of CLIP.

On the modality gap and the contrastive loss in multi-modal representation learning The Double-Ellipsoid Geometry of CLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:7a21b945dc0de847bc5fb299e538d5f575a6526976b1a256320b390b4d649b2c

Observation 32070759-6146-4700-9b57-ee68525bc559 · outbound

This paper cites Understanding the Behaviour of Contrastive Loss , rights =.

On the modality gap and the contrastive loss in multi-modal representation learning Understanding the Behaviour of Contrastive Loss , rights =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:4e1164d6ffcc881f9f744455f46bc3a9a3ce76d1bbc5d0c7408d1690badd51c4

Observation b5b3f29d-ccda-4800-8891-a3422935d2cb · outbound

This paper cites IEEE Transactions on Artificial Intelligence , volume=.

On the modality gap and the contrastive loss in multi-modal representation learning IEEE Transactions on Artificial Intelligence , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:8a971400f2e182181b2904c94677559a5d96a2be09984bf5a59d1a1c4e605651

Observation 0a073013-80ad-4dbe-a74e-0569ad1fee6a · outbound

This paper cites and Lin, Dahua , urldate =.

On the modality gap and the contrastive loss in multi-modal representation learning and Lin, Dahua , urldate =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:6eea7a9c697b71c4bd6b474d44ee21fcb5a1546619ba8814c67a483af2327a95

Observation 3d0ca495-471e-412b-bf37-e81bfdf42e76 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

On the modality gap and the contrastive loss in multi-modal representation learning Representation Learning with Contrastive Predictive Coding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:333402408e759a8bb946ef904d8e0bd513a4fc01272f504f525fd071ac80aa77

Observation 57a1ac9d-0dde-4e01-a5eb-d15445037a9b · outbound

This paper cites Learning Representations by Maximizing Mutual Information Across Views , volume =.

On the modality gap and the contrastive loss in multi-modal representation learning Learning Representations by Maximizing Mutual Information Across Views , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:27aac5004522ff01a709af8a61b9b731c7568c9bfde5022b6c54fcb92fbbe22a

Observation cc14a885-5c21-4bab-9fcc-322636230be4 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

On the modality gap and the contrastive loss in multi-modal representation learning A Simple Framework for Contrastive Learning of Visual Representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:f6c86829dfedb2a2f0d625ac380346d2c684de94e8efa8f065160c353647d31f

Observation 8cb7ef59-3df2-4720-b2df-fd5704b6422f · outbound

This paper cites Momentum.

On the modality gap and the contrastive loss in multi-modal representation learning Momentum

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:73257421b12dbb62a59f74295998ff64628c432027a9b66e125b5fbfa88dbde7

Observation 61cc073c-496b-4d1f-bffa-538097f8b93e · outbound

This paper cites EMNIST: an extension of MNIST to handwritten letters.

On the modality gap and the contrastive loss in multi-modal representation learning EMNIST: an extension of MNIST to handwritten letters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:79a7134ece93c13dbcf2f02e5e2de1a198b3023cf818dece116881ebd455f97b

Observation cc6b0391-478a-43a1-a2c4-47521cd9fbde · outbound

This paper cites Microsoft COCO: Common Objects in Context.

On the modality gap and the contrastive loss in multi-modal representation learning Microsoft COCO: Common Objects in Context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:e145a1bfba5ca12e1eacf80ad74242dcc1b4bae7b144fd3e056b684d904a1cb0

Observation 0a23ab62-74f2-42b3-bd77-b0fe143c1fde · outbound

This paper cites Geodesic Multi-Modal Mixup for Robust Fine-Tuning.

On the modality gap and the contrastive loss in multi-modal representation learning Geodesic Multi-Modal Mixup for Robust Fine-Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-14T10:00:26.872272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:d1edecb977988f551ad66505420993fac552c10839a313f77ebcfe5d0065b9c8

Observation e3e0d41b-9e54-402b-aaf2-72ea9c0969e1 · outbound

This paper cites Second Workshop on Representational Alignment at ICLR 2025 , year=.

On the modality gap and the contrastive loss in multi-modal representation learning Second Workshop on Representational Alignment at ICLR 2025 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:009be6efc427137bb084cb54d995c1e42bb4982b2621377c5f89418f7977b139

Observation 75905059-fca1-40a6-a4c7-1e0a1d2063a8 · outbound

This paper cites 2020 , eprint=.

On the modality gap and the contrastive loss in multi-modal representation learning 2020 , eprint=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:6f35359f7e5d465f4b93c31da5e45d8c56808f4b1a6c9dc41649fd003746e19b

Observation dc5f12f3-0fbd-496d-99db-9cf40ee2ef0c · outbound

This paper cites 2018 , eprint=.

On the modality gap and the contrastive loss in multi-modal representation learning 2018 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:e2bd36cd8c089b0a9f8b80909892aa37eb0c1b0e8e171b10c84496f7a225a652

Observation b8f3fd5a-b823-4315-883d-aa585eb3ac91 · outbound

This paper cites 2015 , eprint=.

On the modality gap and the contrastive loss in multi-modal representation learning 2015 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:c4444f11c0f7bdd513140146bfcccd045b202c5dc6b14b239a5249039016c4e8

Observation 1ea4d842-e4db-40b6-ae32-38ac2ea48ac6 · outbound

This paper cites an unresolved cited work.

On the modality gap and the contrastive loss in multi-modal representation learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:554327c14e3a5d370863432c3ba65c6e5cbf6c31a3d3c7996274342a27ce1e37

Observation 64f25ac6-0a42-447e-9530-f48a44ff3d4c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

On the modality gap and the contrastive loss in multi-modal representation learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:2a31b2b106f6f01091f8c9916c8b1fec9d89f14cca69d3db5fa09c46754cc80e

Pith citing papers

No inbound Pith citation observations are available.