Pith. sign in

Paper Citation Record · LEDGER

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2505.12473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12473 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:39.776348Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:06:08.750038Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:56:38.715423Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25965cb3-610e-4f41-85a3-2e52102d33ba · outbound

This paper cites A kernel method for canonical correlation analysis.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables A kernel method for canonical correlation analysis

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.676329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.676329Z digest=sha256:7c9a1632ab9cce123beb24148ae0fb1d05fc10d0145fcd99054f2a92f75232e5

Observation 5ebf1fa3-c7db-47ce-a0fb-baee357697c7 · outbound

This paper cites Andrew, G., Arora, R., Bilmes, J., and Livescu, K.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Andrew, G., Arora, R., Bilmes, J., and Livescu, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.210667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.681091Z digest=sha256:f70d92b8aabc2ad142a457a0c77ba856ba5bfdbd661055909ea47170ca46cf19

Observation 6ef8bc1a-8ef4-4367-8033-8ced1d52ba6b · outbound

This paper cites an unresolved cited work.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.701819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.701819Z digest=sha256:2551a4acc672a52635d16d36477f61edc4c3797a707be60e8447a4a9053b9ade

Observation ba115941-12b0-4b64-9638-7b5df0b3a6e1 · outbound

This paper cites exp σ(f(X),g (eY )) τ !#) , EP logeqg(Y )|f(X)(V |U) pg(Y )(V ) = 1 τ EP{σ(f(X),g (Y ))}− E eY ( log EX.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables exp σ(f(X),g (eY )) τ !#) , EP logeqg(Y )|f(X)(V |U) pg(Y )(V ) = 1 τ EP{σ(f(X),g (Y ))}− E eY ( log EX

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.128660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.739062Z digest=sha256:748d598f5d1777b6b6a1831733e90862cc729b4e827862128022351e393adf7f

Observation 3acfa6e1-dea7-49d9-8287-60ca3e7334c5 · outbound

This paper cites See proof in Section F.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables See proof in Section F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.170604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.724396Z digest=sha256:cd30a7f9af4fbe894dd4ad1e090cacc03d659bc1a387b5f5c852a45a7ece758a

Observation 9c661668-5e5e-4b97-b718-491dec4c1d7a · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.689343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.689343Z digest=sha256:376626625f03b570adadf430f161750b4c6e284ad2af81fa3f6c74785833125e

Observation 92a01af8-68de-4289-ab38-109701f3434d · outbound

This paper cites an unresolved cited work.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:42:40.160649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.727758Z digest=sha256:0251d2556b44034fb5e40e691358ee06583f814f13715d386f29675514e99563

Observation 7d55e8e8-f07c-47d8-9732-ff58285ce3a1 · outbound

This paper cites 23 See proof in Section F.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables 23 See proof in Section F

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.150490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.731383Z digest=sha256:875dd5c281599ef10e00255d8e44bb8680d61c8e54fe27d75d543c04bf8392ea

Observation d7721b3f-2fd9-4b96-bd55-42a6ced688de · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.693817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.693817Z digest=sha256:63bedd3cacb638b20cfe34def2b907a3dedac9587b3c17c60105d47f5cd80ec5

Observation 2b556839-0d1e-42e6-96f5-8533334fafe5 · outbound

This paper cites Here we defineH(U) as the entropy of random vectorU andUM = idM(U), where id is the identical map on Rd.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Here we defineH(U) as the entropy of random vectorU andUM = idM(U), where id is the identical map on Rd

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.139192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.735786Z digest=sha256:06284bd26564e2473a71b67dd623603bf01e1d8136d6fdf6a98a901834d6c0cd

Observation 44a2d300-b687-46ce-8e51-1d6647899f51 · outbound

This paper cites an unresolved cited work.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:42:40.118246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.742520Z digest=sha256:81f948a359b30bdf17c4b8e1b9b26a9cb0d0fbb37e6229eb0396ca6b62db1750

Observation b5890345-ed4d-4e37-8ff5-fa9405904579 · outbound

This paper cites Pan, Y., Mei, T., Yao, T., Li, H., and Rui, Y.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Pan, Y., Mei, T., Yao, T., Li, H., and Rui, Y

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.201002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.705413Z digest=sha256:23ca9371c3450b9145b7c2f537010101597f405da6d455799a7226fefc715b15

Observation 1d52fc87-e1be-4e3c-87b7-4cc1b718fef5 · outbound

This paper cites exp ⟨f(X),g (eY )⟩ τ E∥f(X)∥E∥g(eY )∥ !#) + E eY ( log EX.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables exp ⟨f(X),g (eY )⟩ τ E∥f(X)∥E∥g(eY )∥ !#) + E eY ( log EX

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.106861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.745788Z digest=sha256:916c569f9ded6f9a83ae047531cd84a8a48abc0aa63d6912a82c74d9a7422966

Observation 2598a32d-eef5-468e-ae53-ca5a5931c5d9 · outbound

This paper cites Particularly, for anyη >0, withτ =ε(η), it holds that lim sup M→+∞ L(fM,gM,ε (η)) + 2I∗ M(H) > lim sup M→+∞ L((f∗)M,gM,ε (η)) + 2I∗ M(H) ≥ 2η, for allη >0.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Particularly, for anyη >0, withτ =ε(η), it holds that lim sup M→+∞ L(fM,gM,ε (η)) + 2I∗ M(H) > lim sup M→+∞ L((f∗)M,gM,ε (η)) + 2I∗ M(H) ≥ 2η, for allη >0

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.096331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.749196Z digest=sha256:bc5a75c8b6bff9e9c5daef33ef3f1e64bb38cfcd6f24f1c770fffc96a7b25510

Observation 93bb4707-dfd3-4d3c-a4de-92cdd20df2d0 · outbound

This paper cites an unresolved cited work.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:42:40.085276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.752985Z digest=sha256:0ced945b37ee248816b79e014ab102f49ddba6b2ccaebeafe45727f93d773fcc

Observation 364098dd-78a6-4c3a-9e29-39b1bfedfe2f · outbound

This paper cites E.1.2 Extension of Proposition 2 We then turn to a general case whereX∈ eBd1 and Y ∈ eBd2, where eBd1 and eBd2 are bounded sets in Rd1 and Rd2, respectively.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables E.1.2 Extension of Proposition 2 We then turn to a general case whereX∈ eBd1 and Y ∈ eBd2, where eBd1 and eBd2 are bounded sets in Rd1 and Rd2, respectively

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.063138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.759626Z digest=sha256:abbab6f5b673b39bfdfb2a84216bed46ffc3d5420fe2b233939ace0f2c2027ec

Observation 070fde9e-4da2-47ce-b96f-9f47c98f2480 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.713324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.713324Z digest=sha256:fd166695c2041e9ef7946b92041dc830aff85f2040ce57e90bf971380a082e30

Observation 165f2d05-de04-46bf-bcf6-b796d0772f90 · outbound

This paper cites 40 Proposition3.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables 40 Proposition3

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.051260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.763073Z digest=sha256:ac0a0211d08dc50a38275399e775dc9d4b3252a0f289b1c32a7db447f054e864

Observation e2563912-40e7-4d33-ab4a-736ffb8a79e0 · outbound

This paper cites E.2 Alignment and uniformity with correctly specified dimension To begin with, recall that for any(f,g )∈A (H), it holds thatf(X)/E∥f(X)∥,g (Y )/E∥g(Y )∥∈S d−1.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables E.2 Alignment and uniformity with correctly specified dimension To begin with, recall that for any(f,g )∈A (H), it holds thatf(X)/E∥f(X)∥,g (Y )/E∥g(Y )∥∈S d−1

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.040151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.766295Z digest=sha256:c8dcb432ebd99d904b49b2edb152db8382ace5a2f984161c05b8457cf6aff282

Observation 78e6023b-7944-4dcc-a6da-94b57aa9e9bd · outbound

This paper cites S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.180907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.720671Z digest=sha256:187bdd8653eb70abef3e2e2e3ca263f4d6cd85cc0bbed0307af53fb9387aa690

Observation 3b9426d7-d8bd-4248-88b1-2bce31d81bc3 · outbound

This paper cites In addition, with the choice ofH∗, the uniformly distributed representation onSd∗−1 maximizes the entropy.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables In addition, with the choice ofH∗, the uniformly distributed representation onSd∗−1 maximizes the entropy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.028014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.769618Z digest=sha256:1b869ac1f4b8daeaa90a9626e8066070bb7b97ffee5d6232b7787ce825dce566

Observation 0f590013-6ee5-4a91-b673-31fa086e1b4f · outbound

This paper cites exp ⟨fM(X),gM(eY )⟩ τ !#) + E eY ( log EX.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables exp ⟨fM(X),gM(eY )⟩ τ !#) + E eY ( log EX

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.013932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.773126Z digest=sha256:12f23b41fb71fe28877a38e9814e3d484d909778b15b9516885e28d27b5ade2a

Observation c6894f33-87a8-4ede-aeff-e4e0158685e4 · outbound

This paper cites This is a photo of a/anlabel.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables This is a photo of a/anlabel

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:42:40.002853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.776348Z digest=sha256:c034503ccca38ccd795c3601fbe7571463952acace011cb3addb857599e41b53

Observation f0f6ec58-9eaa-4508-8075-3118f320cfde · outbound

This paper cites Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.685064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.685064Z digest=sha256:488bc113323f64ae8b2c32c9cdb3db51cce340c6aa2be44dab7b123ccb6551fc

Observation 07a37c9f-a306-466b-92bc-e1666a62cdf2 · outbound

This paper cites and Mei, S.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables and Mei, S

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:42:39.951963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.697731Z digest=sha256:df883ab3f035ab107c9de7ead1520509a09d39b9a3b7eadc666af2fbd4204942

Observation d1e0e41a-4c5f-41a7-8b92-1e08800784ca · outbound

This paper cites On the Importance of Contrastive Loss in Multimodal Learning.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables On the Importance of Contrastive Loss in Multimodal Learning

Reference 1253

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:39.709258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:39.709258Z digest=sha256:56bbeeb9fb01fd32f0c1cf448ef074d1ce13c31c12ae12711da08bfb8ca26c5a

Observation 965c0062-ba7f-478e-b2fa-17175122b60b · outbound

This paper cites 39 E.1.1 Proof of Proposition 2 We prove by showing thatY |=X|f∗(X).

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables 39 E.1.1 Proof of Proposition 2 We prove by showing thatY |=X|f∗(X)

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:40.074426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.756284Z digest=sha256:5a3968a54992c75c0c1d68e6cd365bf94ffe8455d53cfe19ad80eae08495021a

Observation 5f43eb1b-25af-41f9-9de3-36f5b3549bb4 · outbound

This paper cites an unresolved cited work.

Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables Unresolved cited work

Reference 9939

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:42:40.190939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:39.717101Z digest=sha256:dcd0d1b5d7162044f97adb2981e87bcecf1dbb3ecef588dfd35717fcb8047e1f

Pith citing papers

Observation 0f940e70-c63c-4f67-b02a-99258dc01f6f · inbound

Is Dimensionality a Barrier for Retrieval Models? cites this paper.

Is Dimensionality a Barrier for Retrieval Models? Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:56:38.718532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-25T04:55:45.919518Z digest=sha256:3341ebc289119a24ef7ad0850521ccbb398a8516d58bf053516e3162d4795986

Observation 0514509c-44da-48b1-a411-1e314d6c8e4b · inbound

DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing cites this paper.

DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:08.750038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:08.750038Z digest=sha256:9348f65fd3ade416b2ad35d7b048035b71906398a0315178d70ff4c92a0a9eac

Observation b7ce2b79-0041-4c0a-8648-736aa2262c20 · inbound

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data cites this paper.

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:41:28.708834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:41:28.708834Z digest=sha256:00bbe0bdb5a5925a842fc1471913c9c18712c0bf7fc541057dbd00412fea4a0c