Pith. sign in

Paper Citation Record · LEDGER

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 4 inbound Pith citation observations for arXiv:2505.03703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03703 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:49:24.756575Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:39:05.154618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T09:42:30.741220Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b05f7d0c-d197-4e7b-9b58-546959f31903 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.699051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.699051Z digest=sha256:f538141a5a88157b1f28cc076597855394ab9dc77ce7f7c74738fe851df2e814

Observation 25b28392-3b5f-4f04-a5e8-6a963e629f53 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Sigmoid Loss for Language Image Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.704727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.704727Z digest=sha256:9c2d599bbd557c6823c889612f5561b7062d644e01f6c8d7cb86266eaffda34f

Observation fa06d3a8-55a0-4e05-8d22-527827118964 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint 2411.04997, 2024.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint 2411.04997, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.709451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.709451Z digest=sha256:04de013bfc74443f3ac1c71e783b77395e24d7b50d6973c46ee312b1d404b82e

Observation 081861d1-f654-4530-b405-766702e8acd9 · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning ColPali: Efficient Document Retrieval with Vision Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.714057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.714057Z digest=sha256:f23d4cd84940c767edfa18e69727dbe2f6e1b0bdf6d2eee2c96d56dbc15f8d81

Observation 8b75a658-b4a7-4591-b72c-cc076fbd4dc4 · outbound

This paper cites Mind the gap: understanding the modality gap in multi-modal contrastive representation learning.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Mind the gap: understanding the modality gap in multi-modal contrastive representation learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:25.008244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.719757Z digest=sha256:7bca405fcae05dfb178c232a5d349835d66ada6039a208f835b142a099d4d208

Observation c221131b-7b09-4d05-af05-a5623206d7c7 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retriev- ers.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Uniir: Training and benchmarking universal multimodal information retriev- ers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.990060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.724652Z digest=sha256:50dfb5c163a3a164516c96ada3060cd8aa1065b61decbddcd5911ab68c3dcf29

Observation fe76457b-ceb6-483c-be0c-78bbde9e4b48 · outbound

This paper cites an unresolved cited work.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:49:24.974389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.729637Z digest=sha256:2f76539f60060ae551af49e9298c55fd892d09146c5171ff7c12eadf74276904

Observation ae0b1950-e9fc-481f-b128-63189e614000 · outbound

This paper cites A Tutorial on Spectral Clustering.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning A Tutorial on Spectral Clustering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.734226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.734226Z digest=sha256:b191f9a0527a658253423bd914a176036c8b8659297cb3a49d9b2734c3a4e22d

Observation f735136c-af65-493d-908f-9ea012613aab · outbound

This paper cites Regularized discrete optimal transport.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Regularized discrete optimal transport

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.958064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.738839Z digest=sha256:d0ffbcb100086ab81d8cf4104a0c516324d8657a325db7b0af30d113c254291c

Observation 7f3fb7f5-8139-4cb5-ae49-b7827fe6d1c7 · outbound

This paper cites Optimal transport with laplacian regularization.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Optimal transport with laplacian regularization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.942605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.743237Z digest=sha256:8080420d12380bc72f61d4dbe2e9367639fbf724f5e944042dbb017ec875862f

Observation 1e1980cd-8e85-465b-8798-6ecdb1719910 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.926768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.747664Z digest=sha256:2189ce33ae2edc4428791ba487fa3a864397c1886ffa07788ae469b6d31d694d

Observation 7fccb29f-b114-440f-bce8-0473a2caddef · outbound

This paper cites Lawrence Zitnick.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Lawrence Zitnick

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.752120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.752120Z digest=sha256:2d099113b7d291168fb2def8383e43a42b776db8462c21192eaf7f5dd00f491e

Observation 2b385166-435a-4361-9a3d-781166696547 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.901369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:49:24.756575Z digest=sha256:0552941e8e5e3f750389d99924c1eac4270cb8331b9bf19b88a5cb1bbd5faa07

Pith citing papers

Observation 54cde237-0bc4-4ebc-acaf-e3f5582b2e54 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:05.154618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:05.154618Z digest=sha256:720e1c448a5483eaf8e750a34351b77433788ecbc760a52114db8cc33e0bb150

Observation e6954add-8b1d-4272-9f05-16a6c374c73a · inbound

Asynchronous Federated Learning with non-convex client objective functions and heterogeneous dataset cites this paper.

Asynchronous Federated Learning with non-convex client objective functions and heterogeneous dataset Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T05:28:38.325633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:28:38.325633Z digest=sha256:1e59a50219cb549b9ea8c746caf0a9249b8cae75f997532c3ae6ac77a50ce2b3

Observation 372ad32d-ad96-4cd9-9f47-e6bdeb8bf8f6 · inbound

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization cites this paper.

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:42:30.745067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T09:41:38.974413Z digest=sha256:e6fbc292558a898f2d77737e8d0d8ca527b758975b448bf152c429072998a0eb

Observation aaa20988-4bfe-48ae-b506-7093ef8b35f1 · inbound

On the modality gap and the contrastive loss in multi-modal representation learning cites this paper.

On the modality gap and the contrastive loss in multi-modal representation learning Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:53b09ba46b059e74228e087484056eafc099f387cb125f4d1b8d81df4738f83f