Pith. sign in

Paper Citation Record · LEDGER

Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2404.07983.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.07983 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:34:17.947312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:56.194117Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ecae586e-f37f-46c1-93b3-84815eab7a7a · inbound

The Double-Ellipsoid Geometry of CLIP cites this paper.

The Double-Ellipsoid Geometry of CLIP Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.869500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.869500Z digest=sha256:2de7b69221e0849fb80162c5471a1ed13689b2d1276cef8d7fc2982b79f4b8be

Observation b3a268e0-b583-43d8-94c3-96370e224187 · inbound

I0T: Embedding Standardization Method Towards Zero Modality Gap cites this paper.

I0T: Embedding Standardization Method Towards Zero Modality Gap Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:33.149533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:20:33.149533Z digest=sha256:53b3c82c02eb91b4f9e2b6af46139db59c3131cb4db9da63e05126d9ea3525d9

Observation 1ef2bc42-8645-4506-9f70-e9690e01dfe7 · inbound

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion cites this paper.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.302016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.302016Z digest=sha256:89f7bc233e240bb21ebb4177fa10f4d664d95482be28be3028a4c7cad2f90d5a

Observation e761555f-6035-4164-a126-e74f218e673f · inbound

Whitened CLIP as a Likelihood Surrogate of Images and Captions cites this paper.

Whitened CLIP as a Likelihood Surrogate of Images and Captions Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:34:17.947312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:34:17.947312Z digest=sha256:46b6f427840787856a1e50862167fd1d6a147717fdda811a75e36f58a38ae361

Observation 0db09d49-9aa0-466f-8fe7-43051d776125 · inbound

Aligning Multimodal Representations through an Information Bottleneck cites this paper.

Aligning Multimodal Representations through an Information Bottleneck Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:52.101762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:43:52.101762Z digest=sha256:0d45283035f563206bf020c974172beab2518e2228afe8fee27a1792a8e3755f

Observation be98b0ef-1f19-4649-a5ef-38fe0d7df519 · inbound

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval cites this paper.

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:45:06.069155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:45:06.069155Z digest=sha256:0c9cf62bf038d42e9a901ab43f5595668de55d1123fd5b58ee774ace244d8fbc

Observation 75e0a77a-2fa4-4472-ae3c-4ba1f8c1d869 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.310154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.310154Z digest=sha256:febfe3fbe8b25f002e543824dd8687b4b58ced79fe2deb9a1e4c54cd427aa0fd

Observation 3e104e80-3169-4f91-9d14-f6f87cd195d6 · inbound

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation cites this paper.

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:28:42.092582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:28:42.092582Z digest=sha256:4527e292c4fe0b2a19ed5e9d9ff7db0f3fe8e9fd7ebff0335c597575bef53268

Observation 664a590a-6d3d-4696-bbed-00748126c35f · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:29.198405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:29.198405Z digest=sha256:359ac32b5a609e410457291d05bdb715fdc730583c75ac61b4bed871afd0f894

Observation 8d894cb5-c695-425d-b294-db6419b67047 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.255620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:32e284c5307a4dc3148ab0b2d02d866428735a50db9cfe60e89badff31bbcd6d

Observation 8dac79d7-f1df-4eab-ab21-bca0d0161410 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.264059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T01:24:56.236650Z digest=sha256:79dcf3311e2c69bad149df3df7a2f2aa64d2625001317c7e554590ebc54cf782

Observation fbc8683d-1392-4ed2-b25a-0fbc605e5064 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.800091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T19:14:12.092478Z digest=sha256:3d8e51fb2b6c4952e95fffbe60634c81519b3fa7b53467fcc0f0f9fb2bfe76d9

Observation bb8ba82d-434e-4d52-96f4-801f7d839b1e · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.477044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T01:33:55.317107Z digest=sha256:1e751472b4678091cbbf9829449890cf6b05e4d8324c9203aee554af26f1036b

Observation 204c49ec-2a83-4a32-bec3-00acdb7e2416 · inbound

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning cites this paper.

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.543555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:24:22.361181Z digest=sha256:b4fbe6daf323c42b2661c096ddd950e6cc2e8199df919b4d5ff4be6e254074b5

Observation 0604384a-580b-4eea-be5d-350088ca2ec5 · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.083496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:4c7d5e77bdc223bab924c1881f0ce713c450121a1725faa4a1f041e5c6fbbd58

Observation ebe164c0-0317-4617-a594-1be920bbd7b6 · inbound

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models cites this paper.

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.081302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T06:56:00.782128Z digest=sha256:1cea9d51470a25f93d868aef5b9f750e0f0589fd1327b53ad6474cf64413fae5

Observation 05a2ecf6-785e-4c8d-93df-da99bba02c73 · inbound

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model cites this paper.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.196361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:1c0b3b250830b38231ce0573822f322716b5b337fa51165ff0c3d7f0a0d9cacb

Observation 638d44d3-5364-4180-b6bf-c4a6319ab41d · inbound

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization cites this paper.

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:43:47.409552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:43:47.409552Z digest=sha256:0a6cf25a60acbe0e5fb1fc2fc4b0191ca9be5aba002f8685555edb7a69950a51

Observation 00f0db15-6aa4-411f-a033-4e59e2061d5e · inbound

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs cites this paper.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.037233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.037233Z digest=sha256:2713a4ae7c8a4583615d92cfaf98af388fb993a9580c961b0c87cafa885d42ac