Pith. sign in

Paper Citation Record · LEDGER

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.20204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20204 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:54.587472Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 27933465-9277-4982-a242-568cca4b7967 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:54.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:54.587472Z digest=sha256:7488698f1fda525b0b91ed313fcf7cd27572aa24dceddcd97384ed6f6ed13970

Observation 8f096dad-c98c-4e6c-b3e8-2e41080854c1 · inbound

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning cites this paper.

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:24.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:24.201734Z digest=sha256:1594f4470c1961578ee32709a23b78fd33388c4a182c7f587388ac49f275565d

Observation a17f97fd-89cf-4b74-9030-2c605e030027 · inbound

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval cites this paper.

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:02.482363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:02.482363Z digest=sha256:af94b32d89ec0a10aafabbbda2ea59a29e57554ff2d4ba4336c3c1540823fa51

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · inbound

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning cites this paper.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:cf6f8bfd4285179fefec63f7e1d9c26459ef8d6673554b323af202be81ce8c95

Observation 34f2df9d-6b4b-40db-baf5-a60c3f7a1609 · inbound

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL cites this paper.

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:00:58.367741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:51:16.675856Z digest=sha256:67e063dc930f5ae4daafce50b2c7af6274d949f82427e109356027146d3949e5

Observation 7272083d-76f2-4baa-a4ee-b92f93047d00 · inbound

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment cites this paper.

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:58.336606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:28:59.838565Z digest=sha256:838451692e23b6926487eab7f993c15f108a3fa6187d8c54a3153ed8c7bfdea9

Observation d4ea2ad9-f244-488b-a5b1-93850ff9909e · inbound

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models cites this paper.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.685664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:d70ff5b3b11648d669d2d6bc42a13e4993298cc9b597bccedd3dc525cf33d1ca

Observation 99934a97-6fed-4880-9dca-26f56211ef13 · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.071233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:ab0ecbc131c66b5e66c75591b8878ed5e3a2fa08caba9e87018b5128bcd80d78

Observation 0704295e-41db-4e75-98ba-6c0ec752b223 · inbound

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations cites this paper.

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-08T19:09:02.722999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T19:04:40.817413Z digest=sha256:a2c629a1d74c69d57d9ec6fb910153b9f05601c6a8bd73464abcb75fdb24c9ad

Observation 5c59e35f-0378-40f6-96f4-751a7f31a020 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.867871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:71e656a1bd1591457b0460cf85b4900442419d79249f6cae510cbf6f1a34a410

Observation 041fe031-0521-4a9b-9763-085fada59064 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.816288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:bd206413d67b21f299f3b24af92bf9a5ccc1b08753f50e9069f37d363720c61a

Observation 55ed93fa-6ac9-4c67-b0a2-a885ab380a75 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.700545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:043571f52c699854d8997bac3af835ea5a89a8d00c086efb82ad7aa2e6d7441c

Observation 3b2932d2-1ebd-4140-9973-cb608e17bd23 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.462858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:92d87d93f12fd649fb68606ddc5ad064143d9634ecc3168062433c274d1036b8

Observation af075db8-671f-4551-9335-c014c5432b89 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.782048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:1d14cebb12bb18f83f479e350403d7818c64a19d9a9d985949b4150022ccf7cd