Pith. sign in

Paper Citation Record · LEDGER

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.20204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20204 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:54.587472Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 27933465-9277-4982-a242-568cca4b7967 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:54.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:54.587472Z digest=sha256:7488698f1fda525b0b91ed313fcf7cd27572aa24dceddcd97384ed6f6ed13970

Observation 8f096dad-c98c-4e6c-b3e8-2e41080854c1 · inbound

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning cites this paper.

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:24.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:24.201734Z digest=sha256:0fead5e9e245808c9d3865f2e68360854f4958ccd5e0b793e07b933474c2cbdc

Observation a17f97fd-89cf-4b74-9030-2c605e030027 · inbound

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval cites this paper.

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:02.482363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:02.482363Z digest=sha256:af94b32d89ec0a10aafabbbda2ea59a29e57554ff2d4ba4336c3c1540823fa51

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · inbound

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning cites this paper.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:cf6f8bfd4285179fefec63f7e1d9c26459ef8d6673554b323af202be81ce8c95

Observation 34f2df9d-6b4b-40db-baf5-a60c3f7a1609 · inbound

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL cites this paper.

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:00:58.367741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:51:16.675856Z digest=sha256:36ac374699e8dcf8561da43723a60ccbb4e1f58dc00cad685bf0732738bb8586

Observation 7272083d-76f2-4baa-a4ee-b92f93047d00 · inbound

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment cites this paper.

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:58.336606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:28:59.838565Z digest=sha256:3d0524cf3a0494f7cc74bf51708d3bdcda648061a64229d7d1faf4b1ec6b1334

Observation d4ea2ad9-f244-488b-a5b1-93850ff9909e · inbound

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models cites this paper.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.685664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:53a684e12f07cd53b4cde3588e73de068b6bbe117c5c9b447c26e22568928363

Observation 99934a97-6fed-4880-9dca-26f56211ef13 · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.071233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:4c3fa90211f75a32d5e3fe1be11b91a2b89746db307d5cbef7b5059f0b882f42

Observation 0704295e-41db-4e75-98ba-6c0ec752b223 · inbound

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations cites this paper.

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-08T19:09:02.722999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T19:04:40.817413Z digest=sha256:dbc3e1a5d8608e0ad644dc659c4cc3b0569deaec91df5edb9b8a71d16cfe55f6

Observation 5c59e35f-0378-40f6-96f4-751a7f31a020 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.867871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:4e8e02b30b57792f67666e35c9a9ae4f7bc849bb98c74a140f9c6fb1aac7faca

Observation 041fe031-0521-4a9b-9763-085fada59064 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.816288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:33337fe7962036aa1868076ee73f8f9bcc14c753eec91dc98e7fe34bbc7fd703

Observation 55ed93fa-6ac9-4c67-b0a2-a885ab380a75 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.700545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:9783e135f7fd2f2c71e04e0fa0553f6d5b579cfd78360a455487326ad7b77776

Observation 3b2932d2-1ebd-4140-9973-cb608e17bd23 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.462858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:a51efcb1ed346bc9441eb4313fc629ff5e224496f94296de4cd1f9d88f2bef64

Observation af075db8-671f-4551-9335-c014c5432b89 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.782048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:b0165a6d851088414d4192857350a3191a821bf76fa5cf54cf9228c081ac9b54