Pith. sign in

Paper Citation Record · LEDGER

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2412.08802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08802 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:20:01.029650Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.771573Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a37951d5-4c65-4996-8dbf-85523f262b56 · inbound

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation cites this paper.

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T23:20:01.029650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:20:01.029650Z digest=sha256:38a688d43493d5ea096113fd14a1e50130c13b935693c0662285383a7e776fff

Observation 4c233abe-685d-4479-aed1-9faae35b6db7 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.939130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:8dd028ae2efb80f4523cbf0172dbff61b04ebd10fd00a175f50cb66c62200a8e

Observation f5186e79-b65c-4286-91ea-410d98f875a5 · inbound

HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval cites this paper.

HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:02.585194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:22:21.682534Z digest=sha256:81974928bf0f0717f108a4dd77cbd5a1e877c58c8c59b1bb8f07a4acbf64ec7a

Observation e6c5de1f-0bdd-4e19-893a-9e758e94f51b · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.838737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:599892ef258971e7b6879ca41b161494a4f1107e5a31aa66ebdca63836b85978

Observation 8b86ce07-2594-4b42-96d6-f4e48c6c4bd3 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.779441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:b9389ee88426f0f679e712f88be30ed5eef7df50bf148b0776ef82c6ff9d9ee3

Observation fb41aea0-70b2-483a-a861-eaea954a8b99 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.703034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:40f5a85aa817808eadc5e1c69cb3d46255040a6a5dbff17fb2ff03bc1144e010

Observation f1cf3dec-3d74-43a2-8536-ce16658b0faf · inbound

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset cites this paper.

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:23:58.441043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:21:18.369534Z digest=sha256:945159f9b7e0d4e37bdd378f7657bde7839c5fcd506425204923848fcdcdc511

Observation 9f548490-1d6f-4406-8dfd-aa5baf953dfb · inbound

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling cites this paper.

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.771000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:57:15.988771Z digest=sha256:639ac8d6ec5a56a1eee614acbe0649e88bb079abe41eb4a99ff05693cd3acfab

Observation 20b3544d-2409-4544-8762-e46faac2b04a · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:07.773095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:d35b0db80ca11b3c0cfe764dc24e7ca2c6c5603c3abf4e77b00b2f09c1c31b33

Observation 67de105b-4d74-4a96-86cb-b0b461385c3a · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.253079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:d9d117c99898ccd964e0a15373efdccadef451c372fd3c47fa0d3853d29ffa74

Observation 994aec71-e745-4473-8b9b-2de9c6dc08df · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.719907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:6b2d27e625f83395d831ba1ca83f7c1d87907a21a16c4f4eec202a699deea2d0

Observation 5530c45c-12d9-46d5-bb05-c6171f617c16 · inbound

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval cites this paper.

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:09.450205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:09.450205Z digest=sha256:5cba4cb20bdc3a1c87996bbba2078cac809d455f6d61fc98632cb0bc4e35f09e

Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · inbound

Illuminating Visual Identity in Universal Multimodal Embeddings cites this paper.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.254863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.254863Z digest=sha256:327b9325701a2c25374537ee8d0b0bd90eebbeb351a45ca31dcf22c95c86c407