Pith. sign in

Paper Citation Record · LEDGER

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 13 inbound Pith citation observations for arXiv:2506.23115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23115 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:26.368534Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:30:14.527131Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:26:45.637665Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved43
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12acaddd-f71c-4817-9d01-a3f52b1362f7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.406024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.406024Z digest=sha256:f542eaf3bb100d8727ac6cdbede342639737734b7ee1fa6cea7821b97a37c2d1

Observation 82d96836-307e-4b52-9786-ff31030520a3 · outbound

This paper cites Qwen2.5-VL Technical Report.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.501578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.501578Z digest=sha256:3a8585294c565c017e045cad91dfcce5c9a4093b48d3c2bc8e2392ff742f84e8

Observation 909fba20-31a7-404e-aec4-813509e6093f · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.581303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.581303Z digest=sha256:29632084728a31eb0e68c2f72f9f52ac539f30b990648a027c3b81a27f4cc5fa

Observation 435d780e-2ef9-4dae-8ff7-148ce382aeef · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.664520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.664520Z digest=sha256:d1d3d889ccf162ee16b0b4bbb20d85753274c6370a64e4ce85555eae0a9f5786

Observation 8ef96816-7451-403d-b809-34d0de0c171d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.757887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.757887Z digest=sha256:d702d6a75275dba029c00ef7e2c82984cd63cf13e02f1f5444e65e167dcada71

Observation 6bb383f2-3ef2-4592-b394-8b50ba4b9ab0 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.862210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.862210Z digest=sha256:2fe56a5944d799d603e47d104d5dd0bc5fa0557a93054928053b55a603ee47fd

Observation d6831270-7664-4226-a823-c87311c8f726 · outbound

This paper cites UNITER: universal image-text representation learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings UNITER: universal image-text representation learning

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:21.927001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.927001Z digest=sha256:6263e1bae143ed1b10e9c5a2bde42c2f1cdf40b613a77a99f6faab323586d3c1

Observation ba50800f-6021-4c9a-b239-f00b9671e97f · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Reproducible scaling laws for contrastive language-image learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.042338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.042338Z digest=sha256:f3e7383e34ff46d9eb68b528cafd17dbe61ec15c067478f71da93cbb651291e5

Observation 74616016-8654-4317-99d1-9ff2b53422f4 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings BERT: pre-training of deep bidirectional transformers for language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.802175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:22.118312Z digest=sha256:250a29a950d632722b7003e547e0fb525cb6f9024d9af38d0477d3633e2e8747

Observation 8d2775c9-e589-4a2c-9b28-fad78c05067b · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Colpali: Efficient document retrieval with vision language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.661796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:22.378707Z digest=sha256:179ff19393d83b2b97ffd4fdee2787ad30fcb1ac979975a4865dd64540e95d9a

Observation 2cb5873b-4130-4d15-a942-5302520d6146 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.484484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.484484Z digest=sha256:05a833a95618c7959c1d0bbc24f483c885b4e1d7e85751a567a141e98185964c

Observation b1c4f88b-a3a4-4dea-b468-f8ca7bdfed86 · outbound

This paper cites Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.602494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.602494Z digest=sha256:84a527d649369b61279bc2fb23aefbc075bd01279e64a3fb4af1fda4b6067fbd

Observation de573936-c557-4175-a5a9-a506d29443bb · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.687226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.687226Z digest=sha256:8556f5a171444c859268ebc0385703a85b2dcc2d9ff79298e32f7f497ade51a6

Observation e154328b-0d4a-4b0d-b4bc-2725ca9bb854 · outbound

This paper cites Girshick.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Girshick

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.796886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.796886Z digest=sha256:61dcddc0e3a5cbc5c6d32db67d77438d7eff5a3bb50d4ccf22ec56703987fc9f

Observation a6dbd464-fc12-49f0-9ff6-258cd273f447 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.908229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.908229Z digest=sha256:2714e385ba723b6ce1870319eb4f1f31e54f059d1036a4ebb2a07512456c3089

Observation f8732002-6067-48ac-8ef8-632b373dee34 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.138389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.138389Z digest=sha256:03caaa9724ddab5309fd18907fecf1a658ea9f9240459bedca401f746b98069b

Observation 67d2c4ad-3ef7-497a-9315-4477152a8d1f · outbound

This paper cites Continual pre-training of language models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Continual pre-training of language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.199431Z digest=sha256:d8bb725490a2ecfe589de7bda6db4787aedd2116fabff4d355f48ac47a814617

Observation 54ae3e34-60d7-4074-a352-4bab87dc0a52 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Vilt: Vision-and-language transformer without convolution or region supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.295768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291964Z digest=sha256:1bf30f2e88ee4360fc111ce4b06262af5f6a883b575b04c34ec96f5a4cf5ba9a

Observation 43747104-f690-45a4-bb11-b893ba5279eb · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Llave: Large language and vision embedding models with hardness-weighted contrastive learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.463875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.463875Z digest=sha256:913941b8e8680367456cd218b511d9b895a7e7cde8f7a08a8532623283c85e72

Observation 2021fc53-9986-495d-b081-c67fdc160c3b · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673554Z digest=sha256:8e0e750c4a5e3be34c10460a64efa31c02881d83469f5d81173bf5fa5eacd25a

Observation 2cb4484a-1647-4fc6-96ed-38e9d8b51b7c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.115856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.382898Z digest=sha256:70f566a181aa591f5a816e65ca9fbf4fc57bcfe1f9a9ae3596b794d9e43ac296

Observation e88daa49-3aaa-4164-9451-0c643901b11c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.815446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942267Z digest=sha256:04bba6be455dd76f507fce492229be5537bf8a8325b0f7b5aef31bd576884f28

Observation 2fc0f898-7539-423c-9f07-a3f2c902346f · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.616930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.064803Z digest=sha256:9338d3dcdc4655e0fa9197568c8d1fac5ecafb07b78309aefd5f5ef88f32c2ad

Observation 27326480-7e37-4820-8117-a9e3fe2741b6 · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:29.440599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.181614Z digest=sha256:22539d457ad11cb46caeb017f101401233581a1339c4ab0b6e7252ff8445e37a

Observation a198e8bd-8db9-432a-8644-fe2d8b5e600b · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.269954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.269954Z digest=sha256:21759233878a665aca5a982843a146b0a827a3bd3fda6939cb0ba87d778223f9

Observation 467cab76-10a2-4460-b58f-003ed0a09128 · outbound

This paper cites Nv-embed: Improved techniques for training llms as generalist embedding models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Nv-embed: Improved techniques for training llms as generalist embedding models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.950680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851285Z digest=sha256:e416942e3af3a26c50e98ee308bc43d6753445187dd4ca7018b524a6e22f9699

Observation 04a6f182-78a7-4c78-bd37-b931664633a0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Improved baselines with visual instruction tuning, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.483412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.483412Z digest=sha256:bb52fe4a3ca513ac24d293cd7c1568cb103b922faeec98c7fd1562bd315f3382

Observation 40dacb9d-03dc-4323-9b56-d3b19b99167b · outbound

This paper cites Unify- ing multimodal retrieval via document screenshot embedding.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unify- ing multimodal retrieval via document screenshot embedding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.257243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.556343Z digest=sha256:58445e79bc8b77ba1814b61775d72ee4b833902d03cba1bbc38b8122702642a4

Observation de5e1ea6-d8e2-4010-b025-6c62838ceed3 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.627791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.627791Z digest=sha256:0cc2df8f957aa606a29b87d18b78fb74cbb86253cebaab55dd4d497ccdce5293

Observation e4a46cba-2329-4e31-91f6-92d767521cf4 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, April 2024.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Introducing meta llama 3: The most capable openly available llm to date, April 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.007803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672117Z digest=sha256:b7e2c9330fc69b9cb25e19f037b5d1111d233be4ea465a6095fca28dcddf6f8c

Observation 27f972d3-bcaf-4402-ac69-6c91923ce885 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.721367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.721367Z digest=sha256:73e3e8bcbd5e98d0a2622432e89539f442bcc9514e3d21bbd397b4226c9c56b4

Observation b54f5c2c-b25a-4369-8aeb-aafe3724a387 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.419284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.419284Z digest=sha256:38c3377d23b9ca6ee2633214ec0ce8ceeb9d247b37e5f9d3754c479cd07378ea

Observation 33f331cd-9b92-4920-8f2d-4923efaa5809 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.853435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.853435Z digest=sha256:9008a7419ca1ed9da2f8df64e31a86eb4e4665b38e686410ba86d2becc656d85

Observation 4aaaa547-1faf-4365-b5ba-e27ee147e874 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893267Z digest=sha256:ae2fe6b2ce316ceef36d0c1097db6ee5ea6eccaf34933819b2b0764bc69f96fc

Observation 9a1d01eb-cbd1-4ecf-bc7a-68583e134752 · outbound

This paper cites Repetition improves language model embeddings.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Repetition improves language model embeddings

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.720479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962392Z digest=sha256:5490e586f1aeb29fd711b6bce5d6b9ea4bfd2241af23f936e556e3507ad536a8

Observation c75817ac-a642-436b-97f9-5059288eb293 · outbound

This paper cites LXMERT: learning cross-modality encoder representations from transformers.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings LXMERT: learning cross-modality encoder representations from transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.497621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.018644Z digest=sha256:6bb7ec124dee405184fbecc78e52234302a61fda0db9d3d1d09f174067acdd7e

Observation 86f98179-3dd5-419a-85b2-f97c92455e12 · outbound

This paper cites BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.118584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.118584Z digest=sha256:8759857c617c27b10f4465cc36a7b1a0b9fca952b325661052d5d2971ea973bf

Observation 5bafc57e-5de7-4b00-b17c-64d39f540fec · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.800410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.800410Z digest=sha256:96173bdf2b67027d00fd340c1a6714a656694fbce30e3d114a6f89737171a65f

Observation 0f2bbf62-cfd6-4313-b2fd-2552ff1808fc · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.252459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.252459Z digest=sha256:5735dae6b6fa57d805653d76c63a70584f7295c378498a8a9d3efb133ea6e7da

Observation 17a275a2-6222-4a70-afc4-b1203795b1cb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.314513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.314513Z digest=sha256:f1c88c7b84c49179dc309df9fabab6c516a70cd6ac68a9e6e3258cc66389b7bd

Observation 751f62b2-c372-41a7-aa00-ceaba5a439fe · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.380540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.380540Z digest=sha256:f1343c760752d61969298a1db4f2c91d49a9906f5b8c798f8d559af547457723

Observation 1f4645f2-7322-4014-8f1c-438cecd41c9a · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Uniir: Training and benchmarking universal multimodal information retrievers

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:25.446061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.446061Z digest=sha256:b77f46b87d184eae333652c1528b00897d2b9858e45521319ac5fca109e99c19

Observation c1c72e75-c1c5-4c1c-a89c-8b1a88ab05f9 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.514246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.514246Z digest=sha256:c0c382e5b162e92c5bc45667f95c8d22b7a29fdd9954cc0ec5ac4c34dd03fb29

Observation 4a860906-83a2-4e4b-b708-f105f1a163e1 · outbound

This paper cites C-pack: Packed resources for general chinese embeddings.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings C-pack: Packed resources for general chinese embeddings

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.582577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.582577Z digest=sha256:458120a75c86d63025bbf5d3702658a87fd50547fde05d57bfe908a9e916cb54

Observation cb2ad355-3c09-430c-a8e1-a6584a9659b6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.203463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.203463Z digest=sha256:c8d90d25ff986cbbabd107e9c8b4da546568a05dae1b038d5564391fc47faa15

Observation a21ad657-6990-46c3-84f9-2d33bc8c9f1b · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702492Z digest=sha256:29a8353d688b59de0698813d480b78f98bb067473a9373a54c9055cb0f25976d

Observation 5bcfa3ea-7f46-49e2-a258-dc5fa0a2f259 · outbound

This paper cites Sigmoid loss for language image pre-training.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Sigmoid loss for language image pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.706693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.706693Z digest=sha256:8c5c25bec94150b398c637ea0508057cb3470ae4532ef62ee89e5d269cae60bf

Observation e487b7ce-63d3-43cf-9f82-aaed71d2ba84 · outbound

This paper cites Magiclens: Self-supervised image retrieval with open-ended instructions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Magiclens: Self-supervised image retrieval with open-ended instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.972882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.778184Z digest=sha256:68b9b837685324ae17b1b06cdb3986d57f663151c5a4b73f4aba4c83d9366b0d

Observation 38281db6-a1d0-4a2c-9c01-50fb9d8760fe · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.999631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.999631Z digest=sha256:751be33efcf40e68f7895dc6a22e36b3d331c6b4f0fbae8708dba93e45d16bcb

Observation e389c8ad-e413-4311-96c6-dedb6e571120 · outbound

This paper cites QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.204540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.204540Z digest=sha256:6f01dc7d38a3cfe1150aa42745246e7e44843fe160de798aaa94ad70748c8d83

Observation ea48ff13-bf4e-40f4-a2a4-63577caebc06 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:26.368534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.368534Z digest=sha256:83b5e42487e0a40ed5973ea3619a253ab1882fc096a76266d8848661bf79f060

Observation 2f62bcb9-a8b1-4ee3-ba58-d06d0b7c1ae4 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings An image is worth 32 tokens for reconstruction and generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:28.277738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635632Z digest=sha256:b34a73e3213637b2fa80eb7cc9b219f7820ab72ef5205ff00cef792937bc7b6b

Observation 567955fb-40b2-4acc-930a-47ea0fb70693 · outbound

This paper cites URL https://doi.org/10.18653/v1/D19-1514.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings URL https://doi.org/10.18653/v1/D19-1514

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.064163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.064163Z digest=sha256:46a061c2582ad026b98053b33e28f5bc102874003db40b0e2dd12fe172766b42

Observation ecbc4db1-4b25-4932-81d6-0100114c236c · outbound

This paper cites an unresolved cited work.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.026072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.026072Z digest=sha256:6803bd65b6209afd04105feff3fa4f9593caa7a11adc9e7a3bb14f5fea82ea0d

Observation 79d78f86-b5d3-4f1a-9513-935d111541e7 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.355520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.355520Z digest=sha256:7abba702bf3d5a7f8bdbf792d80a9cf3b9915cbfb7a454fc82db18c4b5eb090c

Observation de154433-0b74-42e1-91f1-83e5f4c6ba35 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743202Z digest=sha256:769271da51d189885bd8a68bc47230be969c9fa760e597cbe3bf5ff980eba0bc

Observation e3d2d85c-1d26-4dfb-95ab-6f1377d6f745 · outbound

This paper cites URL https://doi.org/10.48550/arXiv.2503.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings URL https://doi.org/10.48550/arXiv.2503

Reference 2025

Resolution
verified exact
doi, observed 2026-08-06T21:52:27.096202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579127Z digest=sha256:49058f8813c2ab86aaa5dfbd71f48691276f5cc65ad518454812b4b91bb85c67

Observation 338e1be0-732f-4800-ad45-472d56d5bf8a · outbound

This paper cites doi: 10.18653/V1/N19-1423.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings doi: 10.18653/V1/N19-1423

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.242467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.242467Z digest=sha256:544d3b94afb968d8a0d404e3b2e7e503537abbff59f0556803014584db0dc3b0

Pith citing papers

Observation 4040c3b5-7f43-4da2-b596-f0cd0b882796 · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.483851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:20b3503f3b3b3f981590ddbe284f5415ccf98473aad77187379a075cbbd32728

Observation 12d175b6-6b6f-42b9-9584-86eab0bb3b58 · inbound

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories cites this paper.

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:32.688658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:32.688658Z digest=sha256:b1202e102014e6c8de3c831365bb1e2e30b6eee9fd138877f14f65a96d9c8419

Observation 4604f3a0-ff6b-4c15-b94e-32d5feffcd74 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:31.172392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:31.172392Z digest=sha256:e0cf3b5e140d302a09c42c2b6b2701ed2c25e9b953630584c457d16b556468a4

Observation 2346178e-b092-41c7-a44a-ce6802d66eb7 · inbound

PLUME: Latent Reasoning Based Universal Multimodal Embedding cites this paper.

PLUME: Latent Reasoning Based Universal Multimodal Embedding MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.962706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:48:40.722921Z digest=sha256:7c470745621678916183a15a4c86a7bdadfc7ff5cd643f865e08cf46dc90bc4a

Observation 155bb17e-45e4-4b13-bf9e-465976c22e96 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:19.810359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:46:30.827346Z digest=sha256:ab2322b7013ccc3d8330c2b0db2618788c8a111e57ab7edbf28aa7b11e348d39

Observation 3f1d2f26-f696-4c90-9b78-03891cb2e354 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T18:31:37.144968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:31:37.144968Z digest=sha256:b496938cee26f3891bae83bdd9dd2233ba2f5675ce2769127d6c384762cb4f48

Observation 6e2d8627-e61f-471e-8fb1-880581564494 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T15:40:29.103187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:40:29.103187Z digest=sha256:6b9a58bbb8fa074bc4156f8a917e95f6b9c1cdf57f66ab7235d129459aeb1311

Observation 4d7edc2b-ef31-4c3e-b2c7-0e5a977e11f5 · inbound

MINER: Mining Multimodal Internal Representation for Efficient Retrieval cites this paper.

MINER: Mining Multimodal Internal Representation for Efficient Retrieval MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.555032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:40:33.437364Z digest=sha256:64ce149d2aa48755579194161090ab00e9babbdecde75c5bc3c798dcf5d47790

Observation df16cd08-1747-4e39-93d3-9260da89c2f9 · inbound

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation cites this paper.

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:21:16.786986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T08:21:06.959344Z digest=sha256:9c486516887cdf4c205335145269804835bbec4e0acd0d2edf8ac711be5149c3

Observation 175ad642-69e0-43ee-9928-e0b288c9fc5a · inbound

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation cites this paper.

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:59.172275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:45:43.932116Z digest=sha256:99e1b74283e3c5b967b42bc1fbd5a95665ea7fa20924114f18b82322245dcd27

Observation 7ef45a18-3a16-4182-b019-468cb531e09f · inbound

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini cites this paper.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.317846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:2562b068185359ce44b2a4f3728e75f7c420093bdb393bd31da627ed076a90c2

Observation f86d2da2-9aa9-42e3-b9be-d9577ab89c77 · inbound

MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework cites this paper.

MM-Matryoshka: Towards Budget-Elastic Visual Document Retrieval via a 2D Multimodal Matryoshka Training Framework MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:45.639136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T06:59:06.801340Z digest=sha256:3dc7a3c7f138176afcbdb411c9d8deb11904038c7f1552eea941fb96fbb60d36

Observation 7f1dddab-eb44-4019-8c3f-db21b1a3070f · inbound

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval cites this paper.

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:30:14.527131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:30:14.527131Z digest=sha256:4b0d0506d4111fa7302a6cc0ab55f72c54addb34769cc1e0affbb5da8daf69ee