Pith. sign in

Paper Citation Record · LEDGER

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

As of 23 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 12 inbound Pith citation observations for arXiv:2505.19650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19650 v2

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:46.090733Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:30:14.559722Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:49:36.602258Z

Reference resolution

100 of 120 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d25bdf05-a782-4319-83d6-0091c3fc8af7 · outbound

This paper cites Multimodal automated fact-checking: A survey.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Multimodal automated fact-checking: A survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.135801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.135801Z digest=sha256:dfb75eafc41295253bcd24b9c55561606d96473ee55ea4f46e3951b34949b5a0

Observation 62a061c5-05f1-48aa-9003-4f4ff0a5f489 · outbound

This paper cites Localizing moments in video with natural language.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Localizing moments in video with natural language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.194753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.194753Z digest=sha256:aafc21bb25b392204dbcb94216e91025164b84a96c3e1d681a3b618269e3ceb4

Observation 1880f482-8f26-4197-a149-8fc3aa7ba92d · outbound

This paper cites Vqa: Visual question answering.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.236558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.236558Z digest=sha256:30bde44eebfac835823bd28a7610926eeb6ce53f46a31974f409f7db28bd9d45

Observation 613f3f9a-f2c5-4326-be34-a86460cd5d22 · outbound

This paper cites FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.333584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.333584Z digest=sha256:f5c6e5119ddddc475c4b9e9174f23160a24091988589d02f4601d1c9f5dc0321

Observation 47fa31f2-a075-4f07-b00d-96c42d4efb7c · outbound

This paper cites Qwen2.5-VL Technical Report.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.445117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.445117Z digest=sha256:db97394e4d1684234acfb76b7c42e651f852b6952440b39175a1fec6d76adad9

Observation b81a194e-645b-4e3f-a99d-899caf53458c · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.507574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.507574Z digest=sha256:0c8abfd4e781861b99b8b2987b1fb9b8266d12ceaa307ded6c896a60d53036f2

Observation 33c681ea-322c-4d07-9570-bf88a7f1e689 · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Wiki-llava: Hierarchical retrieval-augmented generation for multimodal llms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.583688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.583688Z digest=sha256:bd2d21acc7f045c85be2dbc3c086deae29ee02388367e9d9f7d41e0de46ea51d

Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.686789Z digest=sha256:610263d19dba0581f0f84d4940b0fb09005c097ef76072d57ee8b033f768de4a

Observation 8e2aef62-91dd-4624-9d83-0869ab6c18e7 · outbound

This paper cites Webqa: Multihop and multimodal qa.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Webqa: Multihop and multimodal qa

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.792978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.792978Z digest=sha256:d6588ec4d8afe041c9ea0a3d6b4a96c40edbac2ba338eff3fb1767a3a1a465a5

Observation a17a1232-f62d-4dfb-8d2d-06f2228c2af9 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Collecting highly parallel data for paraphrase evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.877735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.877735Z digest=sha256:5813a45af34617ad6f4623b9caaa7e9376e56880bba8d6b67a92dca18a19fcdb

Observation 14038995-ac3d-40b0-9f9d-34599cad0315 · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.978957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.978957Z digest=sha256:32778d8ee00156f56611b79094d84e1af34844c5131e6087faecd7db7f021c62

Observation 5ce443e8-645b-4c1c-864b-3fc15153c6aa · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sharegpt4v: Improving large multi-modal models with better captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.039664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.039664Z digest=sha256:68f51851cac967153693be7b4516bdeb3933835427f3d8e6c6b94a39e4ed0a76

Observation cb626853-7142-459a-bb3f-07ee6af82e55 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.135154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.135154Z digest=sha256:f66406eceedf6f39e4f422a0e5f4c9115bb8384ea099b918abf714af2ac7dc90

Observation 3abfb845-1d29-48d6-909d-845825214698 · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.241872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.241872Z digest=sha256:bd615434e05f11e9a7f5d71d24d8df3faea0a2f17e2f0ceb3a39cbf068c82b5c

Observation 2ff154db-8bbf-46bf-91b2-fba69349d5dd · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.335797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.335797Z digest=sha256:4a0cad8a3fc63b6dbe7af1c79035bfdec554ee184541ff473dbdec2ad64a547b

Observation 55ade486-db11-43f2-b8c3-e76b36b8c472 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Reproducible scaling laws for contrastive language-image learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.353327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.353327Z digest=sha256:09d2ca851ffc8345de1a3f38164b51667e95bcb530a2010d57a02d6365a60828

Observation ce0df1fa-2f49-48a6-be40-6c7c99cc4813 · outbound

This paper cites Visual dialog.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual dialog

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.487073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.487073Z digest=sha256:190dce6bcd1d5e387396fabd4c67e812b504aadee059e10bdbfa887b2a9609e4

Observation e03ce079-de1e-42e5-bb72-41a602d79443 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.611355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.611355Z digest=sha256:1732bfa3b738c0b335c3ee48a21a9ae61eef1c2ccffc71e0abc7d161548687bc

Observation 20feab99-8379-442e-816a-dc2af501376d · outbound

This paper cites Mmdocir: Benchmarking multi-modal retrieval for long documents.arXiv preprint arXiv:2501.08828,.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mmdocir: Benchmarking multi-modal retrieval for long documents.arXiv preprint arXiv:2501.08828,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.724894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.724894Z digest=sha256:7035fb8fe35e12acf1b4e38fd727ab809589d7ed97f1011b134d5fbe6af6ba96

Observation 6b9a2408-e4ac-48f3-913d-a09e84c5684c · outbound

This paper cites The pascal visual object classes challenge: A retrospective.International Journal of Computer Vision, 111:98–136, 2015.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The pascal visual object classes challenge: A retrospective.International Journal of Computer Vision, 111:98–136, 2015

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.857313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.857313Z digest=sha256:5282301c808b604dad14520b21ed631465f9ee2cd2b6319b5998267316d9c3d9

Observation a452dfd7-f6d5-419d-b6bb-18c1ba7ec49e · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:38.993221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:38.993221Z digest=sha256:9040f71a06b3267ee7ccd8d668d02d7254635a408ee03a10fe39593ef2dfb068

Observation 56c6cc8f-ded3-4f71-a5b4-ad3cb8f4bb2b · outbound

This paper cites Simcse: Simple contrastive learning of sentence embeddings.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Simcse: Simple contrastive learning of sentence embeddings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.124769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.124769Z digest=sha256:812a34d9c093c666ab7294430ad9da6d97b2b73d625a6acaf34b4c2c4720856a

Observation 7f0689e4-5de1-42e4-b919-de06d58519da · outbound

This paper cites Imagebind: One embedding space to bind them all.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Imagebind: One embedding space to bind them all

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.289062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.289062Z digest=sha256:8c23d5c85bbfc2caf09eba0e730aade8d7db7c67f0ee84b97bdc2d4e4455e0b1

Observation 9ada8643-244a-4f40-9648-261e9c3c503a · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.416668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.416668Z digest=sha256:ef6259a44bd8c84b596c57914f107a81842bee15f8feb73d1c39d35e23ffc448

Observation e76bd2ca-1ea7-42db-a61f-afaabef2ee8c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Lora: Low-rank adaptation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.525902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.525902Z digest=sha256:7a4cfe468a213e8b76821e98fbcba3c6e83172ecb3cacf7b211ce0e88ee3bb3b

Observation d97d084b-ec3b-48e1-9ac2-4f3a6cb753ed · outbound

This paper cites Egocvr: An egocentric benchmark for fine-grained composed video retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Egocvr: An egocentric benchmark for fine-grained composed video retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.674376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.674376Z digest=sha256:793fa5a89619412794028589d2e8ac07337091fecee83595c097ca60818ff8ac

Observation adf2eed1-d342-476f-b3c2-12fcc2223410 · outbound

This paper cites Mate: Meet at the embedding-connecting images with long texts.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mate: Meet at the embedding-connecting images with long texts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.787447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.787447Z digest=sha256:eb8b4cc4d3411e2bfb730aef29b73e1f413f705a7dfe4e274e8bef7eff5480bf

Observation 1932084e-9cc0-4812-8e24-96947c515a18 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling up visual and vision-language representation learning with noisy text supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:39.943808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:39.943808Z digest=sha256:36644cde530cd2842dcd182d7ac8c3cf39197457ab549ee13ef0b7a5957291be

Observation d567483b-84af-428a-9f33-3baca666fa76 · outbound

This paper cites Tencent text- video retrieval: hierarchical cross-modal interactions with multi-level representations.IEEE Access, 2022.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tencent text- video retrieval: hierarchical cross-modal interactions with multi-level representations.IEEE Access, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.007625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.007625Z digest=sha256:501121df423ba9633e36e6f253c1382c8694b97621e6375478105b58123a5020

Observation 69ea240f-4921-45fd-bd9a-8d749f023325 · outbound

This paper cites Scaling sentence embeddings with large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Scaling sentence embeddings with large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.062977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.062977Z digest=sha256:140d5d1831e51d42c4a5adce0aa0b94257518dd0973570ff44e142ce9d95a7b4

Observation 7c11c69d-5b64-4c24-bd72-49451fd8d076 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.135649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.135649Z digest=sha256:ded000f5ebd81bfd7de0d7c1134ae27b1e210d55fcc4df3806afff3fcdb040ff

Observation ac76a1ed-3331-45f8-803f-635eead6ea98 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.233513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.233513Z digest=sha256:18c94c431285b115b0c1741b84015f4973e4155ba79d01b8b467d75f63aff1c8

Observation 67ea4878-8040-4681-b778-07b3235699d1 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.292912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.292912Z digest=sha256:cae4239780a2d7dabeb9a922ca7cb16c2c9c37b4ed8b556281d86779341637c5

Observation 92d103e5-e08d-4319-b296-89169c5ffde4 · outbound

This paper cites Visual question answering: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual question answering: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.422494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.422494Z digest=sha256:785c19b312da36ed0b8049c02f90051cb5dd675ed79242a3f8908876191f3b91

Observation 291c51a4-7d5c-46bf-b507-49027c432b08 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Deep visual-semantic alignments for generating image descriptions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.496653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.496653Z digest=sha256:9f6e3109f64e914441247403456c2fe6dd36a681cc5765bc25eebcbb55105866

Observation 0811a0e4-3d0c-4eb8-be38-e8d78780461e · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.596642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.596642Z digest=sha256:3b990a5af143e1c0b232ff8b88ae1d0fffe1c64d5321aa7574d91b567a591c17

Observation e0a6ae9e-06f8-4ccb-b4df-16bb641e9b7d · outbound

This paper cites TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.649758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.649758Z digest=sha256:9bf4b6e093d5135a5d494bfaa7f6628ffb4a286855805ef7358f9aadd5e80b7f

Observation 6e1e8e44-c342-4aad-a0f8-f9dd659c04b4 · outbound

This paper cites Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.710060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.710060Z digest=sha256:00cfafbdbb16a03a3e8d3f8770df5a85476cd7ab1c0887bb65b9ccf37ef9267e

Observation 28885696-1d45-4d01-9b13-455c7adb32cc · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.831458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.831458Z digest=sha256:a892017ccfc6eb773dcabb784fa4f379f7c54171df533385fbd4744f5d100757

Observation 655aff41-ce43-41c3-8937-eba3fd1e8898 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LLaVA-OneVision: Easy Visual Task Transfer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:40.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:40.926888Z digest=sha256:3ab968774e11f1bc819615018d9bf5b14bfde06d02fb1fa20cee45ac7d19febf

Observation 59753726-70b8-44bb-8bef-ba75c75e89e8 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.027258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.027258Z digest=sha256:412dea9ff7c50aa83041d68b34e8eca4331310e7d035cd6901372e7dfb81b4e5

Observation 5a365334-6104-47de-ad03-3b4afc2900d8 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.096391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.096391Z digest=sha256:80774a45e2b42acd1c42d04df16087038cba7ab01a2a6f596e121d4ca7c5e7a3

Observation 9f723ea6-4cd0-4cd6-a9d8-b15993af6316 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.208727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.208727Z digest=sha256:1d8823baa04e7572dcd12173a85ef7a721e780590a197b69cbe0b8b7ad6d3f4e

Observation a6817cad-c2f1-499f-afc9-97037fc12468 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.299337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.299337Z digest=sha256:3b0af5ea22895183b04d20c6cffdba5491df1138a79b09e383e0f7657352b6b8

Observation 67f62fc2-a81e-49fa-a515-6d3472ac5dd0 · outbound

This paper cites Microsoft coco: Common objects in context.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Microsoft coco: Common objects in context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.369793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.369793Z digest=sha256:af48c03827a73b197b5374a8c185aa200f7dd93e6e3990546e997d646b74a971

Observation 551991ec-7975-4641-a0f8-c7238f90f6df · outbound

This paper cites IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.513558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.513558Z digest=sha256:6d1f5cdd56f6e6322d102ec5ec4708e2bf7ebe4304d61177cf21c6b119caeafc

Observation e83cd7ee-0e9e-4b0d-a869-07c9dd7e76c4 · outbound

This paper cites Visual News: Benchmark and Challenges in News Image Captioning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual News: Benchmark and Challenges in News Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.573083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.573083Z digest=sha256:70b829316d5025ecf9900fb3bfc148ab0b0f9830f0872abce7a8967f85fea904

Observation b3c57bc6-8e15-40d9-96aa-0c7f592a4b37 · outbound

This paper cites Improved baselines with visual instruction tuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Improved baselines with visual instruction tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.612473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.612473Z digest=sha256:e7ca47dbd0153d686a2a0fd6a3b60cf63aca7a31145f5d9714184034e9faff60

Observation 659449bb-740b-47ae-9b37-96afcaccccd4 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.651069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.651069Z digest=sha256:1c9fe8fd759963aa4487b4cb4484fb2a9110a222b5da1c0d1dbc6f0519e4b6c5

Observation 8f85b993-5da8-47b3-b77c-2a1152f5dbe7 · outbound

This paper cites Visual instruction tuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Visual instruction tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.752801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.752801Z digest=sha256:eeee8d28edab3411d087a3fa94170ccc7f75e28a9f4c6185912b5f3b2aabd480

Observation 7908fffa-9e17-4bd0-af99-c05ba6b87985 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Llava-plus: Learning to use tools for creating multimodal agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.844748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.844748Z digest=sha256:a5ab42efdec402583f13e506a8533c1cf064ab43546b0544cd72f53d5fab28e9

Observation 923723cb-467f-45db-a2a1-8559ab4dc228 · outbound

This paper cites LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.923753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.923753Z digest=sha256:46e165a8a1490d4752b8502c884cc17f42956283f89d8ea86a57015b3149d8b6

Observation 260b0014-20bc-4880-bd2e-ddd6d1537b45 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Image retrieval on real-life images with pre-trained vision-and-language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:41.986061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:41.986061Z digest=sha256:a673a0080bcfb074e4e7a6de8c3944f9244daeea20ef4ce9b9566043ae14865a

Observation a94f9504-f8e7-4bac-8e61-eff261c60a71 · outbound

This paper cites Generative multi-modal knowledge retrieval with large language models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Generative multi-modal knowledge retrieval with large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.063940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.063940Z digest=sha256:599932f450f6c80f4232849db3ad8fc1151643db59ac3faf37f5f0fa4aaf968f

Observation ee3f2b50-ba0a-43d5-a43c-54a563afc47a · outbound

This paper cites End-to-end knowl- edge retrieval with multi-modal queries.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end knowl- edge retrieval with multi-modal queries

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.139537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.139537Z digest=sha256:811276cbd3972ac4928e642bb0bb9d13a8c67c352cd552f44db16c9446b52d2f

Observation 6eeb2464-83a1-42c9-a2e2-7903dc66dde8 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.196435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.196435Z digest=sha256:b1a36601c4da7e3d5cbec13b6d76c6db89e531680dee433e866809b14d0dc338

Observation 6b2f095a-1e1e-4972-9c8b-953e279d9a02 · outbound

This paper cites Unifying multimodal retrieval via document screenshot embedding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Unifying multimodal retrieval via document screenshot embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.277212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.277212Z digest=sha256:953f945c3d1b6c7a7d1fffbd8b9d7b4951b1cb26c2606b184a6bed6754901f04

Observation d8e194b3-da69-47a2-a2ad-e01cec6593b6 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.381843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.381843Z digest=sha256:56865741779993df5eff1a1bd9ccea49fe1899d4f870dc52eb0624d75c1eed6e

Observation 4ab821c6-af25-475e-8c26-7541f1f79ee4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.469662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.469662Z digest=sha256:e8e3b60593a1d31ae59f9f996ad0e1b2777a8968214fafd32cec1436bc8243df

Observation e2b4ea5b-3fde-419d-a747-7179956369f7 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.534731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.534731Z digest=sha256:bdb2c8d05067c58791f23f672ef5e18dbf4a8c0e3df62d3f53d25b8cc73be538

Observation 023bd09b-026c-499f-8603-90db05775295 · outbound

This paper cites Infographicvqa.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Infographicvqa

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.612365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.612365Z digest=sha256:e76f2ad8b1fcaa5775868e6a0b9e93f6ccb4a74966f62f8835695d589fcfecbd

Observation c140d83b-3342-4e5b-a369-99b7ede44866 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docvqa: A dataset for vqa on document images

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.664676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.664676Z digest=sha256:ca515731b945f9f1ea7497c7715d6aed1a39e4ff1a87204e6967860c26b2f0d1

Observation 4d78713e-fe65-4b1c-b9c1-1e0f20828970 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.493208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:42.732833Z digest=sha256:8980b4ab425ac6e9c3881bda6cd67c0ada5221dedd20a32a8d7c51d068d40beb

Observation ced10b56-ccc6-4d8e-a033-eb1cce907549 · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Docci: Descriptions of connected and contrasting images

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.391113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:42.803646Z digest=sha256:a9610d586ade176bfe98b24ae1c376d241b0918796768e9a177e3394f7ed461b

Observation 49df6ac7-3954-48eb-b100-f5fc6a1d8b9b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Representation Learning with Contrastive Predictive Coding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.931680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.931680Z digest=sha256:a47f96a1e2d99526d3bbcb9a707c0dafdcff514285a7f3f6318762c6f1a5bf26

Observation 0b9d012b-7ae5-462b-905a-907bb39c1e1c · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:42.998277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:42.998277Z digest=sha256:abd551cffc61e72c16dcb19b5f8cb7552c1691151b7e8626c4263c3632ba7c79

Observation d4c51ae5-fb47-48a0-91db-1ed9642c553b · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.208547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.085408Z digest=sha256:fb8c3f3b75116ee59e1b46131701f20fe05507249f4d4bb5aebebddc737b9e3d

Observation 5496358f-8197-43d5-854a-cab5644921b9 · outbound

This paper cites Filtering, distillation, and hard negatives for vision-language pre-training.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Filtering, distillation, and hard negatives for vision-language pre-training

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:55.069133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.158713Z digest=sha256:cd40f87df1d438e787a5915c24ee515b01afde81b869c761fffba77cf24fb3d8

Observation ca105a17-df5a-4e29-84df-50b64bd82dc9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Learning transferable visual models from natural language supervision

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.824769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.236786Z digest=sha256:1d3b4449325406a376c80df9d99712b39c96794273af5fc095a0386d8e378359

Observation e597097e-88f1-46d8-b812-e0d643ff95d8 · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Squad: 100,000+ questions for machine comprehension of text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.517715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.286893Z digest=sha256:97016924271d79ab2641959b621784d2ef22e417f3279a67600beb8795584a87

Observation e021a51d-e594-4c22-9eb9-a07b0d2b105e · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Contrastive Learning with Hard Negative Samples

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:43.344779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:43.344779Z digest=sha256:5a585f53d3dd25e31ca3356daccbde5ff9518ab86157ae1dc7e05b16a41482ea

Observation 465889c8-cf34-48ac-9603-90480d251467 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:54.171773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.400909Z digest=sha256:6db327592fd2720d3d615ac85088c00dcbe12e4046d08472a9ed3aae8ec2b69c

Observation 9c0da33c-6771-4f97-be58-70b508e1a057 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.924513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.456451Z digest=sha256:9e67172022598cc1540c30ee97c9ca40435b7fa523f11507564b11dfec293566

Observation 410094f7-2180-4340-a721-0811d9523df3 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A-okvqa: A benchmark for visual question answering using world knowledge

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.690728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.506063Z digest=sha256:d3292659ad9c379f0c333ad53453de452680ca3baedcc090ed1af9a75262a528

Observation 1393e73d-2486-4a8e-9db2-783e9b418355 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:43.554578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:43.554578Z digest=sha256:6c493f5fc6f247f8c8a094f5cd125a4c425333f962c436cbc747fc8edb1b2386

Observation 4fc5fba1-6b55-444f-83ac-485e825a272f · outbound

This paper cites Composed video retrieval via enriched context and discriminative embeddings.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Composed video retrieval via enriched context and discriminative embeddings

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.527330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.611833Z digest=sha256:15c207bc5079b2b4fe54eef1ce7b8ecb7e5c117f80c40af093fb6bc2ce287ede

Observation 9ddbad45-f0bd-4905-afec-10a0b6c37f36 · outbound

This paper cites Fever: a large-scale dataset for fact extraction and verification.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fever: a large-scale dataset for fact extraction and verification

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.281847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.665331Z digest=sha256:6409d82ec890ab0ad9c04c382e8da67496911bb5511b65841467d0822823b083

Observation 7e49834f-8235-49d4-ab1a-589f9b41a9d2 · outbound

This paper cites Covr-2: Automatic data construction for composed video retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr-2: Automatic data construction for composed video retrieval.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:53.036132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.745704Z digest=sha256:30279caacdd921cdc42dbf79446da3c7d016bc5efbedabb8b013f8291715537a

Observation 43ac363e-3f70-45b1-b50d-cb55d8b92e58 · outbound

This paper cites Covr: Learning composed video retrieval from web video captions.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Covr: Learning composed video retrieval from web video captions

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.799218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:43.885501Z digest=sha256:13baabb4d3485849b1bd0ded9a523674f193c8e4c1adf6c35ede87404c4fd34c

Observation 4165373e-c41e-4f73-bb44-aa18251c6109 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.083171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.083171Z digest=sha256:aaa026f6c375216410399a2917a9c0db4cf942edf2cdac4b5a8fd29b311095be

Observation fa90b388-4267-4933-acb7-be02d1deee11 · outbound

This paper cites A Comprehensive Survey on Cross-modal Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval A Comprehensive Survey on Cross-modal Retrieval

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.181075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.181075Z digest=sha256:721771e53b2eaa7547bee03e8508d78caac5dbd6f7f53d3de68798261e39d802

Observation 112d2fbc-dfce-415c-9cee-36e5875b4c65 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:44.311504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:44.311504Z digest=sha256:8e9f9ba834ffcc39c88aa381f8de65aeb3c6b1e58674692a38daa80c0656fb75

Observation 00e1ee86-7f82-4c3e-9637-2ac103fc5528 · outbound

This paper cites Fvqa: Fact-based visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Fvqa: Fact-based visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(10):2413–2427, 2017

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.620569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:44.452960Z digest=sha256:3de2cd35ff1a61150ede64db65ed7c1ad9238047b2b5461dc777703aef81048d

Observation f7e16bdf-8b49-4887-aea7-5fe6a581e0f8 · outbound

This paper cites Cross- modal retrieval: a systematic review of methods and future directions.Proceedings of the IEEE, 2025.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Cross- modal retrieval: a systematic review of methods and future directions.Proceedings of the IEEE, 2025

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.402528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:44.602734Z digest=sha256:4d8f43f94ed5e3ce16152236f16f3140c28650492404cb571f64686484275f63

Observation 22241213-305e-4750-a912-386dc00fafc9 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:52.210759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:44.751774Z digest=sha256:24f5c1ad63a9aca8b1ba6dbe6511462bf19f14fd1db1497ca7fbf9ea02df968c

Observation 1ccdfecc-be1c-4081-870d-5ebf432c8179 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Internvideo2: Scaling foundation models for multimodal video understanding

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.966932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:44.896788Z digest=sha256:224be0b2366e6ecb29ca1970f6bd55da4be9cf160e517d1cf628719186805f13

Observation c92fe79b-13d1-4bc5-a59a-c97df7637a2d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.013607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.013607Z digest=sha256:105daddaff8f75502a644d8e97b417321e8c744f119c03b6a928d749f54c152e

Observation f54c8e76-bdc0-480b-9fc8-a497a0ec08de · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.104101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.104101Z digest=sha256:2f5ef984249a4b023d81f86bf64b5c8c944a48707909efd09fabccbc09805274

Observation e20025dc-ed00-4245-aed1-d7b7f54c7576 · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.187968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.187968Z digest=sha256:da8354ff3229bdd5170a82c0f3e35598c08926711e197f6949dd42f8a6ebd1b2

Observation 4dc596e4-9b13-4caa-9b6a-29a932e200f5 · outbound

This paper cites N24news: A new dataset for multimodal news classification.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval N24news: A new dataset for multimodal news classification

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.794730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:45.273016Z digest=sha256:7c5401143b0b97c86662145e09410efef4656b9c4a19196c7bd65c5c90180005

Observation 2c6dc2a6-8756-4a3a-bc80-de2df8aaa6c1 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Uniir: Training and benchmarking universal multimodal information retrievers

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.598369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:45.369387Z digest=sha256:81228f1cb5443e5205acee3699bf1d956b77245da2d00bc4e9a3345efc383369

Observation 0ccbf610-017b-47cf-a0cf-93ed9d61fcfd · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Sun database: Large-scale scene recognition from abbey to zoo

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.405208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:45.437690Z digest=sha256:05d6259f26a0a52bfc652efb6fa2c3a3c507a8483d147fe1831b69224c52eac0

Observation b10f7c85-ecde-490e-acd6-857b03f829f8 · outbound

This paper cites Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.558531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.558531Z digest=sha256:eafddeecc00770df6f0ea4e7e42b326dc94b3a0529946d315730da1c58046cc6

Observation fafc94c8-f7eb-405c-bbee-bdbf296ecd29 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.233064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:45.631311Z digest=sha256:3eba74f53a91a967b9735ae8a989e13ea3059d9fe1ae45e823591235eeff5873

Observation 731ad417-6323-44dd-8cd0-da784f786d16 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.703524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.703524Z digest=sha256:416de0cfd1787b70de1bf521b27671b55053a6ad853a4b13f6f6552acd7e7e83

Observation ff5b82ca-a31b-43ae-aebe-5614bf80505d · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.772749Z digest=sha256:5af52b8cb4163059bc131e7cf401cc04424f1814980b1d812049d4a91b73b2e8

Observation 5f71f51c-5bf3-4289-8e15-5670d5225064 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.839239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.839239Z digest=sha256:f9559e682656ed9ffb44f62ff2b23046c3cd72c534b5b0d165dba4b65a676e15

Observation 0da0f99f-033f-4503-b951-6c6d25f52ded · outbound

This paper cites End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:14:51.025254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:14:45.930404Z digest=sha256:c337caafdf9c037443ada16ee49d1d3edbe02cf26a02f55488b6f4a3ab66eb9b

Observation 0b5c142e-69e9-4709-91ce-7bed037176bb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.015018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.015018Z digest=sha256:dd0569f93b93238b3a31c9d943a4d8140bd66fd0a75289cb5a28af943838ddbe

Observation 25d936f7-e057-4d04-b373-1597cff5eec4 · outbound

This paper cites CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.090733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.090733Z digest=sha256:c5a271717b1e1929e6a6876e5d4cc0abb098142be1cff3bf85a17e842a60240d

Pith citing papers

Observation ee1ebd9e-a9c9-45ae-ae6a-38cd48861805 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.563577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:00.563577Z digest=sha256:6ef875f02a017efc67ad2b34ac25495f23dc63f425ab81dd03087a8d7881ab65

Observation edbdef5e-892a-46ba-bff0-03e3b5d5177a · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.464856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:67b461d6148d554dc1cdb2a2efd74e22900f6cabcdd4fca753cbea95213c7413

Observation 913342ff-b8a1-4a73-8d70-b9dd8edb213b · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:23.532492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T13:00:31.952588Z digest=sha256:de862b8cd0c8961f9b46f31b8fee1d096dcc9559184085a26dcbb3746742c950

Observation 53c0bd72-bc9e-42d9-87d1-38ee6fc94016 · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:45.017918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:45.017918Z digest=sha256:3a90d99278063fe3aa85468a92311e866d6f07596d14329374d63945ef450a09

Observation aa897247-7e0b-4148-9f23-44169316f1bc · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:19.736570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T12:46:30.827346Z digest=sha256:8679e645ecd7b5f39bb41bd440c7f17ea04ea788d2844dee1b8125ceb281f644

Observation 8f2ce394-f8b1-45e9-893c-e0de3bc7989f · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T18:31:37.144968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:31:37.144968Z digest=sha256:727fd3fd26e9333cdc992f59cea827784014d7ba551b8b7196f83414e64453a1

Observation 1c769451-754c-4c00-a1c5-adf5685cc380 · inbound

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings cites this paper.

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T15:40:29.161499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:40:29.161499Z digest=sha256:d645e1fe52633c81e814416d8e119bf65a4c839c783b109991d86678088ab01d

Observation 8a9ccc3a-2e9a-4e11-9033-1cd3626b5ee9 · inbound

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval cites this paper.

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:13.551343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T16:56:52.714346Z digest=sha256:f7a79b5d6ad3cfba78e2c0f960fda215487649811fe3ed4cb3c9d41650f04fb6

Observation 8a5273d6-52b2-4e9e-9818-5a94e51accd8 · inbound

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval cites this paper.

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-07-14T18:56:57.352952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:56:57.352952Z digest=sha256:4b62e32210b96b8defb3e7c4cd15fa4099924b6df6e7bae02356ab81e9e2f6a8

Observation 19bd5014-640a-4ada-a4d1-8c375f21155c · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:49:36.604036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:ed59357d8c4285f55537cbb4560ec75c85ba23ba72c492d82cce7a87b322e0dd

Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · inbound

Illuminating Visual Identity in Universal Multimodal Embeddings cites this paper.

Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.106644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.106644Z digest=sha256:362d8b061211afaf82240b450e671b85832fef1de3b30a8fed5a17f502448b62

Observation 8ba9a618-5492-42cf-be31-7c0106c94e0c · inbound

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval cites this paper.

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:30:14.559722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:30:14.559722Z digest=sha256:3fc150e50f0a5a3ab1de0d3fe46a65af52b5ca390d03b9c2075d42e90c98e9b8