Pith. sign in

Paper Citation Record · LEDGER

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.06144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06144 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:31.866869Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:12:07.548447Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:20:53.347635Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19dd6684-388b-4cb7-be7d-caf1c71fb844 · outbound

This paper cites Qwen2.5-vl technical report,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-vl technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.634670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.634670Z digest=sha256:47d4865ccb10efc33476c34ec74c42326de8496ef11a4c5df24315be96726f07

Observation 9c47f90b-1e70-4dbe-9359-ec944a84271d · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.732933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.644806Z digest=sha256:fe5a66e78c9767b27d6bddd74938932ce6cf825539b02bbdc5e689546addcab7

Observation 4d14f39b-27d1-4e85-af09-9a0b17dcb282 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.713163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.652579Z digest=sha256:5b946b46bbafa4287537e40c07a429b108a119af3b47e224ece6817e9de9c9c5

Observation 204d3d2f-313a-435e-bdbf-f9d933524ab9 · outbound

This paper cites V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.702081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.656440Z digest=sha256:972be872e8228537bccadb9617298ad15b3437e2c05bb253daa11e657b9e9e09

Observation 4e6e1dea-47dc-4f43-9209-8bdd96927bd8 · outbound

This paper cites PaLI: A jointly-scaled multilingual language-image model.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval PaLI: A jointly-scaled multilingual language-image model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.691985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.661364Z digest=sha256:99e92823d4a8cce6a71fc194ea66f13086229b1cec4500b3ff13f300ededd1fa

Observation 4041aeb2-69d3-403c-9bbd-41e128cb652d · outbound

This paper cites M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.682596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.666238Z digest=sha256:10f5e357633d90f213056481fed2646364c2b6895238adfcf87314b92946f4af

Observation c196a0ea-c869-4dbc-8144-f7bd9a7b07ff · outbound

This paper cites Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.672817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.675998Z digest=sha256:3220a18528a35432481b32ed8caab9e55b5138cfe0b761de3ad780675fd69f46

Observation 823aed3b-6806-41b2-87b7-e11f5308cda9 · outbound

This paper cites Cormack, Charles L A Clarke, and Stefan Buettcher.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cormack, Charles L A Clarke, and Stefan Buettcher

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.680329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.680329Z digest=sha256:0d3cfbf9c45c1da48efb160bef724d961530369a7cd3d3007c3927f3fc8e945a

Observation fdb88050-ddff-42c2-a84a-9587fc3bd8b6 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval QLoRA: Efficient Finetuning of Quantized LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.684465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.684465Z digest=sha256:40a5141defd76db3a02d1f052665c2e70823bd1adf3ba7bd4de28889a44f8d5d

Observation e94752d0-2164-4e0b-b8ee-8b410a9fdcc2 · outbound

This paper cites Tibshirani.An Introduction to the Bootstrap.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Tibshirani.An Introduction to the Bootstrap

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.662854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.689461Z digest=sha256:01104b353ab37b13cadf27a327d3dd12fa78a9875ca3c5224ec747fdbe4721e9

Observation 503b32ec-c227-4ad7-83b1-328a0819f599 · outbound

This paper cites A hybrid model for multilingual ocr.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval A hybrid model for multilingual ocr

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.694040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.694040Z digest=sha256:9ff9a685bcc4c26ceb29e3447a6fd349a1378007f3de7dfa59cc88f55b2bba8a

Observation 03718175-bbd6-463d-ac4a-df5eb644d7bc · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colpali: Efficient document retrieval with vision language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.652851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.698596Z digest=sha256:dcc0049e63f77edaac7632d54008da9a3eb81887fcaa1a538eee18420270510b

Observation a61972d4-f41b-4aed-9ad7-4321c77f102a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.703431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.703431Z digest=sha256:9d04d81f7fb944ca8b0db4132e1c6248241fc58814ad9f971df56582d0aaa5e9

Observation 16671a50-c346-45eb-bcab-944ba321d299 · outbound

This paper cites Imagebind: One embedding space to bind them all.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Imagebind: One embedding space to bind them all

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.708822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.708822Z digest=sha256:25e3c6b7496f09cca5e82cdd00d52128d38c1d3b15860562221bfca09a2e3ead

Observation 507fc0cc-77dc-41f6-819b-9aa209ddced5 · outbound

This paper cites Cumulated gain-based evaluation of ir techniques.ACM Trans.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cumulated gain-based evaluation of ir techniques.ACM Trans

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.713608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.713608Z digest=sha256:03b9884634c4f724811bd50d0a8dc261800a00ab46bad0fc92e203a7140ff890

Observation 552d9585-d1f9-4858-9c82-fde725678c36 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Scaling up visual and vision-language representation learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.635323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.718360Z digest=sha256:be643c9b232894485f93f4bf0e6fffdca49eda1ce82120979d70456fec7c2785

Observation 2fb790ac-e2f2-40a2-bf1c-aebbb741db7a · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.723331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.723331Z digest=sha256:d5176323dd4b04326a8494a78bad2305c1271ecd785986abc9b063443bd39611

Observation 7ad71b70-7828-4bcd-926a-111f73edcdc4 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Dense passage retrieval for open-domain question answering

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-07T06:05:31.728907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.728907Z digest=sha256:ba27dd2011c2c7c75df3027ae2848301d50a74d23e5a518f8e18de30b6835754

Observation 5d801cd2-d676-4a75-ba06-a2f17672d1ac · outbound

This paper cites Colbert: Efficient and effective passage search via con- textualized late interaction over bert.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colbert: Efficient and effective passage search via con- textualized late interaction over bert

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.734527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.734527Z digest=sha256:3198963ef448bbeb703881ffc11591f1f3b43a0ed76dc29e2b76272417d1fea2

Observation 65563c93-0fa7-45b7-bbe9-24eb6fa33d41 · outbound

This paper cites MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.738189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.738189Z digest=sha256:a6d81af0b0b323bba5d058c347c3c0bdb7d9a2c714268b603caa7e88a09a1b35

Observation 7ba335fd-29dc-4963-95c2-4718223f64d4 · outbound

This paper cites Learning to rank for information retrieval.Found.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning to rank for information retrieval.Found

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.742958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.742958Z digest=sha256:84a5ab6e485b79ddf91b9eda2735fac5e613f75a4221e057629c3c68906a4ca3

Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.747047Z digest=sha256:b50f182f1bbd115dc3ff27d0cf30a5b104ac4330b8c99284ce6650ed50e74c18

Observation 3a544e91-f00c-4ccf-963a-beedaf9bccc3 · outbound

This paper cites Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.751704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.751704Z digest=sha256:705df2552b3d5ab491882ac014285a160f97b6211998fc9e25610eb9039686ba

Observation 3363841f-0da7-4353-aa11-8fee73981574 · outbound

This paper cites Vladva: Discriminative fine-tuning of lvlms,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vladva: Discriminative fine-tuning of lvlms,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.624648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.755787Z digest=sha256:85d501ff85f5b23134c93e4b2ef6de64e9ce35dfe9dc9b15f5e1d1e0324fa10f

Observation 58b2133a-98ee-49e3-921d-dad0b7892035 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.764683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.764683Z digest=sha256:b0d4639eedd382b07acb0c2505283e673a52c36ae5e4c9532c678afb9f73f26e

Observation 642bb0b5-4a42-48eb-aa70-618792b5bc56 · outbound

This paper cites VladVA: Discriminative Fine-tuning of LVLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VladVA: Discriminative Fine-tuning of LVLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:05:32.158935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.760413Z digest=sha256:3a3cb93c1a85d0cd6895eabdf7772d21bdd2fa03e727b8475d96672c29f9638a

Observation 5c79f5b3-dc2d-496f-8114-a8704e143b80 · outbound

This paper cites Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.776301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.776301Z digest=sha256:38078eec2a813ae5d6d2f864a4ce84793ef50d75db2ab572a3da17697bbed1d3

Observation 5c85e58a-1f3e-4e94-a994-50e76e1fc131 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.780127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.780127Z digest=sha256:b796da2df6e81771eda8de1ca653ec4dfe8b873156d688c7a72e5e948a097513

Observation 63cd450c-52ea-467c-b862-0b91c26bfd4f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.772749Z digest=sha256:3e689c5733d933e15b85dc8d9ce589afbb10825f6f31ffcfea1be65de2a305f7

Observation 1ae11b3e-74b7-4d9c-a018-5a43742991c4 · outbound

This paper cites ColBERTv2: Effective and efficient retrieval via lightweight late interaction.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval ColBERTv2: Effective and efficient retrieval via lightweight late interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.790322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.790322Z digest=sha256:6422f6fdcccadf4fe00f7f0646cfe0e52c29a7c6fd69fec713be6b0bdac15929

Observation d37da703-0756-46f2-b5ff-e95065c779c1 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.795910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.795910Z digest=sha256:685e8f50b690de33113462c2ea5ce7eebb69633a10455044cd763bbaff75695d

Observation 2a0a17ba-f607-48f0-887f-96847c25dff2 · outbound

This paper cites MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.784566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.784566Z digest=sha256:16ed42e6f99e315b7e87429c17dd92e83f424e61563effd09c5cfb4a4d6787d3

Observation 3036bb25-e930-48d7-bee1-519b5b577769 · outbound

This paper cites Gemma 3 Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Gemma 3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.804178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.804178Z digest=sha256:1acbbf35870125d33a1038ff88416ee4a05e07736fd188c7bfd4c6f27348c4a8

Observation 326d5b02-42de-418b-bea9-d24b4bf7571a · outbound

This paper cites BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.811569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.811569Z digest=sha256:5125110aa408b172386256ed170d62a5aa17cf997f7aa3df96bbbfb0027f929a

Observation 4bcb1f75-7ef4-4d63-beb6-aa3633bf30c1 · outbound

This paper cites Generative multimodal models are in-context learners.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Generative multimodal models are in-context learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.601741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.800337Z digest=sha256:384a296fb3a207c1f72075a22cd0ecb71a3d391a4d2459946bbf507f8dadd796

Observation 77fa6841-8759-45b8-aff8-227fbedaaf7e · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.586001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.819442Z digest=sha256:60f996ec0af7efb86716bd25e5b1e48fd5a49404fbd540f67a7e66857db3f9ea

Observation bae2896b-f6e0-4548-b8bc-34a2e87c3ace · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.825064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.825064Z digest=sha256:eabebbabed82de786b416e86778a45db3dc917b9f8b0976b31ba3371e8389be8

Observation 6f2d055e-5f99-4982-aa01-344b2833129c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Representation Learning with Contrastive Predictive Coding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.815340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.815340Z digest=sha256:24bf0fb4a2061590c0c6f7b63187a0997483330eec76fd32c70596defea4aa49

Observation 88bb1e5a-a138-4970-8b01-4361cc5500b2 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.569152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.833276Z digest=sha256:7828fd0429b13ac8223340d5893d53f6680e7eb4b53a8fb73791aff560b302b7

Observation 868f1be1-500e-496a-b354-79f63cfaeca3 · outbound

This paper cites Qwen2.5-Omni Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-Omni Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.843978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.843978Z digest=sha256:c6b8fa5d5d1387ecbaa24a78df1b28c826959ca1b1a76af5fe44363110977519

Observation d7dda32d-cafc-4d71-9d8f-2e263aedc925 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.829176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.829176Z digest=sha256:764c617527059683205a170b2b5a7494204a0031579de921e820b857d0206423

Observation 91c94b0e-fd37-4767-bc11-35a9a63c640d · outbound

This paper cites UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.852256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.852256Z digest=sha256:561010e88dd10b6b4ab20270ab60d5ce814a4c28b2da02e78b9495ad7c84304b

Observation 2677d8dd-e083-4e7c-9b36-385150f719b7 · outbound

This paper cites CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.537696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.856797Z digest=sha256:3cdda7fd469b7eb33cc42e49ee422413daf793b1cd5af2973285609812ee97cb

Observation cf2fcc13-053f-43f6-a2b5-69eea1c8d835 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.526800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.861838Z digest=sha256:6f002a4e12f759ab88ffc19bb12baef123146e31789fe9a1322ea224a14b1e37

Observation f2675d97-12d0-44d6-b24c-22b115211bfc · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.548489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.847733Z digest=sha256:8261601be13ed4b3bc1d88a95cf8c3c7217fb508013df0ef20040e1a321db974

Observation 42fd1618-52b1-4a1d-8bde-291cbfa58eaf · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.866869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.866869Z digest=sha256:1ec3d7539d83d458947a536d9a68676b0b5fe74dd05102931950dc20cd9e9dbf

Observation 79d1128f-70a1-45c3-bf22-9e356735e2a4 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.722824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.648707Z digest=sha256:69ace3aeda9d876dd48d7747574000215e5f37ed9e728f6c3eed313131c6f8a4

Observation 245a8727-c197-4289-ba3c-9bed0af1a4b6 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.558569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T06:05:31.839371Z digest=sha256:d43b60326785a4b1e6e554dd918139cc6a81ac5c73f90219075d881048fb3d64

Observation 739bd388-8da6-4992-be2b-18cddfdd5aac · outbound

This paper cites Qwen2.5-VL Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.640938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.640938Z digest=sha256:6469629d27859f462b8832fd8e83b8f97dd6a41e489984f21c8651c263e7d603

Pith citing papers

Observation 1c39de2c-01db-44ed-a743-64a7eb788bcd · inbound

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning cites this paper.

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:53.354526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:12:07.548447Z digest=sha256:18f87972fecc45c95c3d2d8b2c4a9e3ada8ee6010beeb81e594d07dace1805a0