Pith. sign in

Paper Citation Record · LEDGER

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2506.06144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06144 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:31.866869Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:30:52.768388Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:20:53.347635Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19dd6684-388b-4cb7-be7d-caf1c71fb844 · outbound

This paper cites Qwen2.5-vl technical report,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-vl technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.634670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.634670Z digest=sha256:1289dda6a2d59cb136e5baedcedf0eb76d59a053207feb9f20bf926fbf76aba2

Observation 9c47f90b-1e70-4dbe-9359-ec944a84271d · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.732933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.644806Z digest=sha256:bdcc847fd233767b724097bdb67bcfd1ac78fe051b086e4a8fc90cf4a7db2349

Observation 4d14f39b-27d1-4e85-af09-9a0b17dcb282 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Valor: Vision-audio-language omni-perception pretraining model and dataset, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.713163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.652579Z digest=sha256:986f4491dfcc756010b98b4339cb20d81316432ff83c6df180b785cbe4174215

Observation 204d3d2f-313a-435e-bdbf-f9d933524ab9 · outbound

This paper cites V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval V AST: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.702081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.656440Z digest=sha256:ed169bfeba313db276400bb1fb839407474eef8921d52ff7cc8cc1f2f4f8224d

Observation 4e6e1dea-47dc-4f43-9209-8bdd96927bd8 · outbound

This paper cites PaLI: A jointly-scaled multilingual language-image model.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval PaLI: A jointly-scaled multilingual language-image model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.691985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.661364Z digest=sha256:0163299dae2c220f49c6b4074e31d8476335a08547ae6c2731a1800aba9aaed7

Observation 4041aeb2-69d3-403c-9bbd-41e128cb652d · outbound

This paper cites M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval M3docrag: Multi- modal retrieval is what you need for multi-page multi-document understanding, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.682596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.666238Z digest=sha256:b1b43823d4aca4ce2cd4dda86e8166d6f382d7e64b1dc3af21154c52ebf92231

Observation c196a0ea-c869-4dbc-8144-f7bd9a7b07ff · outbound

This paper cites Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Jacolbertv2.5: Optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.672817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.675998Z digest=sha256:5dce16bb9288ba81897fce201345ba69682b3b3d8a5ac08121605ecdad214f30

Observation 823aed3b-6806-41b2-87b7-e11f5308cda9 · outbound

This paper cites Cormack, Charles L A Clarke, and Stefan Buettcher.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cormack, Charles L A Clarke, and Stefan Buettcher

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.680329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.680329Z digest=sha256:81dd3500a4ff9d36f6988d39f00f3fc28ca0213217f74286ebef1a0826dd58b0

Observation fdb88050-ddff-42c2-a84a-9587fc3bd8b6 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval QLoRA: Efficient Finetuning of Quantized LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.684465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.684465Z digest=sha256:086962ea1f7d85cfcdf434f151cf477be41220dbe1e95b9dfcfbcc596e5c1bab

Observation e94752d0-2164-4e0b-b8ee-8b410a9fdcc2 · outbound

This paper cites Tibshirani.An Introduction to the Bootstrap.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Tibshirani.An Introduction to the Bootstrap

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.662854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.689461Z digest=sha256:ad589be7567995dcd663a7888714b3ba706e0179d2ed5092645f89c10d8dce3b

Observation 503b32ec-c227-4ad7-83b1-328a0819f599 · outbound

This paper cites A hybrid model for multilingual ocr.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval A hybrid model for multilingual ocr

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.694040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.694040Z digest=sha256:d7633be6de5e9c9263c8e0a49e372a817ca1ae612401b94b3158a0fa25b9f6a8

Observation 03718175-bbd6-463d-ac4a-df5eb644d7bc · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colpali: Efficient document retrieval with vision language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.652851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.698596Z digest=sha256:addbad037cc77b0519a34190b82a4154558e02addecafbb6cc270137e801e8a7

Observation a61972d4-f41b-4aed-9ad7-4321c77f102a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.703431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.703431Z digest=sha256:fea36c1a0ac09e2e354264f4a36455518616728b907584e3a3b62b76f4dc8abb

Observation 16671a50-c346-45eb-bcab-944ba321d299 · outbound

This paper cites Imagebind: One embedding space to bind them all.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Imagebind: One embedding space to bind them all

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.708822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.708822Z digest=sha256:f9a2c38f70ac04fe4aaa3087340a9e383c1eccf6b8ec7263ab7522ec11e9a06b

Observation 507fc0cc-77dc-41f6-819b-9aa209ddced5 · outbound

This paper cites Cumulated gain-based evaluation of ir techniques.ACM Trans.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Cumulated gain-based evaluation of ir techniques.ACM Trans

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.713608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.713608Z digest=sha256:3e622f6bae011af76ca4658c8b7346aef52546d24a0dce4303999535ee29e89e

Observation 552d9585-d1f9-4858-9c82-fde725678c36 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Scaling up visual and vision-language representation learning with noisy text supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.635323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.718360Z digest=sha256:4be828019a7576350489eafad73114fa0d4c61c12c87a4cd038aafd4dc260902

Observation 2fb790ac-e2f2-40a2-bf1c-aebbb741db7a · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.723331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.723331Z digest=sha256:3dc9fc643e31c9543bbe791baf88f15f9109b39a4dd49f2ce7bc9b0c838c9c0d

Observation 7ad71b70-7828-4bcd-926a-111f73edcdc4 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Dense passage retrieval for open-domain question answering

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-07T06:05:31.728907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.728907Z digest=sha256:fcd094c57a9ee4fe93b01ebcba7ff9cb53d3675a5ea49b441dc5b4a20be5dd9b

Observation 5d801cd2-d676-4a75-ba06-a2f17672d1ac · outbound

This paper cites Colbert: Efficient and effective passage search via con- textualized late interaction over bert.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Colbert: Efficient and effective passage search via con- textualized late interaction over bert

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.734527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.734527Z digest=sha256:fe33d943d70f7db7437d0ddb33499149c2ecc7681b9498e7e0118ba27bcd434f

Observation 65563c93-0fa7-45b7-bbe9-24eb6fa33d41 · outbound

This paper cites MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.738189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.738189Z digest=sha256:86c09792e8e0fefa314e5bae2443cc1f5e02237ba65dd1beb5ca96d2b2945ee9

Observation 7ba335fd-29dc-4963-95c2-4718223f64d4 · outbound

This paper cites Learning to rank for information retrieval.Found.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning to rank for information retrieval.Found

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.742958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.742958Z digest=sha256:bbfb3a6d3389a9cbda0ec3ad4ef8af64e436991057def63e6c536a03e47d8617

Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.747047Z digest=sha256:8ea3ea2dd3ab80e027a7b0e25f514261046620e46763ab41c5019c5fff6ff2df

Observation 3a544e91-f00c-4ccf-963a-beedaf9bccc3 · outbound

This paper cites Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Handwritten optical character recognition (ocr): A comprehensive systematic literature review (slr).IEEE Access, 8: 142642–142668, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.751704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.751704Z digest=sha256:64b8e5f4f3a8c53b7584f4307c76163ddda5c0f7dcbe1e435109a476fd2bbb6d

Observation 3363841f-0da7-4353-aa11-8fee73981574 · outbound

This paper cites Vladva: Discriminative fine-tuning of lvlms,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Vladva: Discriminative fine-tuning of lvlms,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.624648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.755787Z digest=sha256:45c7d4cc8eb746ff0a3e2a322ad045153658814c7f603e41c49396e60a565022

Observation 58b2133a-98ee-49e3-921d-dad0b7892035 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.764683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.764683Z digest=sha256:70e222c73e0feb137c1b489a055e20f5e3be3e3080f0173ead8ad089bb96f548

Observation 642bb0b5-4a42-48eb-aa70-618792b5bc56 · outbound

This paper cites VladVA: Discriminative Fine-tuning of LVLMs.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VladVA: Discriminative Fine-tuning of LVLMs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:05:32.158935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.760413Z digest=sha256:923579c0d1a5d9a26981293c20b0676cd100cedc759e4f58ea32f0191d077e6f

Observation 5c79f5b3-dc2d-496f-8114-a8704e143b80 · outbound

This paper cites Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.776301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.776301Z digest=sha256:423676bb4e7034918a10b7c78d4c0c52a636f9d1598248c5a2ea655d9f82ea74

Observation 5c85e58a-1f3e-4e94-a994-50e76e1fc131 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.780127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.780127Z digest=sha256:67485c704d00f993cf912c78fcb443cd7c8e7e20245479ccd0ff1b62e6153696

Observation 63cd450c-52ea-467c-b862-0b91c26bfd4f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.772749Z digest=sha256:245de38a5b71ba7e2a2fd76e1230609b4a616b12f828985d55a4779e08981481

Observation 1ae11b3e-74b7-4d9c-a018-5a43742991c4 · outbound

This paper cites ColBERTv2: Effective and efficient retrieval via lightweight late interaction.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval ColBERTv2: Effective and efficient retrieval via lightweight late interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.790322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.790322Z digest=sha256:a714154387807c895e55dd9e68c1c431448af6055e2c1938b90add01c1d738c3

Observation d37da703-0756-46f2-b5ff-e95065c779c1 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.795910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.795910Z digest=sha256:597ee2e0c24d27f668602e52876dd431100d15263585133b4d0eab7ad984f467

Observation 2a0a17ba-f607-48f0-887f-96847c25dff2 · outbound

This paper cites MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.784566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.784566Z digest=sha256:35615d99487bbee4d150499bcacb0e7f6de0db0ac44d0fd5c26d106aa1d755b4

Observation 3036bb25-e930-48d7-bee1-519b5b577769 · outbound

This paper cites Gemma 3 Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Gemma 3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.804178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.804178Z digest=sha256:2d5c8cb5048e06b1b0db9ec56fedb9135fd556797f01fca8f12e014a4e5524dd

Observation 326d5b02-42de-418b-bea9-d24b4bf7571a · outbound

This paper cites BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.811569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.811569Z digest=sha256:bdb8eadb965e8e8b1c4a0ec1a4fe2e1ba3441c325a31ae067d2a03919eb1bd42

Observation 4bcb1f75-7ef4-4d63-beb6-aa3633bf30c1 · outbound

This paper cites Generative multimodal models are in-context learners.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Generative multimodal models are in-context learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.601741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.800337Z digest=sha256:6f0a3282467dd7434b9baee3619505e4d255da52024c010f3fd6a919f2e82933

Observation 77fa6841-8759-45b8-aff8-227fbedaaf7e · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.586001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.819442Z digest=sha256:38b7fb1c2881a044650c1ec471fcbaca75a1fc8a694973b3ca5bb7b5e01832ad

Observation bae2896b-f6e0-4548-b8bc-34a2e87c3ace · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.825064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.825064Z digest=sha256:a6d1b86e0bacc2b61358ddb9df016975032171586206d9cac7edc007e834e289

Observation 6f2d055e-5f99-4982-aa01-344b2833129c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Representation Learning with Contrastive Predictive Coding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.815340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.815340Z digest=sha256:3e2e7027fe1f744b777de270049b6863d7ae5fc5b65e215affce743bb47a51eb

Observation 88bb1e5a-a138-4970-8b01-4361cc5500b2 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.569152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.833276Z digest=sha256:1190b8ff9a0c291f5f248f70ebfc12f6efb2b5aae7a49e2876e37ce967573b0e

Observation 868f1be1-500e-496a-b354-79f63cfaeca3 · outbound

This paper cites Qwen2.5-Omni Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-Omni Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.843978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.843978Z digest=sha256:4fec2dae1d928ebb8dfca87cae8583fa2b38c29a88dd0030a2943d70a4a089eb

Observation d7dda32d-cafc-4d71-9d8f-2e263aedc925 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.829176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.829176Z digest=sha256:bc036b6b575a2fbf519a9da74e1c1f922b42d0606b8e747e42df5c958f706d50

Observation 91c94b0e-fd37-4767-bc11-35a9a63c640d · outbound

This paper cites UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.852256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.852256Z digest=sha256:023fb6e636f88f3ce413c73e8634f94cf5e4a62366835ee29a5a2c15b6865371

Observation 2677d8dd-e083-4e7c-9b36-385150f719b7 · outbound

This paper cites CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CREAM: Coarse-to-fine retrieval and multi- modal efficient tuning for document VQA

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.537696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.856797Z digest=sha256:73b6a0a19e8f602290bf1204090ef2a067a3e98769e02df94120b47826bf6bff

Observation cf2fcc13-053f-43f6-a2b5-69eea1c8d835 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.526800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.861838Z digest=sha256:8d30c41de1124b2ccee07bb3f5a0af020bfaf93d90772ef66cad9a9b75e7248a

Observation f2675d97-12d0-44d6-b24c-22b115211bfc · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:05:32.548489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.847733Z digest=sha256:a984fd21395bc4ecb63a0f1aa881953e35d130397c0a339b2bf2894fcc81409a

Observation 42fd1618-52b1-4a1d-8bde-291cbfa58eaf · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.866869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.866869Z digest=sha256:d4f43bc23e464eaa69bf97578bf74103035904c13c3afa422814bf0674b2da67

Observation 79d1128f-70a1-45c3-bf22-9e356735e2a4 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.722824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.648707Z digest=sha256:3df60749ce0946deaad1b981af5991207aadce90a8113f97a9cca22f146de7a9

Observation 245a8727-c197-4289-ba3c-9bed0af1a4b6 · outbound

This paper cites an unresolved cited work.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:32.558569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T06:05:31.839371Z digest=sha256:f9819af8755b897cd477c89763b2e9268708428688d8951b8953fcc148d79736

Observation 739bd388-8da6-4992-be2b-18cddfdd5aac · outbound

This paper cites Qwen2.5-VL Technical Report.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.640938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.640938Z digest=sha256:ec4702af2da6b29ce8408e204196bc129f8ec7cfcbf8d1b36d21f212ba2185c6

Pith citing papers

Observation 1c39de2c-01db-44ed-a743-64a7eb788bcd · inbound

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning cites this paper.

Hidden in the Multiplicative Interaction: Uncovering Fragility in Multimodal Contrastive Learning CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:53.354526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:12:07.548447Z digest=sha256:d682df37841aa205c057e929014e69dfed07cc76b74f599d0abfa43bfa19ad14

Observation b51f727c-9452-4602-b6d9-fa49ca61b69a · inbound

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding cites this paper.

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:30:52.768388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:30:52.768388Z digest=sha256:29b9abda690a55f8549f0e2660c733b79537f816e8f9ce8cb29ff7dfbbc1bea9