Pith. sign in

Paper Citation Record · LEDGER

Recurrence Meets Transformers for Universal Multimodal Retrieval

As of 21 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2509.08897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08897 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:06:09.196053Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T22:42:53.631688Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T22:44:14.780765Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8ca069a-18d3-4f2d-bdde-f7af3d8a7937 · outbound

This paper cites UniIR: Training and Benchmarking Universal Multimodal Information Retrievers,.

Recurrence Meets Transformers for Universal Multimodal Retrieval UniIR: Training and Benchmarking Universal Multimodal Information Retrievers,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:08.992069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:08.992069Z digest=sha256:60480872354a37e0f902eae59f8592b6322ad5cb861d4c7a13f3fd6b99576c46

Observation 1792502f-5eb9-4515-87fa-835e6cbb704f · outbound

This paper cites Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:08.995641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:08.995641Z digest=sha256:4e376d5e8e86dd8beb79bac1a2b921b80018701bd0e737ace795e7332776fed6

Observation d988b9ac-49cb-4cab-ad37-be7dba9ddbf7 · outbound

This paper cites Unsupervised Dense Information Retrieval with Contrastive Learning.

Recurrence Meets Transformers for Universal Multimodal Retrieval Unsupervised Dense Information Retrieval with Contrastive Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:08.998658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:08.998658Z digest=sha256:92f0dc85421300a3242df797a2cb8415c043abdb39d810b0290e0b5cfbaee6ef

Observation 9f52a211-5a86-464f-9de6-8bae366df22a · outbound

This paper cites SIFT meets CNN: A decade survey of instance retrieval,.

Recurrence Meets Transformers for Universal Multimodal Retrieval SIFT meets CNN: A decade survey of instance retrieval,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.001940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.001940Z digest=sha256:7e9767559dc80a3a54316a7ea343d1357ab1f697c6489a6762d4a1c5299b3bde

Observation 4027c289-8b90-4620-8c5e-e6a302f4a09a · outbound

This paper cites Large-scale image retrieval with attentive deep local features,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Large-scale image retrieval with attentive deep local features,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.004839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.004839Z digest=sha256:299897ec16eff7e0e3f4a992517f4e62aeb990de89ef7c602d6439a7c113cc31

Observation 3bcd3642-679d-4a9c-a57a-3db334b70d0e · outbound

This paper cites Microsoft COCO: Common Objects in Context,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Microsoft COCO: Common Objects in Context,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.007711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.007711Z digest=sha256:ca5a24331c499b5588f2b82ae5936c1ca9b9a496f66b3435f5c9a3e10a4fb197

Observation c6229284-f665-4841-a8de-c53b52abc0bc · outbound

This paper cites Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.010579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.010579Z digest=sha256:f81beb357c3b3da1204a780f542b11e51f05e5378cfa7d7bdfa01f07dcfc4c5b

Observation 9c6b7e85-f9ea-489c-b3d6-381ce3ebfb15 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models,.

Recurrence Meets Transformers for Universal Multimodal Retrieval LAION-5B: An open large-scale dataset for training next generation image-text models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.013128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.013128Z digest=sha256:0d6cc326c0a793ad78cca8ac294bd27d0ea5f7d7b606cc14a5ca39b6e9475053

Observation 52034a3c-ed93-4ad3-b8ff-0d46d790feb4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Learning Transferable Visual Models From Natural Language Super- vision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.015687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.015687Z digest=sha256:024f04e28772ade4a351c1bd7ed60286bfa16592ce786e6df254c0ca57f7c302

Observation bcc40c18-7530-4917-a83a-428a3dd8e3aa · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.018250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.018250Z digest=sha256:8e6f4e2b14ef564d90a8bddac50f1bd4a7bc5b086746396b7a22b641f386b7bf

Observation 582e9aad-f2a4-4d20-b248-f15f3818faf2 · outbound

This paper cites Reproducible scaling laws for contrastive language-image Learning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Reproducible scaling laws for contrastive language-image Learning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.020900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.020900Z digest=sha256:3f8cc95d98097178e10194b49daae3d35a5458cf87c42b63c628a257b6f56cd2

Observation b68572ae-dfbe-441f-b4fa-60882edb52e3 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Sigmoid Loss for Language Image Pre-Training,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.023616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.023616Z digest=sha256:8e95557b2ca43b7a3ce3862a5640668ae62dbaaf92d5f13109bcc08e7f134aae

Observation 20c5bc2a-2a8d-4ee2-b109-de938fac207d · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Recurrence Meets Transformers for Universal Multimodal Retrieval SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.026112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.026112Z digest=sha256:81afc93c47d8ff4e361dc9f08b8802a5b5dc3b8becafbf544de6eabe9fa6e334

Observation 73f81a17-c6d3-4a28-91d2-658463a6484e · outbound

This paper cites The Revolution of Multimodal Large Language Models: A Survey,.

Recurrence Meets Transformers for Universal Multimodal Retrieval The Revolution of Multimodal Large Language Models: A Survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.029205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.029205Z digest=sha256:8b32bdcb2165dc4321d533b486e0c95cc27259966afaf24a1572ef75ef9a6981

Observation a126f73a-3fb9-487d-9bb9-5a99c99a4093 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Improved Baselines with Visual Instruction Tuning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.031635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.031635Z digest=sha256:6530705f74fb124e3b54370843b4401e2065e8e0aee33d06422d2baec208974f

Observation 8e9a095f-918e-449f-a442-2f29cfd86a5e · outbound

This paper cites LLaV A-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval LLaV A-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.034075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.034075Z digest=sha256:eb8bc9ef55cbcd1b970ab1df922ca4f73edad063f72eb961e4a7f20b7b052d06

Observation 5bdd6a16-d827-4f17-80e0-f86f9a32aa82 · outbound

This paper cites Qwen2.5-VL Technical Report.

Recurrence Meets Transformers for Universal Multimodal Retrieval Qwen2.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.036465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.036465Z digest=sha256:3e1257a6823b96c0c5e77c6f2691b28e085f595be51cb36f43445c3fa85897c0

Observation b322d1b4-7979-4ddc-a319-2655e934f2be · outbound

This paper cites PreFLMR: Scaling Up Fine- Grained Late-Interaction Multi-modal Retrievers,.

Recurrence Meets Transformers for Universal Multimodal Retrieval PreFLMR: Scaling Up Fine- Grained Late-Interaction Multi-modal Retrievers,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.039420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.039420Z digest=sha256:f0b7441f4c30124e96b9b960d69b3f505861d3fd8d2bacf58ea1f12ebb6a9912

Observation b593bbeb-fd40-4846-a029-dc65e6e78011 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

Recurrence Meets Transformers for Universal Multimodal Retrieval Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.041750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.041750Z digest=sha256:837cc5ff4199593261c97dcb0fa923b7506f92d1ac46bd3f20cd415b660acc5f

Observation 331810db-6cb3-4ab9-9e6f-28a98cdfe718 · outbound

This paper cites Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.044289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.044289Z digest=sha256:730fe8df35b1fc1263271292f3b074ead93d6797e22b5597f52db5d913763134

Observation d1f78b14-bd6b-433d-a01e-279da814e1f9 · outbound

This paper cites Long Short-Term Memory,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Long Short-Term Memory,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.046723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.046723Z digest=sha256:98c7d1798a227749e9071d1f08d82a437ddd53c4860eace8760b511b15009d91

Observation 7d6f6d1d-d359-4f7c-a9b7-b837ce3adae3 · outbound

This paper cites Open-domain Visual Entity Recog- nition: Towards Recognizing Millions of Wikipedia Entities,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Open-domain Visual Entity Recog- nition: Towards Recognizing Millions of Wikipedia Entities,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.049120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.049120Z digest=sha256:4ceab5b2054f1501acd9724e224642ae1498c06dc3003341c8854f501c40470a

Observation 1a048ac7-65df-49ee-948d-709bc298ce4d · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge,.

Recurrence Meets Transformers for Universal Multimodal Retrieval OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.051455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.051455Z digest=sha256:3b31bff07407cd74ee9861e411f5ad5984106e7de71cec65a06722f05515ca08

Observation 91419f0c-a081-47d8-a75b-c9473fdbcc9b · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Recurrence Meets Transformers for Universal Multimodal Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.054038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.054038Z digest=sha256:78c824d60e0bf063aa2edd21eb9c822d5bb4ed85990f998dff06ab487e6d104d

Observation 950f727c-63a1-4901-9867-b64af77b98f7 · outbound

This paper cites Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.056915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.056915Z digest=sha256:4d22a4eb63d709f7145740332df8e61ad0fdfd10940e98817185386c7f1d075a

Observation ed794539-b2af-427d-8e46-55940d9d77e1 · outbound

This paper cites Image Retrieval on Real-Life Images With Pre-Trained Vision-and-Language Models,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Image Retrieval on Real-Life Images With Pre-Trained Vision-and-Language Models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.059565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.059565Z digest=sha256:274bb5f07aff8bdf6896ca4f3a36cfeef96232da209ac32a5410db752dd5040a

Observation 5d1b67ac-a3ea-4d46-8727-c550582c3394 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Long-CLIP: Unlocking the Long-Text Capability of CLIP,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.062053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.062053Z digest=sha256:8edce43ba5bebd04cf257973059e26f0efa35f651c6910a77879a256494be656

Observation b273f406-2d24-4385-a4b2-1b714b2a8435 · outbound

This paper cites Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With Transformers,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With Transformers,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.064447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.064447Z digest=sha256:2c3934ab9a515b4a0ad0807ec27bee01312183d94500bbd9209c08c7338aa610

Observation 4e0da6de-bfea-4d8a-9317-224bae26ae74 · outbound

This paper cites Smooth-AP: Smoothing the Path Towards Large-Scale Image Retrieval,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Smooth-AP: Smoothing the Path Towards Large-Scale Image Retrieval,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.066886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.066886Z digest=sha256:54bdb85c333b5b5e750efaa11cf4cdc96be1ef2f135d7803d4c356e7a18324b1

Observation 2861b53c-a93a-469a-8a3b-8398139d3926 · outbound

This paper cites Conditioned and Composed Image Retrieval Combining and Partially Fine-Tuning CLIP-Based Features,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Conditioned and Composed Image Retrieval Combining and Partially Fine-Tuning CLIP-Based Features,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.069339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.069339Z digest=sha256:d1b7aa2f05d4b487ce5e914a91b8916e0c0800dc65cd7500026f7a7a6aef409b

Observation c586d371-e594-4e4a-8495-2341acf003ff · outbound

This paper cites BLIP: Bootstrapping Language- Image Pre-training for Unified Vision-Language Understanding and Generation,.

Recurrence Meets Transformers for Universal Multimodal Retrieval BLIP: Bootstrapping Language- Image Pre-training for Unified Vision-Language Understanding and Generation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.071872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.071872Z digest=sha256:1274fcd03adca908145cfc144e4d958b1eab7e1e3c0b6656300362c6e28f8c0b

Observation 67842716-390b-4f85-8040-c282f7d21289 · outbound

This paper cites GENIUS: A Generative Framework for Universal Multimodal Search,.

Recurrence Meets Transformers for Universal Multimodal Retrieval GENIUS: A Generative Framework for Universal Multimodal Search,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.074201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.074201Z digest=sha256:65b1eb7796e830fdb78273827c7b88fb9003bc4176baaa940d264aade2f088ec

Observation d76ecf55-2ab5-4a1b-a724-74536f0cabbb · outbound

This paper cites Fine-grained Late- interaction Multi-modal Retrieval for Retrieval Augmented Visual Ques- tion Answering,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Fine-grained Late- interaction Multi-modal Retrieval for Retrieval Augmented Visual Ques- tion Answering,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.076615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.076615Z digest=sha256:323b535d4614d50ddaeb6a2f299ade9a037d6a326f599a57748ce6c8a9372d39

Observation 854b0c7c-1671-4879-9d49-2ddb5b578578 · outbound

This paper cites ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT,.

Recurrence Meets Transformers for Universal Multimodal Retrieval ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.078945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.078945Z digest=sha256:a6ce703fa2565376ba16db195edfc1bdc4e91f7333aa96108a04ef079e6743df

Observation ca31e2cc-583e-4813-b8c2-6bd7e9ec1f37 · outbound

This paper cites Language models are few-shot learners,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Language models are few-shot learners,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.081218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.081218Z digest=sha256:a100666abff18ec7bd23289281d39d8bc026383237f1626f8f45056442f549ff

Observation 54201e63-26b7-488c-a608-6e3c4b8df992 · outbound

This paper cites The Llama 3 Herd of Models.

Recurrence Meets Transformers for Universal Multimodal Retrieval The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.083759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.083759Z digest=sha256:7746b0d874ba4a3f7ab75fd4d326ba3eac49f5bf23f4e53aa89bad0337034f06

Observation dee484a4-50ec-4804-97e7-5fab419e8214 · outbound

This paper cites LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant,.

Recurrence Meets Transformers for Universal Multimodal Retrieval LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.086532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.086532Z digest=sha256:3644582e4e7bb6b5d50d44db0ced253d979c05b1fe5f417f7bcb44e5510551e7

Observation 98ef7cc0-2c0e-421c-82d9-a709493efacd · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs,.

Recurrence Meets Transformers for Universal Multimodal Retrieval MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.089128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.089128Z digest=sha256:3e3a0c315426d4b6ef801b29bc0effba9b0bb4ecb4de5bfa62c82b51b15b25f1

Observation c6fa9f74-f7c8-46f9-ac62-20150874a583 · outbound

This paper cites PUMA: Layer- Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval PUMA: Layer- Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.091490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.091490Z digest=sha256:ab9048a7d156a8732626654b625adc02206bf65316b47760ea4dd9d71e1b75e4

Observation a0c256f8-118b-46ad-8321-8af2be4a9b0c · outbound

This paper cites Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up.

Recurrence Meets Transformers for Universal Multimodal Retrieval Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.094045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.094045Z digest=sha256:8c64ddf7095520b126d4ba1d831ff63be10dc2bfe344332d87e08a76139af3e4

Observation 8045a789-af75-4125-af60-5087d9a00c38 · outbound

This paper cites Attention Is All You Need,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Attention Is All You Need,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.096789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.096789Z digest=sha256:def889c6d4340fe6e07630dcd907be8299010f6f4d3de20ad1797a4ce0daeee0

Observation 6b72423c-abb8-47e8-8258-f3dd2b7b2bb0 · outbound

This paper cites Transformers: State- of-the-Art Natural Language Processing,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Transformers: State- of-the-Art Natural Language Processing,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.099531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.099531Z digest=sha256:12bc03825b21782ae3bfcce3e20b56536eec8998358975eb81ecb194410abaf9

Observation 7dde63b2-f6c1-4adf-a628-cf4eacd7ff5f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Recurrence Meets Transformers for Universal Multimodal Retrieval LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.101907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.101907Z digest=sha256:8bed6e266343be14e86c4b7be154d23b7318400f81d49261a7d4ca69221d23ff

Observation 2ffab2c6-b60d-462a-b859-9b3d819acbe1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,.

Recurrence Meets Transformers for Universal Multimodal Retrieval An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.105110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.105110Z digest=sha256:2f20031ae5131cdb146e04435795b4dc618c99797499c9c1c7712e9faa8c02d1

Observation 1aac34ea-134c-4dbe-9b3d-69b3ab4308ba · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Training data-efficient image transformers & distillation through attention,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.107756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.107756Z digest=sha256:e52445e31fd63a313caf6b8383f61a5415e95d3d55d0d2fe4a708710cdb628e5

Observation c43b9f68-3914-4ed1-9cd5-9115b16912ea · outbound

This paper cites Scaling Vision Transformers,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Scaling Vision Transformers,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.110276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.110276Z digest=sha256:39c95dc1d1de9787375794f077e9a2207ec4f15ed7a8aeb41da3a142862e2221

Observation f26dd060-5577-4bd0-ae7b-64ca67bee48f · outbound

This paper cites Transformers in Vision: A Survey,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Transformers in Vision: A Survey,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.112765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.112765Z digest=sha256:aa16056911119a0f1ee3c2b3cbc393e168ec57fc515751441867fdf0fe66eda0

Observation 58536b43-ec6f-4ba1-8eba-aa751a0e78ba · outbound

This paper cites When attention meets fast recurrence: Training language models with reduced compute,.

Recurrence Meets Transformers for Universal Multimodal Retrieval When attention meets fast recurrence: Training language models with reduced compute,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.115088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.115088Z digest=sha256:94dfcbf891487125ae6be64bbae3644fb7861c7a069428f82b650dbf0d03615a

Observation 856e754f-dd00-483e-baa9-45b11ab4d7d7 · outbound

This paper cites Simple recurrent units for highly parallelizable recurrence,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Simple recurrent units for highly parallelizable recurrence,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.117483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.117483Z digest=sha256:49999e626f15e7e0cc9f2b18b031c9c5af218aa9e1d393ae06a16f7af08b25c0

Observation ef0b1f1f-69ff-4c99-bd98-0b5456574a66 · outbound

This paper cites The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation,.

Recurrence Meets Transformers for Universal Multimodal Retrieval The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.119939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.119939Z digest=sha256:5e05b5fcc3d99cccad0368a50ef9046a219de4f21810aa92aae76269f2565a7d

Observation ee4debea-d943-4f43-afac-92874c33d6ba · outbound

This paper cites R-Transformer: Recurrent Neural Network Enhanced Transformer.

Recurrence Meets Transformers for Universal Multimodal Retrieval R-Transformer: Recurrent Neural Network Enhanced Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.122432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.122432Z digest=sha256:e39ec1953732d172346aaf5db718f14aa8c22865f670254cf198c0381bb9b0e7

Observation ace8f110-196c-44a2-bd84-36548fc19836 · outbound

This paper cites Block- Recurrent Transformers,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Block- Recurrent Transformers,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.125375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.125375Z digest=sha256:c7fa33d5e7f072cfbfa6e53208a692890a6beccdba6d87e88b59656c9d5581fa

Observation c4ba4dda-e624-4987-b6e9-451b7b691555 · outbound

This paper cites ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction,.

Recurrence Meets Transformers for Universal Multimodal Retrieval ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.127859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.127859Z digest=sha256:ff786d6c0f14f7be022299898799a95528b935458f8683a05d0b7137f25b3b71

Observation 2eaf6806-7ed9-43a5-bf4d-23dbb1ae284b · outbound

This paper cites Lambda-Skip Connections: the Architectural Component that Prevents Rank Collapse,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Lambda-Skip Connections: the Architectural Component that Prevents Rank Collapse,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.130438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.130438Z digest=sha256:9d84f6568934ec2629631d663e3eaf6d483c92f77cb94611a42877156c0ea745

Observation eec145b5-2140-498b-8557-18373cbf9f90 · outbound

This paper cites Layer Normalization.

Recurrence Meets Transformers for Universal Multimodal Retrieval Layer Normalization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.132854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.132854Z digest=sha256:55f44815bbdbbf7527224fe202c190f6a12ad19275549ecec90aa67546565cc9

Observation 4a6ee3a3-927f-4a83-8e17-90a6c7aa00ea · outbound

This paper cites WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.135440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.135440Z digest=sha256:415bf610efb91660067086bb203b70234e783d18bcea318fc1825b9ac82844f7

Observation 1c34b95a-4a30-4a74-a8c3-80eba7e986e5 · outbound

This paper cites IGLUE: A Benchmark for Transfer Learning Across Modalities, Tasks, and Languages,.

Recurrence Meets Transformers for Universal Multimodal Retrieval IGLUE: A Benchmark for Transfer Learning Across Modalities, Tasks, and Languages,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.137956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.137956Z digest=sha256:36e17d224361b989dd173a40aa2c54aff85cad8f21d7b0cb1b6b453b8a716b97

Observation 65d7283c-b9be-4510-b79e-69cdabb5db47 · outbound

This paper cites KVQA: Knowledge- Aware Visual Question Answering,.

Recurrence Meets Transformers for Universal Multimodal Retrieval KVQA: Knowledge- Aware Visual Question Answering,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.140358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.140358Z digest=sha256:5415564585d3037e3ffc46b56a261dda2235cd58d1c654c9e2e606073bfc7d38

Observation 5d43dffc-4a4a-4ba7-8a8e-57c57a64f08c · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset,.

Recurrence Meets Transformers for Universal Multimodal Retrieval MS MARCO: A Human Generated MAchine Reading COmprehension Dataset,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.142678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.142678Z digest=sha256:665bf8bb32860a74a722bde4455d5883bc9c00754ca5bff29e32f96f6b5787d1

Observation c8383d7b-6ff1-4845-9e57-568168d14655 · outbound

This paper cites Visual Instruction Tuning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Visual Instruction Tuning,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.145450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.145450Z digest=sha256:2db46f22a52bb926d970f8c07f2285bb5d0652f3157c31d580e99ef8fe80221b

Observation 635b5415-afb3-4dce-9e2b-8540313b3e6a · outbound

This paper cites EDIS: Entity- Driven Image Search over Multimodal Web Content,.

Recurrence Meets Transformers for Universal Multimodal Retrieval EDIS: Entity- Driven Image Search over Multimodal Web Content,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.147996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.147996Z digest=sha256:228f055025c2be2b7e3700cc7965e3a202b9d5822e3a4434273e6830f0b2e131

Observation 71e30fde-2f50-40c1-ac9e-1ccd0e201500 · outbound

This paper cites Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.150370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.150370Z digest=sha256:78970c7831ede28cfc74f7d89d2aa5a8ff4fc5fe54761001b83f6c8a014f57d3

Observation 7b94b925-1f30-4816-9ccd-b347f679f5a6 · outbound

This paper cites Automatic Spatially-Aware Fashion Concept Discovery,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Automatic Spatially-Aware Fashion Concept Discovery,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.152908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.152908Z digest=sha256:f8bef23d99daea062076f2085aa3bd93ac9734396280b4d3823b55858a0f535f

Observation b1318d12-c2cf-44ba-861b-07f92d83509b · outbound

This paper cites Visual News: Benchmark and Challenges in News Image Captioning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Visual News: Benchmark and Challenges in News Image Captioning,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.155312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.155312Z digest=sha256:9d22811f26aafda556632c2668b94bd111d6705402861b8d185e12b2765c16f7

Observation 65ec9ad0-dfc7-4bcf-ae02-3a1eb54ba887 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data,.

Recurrence Meets Transformers for Universal Multimodal Retrieval DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.157984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.157984Z digest=sha256:73b6baddd126e923e3eabdda07a397cb14371d95bf5389da90632578d53f8469

Observation b6ad4bbe-f9a0-4346-a4ee-8ea6a3abf05f · outbound

This paper cites ADAM: a Method for Stochastic Optimiza- tion,.

Recurrence Meets Transformers for Universal Multimodal Retrieval ADAM: a Method for Stochastic Optimiza- tion,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.160539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.160539Z digest=sha256:08063fa51f7556d7d7fdd04363f8e33e3d3da68ecc69c070e08cdba60ab4535f

Observation 55204586-01e9-45d9-a426-4000ae0ad663 · outbound

This paper cites Billion-Scale Similarity Search with GPUs,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Billion-Scale Similarity Search with GPUs,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.162988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.162988Z digest=sha256:fa181f020c46835a39c183357666307b9ef7f56a0b05478d9b56d6f208fc2b92

Observation abf2788a-b98e-49be-9b54-5802401b71c7 · outbound

This paper cites Multi- Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Multi- Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.165771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.165771Z digest=sha256:9a4c52d0adb82f9599524c40576227d148f70fc784cd9e9a8fc0df3887d6e0a1

Observation e120b265-e3e4-4676-a724-883245c14fe4 · outbound

This paper cites Wiki-LLaV A: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Wiki-LLaV A: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.168123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.168123Z digest=sha256:8dc4bd7e67f4c33618e0c4afac100a60c7b94f09a7ef4f97778a33076ced95e9

Observation 5e67fdea-71c0-43f0-96f4-cdbef572032b · outbound

This paper cites Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Ques- tion Answering Evaluation,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Ques- tion Answering Evaluation,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.170486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.170486Z digest=sha256:7ffd482d02ac03aa8bde3fd61379e61d0e757b3d939963268304468ffb3ce91a

Observation 19c4cf41-1285-4542-967d-ddbd300ddf19 · outbound

This paper cites Aug- menting Multimodal LLMs with Self-Reflective Tokens for Knowledge- based Visual Question Answering,.

Recurrence Meets Transformers for Universal Multimodal Retrieval Aug- menting Multimodal LLMs with Self-Reflective Tokens for Knowledge- based Visual Question Answering,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.172726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.172726Z digest=sha256:e68e81308c10c7ae5894eb7c1b31fcda28e5818f48f0a4dad3970dd8f2da63fd

Observation ada187f9-6ea1-4a7e-8091-d71603d7fb3d · outbound

This paper cites RoRA-VLM: Robust Retrieval-Augmented Vision Language Models.

Recurrence Meets Transformers for Universal Multimodal Retrieval RoRA-VLM: Robust Retrieval-Augmented Vision Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.175173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.175173Z digest=sha256:09a8be4d42b9e77ae0d4ca2e9b15c1285360f8aff6058d0030fe90da18fb0a8a

Observation ccaa28e4-9ba2-4657-b4da-72f4b444d3bd · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge,.

Recurrence Meets Transformers for Universal Multimodal Retrieval EchoSight: Advancing Visual-Language Models with Wiki Knowledge,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.177829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.177829Z digest=sha256:68b41db4833e4a42d47642d43ea9f26adc3798da56a89fd1b233707ee32f6e27

Observation d2222c44-7745-4fec-99f5-5a4fc3d02b47 · outbound

This paper cites Towards General Continuous Memory for Vision-Language Models.

Recurrence Meets Transformers for Universal Multimodal Retrieval Towards General Continuous Memory for Vision-Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.180245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.180245Z digest=sha256:c7c2f5f5c90c8450594304b0d4babf0227db98ae6cfae14d21cb3eb79136f9ae

Observation c3f50af2-666a-4c96-9b4f-03dc62a0dc58 · outbound

This paper cites mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA.

Recurrence Meets Transformers for Universal Multimodal Retrieval mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.182830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.182830Z digest=sha256:4f2c333a0b9a74a96c3664b5a7fba05de0e5ae00e044f140d5260c29142008dc

Observation d15fa04c-1442-4229-8cb8-15a2dbf51316 · outbound

This paper cites BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models,.

Recurrence Meets Transformers for Universal Multimodal Retrieval BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.185613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.185613Z digest=sha256:02684e00318c3ad58807abeac8cdf5e877518fb2fabd11122ab63f27731cb715

Observation 8becf586-4688-4168-bb6c-96adbf4464fb · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning,.

Recurrence Meets Transformers for Universal Multimodal Retrieval InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.188123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.188123Z digest=sha256:8e86cd5829de5cdf9089e2ac6774cf861f917f2547c2979c3a53f2e552b4de0a

Observation 0e253ce2-8f5e-42ec-a572-d779bd89a12c · outbound

This paper cites WebQA: Multihop and Multimodal QA,.

Recurrence Meets Transformers for Universal Multimodal Retrieval WebQA: Multihop and Multimodal QA,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.190546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.190546Z digest=sha256:8cdd0cd92514992cceb82413cdb85252ffe6c234d5c24f652ccd1e27fb44ce26

Observation ff4ca8b7-b9c0-4d9c-a4d1-7e79e43f8e12 · outbound

This paper cites The Eleven.

Recurrence Meets Transformers for Universal Multimodal Retrieval The Eleven

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.193095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.193095Z digest=sha256:449655644efea8178679d89869aab347fc217e7275a4605bfff11f67ee24f64a

Observation 92eb5617-04a8-47f6-ab27-745cf85812f7 · outbound

This paper cites Eu- phorbia pulcherrima.

Recurrence Meets Transformers for Universal Multimodal Retrieval Eu- phorbia pulcherrima

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.196053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.196053Z digest=sha256:5a50bd7556105806ca4bf06d41453a62e411937d716accb5b865e1b55415f6b6

Pith citing papers

Observation ac02a269-313c-4e29-8430-233915e576f7 · inbound

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval cites this paper.

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval Recurrence Meets Transformers for Universal Multimodal Retrieval

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:14.781939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T22:42:53.631688Z digest=sha256:bc2b7cf1cc271efebc300afa93ebb6142e708d3d4cbfabfd8ef85c162fb8b21e