Pith. sign in

Paper Citation Record · LEDGER

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.14445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14445 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:18.773662Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:15.642048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.097149Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · outbound

This paper cites Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:15.642048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:15.642048Z digest=sha256:21c364f996c5bacc66758e3460860b8e78ce69f880ab72efe965f5f51540e9f7

Observation 0d7d56aa-8fac-4490-b5b9-b7fe1844e710 · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.872287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:15.734690Z digest=sha256:8c8100c747a62d144a03b78bee9006a28e7b88dddb89dd1f7d5fd4171b96753f

Observation 6de713e2-bfdf-48a2-bedc-8b0e509dd27e · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.849966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:15.849949Z digest=sha256:f45cfff521d1972bc2eba9efbd8695768a72c10619aa9ea96ccbc7a23646b37d

Observation a587ad28-d860-45e0-a8c9-83d37b4d677f · outbound

This paper cites (1) In thetrainingstage, by unifying multimodal representations into the same embedding space, Vela improves multimodal embeddings usingonly contrastive learning on text pairs.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval (1) In thetrainingstage, by unifying multimodal representations into the same embedding space, Vela improves multimodal embeddings usingonly contrastive learning on text pairs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.827814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:15.952336Z digest=sha256:99496268076d61e46ba99747c1f9259eca9d44d72e82a18708960fe714d20400

Observation 36072b47-8846-45f2-b5f4-cc95fe952309 · outbound

This paper cites in one word.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.790304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.044484Z digest=sha256:a55686c759eb7a1dfe42d674d13d927bb9bc2b465112b8c9c2b51d373ae24966

Observation cbb6639b-791e-494a-bc7f-299299a3f86e · outbound

This paper cites Datasets For the training data, we use NLI [18], which contains approx- imately 273k sentence pairs.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Datasets For the training data, we use NLI [18], which contains approx- imately 273k sentence pairs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.766028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.123834Z digest=sha256:3c1b3242ace4b8fef12c33fdba0bd613d432a26fe438ee2b4b04b7a24bc9473d

Observation 133ae229-6bb5-4a42-b85a-323034cf5a79 · outbound

This paper cites in one word.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:23:19.727739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.218939Z digest=sha256:7de47abf1cbdd51125f905be5c27f3fbef1f78c97663bdc91ddf752fbfea67b6

Observation 58f87e0e-9792-458c-b02e-1d7cd375e1a1 · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.688912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.303912Z digest=sha256:8ca967fad539c880e1b6165cb2b7b4df7095435a2a628f7b8a3471b528f2957c

Observation 6076796a-ac42-411b-95aa-81896bc425c8 · outbound

This paper cites Pioneer” and “Lead- ing Goose.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Pioneer” and “Lead- ing Goose

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.659669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.433958Z digest=sha256:fd86111dcbe8128a8db64b5bd457d39242cdefa42236bb1b111156b08a22ec8a

Observation 56e5fd14-df51-4958-83a5-9fbf72c1ac10 · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Natural Language Supervision for General-Purpose Audio Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:16.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:16.501373Z digest=sha256:8ca48ca46040b6bf07036847af9c167765ba69eebb255914f08e59fa570d064c

Observation 8861ca3d-b434-43fd-9a37-004cdc0e5959 · outbound

This paper cites Large-scale contrastive language-audio pretrain- ing with feature fusion and keyword-to-caption augmentation,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Large-scale contrastive language-audio pretrain- ing with feature fusion and keyword-to-caption augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.620030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.593799Z digest=sha256:12f08eaec529041f0467fa82ed3292d455ab4c9e40afa072063c3daa6f122222

Observation 63575974-ea51-4c1e-b2d0-1f91f3a867ce · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.588446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.724661Z digest=sha256:9aa3f818c588e86fd7de2c12e38b6f04e94b25629a2b20bbd35b5a5ab696112c

Observation d4795b6c-b14e-4649-ba61-c559c05a85d5 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.253506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.841565Z digest=sha256:33fbe1634c2f54187c2aeb14f7aa64ddb253186c99c647be7fa1d297c5d6362d

Observation e0db3dfc-e0cf-46ae-90ce-2fd5d84c5395 · outbound

This paper cites Auto-acd: A large-scale dataset for audio-language representation learning,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Auto-acd: A large-scale dataset for audio-language representation learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.566004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:16.970134Z digest=sha256:09ad3b59fb9faa86c051aa05f28f0236c3ca63e882cac5f29ef87d750dbbceac

Observation ae836ea5-c0ba-4d47-ae03-83dd62965443 · outbound

This paper cites Grounding language models for visual entity recog- nition,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Grounding language models for visual entity recog- nition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.540632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:17.108875Z digest=sha256:538513b53060a4a6e99fe76e77d0b82b6e06a116e5fd537ef7e2da17f9cd8ff3

Observation 44f06b76-705b-4f28-87f0-59fca8532628 · outbound

This paper cites Improving clip training with language rewrites,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Improving clip training with language rewrites,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.516980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:17.191597Z digest=sha256:c6b0880e986263461850afe86c6e9e50d5709837de70d405137a3c9b91907bcf

Observation 170d5e1b-d45d-4bd5-b171-84c2eb50ed2b · outbound

This paper cites MATE: Meet At The Embedding -- Connecting Images with Long Texts.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MATE: Meet At The Embedding -- Connecting Images with Long Texts

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.214601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:17.240558Z digest=sha256:fcec5f331f946b8b5e11b5109fdb3d812017e0ef04c377754d0953660e279ea1

Observation 6490dd53-93a4-4064-b166-ea18c5020fc0 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.301298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.301298Z digest=sha256:230c80d6e2a0b35b6f7ca9fdd06a2c78cff7d912cbfb9271cabd63924fd3160c

Observation f55daae4-eb1e-4855-b42d-edbaa6ebe0d3 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.368291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.368291Z digest=sha256:3a42db3d14647941c46f9868646eba5c04e1ccba53d49d0baa665a09d1fa04e9

Observation 01cbf8ab-f77c-469a-b7aa-740d2e687b6a · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.464515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.464515Z digest=sha256:4554fd6e7434c6b17af0f8cf0df6d8ba3dbe6546ed524403c44052e7d976c678

Observation 5b4ba41e-5177-4dae-9255-b868ae0bb404 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.617017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.617017Z digest=sha256:c23009c109a50cba9f7d5ad0bb22219eb9121cf2e75685fb4f8b0c798c8c4e0c

Observation 0f15b79a-3f1c-400c-9acc-0ca07c27cef8 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.475811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:17.715571Z digest=sha256:7c0a5dbd224f8631dd18e7952a3bb9bba65547b854117c6b573f48fdfc10cad8

Observation b30d69b7-e188-43b7-a9ef-617c3899c70a · outbound

This paper cites DefSent: Sentence Embeddings using Definition Sentences.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval DefSent: Sentence Embeddings using Definition Sentences

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.089736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:17.848804Z digest=sha256:fd2dc0bb2523921e33ec8229adbd5f5be8820ba0095636b156dda43477de9dc4

Observation 9bcfe10a-71cd-4a1b-ad24-1104ed40c164 · outbound

This paper cites GPT-4 Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.933544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.933544Z digest=sha256:25bbd1cb08754da55d3f132b28810de43cbbb1a2971a83961229504b3848a52b

Observation a2047d02-e3f2-4446-81cb-9199c939b801 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.989473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.989473Z digest=sha256:8ed48a0056e14b40e40c386a42ced9d5f5d1cee1bd2ae986be5ba74cfc7c15a0

Observation d8427953-163d-46a7-909c-1a0335f0905f · outbound

This paper cites Breaking the length barrier: Llm-enhanced ctr pre- diction in long textual user behaviors,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Breaking the length barrier: Llm-enhanced ctr pre- diction in long textual user behaviors,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.450427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:18.114021Z digest=sha256:d5a6b376d033cd768f856f6fdccaeff01e17ce94ab6764ff5efbea06427db756

Observation 9398214e-f5a2-49e6-b4a4-c6df3d59032e · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.247673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.247673Z digest=sha256:ea4460d04df767fddff5eecad9a217051ea886bf1c2bd84f9b79f213d26687c0

Observation 2a539ff0-0b44-4f33-989c-e56e5cffcd67 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Clotho: An audio cap- tioning dataset,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.346676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.346676Z digest=sha256:13658ef30fbd5f676ee1e7a55798a7f237072b4e40309b84eaaff032fd9b5fa6

Observation 380a3921-f767-4491-8122-9e123b36e99a · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Fsd50k: an open dataset of human-labeled sound events,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.465156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.465156Z digest=sha256:f3ae2aaad7490f6feb8f1380630b1436ffe4027bd45816845ab2a3ca78b9d014

Observation 1d21cce4-bdcb-4393-9b1c-4332cefd3df7 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Audiocaps: Generat- ing captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.394524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:18.565474Z digest=sha256:489584b7534b8cffba82cdaef53cdb16373b4f1c8897af07166417c371924585

Observation 71e4145e-c544-4d9b-8c55-318401a0c01d · outbound

This paper cites Qwen2-Audio Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen2-Audio Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.677734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.677734Z digest=sha256:928a45dcc50a1e8251b1acc0cac51274d40d4ea2dc2234572787dd8c3be2eef7

Observation 825f3886-9ccf-4810-a33d-33846519efae · outbound

This paper cites Qwen Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.744126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.744126Z digest=sha256:4921683d3eb9c2ce551c2d71c1de1c6c1fbaa71dddeee38ec250c31f8b81a835

Observation 6bf11e42-7a0b-4500-8708-691084e9f4e8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Robust speech recognition via large-scale weak supervision,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.754048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.754048Z digest=sha256:3ac5904c15671ae5ef96cf4d9f1de9777f1a3683634962feb159851631bf6699

Observation eb102d5b-958d-461e-b25e-492a2fa22f31 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qlora: Efficient finetuning of quantized llms,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.344994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:23:18.765749Z digest=sha256:0c73955baf97661595e4f9c4031bd74cf8b92084866141395fb17970ba290525

Observation 6a298787-e643-4b26-9dbb-44ad3a01906d · outbound

This paper cites From Matching to Generation: A Survey on Generative Information Retrieval.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval From Matching to Generation: A Survey on Generative Information Retrieval

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.773662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.773662Z digest=sha256:9f04eb8f208808a2b1e3dea326128dd7ae0953016730b4185cd6812ba3f688f9

Pith citing papers

Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · inbound

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval cites this paper.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:15.642048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:15.642048Z digest=sha256:21c364f996c5bacc66758e3460860b8e78ce69f880ab72efe965f5f51540e9f7

Observation 521088c5-f028-49b5-a8b3-64f85fbcaaa5 · inbound

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark cites this paper.

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.098727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:09:41.068000Z digest=sha256:2edc8a39b768071e37f3f974746a229ca691362c1675866df206b191a8adc6ed