Pith. sign in

Paper Citation Record · LEDGER

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.14445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14445 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:18.773662Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:15.642048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.097149Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · outbound

This paper cites Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:15.642048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:15.642048Z digest=sha256:4f1ad54c10f654d47337bd4230b094c1620c7193431169151650ca13e68c32ef

Observation 0d7d56aa-8fac-4490-b5b9-b7fe1844e710 · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.872287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:15.734690Z digest=sha256:988305b8929687cb7ec1f3e2d84b6a3b3e24fced11f63d86641eafac71f29176

Observation 6de713e2-bfdf-48a2-bedc-8b0e509dd27e · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.849966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:15.849949Z digest=sha256:af7f0e925ed9701ef369ca3ff1f45552feab673916bf059e982a335ae232f8d0

Observation a587ad28-d860-45e0-a8c9-83d37b4d677f · outbound

This paper cites (1) In thetrainingstage, by unifying multimodal representations into the same embedding space, Vela improves multimodal embeddings usingonly contrastive learning on text pairs.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval (1) In thetrainingstage, by unifying multimodal representations into the same embedding space, Vela improves multimodal embeddings usingonly contrastive learning on text pairs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.827814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:15.952336Z digest=sha256:5d0ccd8663d5946ed8cbc8352bd75b08b5bec6e3c3201b7df8c80eef5f534b29

Observation 36072b47-8846-45f2-b5f4-cc95fe952309 · outbound

This paper cites in one word.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.790304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.044484Z digest=sha256:266d2c9eafcf9d115f3d620f6e064c17d7244b3e790d0db4174579229a7487ea

Observation cbb6639b-791e-494a-bc7f-299299a3f86e · outbound

This paper cites Datasets For the training data, we use NLI [18], which contains approx- imately 273k sentence pairs.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Datasets For the training data, we use NLI [18], which contains approx- imately 273k sentence pairs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.766028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.123834Z digest=sha256:b76d036114714b23ea9bd0a7edaaa15d5304c489a8f0c8658836b8f07e74d75e

Observation 133ae229-6bb5-4a42-b85a-323034cf5a79 · outbound

This paper cites in one word.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:23:19.727739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.218939Z digest=sha256:27fae6eabf41cc9848dbfde12eda3e2fa3b0eba48aa8bd06f28ea30fb61fd16a

Observation 58f87e0e-9792-458c-b02e-1d7cd375e1a1 · outbound

This paper cites an unresolved cited work.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:19.688912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.303912Z digest=sha256:d0283fff39789acfd2d1f2f9b74510de8a91752fe52e0bf537ba75f141d142c1

Observation 6076796a-ac42-411b-95aa-81896bc425c8 · outbound

This paper cites Pioneer” and “Lead- ing Goose.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Pioneer” and “Lead- ing Goose

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.659669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.433958Z digest=sha256:58b8cdfca540fe3bd850860bb67a996f6661086bb088cc28239db0940e6e7e93

Observation 56e5fd14-df51-4958-83a5-9fbf72c1ac10 · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Natural Language Supervision for General-Purpose Audio Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:16.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:16.501373Z digest=sha256:92c34ddf3de1d0e5294e1bf1af01e39c2223eae4058ea5b957e7c2e1ad904b29

Observation 8861ca3d-b434-43fd-9a37-004cdc0e5959 · outbound

This paper cites Large-scale contrastive language-audio pretrain- ing with feature fusion and keyword-to-caption augmentation,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Large-scale contrastive language-audio pretrain- ing with feature fusion and keyword-to-caption augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.620030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.593799Z digest=sha256:b026a65c3ee01b4b69501faee95d3ec810c60917221b248e85b52c90b94f68d8

Observation 63575974-ea51-4c1e-b2d0-1f91f3a867ce · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.588446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.724661Z digest=sha256:18c041300d1f4dac4d22896e7897043695e02b28be1503d7f4e2d3a94b967ce0

Observation d4795b6c-b14e-4649-ba61-c559c05a85d5 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.253506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.841565Z digest=sha256:5153f1caf333b4dc6c07adefd7ccdac63f307552047c72a740a3c0c221b10c2e

Observation e0db3dfc-e0cf-46ae-90ce-2fd5d84c5395 · outbound

This paper cites Auto-acd: A large-scale dataset for audio-language representation learning,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Auto-acd: A large-scale dataset for audio-language representation learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.566004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:16.970134Z digest=sha256:e51bf68fff104ff635adf3b13ba946f215ef1c24a279d8e5d6dd6803aaca9ebc

Observation ae836ea5-c0ba-4d47-ae03-83dd62965443 · outbound

This paper cites Grounding language models for visual entity recog- nition,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Grounding language models for visual entity recog- nition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.540632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:17.108875Z digest=sha256:b1af389b1986ccfa488912b385df60d9c4482bbaac817dcf7d46f2b97c342e45

Observation 44f06b76-705b-4f28-87f0-59fca8532628 · outbound

This paper cites Improving clip training with language rewrites,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Improving clip training with language rewrites,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.516980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:17.191597Z digest=sha256:2cb718304223694cb08ae2290cf7309e589bd77174535864a114e13a7ff33bf7

Observation 170d5e1b-d45d-4bd5-b171-84c2eb50ed2b · outbound

This paper cites MATE: Meet At The Embedding -- Connecting Images with Long Texts.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MATE: Meet At The Embedding -- Connecting Images with Long Texts

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.214601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:17.240558Z digest=sha256:3a3581b1be35c29e1dda57ba552da21302a6ffcdc6775eac5a79d2583d83afff

Observation 6490dd53-93a4-4064-b166-ea18c5020fc0 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.301298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.301298Z digest=sha256:c77c1f73f069b89b3cedcf30cd78f5508f9b6737415c7ba2d3ad079536093bdb

Observation f55daae4-eb1e-4855-b42d-edbaa6ebe0d3 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.368291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.368291Z digest=sha256:4f0cbd63d304ab4d8b30bf8ed9a80ee6a700fa4239486950a06d7aa3559e097d

Observation 01cbf8ab-f77c-469a-b7aa-740d2e687b6a · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.464515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.464515Z digest=sha256:d7bd17ab67d7b552bb74aadacf12629bdb865e93fe676962ee7b7143f9670cc7

Observation 5b4ba41e-5177-4dae-9255-b868ae0bb404 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.617017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.617017Z digest=sha256:694481e611658ed3fd4642479e1d3e61270cbcc39d1ff6b2af13dfeb7d643954

Observation 0f15b79a-3f1c-400c-9acc-0ca07c27cef8 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.475811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:17.715571Z digest=sha256:200a09e0389d901ede6499bd7adae899e1ccec4a0f0544822bd5b8d7b23d38e8

Observation b30d69b7-e188-43b7-a9ef-617c3899c70a · outbound

This paper cites DefSent: Sentence Embeddings using Definition Sentences.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval DefSent: Sentence Embeddings using Definition Sentences

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:19.089736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:17.848804Z digest=sha256:6323c76c5a4e19f0e9f2dadd5c329eeda470b88117c9bbd2a02c9baa6d692f0a

Observation 9bcfe10a-71cd-4a1b-ad24-1104ed40c164 · outbound

This paper cites GPT-4 Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.933544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.933544Z digest=sha256:76563b86edd1154d5e5030102dbecc0c683c4116417ee49618a5f5314a753a42

Observation a2047d02-e3f2-4446-81cb-9199c939b801 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.989473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.989473Z digest=sha256:3f9bbd1178747305ad1ca4f8beb544b9fe2e3e03f8e90a0721b2f403f5454bda

Observation d8427953-163d-46a7-909c-1a0335f0905f · outbound

This paper cites Breaking the length barrier: Llm-enhanced ctr pre- diction in long textual user behaviors,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Breaking the length barrier: Llm-enhanced ctr pre- diction in long textual user behaviors,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.450427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:18.114021Z digest=sha256:bc0f42208766c3278f71fd97962f0170aee2cccac7bf5f70f75ff8641f1508bb

Observation 9398214e-f5a2-49e6-b4a4-c6df3d59032e · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.247673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.247673Z digest=sha256:c8d51ccd4be2dc1759252b1cfd1347da6c5abd3d8772e2e427b53af2cc14309e

Observation 2a539ff0-0b44-4f33-989c-e56e5cffcd67 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Clotho: An audio cap- tioning dataset,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.346676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.346676Z digest=sha256:de9e463a142c624ac50f294e8f6fc56f9ebaad789f00a10588ca17738fb440d9

Observation 380a3921-f767-4491-8122-9e123b36e99a · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Fsd50k: an open dataset of human-labeled sound events,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.465156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.465156Z digest=sha256:937e9a432eaba65a763f6fd76d734e01ccec9ca5efdb80a800bbc0c8a0792ce6

Observation 1d21cce4-bdcb-4393-9b1c-4332cefd3df7 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Audiocaps: Generat- ing captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.394524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:18.565474Z digest=sha256:ae06113bce83d9f1fb405f846625cf7503f5cdfe8361556e789c00f2f8a84381

Observation 71e4145e-c544-4d9b-8c55-318401a0c01d · outbound

This paper cites Qwen2-Audio Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen2-Audio Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.677734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.677734Z digest=sha256:58b2568e2f01db26052d77621e521f7fc32ed451ef715c1088adf4ad8af02943

Observation 825f3886-9ccf-4810-a33d-33846519efae · outbound

This paper cites Qwen Technical Report.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.744126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.744126Z digest=sha256:f74ee576cff8b7f55e886f4e024e082f31dc1a058693e51e7be7581d9f724b29

Observation 6bf11e42-7a0b-4500-8708-691084e9f4e8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Robust speech recognition via large-scale weak supervision,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.754048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.754048Z digest=sha256:e4095cd1a4ee77b24072065835ea4605d693ee67d39cad172d910caa4e836557

Observation eb102d5b-958d-461e-b25e-492a2fa22f31 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qlora: Efficient finetuning of quantized llms,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:19.344994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:18.765749Z digest=sha256:8e920b7d83a0effbdb14e4fc03fe1053234d5c972e788dcc9f3693744b0889ed

Observation 6a298787-e643-4b26-9dbb-44ad3a01906d · outbound

This paper cites From Matching to Generation: A Survey on Generative Information Retrieval.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval From Matching to Generation: A Survey on Generative Information Retrieval

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.773662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.773662Z digest=sha256:ce9c99ecea9c4b16cf3bc5d34efc97167ce5f9a6f9f6722147bd95fe4efdc60d

Pith citing papers

Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · inbound

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval cites this paper.

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:15.642048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:15.642048Z digest=sha256:4f1ad54c10f654d47337bd4230b094c1620c7193431169151650ca13e68c32ef

Observation 521088c5-f028-49b5-a8b3-64f85fbcaaa5 · inbound

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark cites this paper.

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.098727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:09:41.068000Z digest=sha256:2238ff95f3dcbd2929f6a1d8fbae814dc7eb3fe855d19b2c2329aa4852114144