Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:18.773662Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.14445.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:18.773662Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:15.642048Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:13:15.097149Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7d56aa-8fac-4490-b5b9-b7fe1844e710 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6de713e2-bfdf-48a2-bedc-8b0e509dd27e · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a587ad28-d860-45e0-a8c9-83d37b4d677f · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval (1) In thetrainingstage, by unifying multimodal representations into the same embedding space, Vela improves multimodal embeddings usingonly contrastive learning on text pairs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36072b47-8846-45f2-b5f4-cc95fe952309 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbb6639b-791e-494a-bc7f-299299a3f86e · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Datasets For the training data, we use NLI [18], which contains approx- imately 273k sentence pairs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 133ae229-6bb5-4a42-b85a-323034cf5a79 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval in one word
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58f87e0e-9792-458c-b02e-1d7cd375e1a1 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6076796a-ac42-411b-95aa-81896bc425c8 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Pioneer” and “Lead- ing Goose
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56e5fd14-df51-4958-83a5-9fbf72c1ac10 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Natural Language Supervision for General-Purpose Audio Representations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8861ca3d-b434-43fd-9a37-004cdc0e5959 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Large-scale contrastive language-audio pretrain- ing with feature fusion and keyword-to-caption augmentation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63575974-ea51-4c1e-b2d0-1f91f3a867ce · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4795b6c-b14e-4649-ba61-c559c05a85d5 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0db3dfc-e0cf-46ae-90ce-2fd5d84c5395 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Auto-acd: A large-scale dataset for audio-language representation learning,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae836ea5-c0ba-4d47-ae03-83dd62965443 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Grounding language models for visual entity recog- nition,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44f06b76-705b-4f28-87f0-59fca8532628 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Improving clip training with language rewrites,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 170d5e1b-d45d-4bd5-b171-84c2eb50ed2b · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MATE: Meet At The Embedding -- Connecting Images with Long Texts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6490dd53-93a4-4064-b166-ea18c5020fc0 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55daae4-eb1e-4855-b42d-edbaa6ebe0d3 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01cbf8ab-f77c-469a-b7aa-740d2e687b6a · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval E5-V: Universal Embeddings with Multimodal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4ba41e-5177-4dae-9255-b868ae0bb404 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Learning transferable visual models from natural language supervision,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f15b79a-3f1c-400c-9acc-0ca07c27cef8 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Mind the gap: Understanding the modality gap in multi-modal con- trastive representation learning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b30d69b7-e188-43b7-a9ef-617c3899c70a · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval DefSent: Sentence Embeddings using Definition Sentences
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bcfe10a-71cd-4a1b-ad24-1104ed40c164 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval GPT-4 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2047d02-e3f2-4446-81cb-9199c939b801 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Gemini: A Family of Highly Capable Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8427953-163d-46a7-909c-1a0335f0905f · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Breaking the length barrier: Llm-enhanced ctr pre- diction in long textual user behaviors,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9398214e-f5a2-49e6-b4a4-c6df3d59032e · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval SimCSE: Simple Contrastive Learning of Sentence Embeddings
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a539ff0-0b44-4f33-989c-e56e5cffcd67 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Clotho: An audio cap- tioning dataset,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380a3921-f767-4491-8122-9e123b36e99a · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Fsd50k: an open dataset of human-labeled sound events,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d21cce4-bdcb-4393-9b1c-4332cefd3df7 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Audiocaps: Generat- ing captions for audios in the wild,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 71e4145e-c544-4d9b-8c55-318401a0c01d · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen2-Audio Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825f3886-9ccf-4810-a33d-33846519efae · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qwen Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf11e42-7a0b-4500-8708-691084e9f4e8 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Robust speech recognition via large-scale weak supervision,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb102d5b-958d-461e-b25e-492a2fa22f31 · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Qlora: Efficient finetuning of quantized llms,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a298787-e643-4b26-9dbb-44ad3a01906d · outbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval From Matching to Generation: A Survey on Generative Information Retrieval
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d2b08f-0ae8-4e2a-a849-14f74b4de845 · inbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 521088c5-f028-49b5-a8b3-64f85fbcaaa5 · inbound
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.