Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:23.625101Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2411.18953.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:23.625101Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:36:19.810821Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:23:19.246908Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5412e09b-4750-44a0-81b6-866377de6ce9 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fda6a710-5d53-4d4f-8cb7-a5448786865f · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Pengi: An audio language model for audio tasks,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1cce0362-de32-4d59-aae4-00b756525d12 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A Survey of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f309895-9776-45e8-bca3-ad3d57ba8667 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Prompting large language models with speech recognition abilities,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1c6c905b-d0bd-4504-827b-2fdac303ae4e · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioLDM: Text-to-audio generation with latent diffusion models,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6111ccf1-dcdc-4d26-bb53-b98f14742a81 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be342558-9bb6-4582-9953-923f192bc141 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio retrieval with natural language queries: A benchmark study,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 153faeac-14dd-41d7-907e-b1f231193a76 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audiolog: LLMs-powered long audio logging with hybrid token- semantic contrastive learning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb9a2217-7c62-492d-87ad-1945a3870248 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c3c30a-c2b8-40b7-8c22-efc1038db471 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Sparks of Large Audio Models: A Survey and Outlook
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70798fde-685b-4eb8-8f9d-94e74084282b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio-Language Datasets of Scenes and Events: A Survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4c5665-ff4c-4ec6-a60a-1029d8409f5d · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4d9f7c2f-491f-4c55-9202-38437eb3e36b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CLAP learning audio concepts from natural language supervision,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c452f5cb-cd64-4616-8bbc-94f3eefa1848 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8187fd51-6258-45fb-b0c5-6d2453337a10 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Auto-ACD: A large-scale dataset for audio-language representation learning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ee79af6-7f1e-4d93-86f8-67acf816466b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34df367f-df9c-4851-9380-ba9f08c25cf3 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Listen, think, and understand,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37b41303-ac18-4d8b-9192-6c210e4aa9a3 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Joint audio and speech understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 19b40a60-6b0d-4cc3-af48-8d55b32759de · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7598409f-d473-4ab0-a335-7a5eab74f37c · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio Set: An ontology and human-labeled dataset for audio events,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2cd386a-2a6b-4180-ae51-059a378cbc13 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AudioCaps: Generating captions for audios in the wild,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 512e81bb-4ea8-4c8b-a1af-67228eb499cb · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Clotho: an audio captioning dataset,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b82374e-64cf-48ec-ab1b-32d3124a7b24 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Diversity and bias in audio cap- tioning datasets,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6087bed5-8891-48d3-83c8-9abcdd51d92e · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Mistral 7B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815b8c94-99bb-43bc-9be0-a4e1b8331bb4 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb17653-531d-4f7c-8172-406fa979f9af · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd9fd254-456a-48c5-b36e-da9558d15077 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2cf8811-6424-4476-8026-61e30ce46a5c · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models AI Chains: Transparent and controllable human-ai interaction by chaining large language model prompts,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 748f18a3-2bbb-45e2-b8c8-d796dbf03383 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Classify first, and then extract: Prompt chaining technique for information extraction,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e668f465-a80e-4cc5-982e-254274855cce · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a600e20-b37a-4f8e-aefe-129ce631f31a · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automated audio captioning with recurrent neural networks,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e55bcf0f-efbe-4541-b911-c6493ac10e93 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automated audio cap- tioning: An overview of recent progress and new challenges,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5cdef9eb-96ed-4ca9-b7e7-bacbc9f3636d · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Beyond the Status Quo: A contemporary survey of advances and challenges in audio captioning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6f33fba-db54-44e9-ae5e-222ca397ce3b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa36887d-76ef-4160-944d-771faaf6b155 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BART: Denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85acfd53-8061-4388-8457-9f498f86c114 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BLEU: a method for automatic evaluation of machine translation,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3e43268d-f253-4f4b-8beb-59223e87714c · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models ROUGE: A package for automatic evaluation of summaries,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae0e2482-55a8-46d8-9cf0-f0a94091ee2b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99f8ee5e-7073-4463-a19d-8dd8ec3bdf39 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CIDEr: Consensus- based image description evaluation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7effa66e-3450-4e30-a92b-6431f4de1500 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models SPICE: Semantic propositional image caption evaluation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0defaa85-deb6-4f3b-94db-e971433ea7c4 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Improved image captioning via policy gradient optimization of SPIDEr,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b17b16cf-a201-4da7-9d4a-8e4c5e4a9e3b · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio cap- tioning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea6a31d0-8fb8-451d-babe-cc4b4977a19c · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ca402966-e9f1-484e-8561-a58021840176 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Taming Data and Transformers for Audio Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b14bb3e-303a-4526-85d9-a59f16d7b0c3 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Bridging Language Gaps in Audio-Text Retrieval
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47cbc6da-e85c-41e2-8fe3-0814f149b2bf · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Language-based Audio Retrieval Task in DCASE 2022 Challenge,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 716f0ad4-8b3a-4022-81cd-22d38ca135eb · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models On metric learning for audio-text cross-modal retrieval,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c239b82e-ec42-407c-a247-6a915d972e0a · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Audio-text retrieval in context,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b09eaca3-bdf7-4705-bced-ae5a5300e589 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Momentum contrast for unsupervised visual representation learning,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee41d20-b763-4d4b-b797-5cdf188e1bfe · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models BERT: Pre- training of deep bidirectional transformers for language understanding,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc61dd2-70c1-4e9a-b933-87385846af75 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c987fde5-e302-4034-b548-a4eb27d8b031 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CED: Consistent ensemble distillation for audio tagging,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ed14750-33cc-4f36-a9a6-d3f5d8ebf5de · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models SONAR: Sentence-Level Multimodal and Language-Agnostic Representations
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee550096-6068-4f51-96ee-bfd8a58a5077 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Zero-shot audio classification based on class label embeddings,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ec1a6be6-1f34-41c2-9881-b14f3e405b75 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Zero-shot audio classification via semantic embeddings,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c63467ed-dff0-4033-91c6-9eb6d43fd350 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models A dataset and taxonomy for urban sound research,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a70be6d-78e9-42d0-829f-28064a60410f · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models ESC: Dataset for environmental sound classification,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fac8381-5f92-4251-bfee-977d69e70d3d · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Common V oice: A massively-multilingual speech corpus,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4f10cf4-21ac-40f6-b90d-aa60b18d2462 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models CREMA-D: Crowd-sourced emotional multimodal actors dataset,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eff404b-cbdf-4265-84e2-48aa01e9330f · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c6807c65-e7c8-47e2-9b09-65f31a822756 · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Automatic musical genre clas- sification of audio signals,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8237517-42ab-46ae-8f75-d837defccb9f · outbound
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models OpenMIC-2018: An open data-set for multiple instrument recognition,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e0dc5f4c-f6ea-4aff-9abf-fa28969b7b5a · inbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 137
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c2d147-4f94-4224-a045-a88b69bd12b0 · inbound
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4795b6c-b14e-4649-ba61-c559c05a85d5 · inbound
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 03c50d2d-29da-42ee-90ed-d6ba7bb606b0 · inbound
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.