Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:18.684100Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2505.13880.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:18.684100Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:18.478312Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:16:56.616742Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b801ed0-5f3c-4dd9-9a3e-55e117ebf0dd · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Their ability to reason and generate coherent outputs stems from ex- tensive pretraining on large-scale text data
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa18aeb5-ef61-4179-908c-28c56e5dcd78 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 949008e4-e063-4e42-8ca0-7ab641bd9744 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Data Specifications U-SAM is trained and evaluated on diverse datasets across mul- tiple tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 72900f29-989a-40e0-a6ef-15acc96fef61 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding By lever- aging LoRA for efficient fine-tuning and incorporating TAPM and SACLM, U-SAM achieves robust audio-text alignment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2858148a-4edf-44ef-9f2c-f134edaedd8c · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding A Survey of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d7db24-b4c0-4d93-8c14-49a199065a63 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Flamingo: a visual language model for few-shot learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1a50dbf1-11b9-4e69-a8af-e4c917f840dc · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4463d5-c135-45f6-9890-36099b6f7009 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f082a78-9ad0-464e-b182-188f91b5ef80 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Clap learning audio concepts from natural language supervision,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e551f007-3d86-4269-aaf7-c26bff2feb44 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audio Retrieval with WavText5K and CLAP Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c8e2ed-226a-4d1d-9733-a37bb13c48b6 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cffd971d-ed18-4443-8268-d738af0d7b82 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf46e0a1-1032-471e-8294-265091980d60 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69549a36-397c-4536-826f-04325f318506 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Qwen2-Audio Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a125b666-8ec8-4d97-9aa1-50b8fcb6568f · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Robust Speech Recognition via Large-Scale Weak Supervision
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25bff77-2590-460e-a26a-3723851efe18 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Pengi: An audio language model for audio tasks,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94127408-9008-48ba-b05f-1a1bce4b8d67 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Listen, Think, and Understand
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2a38829-a7f6-4068-bbc6-558431cac7c4 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8802c7f4-0186-476a-b0fe-cbb3e645fb87 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding AST: Audio Spectrogram Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211f1bba-6f45-45f6-8c2e-f5268134fef3 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a362eff2-7e00-406b-9f0a-7a568ed037f8 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35cb5431-fdf8-4f45-8f21-716181c6a114 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Why do speech language models fail to generate semantically coherent outputs? a modality evolving perspective,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5296f60-2df1-47ed-b05a-5e2a3ae8e6cf · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2414703c-6999-4012-8c63-c22a39ae5c14 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Linguistic-aware patch slimming framework for fine-grained cross-modal alignment,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be269beb-3836-495f-8c8b-4e0ce6276e12 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Ced: Con- sistent ensemble distillation for audio tagging,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c6e7de2-3746-4edb-a482-73574fd05956 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Mert: Acoustic music under- standing model with large-scale self-supervised training,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 12cce0b4-d670-4985-b429-c53551101ef2 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding A Survey on Mixture of Experts in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6533d0b6-0cca-4108-9ab5-0eff44eb098d · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding LoRA: Low-Rank Adaptation of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7167e6a3-e5cf-421a-aa76-011e4b1f039a · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Triplet loss in siamese network for object tracking,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e29a17fd-129a-4302-92db-4a12a8bc118a · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Lib- rispeech: an asr corpus based on public domain audio books,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553cc07a-e3be-4f57-9c42-2545c5dff3fd · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a53424c-e508-457b-9fbd-452912839637 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audiocaps: Generat- ing captions for audios in the wild,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9967068-d8f7-4167-aad6-c684e2703bf4 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Clotho: An audio cap- tioning dataset,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7537a4dd-5b57-47b2-a7c4-b6f97feb39bc · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a889230-92b2-48fc-a6c2-4d9c44fb1199 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding MusicLM: Generating Music From Text
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10244416-cce3-4e67-b80b-017496249c48 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Decoupled Weight Decay Regularization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13d6ade-9802-4d6d-aa2e-38110ddcb4c5 · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding LP-MusicCaps: LLM-Based Pseudo Music Captioning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75d3033-feb7-48f1-9d46-f9d3783d86bc · outbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding Audiogpt: Understanding and generating speech, music, sound, and talking head,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa18aeb5-ef61-4179-908c-28c56e5dcd78 · inbound
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · inbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.