Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.533962Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2506.01111.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.533962Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.152062Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:39:38.511905Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc370d3f-ef9f-49e8-87e9-f98a4d298af9 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d7254e-d4ca-4bdf-80a0-79f7f82527d3 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d820f451-eaf9-45fe-9417-c12eb40bc215 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen2-Audio Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78b85a1-16ac-4e6c-a2d5-bd3926cfae93 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Clotho: An audio captioning dataset, 2019
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5eecb123-44b0-4203-b7d2-969117562a76 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion AudioCaps: Gen- erating captions for audios in the wild
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b35f10ac-e745-4208-a2a2-9fcf797c9f94 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Laion-audio-630k dataset, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2451f32d-b313-47aa-b839-3bddeac183c0 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, Yuexian Zou, and Wenwu Wang
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a50eb9d-dd42-4b28-b624-f47d9fb5e24e · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, Woon- Seng Gan, and Jianfeng Chen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fc93b794-3e01-472c-876e-f7598d659ac8 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Auto-acd: A large-scale dataset for audio-language representation learning, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6644e575-145e-4f33-a8c6-998eed237ab5 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Plumbley, and Wenwu Wang
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2595efec-ca1b-4177-baf7-4db62f5eb713 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Air-bench: Benchmarking large audio-language models via generative comprehension, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7526a4a6-4242-4a4a-92b1-275753e75560 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 623cae2a-1a2d-484e-bd27-603b159ac3dc · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Logothetis, and Stefano Panzeri
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 66fa02db-07a5-46a8-b7b0-df63937ee244 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Ernst and Heinrich H
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f5478d8a-f51b-41a3-9ce4-9edf31aa7945 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Bregman.Auditory scene analysis: The perceptual organization of sound
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation db545bcc-c8e5-44d6-b9d6-a62097f4f62a · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fcc38876-3201-4953-8df3-a34d43f5d720 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Robust speech recognition via large-scale weak supervision, 2022
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7114b2-1ba5-47c6-a4ce-75b965456850 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion OpenMU: Your Swiss Army Knife for Music Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05f132e-b094-48e9-842c-ed8375b1003c · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen2.5-VL Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246d2c6d-8a8a-4502-adc1-9c8eb70c8476 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b530e6b-57c4-448c-a732-c3f507502ff5 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Elizalde, S
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 96a26cd4-5438-442b-8501-83c7f5e3ccbf · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 21d36ae2-4189-491f-a8d3-e1e5c3946cab · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Yeh, P.-Y
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 93b38794-9dc5-4c3f-b4e9-219efef4dac9 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934e5a04-3a41-46a3-a014-3186c6fd475d · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ee85056-7646-4160-aa5d-1a7aa678299e · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Pengi: An audio language model for audio tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 49ce8f76-5915-466e-9e36-442b9caf9ff3 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf48656-4821-4cf4-a23e-f436705725c2 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72eafdcd-927d-4425-8bee-60ce0ed98965 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Hybrid transformers for music source separation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02d9edf9-e02d-4217-9fc5-4aff04ff41d0 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Yamnet: Audio event classification
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 30d3e3b0-1c19-4c46-96f8-6d74b5b4b9b1 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Gemmeke, Daniel P
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c2a87ce3-ce16-4470-ab22-7d48c573fdd7 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Visualizing data using t-sne.Journal of Machine Learning Research, 9(86):2579–2605, 2008
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cdb47c3-095c-49ba-9dea-11f5b8c8a35a · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Hts- at: A hierarchical token-semantic audio transformer for sound classification and detection
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5a1aa497-cdd8-43e2-af63-140883b60f41 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5f3d0f4a-b6a0-48ba-a7a0-2de3f300d16d · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion • If the intensity or emotional context of the sound is conveyed (e.g., the dog barking intensely or the doorbell ringing in a rapid succession)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f76d6af2-fcf9-4a94-948d-b022ba748b5e · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion trumpets,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 942d7622-30eb-4367-ad85-108b4391203d · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 887a1896-de18-4dae-81fe-055598b02891 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2f4dff6b-52f8-4c4b-be0b-301752e4994c · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 647bcb52-c431-4d05-bcc4-bf9144e0205b · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 84987096-14f7-4a6a-93c5-df9dea7ea3c5 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a2bcbc8d-1215-46ca-a725-a5dbb990fb6d · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7fe5bfa9-5d6a-423e-bd72-555cdececd27 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion instrument
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ff0f6336-9620-4f16-b60c-9a53810b553d · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 98da82d0-c3fe-4687-b349-4a6ed02d6fe9 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7ad94d4-13eb-415c-8f32-d6641bdda7f4 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion A car’s engine roars as it accelerates
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fdc345c5-9d2c-405e-9690-b471e65f7f27 · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Audio Description
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation afbe7848-d998-486f-aed2-0171e0a7c2ae · outbound
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dca371f-5f41-4521-8e81-5b79e200dd2f · inbound
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27bde799-b3b5-4866-86e4-ecf1e2a11ab2 · inbound
EvA: An Evidence-First Audio Understanding Paradigm for LALMs FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3d27f5-dfab-44d1-80d7-b2c6cfffbafe · inbound
Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · inbound
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.