Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:16.670571Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2504.14076.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:16.670571Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T23:50:12.369547Z
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69f52435-176f-4775-a41d-00be433644ec · outbound
Transformation of audio embeddings into interpretable, concept-based representations PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0cedf91b-16a7-4d26-9836-1573cc788219 · outbound
Transformation of audio embeddings into interpretable, concept-based representations HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91270124-cbe5-43d5-b098-362617de1c58 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Pengi: An audio language model for audio tasks,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ff7f4efe-477f-4132-9490-ca7aafdd242f · outbound
Transformation of audio embeddings into interpretable, concept-based representations Learning transferable visual models from natural language supervision,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 550fb7bd-3af7-4156-b93c-7edadac2c65a · outbound
Transformation of audio embeddings into interpretable, concept-based representations CLAP: learning audio concepts from natural language supervision,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84048c50-14ca-4591-ad75-e06e42feaa43 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59e8959c-14d4-4413-8a69-0c4075e94cfd · outbound
Transformation of audio embeddings into interpretable, concept-based representations Natural language supervision for general-purpose audio representations,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4ba268b-eeee-4b15-92ac-25bf84a46d74 · outbound
Transformation of audio embeddings into interpretable, concept-based representations European union regulations on algorith- mic decision making and a “right to explanation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5cbd2d83-aaae-4171-81a3-a02f51c450bd · outbound
Transformation of audio embeddings into interpretable, concept-based representations Interpreting CLIP’s image representation via text-based decomposition,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c40411c-8aff-412d-a7b7-5d0ac3f9da7d · outbound
Transformation of audio embeddings into interpretable, concept-based representations Interpreting CLIP with sparse linear concept embeddings (SpLICE),
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e1a20157-e0fc-47a2-a996-3a6bc1c21024 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Disentangling visual and written concepts in CLIP,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91b2dbbf-5d76-4c88-b4a6-7b13a01251e5 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Toward interpretable music tagging with self-attention,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 226d05b9-2013-4288-ac8a-6d0e97a6676f · outbound
Transformation of audio embeddings into interpretable, concept-based representations Interpreting and explaining deep neural networks for classification of audio signals,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c10b0f4-9cad-4e6e-9ef8-afc14b8ad4b1 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Why are speech spectrograms hard to read?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 430f39fb-4d59-4db8-9532-61908fefe85f · outbound
Transformation of audio embeddings into interpretable, concept-based representations SPES: Spectrogram perturbation for explainable speech-to-text generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae3003f0-8f51-4b10-b8e1-a949cf023418 · outbound
Transformation of audio embeddings into interpretable, concept-based representations ”Why should I trust you?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eeefddc4-e246-42a9-aa3e-a24ba9493060 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Local interpretable model- agnostic explanations for music content analysis,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35ed7973-de74-4cdc-9bf2-f07085a7a77b · outbound
Transformation of audio embeddings into interpretable, concept-based representations audioLIME: Listenable Explanations Using Source Separation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59209025-de0d-4587-9643-f0021f9bf11b · outbound
Transformation of audio embeddings into interpretable, concept-based representations Listen to interpret: Post-hoc interpretability for audio networks with NMF,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 827e0a37-5354-4e5b-a981-cc9c000b7ec6 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Towards automatic concept-based explanations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e36c3f0-e4d3-48e7-bf88-2dd38fbf217b · outbound
Transformation of audio embeddings into interpretable, concept-based representations Concept bottleneck models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e68192a7-ce59-442d-b1f7-43edcb5081d5 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Label-free concept bottleneck models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f77d455e-5339-409f-8938-0830d8c3971c · outbound
Transformation of audio embeddings into interpretable, concept-based representations Post-hoc concept bottleneck models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9d5246df-b676-449f-b015-d3eec045d3fc · outbound
Transformation of audio embeddings into interpretable, concept-based representations Information maximization perspective of orthogonal matching pursuit with applications to explain- able AI,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2ce901f-a1da-4c0f-856b-7c844eb110d9 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Sparse Linear Concept Discovery Models ,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ec2f441-fb6f-4ffe-ae29-be537c52cf8c · outbound
Transformation of audio embeddings into interpretable, concept-based representations CLIP-Dissect: Automatic description of neuron representations in deep vision networks,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a460c868-749e-4ece-9c68-88183b6bbf72 · outbound
Transformation of audio embeddings into interpretable, concept-based representations FSD50K: An open dataset of human-labeled sound events,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f5e55ee-4fbe-4574-8d38-2cfda4822559 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Concept-based explanations using non- negative concept activation vectors and decision tree for CNN models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 42805c1e-0f0c-47cb-9951-27485a14a83e · outbound
Transformation of audio embeddings into interpretable, concept-based representations Invertible concept-based explanations for CNN models with non- negative concept activation vectors,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39e9e94a-2cd4-4391-bc0b-a2cba171151e · outbound
Transformation of audio embeddings into interpretable, concept-based representations A dataset and taxonomy for urban sound research,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ece9a44-34ee-460d-8aed-12d991c85e2e · outbound
Transformation of audio embeddings into interpretable, concept-based representations Sound event detection in the DCASE 2017 challenge,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6924f4d-62cf-40f7-a96c-e2e0661676b2 · outbound
Transformation of audio embeddings into interpretable, concept-based representations ESC: Dataset for Environmental Sound Classification,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c45c692-7593-454c-ba23-905b89f8cbce · outbound
Transformation of audio embeddings into interpretable, concept-based representations Audio Set: An ontology and human-labeled dataset for audio events,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation da1718ea-4e25-4902-b11c-fa922fb2482c · outbound
Transformation of audio embeddings into interpretable, concept-based representations V ocalsound: A dataset for improving human vocal sounds recognition,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6bf75b29-0397-4e47-ad26-79d70c49b43d · outbound
Transformation of audio embeddings into interpretable, concept-based representations WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 92713099-43f7-49a9-8ea4-29ed301e4eea · outbound
Transformation of audio embeddings into interpretable, concept-based representations LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae084b69-62e5-4e23-a599-39ce33734332 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Investigating the emergent audio classification ability of ASR foundation models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4eddcf9c-0342-480f-8c4a-9b393ed4011e · outbound
Transformation of audio embeddings into interpretable, concept-based representations Clotho: an audio captioning dataset,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f20a72a0-560c-4a22-be47-5d408a320910 · outbound
Transformation of audio embeddings into interpretable, concept-based representations AudioCLIP: Extending CLIP to image, text and audio,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e54742d7-0b0f-4f03-87e8-79f2eb51a9e2 · outbound
Transformation of audio embeddings into interpretable, concept-based representations OmniVec2 - a novel transformer based network for large scale multimodal and multitask learning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b5300b57-01e1-4f7d-8ea6-f297351a2298 · outbound
Transformation of audio embeddings into interpretable, concept-based representations Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc246ef7-dc97-4383-9a25-f9ad0bdfbd67 · inbound
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings Transformation of audio embeddings into interpretable, concept-based representations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.