Pith. sign in

Paper Citation Record · LEDGER

Transformation of audio embeddings into interpretable, concept-based representations

As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2504.14076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14076 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:16.670571Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T23:50:12.369547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69f52435-176f-4775-a41d-00be433644ec · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,.

Transformation of audio embeddings into interpretable, concept-based representations PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.293665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.494155Z digest=sha256:84a7bef88bdd49e708aefd170b1c4a7f1e41e112f5f02de71c05a7be4a44d05c

Observation 0cedf91b-16a7-4d26-9836-1573cc788219 · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

Transformation of audio embeddings into interpretable, concept-based representations HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.279122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.500926Z digest=sha256:ced373200bed49b678389f401d5bb5724c482233061b9bf977afaae59487e87e

Observation 91270124-cbe5-43d5-b098-362617de1c58 · outbound

This paper cites Pengi: An audio language model for audio tasks,.

Transformation of audio embeddings into interpretable, concept-based representations Pengi: An audio language model for audio tasks,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.264756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.505891Z digest=sha256:77115c218fbb2b1bff874d584ff065279da6058dd7834797da71e6ed8e5d904c

Observation ff7f4efe-477f-4132-9490-ca7aafdd242f · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Transformation of audio embeddings into interpretable, concept-based representations Learning transferable visual models from natural language supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.249086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.510528Z digest=sha256:58d53916c8be69925e46cd72545b912d72135abe508cb714b2c76761e761524e

Observation 550fb7bd-3af7-4156-b93c-7edadac2c65a · outbound

This paper cites CLAP: learning audio concepts from natural language supervision,.

Transformation of audio embeddings into interpretable, concept-based representations CLAP: learning audio concepts from natural language supervision,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.234572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.515116Z digest=sha256:28ef48e31bed482c70b43c2f8c4a1ca6b4728b1a52811d43d35a0fdb078203b8

Observation 84048c50-14ca-4591-ad75-e06e42feaa43 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Transformation of audio embeddings into interpretable, concept-based representations Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.219440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.519845Z digest=sha256:bd9aa91a8bf03bcab97e4fa0dae43bfe6419943d99fe705835f0844b4a45659c

Observation 59e8959c-14d4-4413-8a69-0c4075e94cfd · outbound

This paper cites Natural language supervision for general-purpose audio representations,.

Transformation of audio embeddings into interpretable, concept-based representations Natural language supervision for general-purpose audio representations,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.204407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.525028Z digest=sha256:1cb7c1b5c03c1b9ca0176bfd2b2b13c269039900f442f3e473fe02122ade71e0

Observation e4ba268b-eeee-4b15-92ac-25bf84a46d74 · outbound

This paper cites European union regulations on algorith- mic decision making and a “right to explanation.

Transformation of audio embeddings into interpretable, concept-based representations European union regulations on algorith- mic decision making and a “right to explanation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.189034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.529188Z digest=sha256:f6d3f8b203fe5c08e2a626c602d1a0e23df12059e45cd02d3a8d4f9d8427d2ec

Observation 5cbd2d83-aaae-4171-81a3-a02f51c450bd · outbound

This paper cites Interpreting CLIP’s image representation via text-based decomposition,.

Transformation of audio embeddings into interpretable, concept-based representations Interpreting CLIP’s image representation via text-based decomposition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.174597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.533714Z digest=sha256:b80109dd8b4131ac0b090ae3a830b6ec4dd41720beb6689cf356ff9cc988bcc8

Observation 2c40411c-8aff-412d-a7b7-5d0ac3f9da7d · outbound

This paper cites Interpreting CLIP with sparse linear concept embeddings (SpLICE),.

Transformation of audio embeddings into interpretable, concept-based representations Interpreting CLIP with sparse linear concept embeddings (SpLICE),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.160399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.537920Z digest=sha256:02031fb67447034a51d0250123abe59e2ce9d7cddaebf9b7345ed639803e5565

Observation e1a20157-e0fc-47a2-a996-3a6bc1c21024 · outbound

This paper cites Disentangling visual and written concepts in CLIP,.

Transformation of audio embeddings into interpretable, concept-based representations Disentangling visual and written concepts in CLIP,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.146374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.542255Z digest=sha256:c3e63eac83d04a6a5169da3c294f92cc20404f27d65c4fdcee8d84f14a3f792b

Observation 91b2dbbf-5d76-4c88-b4a6-7b13a01251e5 · outbound

This paper cites Toward interpretable music tagging with self-attention,.

Transformation of audio embeddings into interpretable, concept-based representations Toward interpretable music tagging with self-attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.131675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.546669Z digest=sha256:e2fea918c2d52bf74b59aa50ee8349754e7bf8e574520c512edac8238a38f161

Observation 226d05b9-2013-4288-ac8a-6d0e97a6676f · outbound

This paper cites Interpreting and explaining deep neural networks for classification of audio signals,.

Transformation of audio embeddings into interpretable, concept-based representations Interpreting and explaining deep neural networks for classification of audio signals,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.117133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.550788Z digest=sha256:fca6c11583d19bf0933491e4b7621c72cd9a215373151eff2ba7c79fd75ca18e

Observation 3c10b0f4-9cad-4e6e-9ef8-afc14b8ad4b1 · outbound

This paper cites Why are speech spectrograms hard to read?.

Transformation of audio embeddings into interpretable, concept-based representations Why are speech spectrograms hard to read?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.103234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.555096Z digest=sha256:0916d0f0df5fdde29bd9a80473272fa53f36314e4c2ad03fd25fd5e5eea3b6f6

Observation 430f39fb-4d59-4db8-9532-61908fefe85f · outbound

This paper cites SPES: Spectrogram perturbation for explainable speech-to-text generation,.

Transformation of audio embeddings into interpretable, concept-based representations SPES: Spectrogram perturbation for explainable speech-to-text generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.089536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.559439Z digest=sha256:b17df2001c20d9165a482d8d06d52f62c996320f8ba60240e102782224538a33

Observation ae3003f0-8f51-4b10-b8e1-a949cf023418 · outbound

This paper cites ”Why should I trust you?.

Transformation of audio embeddings into interpretable, concept-based representations ”Why should I trust you?

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.075134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.563563Z digest=sha256:ef845331b3e4841c0dd6724a91e73fc7bad4a0755dcaa73402692109fe0e669a

Observation eeefddc4-e246-42a9-aa3e-a24ba9493060 · outbound

This paper cites Local interpretable model- agnostic explanations for music content analysis,.

Transformation of audio embeddings into interpretable, concept-based representations Local interpretable model- agnostic explanations for music content analysis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.060990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.567552Z digest=sha256:93faf617a71622d48617eea7dd2d50473cd35612d24405a5b85ece1a59986451

Observation 35ed7973-de74-4cdc-9bf2-f07085a7a77b · outbound

This paper cites audioLIME: Listenable Explanations Using Source Separation,.

Transformation of audio embeddings into interpretable, concept-based representations audioLIME: Listenable Explanations Using Source Separation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.046797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.571959Z digest=sha256:916ba64ca540ae86566c05693edaf436dd3908ef8976e1bb290505022269ee2b

Observation 59209025-de0d-4587-9643-f0021f9bf11b · outbound

This paper cites Listen to interpret: Post-hoc interpretability for audio networks with NMF,.

Transformation of audio embeddings into interpretable, concept-based representations Listen to interpret: Post-hoc interpretability for audio networks with NMF,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.032201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.576151Z digest=sha256:57b8eda5c8dc6d2a1db588e83da3573ddbcbaa0e8de0ee4415adca58368cfda0

Observation 827e0a37-5354-4e5b-a981-cc9c000b7ec6 · outbound

This paper cites Towards automatic concept-based explanations,.

Transformation of audio embeddings into interpretable, concept-based representations Towards automatic concept-based explanations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.018001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.580387Z digest=sha256:e0f43819cca3716d7cd74c650a61836cc40548bf6b8f17f4b3e4d22ce3e6e778

Observation 4e36c3f0-e4d3-48e7-bf88-2dd38fbf217b · outbound

This paper cites Concept bottleneck models,.

Transformation of audio embeddings into interpretable, concept-based representations Concept bottleneck models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:17.003041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.584696Z digest=sha256:6e92817ea649e6c8118038e8ef4f12b9fa3d4f81c1ddf477b4b957c8e4ecd270

Observation e68192a7-ce59-442d-b1f7-43edcb5081d5 · outbound

This paper cites Label-free concept bottleneck models,.

Transformation of audio embeddings into interpretable, concept-based representations Label-free concept bottleneck models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.989307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.589131Z digest=sha256:5aa0e84d66e16e9674ecde671f5493cd44eb7f0197ca8a88b081b81ebd32939e

Observation f77d455e-5339-409f-8938-0830d8c3971c · outbound

This paper cites Post-hoc concept bottleneck models,.

Transformation of audio embeddings into interpretable, concept-based representations Post-hoc concept bottleneck models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.975330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.593106Z digest=sha256:2d27e6246bed29313819deb222b4bc39ec646ae2a370a99f32697a7a11addae0

Observation 9d5246df-b676-449f-b015-d3eec045d3fc · outbound

This paper cites Information maximization perspective of orthogonal matching pursuit with applications to explain- able AI,.

Transformation of audio embeddings into interpretable, concept-based representations Information maximization perspective of orthogonal matching pursuit with applications to explain- able AI,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.960873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.597713Z digest=sha256:815844074e103d96bdcf9694f240f7e648869666b609dc5d3d9857df0f781d40

Observation c2ce901f-a1da-4c0f-856b-7c844eb110d9 · outbound

This paper cites Sparse Linear Concept Discovery Models ,.

Transformation of audio embeddings into interpretable, concept-based representations Sparse Linear Concept Discovery Models ,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.946539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.602073Z digest=sha256:54eeeb60465d60c24953184dd5408b058d82b96da853b396a360f242b3c4e972

Observation 0ec2f441-fb6f-4ffe-ae29-be537c52cf8c · outbound

This paper cites CLIP-Dissect: Automatic description of neuron representations in deep vision networks,.

Transformation of audio embeddings into interpretable, concept-based representations CLIP-Dissect: Automatic description of neuron representations in deep vision networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.932010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.606172Z digest=sha256:35b9589ca3be991410610b9d23995005a5961babd0cceba878b6a3af7f7f640b

Observation a460c868-749e-4ece-9c68-88183b6bbf72 · outbound

This paper cites FSD50K: An open dataset of human-labeled sound events,.

Transformation of audio embeddings into interpretable, concept-based representations FSD50K: An open dataset of human-labeled sound events,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.918035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.610210Z digest=sha256:ab754a8adf4d891448008fcbbe793531be5e7c328c3e6b46281e7616dad27f77

Observation 2f5e55ee-4fbe-4574-8d38-2cfda4822559 · outbound

This paper cites Concept-based explanations using non- negative concept activation vectors and decision tree for CNN models,.

Transformation of audio embeddings into interpretable, concept-based representations Concept-based explanations using non- negative concept activation vectors and decision tree for CNN models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.903926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.614387Z digest=sha256:7426096731bfcd3c8641d0c233cdd355d03bc37b9dba9d1d401b9122576627ff

Observation 42805c1e-0f0c-47cb-9951-27485a14a83e · outbound

This paper cites Invertible concept-based explanations for CNN models with non- negative concept activation vectors,.

Transformation of audio embeddings into interpretable, concept-based representations Invertible concept-based explanations for CNN models with non- negative concept activation vectors,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.889262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.618368Z digest=sha256:bf95d6e732131641ecaf57fd6b5c9bc1cfb9809d48a2c98a4ab53a1fc92a36d5

Observation 39e9e94a-2cd4-4391-bc0b-a2cba171151e · outbound

This paper cites A dataset and taxonomy for urban sound research,.

Transformation of audio embeddings into interpretable, concept-based representations A dataset and taxonomy for urban sound research,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.874726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.622608Z digest=sha256:d154329b171a03711238973fb24987d6045e82262851af7df378e46e1f3cee3d

Observation 3ece9a44-34ee-460d-8aed-12d991c85e2e · outbound

This paper cites Sound event detection in the DCASE 2017 challenge,.

Transformation of audio embeddings into interpretable, concept-based representations Sound event detection in the DCASE 2017 challenge,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.859860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.626582Z digest=sha256:b71c31874ee440537562cf51804702e0fcc174489574915164a10398471764ab

Observation e6924f4d-62cf-40f7-a96c-e2e0661676b2 · outbound

This paper cites ESC: Dataset for Environmental Sound Classification,.

Transformation of audio embeddings into interpretable, concept-based representations ESC: Dataset for Environmental Sound Classification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.845826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.631439Z digest=sha256:2b676c0b05daec78ed3c8ae1bda9daae5fa37007f600d3b717965c1349ab9950

Observation 7c45c692-7593-454c-ba23-905b89f8cbce · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Transformation of audio embeddings into interpretable, concept-based representations Audio Set: An ontology and human-labeled dataset for audio events,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.831602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.636187Z digest=sha256:dceb05818a29092916772dd8ddc22b83b3ae126eef5a22415caa098b05d1c12a

Observation da1718ea-4e25-4902-b11c-fa922fb2482c · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition,.

Transformation of audio embeddings into interpretable, concept-based representations V ocalsound: A dataset for improving human vocal sounds recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.816927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.640300Z digest=sha256:680099187b903cc0322fa063a7141d857b01f4ad05acc971133fac5e566ab544

Observation 6bf75b29-0397-4e47-ad26-79d70c49b43d · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

Transformation of audio embeddings into interpretable, concept-based representations WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.801445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.644835Z digest=sha256:51ae8763893141ea7d0a8a04728bdbd870e4bc0d1076a7d3dd1e874441ace08b

Observation 92713099-43f7-49a9-8ea4-29ed301e4eea · outbound

This paper cites LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,.

Transformation of audio embeddings into interpretable, concept-based representations LanguageBind: Extending video-language pretraining to n-modality by language-based semantic alignment,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.784892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.649281Z digest=sha256:681b7eb1e6df1036703f547382000548c4e3d5a0b85d8b27c0479115cff037a8

Observation ae084b69-62e5-4e23-a599-39ce33734332 · outbound

This paper cites Investigating the emergent audio classification ability of ASR foundation models,.

Transformation of audio embeddings into interpretable, concept-based representations Investigating the emergent audio classification ability of ASR foundation models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.769299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.653524Z digest=sha256:4ec716678de3c1732300324566ebd4101999adc044f1f6b603b74740a2b1555f

Observation 4eddcf9c-0342-480f-8c4a-9b393ed4011e · outbound

This paper cites Clotho: an audio captioning dataset,.

Transformation of audio embeddings into interpretable, concept-based representations Clotho: an audio captioning dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.754656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.658053Z digest=sha256:720a28c26423b9d7a5cf1ce24c60cbf8c32640527fc304bc7316710304c04261

Observation f20a72a0-560c-4a22-be47-5d408a320910 · outbound

This paper cites AudioCLIP: Extending CLIP to image, text and audio,.

Transformation of audio embeddings into interpretable, concept-based representations AudioCLIP: Extending CLIP to image, text and audio,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.739768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.662104Z digest=sha256:18867f062fbbf08fe64094d5394372c8efb7542dee86f29d0bbef562a867703b

Observation e54742d7-0b0f-4f03-87e8-79f2eb51a9e2 · outbound

This paper cites OmniVec2 - a novel transformer based network for large scale multimodal and multitask learning,.

Transformation of audio embeddings into interpretable, concept-based representations OmniVec2 - a novel transformer based network for large scale multimodal and multitask learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:01:16.724370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:01:16.666251Z digest=sha256:6cd22025448b62b2977fad9110dd8b5df1d5f2c360ec7c36b1c9c05e3caa2a3d

Observation b5300b57-01e1-4f7d-8ea6-f297351a2298 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Transformation of audio embeddings into interpretable, concept-based representations Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:16.670571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:16.670571Z digest=sha256:02a395930c424f84ea2c3be0e3415956fe2c4e109aa525896a964989dbea35d6

Pith citing papers

Observation fc246ef7-dc97-4383-9a25-f9ad0bdfbd67 · inbound

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings cites this paper.

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings Transformation of audio embeddings into interpretable, concept-based representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T23:50:12.369547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T23:50:12.369547Z digest=sha256:fd26fc4bfa1e3f1c091dc31a18e717c306381f818e24f41ef1e6ed134a0a10db