Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2502.02366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02366 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:27:19.487461Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact8
  • verified fuzzy8
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f22b93c6-b266-49fe-8059-01aa1f00697d · outbound

This paper cites PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.256603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.256603Z digest=sha256:445386281fd5eb1bca41b7626fb7ba7d9adae96f3d8e46565538f7e549fbac79

Observation 42358cba-21a7-4453-9d26-22cddc5d4bff · outbound

This paper cites Embeddings for up to 2000 random samples from the validation partition of select datasets representing speech, non-speech and VAD audio domains.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Embeddings for up to 2000 random samples from the validation partition of select datasets representing speech, non-speech and VAD audio domains

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.159538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.250437Z digest=sha256:7734b81ef031de198c152c9468636b081ef17b901709a361bc06bdb1b7987034

Observation e57ecd9a-7554-4c9d-bcd5-aafea976dbb3 · outbound

This paper cites Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer’s Disease and Assess its Severity,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer’s Disease and Assess its Severity,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.271776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.271776Z digest=sha256:c2f8eb23ee6cedc9feb22c7f138b5c255a3e126c99a12e554ad10d6eaa067750

Observation dc4165f4-1aa9-41f7-bc33-719a3bbcde75 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.143137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.261659Z digest=sha256:41d3af93c016c4e6365b924631d389f88b0afc7c007459ea04ebf8c902650664

Observation d362b1a4-41e5-4904-8248-60c7a41cf417 · outbound

This paper cites Characterizing soundscapes across diverse ecosystems using a universal acoustic feature set,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Characterizing soundscapes across diverse ecosystems using a universal acoustic feature set,

Reference 5

Resolution
verified exact
doi, observed 2026-08-09T12:27:20.051458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.282694Z digest=sha256:921cab3b24d31a064f5ac59620ce8eb0aec2f498d29d0231dd5d3c02d3332dfd

Observation 26e6edd9-68e2-4b4d-b8e1-86c32ed79e01 · outbound

This paper cites Soundscapes and deep learning enable tracking biodiversity recovery in tropical forests,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Soundscapes and deep learning enable tracking biodiversity recovery in tropical forests,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.287723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.287723Z digest=sha256:3757b4b1582c1d6bc1dac881831fa5384a88d3fb97df797271bfb4e9dad55b52

Observation 0ed909f8-a55a-4391-8c4e-b41ddafeec3d · outbound

This paper cites Using X-Vectors to Automatically Detect Parkinson’s Disease from Speech,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Using X-Vectors to Automatically Detect Parkinson’s Disease from Speech,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.277150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.277150Z digest=sha256:a913d3e8b23a36dd910314b71e1e0f867226a994d12a6d79c9a22ac4095dc286

Observation 249558b5-4a6a-4e15-af5e-751f385b1726 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Representation Learning with Contrastive Predictive Coding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.297365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.297365Z digest=sha256:bff3c382320b9e56abbcefe3eea9d6eb99372c9811cb7ae3351656991fd1f2c0

Observation 53867186-9827-4bb7-a66c-f5990b2fb264 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning for Speech Recognition.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unsupervised Cross-lingual Representation Learning for Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.302520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.302520Z digest=sha256:9804e1e942fa40e0d8ec99ffa7ae4ecc6639a916480525cd949f95623fd8b65f

Observation 6293db8b-5677-47ad-92ae-2d5f5b9d0681 · outbound

This paper cites Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.292513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.292513Z digest=sha256:abd23926a151991ed4fb984b0655c770aaaf0b694513c40661544e4b1b36d8d2

Observation 7cf3e73f-78d9-450d-b054-e0a62a91c99a · outbound

This paper cites BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.312022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.312022Z digest=sha256:0b763122778b33be3f5954f971a2632e21a5340396bf43ff4448467576fb4e16

Observation fd418345-df05-42d9-9378-9c66edf65639 · outbound

This paper cites BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.316685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.316685Z digest=sha256:d97fc2f0deb0329f70cf2932d71c129d5c5f14cf231ca41ae2f14ea5a1f743b8

Observation ff44cd45-6874-4b15-ba76-f18e5df4f0e0 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 13

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.973418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.307198Z digest=sha256:bb81b12cdb6b86520d2037a704917fa56f6790438f4f5f946462e5949e262d12

Observation 1f9be9d3-9ac4-4bfe-8c11-f17534111143 · outbound

This paper cites The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.333240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.333240Z digest=sha256:e82795c63294919d924b8e708a8bc843312b0c699ff515a99a12b1a3a31b5332

Observation e7bbcfef-e100-4020-833b-67ce9bf0711c · outbound

This paper cites Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,

Reference 15

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.857230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.339153Z digest=sha256:0fdcce959fe133935e6ac58fa9b94405fb2608e47d803c81b7f5202d93f0e3c8

Observation ff13feae-5940-4199-94fc-d8c14ac8c2c5 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Common Voice: A Massively-Multilingual Speech Corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.108393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.322091Z digest=sha256:2a90636cd658394675191b3829f0ff39b1916bdfe0b674582f405e36e5508539

Observation 0b8535c4-248a-4dd8-b80e-e720f53b3633 · outbound

This paper cites Available: https://aclanthology.org/2020.lrec-1.520.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Available: https://aclanthology.org/2020.lrec-1.520

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.092504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.326958Z digest=sha256:00987d88a7f71495dfdf88ef5f0d7cb0332bf56a10a019e7441a3744cc4046cf

Observation 520a4489-9a03-4f98-a9f0-79822b649030 · outbound

This paper cites Recognition and understanding of meetings the AMI and AMIDA projects,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Recognition and understanding of meetings the AMI and AMIDA projects,

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T12:27:20.657531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.353246Z digest=sha256:cf71cd0afa10f882913c433385a46c8de3e8448800361778edad1be58b250a52

Observation 74dc6609-6f2d-4697-a419-2495e7c30850 · outbound

This paper cites Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.075801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.358039Z digest=sha256:b8b849354c37496a68d6da38a858ee93c3ae53ff327b9bf47a7562f0ff900449

Observation 314f1b96-be1f-498b-bb11-68e5e25b1745 · outbound

This paper cites VoxCeleb: A Large-Scale Speaker Identification Dataset,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study VoxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.343966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.343966Z digest=sha256:e5e9e187ba048764310ffc2fee22fd6a882ca75d69fc045b392ce7041b760d5d

Observation b2a82536-1475-4947-9d6a-5802f59c35bd · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Librispeech: An ASR corpus based on public domain audio books,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.348655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.348655Z digest=sha256:793fdc87a9af87a15dcddcb0213cf3576f1b00bb5679cc8a184c565783df7f67

Observation c0bbf53d-9141-497d-9450-cb929aabf208 · outbound

This paper cites General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.373247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.373247Z digest=sha256:b7ae451a1a7f2f251186598f651e59c9ec782dc4283bf2f2247b5f6bd6cdba8b

Observation 86254a16-ab1c-4f01-99ba-f180d0b9b32a · outbound

This paper cites Audio tagging with noisy labels and minimal supervision.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Audio tagging with noisy labels and minimal supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.765128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.378879Z digest=sha256:3c1720690c166cf11e2e062757bdc9608e7f62e5aff11ded0267e85208f37da5

Observation 36c1e223-ea52-4c08-8b1e-ae6e318c66bd · outbound

This paper cites The Multilingual TEDx Corpus for Speech Recognition and Translation.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Multilingual TEDx Corpus for Speech Recognition and Translation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.362872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.362872Z digest=sha256:d8df25e61a4409471078dea49e58d8e441759192b6010b2f6d5b97c6c4ec0af8

Observation 2d3bd654-f1e8-4874-a971-b446dc5715f2 · outbound

This paper cites SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.367959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.367959Z digest=sha256:5c28080bcfbb98bb4a78edf0c1e2d79ecc64e213254ace6dd0d85d288a1d15ea

Observation c54c481e-8447-4f99-aa94-9734d26c196a · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study MUSAN: A Music, Speech, and Noise Corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.393622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.393622Z digest=sha256:bdd63476a03ec06a6bfba66c7bbc4739737dca0847f290f3bee46fd09e49b249

Observation 14668544-0e06-48e2-a15d-8450110aadbb · outbound

This paper cites An open dataset for research on audio field recording archives: freefield1010.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study An open dataset for research on audio field recording archives: freefield1010

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.720639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.399285Z digest=sha256:c9a527235f47b6c9dec45e938d65fb5dc541272d4e98134350971bd63730f494

Observation 110d69d4-f231-4a14-aa5d-bfa7b7d19866 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study FSD50K: An Open Dataset of Human-Labeled Sound Events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.384159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.384159Z digest=sha256:62cc434cd1d9773ff08aa61761fc5911420ffbbc7a817894d1f98402e1c4d2b3

Observation 1b64d332-ca61-423a-9c36-6ce02c484be4 · outbound

This paper cites an unresolved cited work.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:27:21.176477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.244016Z digest=sha256:6994c51523281777762984d107d9ded6cea59c0745ea744f41cb6332cd85d889

Observation 9b592abc-a6af-4bd6-a354-5a2856c81cb5 · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Audio Set: An ontology and human-labeled dataset for audio events,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.388827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.388827Z digest=sha256:128d690c9402403b2d27aeee475435a4f3304160bf15ca4b0d8e0eeeb52c751b

Observation 20a62922-1eda-4a1f-ae5d-c67e70a5b59f · outbound

This paper cites CNN architectures for large-scale audio classification,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study CNN architectures for large-scale audio classification,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.418526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.418526Z digest=sha256:05378943528cf374b301a283d47778a33c2ab63e58c55cb4d522f8c4dcffebf8

Observation fb81c49d-4cbf-4b1e-9247-254b030e661e · outbound

This paper cites WHAM!: Extending Speech Separation to Noisy Environments.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study WHAM!: Extending Speech Separation to Noisy Environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.404632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.404632Z digest=sha256:518c797ca232c260c98fc3cd1fb0c8fe50a4335ea1f18390053e0ae4867727e9

Observation 8433630e-532e-4569-8d87-47c260851edb · outbound

This paper cites Bootstrap your own latent a new approach to self-supervised learning,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Bootstrap your own latent a new approach to self-supervised learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.057812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.409636Z digest=sha256:5d3e492069be84e815f2cc4a4a28275b9e6d4678435090ff7ac852ed1182d328

Observation ec4dd20b-a7e9-4b7e-95c4-6c7677caec9e · outbound

This paper cites MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.413855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.413855Z digest=sha256:042e38e92804d2077599b3b5546976da16cee312044acf173cb8346506522bf1

Observation 074a7df4-19ee-43f9-94eb-87bcdb907136 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.436477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.436477Z digest=sha256:570fda91b9ff78d2db3702e30331d9d3888d432e323e33ace76ded7a82632193

Observation acd3b149-c1c7-4a2a-ab38-61d016f7b473 · outbound

This paper cites Environmental sound classification with convolutional neural networks,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Environmental sound classification with convolutional neural networks,

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T12:27:20.321447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.422945Z digest=sha256:d20519f3d85752693b68b2c05fb4dbbc636d8d3b57e94f3700b10c700255d94b

Observation fe80b230-89ce-445f-9f8f-7be5761927af · outbound

This paper cites A Dataset and Taxonomy for Urban Sound Research,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study A Dataset and Taxonomy for Urban Sound Research,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.427052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.427052Z digest=sha256:48f8cc5e058b521fa287a1f69176a9d385dc73cf28d22b171a672baa317e34b2

Observation 0ca4c5ec-044d-4176-8293-d70414a0651d · outbound

This paper cites Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.040629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.431944Z digest=sha256:c34896ca116cbe749fcdddb59a060db672113a540340bbbf59e47950c83d3648

Observation 53af1c20-6702-4b24-af27-330b3785cc1b · outbound

This paper cites CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.441197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.441197Z digest=sha256:00bc9df7509d1880d90930a3905fab8cd4c90340d73ac48e13aebae7088e8ba3

Observation 4c5150ae-374a-4d78-a102-d8f1a92e5417 · outbound

This paper cites AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.631797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.446594Z digest=sha256:1fc363c3722ea9d1fb8d142b4936180d36023cb5a73b38116120d0d5ea27bede

Observation 956c6d5a-f35a-423e-912e-b1bfa9ea51ad · outbound

This paper cites Representational geometry: integrating cognition, computation, and the brain,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Representational geometry: integrating cognition, computation, and the brain,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.451806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.451806Z digest=sha256:8b71b8660ec5e8d6a5ec52d9287d09831f96227d4e8002cafbc164a3d275e7de

Observation 9a4ce4d0-467c-4015-a675-305a16683bf3 · outbound

This paper cites The Timbre Toolbox: Extracting audio descriptors from musical signals,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Timbre Toolbox: Extracting audio descriptors from musical signals,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.456940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.456940Z digest=sha256:b0198c28baf9a9591e997ce0696c0b4eff57aafe1653e45ccd25b68079666d78

Observation d43fedd3-434b-4923-a10c-709ec2a3479b · outbound

This paper cites The Modulation Transfer Function for Speech Intelligibility,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Modulation Transfer Function for Speech Intelligibility,

Reference 44

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.586295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.462136Z digest=sha256:0e773ad084e595f8f95485f9b862a6ad6073124b3c29fd86652d74ae882cb1a9

Observation 8992706d-8d5d-49cc-bf50-e2ddb1370209 · outbound

This paper cites YIN, a fundamental frequency estimator for speech and music,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study YIN, a fundamental frequency estimator for speech and music,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.467070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.467070Z digest=sha256:b2b6ccc9876dde0c614654e5d799f91f31032b4fa395d9537ba5cc44c25bd46f

Observation 6fdc6041-464b-43b8-8394-9cc7d6ab7f27 · outbound

This paper cites Acoustic Event Detection Using Speaker Recognition Techniques: Model Optimization and Explainable Features,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Acoustic Event Detection Using Speaker Recognition Techniques: Model Optimization and Explainable Features,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.472186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.472186Z digest=sha256:4ddb8e15500a2b4ee4bbd9926f671dc2f5aa3becd2c5ae06dcc70f540fb5a11b

Observation 4323c434-8026-43df-94e6-1aba819f9913 · outbound

This paper cites Acoustic Correlates of Auditory Object and Event Perception: Speakers, Musical Timbres, and Environmental Sounds,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Acoustic Correlates of Auditory Object and Event Perception: Speakers, Musical Timbres, and Environmental Sounds,

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-09T12:27:20.150848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.477019Z digest=sha256:2f879d6afb836fbc5156709ef1013fba76bcbff8037683c9e2fef208430d6a11

Observation aca9b87a-2b79-439c-849e-90e5273d74cb · outbound

This paper cites AVES: Animal Vocalization Encoder based on Self-Supervision.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study AVES: Animal Vocalization Encoder based on Self-Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.482197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.482197Z digest=sha256:0295f15f4ce270a5acd7025361402c839b18bd1619f5ec8554cfbfbed01b991b

Observation 1d76a4ec-b7cf-4bbe-82ce-6e655527bbe9 · outbound

This paper cites SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.487461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.487461Z digest=sha256:6d6073a626a38dee929b2a39d856137aa1faeac0aeda50906b43245e62c9640c

Observation a6dbf56d-3d40-4535-a098-7f095a33addc · outbound

This paper cites Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.126288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T12:27:19.266785Z digest=sha256:a4a5ff611a8ec993690043805489f5749fadd3b0eb1850a8b6ff03da0d3882f4

Pith citing papers

No inbound Pith citation observations are available.