Pith. sign in

Paper Citation Record · LEDGER

USAD: Universal Speech and Audio Representation via Distillation

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2506.18843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18843 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:26.977695Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:27:05.187538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:41:06.244426Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9803bd2d-a80e-4467-8fd1-1bf7fce1c619 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

USAD: Universal Speech and Audio Representation via Distillation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:20.950820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:20.950820Z digest=sha256:12bc09e36445ad7d8e990b9d01b7750281fefd9fbc83d169e93ff226c0eafa56

Observation 2857be56-1e71-4949-81eb-aaefadcbb708 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

USAD: Universal Speech and Audio Representation via Distillation Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:35.149503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.019181Z digest=sha256:f9db921ae7f2325076b71297a78d0aceb973ff180b7c1adbd178c8772061ad15

Observation 9580200e-b490-45ad-a885-d27f022ad6f8 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

USAD: Universal Speech and Audio Representation via Distillation Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.969084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.120019Z digest=sha256:1dc58fd3279276b18632d386527d3637f13c70179a53fe0f9be7ec9fd3dea66d

Observation bd83d77f-d534-4744-998f-41adca1bfe1c · outbound

This paper cites Ssast: Self-supervised audio spectrogram transformer,.

USAD: Universal Speech and Audio Representation via Distillation Ssast: Self-supervised audio spectrogram transformer,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.783760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.326189Z digest=sha256:f36d3de723f1426e7c6d6dfe2a401950f269b2bd9518e0d650ce85e735b46484

Observation 96f93f5c-1c38-446e-9399-342e15204231 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

USAD: Universal Speech and Audio Representation via Distillation Beats: Audio pre-training with acoustic tokenizers,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.540726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.493407Z digest=sha256:434ecd214e1e6bee646a9a6748f3b3cf15b9dfd46111a884391d195d747b5bfd

Observation 4aa95dd6-9ff3-465d-bed9-f5714ff419ff · outbound

This paper cites Mert: Acoustic music understanding model with large-scale self-supervised training,.

USAD: Universal Speech and Audio Representation via Distillation Mert: Acoustic music understanding model with large-scale self-supervised training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.366351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.664948Z digest=sha256:69d799511b05e4921005b069bd3948c55291604159de50b795f920fdf7d935e9

Observation 0ac23d43-ed8b-4c66-b5d7-b980a7db7d3a · outbound

This paper cites Listen, think, and understand,.

USAD: Universal Speech and Audio Representation via Distillation Listen, think, and understand,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:34.178286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.782308Z digest=sha256:9a22ff62a2d606dbd71b7e2891a961c8dc90393a2962ecc7b2aa439af7b68111

Observation 2797e9a3-3d69-4259-a68d-31139cf72acc · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

USAD: Universal Speech and Audio Representation via Distillation SALMONN: Towards generic hearing abilities for large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.985657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.848096Z digest=sha256:d0ea1c74e0a3f8fad6701e34e20e2a459ba680d7769c0f76f82fbcde2dacc2c5

Observation adf77828-8d75-4f97-ab06-0a5b90908caa · outbound

This paper cites Qwen2-audio technical report,.

USAD: Universal Speech and Audio Representation via Distillation Qwen2-audio technical report,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.868863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:21.989905Z digest=sha256:4ba9a078c36d25a92ae35e61f8f917861c38528030e6e032a329c0c56b568b5a

Observation 830cc216-8902-46d8-85d0-feafef981276 · outbound

This paper cites Gama: A large audio- language model with advanced audio understanding and complex rea- soning abilities,.

USAD: Universal Speech and Audio Representation via Distillation Gama: A large audio- language model with advanced audio understanding and complex rea- soning abilities,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.707482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.121921Z digest=sha256:9165178020c2698fb305e67982c485aa2f512e4c02d6dfea17357247b0565d60

Observation ed7d1c45-9811-4b44-876c-b7527602b483 · outbound

This paper cites Google usm: Scaling automatic speech recognition beyond 100 languages,.

USAD: Universal Speech and Audio Representation via Distillation Google usm: Scaling automatic speech recognition beyond 100 languages,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.583245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.232038Z digest=sha256:7f2a9c73b14edce3f0195eef7248366ea11e44d7a822240a18710ab68f505276

Observation 9b99dd78-a658-4556-8360-1871eac56571 · outbound

This paper cites Speechtokenizer: Unified speech tokenizer for speech language models,.

USAD: Universal Speech and Audio Representation via Distillation Speechtokenizer: Unified speech tokenizer for speech language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.479959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.385294Z digest=sha256:adf34319f7226c510e4d2ddddf88cc2c9df6d4144c65e73064a626ea513c0224

Observation 1aeebdb9-51bb-4a6d-9d2e-f833a35391ac · outbound

This paper cites Soundstorm: Efficient parallel audio generation,.

USAD: Universal Speech and Audio Representation via Distillation Soundstorm: Efficient parallel audio generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.361483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.542181Z digest=sha256:d38786e2b8ebde109fe2160e9997d3964cac6a05e230a9b92784a6eddca55b46

Observation f1cb00c9-0acf-4673-8a29-054a36ec84dd · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

USAD: Universal Speech and Audio Representation via Distillation Moshi: a speech-text foundation model for real-time dialogue,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.210165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.709397Z digest=sha256:04b492dee1ed96db484921cdaee58f1cbf6a5f63989fa052fdb73403d871e89c

Observation d9e96910-e0ac-4769-896e-4ecfa949cbbd · outbound

This paper cites Dc-spin: A speaker-invariant speech tokenizer for spoken language models,.

USAD: Universal Speech and Audio Representation via Distillation Dc-spin: A speaker-invariant speech tokenizer for spoken language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:33.081468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.906742Z digest=sha256:f2deedc7e0877c701a89501b8c7ecd29baf981ca80e517e32c19bea69db5f93f

Observation a0c42ab5-3310-4048-9006-6d2db6be8ec1 · outbound

This paper cites Joint audio and speech understanding,.

USAD: Universal Speech and Audio Representation via Distillation Joint audio and speech understanding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.962859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:22.997389Z digest=sha256:519566f4d661c45947658e3532939710f26f4f63cd341ae4b99872b71bb72c3a

Observation af58277e-7e2c-4572-be75-91a401135ca4 · outbound

This paper cites U-sam: An audio language model for unified speech, audio, and music understanding,.

USAD: Universal Speech and Audio Representation via Distillation U-sam: An audio language model for unified speech, audio, and music understanding,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.826324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.094069Z digest=sha256:d3097c32e4c4df312e170ebae4951df93f164f01213cb0344c06a38375fb0428

Observation c3385d0f-2302-49f3-b9cb-942180b51746 · outbound

This paper cites CoLLD: Contrastive layer-to-layer distillation for compressing multi- lingual pre-trained speech encoders,.

USAD: Universal Speech and Audio Representation via Distillation CoLLD: Contrastive layer-to-layer distillation for compressing multi- lingual pre-trained speech encoders,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.702387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.166044Z digest=sha256:bf3774623cc31dac433451891674b274bee0781bbc2cf169c671820fa11a79b9

Observation b2340a75-f99a-4c4b-9cc9-130ac425ffb7 · outbound

This paper cites Mae-ast: Masked autoencoding audio spectrogram transformer,.

USAD: Universal Speech and Audio Representation via Distillation Mae-ast: Masked autoencoding audio spectrogram transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.622842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.277080Z digest=sha256:b668c9bf6973681fdbe2120773ddbb75a898700cd9a31720e740677489693ec1

Observation 21f37735-0ccd-47d3-82be-a112d6de1950 · outbound

This paper cites Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,.

USAD: Universal Speech and Audio Representation via Distillation Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.506691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.390695Z digest=sha256:297bc35629503ff3eb5b950d20ff2f57cab4091659de7c1ab9b34a34b30dd789

Observation f0bd5528-3a08-4a0c-8b9b-690bcee749fe · outbound

This paper cites Masked autoencoders that listen,.

USAD: Universal Speech and Audio Representation via Distillation Masked autoencoders that listen,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:23.535235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:23.535235Z digest=sha256:3380a1c59e77c4d217ea296fa64ac60227492ac8de09f97a473fde1af9489a00

Observation b12c6f02-375e-4ffc-9089-6363a8b22cf3 · outbound

This paper cites data2vec: A general framework for self-supervised learning in speech, vision and language,.

USAD: Universal Speech and Audio Representation via Distillation data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.295570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.677938Z digest=sha256:f964efd0ab82bdb4591c2e2774f9f03c639853c8b0f77aa01a6c2c6aaf379a93

Observation e8a8dcea-19d2-47e1-980f-e119fa9246b5 · outbound

This paper cites Efficient self-supervised learning with contextualized target representations for vision, speech and language,.

USAD: Universal Speech and Audio Representation via Distillation Efficient self-supervised learning with contextualized target representations for vision, speech and language,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.234671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.826605Z digest=sha256:b049d5b57f171c62b1fc410c881f4b1ae5c14d45440ebe5162ce8d1876405635

Observation cb9532be-b1a9-42b2-bd44-e3c451245d5b · outbound

This paper cites Dinosr: Self-distillation and online clustering for self-supervised speech repre- sentation learning,.

USAD: Universal Speech and Audio Representation via Distillation Dinosr: Self-distillation and online clustering for self-supervised speech repre- sentation learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.101137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:23.957169Z digest=sha256:72aa5350ff5ce2c06eb7c032de155c8feacf9ed712bdb3e6d1846920dbd3a45c

Observation 2681e63e-d579-4757-b997-a9e21f2a0df1 · outbound

This paper cites Eat: Self-supervised pre-training with efficient audio transformer,.

USAD: Universal Speech and Audio Representation via Distillation Eat: Self-supervised pre-training with efficient audio transformer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:32.015670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.031734Z digest=sha256:3f90c5dc21a1404dcbf10fdbda8514ba42b0fe7c5d1dbb004a609f89d4a93342

Observation 59f5c611-2488-44ba-9586-f5601e530526 · outbound

This paper cites Sslam: Enhancing self-supervised models with audio mixtures for polyphonic soundscapes,.

USAD: Universal Speech and Audio Representation via Distillation Sslam: Enhancing self-supervised models with audio mixtures for polyphonic soundscapes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.851676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.185064Z digest=sha256:e25d40b95d344379c28c808a1808aa3f18cf84b5fea9af563810e0c1adabf7d1

Observation 5afdf0c6-b789-492e-9204-ea3c020b92c7 · outbound

This paper cites Byol for audio: Self-supervised learning for general-purpose audio representation,.

USAD: Universal Speech and Audio Representation via Distillation Byol for audio: Self-supervised learning for general-purpose audio representation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.704745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.281964Z digest=sha256:3cd148a98dee7030eef6256c4a254c9b4f9de823dc326c153a69472064c80bb3

Observation 0f8249ce-4a93-4bc8-a80e-58a6b4792f51 · outbound

This paper cites Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,.

USAD: Universal Speech and Audio Representation via Distillation Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.535281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.398795Z digest=sha256:1f168b62d4264ee0344f3524739e2d7b517fc2e3727e2dc7fd1df67c1fb6d569

Observation 4c8d1068-292d-4fbe-b655-ae6e422b240e · outbound

This paper cites Masked modeling duo: Learning representations by encouraging both networks to model the input,.

USAD: Universal Speech and Audio Representation via Distillation Masked modeling duo: Learning representations by encouraging both networks to model the input,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.423539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.493746Z digest=sha256:97869e679942652707e03b98b3a5280379805c98688f52d0a11f5e6f9d66fc3b

Observation 349d3f1f-58f9-4cf5-b2b9-7b3a0324ee33 · outbound

This paper cites DistilHuBERT: Speech rep- resentation learning by layer-wise distillation of hidden-unit bert,.

USAD: Universal Speech and Audio Representation via Distillation DistilHuBERT: Speech rep- resentation learning by layer-wise distillation of hidden-unit bert,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.310082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.560928Z digest=sha256:d7f78bf28ad52450f6b263ec9792a4c9b11c1dcb73aa3fafc747a13f29b59ff2

Observation 13ca0c5e-f154-42ec-af26-02292cf4c903 · outbound

This paper cites Dphubert: Joint dis- tillation and pruning of self-supervised speech models,.

USAD: Universal Speech and Audio Representation via Distillation Dphubert: Joint dis- tillation and pruning of self-supervised speech models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.214849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.625963Z digest=sha256:90b2ede541bc3509ca8b3f41a43b76cf949629a76c0f0c42e2f6265423fc2020

Observation 8ce43848-a607-4024-8e75-8d684c6286a3 · outbound

This paper cites Dass: Distilled audio state space models are stronger and more duration- scalable learners,.

USAD: Universal Speech and Audio Representation via Distillation Dass: Distilled audio state space models are stronger and more duration- scalable learners,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:24.692213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:24.692213Z digest=sha256:db131ccc1259817890618574ab88d1dc51a541e34057f168cfb71587aee3b85c

Observation 03653dc8-1f29-4cfd-b528-f408fdcac4a2 · outbound

This paper cites Ensemble knowledge distillation of self- supervised speech models,.

USAD: Universal Speech and Audio Representation via Distillation Ensemble knowledge distillation of self- supervised speech models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:31.064742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.808046Z digest=sha256:4fe38436253807bd1c2815080da4936c6b9be2d79ad85c37afa1ff8154aa58aa

Observation 71239d75-3b5b-4d83-bd53-4739d01a2d80 · outbound

This paper cites Distilling a speech and music encoder with task arithmetic,.

USAD: Universal Speech and Audio Representation via Distillation Distilling a speech and music encoder with task arithmetic,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.875657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.906549Z digest=sha256:dd87c3b3acb7d524554028ce756e7d641c159675af08fc15f0701a7de4c14f96

Observation d0d6e2be-b8f1-4a92-a866-ce4130a97cdd · outbound

This paper cites Attention is all you need,.

USAD: Universal Speech and Audio Representation via Distillation Attention is all you need,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.686734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:24.997035Z digest=sha256:1588809d27b5de70a98b024fc4f7ccbf00aff74ed39516ac46d047fa8c5ede0c

Observation 72a0f1f2-8997-4bd1-9b7f-1fa8c5e77264 · outbound

This paper cites Layer-wise analysis of a self- supervised speech representation model,.

USAD: Universal Speech and Audio Representation via Distillation Layer-wise analysis of a self- supervised speech representation model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.422581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.069275Z digest=sha256:db1593dd7a0c778f9851dc450a7691cade3229cf8c146fd929a7d17f34c5cb1d

Observation 29cf1ab6-5aec-43ce-ae48-0e3f82afb655 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

USAD: Universal Speech and Audio Representation via Distillation Robust speech recognition via large-scale weak super- vision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:25.144389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:25.144389Z digest=sha256:d3a3c3be403174db409c73a1c63633a2d97fa4bbe4b03fbb906b7c00b7576c8e

Observation aa963bdb-c1a6-4c9b-b94b-0cc331f22410 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

USAD: Universal Speech and Audio Representation via Distillation Librispeech: An ASR corpus based on public domain audio books,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.306828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.242305Z digest=sha256:7956eba7d67a047282d034e0e5363bd2eddd69deef1321597a2b5fe5a8febd1c

Observation 5b88cdb2-f84d-4909-a1e6-5f15722888ad · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

USAD: Universal Speech and Audio Representation via Distillation Libri-light: A benchmark for asr with limited or no supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:30.157667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.364102Z digest=sha256:c7ea41ca54cef771c2ff52bdcf1e147543da5a2476dd838e0b3eb1383f817484

Observation c24b73c5-4638-4bc6-8cc4-095b73dbadd2 · outbound

This paper cites Mls: A large-scale multilingual dataset for speech research,.

USAD: Universal Speech and Audio Representation via Distillation Mls: A large-scale multilingual dataset for speech research,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.984380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.477593Z digest=sha256:1a4927d13a6604ac5a3a9b71a5c2c939130303aa9001d282d7490761a77a3005

Observation 70e7d4af-fe2e-449f-ab09-05f17da6fb30 · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

USAD: Universal Speech and Audio Representation via Distillation V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.730479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.592409Z digest=sha256:e56f0e0ae3f56d4def9570f0291319811080a70f61a8cc00f9f000e03164c706

Observation d3bd4328-5d1c-4bef-9692-8674e551f07c · outbound

This paper cites Gigaspeech: An evolving, multi- domain asr corpus with 10,000 hours of transcribed audio,.

USAD: Universal Speech and Audio Representation via Distillation Gigaspeech: An evolving, multi- domain asr corpus with 10,000 hours of transcribed audio,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.469585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.710394Z digest=sha256:7e6f6891d498940101d08e4ebcf3ad4622bf2d7c0dd4819add421833f09dbe9b

Observation 7a3a4c6c-d745-42b7-8974-4864d027eae2 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

USAD: Universal Speech and Audio Representation via Distillation Common voice: A massively-multilingual speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.309280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.805425Z digest=sha256:97a4617dabb2b5855122161f743fbee2f2c9b24f5660dac8fd840c86e16e308a

Observation 3c27e679-e562-4d1b-a244-f98e5a118880 · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

USAD: Universal Speech and Audio Representation via Distillation The fisher corpus: A resource for the next generations of speech-to-text

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:29.137544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:25.937214Z digest=sha256:8fae27ff6e0a35a3d762d8eeeca7addbce32b92faf09ced9d71a9dbe47963672

Observation 287d4539-4ada-4f48-945e-e56defe31875 · outbound

This paper cites V oxlingua107: a dataset for spoken language recognition,.

USAD: Universal Speech and Audio Representation via Distillation V oxlingua107: a dataset for spoken language recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.955667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.051603Z digest=sha256:36cf03a7c526ece07d8cd0842d8fdbead6b34230d98cb39b806d7f2caeb6de42

Observation ec99ff25-39c1-479f-af6c-f5297ae458f4 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

USAD: Universal Speech and Audio Representation via Distillation Audio set: An ontology and human-labeled dataset for audio events,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:26.141172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:26.141172Z digest=sha256:a65e44a1133a1539aedee6d143c311576c459d81a914f6fb66c76a6e5d20f21e

Observation 5bfdbea0-dfaa-491d-b550-6c3049143255 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video,.

USAD: Universal Speech and Audio Representation via Distillation Soundnet: Learning sound representations from unlabeled video,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.794265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.187046Z digest=sha256:00490aafcf06d31de78999dc111ac874ef346fdddf25ec3c1be316cd4159d5a4

Observation cde9ede7-e5e0-4ccd-87d9-c4566b585fc0 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

USAD: Universal Speech and Audio Representation via Distillation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:26.234805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:26.234805Z digest=sha256:a536363418fa20cfdd0510c9d6bea6f94bce60af7713e05d051c5afc5270aaac

Observation 10effc48-0b51-4139-aa79-29293a9bee66 · outbound

This paper cites Music4all: A new music database and its applications,.

USAD: Universal Speech and Audio Representation via Distillation Music4all: A new music database and its applications,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.652831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.272726Z digest=sha256:52ab79ec91f1902d037980057451f75402091f2b1991cf4259843a95535e817e

Observation 38c9d75c-da90-4391-b0db-e86f3fa19f13 · outbound

This paper cites fairseq: A fast, extensible toolkit for sequence modeling,.

USAD: Universal Speech and Audio Representation via Distillation fairseq: A fast, extensible toolkit for sequence modeling,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.543437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.352239Z digest=sha256:90610639cb9c6ca91e414c46a63288b6767250e29cb3165dd1adfc9aa34fd4c1

Observation f5a81b72-a1be-4027-a593-5ac2b613813c · outbound

This paper cites Layer normalization,.

USAD: Universal Speech and Audio Representation via Distillation Layer normalization,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.411742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.409954Z digest=sha256:83479ad3678cd0ef022d6e5d6ebfa6a6ee80ed68cd23b78d2c7f2fe4a21f9a49

Observation b6dc93b6-4b9d-40c9-afe1-92e3732586c1 · outbound

This paper cites Self-attention with relative position representations,.

USAD: Universal Speech and Audio Representation via Distillation Self-attention with relative position representations,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.264759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.471051Z digest=sha256:d7e6008fee71f3c7bdcc00bc71ae18e49c3620c7b63011bce1b505d8a9d6be27

Observation 8f1808d9-5252-4f93-bbdd-6ed5c959396b · outbound

This paper cites Speech commands: A dataset for limited-vocabulary speech recognition,.

USAD: Universal Speech and Audio Representation via Distillation Speech commands: A dataset for limited-vocabulary speech recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:28.140213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.529521Z digest=sha256:1e929da54d89f2f075a85ccb088515b38f14d7904797dccb01d695acabef9c17

Observation b16ec7cc-281d-45e9-b90e-a5bac8a52c51 · outbound

This paper cites SUPERB: Speech processing universal performance benchmark,.

USAD: Universal Speech and Audio Representation via Distillation SUPERB: Speech processing universal performance benchmark,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.959384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.600823Z digest=sha256:b58cd0cd962bf3e3083dbe96504257bde61e42e00c75e642cc1dfa01a96f0ad7

Observation 4926ca56-4405-4134-acab-f5b787454b07 · outbound

This paper cites SUPERB-SG: Enhanced speech processing universal PERformance benchmark for semantic and generative capabilities,.

USAD: Universal Speech and Audio Representation via Distillation SUPERB-SG: Enhanced speech processing universal PERformance benchmark for semantic and generative capabilities,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.723306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.658300Z digest=sha256:cf48361744a100efb58b47b7671a715690c0074a7136cc14dbdfa8f0c98acf9d

Observation eb20a4d8-dfbf-478c-834c-61c39d3e8eea · outbound

This paper cites A large-scale evaluation of speech foundation models,.

USAD: Universal Speech and Audio Representation via Distillation A large-scale evaluation of speech foundation models,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.603002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.722253Z digest=sha256:fc9ec2d0ee78aaeae77a59fd900962e447cf4b5b8d02db8992931b2f576d5660

Observation 3ef690ab-d0ea-4524-b326-34161c0a3c03 · outbound

This paper cites Hear: Holistic evaluation of audio representations,.

USAD: Universal Speech and Audio Representation via Distillation Hear: Holistic evaluation of audio representations,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.479424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.790110Z digest=sha256:777b28a238504b6dde67c79deec42ca7e10a5f0f32c0c6b141d3806ef5927d25

Observation 20dcce66-1464-49a9-a821-5d5ab6a1bc5f · outbound

This paper cites ESC: Dataset for Environmental Sound Classification,.

USAD: Universal Speech and Audio Representation via Distillation ESC: Dataset for Environmental Sound Classification,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.325179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.882696Z digest=sha256:16a54ae3a3519edfe6def5b5d6a78668e0c237c653f75baa6e9225f09f7b78b9

Observation 1c975ef0-3fe5-43a6-bd22-1e4d3b6f7699 · outbound

This paper cites Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,.

USAD: Universal Speech and Audio Representation via Distillation Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:27.152485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:20:26.977695Z digest=sha256:5640f8103b2bf7e439abaf32b368d1df14b11a41f84169b4535349a998fc1390

Pith citing papers

Observation 509a8141-b875-4dfd-a562-9bf5547bf3f3 · inbound

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations cites this paper.

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations USAD: Universal Speech and Audio Representation via Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T07:27:05.187538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:27:05.187538Z digest=sha256:d63d0c87deb1c6429fd90636297194259952e78f1dfb703cc8b1c9ba27e2607c

Observation b3c39e7d-ebbd-417a-bfa5-7ab57dcca0a4 · inbound

Alethia: A Foundational Encoder for Voice Deepfakes cites this paper.

Alethia: A Foundational Encoder for Voice Deepfakes USAD: Universal Speech and Audio Representation via Distillation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:33.721376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T19:27:59.124425Z digest=sha256:20c3d03be3db4d158e6dcce66811cf323b2bc707fd57e6e7fb4bd78d47775964

Observation adc33ca8-be64-4795-a30f-20f8f644cdb5 · inbound

Stage-adaptive audio diffusion modeling cites this paper.

Stage-adaptive audio diffusion modeling USAD: Universal Speech and Audio Representation via Distillation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:06.248596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T17:21:34.140699Z digest=sha256:57b6cee63cb596c7ee02010b17fd52dedac4d0b570bb1deddf7a2a45d992691e