Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:26.977695Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2506.18843.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:26.977695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:27:05.187538Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T17:41:06.244426Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9803bd2d-a80e-4467-8fd1-1bf7fce1c619 · outbound
USAD: Universal Speech and Audio Representation via Distillation wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2857be56-1e71-4949-81eb-aaefadcbb708 · outbound
USAD: Universal Speech and Audio Representation via Distillation Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9580200e-b490-45ad-a885-d27f022ad6f8 · outbound
USAD: Universal Speech and Audio Representation via Distillation Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd83d77f-d534-4744-998f-41adca1bfe1c · outbound
USAD: Universal Speech and Audio Representation via Distillation Ssast: Self-supervised audio spectrogram transformer,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 96f93f5c-1c38-446e-9399-342e15204231 · outbound
USAD: Universal Speech and Audio Representation via Distillation Beats: Audio pre-training with acoustic tokenizers,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4aa95dd6-9ff3-465d-bed9-f5714ff419ff · outbound
USAD: Universal Speech and Audio Representation via Distillation Mert: Acoustic music understanding model with large-scale self-supervised training,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0ac23d43-ed8b-4c66-b5d7-b980a7db7d3a · outbound
USAD: Universal Speech and Audio Representation via Distillation Listen, think, and understand,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2797e9a3-3d69-4259-a68d-31139cf72acc · outbound
USAD: Universal Speech and Audio Representation via Distillation SALMONN: Towards generic hearing abilities for large language models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation adf77828-8d75-4f97-ab06-0a5b90908caa · outbound
USAD: Universal Speech and Audio Representation via Distillation Qwen2-audio technical report,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 830cc216-8902-46d8-85d0-feafef981276 · outbound
USAD: Universal Speech and Audio Representation via Distillation Gama: A large audio- language model with advanced audio understanding and complex rea- soning abilities,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed7d1c45-9811-4b44-876c-b7527602b483 · outbound
USAD: Universal Speech and Audio Representation via Distillation Google usm: Scaling automatic speech recognition beyond 100 languages,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9b99dd78-a658-4556-8360-1871eac56571 · outbound
USAD: Universal Speech and Audio Representation via Distillation Speechtokenizer: Unified speech tokenizer for speech language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1aeebdb9-51bb-4a6d-9d2e-f833a35391ac · outbound
USAD: Universal Speech and Audio Representation via Distillation Soundstorm: Efficient parallel audio generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1cb00c9-0acf-4673-8a29-054a36ec84dd · outbound
USAD: Universal Speech and Audio Representation via Distillation Moshi: a speech-text foundation model for real-time dialogue,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9e96910-e0ac-4769-896e-4ecfa949cbbd · outbound
USAD: Universal Speech and Audio Representation via Distillation Dc-spin: A speaker-invariant speech tokenizer for spoken language models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a0c42ab5-3310-4048-9006-6d2db6be8ec1 · outbound
USAD: Universal Speech and Audio Representation via Distillation Joint audio and speech understanding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af58277e-7e2c-4572-be75-91a401135ca4 · outbound
USAD: Universal Speech and Audio Representation via Distillation U-sam: An audio language model for unified speech, audio, and music understanding,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3385d0f-2302-49f3-b9cb-942180b51746 · outbound
USAD: Universal Speech and Audio Representation via Distillation CoLLD: Contrastive layer-to-layer distillation for compressing multi- lingual pre-trained speech encoders,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b2340a75-f99a-4c4b-9cc9-130ac425ffb7 · outbound
USAD: Universal Speech and Audio Representation via Distillation Mae-ast: Masked autoencoding audio spectrogram transformer,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 21f37735-0ccd-47d3-82be-a112d6de1950 · outbound
USAD: Universal Speech and Audio Representation via Distillation Masked spectrogram modeling using masked autoencoders for learning general-purpose audio representation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0bd5528-3a08-4a0c-8b9b-690bcee749fe · outbound
USAD: Universal Speech and Audio Representation via Distillation Masked autoencoders that listen,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12c6f02-375e-4ffc-9089-6363a8b22cf3 · outbound
USAD: Universal Speech and Audio Representation via Distillation data2vec: A general framework for self-supervised learning in speech, vision and language,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8a8dcea-19d2-47e1-980f-e119fa9246b5 · outbound
USAD: Universal Speech and Audio Representation via Distillation Efficient self-supervised learning with contextualized target representations for vision, speech and language,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb9532be-b1a9-42b2-bd44-e3c451245d5b · outbound
USAD: Universal Speech and Audio Representation via Distillation Dinosr: Self-distillation and online clustering for self-supervised speech repre- sentation learning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2681e63e-d579-4757-b997-a9e21f2a0df1 · outbound
USAD: Universal Speech and Audio Representation via Distillation Eat: Self-supervised pre-training with efficient audio transformer,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 59f5c611-2488-44ba-9586-f5601e530526 · outbound
USAD: Universal Speech and Audio Representation via Distillation Sslam: Enhancing self-supervised models with audio mixtures for polyphonic soundscapes,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5afdf0c6-b789-492e-9204-ea3c020b92c7 · outbound
USAD: Universal Speech and Audio Representation via Distillation Byol for audio: Self-supervised learning for general-purpose audio representation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f8249ce-4a93-4bc8-a80e-58a6b4792f51 · outbound
USAD: Universal Speech and Audio Representation via Distillation Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c8d1068-292d-4fbe-b655-ae6e422b240e · outbound
USAD: Universal Speech and Audio Representation via Distillation Masked modeling duo: Learning representations by encouraging both networks to model the input,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 349d3f1f-58f9-4cf5-b2b9-7b3a0324ee33 · outbound
USAD: Universal Speech and Audio Representation via Distillation DistilHuBERT: Speech rep- resentation learning by layer-wise distillation of hidden-unit bert,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13ca0c5e-f154-42ec-af26-02292cf4c903 · outbound
USAD: Universal Speech and Audio Representation via Distillation Dphubert: Joint dis- tillation and pruning of self-supervised speech models,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ce43848-a607-4024-8e75-8d684c6286a3 · outbound
USAD: Universal Speech and Audio Representation via Distillation Dass: Distilled audio state space models are stronger and more duration- scalable learners,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03653dc8-1f29-4cfd-b528-f408fdcac4a2 · outbound
USAD: Universal Speech and Audio Representation via Distillation Ensemble knowledge distillation of self- supervised speech models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 71239d75-3b5b-4d83-bd53-4739d01a2d80 · outbound
USAD: Universal Speech and Audio Representation via Distillation Distilling a speech and music encoder with task arithmetic,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d0d6e2be-b8f1-4a92-a866-ce4130a97cdd · outbound
USAD: Universal Speech and Audio Representation via Distillation Attention is all you need,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 72a0f1f2-8997-4bd1-9b7f-1fa8c5e77264 · outbound
USAD: Universal Speech and Audio Representation via Distillation Layer-wise analysis of a self- supervised speech representation model,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 29cf1ab6-5aec-43ce-ae48-0e3f82afb655 · outbound
USAD: Universal Speech and Audio Representation via Distillation Robust speech recognition via large-scale weak super- vision,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa963bdb-c1a6-4c9b-b94b-0cc331f22410 · outbound
USAD: Universal Speech and Audio Representation via Distillation Librispeech: An ASR corpus based on public domain audio books,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b88cdb2-f84d-4909-a1e6-5f15722888ad · outbound
USAD: Universal Speech and Audio Representation via Distillation Libri-light: A benchmark for asr with limited or no supervision,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c24b73c5-4638-4bc6-8cc4-095b73dbadd2 · outbound
USAD: Universal Speech and Audio Representation via Distillation Mls: A large-scale multilingual dataset for speech research,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 70e7d4af-fe2e-449f-ab09-05f17da6fb30 · outbound
USAD: Universal Speech and Audio Representation via Distillation V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3bd4328-5d1c-4bef-9692-8674e551f07c · outbound
USAD: Universal Speech and Audio Representation via Distillation Gigaspeech: An evolving, multi- domain asr corpus with 10,000 hours of transcribed audio,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7a3a4c6c-d745-42b7-8974-4864d027eae2 · outbound
USAD: Universal Speech and Audio Representation via Distillation Common voice: A massively-multilingual speech corpus,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3c27e679-e562-4d1b-a244-f98e5a118880 · outbound
USAD: Universal Speech and Audio Representation via Distillation The fisher corpus: A resource for the next generations of speech-to-text
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 287d4539-4ada-4f48-945e-e56defe31875 · outbound
USAD: Universal Speech and Audio Representation via Distillation V oxlingua107: a dataset for spoken language recognition,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ec99ff25-39c1-479f-af6c-f5297ae458f4 · outbound
USAD: Universal Speech and Audio Representation via Distillation Audio set: An ontology and human-labeled dataset for audio events,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bfdbea0-dfaa-491d-b550-6c3049143255 · outbound
USAD: Universal Speech and Audio Representation via Distillation Soundnet: Learning sound representations from unlabeled video,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cde9ede7-e5e0-4ccd-87d9-c4566b585fc0 · outbound
USAD: Universal Speech and Audio Representation via Distillation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10effc48-0b51-4139-aa79-29293a9bee66 · outbound
USAD: Universal Speech and Audio Representation via Distillation Music4all: A new music database and its applications,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38c9d75c-da90-4391-b0db-e86f3fa19f13 · outbound
USAD: Universal Speech and Audio Representation via Distillation fairseq: A fast, extensible toolkit for sequence modeling,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f5a81b72-a1be-4027-a593-5ac2b613813c · outbound
USAD: Universal Speech and Audio Representation via Distillation Layer normalization,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6dc93b6-4b9d-40c9-afe1-92e3732586c1 · outbound
USAD: Universal Speech and Audio Representation via Distillation Self-attention with relative position representations,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8f1808d9-5252-4f93-bbdd-6ed5c959396b · outbound
USAD: Universal Speech and Audio Representation via Distillation Speech commands: A dataset for limited-vocabulary speech recognition,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b16ec7cc-281d-45e9-b90e-a5bac8a52c51 · outbound
USAD: Universal Speech and Audio Representation via Distillation SUPERB: Speech processing universal performance benchmark,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4926ca56-4405-4134-acab-f5b787454b07 · outbound
USAD: Universal Speech and Audio Representation via Distillation SUPERB-SG: Enhanced speech processing universal PERformance benchmark for semantic and generative capabilities,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eb20a4d8-dfbf-478c-834c-61c39d3e8eea · outbound
USAD: Universal Speech and Audio Representation via Distillation A large-scale evaluation of speech foundation models,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3ef690ab-d0ea-4524-b326-34161c0a3c03 · outbound
USAD: Universal Speech and Audio Representation via Distillation Hear: Holistic evaluation of audio representations,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 20dcce66-1464-49a9-a821-5d5ab6a1bc5f · outbound
USAD: Universal Speech and Audio Representation via Distillation ESC: Dataset for Environmental Sound Classification,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1c975ef0-3fe5-43a6-bd22-1e4d3b6f7699 · outbound
USAD: Universal Speech and Audio Representation via Distillation Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 509a8141-b875-4dfd-a562-9bf5547bf3f3 · inbound
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations USAD: Universal Speech and Audio Representation via Distillation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c39e7d-ebbd-417a-bfa5-7ab57dcca0a4 · inbound
Alethia: A Foundational Encoder for Voice Deepfakes USAD: Universal Speech and Audio Representation via Distillation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation adc33ca8-be64-4795-a30f-20f8f644cdb5 · inbound
Stage-adaptive audio diffusion modeling USAD: Universal Speech and Audio Representation via Distillation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.