Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:16.177062Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2505.16369.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:16.177062Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:12.925006Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T14:17:03.115400Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c036217-80f1-4122-bc8b-f229c44341d8 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbca00b5-4be5-4ff5-8913-45f65a97eef7 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4a2c8fe-1142-42e7-879a-c8299e39c340 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1232899-2511-4dea-b526-3db91cf0a0a5 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 704ada23-c0ac-4b4b-86a8-cd5792e6c03c · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f11c02f-2ad4-4b77-9ca6-58eae6c42276 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993755cc-6b45-4341-9d80-3e9e0a935380 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc18be9a-6119-47be-887e-83e6eed81894 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Speech tasks in X-ARES assess both linguistic content (e.g., speech content, word spotting) and paralinguistic features (e.g., emotion, speaker identity, accent)
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dfd6b30-c20b-4582-962b-d9c6eec7a980 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Task-Specific Metrics A summary of all metrics used in X-ARES is provided in Ta- ble 1
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e4da20e-5f83-48ab-8ee3-207cda013197 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance First, speech encoders such as data2vec [3], HuBERT [35], wav2vec2- large [2], WavLM [36] and Whisper-base [37]
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19e6f7d6-df6b-4ca8-86f1-4b048b78bcca · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1527872a-4813-475c-ae7b-70f75ee2c235 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Scal- ing up masked audio encoder learning for general audio classifica- tion,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06656f89-cb63-4e21-a1d1-0c6d306a15f6 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ed79dd-3a94-4892-9e2d-c850a7aee805 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Data2vec: A general framework for self-supervised learning in speech, vision and language,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439a7da9-d5ea-410b-b0ae-9e61248b5594 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac3bb46-8cea-453b-a5b0-f256e9e01d9a · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Moshi: a speech-text foundation model for real-time dialogue,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5f248cf-2e0c-4431-ace5-6751d1a27ea2 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e52c619-c6af-46ad-b5c9-5343f2bac86d · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance HEAR: Holistic evaluation of audio representations,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c729de9a-d587-4aad-9f1a-e3fed0156598 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance SUPERB: Speech processing universal performance benchmark,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dcec1deb-46c5-4495-bcb4-7b79390f6993 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance DASB - Discrete Audio and Speech Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c21023-6c9d-4848-b8e6-0289d70ae5fa · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Webdataset: A library for efficient loading of large-scale datasets,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73c8ef13-d7cc-4755-9169-6b8a683ff4b4 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771974a9-2a03-491a-86a7-4004fae0a578 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Auto- matic speaker verification spoofing and countermeasures challenge (asvspoof 2015) database,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a866a237-3e77-4dbc-898b-87660ba7ee8c · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Crema-d: Crowd-sourced emotional multimodal actors dataset,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5e70fa5-877b-462d-8c31-0932f4dfd7eb · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Speech Model Pre-training for End-to-End Spoken Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ffa220-145d-49d7-b987-8b9d6fc74cfa · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Libricount, a dataset for speaker count estimation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a8b808b-fa75-4ebb-b7db-3ae814e7ac6e · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Librispeech: an asr corpus based on public domain audio books,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 965c1ea6-0086-4115-af28-0b07ec0223a9 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, mul- timodal set of facial and vocal expressions in north american en- glish,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa3a3cd7-6b1c-4d24-acea-f2532c1d4aaa · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance V ocalsound: A dataset for improving human vocal sounds recognition,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de60e614-7968-4087-ac1d-bcad00f2d2b5 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed14a54-584f-4ce3-936f-e8b78d8d23da · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance V oxceleb: Large-scale speaker verification in the wild,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a978da9-62e7-4c3f-aadd-e26ad02fa055 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance V oxlingua107: a dataset for spoken lan- guage recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c1653dc-7cf2-4660-8f5b-24600d37b9a4 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Clotho: An audio cap- tioning dataset,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5aa12670-3ac5-4565-b519-cc23cd978ad0 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f338b51a-d029-47c9-b170-a18b304474ff · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Esc: Dataset for environmental sound classification,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 634f069f-c613-4450-af48-f650ef6c26d0 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81afb060-0127-4091-84d4-bcf206baac7b · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Fsd50k: an open dataset of human-labeled sound events,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f14be62-c8e4-40b2-a172-651dbe937d78 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance A dataset and taxonomy for urban sound research,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68e2ddd3-4fc4-4647-8672-3c4caa763074 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance V ocal imitation set: a dataset of vocally imitated sound events using the audioset ontol- ogy
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed313b6f-be7c-49e9-9de2-3c4aff4b2c79 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance FMA: A Dataset For Music Analysis
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cab7c97-93df-4044-a300-c21eb5f64746 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08233fbb-8c66-45fb-9510-c75134e287b1 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Enabling factorized piano music modeling and generation with the MAESTRO dataset,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92b1ddae-ce78-4abb-be0c-85afeef3dcd3 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Neural audio synthesis of musical notes with wavenet autoencoders,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79a4b8a-b26f-42a1-b1b4-ab7a545d4d09 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d4c6043-9121-43c0-bad2-bdb496a44859 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Qwen2.5 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499d38d2-16dd-4216-a0c9-77f8f2d873a0 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2369ea91-14c9-4942-917f-7a350be7f87a · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38afb6aa-897c-41a8-8494-e6cc89f06562 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Robust speech recognition via large-scale weak supervision,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7beb06ff-c36e-4a0f-a60f-e687ac17af26 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6b32bd-cf43-428a-b26f-4614cd7e9644 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Byol-s: Learning self-supervised speech representations by bootstrapping,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d91046c1-1e72-4745-9417-b1265cd2227b · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Ced: Consis- tent ensemble distillation for audio tagging,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 042f13fe-47ea-4425-8537-3c102c889efe · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c617c87-c424-4433-88e4-a67a9575db48 · outbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Masked modeling duo: Learning representations by encouraging both networks to model the input,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f11c02f-2ad4-4b77-9ca6-58eae6c42276 · inbound
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054cc496-1d59-4e3e-8e16-1c7faa9956a9 · inbound
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50834e18-e74a-4b31-a36e-a43cb59cea13 · inbound
Probing Spatial Structure in Pretrained Audio Representations X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed45fb58-73c8-4f9a-baa7-ded281058537 · inbound
Probing Spatial Structure in Pretrained Audio Representations X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.