Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:21:50.792032Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 12 inbound Pith citation observations for arXiv:2501.13306.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:21:50.792032Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:51.939607Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:49:50.788200Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c4845ecc-ab65-4973-bb13-35ee447b6f80 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data Products , 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0a0579ea-5f4d-4fa8-a1e7-9c9f4ee76847 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26405e1f-01c6-4177-9c8e-34b8d827b0a4 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2f7d2d-bf49-4319-9d59-0b74ce33614c · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-1 : An open-source mandarin speech corpus and a speech recognition baseline
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46728167-1fc3-45b8-b624-2198cb239ab6 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia IEMOCAP : Interactive emotional dyadic motion capture database
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 52440562-690e-4e16-84cb-b880bc7a35ee · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MSP-IMPROV : An acted corpus of dyadic interactions to study emotion perception
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 15110fa2-db69-4e3d-a69b-0a745922a32d · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953d4eb8-70f3-4527-b300-24d2dec5e61f · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen2-Audio Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3e2324-61fc-46a6-8ea9-9c3b754f2a8b · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0013ddfe-9b79-4433-88e9-fedde57b0f45 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef476ca6-1f6f-4dad-8424-4384230db56e · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02eea089-deb6-4bff-ab24-2e277fa8f5e9 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Gemmeke, Daniel P
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b763e6cd-8fb7-45a2-ac99-ef753df47ed6 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Vocalsound: A dataset for improving human vocal sounds recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f7be6a17-c051-4edf-a4c8-4012dc488c1a · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia LoRA : Low -rank adaptation of large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2030894-3781-4c39-91b1-699135cd2541 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Datasets, 2017
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b228b375-8c03-4806-9680-93ae9f01d650 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Schuller, and Jianhua Tao
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f66d6d84-5633-41ca-9be7-69b3c0021db3 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Emotion2vec: Self-supervised pre-training for speech emotion representation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fdfd17ed-1cf7-4c0e-be38-c59aaeadb688 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The MSP -conversation corpus
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9167552-5a2f-4ed4-9def-0c061fd591f3 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MAGICDATA mandarin Chinese read speech corpus, 2019
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a9ba5083-904c-4640-9995-ae09ad4aef9e · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Librispeech: An ASR corpus based on public domain audio books
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 40c2663a-8e81-4e98-b77a-ac501ac14610 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Reproducing whisper-style training using an open-source toolkit and publicly available data
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5df2899-512f-413b-933e-4b6cfa6d6f51 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0c1bfbc9-4cb3-435b-9387-85f3af8b9463 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MELD : A multimodal multi-party dataset for emotion recognition in conversations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4678dbb-125c-49d6-8e58-cd612f3a968a · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Robust speech recognition via large-scale weak supervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 495dd3a5-d28c-4dd6-a54d-f0e27ab3d062 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Nonspeech7k dataset: Classification and analysis of human non-speech sound
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64e3ad61-25f5-4207-90f2-ffa3e92de62b · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20f0866-7a74-4ce1-8c57-5d36fc155563 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Achieving timestamp prediction while recognizing with non-autoregressive end-to-end ASR model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation afe57756-a2d3-470d-80f4-fd23a3dffd5b · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 84ec1fc9-4e59-48ac-8f98-fc5b362f883e · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia PandaGPT : One model to instruction-follow them all
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5fcabe55-126a-48c5-bc0c-739cc90a3f2d · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia SALMONN : Towards generic hearing abilities for large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f9dd831-0c14-453c-9f29-ec7c9ad4fe32 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Kespeech: An open source speech dataset of Mandarin and its eight subdialects
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5eef7f39-a106-4968-9c33-982a1ac3b2c2 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 52d13f61-9502-479a-936d-b5b890a9a7ae · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Attention is all you need
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation caff80b2-1551-4bc4-b5e7-265f7e582f6b · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia A large-scale Chinese short-text conversation dataset
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 75834008-8b63-4873-963e-aa9b98735c35 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia WENETSPEECH : A 10000+ hours multi-domain Mandarin corpus for speech recognition
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98512b3b-063f-4236-be2d-d25226d42cf0 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia M 3 ED : Multi -modal multi-scene multi-label emotional dialogue database
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26a6b988-ecf9-42c2-8381-48f36e2c8c64 · outbound
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 203b544f-3836-4423-a139-4e5a5f95024d · inbound
Kimi-Audio Technical Report OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 314f9479-bc79-4f09-bcb5-5fa240649b68 · inbound
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27882c48-9fd5-4dce-80ee-c7519ae63114 · inbound
BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 787c8920-5478-449b-bb67-7791784322b4 · inbound
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93284ac9-394a-42dc-9ced-1d4b75bc47dd · inbound
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 267ba0df-de4f-406e-956d-5b4178a6c813 · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196a7acf-ff2b-4f85-a4b7-7f2990d536e4 · inbound
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6ead0d2e-a824-4f9f-951d-49a72e4765b3 · inbound
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2db76cc-80e9-4dfd-93d8-fbf512f19a12 · inbound
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae99d2f1-9b81-426c-9ce1-96113b53d4bc · inbound
Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 760bc7e5-e472-4085-b43d-a2c487a220c8 · inbound
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 459daaf3-3a19-4091-a8cc-2608f1eb74b1 · inbound
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.