Pith. sign in

Paper Citation Record · LEDGER

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 12 inbound Pith citation observations for arXiv:2501.13306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13306 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:21:50.792032Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:51.939607Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:49:50.788200Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4845ecc-ab65-4973-bb13-35ee447b6f80 · outbound

This paper cites Data Products , 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data Products , 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.254183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.641410Z digest=sha256:9838337f05405cba535c073c763eac3600dc2e3cdbb935ba6e9e6c66a13e3884

Observation 0a0579ea-5f4d-4fa8-a1e7-9c9f4ee76847 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.645939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.645939Z digest=sha256:720811d7dd8f52a385761317ee30709d443515a3b46e8b203a54d05f91709bad

Observation 26405e1f-01c6-4177-9c8e-34b8d827b0a4 · outbound

This paper cites Qwen Technical Report.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.650773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.650773Z digest=sha256:4f53219f004b3fa7743a8f9d12f141ffe48c9ddce9bcb4d248465b282d88d65a

Observation 1d2f7d2d-bf49-4319-9d59-0b74ce33614c · outbound

This paper cites AISHELL-1 : An open-source mandarin speech corpus and a speech recognition baseline.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-1 : An open-source mandarin speech corpus and a speech recognition baseline

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.241788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.655249Z digest=sha256:bfa8115ecf48c459c735ec9476e725238654da6cfab9b90432b5a072c1766fe9

Observation 46728167-1fc3-45b8-b624-2198cb239ab6 · outbound

This paper cites IEMOCAP : Interactive emotional dyadic motion capture database.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia IEMOCAP : Interactive emotional dyadic motion capture database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.229299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.660136Z digest=sha256:b2815a35e2c71ffda11903099b96a49314237d4672c99efee0f2822b46d0d1fa

Observation 52440562-690e-4e16-84cb-b880bc7a35ee · outbound

This paper cites MSP-IMPROV : An acted corpus of dyadic interactions to study emotion perception.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MSP-IMPROV : An acted corpus of dyadic interactions to study emotion perception

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.217573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.664124Z digest=sha256:3a53bf5474b7b8162815c1543e8f67ff04e0e14e7b180d07405f6e6e1db1bbf1

Observation 15110fa2-db69-4e3d-a69b-0a745922a32d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.668642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.668642Z digest=sha256:650b132e0df6b0c7c1c6f97751fd7b7539f5354e77df0fe44bececad8011897f

Observation 953d4eb8-70f3-4527-b300-24d2dec5e61f · outbound

This paper cites Qwen2-Audio Technical Report.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Qwen2-Audio Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.672430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.672430Z digest=sha256:6efc5362d56b23d6f2dbf627b04a6adc4106db4ae3fc9d1eaa519f76c35ec993

Observation 2b3e2324-61fc-46a6-8ea9-9c3b754f2a8b · outbound

This paper cites Data products, 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.205841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.676243Z digest=sha256:a32e028822c146c0d364668e7c870c5e0814675dd0f793eefdb451affaa4adf1

Observation 0013ddfe-9b79-4433-88e9-fedde57b0f45 · outbound

This paper cites Data products, 2024.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Data products, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.194923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.679557Z digest=sha256:6363c74ba36e135e62e3b335de2e997e18ba51eb1e19cdaecb01bbeba9097d65

Observation ef476ca6-1f6f-4dad-8424-4384230db56e · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.682871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.682871Z digest=sha256:a5554bcd0dea959b5fb771933b9d492b37a9abc00e24d225a9e7db2a9fbc7e60

Observation 02eea089-deb6-4bff-ab24-2e277fa8f5e9 · outbound

This paper cites Gemmeke, Daniel P.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Gemmeke, Daniel P

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.184074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.688136Z digest=sha256:de25cd19a259fe1a43f26353ecac269d9934350cb3f475fe9b03361f1c9c5e21

Observation b763e6cd-8fb7-45a2-ac99-ef753df47ed6 · outbound

This paper cites Vocalsound: A dataset for improving human vocal sounds recognition.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Vocalsound: A dataset for improving human vocal sounds recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.173901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.692800Z digest=sha256:08320f3557c755c6186e9514ff854a68345a7e8e8146583b19e927b3bbb13147

Observation f7be6a17-c051-4edf-a4c8-4012dc488c1a · outbound

This paper cites LoRA : Low -rank adaptation of large language models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia LoRA : Low -rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.163633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.697178Z digest=sha256:33dd0b3ff421d4ddb877bcc8d6a68aa1bf9c583fd774697be27bb3346cae8a88

Observation d2030894-3781-4c39-91b1-699135cd2541 · outbound

This paper cites Datasets, 2017.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Datasets, 2017

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.153105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.701458Z digest=sha256:cb7a25574fc63345ff52a832479383f2739aaed81365ce44de8be834a7172161

Observation b228b375-8c03-4806-9680-93ae9f01d650 · outbound

This paper cites Schuller, and Jianhua Tao.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Schuller, and Jianhua Tao

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.142135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.705820Z digest=sha256:df6d22fa40781cafd2fb126d9cde6cd8d527df9cc701342beb9abd7a5b2ec2bd

Observation f66d6d84-5633-41ca-9be7-69b3c0021db3 · outbound

This paper cites Emotion2vec: Self-supervised pre-training for speech emotion representation.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.131662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.709957Z digest=sha256:6661def6745bec4d3542b078b5f0c90c96faa29be6f574d907d7be080549d873

Observation fdfd17ed-1cf7-4c0e-be38-c59aaeadb688 · outbound

This paper cites The MSP -conversation corpus.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The MSP -conversation corpus

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.120340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.714460Z digest=sha256:bab277ec54b40b1d7392587dd2c5589d57af7f58c033827243aac71c577a1eea

Observation f9167552-5a2f-4ed4-9def-0c061fd591f3 · outbound

This paper cites MAGICDATA mandarin Chinese read speech corpus, 2019.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MAGICDATA mandarin Chinese read speech corpus, 2019

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.107988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.718846Z digest=sha256:cae797937dae2c070340f605b70684a04d6c125bdba32ed5b7661a61786cf448

Observation a9ba5083-904c-4640-9995-ae09ad4aef9e · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Librispeech: An ASR corpus based on public domain audio books

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.096486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.724151Z digest=sha256:4533bcdf7a3f149ee20e56eb9b46a23c7814baaa77adccb90e3a8d6eeb45b757

Observation 40c2663a-8e81-4e98-b77a-ac501ac14610 · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Reproducing whisper-style training using an open-source toolkit and publicly available data

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.084881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.728388Z digest=sha256:c183dee055442547c5f9b583516464e11789fe4e75a030d0b1faf9bd63b78030

Observation a5df2899-512f-413b-933e-4b6cfa6d6f51 · outbound

This paper cites an unresolved cited work.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:21:51.072413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.731940Z digest=sha256:9121672515e18322155645a402f4a28338a64aca32c339a9ca7e666097b464f4

Observation 0c1bfbc9-4cb3-435b-9387-85f3af8b9463 · outbound

This paper cites MELD : A multimodal multi-party dataset for emotion recognition in conversations.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia MELD : A multimodal multi-party dataset for emotion recognition in conversations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.061704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.735946Z digest=sha256:485dbc4a4d3ff22a834dc6c492ce64a4eef102f9031ddaa3a6eaa13bc706deb4

Observation c4678dbb-125c-49d6-8e58-cd612f3a968a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Robust speech recognition via large-scale weak supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.051351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.739872Z digest=sha256:7f77f5e5f576e1fe41b9d1eba74d8013f22445f92951e3e62b539404b5734afa

Observation 495dd3a5-d28c-4dd6-a54d-f0e27ab3d062 · outbound

This paper cites Nonspeech7k dataset: Classification and analysis of human non-speech sound.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Nonspeech7k dataset: Classification and analysis of human non-speech sound

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.039921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.743915Z digest=sha256:f38c5b2800e9b4322d1dd21e0933f9bca2f30b0577b172232a5b61ac1a6343d5

Observation 64e3ad61-25f5-4207-90f2-ffa3e92de62b · outbound

This paper cites The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:21:50.748041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:21:50.748041Z digest=sha256:e989672191864dc6e1b5c2889eb10780d20dbfb47a3cfef2af90970aa35b697e

Observation e20f0866-7a74-4ce1-8c57-5d36fc155563 · outbound

This paper cites Achieving timestamp prediction while recognizing with non-autoregressive end-to-end ASR model.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Achieving timestamp prediction while recognizing with non-autoregressive end-to-end ASR model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.026269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.751915Z digest=sha256:7433ad1d8fd724a5049f7cf3a4b1466df0f555d464ecaa1847e7f9a5b0217703

Observation afe57756-a2d3-470d-80f4-fd23a3dffd5b · outbound

This paper cites TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:21:50.833353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.755899Z digest=sha256:394426e53a93ccd9ec8ecbe556140fed8aa15de7b2a30b78a035c50172126a5c

Observation 84ec1fc9-4e59-48ac-8f98-fc5b362f883e · outbound

This paper cites PandaGPT : One model to instruction-follow them all.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia PandaGPT : One model to instruction-follow them all

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.014485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.760638Z digest=sha256:9d4e2ad5374c9e1ac3fcd58ed0dcebd437fc0082daa23a8c99abaa24719fd733

Observation 5fcabe55-126a-48c5-bc0c-739cc90a3f2d · outbound

This paper cites SALMONN : Towards generic hearing abilities for large language models.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia SALMONN : Towards generic hearing abilities for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:51.002925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.764260Z digest=sha256:849513b9d2f7253f89d057321488178f8c237afbc6448429f6c4a995cd01df98

Observation 2f9dd831-0c14-453c-9f29-ec7c9ad4fe32 · outbound

This paper cites Kespeech: An open source speech dataset of Mandarin and its eight subdialects.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Kespeech: An open source speech dataset of Mandarin and its eight subdialects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.992013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.767988Z digest=sha256:70e6a51135ab2e491c9afb710e42cc5b7816d9455fda4d46c3b2641f57410f30

Observation 5eef7f39-a106-4968-9c33-982a1ac3b2c2 · outbound

This paper cites Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.980216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.771903Z digest=sha256:441a376955343551e9d3e03bd372ac455913612cea4ea8343aa1081743e2ba99

Observation 52d13f61-9502-479a-936d-b5b890a9a7ae · outbound

This paper cites Attention is all you need.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.967826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.775557Z digest=sha256:0ea3ef8c4fffec045f8731e1cc338df19001cb1e6381def2e00efcf507d3f2b0

Observation caff80b2-1551-4bc4-b5e7-265f7e582f6b · outbound

This paper cites A large-scale Chinese short-text conversation dataset.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia A large-scale Chinese short-text conversation dataset

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.955944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.779084Z digest=sha256:5ea88563581616bb38526b8df5f6041403a91f483798b736d6e47f3df4c263d7

Observation 75834008-8b63-4873-963e-aa9b98735c35 · outbound

This paper cites WENETSPEECH : A 10000+ hours multi-domain Mandarin corpus for speech recognition.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia WENETSPEECH : A 10000+ hours multi-domain Mandarin corpus for speech recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.943291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.783589Z digest=sha256:c7b7b7f4b772c69386ee24c8b70b40918dc6156c4ff18f650c885be5ef57b96c

Observation 98512b3b-063f-4236-be2d-d25226d42cf0 · outbound

This paper cites M 3 ED : Multi -modal multi-scene multi-label emotional dialogue database.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia M 3 ED : Multi -modal multi-scene multi-label emotional dialogue database

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.930655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.788155Z digest=sha256:efbc09cae5bf68b41a353432c4f83c92862e9a824c2a7d8f1f8c3e2e705cc134

Observation 26a6b988-ecf9-42c2-8381-48f36e2c8c64 · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset.

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:21:50.916362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T16:21:50.792032Z digest=sha256:7b56661fc33befd9e8b202cc911fc467bfc6e891cab10693e1bd2bf9d7917b04

Pith citing papers

Observation 203b544f-3836-4423-a139-4e5a5f95024d · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.283015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:93474420664a71c4f68c1e629710f3b7d44bc4e8eaabf8538d4966548ddbfde2

Observation 314f9479-bc79-4f09-bcb5-5fa240649b68 · inbound

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models cites this paper.

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.939607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.939607Z digest=sha256:22f3af362dcb96a8fc6b2a69caa061c68e1c72b9d53cb2b28cfd06ab0c44afe0

Observation 27882c48-9fd5-4dce-80ee-c7519ae63114 · inbound

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation cites this paper.

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T20:59:46.041859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:59:46.041859Z digest=sha256:1dc95b9617b943c8717e8b24c47f8af733ec6af4096f4501cc0ba8f36f99780c

Observation 787c8920-5478-449b-bb67-7791784322b4 · inbound

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue cites this paper.

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T21:03:09.640862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:03:09.640862Z digest=sha256:4f0b8fec45e006db01b0a71cf19a364db4188a8e5c82ae35a9cbe06b2e80b926

Observation 93284ac9-394a-42dc-9ced-1d4b75bc47dd · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.228057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:cdb2b5aa51947860a9da1cc7cd957c72887fc576d0aaac41cd4c108e7056244e

Observation 267ba0df-de4f-406e-956d-5b4178a6c813 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:38.479016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:38.479016Z digest=sha256:2c110d7678182ebd416543ac641605af4aac568863598f7b21c3f45db17a470a

Observation 196a7acf-ff2b-4f85-a4b7-7f2990d536e4 · inbound

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models cites this paper.

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.898403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:24:27.118694Z digest=sha256:4096c628959dea455f20a7b58e17e15caf8aabb174bc1cb5fef4630bf3844e40

Observation 6ead0d2e-a824-4f9f-951d-49a72e4765b3 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.342560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:acfd24152e7fd26213f6283f440ab8cbd8632e0e23fac026986a48493a323051

Observation d2db76cc-80e9-4dfd-93d8-fbf512f19a12 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:aa5c955f7a27dd461dd32e216c8e62876f70492ad06b2060f4d21165a1588b53

Observation ae99d2f1-9b81-426c-9ce1-96113b53d4bc · inbound

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages cites this paper.

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.006517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:25:37.725216Z digest=sha256:26345af2e22ebf4b711d62022fa0f1b6d4821679112119ded38a25a1b211c096

Observation 760bc7e5-e472-4085-b43d-a2c487a220c8 · inbound

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios cites this paper.

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:49:50.790654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T07:36:34.307652Z digest=sha256:22e90509c675e1c9d753fa0fa28d2d1327ac96476a61b5096e4cdfe0c52a99a8

Observation 459daaf3-3a19-4091-a8cc-2608f1eb74b1 · inbound

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond cites this paper.

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T05:50:29.618737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:50:29.618737Z digest=sha256:4f12c1038a13e2a546ec5844cd85d7dd931891efb8cf785591437093982c7aca