REVIEW 6 cited by
OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced SULMs are developed by the industry, leveraging large-scale datasets and computational resources that are not readily available to the academic community. Moreover, the lack of transparency in training details creates additional barriers to further innovation. In this study, we present OSUM, an Open Speech Understanding Model designed to explore the potential of training SLUMs under constrained academic resources. The OSUM model combines a Whisper encoder with a Qwen2 LLM and supports a wide range of speech tasks, including speech recognition (ASR), speech recognition with timestamps (SRWT), vocal event detection (VED), speech emotion recognition (SER), speaking style recognition (SSR), speaker gender classification (SGC), speaker age prediction (SAP), and speech-to-text chat (STTC). By employing an ASR+X training strategy, OSUM achieves efficient and stable multi-task training by simultaneously optimizing ASR alongside target tasks. Beyond delivering strong performance, OSUM emphasizes transparency by providing openly available data preparation and training methodologies, offering valuable insights and practical guidance for the academic community. By doing so, we aim to accelerate research and innovation in advanced SULM technologies.
Forward citations
Cited by 6 Pith papers
-
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
Self-distillation of Qwen3-Omni-Thinking on 545k Cogito-Pipe audio reasoning traces yields the best open-source MMAR CoT scores and top-tier challenge ranking.
-
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.
-
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
A three-stage trained speech-to-speech chatbot with an explicit think step transfers paralinguistic understanding into empathetic responses, outperforming prior end-to-end spoken dialogue systems on a new LLM-scored b...
-
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.
-
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
MEUSLI, a released family of linear projectors, enables open-source LLM-based ASR across 28 European languages and supports few-hour adaptation to unseen languages and extra speech tasks.
-
BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation
A teacher-assistant distillation framework bridges the representation gap between LiDAR-camera and camera-only BEV segmentation, improving camera-only mIoU by 4.2% on nuScenes.
Discussion (0). Continue with ORCID to comment.