{"work":{"id":"7cd6d289-dca2-414f-99e0-809f37c065fa","openalex_id":"https://openalex.org/W4400434308","doi":"10.48550/arxiv.2407.04051","arxiv_id":"2407.04051","raw_key":null,"title":"FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs","authors":null,"authors_text":"K","year":2024,"venue":"cs.SD","abstract":"This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https://fun-audio-llm.github.io, and the code can be accessed at https://github.com/FunAudioLLM.","external_url":"https://arxiv.org/abs/2407.04051","cited_by_count":7,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2407.04051","created_at":"2026-05-10T05:20:54.953794+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Funaudiollm: Voice understanding and generation foundation models for natural interaction between humans and llms.arXiv preprint arXiv:2407.04051","render_title":"Funaudiollm: Voice understanding and generation foundation models for natural interaction between humans and llms.arXiv preprint arXiv:2407.04051"},"hub":{"state":{"work_id":"7cd6d289-dca2-414f-99e0-809f37c065fa","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":26,"external_cited_by_count":7,"distinct_field_count":8,"first_pith_cited_at":"2024-12-03T17:41:24+00:00","last_pith_cited_at":"2026-07-07T17:43:36+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T03:59:30.374873+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4},{"context_role":"method","n":3}],"polarity_counts":[{"context_polarity":"background","n":4},{"context_polarity":"use_method","n":3}],"runs":{},"summary":{},"graph":{},"authors":[]}}