Pith. sign in

REVIEW 12 cited by

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.06282 v1 pith:X2MB53XE submitted 2025-01-10 cs.CL cs.AIcs.HCcs.SDeess.AS

classification cs.CLcs.AIcs.HCcs.SDeess.AS
keywords minmomodelsvoicespeechalignmentmultimodalalignedapproximately
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice interactions are categorized as native and aligned. Native models integrate speech and text processing in one framework but struggle with issues like differing sequence lengths and insufficient pre-training. Aligned models maintain text LLM capabilities but are often limited by small datasets and a narrow focus on speech tasks. In this work, we introduce MinMo, a Multimodal Large Language Model with approximately 8B parameters for seamless voice interaction. We address the main limitations of prior aligned multimodal models. We train MinMo through multiple stages of speech-to-text alignment, text-to-speech alignment, speech-to-speech alignment, and duplex interaction alignment, on 1.4 million hours of diverse speech data and a broad range of speech tasks. After the multi-stage training, MinMo achieves state-of-the-art performance across various benchmarks for voice comprehension and generation while maintaining the capabilities of text LLMs, and also facilitates full-duplex conversation, that is, simultaneous two-way communication between the user and the system. Moreover, we propose a novel and simple voice decoder that outperforms prior models in voice generation. The enhanced instruction-following capabilities of MinMo supports controlling speech generation based on user instructions, with various nuances including emotions, dialects, and speaking rates, and mimicking specific voices. For MinMo, the speech-to-text latency is approximately 100ms, full-duplex latency is approximately 600ms in theory and 800ms in practice. The MinMo project web page is https://funaudiollm.github.io/minmo, and the code and models will be released soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  2. A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff

    cs.IT 2026-04 unverdicted novelty 6.0 of 10

    Synset-based reconstruction and synonymous variational inference are claimed to derive the distributional divergence in RDP and unify it with classical rate-distortion theory.

  3. Differentiable Reward Optimization for LLM based TTS system

    cs.SD 2025-07 conditional novelty 6.0 of 10

    DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.

  4. Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Using Mimi neural codec features with label-delayed training reduces endpoint cutoff errors by 42.7% (single-stream) and 37.5% (two-stream) at 160 ms median latency.

  5. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  6. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.

  7. Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Qwen-Audio-3.0-TTS claims state-of-the-art controllable multilingual text-to-speech across 16 languages and 20 Chinese dialects, using a 12.5 Hz tokenizer and multi-stage RL.

  8. Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

    cs.CV 2026-02 conditional novelty 5.0 of 10

    Ex-Omni is an OLLM that natively generates speech and ARKit-52 3D facial animation in one pass by using discrete speech units as temporal scaffolding and gated semantic injection.

  9. FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A full-duplex voice system with streaming personalized VAD and semantic end-of-turn detection reports fewer false barge-ins and latencies near commercial benchmarks.

  10. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  11. RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

    cs.AI 2025-06 conditional novelty 5.0 of 10

    RoboEgo combines a full-duplex audio/text backbone with vision and action heads to achieve an 80 ms theoretical response granularity, reporting competitive quality and better responsiveness than Qwen2.5-Omni in a smal...

  12. SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge

    cs.CL 2025-07 conditional novelty 4.0 of 10

    SHNU-mASR, a parallel-encoder LLM system, achieves 11.76% CER/WER on the MLC-SLM blind eval set, 8.41 points better than the official baseline.

Pith tools