Pith. sign in

REVIEW 9 cited by

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18138 v1 pith:3ZZHWCZY submitted 2024-11-27 eess.AS cs.AIcs.CLcs.SD

classification eess.AScs.AIcs.CLcs.SD
keywords speechgenerationsalmonn-omnifull-duplexunderstandingcodec-freemodelacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations. Unlike traditional modularised conversational AI systems, which separate speech recognition, understanding, and text-to-speech generation into distinct components, multimodal LLMs operate as single end-to-end models. This streamlined design eliminates error propagation across components and fully leverages the rich non-verbal information embedded in input speech signals. We introduce SALMONN-omni, a codec-free, full-duplex speech understanding and generation model capable of simultaneously listening to its own generated speech and background sounds while speaking. To support this capability, we propose a novel duplex spoken dialogue framework incorporating a ``thinking'' mechanism that facilitates asynchronous text and speech generation relying on embeddings instead of codecs (quantized speech and audio tokens). Experimental results demonstrate SALMONN-omni's versatility across a broad range of streaming speech tasks, including speech recognition, speech enhancement, and spoken question answering. Additionally, SALMONN-omni excels at managing turn-taking, barge-in, and echo cancellation scenarios, establishing its potential as a robust prototype for full-duplex conversational AI systems. To the best of our knowledge, SALMONN-omni is the first codec-free model of its kind. A full technical report along with model checkpoints will be released soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

    cs.AI 2026-03 conditional novelty 6.5 of 10

    SocialOmni jointly evaluates speaker ID, turn-entry timing, and interruption phrasing on 12 OLMs and finds perception accuracy decouples from socially appropriate generation.

  2. Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Spoken math models that emit a 40%-compressed reasoning trace between question and answer beat full-reasoning baselines by ~3 accuracy points while using roughly one third of the text tokens.

  3. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  4. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.

  5. Sharp spectral estimates for free boundary problems arising in plasma physics

    math.AP 2026-04 unverdicted novelty 5.0 of 10

    For a constrained superlinear free-boundary plasma model, the non-local first eigenvalue σ₁ is always positive on balls in every dimension N≥2, despite lacking a general Faber–Krahn property.

  6. Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Stream-Omni uses CTC-based layer-dimension mapping to align speech with text, achieving vision, speech, and text interaction in one 8B model trained on 23,000 hours of speech.

  7. Towards a Japanese Full-duplex Spoken Dialogue System

    cs.CL 2025-06 conditional novelty 5.0 of 10

    J-Moshi, the first public Japanese full-duplex spoken dialogue model, is built from Moshi and outperforms a Japanese dGSLM baseline on naturalness and meaningfulness.

  8. X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance

    cs.SD 2025-05 conditional novelty 5.0 of 10

    X-ARES evaluates 13 audio encoders on 22 speech, sound, and music tasks using linear probing and nearest-neighbor classifiers, revealing strong domain-dependent performance differences.

  9. S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models

    cs.SD 2025-05 conditional novelty 4.0 of 10

    S2SBench quantifies intelligence degradation in speech-input LLMs by comparing perplexity between plausible and implausible continuations, and shows two-stage training of Baichuan-Audio reduces this degradation.

Pith tools