Pith. sign in

REVIEW 8 cited by

Spirit LM: Interleaved Spoken and Written Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05755 v2 pith:UN3NDKLH submitted 2024-02-08 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechmodeltextspiritunitslanguagemodelsabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Spirit LM, a foundation multimodal language model that freely mixes text and speech. Our model is based on a 7B pretrained text language model that we extend to the speech modality by continuously training it on text and speech units. Speech and text sequences are concatenated as a single stream of tokens, and trained with a word-level interleaving method using a small automatically-curated speech-text parallel corpus. Spirit LM comes in two versions: a Base version that uses speech phonetic units (HuBERT) and an Expressive version that models expressivity using pitch and style units in addition to the phonetic units. For both versions, the text is encoded with subword BPE tokens. The resulting model displays both the semantic abilities of text models and the expressive abilities of speech models. Additionally, we demonstrate that Spirit LM can learn new tasks in a few-shot fashion across modalities (i.e. ASR, TTS, Speech Classification). We make available model weights and inference code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Global token perplexity mis-ranks spoken language models; localized/normalized likelihood scores and an embedding judge track human MOS better and make the best model look much closer to human.

  2. Autoregressive Speech Enhancement via Acoustic Tokens

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.

  3. OpusLM: A Family of Open Unified Speech Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.

  4. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  5. Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Using 80 ms speech segments and 16,384 sound tokens improves zero-shot spoken language understanding and cuts training cost by up to 70%.

  6. An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Smaller discrete vocabularies (k = 125 to 1,000), WavLM units, and larger models give the lowest negative log-likelihood in speech language model pre-training.

  7. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  8. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.

Pith tools