Pith. sign in

REVIEW 9 cited by

BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.00916 v2 pith:PAP3QSG4 submitted 2023-09-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechalignmentmodalityapproachbehaviorlanguagellmstext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of large language models (LLMs) has sparked significant interest in extending their remarkable language capabilities to speech. However, modality alignment between speech and text still remains an open problem. Current solutions can be categorized into two strategies. One is a cascaded approach where outputs (tokens or states) of a separately trained speech recognition system are used as inputs for LLMs, which limits their potential in modeling alignment between speech and text. The other is an end-to-end approach that relies on speech instruction data, which is very difficult to collect in large quantities. In this paper, we address these issues and propose the BLSP approach that Bootstraps Language-Speech Pre-training via behavior alignment of continuation writing. We achieve this by learning a lightweight modality adapter between a frozen speech encoder and an LLM, ensuring that the LLM exhibits the same generation behavior regardless of the modality of input: a speech segment or its transcript. The training process can be divided into two steps. The first step prompts an LLM to generate texts with speech transcripts as prefixes, obtaining text continuations. In the second step, these continuations are used as supervised signals to train the modality adapter in an end-to-end manner. We demonstrate that this straightforward process can extend the capabilities of LLMs to speech, enabling speech recognition, speech translation, spoken language understanding, and speech conversation, even in zero-shot cross-lingual scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.

  2. Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Language-aware query selection with a gated query bank improves multilingual, ASR-only-distilled speech LLMs on instruction following and spoken QA.

  3. Towards Reliable Large Audio Language Model

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Training a large audio language model to say 'I don't know' on one audio type (speech, music, or sound) makes it more likely to refuse uncertain questions on the other types.

  4. Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A contrastive-style adapter trained on LLM-generated positive and negative audio descriptions improves audio hallucination accuracy to 77.5 percent and audio question answering to 84.3 percent, without changing the fr...

  5. TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TESU-LLM shows that a frozen LLM can answer spoken queries after training only a 13M-parameter projector on text, using SeamlessM4T's shared speech-text encoder.

  6. Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Multi-task behavior imitation with speech-text interleaving improves speech LLM generalization on prompts and zero-shot tasks using only paired speech and transcripts.

  7. SparQLe: Speech Queries to Text Translation Through LLMs

    cs.CL 2025-02 conditional novelty 5.0 of 10

    SparQLe uses a Q-Former adapter to bridge frozen HuBERT speech features to Llama-3, achieving BERTScore gains over IWSLT 2022 baselines on English-French and zero-shot English-German speech translation.

  8. DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality

    cs.HC 2025-08 conditional novelty 4.0 of 10

    DESAMO runs the Qwen2.5-Omni audio LLM on a Jetson Orin Nano to classify elderly users' voice commands and ambient sounds on-device, reaching 98% intent accuracy on a 300-sample pilot.

  9. Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    Under matched settings, continuous SSL speech features generally outperform discrete tokens on six spoken language understanding tasks in SpeechLLMs.

Pith tools