Pith. sign in

REVIEW 4 cited by

Chain-of-Thought Prompting for Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11538 v2 pith:PBSUR7KS submitted 2024-09-17 cs.CL

classification cs.CL
keywords speechpromptingtranscriptsmodelmodelsperformancespeech-llmtranslation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance in automatic speech recognition (ASR) and automatic speech translation (AST). In this work, we propose a novel approach to leverage ASR transcripts as prompts for AST in a Speech-LLM built on an encoder-decoder text LLM. The Speech-LLM model consists of a speech encoder and an encoder-decoder structure Megatron-T5. By first decoding speech to generate ASR transcripts and subsequently using these transcripts along with encoded speech for prompting, we guide the speech translation in a two-step process like chain-of-thought (CoT) prompting. Low-rank adaptation (LoRA) is used for the T5 LLM for model adaptation and shows superior performance to full model fine-tuning. Experimental results show that the proposed CoT prompting significantly improves AST performance, achieving an average increase of 2.4 BLEU points across 6 En->X or X->En AST tasks compared to speech prompting alone. Additionally, compared to a related CoT prediction method that predicts a concatenated sequence of ASR and AST transcripts, our method performs better by an average of 2 BLEU points.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    MAVL is a five-language lyrics benchmark pairing animated-song lyrics with audio and video, and the SylAVL-CoT prompt method improves syllable-count fit and user-rated singability over text-only translation baselines.

  2. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.

  3. LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

    cs.AI 2025-05 conditional novelty 5.0 of 10

    LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.

  4. XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion

    cs.CV 2025-02 reject novelty 4.0 of 10

    XiHeFusion is a Qwen2.5-14B model fine-tuned on 1.2 million fusion knowledge pairs to answer nuclear fusion questions for science communication.

Pith tools