REVIEW 4 cited by
Chain-of-Thought Prompting for Speech Translation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance in automatic speech recognition (ASR) and automatic speech translation (AST). In this work, we propose a novel approach to leverage ASR transcripts as prompts for AST in a Speech-LLM built on an encoder-decoder text LLM. The Speech-LLM model consists of a speech encoder and an encoder-decoder structure Megatron-T5. By first decoding speech to generate ASR transcripts and subsequently using these transcripts along with encoded speech for prompting, we guide the speech translation in a two-step process like chain-of-thought (CoT) prompting. Low-rank adaptation (LoRA) is used for the T5 LLM for model adaptation and shows superior performance to full model fine-tuning. Experimental results show that the proposed CoT prompting significantly improves AST performance, achieving an average increase of 2.4 BLEU points across 6 En->X or X->En AST tasks compared to speech prompting alone. Additionally, compared to a related CoT prediction method that predicts a concatenated sequence of ASR and AST transcripts, our method performs better by an average of 2 BLEU points.
Forward citations
Cited by 4 Pith papers
-
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation
MAVL is a five-language lyrics benchmark pairing animated-song lyrics with audio and video, and the SylAVL-CoT prompt method improves syllable-count fit and user-rated singability over text-only translation baselines.
-
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
A speech-to-speech language model uses channel fusion of a streaming encoder and codec tokens to handle barge-in and turn-taking without speech pretraining, showing improved metrics over Moshi at 0.6 kbps.
-
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.
-
XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion
XiHeFusion is a Qwen2.5-14B model fine-tuned on 1.2 million fusion knowledge pairs to answer nuclear fusion questions for science communication.
Discussion (0). Continue with ORCID to comment.