REVIEW 5 cited by
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we present Cross Language Agent -- Simultaneous Interpretation, CLASI, a high-quality and human-like Simultaneous Speech Translation (SiST) System. Inspired by professional human interpreters, we utilize a novel data-driven read-write strategy to balance the translation quality and latency. To address the challenge of translating in-domain terminologies, CLASI employs a multi-modal retrieving module to obtain relevant information to augment the translation. Supported by LLMs, our approach can generate error-tolerated translation by considering the input audio, historical context, and retrieved information. Experimental results show that our system outperforms other systems by significant margins. Aligned with professional human interpreters, we evaluate CLASI with a better human evaluation metric, valid information proportion (VIP), which measures the amount of information that can be successfully conveyed to the listeners. In the real-world scenarios, where the speeches are often disfluent, informal, and unclear, CLASI achieves VIP of 81.3% and 78.0% for Chinese-to-English and English-to-Chinese translation directions, respectively. In contrast, state-of-the-art commercial or open-source systems only achieve 35.4% and 41.6%. On the extremely hard dataset, where other systems achieve under 13% VIP, CLASI can still achieve 70% VIP.
Forward citations
Cited by 5 Pith papers
-
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.
-
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
A unified large speech-language model uses speech chain-of-thought to jointly perform segmentation, generation-policy decisions, and streaming translation.
-
SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
SeqPO-SiMT uses sequential policy optimization with a combined quality-and-latency reward to improve simultaneous machine translation, beating supervised fine-tuning on six En-Zh and Zh-En datasets.
-
PHRASED: Phrase Dictionary Biasing for Speech Translation
Phrase dictionary biasing, which matches source phrases in intermediate ASR text and then boosts or prompts the matching target phrases, improves phrase recall in streaming and LLM-based speech translation.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
Discussion (0). Continue with ORCID to comment.