Pith. sign in

REVIEW 5 cited by

Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21646 v2 pith:RSAAD6FQ submitted 2024-07-31 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords translationclasihumaninformationachievesimultaneoussystemsagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present Cross Language Agent -- Simultaneous Interpretation, CLASI, a high-quality and human-like Simultaneous Speech Translation (SiST) System. Inspired by professional human interpreters, we utilize a novel data-driven read-write strategy to balance the translation quality and latency. To address the challenge of translating in-domain terminologies, CLASI employs a multi-modal retrieving module to obtain relevant information to augment the translation. Supported by LLMs, our approach can generate error-tolerated translation by considering the input audio, historical context, and retrieved information. Experimental results show that our system outperforms other systems by significant margins. Aligned with professional human interpreters, we evaluate CLASI with a better human evaluation metric, valid information proportion (VIP), which measures the amount of information that can be successfully conveyed to the listeners. In the real-world scenarios, where the speeches are often disfluent, informal, and unclear, CLASI achieves VIP of 81.3% and 78.0% for Chinese-to-English and English-to-Chinese translation directions, respectively. In contrast, state-of-the-art commercial or open-source systems only achieve 35.4% and 41.6%. On the extremely hard dataset, where other systems achieve under 13% VIP, CLASI can still achieve 70% VIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.

  2. StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A unified large speech-language model uses speech chain-of-thought to jointly perform segmentation, generation-policy decisions, and streaming translation.

  3. SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SeqPO-SiMT uses sequential policy optimization with a combined quality-and-latency reward to improve simultaneous machine translation, beating supervised fine-tuning on six En-Zh and Zh-En datasets.

  4. PHRASED: Phrase Dictionary Biasing for Speech Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Phrase dictionary biasing, which matches source phrases in intermediate ASR text and then boosts or prompts the matching target phrases, improves phrase recall in streaming and LLM-based speech translation.

  5. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

Pith tools