Pith. sign in

REVIEW 1 cited by

Cross-modality Data Augmentation for End-to-End Sign Language Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11096 v4 pith:FUUFXFKV submitted 2023-05-18 cs.CL

classification cs.CL
keywords languagesigntranslationcross-modalitydataend-to-endxmdageneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end sign language translation (SLT) aims to convert sign language videos into spoken language texts directly without intermediate representations. It has been a challenging task due to the modality gap between sign videos and texts and the data scarcity of labeled data. Due to these challenges, the input and output distributions of end-to-end sign language translation (i.e., video-to-text) are less effective compared to the gloss-to-text approach (i.e., text-to-text). To tackle these challenges, we propose a novel Cross-modality Data Augmentation (XmDA) framework to transfer the powerful gloss-to-text translation capabilities to end-to-end sign language translation (i.e. video-to-text) by exploiting pseudo gloss-text pairs from the sign gloss translation model. Specifically, XmDA consists of two key components, namely, cross-modality mix-up and cross-modality knowledge distillation. The former explicitly encourages the alignment between sign video features and gloss embeddings to bridge the modality gap. The latter utilizes the generation knowledge from gloss-to-text teacher models to guide the spoken language text generation. Experimental results on two widely used SLT datasets, i.e., PHOENIX-2014T and CSL-Daily, demonstrate that the proposed XmDA framework significantly and consistently outperforms the baseline models. Extensive analyses confirm our claim that XmDA enhances spoken language text generation by reducing the representation distance between videos and texts, as well as improving the processing of low-frequency words and long sentences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Fusing spatio-temporal RGB and OpenPose features via intra- and cross-modal temporal modeling plus contrastive LLM fine-tuning yields new gloss-free SOTA on PHOENIX14T (BLEU-4 27.10) and CSL-Daily (BLEU-4 25.85).

Pith tools