Pith. sign in

REVIEW 1 cited by

Diverse Sign Language Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19586 v1 pith:T6K23GTL submitted 2024-10-25 cs.MM cs.CV

classification cs.MMcs.CV
keywords languagetranslationdiversesigndivsltmultipletaskaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in the case of limited data. In this work, we introduce a Diverse Sign Language Translation (DivSLT) task, aiming to generate diverse yet accurate translations for sign language videos. Firstly, we employ large language models (LLM) to generate multiple references for the widely-used CSL-Daily and PHOENIX14T SLT datasets. Here, native speakers are only invited to touch up inaccurate references, thus significantly improving the annotation efficiency. Secondly, we provide a benchmark model to spur research in this task. Specifically, we investigate multi-reference training strategies to enable our DivSLT model to achieve diverse translations. Then, to enhance translation accuracy, we employ the max-reward-driven reinforcement learning objective that maximizes the reward of the translated result. Additionally, we utilize multiple metrics to assess the accuracy, diversity, and semantic precision of the DivSLT task. Experimental results on the enriched datasets demonstrate that our DivSLT method achieves not only better translation performance but also diverse translation results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition

    cs.CV 2025-02 conditional novelty 2.0 of 10

    An ensemble of three Video Swin Transformer sizes with RGB and depth fusion achieves 20.29% RGB and 24.53% RGB-D top-1 accuracy, ranking third in the CV-ISLR challenge.

Pith tools