Pith. sign in

REVIEW 3 cited by

Everybody Sign Now: Translating Spoken Language to Photo Realistic Sign Language Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.09846 v4 pith:RJTSNX5T submitted 2020-11-19 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords languagesignphoto-realisticposespokenvideodeafdirectly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To be truly understandable and accepted by Deaf communities, an automatic Sign Language Production (SLP) system must generate a photo-realistic signer. Prior approaches based on graphical avatars have proven unpopular, whereas recent neural SLP works that produce skeleton pose sequences have been shown to be not understandable to Deaf viewers. In this paper, we propose SignGAN, the first SLP model to produce photo-realistic continuous sign language videos directly from spoken language. We employ a transformer architecture with a Mixture Density Network (MDN) formulation to handle the translation from spoken language to skeletal pose. A pose-conditioned human synthesis model is then introduced to generate a photo-realistic sign language video from the skeletal pose sequence. This allows the photo-realistic production of sign videos directly translated from written text. We further propose a novel keypoint-based loss function, which significantly improves the quality of synthesized hand images, operating in the keypoint space to avoid issues caused by motion blur. In addition, we introduce a method for controllable video generation, enabling training on large, diverse sign language datasets and providing the ability to control the signer appearance at inference. Using a dataset of eight different sign language interpreters extracted from broadcast footage, we show that SignGAN significantly outperforms all baseline methods for quantitative metrics and human perceptual studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A vector-quantized autoencoder with parameter-disentangled streams and phonological semi-supervision reconstructs and identifies unseen ASL signs better than a standard VQ-VAE on the Sem-Lex benchmark.

  2. Towards AI-driven Sign Language Generation with Non-manual Markers

    cs.HC 2025-02 conditional novelty 6.0 of 10

    The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated genera...

  3. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

Pith tools