REVIEW 7 cited by
Uni-Sign: Toward Unified Sign Language Understanding at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose Uni-Sign, a unified pre-training framework that eliminates the gap between pre-training and downstream SLU tasks through a large-scale generative pre-training strategy and a novel fine-tuning paradigm. First, we introduce CSL-News, a large-scale Chinese Sign Language (CSL) dataset containing 1,985 hours of video paired with textual annotations, which enables effective large-scale pre-training. Second, Uni-Sign unifies SLU tasks by treating downstream tasks as a single sign language translation (SLT) task during fine-tuning, ensuring seamless knowledge transfer between pre-training and fine-tuning. Furthermore, we incorporate a prior-guided fusion (PGF) module and a score-aware sampling strategy to efficiently fuse pose and RGB information, addressing keypoint inaccuracies and improving computational efficiency. Extensive experiments across multiple SLU benchmarks demonstrate that Uni-Sign achieves state-of-the-art performance across multiple downstream SLU tasks. Dataset and code are available at github.com/ZechengLi19/Uni-Sign.
Forward citations
Cited by 7 Pith papers
-
Attention-Steered Vision-Language Models for Sign Language Translation
AttnSign adds spatial attention supervision and motion-cadence reinforcement learning to a VLM, improving sign language translation accuracy on How2Sign and OpenASL.
-
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Sign Language QA benchmarks are introduced from PHOENIX14T and CSL-Daily via template-generated questions, and a question-conditioned baseline outperforms video-language and cascaded baselines.
-
DESign: Dynamic Context-Aware Convolution and Efficient Subnet Regularization for Continuous Sign Language Recognition
A sign language recognition model using context-aware dynamic convolutions and subnetwork CTC regularization reports new state-of-the-art word error rates on PHOENIX14, PHOENIX14-T, and CSL-Daily.
-
Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation
LLM-generated pseudo glosses, reordered via weak video supervision, enable sign language translation that rivals gloss-supervised models while needing only 30 gloss examples.
-
SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation
SAGE uses a frozen sign-segmentation model to turn sign videos into about half as many visual tokens as prior methods, then aligns those tokens with a language model to reach BLEU-4 of 24.10 on PHOENIX14T.
-
Sign Spotting Disambiguation using Large Language Models
LLM-based beam search disambiguation improves dictionary sign spotting WER from 47.2% to 44.4% on an internal BSL dataset.
-
Using Sign Language Production as Data Augmentation to enhance Sign Language Translation
Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.
Discussion (0). Continue with ORCID to comment.