REVIEW 6 cited by
Quantitative Survey of the State of the Art in Sign Language Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work presents a meta study covering around 300 published sign language recognition papers with over 400 experimental results. It includes most papers between the start of the field in 1983 and 2020. Additionally, it covers a fine-grained analysis on over 25 studies that have compared their recognition approaches on RWTH-PHOENIX-Weather 2014, the standard benchmark task of the field. Research in the domain of sign language recognition has progressed significantly in the last decade, reaching a point where the task attracts much more attention than ever before. This study compiles the state of the art in a concise way to help advance the field and reveal open questions. Moreover, all of this meta study's source data is made public, easing future work with it and further expansion. The analyzed papers have been manually labeled with a set of categories. The data reveals many insights, such as, among others, shifts in the field from intrusive to non-intrusive capturing, from local to global features and the lack of non-manual parameters included in medium and larger vocabulary recognition systems. Surprisingly, RWTH-PHOENIX-Weather with a vocabulary of 1080 signs represents the only resource for large vocabulary continuous sign language recognition benchmarking world wide.
Forward citations
Cited by 6 Pith papers
-
Bringing Balance to Hand Shape Classification: Mitigating Data Imbalance Through Generative Models
Pre-training an EfficientNet classifier on GAN-generated balanced hand images, then fine-tuning on real data, raises accuracy on the imbalanced RWTH handshape benchmark from 80.6% to 85.3%.
-
The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?
On a 12-gloss German Sign Language dataset, mouth and full-face crops improve isolated sign recognition more than eyes do, with mouth matching full face when fused with body input.
-
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
Sign-IDD generates sign language poses by disentangling joint coordinates into bone direction and length attributes inside a gloss-conditioned diffusion model, reporting state-of-the-art scores on PHOENIX14T and USTC-CSL.
-
Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
Synthetic depth images from Depth Anything V2 can support point-cloud sign language recognition with accuracies close to, and in some models above, real depth data.
-
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
StgcDiff generates sign language transition frames with a graph-conditioned diffusion model, improving temporal coherence over prior methods on PHOENIX14T, USTC-CSL100, and USTC-SLR500.
-
Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observation
Sign-D2C uses a conditional diffusion model with random masking and linear interpolation padding to generate transition poses between discrete signs, but its evaluation is largely a same-task reconstruction rather tha...
Discussion (0). Continue with ORCID to comment.