Pith. sign in

REVIEW 8 cited by

BBC-Oxford British Sign Language Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.03635 v1 pith:3RGXGUCX submitted 2021-11-05 cs.CV

classification cs.CV
keywords datasetsignlanguagebobslbritishavailablebbc-oxforddata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we introduce the BBC-Oxford British Sign Language (BOBSL) dataset, a large-scale video collection of British Sign Language (BSL). BOBSL is an extended and publicly released dataset based on the BSL-1K dataset introduced in previous work. We describe the motivation for the dataset, together with statistics and available annotations. We conduct experiments to provide baselines for the tasks of sign recognition, sign language alignment, and sign language translation. Finally, we describe several strengths and limitations of the data from the perspectives of machine learning and linguistics, note sources of bias present in the dataset, and discuss potential applications of BOBSL in the context of sign language technology. The dataset is available at https://www.robots.ox.ac.uk/~vgg/data/bobsl/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Isharah is a new 30,000-clip, multi-scene Saudi Sign Language dataset with gloss and translation annotations, plus signer-independent and unseen-sentence benchmarks for continuous sign language recognition and translation.

  2. Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Hard negatives selected by visual confusability in sign embeddings, not linguistic similarity, substantially raise fine-grained sign-language retrieval accuracy without collapsing coarse performance.

  3. SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

    cs.CV 2026-03 accept novelty 6.0 of 10

    Sparse keyframe-conditioned Conditional Flow Matching produces fluid, articulate 3D sign language motion across four languages while enabling precise Keyframe-to-Pose editing.

  4. Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dual visual encoder with contrastive visual-text pretraining achieves the best reported BLEU-4 score among gloss-free sign language translation methods on Phoenix-2014T.

  5. iLSU-T: an Open Dataset for Uruguayan Sign Language Translation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    iLSU-T is a 187-hour Uruguayan Sign Language video dataset with Spanish text, 18 interpreters, and first baseline translation results.

  6. Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LLM-generated pseudo glosses, reordered via weak video supervision, enable sign language translation that rivals gloss-supervised models while needing only 30 gloss examples.

  7. Sign Spotting Disambiguation using Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LLM-based beam search disambiguation improves dictionary sign spotting WER from 47.2% to 44.4% on an internal BSL dataset.

  8. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

Pith tools