{"id":"6f2a0527-e4bb-4b57-87e3-bdf9a683ad06","arxiv_id":"2508.05374","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A corpus of 12 sign languages with spoken-language subtitles: 1,300+ hours across 4,381 videos and 14M subtitle tokens, including the first consistent parallel corpora for 8 Latin American sign languages.","lead":"This paper describes a new collection of more than 1,300 hours of sign language videos with spoken-language subtitles, covering 12 sign languages, including the first substantial parallel corpora for 8 Latin American sign languages. It matters because large, consistent parallel data is the main missing ingredient for sign language machine translation and recognition.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The corpus's 'parallel' claim rests on the assumption that subtitles translate signed content; broadcast subtitles often transcribe the audio instead, so the central claim is unverified.","rationale":"The reader's weakest assumption is exactly the one I identify: the subtitles may not be parallel to the signed content. The central claim of 'parallel corpora' for 12 sign languages depends on this assumption, and the abstract provides no methodology to verify it. In broadcast data, subtitles are commonly intralingual captions of the audio track, not translations of the signed message; this is a well-known problem in sign-language corpus work. The concern is therefore not speculative in an empty sense—it targets a specific, testable property of the data that determines whether the resource supports the claimed MT and linguistic uses. I found no internal inconsistency or ad hominem issue. The concrete test would resolve the concern: a human evaluation of a stratified sample. If the test passes, the abstract's claim is supported; if it fails, the corpus would need to be reframed or re-annotated. Because the reader already assigned UNVERDICTED precisely on the basis of insufficient verification of the data claims, my concern reinforces that verdict rather than changing it. I would not move to ACCEPT or REJECT on the current evidence, since the corpus itself may still be valuable even if the subtitles are not perfectly parallel. Thus, the verdict remains UNCHANGED, with the caveat that the suggested test is a prerequisite for any future ACCEPT verdict once full text and data access are available.","tokens_in":893,"tokens_out":2859,"duration_ms":30887,"concrete_test":"Draw a stratified random sample of 100 video-subtitle pairs (at least 8 per language). For each pair, have two independent fluent signers of that sign language watch the video with subtitles and classify every subtitle segment as: (a) a faithful translation of the signed content; (b) a transcription of the audio that diverges from the signed content; or (c) unrelated. Report inter-annotator agreement and the per-language proportion of (a). If the proportion of (a) is below, say, 95%—or if annotators cannot reliably distinguish—then the corpus is not demonstrably parallel, and the abstract should be revised to 'signed videos with synchronous subtitles' pending manual alignment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's headline claim is that the collection provides 'consistent parallel corpora' for 12 sign languages with spoken-language subtitles. The load-bearing premise is that each subtitle is a translation of the signed content, not merely a verbatim caption of the audio track. In broadcast settings—news, government, education—subtitles are frequently intralingual captions of the spoken audio, which can diverge substantially from the signed interpretation (e.g., the signer may summarize, expand, or re-order; the audio may describe off-screen content). If the subtitles are not semantic counterparts of the signing, then downstream sign-to-text MT trained on these pairs would learn a weak or misleading mapping, and the 'parallel' characterization is overstated. The abstract gives no evidence for the alignment (e.g., manual validation, timing alignment, or language-identification checks), and no full text is available to inspect methodology. This is not an internal inconsistency but an empirically unsupported assumption at the center of the contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes a collection of parallel sign-language corpora in video format with spoken-language subtitles. It reports over 1,300 hours across 4,381 video files, 1.3 million subtitles and 14 million tokens, covering 12 sign languages. The main selling points are the first consistent parallel corpora for 8 Latin American sign languages and a roughly tenfold increase in the size of German Sign Language corpora relative to previous resources. The collection is based on online broadcast, government, and educational videos, with a processing pipeline including scraping, cropping, consent-seeking, and statistics reporting. Only the abstract was made available for this review; consequently, my assessment is limited to the claims in the abstract and cannot confirm the validity of the headline numbers.","tokens_in":970,"tokens_out":3625,"duration_ms":33676,"significance":"If the claims are accurate, this is a potentially valuable public resource for sign-language NLP and linguistic research, particularly for underserved Latin American sign languages. The reported scale is an order of magnitude beyond previous DGS resources, and the explicit consent-seeking step is commendable. However, the central claim that the corpus is 'parallel' depends on the relationship between the signed content and the subtitles. Because the abstract provides no evidence of alignment quality, language identification, or annotation validation, the significance can only be assessed provisionally. The strength of the contribution would be substantially enhanced by reported alignment-validation statistics and a clear release protocol.","major_comments":[{"comment":"The abstract characterizes the resource as 'consistent parallel corpora' and states that subtitles in dominant spoken languages accompany the videos. This is the load-bearing claim for downstream sign-to-text MT. In broadcast settings, subtitles are frequently verbatim captions of the audio track rather than translations of the signed interpretation, and the signed content may summarize or restructure the spoken content. The abstract provides no evidence about how subtitle-sign alignment was established (e.g., manual validation, inter-annotator agreement, timing alignment, or a filtering protocol). Without such evidence, the 'parallel' claim is unverified. I ask the authors to provide in the paper a precise definition of 'consistent parallel' and the alignment-validation protocol.","section":"Abstract"},{"comment":"The comparative claims—'first consistent parallel corpora for 8 Latin American sign languages' and 'ten times the size' of prior German Sign Language corpora—depend on the completeness and fairness of the external baselines. The abstract names no prior corpora, no search methodology, and no inclusion/exclusion criteria. To make these claims falsifiable, the full paper must specify the baseline collection used for comparison and the per-language statistics (hours, files, tokens, sign-language variety) that justify the 'first' and 'ten times' statements.","section":"Abstract"},{"comment":"The abstract does not report a language-identification or label-verification protocol. Because sign languages have varieties and the videos come from heterogeneous online sources, it is not clear how each video was assigned to a specific sign language and whether the spoken-language subtitle language was independently verified. This is particularly important for multilingual countries where the 'dominant spoken language' is not unique. The paper should report per-language and per-source counts and any automatic or manual checks applied.","section":"Abstract"}],"minor_comments":[{"comment":"The notation '1,3~M subtitles' appears to be a typographical slip; it should read '1.3M subtitles'.","section":"Abstract"},{"comment":"The phrase 'ten times the size of the previously available corpora' does not specify the metric (hours, files, or tokens). Please clarify which quantity is being compared.","section":"Abstract"},{"comment":"The term 'consistent parallel' is not defined in the abstract. Adding a parenthetical such as 'with human-verified subtitle-sign correspondence' would help readers judge the claim.","section":"Abstract"},{"comment":"The abstract does not state where the corpus will be released or under what access terms. A URL or a statement of availability would be useful in the abstract.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based on the abstract only, as the full text was not provided. The central claims—particularly the parallel alignment and the comparative 'first'/'ten times' statements—cannot be verified. I recommend soliciting the full manuscript before a final decision. My 'uncertain' verdict is a request for evidence, not an accusation of error."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. The pitch is straightforward: 1,300+ hours of signed video for 12 sign languages, with spoken-language subtitles, and the first consistent parallel corpora for 8 Latin American sign languages plus a 10x scale-up for German Sign Language. If the data is as described, it would be a real asset for sign-to-text MT and recognition, especially for languages that have almost nothing. The corpus-construction pipeline — collection, creator approvals, scraping, cropping — is the right way to do this, and the paper's stated goal of providing statistics is sensible.\n\nThe soft spot is also the load-bearing one. The word 'parallel' means the subtitles have to be translations or semantic counterparts of the signed content. Broadcast subtitles are often captions of the audio track, and the signed interpretation can diverge — the signer may summarize or re-order, or the audio may describe things the signer does not convey. The abstract gives no evidence of alignment validation: no manual checks, no timing alignment, no language-ID protocol. The claim is not internally contradictory, but it is unverified. Same goes for per-file language labels: there is no stated protocol for identifying which sign language is actually in each video, which matters when mixing sources across 12 languages.\n\nI also cannot check the 'first' and 'ten times' claims from the abstract alone. Those are benchmarked against prior corpora, and the authors may well be right, but the evidence is not here.\n\nNone of this is a fatal flaw. It is an abstract-only submission, and the weaknesses are all of the form 'not shown yet,' not 'shown to be wrong.' The right outcome is to send the full paper to a referee who can inspect the statistics section, the alignment-quality checks, and the data-access terms. If the paper reports even modest validation of subtitle-sign alignment, it deserves to be published and used. If it does not, the authors should be asked to add that analysis before release, because downstream MT trained on misaligned pairs would be misleading.\n\nVerdict: I'd take the full paper seriously. The resource, if real, is important; the risk is in the alignment assumption, which is exactly what a referee should probe.","headline":"Potentially major sign-language resource, but the abstract leaves the central 'parallel' claim unverified; worth refereeing on the full paper.","tokens_in":1632,"tokens_out":1603,"would_cite":true,"duration_ms":16026,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new collection provides parallel video corpora for 12 sign languages, including first consistent data for 8 Latin American sign languages.","keywords":["sign language corpora","parallel corpora","Latin American sign languages","German Sign Language","video subtitles","broadcast data","low-resource languages","machine translation data"],"falsifier":"Take a random sample of videos across all 12 languages, have fluent signers watch the videos without sound and write what is signed, then compare those transcripts to the subtitle text. Low alignment rates in any language would falsify the parallel-corpus claim for that language. Additionally, running automatic sign-language identification on a sample could test whether each file is genuinely in the stated sign language.","tokens_in":674,"feed_emoji":"🤟","tokens_out":2839,"duration_ms":29063,"temperature":0.7,"pith_summary":"The paper presents a collection of parallel sign-language corpora: 12 sign languages in video form, paired with subtitles in the dominant spoken languages of the corresponding countries. The authors aim to establish that this is the first consistent parallel corpus for 8 Latin American sign languages, and that the German Sign Language portion is ten times larger than previously available data. The full collection totals more than 1,300 hours across 4,381 video files, with 1.3 million subtitles containing 14 million tokens. Because sign-language research has been held back by scarce, inconsistent data, a large parallel video-subtitle corpus is a step toward training translation models and doing corpus linguistics for these languages at scale. The paper contributes the collection itself, statistics about it, and a description of the data-collection pipeline.","feed_headline":"1,300 hours of signed video, 12 languages, now parallel","feed_subtitle":"First consistent parallel corpora for 8 Latin American sign languages; German Sign Language data ten times larger than before.","key_machinery":"The central object is a parallel sign-language corpus: video recordings of signing paired with subtitle text in the surrounding spoken language, processed into a uniform collection. The key work performed by this object is to turn scattered broadcast material into a comparable, machine-readable resource where each sign language has the same kind of video-subtitle pairing. That consistency and scale is what allows the eight Latin American sign languages and German Sign Language to be treated as parallel resources rather than as isolated clips.","core_discovery":"The central claim is that the authors have assembled a parallel corpus of 12 sign languages in video form, with subtitles in the dominant spoken languages of the corresponding countries, and that this collection is an order-of-magnitude advance for at least nine of those languages. Specifically, it provides the first consistent parallel corpora for 8 Latin American sign languages, and it makes the German Sign Language portion ten times larger than the previously available corpus. The collection totals more than 1,300 hours in 4,381 videos, 1.3 million subtitles, and 14 million tokens. The data were gathered from online broadcast, governmental, and educational sources through a pipeline that","pith_inferences":["The same collection pipeline could likely be extended to other broadcast-available sign languages beyond the 12, since the method is not language-specific.","If the subtitles turn out to be transcripts of the audio rather than faithful translations of the signing, the 'parallel' property would be weaker than claimed; this is testable by human judgment on a sample.","Because language labels appear to come from source channels rather than from sign-language identification, dialectal and regional variation within each labeled language may be conflated in the corpus."],"forward_implications":["Machine translation for the eight Latin American sign languages can be trained on consistent parallel data for the first time.","German Sign Language models gain ten times more training data, which should yield substantially better translation quality.","The 14 million subtitle tokens provide a volume that makes statistical and neural approaches to sign-language translation viable.","The collection's method, including consent-seeking and cropping, offers a template for building similar resources for other sign languages.","The published statistics give the field a baseline for measuring future growth of sign-language corpora."],"supporting_citations":[],"fun_headline_variants":["12 sign languages, 1,300+ hours of parallel video","First parallel corpora for 8 Latin American sign languages","German Sign Language data now 10 times larger","1,300+ hours: new parallel sign language corpus","14M tokens: sign language corpus spans 12 languages"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The central claim depends on the premise that the subtitles are genuine translations of what is signed in each video and that each video is correctly attributed to the stated sign language; if subtitles merely transcribe the audio track, the corpus is not truly parallel.","fun_headline_variants_meta":{"raw":{"variants":["12 sign languages, 1,300+ hours of parallel video","First parallel corpora for 8 Latin American sign languages","German Sign Language data now 10 times larger","1,300+ hours: new parallel sign language corpus","14M tokens: sign language corpus spans 12 languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1297,"prompt_tokens":674,"completion_tokens":623,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":556}},"tokens_in":418,"tokens_out":623,"duration_ms":6365,"temperature":1.0,"reasoning_tokens":556,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:23:01.069731+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of videos across all 12 languages, have fluent signers watch the videos without sound and write what is signed, then compare those transcripts to the subtitle text. Low alignment rates in any language would falsify the parallel-corpus claim for that language. Additionally, running automatic sign-language identification on a sample could test whether each file is genuinely in the stated sign language.","supporting_citations":[],"review_version":1}