{"id":"0495d151-5078-4cad-938d-003ce7892a51","arxiv_id":"2506.14677","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"A real-time speech-to-sign system with an editable JSON layer and a local resampling hook claims large usability gains for deaf users, though several headline numbers conflict with the paper's tables.","lead":"This paper describes a system that turns spoken English into animated sign language, with a machine-readable JSON layer that lets deaf users inspect and fix individual sign segments. The intended use is real-time accessibility: users edit the animation, and the system is supposed to learn from their edits over time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline Edit-mode gains in Table 6 are not causally attributable to the editor: the reported Latin square balances task types, not Auto/Edit block order, so practice effects are a viable alternative explanation for every significant improvement.","rationale":"The reader's weakest assumption correctly identifies the most load-bearing threat to the central claim: the Edit-versus-Auto user study is not protected against order effects. The paper's abstract and introduction make causal claims about edit-in-the-loop improvements, and Table 6 is the primary evidence. Because the reported Latin square only balances task types, not condition order, any practice or familiarity effect accumulated during the Auto block fully contaminates the Edit block means. This is not a matter of choosing a different evaluation metric; it is a structural flaw in the experimental design that makes the headline numbers uninterpretable as evidence for the editor's value. The Appendix G admission that the HITL fine-tuning logs are placeholders further weakens the paper, but even setting that aside, the user-study confound alone invalidates the central human-centered contribution. I therefore agree with the reader's verdict and recommend no change: the paper should be rejected as it stands, because the main empirical claim is not supported by the reported experimental design. The proposed replication or re-analysis would settle the concern directly and could be used to reconsider the result if it is ever performed.","tokens_in":18018,"tokens_out":4832,"duration_ms":52732,"concrete_test":"Extract per-participant block order and raw scores from the session logs. If any participants received Edit before Auto, recompute Table 6 with block order as a factor; if none did, as the reported design suggests, run a preregistered fully counterbalanced replication with N=25, half Auto-first and half Edit-first. If the Edit-minus-Auto deltas shrink to non-significance or reverse when Edit is presented first, the claimed editing benefits are order artifacts rather than properties of the JSON editor and Resampling Hook.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that the JSON editor and Resampling Hook produce large, significant gains in comprehension, SUS, trust, agency, and cognitive load (Table 6, abstract). Those gains come from a within-subject comparison between Auto and Edit blocks. The only counterbalancing described is a 4x4 Latin square over four task types (Appendix C, Table A-4; Figure 3); no mechanism balances the order of the Auto and Edit blocks. As reported, the procedure states that each participant completed two blocks, Auto-generation (Auto) and Generation + Editing (Edit), and gives no order randomization. If the blocks were always administered in that listed order, every Edit condition follows every Auto condition, so all six Table 6 deltas—+28% comprehension, +24% naturalness, +19% SUS (elsewhere +13 points), +34% trust, -16% TLX, and +38% agency—are aliased with practice, interface familiarization, and task learning. This is an internal design problem, not a disagreement over evaluation conventions: the experimental design as documented cannot support the causal wording used in the abstract and introduction. A separate but compounding issue is Appendix G, which states all HITL fine-tuning logs 'will be replaced by real data when available,' so the continuous-adaptation contribution is also unverified; however, the order confound alone undermines the headline human-centered claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a real-time speech-to-sign-language animation system built from a streaming Conformer encoder, an autoregressive Transformer-MDN decoder, an editable JSON intermediate representation, and a human-in-the-loop fine-tuning loop. It reports benchmark results on WLASL100 against three external baselines and a user study with 20 deaf signers and 5 professional interpreters, claiming substantial gains in comprehension, usability, trust, and cognitive load in an edit-in-the-loop mode. The abstract frames the contributions as low-latency generation, user controllability, and continuous adaptation.","tokens_in":18146,"tokens_out":6016,"duration_ms":56124,"significance":"If the reported results held, the system would be a meaningful step: a 103 ms end-to-end sign-language animation pipeline with user editing and continuous adaptation, evaluated with external baselines and independent human raters. The choice of WLASL100, the inclusion of three external baselines, and the use of human participants are all appropriate and are strengths of the study design. However, the manuscript's central empirical claims are undermined by internal inconsistencies in the headline numbers, a benchmark conclusion that contradicts the paper's own table, a confounded user-study design, and an appendix that explicitly labels the adaptation data as placeholder. Consequently, the contribution cannot be assessed from the evidence as presented.","major_comments":[{"comment":"The Auto-versus-Edit comparison in Table 6 is confounded with block order. The procedure states that 'Each participant completed two blocks—Auto-generation (Auto) and Generation + Editing (Edit)' and does not describe any randomization or counterbalancing of the block order; the Latin-square scheme in Appendix C (Table A-4, Figure 3) balances only the four task types (G, I, T, E) across groups. Because every participant completed the Auto block before the Edit block, all seven deltas in Table 6 (+28% comprehension, +24% naturalness, +19% SUS, +34% trust, -16% TLX, +38% agency, -46% error recovery) are aliased with practice, interface familiarization, and task learning. The causal wording in the abstract and introduction ('edit-in-the-loop approach increased comprehension by 28%') is therefore not supported by the experimental design.","section":"Evaluation Methods and Results / Appendix C, Table 6"},{"comment":"The headline usability gain is not reproducible from the paper's own data. The abstract and introduction claim a '+13 point SUS improvement' and 'SUS +13'; Table 6 reports SUS 73.5±8.1 (Auto) versus 81.3±6.4 (Edit), a difference of 7.8 points. The table's '+19%' label is also inconsistent with the relative change, which is approximately +10.6%. These discrepancies mean the paper's most prominent quantitative claim must be corrected and reconciled with the underlying results.","section":"Abstract / Table 6"},{"comment":"The benchmark claims contradict Table 3. The Findings paragraph states that the system achieves 'best visual realism,' but FID is lower-better and SignDiff's FID (52) is lower than Ours (54), so SignDiff is better on this metric. The same paragraph says the system has 'highest understandability (+6.2 pp over the best baseline Fast-SLP)'; however, SignDiff achieves 57.0 SLR-Acc, so the best baseline is SignDiff and Ours is +0.4 pp above it, not +6.2. These errors invalidate the stated benchmark conclusions.","section":"Table 3 / Benchmark Comparison and Findings"},{"comment":"Appendix G explicitly states that all fine-tuning logs 'are generated based on typical throughput of a single RTX 5090 and i9-14900K workstation, and will be replaced by real data when available.' This means the week-by-week triplet accumulation (Table A-5), the hyper-parameter schedule (Table A-6), and the sample fine-tuning log (G.4) are simulated placeholder data, not experimental evidence. The continuous human-in-the-loop adaptation—one of the three stated key contributions—therefore has no empirical validation in the manuscript.","section":"Appendix G"}],"minor_comments":[{"comment":"The caption says 'TRT-INT8 latency for K=5,D=128 = 13 ms≈77 fps,' but the FPS column lists 24 for that row; please clarify whether the FPS column reports FP32 throughput and the latency column reports INT8 latency, and define the relationship between the two.","section":"Table 2"},{"comment":"The abstract reports a '13 ms average frame-inference time,' while the system architecture section decomposes end-to-end latency into audio, encoder, decoder, IK, and rendering components; please specify which quantity the abstract refers to.","section":"Abstract / System Architecture"},{"comment":"Table 7 reports Demographic Gap and Energy/frame reductions without significance values, although the text states 'ANOVA, p < 0.05' for the demographic gaps; please supply the test statistics or remove the claim.","section":"Table 7"},{"comment":"The Error-Recovery row in Table 6 has no p-value in the significance column, despite the text citing p < .001; please reconcile the table and the text.","section":"Table 6"},{"comment":"The evaluation section describes a 'NASA-TLX simplified version (C31–C34)' with four items, while Appendix D.7 describes the full six-dimensional weighted NASA-TLX procedure; please clarify which instrument was actually administered.","section":"Evaluation Methods and Results / Appendix D"}],"recommendation":"reject","confidential_remarks":"The manuscript contains several self-undermining features, notably the placeholder data in Appendix G and the mismatch between the abstract's SUS claim and Table 6. The user-study design flaw around block order is not fixable by a text revision; it requires a new experiment. If the authors run a properly counterbalanced study and reconcile all quantitative claims, the core idea may merit a fresh submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read of arXiv:2506.14677. The core idea is worth a look: an editable JSON intermediate representation for sign-language animation, with a local resampling hook that re-synthesizes only the edited segment, is genuinely new relative to the cited baselines. The drag-and-drop editing layer and the live MDN uncertainty heatmap are clever. If the system worked as described, it would be a real contribution for accessible, user-controllable signing avatars.\n\nBut the evidence as reported doesn't support the claims. The most serious problem is the user-study design: the procedure describes two blocks, Auto and Edit, with no randomization of block order; the Latin square only balances task types across groups. Since Edit always follows Auto, the large Table 6 gains (comprehension +28%, SUS +7.8, agency +38%) are confounded with practice and familiarization. That alone invalidates the causal wording in the abstract.\n\nThere are also internal contradictions that should have been caught before submission. The abstract promises a '+13 point SUS improvement'; Table 6 shows 81.3 vs 73.5, which is 7.8 points and a 10.6% relative gain, not the '+19%' in the table. The paper claims 'best visual realism' but Table 3 shows SignDiff's FID (52) is better than yours (54). Participant counts differ between the main text (20 deaf + 5 interpreters) and Appendix B (13 deaf + 7 interpreters + 5 linguists). And Appendix G states that all the HITL fine-tuning logs were 'generated based on typical throughput' and will be replaced by real data, so the paper's most novel contribution is currently backed by synthetic numbers, not measurements. No code, data, or model is released.\n\nThe architecture itself is described in enough detail that an expert could reconstruct it, and the editable-intermediate-layer idea deserves follow-up. But this version is not close to publishable. The right move is to fix the experiment (randomize Auto/Edit order, report the actual SUS difference), run the HITL loop on real user data, and re-check every table against the prose. If the authors do that, the core concept may well be a solid contribution.\n\nWho is this for? Researchers working on sign-language production or user-controllable generative animation, especially those interested in HCI-driven editing of continuous outputs. I would not cite it in its current form. For peer review, I'd say send it out rather than desk-reject: the novelty is real, and a good referee can give the authors a concrete path to fix the evaluation. With the evidence as it stands, though, the paper should be rejected.\n\nBest.","headline":"Genuinely novel editable JSON + resampling-hook idea, but the empirical evaluation is confounded and internally inconsistent; the paper should not be accepted in its current form.","tokens_in":18856,"tokens_out":4457,"would_cite":false,"duration_ms":44441,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An editable, real-time speech-to-sign animation system with a transparent JSON layer and local window resampling improves comprehension, usability, trust, and naturalness for deaf users and interpreters.","keywords":["speech-to-sign language generation","human-in-the-loop","editable JSON intermediate representation","mixture density network","streaming Conformer","sign language avatar animation","real-time inference","user study"],"falsifier":"A controlled replication with participant order counterbalanced—half doing Edit before Auto, with comparable tasks and blind scoring—would settle it: if comprehension, SUS, trust, and cognitive-load differences between modes shrink to non-significance when Auto is second, then practice or learning effects, not the editing mechanism, produced the reported gains.","tokens_in":17600,"feed_emoji":"🤟","tokens_out":10288,"duration_ms":94992,"temperature":0.7,"pith_summary":"This paper tries to establish that sign-language animation from live speech can be made real-time, editable, and trustworthy by opening the internal generation to users rather than treating it as a black box. The proposed system pairs a streaming Conformer encoder with an autoregressive Transformer-mixture-density decoder to generate upper-body and facial motion, translates each sign segment into an inspectable six-field JSON record, and lets users edit any field while a Resampling Hook re-synthesizes only the affected 50-frame window in about 75 ms. A human-in-the-loop loop then fine-tunes the model on logged edits and ratings. If the claims hold, a working assistive technology exists: 103 ms end-to-end latency on an RTX 4070, and user studies with 20 deaf signers and 5 professional interpreters reporting a +13-point System Usability Score improvement, a 6.7-point cognitive-load reduction, and significant gains in naturalness, comprehension, and trust for the editable mode over fully automatic generation.","feed_headline":"Real-time sign-language AI lets users edit, gaining 28% comprehension","feed_subtitle":"Deaf signers and interpreters rate the editable mode 13 SUS points higher, with 6.7-point lower cognitive load.","key_machinery":"The load-bearing machinery is the editable JSON action-structure plus the Resampling Hook: a transparent six-field schema (gloss identifier, handshape, trajectory, duration, non-manual markers, emphasis) that users can edit directly, and an inference-time module that, on any edit, patches the affected frames and performs a local forward pass of the autoregressive Transformer-mixture-density decoder (K=5 components on a 128-dimensional VAE latent) over a window of at most 50 frames, using k=8-12 context frames for continuity. This turns a single heavy end-to-end inference into many cheap local refreshes and gives the human a precise handle on sign identity, handshape, timing, facial markers, and emphasis. The supporting machinery is the streaming 6-layer Conformer encoder with causal state caching (30 ms per second of audio under TensorRT+INT8), the VAE-compressed latent space that preserves 99.3% of motion variance, and the human-in-the-loop fine-tuning loop that optimizes a KL-regularized PPO-style reward from user and expert ratings.","core_discovery":"The paper's central claim is that an editable intermediate representation is not a side feature but the mechanism that makes end-to-end sign-language generation usable: exposing signs as JSON with fields for gloss identifier, handshape, trajectory, duration, non-manual markers, and emphasis lets deaf users and interpreters inspect, correct, and personalize each segment, while the Resampling Hook re-runs the mixture-density decoder only over the edited window (at most 50 frames) with 8-12 context frames for smooth blending. This local re-synthesis, taking 75±9 ms, keeps the whole speech-to-avatar pipeline at 103±6 ms on an RTX 4070 and 13 ms per frame after pruning, INT8 quantization, and TensorRT acceleration. In benchmark comparisons on WLASL100 the system obtains the highest sign-language recognition accuracy among the compared generators at 57.4% and the best FID of 54 while running 1.2-3× faster. In the user study, Edit mode beats Auto mode on comprehension (+28%), naturalness (+24%), SUS (+13 points), trust (+34%), and error-recovery time (−46%), with model uncertainty displayed as a live MDN-weight heatmap, and accumulated edits feed a KL-regularized PPO fine-tuning loop that adapts the model every two weeks.","pith_inferences":["The six-field JSON schema is defined independently of the decoder weights, so it could be reused as a portable correction and personalization format across other sign-language production systems; a stable schema would let one user's edit transfer between engines.","The Resampling Hook is a general pattern for streaming generative models that need human correction: patch the latent and re-infer a local window rather than regenerating the whole sequence. A direct test would apply the same windowing idea to other long-form motion or speech generation tasks with similar latency budgets.","Because the MDN-weight heatmap is the only visual uncertainty cue, one can isolate its contribution by running the Edit condition with the heatmap hidden; the paper's trust and error-recovery gains would be expected to shrink if the heatmap is doing causal work.","The paper states in its appendix that the weekly triplet and fine-tuning logs are generated placeholder numbers to be replaced by real data, so the continuous-adaptation component is currently an architectural claim rather than an observed effect; logging real edits over several weeks would test it."],"forward_implications":["The system's 103±6 ms end-to-end latency places it under the 150 ms real-time threshold, making live speech-to-sign animation feasible for assistive dialogue rather than offline video production.","Because the Resampling Hook re-synthesizes only the edited window, corrections take about 75 ms and error-recovery time drops from 5.4 s to 2.9 s, suggesting the edit loop can support conversational pacing.","If the user-study results generalize, editable generation increases comprehension by 28% and perceived trust by 34% over automatic output, with a 6.7-point reduction in NASA-TLX cognitive load.","Accumulated JSON diffs and ratings can be converted into continual fine-tuning data, allowing the model to adapt to individual signers and terminology without full retraining.","On constrained hardware the pruned INT8 system maintains 13-24 FPS on typical notebook CPUs, extending deployment beyond dedicated GPUs."],"supporting_citations":[{"why":"Supplies the mixture-density-network formulation used by the decoder and the multimodal sampling scheme the Resampling Hook re-runs.","marker":"Saunders, Camgoz, and Bowden 2021"},{"why":"Provides the WLASL dataset whose WLASL100 split is the benchmark and fine-tuning vocabulary.","marker":"Li et al. 2020"},{"why":"SignVQNet, the discrete-token baseline whose accuracy and latency are compared under the same runtime.","marker":"Hwang, Lee, and Park 2024"},{"why":"Fast-SLP, the non-autoregressive baseline that sets the speed and quality bar for the comparison.","marker":"Huang et al. 2021"},{"why":"SignDiff, the diffusion baseline used to benchmark realism and the 3x latency penalty.","marker":"Fang et al. 2025"},{"why":"Supplies the PPO-style objective and KL-regularization used in the human-in-the-loop fine-tuning loop.","marker":"Schulman et al. 2017"},{"why":"Articulates the human-centered AI principles of transparency, controllability, and trust that motivate the editable JSON layer.","marker":"Shneiderman 2022"},{"why":"Establishes participatory avatar design with Deaf users, the precedent for the co-creation workshops that shaped the JSON schema.","marker":"Dimou et al. 2022"}],"fun_headline_variants":["Editable sign-language AI lifts comprehension 28%","User-edited sign language: 13-point SUS gain, 6.7 lower load","Real-time speech-to-sign with JSON editing: 103 ms","Human-in-the-loop sign language: +28% comprehension, −46% error time","Sign-language AI adapts to user edits, boosting trust 34%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the user-study assumption that the Edit-mode gains were caused by the editor and Resampling Hook rather than by the fixed Auto-block-before-Edit-block ordering described in the paper's Appendix C; the Latin-square there rotates task types within each block but does not counterbalance which block comes first.","fun_headline_variants_meta":{"raw":{"variants":["Editable sign-language AI lifts comprehension 28%","User-edited sign language: 13-point SUS gain, 6.7 lower load","Real-time speech-to-sign with JSON editing: 103 ms","Human-in-the-loop sign language: +28% comprehension, −46% error time","Sign-language AI adapts to user edits, boosting trust 34%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000598,"raw_usage":{"total_tokens":2859,"prompt_tokens":1073,"completion_tokens":1786,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":689,"completion_tokens_details":{"reasoning_tokens":1689}},"tokens_in":689,"tokens_out":1786,"duration_ms":12752,"temperature":1.0,"reasoning_tokens":1689,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:50:14.312624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled replication with participant order counterbalanced—half doing Edit before Auto, with comparable tasks and blind scoring—would settle it: if comprehension, SUS, trust, and cognitive-load differences between modes shrink to non-significance when Auto is second, then practice or learning effects, not the editing mechanism, produced the reported gains.","supporting_citations":[],"review_version":1}