{"id":"37f09dd6-8d73-447f-972c-6ddfaea4b621","arxiv_id":"2509.09889","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Pepper can produce a subset of Italian Sign Language signs that LIS users recognize at the single-sign level, but sentence-level comprehension largely fails (about 8 percent correct).","lead":"Researchers programmed 52 Italian Sign Language (LIS) signs onto the Pepper robot with help from a Deaf student and his interpreter. In a 12-participant online study, most isolated signs were recognized, but full signed sentences were usually misunderstood, pointing to the robot's physical limits.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'substantial number' conclusion rests on a curated 15-sign test set selected by the co-designers; the unselected 37 signs are never evaluated, so the main generalization is unsupported.","rationale":"The reader's conditional verdict already identifies the main weak spot: the hand-picked sign subset and the overlap between those who chose signs and those who designed them. My reading confirms this is the load-bearing threat. The paper is honest about failures, publishes videos, and reports per-sign counts, which is genuine evidence that at least some LIS signs are intelligible on Pepper. The aggregate 119/180 recognition rate and the very small p-values for several individual signs make pure chance an implausible explanation. What is not secured is the jump from 'the 15 selected signs were mostly recognized' to 'a substantial number of LIS signs are intelligible.' That inference requires that the selected signs represent the 52 implemented signs or feasible LIS signs generally, and the design undermines that assumption. The sentence-level results are also weakened by the acknowledged omission of spatial accord, but they are not the main support for the sign-level claim. Because the concern is about generalizability rather than internal validity, CONDITIONAL remains the appropriate verdict; the proposed full-set or random-sample evaluation would settle whether the concern actually lands.","tokens_in":10616,"tokens_out":20711,"duration_ms":232792,"concrete_test":"Run the same forced-choice protocol on all 52 implemented signs, or on a pre-registered random sample stratified by one-handed/two-handed and semantic category, with a fresh panel of LIS signers who had no role in sign design. Pre-register an intelligibility criterion (e.g., recognition rate significantly above 25% after Holm correction and ≥75% correct). If the proportion of intelligible signs in the full/random set is materially lower than the 10/15 observed on the curated set, the Section 5 'substantial number' claim should be revised to enumerate the specific intelligible signs instead of generalizing to the implemented vocabulary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states that the 15 questionnaire signs were chosen 'with the help of the LIS interpreter and the Deaf student' — the same people who co-designed the 52 signs and, according to the Acknowledgments, discarded signs that Pepper did not perform correctly. The tested set is therefore a curated sample of likely-good signs, not a random or representative sample of the 52 implemented signs or of feasible LIS signs. Table 1 shows 10/15 signs reaching nominal significance in uncorrected one-tailed binomial tests; after a multiple-comparison correction the per-sign count drops to about 6/15, though the aggregate 119/180 correct remains far above chance. The data support 'some sign stimuli produced by Pepper are intelligible,' but Section 5's 'substantial number of LIS signs with a good level of intelligibility' requires evidence about the untested 37 signs. The paper's own limitation paragraph addresses only sample size, not this selection channel; the Section 4 claim that omitting spatial accord does not affect findings is also asserted without support and confounds the sentence-level interpretation, but the sign-level claim is the central one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether the Pepper humanoid robot can produce intelligible isolated signs and short sentences of Italian Sign Language (LIS). The authors co-designed 52 signs with a Deaf student and his LIS interpreter, using both manual animation and a semi-automated inverse-kinematics pipeline, and then ran an online questionnaire with 12 LIS-proficient participants: 15 four-alternative forced-choice sign recognition items and 2 open-ended sentence interpretation items. They report 119/180 correct sign choices overall, 10/15 signs significant at the nominal 5% level, and much lower sentence-level comprehension. The paper concludes that Pepper can perform a substantial number of LIS signs with good intelligibility, while acknowledging technical limitations and the exploratory nature of the study.","tokens_in":10919,"tokens_out":5886,"duration_ms":67731,"significance":"If the claims are accepted, the paper provides a useful data point for inclusive HRI: a widely available commercial robot can be made to produce a non-trivial subset of signs that LIS users can recognize, with honest reporting of per-sign failures (e.g., Profumo 0%, Insegnare 33%). The participatory co-design and public availability of sign videos are strengths, as is the reported time saving of the IK toolchain. However, the main generalization to 'a substantial number of LIS signs' goes beyond the evaluated 15-sign subset, and the sentence-level result is confounded by the deliberate omission of spatial accord. The contribution is therefore best framed as a feasibility study of a curated subset, not as an established capability for general LIS communication.","major_comments":[{"comment":"The 15 questionnaire signs were chosen with the help of the LIS interpreter and the Deaf student — the same individuals who co-designed the 52 signs and, per the Acknowledgments, discarded signs that Pepper did not perform correctly. The remaining 37 implemented signs were never evaluated. Thus the aggregate 119/180 and the per-sign p-values support intelligibility of this curated, likely easy-to-animate subset, not the Section 5 claim that Pepper 'is capable of performing a substantial number of LIS signs with a good level of intelligibility.' In addition, the 15 binomial tests are uncorrected; with a Bonferroni threshold (0.05/15 ≈ 0.0033) only 6 of 15 signs remain individually significant. Please either test a representative/random sample of the 52 signs, or substantially qualify the conclusion to the evaluated subset; report multiplicity-corrected results. The limitation paragraph in","section":"Section 4.2, Table 1, Section 5"},{"comment":"The feasibility claim for the automated pipeline (38/52 signs) relies on the assertion that the elbow-movement difference between manual and automated signs 'had no impact on the recognition ability of the Deaf student and his sign language interpreter.' This is an anecdotal assessment by the two co-designers, not a measurement from the 12-participant study. Table 1 does not report which evaluated signs were manual vs. automated, so the user study cannot confirm equivalence. Please provide the manual/automated breakdown of the 15 evaluated signs, or temper the claim to anecdotal evidence from the co-design phase.","section":"Section 3.2"},{"comment":"The sentence stimuli omit LIS spatial accord by design, and the authors assert this simplification 'so not impact the findings' without supporting evidence. Spatial accord is a grammatical device for marking syntactic roles; removing it changes the linguistic stimulus and may directly lower sentence comprehension. Consequently, the Section 5 conclusion that Pepper has limited capacity to convey temporal and syntactic structure is not supported: the low sentence recognition (2/24 correct) could reflect the missing grammatical device, lexical confusions (e.g., mela/acqua), or timing, rather than Pepper's embodiment per se. A control condition with correctly spatially-marked sentences, or a human-signing baseline, is needed before attributing the deficit to the robot.","section":"Section 4, open-ended questions"}],"minor_comments":[{"comment":"There are numerous OCR-style typos and formatting artifacts (e.g., 'Communica7on', 'Assis3ve', '0me', 'An.pa.co', 'Fa_o'); please proofread the final manuscript carefully.","section":"Throughout"},{"comment":"Grammar: 'a exploratory user study' should be 'an exploratory user study'; 'Results shows' should be 'Results show'.","section":"Abstract, Section 5"},{"comment":"The column header 'Binomial value p-' is unclear; also the note for Spaventarsi ('None of the above answer, 50% selected it correctly') is ambiguous and should be rewritten.","section":"Table 1"},{"comment":"In the description of the Andare sign, 'the leg arm remains stationary' appears to be a typo for 'the left arm remains stationary'.","section":"Section 3.1"},{"comment":"Clarify why the 10 'asymmetric or body-involving' signs could not be reproduced by the pipeline — is this due to the single-arm URDF model? This is relevant for assessing the generality of the toolchain.","section":"Section 3.2"},{"comment":"The GitHub repository currently contains videos of signs. Consider also releasing the qianim animation files and the IK exporter script to support reproducibility and reuse.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The honest per-sign reporting is commendable; the main risk is overgeneralization from a curated 15-sign subset to the full 52-sign set. With the conclusions reframed and the requested analyses (representativeness check or qualification, multiplicity correction, manual/automated breakdown, and a spatial-accord control or explicit confound acknowledgment), the paper would be publishable as an exploratory feasibility study. The sampling issue is an under-specified protocol rather than a sign of misconduct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this is a legitimate, honestly reported exploratory study of whether a commercial Pepper can produce intelligible isolated LIS signs. The novel bit is the combination—LIS specifically, Pepper specifically, a co-design process with a Deaf student and interpreter, plus a small MATLAB-IK toolchain that auto-generates admissible animations. The sign-level data are real: 10/15 signs above chance in uncorrected one-tailed binomial tests, aggregate 119/180 correct, and the paper reports failures like Profumo 0% and Insegnare 33% without spin. The sentence-level results are poor, and the paper says so. That honesty counts.\n\nThe soft spots are real but proportionate. The central claim—\"substantial number of LIS signs with good intelligibility\"—rests on a curated 15-sign subset chosen with the same co-designers who built and filtered the 52 signs, and the untested 37 signs are never evaluated. So the data support \"some signs are intelligible,\" not a broad generalization about feasible LIS on Pepper. The multiple-comparison issue is real: after correction, roughly 6/15 signs remain individually significant, though the aggregate still looks far above chance. The spatial-accord omission in the sentence test is asserted as harmless without support, and it confounds the sentence-level interpretation; the sign-level claim is the load-bearing one and it mostly holds. Also, the sentence \"elbow differences had no impact on recognition\" was checked by the same two co-designers, not by the 12 participants—mild circularity, not fatal.\n\nThe reader's stress-test note is largely right: the selection channel is the strongest reason to cap the generalization. But I would not call this a load-bearing flaw, because the paper repeatedly frames the study as exploratory and limits its own ambition. The abstract says \"subset of LIS signs,\" not \"LIS.\" The methodology section is careful about Pepper's constraints. So the main conclusion, read charitably, is defensible.\n\nWho gets value from this: people working on assistive HRI, robot sign production, and participatory design with Deaf communities. It extends prior robot-signing work (NAO, RASA, TEO) to LIS/Pepper and adds a modest engineering contribution. It is not a paradigm shift, and it does not resolve an open scientific question, but it is a solid stepping stone with reproducible materials (videos on GitHub).\n\nMy recommendation: send it to peer review. It deserves referee time, especially for the HRI community where this kind of application-oriented, honestly-reported empirical work is useful. The authors should be asked to address the sign-selection bias, correct for multiple comparisons or soften the per-sign claims, and either justify or remove the spatial-accord assertion. But these are revision issues, not rejection issues.\n\nBest.","headline":"Honest limited empirical step for LIS on Pepper, with a clear selection caveat that caps the generalization but does not sink the modest claim.","tokens_in":11381,"tokens_out":689,"would_cite":true,"duration_ms":10507,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A standard Pepper robot can produce intelligible Italian Sign Language signs for a limited, co-designed vocabulary, but sentence-level signing remains beyond it.","keywords":["social robots","Italian Sign Language","LIS","Pepper robot","human-robot interaction","sign language recognition","inverse kinematics","accessibility"],"falsifier":"Run the same recognition questionnaire on a set of 15 signs drawn at random from the 52 (or from the dictionary) without co-designer filtering, and also readminister the sentence test with proper spatial accord added. If random-sign recognition falls to chance, or if accord-full sentences remain near-zero, the paper's 'good intelligibility' and 'sentences fail' claims have to be restated.","tokens_in":10514,"feed_emoji":"🤖","tokens_out":4042,"duration_ms":42977,"temperature":0.7,"pith_summary":"This paper investigates whether a commercially available social robot, Pepper, can produce intelligible Italian Sign Language (LIS) signs. Working with a Deaf student and his interpreter, the authors co-designed and implemented 52 signs, then asked 12 proficient LIS users to recognize 15 isolated signs and 2 short signed sentences. The majority of isolated signs were recognized correctly—10 of the 15 significantly above the 25% chance level—while full-sentence recognition was very poor, with only 2 of 24 responses correct. The paper's central claim is that even a mass-produced robot can convey a limited LIS vocabulary intelligibly, but cannot yet support sentence-level signed communication. The authors argue this opens a path toward more inclusive human-robot interaction in public settings, provided multimodal support and participatory design are added.","feed_headline":"Pepper robot signs a subset of LIS clearly—sentences still fail","feed_subtitle":"Deaf and hearing LIS users recognized most isolated signs from a commercial robot, yet full sentences were nearly unintelligible.","key_machinery":"The central object is the Pepper humanoid's constrained embodiment—its five fingers move only together, its wrist rotation and elbow mobility are limited, and a chest-mounted tablet blocks torso-close gestures—together with the two sign-authoring pipelines used: a manual keyframe Animation Editor and a MATLAB numerical inverse-kinematics solver that converts 3D hand trajectories into joint values and .qianim animation files. The co-design loop with a Deaf student and interpreter is what filters the 52 signs: each sign is implemented, shown, and iteratively revised until the Deaf collaborators accept it, so the sign set is shaped by both linguistic validity and Pepper's physical affordances.","core_discovery":"On the paper's own terms, the discovery is that a standard Pepper robot, whose fingers can only open and close as a group, can nonetheless perform a substantial subset of LIS signs with good intelligibility. In a recognition test with 12 Deaf and hearing LIS users, signs such as 'forget,' 'done,' 'shampoo,' and 'university' were recognized perfectly, raising the question of which articulatory features are truly load-bearing for sign recognition. At the same time, combining signs into short sentences dropped comprehension near chance, showing that lexical intelligibility does not by itself scale to utterance-level communication. The authors also show that an inverse-kinematics pipeline can ge","pith_inferences":["Because the 15 tested signs were selected with the Deaf co-designers for Pepper-compatibility, the reported 83% isolated-sign accuracy likely overstates how well Pepper would do on a random LIS sample; a random-sample test would separate selection effects from genuine intelligibility.","The near-zero sentence performance may be largely attributable to the deliberate omission of spatial accord, which assigns grammatical roles; restoring it would likely lower sentence recognition still further, meaning the robot's communicative ceiling is below what the paper's favorable framing implies.","The success of the IK pipeline suggests that motion-capture or avatar-generated trajectories could be retargeted to Pepper at scale, but only for the open/closed finger subset; robots with independently articulable fingers would be the natural next step to test whether sentence-level signing becomes feasible.","If a robot is to serve Deaf users in public spaces, the results imply that its signing channel should be considered a fallback or supplement to the screen, not a standalone language channel—a design constraint that argues for co-locating sign output with visual text."],"forward_implications":["A commercial social robot can act as a partial LIS communicator for isolated words in domains like schools, museums, or information points.","Sentence-level signed interaction on this class of hardware is not achievable with arm movements alone; multimodal channels (tablet, subtitles, avatars) will be required.","Automated inverse-kinematics generation makes sign authoring fast enough to expand the tested vocabulary, but only for signs whose handshapes are all-fingers-open or all-closed.","Recognition of iconic, simplified signs is high while articulation-ambiguous signs fail, so robot sign sets must be selected by Deaf users, not by dictionary alone."],"fun_headline_variants":["Pepper robot says LIS words clearly, sentences fall flat","Sign-language test: Pepper passes words, flunks sentences","Pepper's LIS signing: words intelligible, sentences unintelligible","Robot delivers clear LIS signs, but phrase comprehension crashes","Pepper signs LIS: good single words, poor sentence flow"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conclusion rests on the assumption that the 15 signs hand-picked by the Deaf student and interpreter for the test fairly represent the 52 implemented signs and what Pepper can likely sign, and that omitting LIS spatial accord does not change the sentence-difficulty result.","fun_headline_variants_meta":{"raw":{"variants":["Pepper robot says LIS words clearly, sentences fall flat","Sign-language test: Pepper passes words, flunks sentences","Pepper's LIS signing: words intelligible, sentences unintelligible","Robot delivers clear LIS signs, but phrase comprehension crashes","Pepper signs LIS: good single words, poor sentence flow"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1431,"prompt_tokens":787,"completion_tokens":644,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":556}},"tokens_in":531,"tokens_out":644,"duration_ms":7566,"temperature":1.0,"reasoning_tokens":556,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:30:50.865958+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same recognition questionnaire on a set of 15 signs drawn at random from the 52 (or from the dictionary) without co-designer filtering, and also readminister the sentence test with proper spatial accord added. If random-sign recognition falls to chance, or if accord-full sentences remain near-zero, the paper's 'good intelligibility' and 'sentences fail' claims have to be restated.","supporting_citations":[],"review_version":1}