{"id":"294d6f0e-68c7-4836-bc1a-b9878374f745","arxiv_id":"2508.05358","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding user-adjustable settings to a German Sign Language avatar on HoloLens 2 did not improve comprehension or UX for expert signers, despite being preferred, because core animation quality was the limiting factor.","lead":"The paper tested an adjustable sign language avatar on the HoloLens 2 headset with expert German Sign Language users and found that letting users tweak the avatar did not improve comprehension or user experience, which stayed low. It matters because sign language avatars are increasingly used in public services, and this result says design effort should go into basic animation quality before personalization.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The null result may reflect a limited adjustability implementation rather than the inadequacy of personalisation; the abstract does not show that the adjustable parameters included mouthings/facial expressions.","rationale":"The reader's UNVERDICTED verdict is appropriate given abstract-only availability. My concern complements rather than replaces the reader's weakest_assumption: both identify that implementation or operationalization issues could explain the null result. The reader focused on interface defects (hand shapes, feedback, menu position); I focus on whether the adjustable parameter set even targeted the missing SL elements (mouthings, facial expressions). These are distinct but related construct-validity threats. Since the abstract lacks the necessary details, the conclusion cannot be verified or falsified from the abstract alone, so the verdict remains UNVERDICTED. I do not find an internal inconsistency; the study may be well-executed for what it tested, but the breadth of the conclusion exceeds what the abstract supports. The proposed test—checking the adjustable parameter list—would settle whether the concern lands. If the parameters did include mouthing/facial expression controls, my concern would be weakened; if not, the central claim would need to be narrowed. No credit issues: the study appears to be a genuine empirical HCI evaluation, and the abstract honestly reports limitations. However, the absence of quantitative details and the full instrument makes a definitive assessment impossible.","tokens_in":841,"tokens_out":2338,"duration_ms":25234,"concrete_test":"Inspect the full paper's method section to list the adjustable parameters. If mouthing/facial expression controls were not available to users, the central claim is unsupported as stated. Alternatively, check whether the adjustable condition differed only in interface affordances unrelated to linguistic content (e.g., speed, size, angle). If so, recommend a follow-up study with an adjustable avatar that includes non-manual feature controls; a significant improvement there would falsify the 'personalisation alone is insufficient' conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that personalisation alone cannot compensate for missing signing quality. For this to hold, the adjustable condition must have allowed users to actually adjust the elements whose absence drove low comprehensibility—mouthings and facial expressions—or at least plausibly compensate for them. The abstract reports that comprehensibility remained low 'amid missing SL elements (mouthings and facial expressions) and implementation issues.' It does not state what parameters were user-adjustable. If the adjustment set was confined to e.g., avatar size, speed, viewpoint, or menu-related presentation, then the study never tested whether user-controlled customization of core linguistic signal (mouthing, non-manual markers) could improve comprehension. The higher stress and frustration in the adjustable condition also suggest the adjustment UI itself may have imposed cognitive load, potentially masking any benefit. The conclusion 'personalisation alone is insufficient' requires the absence of an improvement in a condition where personalisation was meaningfully possible; without evidence that the adjustable parameters included the deficient channels, the result is compatible with the weaker claim that this particular adjustment UI was ineffective or counterproductive. This is not an internal inconsistency but a construct-validity threat: the operationalization of 'personalisation' may be too narrow for the inferential claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a user study in which expert German Sign Language (DGS) users interacted with a sign language avatar on a Microsoft Hololens 2 device, in both adjustable and non-adjustable versions. The abstract claims that, despite user preference for adjustable settings, no significant improvements in comprehensibility or user experience were observed; comprehensibility remained low, with missing mouthings and facial expressions and with implementation issues such as indistinct hand shapes, lack of feedback, and menu positioning. Stress was higher in the adjustable condition, and hedonic quality exceeded pragmatic quality. The authors conclude that personalisation alone is insufficient and that sign language avatars must be comprehensible by default, recommending improvements to mouthing/facial animation, interfaces, and participatory design.","tokens_in":1088,"tokens_out":3063,"duration_ms":34850,"significance":"If the finding is robust, it provides a valuable negative result for sign language avatar design: adding user customization is not a substitute for baseline signing quality. The abstract is appropriately hedged and transparently acknowledges implementation defects, which is a strength. The practical recommendations follow from the reported data. However, the abstract alone does not carry the inferential load of the central conclusion, because neither the adjustable parameters nor the statistical support are described. The result would be a useful contribution if the full manuscript supplies the missing operationalization and data.","major_comments":[{"comment":"The central conclusion that 'personalisation alone is insufficient' depends on the adjustable condition giving users meaningful control over the channels reported as deficient: mouthings and facial expressions, which the abstract identifies as missing SL elements. The abstract does not state which parameters were user-adjustable. If the adjustable set was limited to, e.g., menu-based presentation or viewpoint, the study did not actually test whether user control over core linguistic signal can improve comprehension. The result would then concern this particular adjustment interface, not personalisation as a general design concept. Please specify the adjustable parameters and justify that they constitute the kind of personalisation the conclusion targets.","section":"Abstract, findings"},{"comment":"The abstract reports a null result for UX and comprehensibility without providing sample size, p-values, effect sizes, or confidence intervals. A claim of 'no significant improvement' requires adequate statistical power or at least a quantitative comparison of effect magnitudes. Without this information, the reader cannot distinguish a true absence of effect from an underpowered study. Please include the number of participants and the relevant inferential statistics in the abstract or, if this is an abstract-only artifact of the review, ensure they are prominent in the reported results.","section":"Abstract, 'no significant improvements observed'"},{"comment":"The adjustable condition produced higher stress, described as lower performance, greater effort, and more frustration. The authors list implementation issues (lack of feedback, menu positioning) that could themselves cause these outcomes. This is a serious construct-validity threat: the negative result may reflect the usability of the adjustment interface rather than the ineffectiveness of personalisation. To support the loading-bearing generalization, the paper must argue or demonstrate that the observed null effect is not attributable to the specific UI confounds. Please discuss how the design or analysis separates the adjustment concept from the implementation.","section":"Abstract, stress and implementation issues"}],"minor_comments":[{"comment":"Please define 'adjustment features' or 'adjustable settings' explicitly; the current wording leaves unclear what users were able to change.","section":"Abstract"},{"comment":"The phrase 'no significant improvements' should be accompanied by the significance criterion and directional information; otherwise it is easily misread as evidence of equivalence.","section":"Abstract"},{"comment":"The terms 'personalisation' and 'adjustability' are used interchangeably; consider consistent terminology.","section":"Abstract"},{"comment":"The statement 'user preference for adjustable settings' is qualitative; specify how preference was measured and whether it was statistically assessed.","section":"Abstract"},{"comment":"A brief note on generalizability beyond DGS and the specific use case would help readers calibrate the scope of the recommendations.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"This is an abstract-only review, so the main technical verdict rests on what the abstract does and does not report. The most important concern is construct validity: if the full manuscript already specifies the adjustable parameters and provides complete inferential statistics, the major points may be resolvable by reporting the missing details. If the adjustable parameters did not include the deficient linguistic channels, the central conclusion needs to be reframed. The authors' explicit acknowledgement of implementation defects is commendable, but it also underscores the confound. I would not reject the paper—the negative result is potentially useful—but it needs stronger operational disclosure and statistical transparency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: you can send this to peer review. The authors ran an honest, small-scale user study on HoloLens 2 comparing adjustable vs non-adjustable German Sign Language avatars, with expert signers. Their headline result—that adjustability didn't improve comprehensibility or UX, and stress went up—is a useful null result that challenges the 'just let users tweak it' reflex in sign language avatar design. They also openly list implementation defects, which is more than many papers do.\n\nWhat's genuinely new: a controlled comparison on an AR headset with expert users, and a specific negative outcome for personalisation. That's a real data point for a specialized community.\n\nWhat's soft: the abstract doesn't tell us what parameters the adjustable condition actually controlled. The conclusion 'personalisation alone is insufficient' requires that the adjustable version gave users a chance to modify the missing linguistic content—mouthings and facial expressions. If the adjustment menu only covered size, speed, or viewpoint, then the study tested a narrow version of personalisation. The higher stress and frustration in the adjustable condition also suggest the UI itself may have masked any benefit. So the construct validity of 'personalisation' is the load-bearing assumption, and it isn't visible in the abstract. That doesn't kill the paper, but it limits how far the general conclusion can travel.\n\nAlso, the abstract has no numbers: no sample size, no effect sizes, no test statistics. That's standard for an abstract, but it means we can't evaluate the null from what we have.\n\nThe self-evaluation angle is present but the negative result for the home-built adjustable feature actually cuts against self-promotion, so I'm not worried there.\n\nWho this is for: people building sign language avatars, especially in accessibility and HCI. They'll get practical guidance on where to spend effort (core signing quality before personalisation). A broader NLP audience probably won't need to read it.\n\nBottom line: I'd let it through to referees. The abstract shows a real study with an honest negative result; the full text likely contains the parameter details and statistical analysis that would settle whether the conclusion is warranted. If I were reviewing, I'd ask the authors to make explicit what parameters users could adjust, and to soften the general claim if the adjustable set excluded mouthing and facial expressions.","headline":"Honest null result on avatar adjustability, but the abstract doesn't prove the adjustable condition actually let users fix the missing linguistic content—so the headline conclusion is one step ahead of the evidence.","tokens_in":1565,"tokens_out":2116,"would_cite":false,"duration_ms":21342,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that letting users adjust a German Sign Language avatar on a HoloLens 2 does not improve comprehensibility or user experience, because the avatar's baseline signing quality—especially mouthings and facial expressions—is wh","keywords":["sign language avatar","German Sign Language","comprehensibility","user experience","personalisation","HoloLens 2","mixed reality","accessibility"],"falsifier":"A controlled study with an adjustable sign-language avatar whose interface fixes the identified implementation problems (clear hand shapes, proper feedback, well-placed menu) and whose baseline animation quality is varied, showing that adjustment yields significant comprehension gains at higher baseline quality, would directly contradict the paper's claim that personalisation alone is insufficient.","tokens_in":764,"feed_emoji":"🤟","tokens_out":3488,"duration_ms":33956,"temperature":0.7,"pith_summary":"The paper asks whether giving users control over a sign language avatar's settings can compensate for weaknesses in how the avatar signs. Working with expert German Sign Language users on a mixed-reality headset, it finds that an adjustable version was preferred in principle but produced no measurable gain in comprehensibility or user experience. The authors conclude that personalisation is not a substitute for making the avatar comprehensible by default, and that missing elements such as mouthings, facial expressions, and distinct hand shapes determine whether signing is understood. This matters for anyone building sign-language technology because it redirects effort from customization features toward core animation quality and participatory design.","feed_headline":"Adjustability won't fix sign avatar comprehension","feed_subtitle":"Users liked the idea of adjustability, but comprehension stayed low without proper mouthings and facial expressions.","key_machinery":"The central object is the adjustable sign-language avatar: a baseline avatar plus a set of user-modifiable parameters presented on a HoloLens 2 headset. The experimental machinery is a user study with expert DGS users that measures objective comprehensibility, subjective UX (hedonic vs pragmatic quality), stress, and acceptability, and triangulates these with interaction analysis. The load-bearing comparison is between the adjustable and non-adjustable conditions as a test of whether personalisation can compensate for baseline signing quality.","core_discovery":"With expert DGS users on a Microsoft HoloLens 2, the study compares an existing sign language avatar with and without adjustment features. The headline result is a null result: despite users reporting that they liked having adjustable settings, the adjustable avatar did not significantly improve either UX or comprehensibility, and comprehensibility stayed low for both conditions. Users rated the system's hedonic quality higher than its pragmatic quality, meaning they found it emotionally or aesthetically appealing but not functionally useful. Stress measures were higher with the adjustable version, consistent with greater effort and frustration, and the headset's adjustment gestures were que","pith_inferences":["A cleaner implementation might change the result; the paper's own description of defects suggests the adjustable condition was disadvantaged beyond the concept itself.","The hedonic-pragmatic gap hints that users may like a system they cannot rely on; future work could test whether positive affect decays after real-world use.","The conclusion generalises only if the parameters users were allowed to adjust cover the dimensions they actually need; a wider or more meaningful parameter set might yield different outcomes.","A testable extension would be to vary baseline animation quality as an independent factor alongside adjustability, mapping their interaction directly."],"forward_implications":["Designers of SL avatars should treat comprehensibility as a default requirement, not something personalisation can fix later.","User preference for adjustable settings does not by itself predict better comprehension or UX; objective performance measures matter.","Missing mouthings and facial expressions appear to be a principal bottleneck for DGS avatar comprehension.","Adding adjustment features without fixing baseline animation risks increasing user effort and stress.","Acceptability of adjustability is conditional on usability and animation quality, so feature work should follow core quality work."],"supporting_citations":[],"fun_headline_variants":["Sign avatar adjustability fails to boost comprehension","Adjustable sign avatars: liked but not understood","Sign avatar tweaks don't fix comprehension gap","Sign avatar users like options, but clarity stays low","Adjustable sign avatars: more stress, no comprehension gain"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The conclusion that personalisation is insufficient rests on the assumption that the implementation flaws in the adjustable interface—indistinct hand shapes, lack of feedback, awkward menu placement—did not themselves cause the null result; a cleaner adjustable interface might have been a fairer test.","fun_headline_variants_meta":{"raw":{"variants":["Sign avatar adjustability fails to boost comprehension","Adjustable sign avatars: liked but not understood","Sign avatar tweaks don't fix comprehension gap","Sign avatar users like options, but clarity stays low","Adjustable sign avatars: more stress, no comprehension gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000552,"raw_usage":{"total_tokens":2469,"prompt_tokens":745,"completion_tokens":1724,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":1649}},"tokens_in":489,"tokens_out":1724,"duration_ms":12093,"temperature":1.0,"reasoning_tokens":1649,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:22:54.919436+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study with an adjustable sign-language avatar whose interface fixes the identified implementation problems (clear hand shapes, proper feedback, well-placed menu) and whose baseline animation quality is varied, showing that adjustment yields significant comprehension gains at higher baseline quality, would directly contradict the paper's claim that personalisation alone is insufficient.","supporting_citations":[],"review_version":1}