{"id":"be2b1044-9b33-43e8-ab7f-58a992307e0c","arxiv_id":"2608.08200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Personalized, context-dependent motion scaling improves delayed telemanipulation performance in simulation and transfers to a real surgical robot with reduced benefit.","lead":"A 20-person study finds that motion scaling gains personalized to each user's delay, direction, and reach distance reduce overshoot and improve smoothness in delayed telemanipulation. The gains transfer from simulation to a real surgical robot, though benefits shrink.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Order confound in unassisted vs. assisted comparison: all unassisted delayed trials precede all assisted trials, so practice effects may drive the reported 20-25% gains.","rationale":"The reader's verdict correctly identifies the need for artifacts and alignment of claims. The most load-bearing threat to the central claim is not the stability of the fitted gain per se but the experimental architecture that makes the unassisted-to-assisted comparison a between-block, always-unassisted-first comparison. In SimOnly, Phase 2 (unassisted delayed) always precedes Phase 4 (assisted delayed); the baseline phase in between does not provide a delayed control. In Sim2Real, the control block always precedes the two evaluation blocks. This order confound means practice, delay adaptation, or fatigue can masquerade as an assistance benefit. The 20-25% gains and the growing benefit with delay could reflect a time-on-task trend. The paper states that catch trials are embedded in both phases, so the data to estimate learning effects exist but are not analyzed. A simple re-analysis comparing no-delay catch trials between phases would test this. If no-delay performance improves, the effect sizes must be adjusted. This is a fixable issue, hence the verdict remains conditional, but it is more fundamental than the gain-transfer assumption because it affects the primary causal claim. We therefore recommend the reader's conditional verdict be retained, with the added condition that the authors re-analyze the catch trials and report order-adjusted estimates.","tokens_in":14509,"tokens_out":8158,"duration_ms":74815,"concrete_test":"Re-analyze the embedded no-delay catch trials from SimOnly Phase 2 (Gain ID) and Phase 4 (Evaluation) for each participant and condition, and similarly from the Sim2Real control and evaluation blocks if catch trials exist. Fit the same linear mixed model used for the main metrics (Section III-I) with an added phase or trial-index covariate on overshoot, endpoint error, and nCET. If no-delay performance significantly improves from Phase 2 to Phase 4, the unassisted-to-assisted comparison is confounded with learning; re-estimate the assistance effect including trial index as a covariate to determine whether the reported 20-25% gains survive. Report the learning-corrected effect sizes and p-values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Both experimental sessions confound assistance with trial order. In SimOnly (Section III-E), Phase 2 (Gain ID) collects all unassisted delayed trials, Phase 3 collects a no-delay baseline, and Phase 4 (Evaluation) collects assisted delayed trials. In Sim2Real, the control block (Phase 3) always precedes the two evaluation blocks (Phases 4 & 5). Thus the unassisted and assisted delayed conditions are not interleaved or counterbalanced; any monotonic change in performance over the session (learning, adaptation to delay, fatigue, strategic shifts) is attributed to motion scaling. The no-delay catch trials embedded in both phases (by the stated 'same structure' design) could provide a practice control, but they are not analyzed. Even if the gain fitted by Eq. 1 is perfectly stable and transfers, the outcome comparison remains biased by this design. This threatens the central claim of consistent assistance benefits and the delay-dependent increase in benefit (e.g., Table I Asst:Delay interactions). The reader's concern about overshoot stability is secondary: the primary comparison cannot distinguish assistance from time.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a personalized, context-dependent motion-scaling method for delayed telemanipulation. In a simulator session, each participant's scaling gain is fitted as a ratio of baseline and delayed peak reach distances for each combination of delay, distance, and direction (Eq. 1); the fitted gains are then evaluated in delayed reaching trials in simulation and in a physical dVRK peg-transfer task, with a generalized population gain used as an additional comparator in the sim-to-real session. The authors report consistent benefits of motion scaling on overshoot, smoothness, movement economy, and a combined error-time metric, with larger benefits at longer delays, and more limited or even negative effects on endpoint error in the real task. Personalization showed specific benefits for inward reaching and for nCET at 400 ms delay compared with the generic gain. The manuscript includes a detailed experimental protocol, linear mixed-effects analyses, and honest reporting of conditions where assistance did not help.","tokens_in":14739,"tokens_out":6483,"duration_ms":68264,"significance":"If the central comparison is valid, the paper makes a useful contribution: it extends prior single-DOF and fixed-gain motion-scaling results to a multi-DOF, context-dependent personalization framework and provides a sim-to-real transfer evaluation on a physical surgical robot. Strengths include the within-subject design across delays, distances, and directions; the use of multiple performance metrics; the transparent statistical modeling; and the explicit discussion of conditions where personalization or assistance did not help. The reported simulation benefits, if supported, would be practically relevant for delay-mitigating assistance layers. However, the manuscript currently provides no data or code, and the main assistance-versus-unassisted comparison is compromised by a systematic order confound; the latter is load-bearing for the paper's central claim.","major_comments":[{"comment":"The unassisted and assisted delayed trials are not interleaved or counterbalanced in either session. In SimOnly, all unassisted delayed trials (Phase 2, Gain ID) precede all assisted delayed trials (Phase 4, Evaluation). In Sim2Real, the unassisted control block (Phase 3) always precedes the two assisted evaluation blocks (Phases 4 and 5), with only the order of the two assisted blocks counterbalanced. Any monotonic practice, fatigue, or delay-adaptation trend over the session is therefore attributed to assistance. The no-delay catch trials embedded in both phases could provide a practice control, but they are not analyzed. This confound threatens the main assistance main effects and the Asst:Delay interactions in Table I, as well as the abstract's central claim of consistent 20–25% performance gains.","section":"§III-E, §IV (Tables I and reported contrasts)"},{"comment":"The smoothness metric sign convention is internally inconsistent. Section III-G.3 states that lower SAL values indicate more erratic or segmented motion and higher values correspond to fluid movements, but Section IV-A.3 reports that personalized assistance 'reduced SAL values' while claiming increased smoothness. Under the standard spectral arc length definition (which is negative, with higher values indicating smoother movement), these statements cannot both be true. Please state the exact sign convention, verify whether the contrasts are increases or decreases in SAL, and re-report the smoothness results and any associated table entries if the sign was reversed.","section":"§III-G.3 and §IV-A.3"},{"comment":"The design is described as having three reaching directions (lateral, longitudinal, vertical), but the statistical model uses a two-level factor classified as inward (adduction) versus outward (abduction). The manuscript does not explain how the vertical direction and the two horizontal directions map onto this binary factor, or whether vertical trials were excluded from the analysis. Please specify the mapping or report a three-level/direction-specific analysis; otherwise the Direction effects in Table I and the direction-related contrasts are ambiguous.","section":"§III-C and §III-I"},{"comment":"Equation (1) defines the gain as the ratio of mean peak reach distances in baseline and delayed trials, not as a ratio of overshoots. The text and abstract state that scaling gains were computed to minimize mean overshoot, but peak reach distance equals target distance plus overshoot along the movement axis. The ratio of peak distances is therefore not equivalent to an overshoot-minimizing gain and will be biased toward unity when baseline overshoot is nonzero. Please define the objective precisely and justify the peak-distance ratio, or compute the gain from the overshoot values defined in Eq. (2).","section":"§III-F, Eq. (1)"},{"comment":"The abstract's claim that 'Motion scaling consistently improved performance relative to unassisted trials' is stronger than the reported Sim2Real endpoint-error results. Section IV-B.2 states that assistance did not produce a consistent endpoint-error reduction and that the only significant contrast was a decrease in accuracy due to assistance at 250 ms delay. Please qualify the abstract so that it reflects this important exception, for example by specifying that improvements were consistent across most metrics but not endpoint error in the transfer task.","section":"§IV-B.2 and Abstract"}],"minor_comments":[{"comment":"The sentence 'participants experienced personalized compensation gains for each non-delayed trial condition' appears to say 'delayed' rather than 'non-delayed'; the evaluation block applies compensation to delayed trials.","section":"§III-E, Phase 4"},{"comment":"The manuscript does not mention data or code availability. Given the scale of the dataset and the importance of the human-factors results, providing processed data or analysis code would materially improve reproducibility.","section":"General"},{"comment":"The same reference appears twice: Richter, Orosco, and Yip, ICRA 2019, is listed as both [9] and [30]. Please merge the duplicate.","section":"References"},{"comment":"Equation (1) defines gains over (δ, d, θ), but the GAM in Section III-J uses delay as the sole predictor and excludes distance. Please clarify whether the evaluation used the discrete three-dimensional gain array or the direction-only GAM, and whether the reported personalization effects are based on the distance-conditioned gains or the reduced model.","section":"§III-J"},{"comment":"The effective target width W in the Fitts' law index of difficulty is not defined for either the shape-matching or peg-transfer task. Please specify how W was measured or set, as this affects the nCET normalization.","section":"§III-G.5, Eq. (7)"},{"comment":"The term 'gain tensor' is acknowledged as informal; consider using 'gain array' throughout to avoid confusion with tensor algebra in a robotics readership.","section":"§III-F"}],"recommendation":"major_revision","confidential_remarks":"The order confound in the assistance-versus-unassisted comparison is the main substantive risk. If the authors can reanalyze the embedded catch trials as a practice-control, or provide additional counterbalanced data, the central claim could become defensible. The SAL sign inconsistency and the direction-factor ambiguity are also fixable but should be corrected before publication. I would encourage the authors to share de-identified processed data, as the study is large and the field would benefit from reanalysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. This is a careful, well-written study that systematically fits per-participant motion-scaling gains as a function of delay, distance, and direction, and then tests them in simulation and on a real dVRK. That conditioning is genuinely new relative to Nisky and Richter/Lim; it is an extension of known motion-scaling work, but the mapping itself and the sim2real transfer test are useful. The authors also report where personalization did not help, which is honest.\n\nThe soft spot is the design. In both sessions, all unassisted delayed trials come before all assisted delayed trials. In SimOnly, Phase 2 (gain ID) is unassisted, Phase 4 assisted; in Sim2Real, the control block always precedes the evaluation blocks. That means practice, fatigue, or strategic shifts over the session are aliased with assistance. The embedded no-delay catch trials could have served as a practice control, but they are not analyzed. The stress-test note is correct: this threatens the central claim of consistent 20–25% gains and the delay-dependent increase in benefit. The reader's concern about overshoot stability is more secondary.\n\nThe lack of data and code makes it impossible to check the LMM results independently, and the abstract's phrasing (relative to non-delayed baseline in the intro hypothesis; '20-25% gains' in the abstract) overshoots the actual contrasts.\n\nThat said, the paper deserves a serious referee. The idea is sound, the experiments are substantial, and the order confound can be addressed by reanalyzing the catch trials or collecting a short interleaved control. If the benefit survives that reanalysis, this is a solid contribution to teleoperation. If not, the qualitative conclusions about direction-dependent personalization may still hold. I would not cite it until the data are out, but I would be glad to see it in review.","headline":"A careful, substantial study of context-conditioned motion scaling whose main assistance comparison is confounded by trial order; worth reviewing with data and a reanalysis.","tokens_in":15264,"tokens_out":2110,"would_cite":false,"duration_ms":21755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that personalized motion scaling conditioned on delay, reach distance, and movement direction improves delayed telemanipulation, yielding up to 20-25% performance gains in key metrics, with the largest effects at longer…","keywords":["delayed telemanipulation","motion scaling","personalized assistance","overshoot compensation","telesurgery","haptic teleoperation","sim-to-real transfer"],"falsifier":"Re-run the gain identification block twice on the same participants at the same delays, distances, and directions, with no assistance in either block, and compute the fitted gain each time; if the typical within-person difference between the two fitted gains is larger than the mean personalized-versus-generic gain difference the paper reports, the fitted gains are not stable enough to carry the evaluation.","tokens_in":14298,"feed_emoji":"🎯","tokens_out":10888,"duration_ms":97937,"temperature":0.7,"pith_summary":"The paper is trying to show that the right response to communication delay in teleoperation is not a one-size-fits-all gain but a scaling factor fitted to each operator and to each delay, distance, and direction. Twenty participants reached for targets in a simulator while the system computed, for every condition, the ratio of their no-delay peak reach to their delayed peak reach, and applied that ratio as a motion-scaling gain. Across the tested delays, gains reduced overshoot and improved smoothness, path economy, and a combined error-time score by up to 20-25%, with benefits growing at longer delays; gains fit in simulation also transferred partially to a physical telesurgical robot. If this is right, delay compensation can be initialized from population data and then tuned per person, which matters for telesurgery and other time-critical teleoperation.","feed_headline":"Per-user motion scaling cuts delayed surgery errors by 25%","feed_subtitle":"Gains tuned to each operator's overshoot improve delayed teleoperation and transfer, partly, to a real surgical robot","key_machinery":"The load-bearing object is the 'gain tensor,' a 3-D array of scalar gains indexed by delay $\\delta$, reach distance $d$, and movement direction $\\theta$. Each entry is computed as $G(\\delta,d,\\theta)=\\mathbb{E}_i[P_0^{(i)}(d,\\theta)]\\,/\\,\\mathbb{E}_j[P_\\delta^{(j)}(d,\\theta)]$, where $P_0$ and $P_\\delta$ are peak reach distances in no-delay and delayed trials under the same spatial conditions. This ratio converts a person's delay-induced overshoot into a multiplicative scaling correction, and its dependence on direction and distance is what carries the paper's context-dependence claim. Delay-conditioned gain curves fitted with splines then turn the discrete fitted values into continuous per-direction functions of delay.","core_discovery":"On the paper's own terms, the central claim is that an individually fitted, context-conditioned motion-scaling gain improves delayed teleoperation relative to unassisted operation. The gain is computed as the ratio of the mean peak reach distance in no-delay baseline trials to the mean peak reach distance in delayed trials, separately for each combination of delay, reach distance, and movement direction, and is applied as a multiplicative modifier of the base input scaling. In simulation, personalized assistance reduced initial reaching error at every delay (roughly 0.035 mm at 100 ms, 0.068 mm at 250 ms, and 0.122 mm at 400 ms), improved smoothness and movement economy, and lowered a Fitts-normalized error-time score; the largest gains were at 400 ms. Personalization's extra benefit over a generic gain appeared mainly for inward reaches at short distance under moderate delay. When the same fitted gains were applied to a physical telesurgical robot in a peg-transfer task, smoothness, economy, and error-time improved, but endpoint error did not, indicating partial but not full transfer.","pith_inferences":["Editorial inference: the same peak-overshoot ratio could be re-estimated continuously from recent trials, turning the one-time calibration into an online adaptive gain without changing the cost function.","Editorial inference: the inward-reach advantage suggests arm posture relative to the body modulates the right gain; holding reach distance fixed while varying starting shoulder position would test this directly.","Editorial inference: the weaker sim-to-real benefit may come from the real task's loose success criterion (ring on peg rather than centered), so a real task that scores final centering would separate platform transfer loss from task-objective loss."],"forward_implications":["Personalized motion scaling improves overshoot, smoothness, path economy, and a normalized error-time score relative to unassisted reaching, and the benefit grows with delay; key-metric gains reach 20-25%.","Personalized gains outperform the unassisted baseline at every tested delay in simulation, and they outperform a generic population gain mainly for inward, short-distance reaches under moderate delay.","Gains fitted in simulation transfer to a physical telesurgical robot for smoothness, economy, and error-time, but not for endpoint error, so transfer is real but partial.","Most participants need sub-unity gains that decrease with delay, while some show hypometria at 100 ms and need gains above unity, so a fixed gain cannot capture the range.","Personalized assistance lowers reported workload relative to unassisted operation in simulation and is perceived as lighter than generic assistance in the real-robot session."],"supporting_citations":[{"why":"Provides the state-based delay representation that motivates treating delay as a gain leading to overshoot.","marker":"[5]"},{"why":"Shows sub-unity motion scaling mitigates delay-induced overshoot in needle insertion, the basis for this approach.","marker":"[6]"},{"why":"Reports optimized motion scaling gains and their variation across operators, motivating personalization.","marker":"[7]"},{"why":"Demonstrates motion-scaling benefits in multi-DOF surgical teleoperation under high delay.","marker":"[9]"},{"why":"Supplies the telesurgical robot platform used in both simulation and physical transfer.","marker":"[14]"},{"why":"Provides the peg-transfer task design adapted for this study's reaching tasks.","marker":"[15]"},{"why":"Defines the workload measure used in both sessions.","marker":"[16]"},{"why":"Supplies the spectral arc length metric used to quantify movement smoothness.","marker":"[18]"},{"why":"Source of the combined error-time (CET) performance metric.","marker":"[23]"}],"fun_headline_variants":["Personalized motion scaling improves delayed teleoperation by 25%","Context-aware scaling boosts delayed telerobotics precision","Per-user gains reduce overshoot in delayed teleoperation","Adaptive motion gains enhance delayed surgical teleop","Custom scaling aids delayed telemanipulation, up to 25% better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the delay-induced peak overshoot a person shows in five calibration trials is stable enough to predict their overshoot in later evaluation trials, and that the same fitted gain remains appropriate when the task moves to the physical robot.","fun_headline_variants_meta":{"raw":{"variants":["Personalized motion scaling improves delayed teleoperation by 25%","Context-aware scaling boosts delayed telerobotics precision","Per-user gains reduce overshoot in delayed teleoperation","Adaptive motion gains enhance delayed surgical teleop","Custom scaling aids delayed telemanipulation, up to 25% better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1761,"prompt_tokens":983,"completion_tokens":778,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":696}},"tokens_in":599,"tokens_out":778,"duration_ms":8300,"temperature":1.0,"reasoning_tokens":696,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:16:52.076885+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the gain identification block twice on the same participants at the same delays, distances, and directions, with no assistance in either block, and compute the fitted gain each time; if the typical within-person difference between the two fitted gains is larger than the mean personalized-versus-generic gain difference the paper reports, the fitted gains are not stable enough to carry the evaluation.","supporting_citations":[{"cited_title":"State-based delay representation and its transfer from a game of pong to reaching and tracking,","cited_arxiv_id":null,"evidence_quote":"Provides the state-based delay representation that motivates treating delay as a gain leading to overshoot."},{"cited_title":"Perception and action inteleoperated needle insertion,","cited_arxiv_id":null,"evidence_quote":"Shows sub-unity motion scaling mitigates delay-induced overshoot in needle insertion, the basis for this approach."},{"cited_title":"Optimal Motion Scaling for Delayed Telesurgery","cited_arxiv_id":"2506.21689","evidence_quote":"Reports optimized motion scaling gains and their variation across operators, motivating personalization."},{"cited_title":"An open-source research kit for the da vinci surgical system,","cited_arxiv_id":null,"evidence_quote":"Supplies the telesurgical robot platform used in both simulation and physical transfer."},{"cited_title":"Fundamentals of laparoscopic surgery (fls) and of endoscopic surgery (fes),","cited_arxiv_id":null,"evidence_quote":"Provides the peg-transfer task design adapted for this study's reaching tasks."},{"cited_title":"Development of NASA-TLX (task load index): Results of empirical and theoretical research,","cited_arxiv_id":null,"evidence_quote":"Defines the workload measure used in both sessions."},{"cited_title":"A robust and sensitive metric for quantifying movement smoothness,","cited_arxiv_id":null,"evidence_quote":"Supplies the spectral arc length metric used to quantify movement smoothness."},{"cited_title":"Haptic guidance and haptic error amplification in a virtual surgical robotic training environment,","cited_arxiv_id":null,"evidence_quote":"Source of the combined error-time (CET) performance metric."}],"review_version":1}