{"id":"dc9f874d-8c57-4d2e-bfb6-56593aaa7331","arxiv_id":"2412.00541","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A reservoir-computing model, CESN+, adds prediction-interval confidence to trajectory generation and uses it to adaptively share control, reducing human effort in simulation.","lead":"Researchers add prediction intervals to a lightweight learning-from-demonstration model, letting a robot estimate its own confidence while generating movement trajectories. In a simulated human-robot shared control task, the confidence signal is used to adaptively shift control between human and robot, which the authors report cuts measured human effort.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Adaptive weight condition lacks a matched fixed-weight control; the reported human-effort reduction may reflect more robot authority, not confidence.","rationale":"I read the paper's central claim as a causal statement: the prediction-interval-derived confidence signal is what lowers human effort. The experiment compares adaptive confidence-based omega against a constant 0.5. This confounds the information content of the confidence signal with the average level of robot authority. A matched fixed-weight baseline is the standard way to disentangle these. The reader's concern about PI calibration (Eq. 9) is real, but even a miscalibrated interval could serve as a useful relative confidence if it correlates with error; the missing control is a more direct threat to the claimed benefit. I therefore partially agree with the reader: the PI assumption is a weak spot, but the experimental design is the load-bearing one. The concern is addressable with an additional baseline, so the verdict remains CONDITIONAL.","tokens_in":10590,"tokens_out":3312,"duration_ms":32044,"concrete_test":"Rerun the shared-control experiment with a matched fixed-weight condition: set omega at each time step to the mean omega (or a randomly permuted replica of the adaptive omega trajectory) while keeping all other factors identical. If human effort under this matched fixed-weight condition is not significantly greater than under the adaptive condition, the benefit is due to average robot authority, not confidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that prediction confidence improves shared control rests on the comparison in Section V between 'Adaptive Weight Sharing' and 'Fixed Weight Sharing' with omega=0.5. In the adaptive condition, omega is set from the normalized prediction interval (Section V.A.3), so when the robot is confident, omega is likely small (more robot control). If the adaptive policy's average omega is below 0.5, human effort will drop even if the confidence signal carries no information about trajectory quality; any rule that reduces average human authority would show the same effect. The paper does not report the mean omega in the adaptive condition, nor does it include a baseline with fixed omega equal to that mean. Without such a control, the observed significant reduction in human effort (Table I) cannot be attributed to the confidence signal. The reported p-values are also computed over 1176 pooled time-step samples per condition from 14 trials, violating independence and inflating significance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CESN+, an extension of the context-based echo state network that augments reservoir-computing trajectory generation with prediction intervals computed from linear-regression theory. The authors compare CESN+ with Conditional Neural Movement Primitives (CNMP) on simple trajectory-generation tasks, claiming better accuracy and more reliable confidence signals in interpolation and extrapolation cases, and then apply CESN+ in a simulated Franka Emika shared-control task in which the normalized prediction interval sets the sharing weight between human joystick commands and robot-generated commands. They report that adaptive weight sharing significantly reduces human effort relative to a fixed 50-50 sharing policy.","tokens_in":1433,"tokens_out":1917,"duration_ms":110101,"significance":"If the central claims held, CESN+ would be a valuable lightweight alternative to Bayesian or neural-process models for producing uncertainty-aware movement predictions and using those uncertainties for arbitration in shared control. The strengths of the paper are the simplicity of the model, the explicit use of frequentist prediction intervals rather than ad hoc variance heuristics, and the concrete proof-of-concept setup with a simulated robot. However, the current evidence is not sufficient: the shared-control result lacks the control condition needed to attribute the effect to the confidence signal, the prediction intervals are not checked for calibration, and the comparison with CNMP is qualitative. The idea is promising and the experiments are reproducible in principle, but the load-bearing empirical claims need additional controls and quantitative validation.","major_comments":[{"comment":"The adaptive condition sets the sharing weight from the normalized prediction interval, while the fixed condition uses omega = 0.5. Because higher robot authority (lower omega) tends to reduce the human's commanded input, the significant reduction in effort reported in Table I could be explained entirely by the average omega being below 0.5 in the adaptive condition, regardless of whether the prediction interval carries information. The paper does not report the mean or distribution of omega in the adaptive condition, nor does it include a fixed-weight baseline with omega equal to that average. Without that matched control, the conclusion that confidence-based adaptation reduces human effort is not supported. Please add a fixed-omega baseline matched to the mean adaptive omega and report the omega trajectories.","section":"Section V.A.3 and V.B (Table I)"},{"comment":"The statistical tests treat each time step as an independent sample (N = 1176 per condition from 14 trials). Human joystick input at 10 Hz is strongly autocorrelated within a trial, so the effective sample size is much smaller than 1176 and the reported p-values are spuriously small. Analyze the data at the trial level (for example, mean effort per trial) or use a mixed-effects model with trial as a random effect, and report effect sizes and confidence intervals. The number of human operators should also be stated, since all 14 trials may come from a single operator.","section":"Section V.B (Table I)"},{"comment":"The prediction-interval formula assumes independent, homoscedastic residuals around the fitted linear readout. Reservoir state sequences are autocorrelated and the readout is a linear approximation of a nonlinear dynamical map, so the nominal coverage of these intervals is not guaranteed. No calibration check is reported anywhere in the paper; examples would be empirical coverage of the nominal 95% interval on held-out trajectories, or the correlation between interval width and absolute prediction error. Because the adaptive weight in Section V.A.3 is set directly from this interval, the shared-control claim is conditional on the intervals being meaningful. Please add such calibration evidence.","section":"Section III.C (Eq. 9)"},{"comment":"The performance comparison between CESN+ and CNMP is qualitative. Statements such as 'CESN+ exhibits superior performance' and 'CNMP's confidence metric is misleading' are based on visual inspection of a small number of trajectories. Please provide quantitative error metrics, such as trajectory RMSE and final-point error for known, interpolation, and extrapolation conditions, and for the confidence comparison use a quantitative measure such as calibration error or the correlation between confidence and prediction error. Multiple random seeds for both models would also strengthen the comparison.","section":"Section IV (Figs. 2-4)"},{"comment":"The mapping from the prediction interval to the sharing weight omega is not specified. The text says only that the 'normalized value of the prediction interval' determines omega; there is no equation, normalization constant, or clipping rule. This makes the experiment irreproducible and prevents assessment of how sensitive the result is to the chosen mapping. Please define the exact functional relationship and report the resulting range of omega values.","section":"Section V.A.3"}],"minor_comments":[{"comment":"The phrase 'with desirable properties such fast training' should read 'with desirable properties such as fast training'.","section":"Abstract"},{"comment":"'One challenge faced CNMPs' should be 'One challenge faced by CNMPs' or 'One challenge that CNMPs face'.","section":"Section II.A"},{"comment":"'Thedesired end-effector pose' contains a missing space and should be 'The desired end-effector pose'.","section":"Section V.A.1"},{"comment":"The matrix X in Eq. (9) is not defined precisely; specify whether a bias or constant column is included and state the dimensions of X and X_pred.","section":"Section III.C (Eq. 9)"},{"comment":"The reservoir hyperparameters (spectral radius, leaking rate, input scaling) are not reported beyond the reservoir size; these values are needed for reproducibility and for assessing the fairness of the comparison with CNMP.","section":"Section III.B"},{"comment":"References [2] and [47] appear to be the same paper and should be consolidated.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The shared-control experiment needs a matched fixed-weight control and trial-level statistical analysis, and the prediction intervals need calibration checks before the central claim can be accepted. The proof-of-concept nature is acceptable for a conference venue, but the current text states conclusions more strongly than the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper is a reasonable extension of the authors' own CESN—attach textbook linear-regression prediction intervals to the reservoir readout and use the interval width to set a human-robot sharing weight. That combination is new enough, and the model is genuinely lightweight and fast to train. The CNMP comparison is only qualitative, but the extrapolation figures do show the CESN+ tracking conditions better than CNMP, so the modeling claim is plausible.\n\nWhat the paper does well: it is clearly written, the PI formula is standard and the implementation is straightforward, and the shared-control task is a sensible proof-of-concept. The authors also note limitations of the fixed-weight comparison, though they do not address the most serious one.\n\nThe soft spots are real and they matter. First, the prediction intervals are never checked for calibration. Equation (9) assumes homoscedastic, independent residuals around a linear fit; reservoir states are highly autocorrelated and the readout is only an approximation, so the intervals may not have the claimed coverage. If the intervals are miscalibrated, the adaptive weight is not a sound confidence signal. Second, and more damaging, the shared-control experiment does not control for the average robot authority. In the adaptive condition, omega is set from the normalized PI, so when the robot is confident omega is small (more robot control). The paper does not report the mean omega in the adaptive condition, nor a fixed-weight baseline matched to that mean. Any rule that reduces average human authority would show the same reduction in human effort, even if the confidence signal carried no information about trajectory quality. The stress-test note is right. Third, the significance test pools 1176 time-step samples from 14 trials into a two-sample test; those samples are temporally correlated, so the p-values are inflated. The raw effect size may still exist, but the statistics do not establish it.\n\nWho this is for: readers working on lightweight LfD or confidence-based arbitration in HRI will find the CESN+ idea useful and the experiment suggestive. It deserves a serious referee, but the revision needs a matched fixed-weight control, calibration checks on the intervals, and a proper nested analysis of the trials. I would not cite it in its current form.\n\nRecommendation: send to peer review, require major revision.","headline":"CESN+ adds standard prediction intervals to an existing reservoir-computing LfD model, but the shared-control experiment does not isolate the confidence signal.","tokens_in":11259,"tokens_out":1694,"would_cite":false,"duration_ms":17578,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A lightweight reservoir-computing model can learn movement trajectories from demonstrations, output prediction intervals for its own forecasts, and use those intervals to automatically adjust how much control a human keeps in a…","keywords":["echo state networks","reservoir computing","prediction intervals","learning from demonstration","shared control","human-robot interaction","uncertainty quantification","movement primitives"],"falsifier":"Run CESN+ on held-out trajectories and compare the empirical frequency with which the true trajectory falls inside the reported $t_{\\alpha/2}$ interval against the nominal level (e.g., 95%): if the coverage is far below the nominal rate, especially for extrapolated conditions, then the interval is miscalibrated and the adaptive weight is not a reliable confidence signal. A second check is to inspect the autocorrelation of readout residuals; strong serial correlation directly violates the assumption behind Eq. (9).","tokens_in":10401,"feed_emoji":"🤖","tokens_out":4348,"duration_ms":37654,"temperature":0.7,"pith_summary":"CESN+ is a learning-from-demonstration model built on a fixed random recurrent reservoir with a linear readout, trained in one pass on a handful of example trajectories. The paper claims that by adding the standard linear-regression prediction interval to the readout, the model can output both a context-conditioned trajectory and a defensible confidence band for it, and that this confidence signal can be used online to set the human share of control. In a simulated pick-and-place task with obstacles, using the prediction interval to adapt the sharing weight significantly reduced measured human effort compared with a fixed 50-50 split. If the claim holds, a lightweight, non-Bayesian model can supply the uncertainty information needed for adaptive human-robot arbitration.","feed_headline":"Reservoir network with confidence cuts human effort in shared control","feed_subtitle":"A lightweight model learns from a few demonstrations and uses prediction intervals to decide when the human should retake control.","key_machinery":"The load-bearing object is the Context-based Echo State Network plus prediction interval (CESN+). A reservoir of fixed random recurrent units with leaking rate $\\alpha$ is driven by a context-augmented input; the readout $y(t)=W^{\\mathrm{out}}[1;x(t)]$ is fit by linear regression to demonstrated trajectories. Confidence comes from the standard multivariate prediction interval $\\hat{Y}_{\\mathrm{pred}} \\pm t_{\\alpha/2}\\, s\\, \\sqrt{1 + X_{\\mathrm{pred}}^T (X^T X)^{-1} X_{\\mathrm{pred}}}$ (Eq. 9), where $X$ collects training reservoir states and $s$ is the residual standard error. In shared control, the normalized width of this interval sets the human weight $\\omega$, so the robot yields control when its own forecast is uncertain. The model needs no probabilistic parameters over the reservoir; the interval is computed from linear-regression statistics around the readout.","core_discovery":"The central claim is that a reservoir readout trained by ordinary linear regression can be upgraded with prediction intervals (Eq. 9) to yield a practical confidence-aware movement generator. The authors argue that CESN+ generates trajectories that satisfy user-specified context points at least as accurately as Conditional Neural Movement Primitives (CNMP), and that it does better in extrapolation regimes where CNMP degrades; moreover, its confidence interval widens when its prediction is poor, whereas CNMP can report high confidence for badly wrong trajectories. In the shared-control experiment, the robot's control share is set from the normalized prediction interval after a checkpoint, and the adaptive scheme produces a statistically significant reduction in human command magnitude across 14 trials, with $p<0.0001$ on both a $t$-test and a Mann-Whitney $U$ test. The intended conclusion is that uncertainty from a frequentist prediction interval is sufficient to arbitrate shared control, without Bayesian machinery.","pith_inferences":["Because the prediction interval is derived from local regression geometry, the method could be extended to conformal prediction to obtain distribution-free coverage guarantees; the paper does not attempt this.","The single-checkpoint design could be generalized to continuous re-conditioning at multiple checkpoints; the authors list this as future work, and the results suggest each re-conditioning would refresh the context and narrow the interval.","The human-effort reduction may partly reflect that the adaptive scheme trusts the robot more in easy segments and cedes control earlier; a human-factors study with real operators would be needed to confirm that the reduced joystick effort is experienced as reduced workload."],"forward_implications":["A robot controller can use the prediction interval from a reservoir readout as a real-time authority signal, letting the human take over only when the model is uncertain.","Since training is a single linear-regression fit, the approach can be re-trained online from a few demonstrations, making confidence-aware control feasible on low-power hardware.","The comparison with CNMP suggests that frequentist intervals can expose model failure modes (overconfidence) that a learned variance output may hide.","The adaptive arbitration scheme should transfer to other continuous shared-control tasks where a goal or obstacle configuration can be expressed as a context."],"supporting_citations":[{"why":"Conditional Neural Movement Primitives, the baseline model that CESN+ is compared against.","marker":"[10]"},{"why":"Conditional Neural Processes, the foundation on which CNMP is built.","marker":"[8]"},{"why":"Context-based Echo State Networks, the prior model that CESN+ extends with prediction confidence.","marker":"[11]"},{"why":"Echo State Networks, the reservoir-computing architecture underlying the model.","marker":"[6]"},{"why":"Practical confidence and prediction intervals, the statistical framework used for uncertainty quantification.","marker":"[39]"},{"why":"Source of the prediction-interval formula in Eq. (8) and its multivariate extension in Eq. (9).","marker":"[46]"},{"why":"Prior adaptive shared control work that supplies the experimental paradigm and CNMP architecture choices.","marker":"[47]"},{"why":"Documents the extrapolation limitation of CNMP that motivates the comparison.","marker":"[44]"}],"fun_headline_variants":["Echo state network with confidence cuts human effort in shared control","Confidence-aware echo state network learns faster and extrapolates better","Uncertainty-guided echo state network improves human-robot shared control","Prediction intervals from reservoir net reduce human load in shared control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The prediction-interval formula assumes that the errors of the readout around the fitted trajectory are independent and have constant variance, but reservoir states over time are strongly autocorrelated and the linear readout is only an approximation, so the stated coverage of the intervals may not hold.","fun_headline_variants_meta":{"raw":{"variants":["Echo state network with confidence cuts human effort in shared control","Confidence-aware echo state network learns faster and extrapolates better","Uncertainty-guided echo state network improves human-robot shared control","Prediction intervals from reservoir net reduce human load in shared control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0008,"raw_usage":{"total_tokens":3551,"prompt_tokens":1009,"completion_tokens":2542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":2469}},"tokens_in":625,"tokens_out":2542,"duration_ms":17715,"temperature":1.0,"reasoning_tokens":2469,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:14:11.775930+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CESN+ on held-out trajectories and compare the empirical frequency with which the true trajectory falls inside the reported $t_{\\alpha/2}$ interval against the nominal level (e.g., 95%): if the coverage is far below the nominal rate, especially for extrapolated conditions, then the interval is miscalibrated and the adaptive weight is not a reliable confidence signal. A second check is to inspect the autocorrelation of readout residuals; strong serial correlation directly violates the assumption behind Eq. (9).","supporting_citations":[{"cited_title":"Conditional neural movement primitives","cited_arxiv_id":null,"evidence_quote":"Conditional Neural Movement Primitives, the baseline model that CESN+ is compared against."},{"cited_title":"Conditional neural processes,","cited_arxiv_id":null,"evidence_quote":"Conditional Neural Processes, the foundation on which CNMP is built."},{"cited_title":"Context based echo state networks for robot movement primitives,","cited_arxiv_id":null,"evidence_quote":"Context-based Echo State Networks, the prior model that CESN+ extends with prediction confidence."},{"cited_title":"Echo state network,","cited_arxiv_id":null,"evidence_quote":"Echo State Networks, the reservoir-computing architecture underlying the model."},{"cited_title":"Practical confidence and prediction intervals,","cited_arxiv_id":null,"evidence_quote":"Practical confidence and prediction intervals, the statistical framework used for uncertainty quantification."},{"cited_title":"Confidence intervals and prediction intervals for feed-forward neural networks","cited_arxiv_id":null,"evidence_quote":"Source of the prediction-interval formula in Eq. (8) and its multivariate extension in Eq. (9)."},{"cited_title":"Adaptive shared control with human intention estimation for human agent collabora- tion,","cited_arxiv_id":null,"evidence_quote":"Prior adaptive shared control work that supplies the experimental paradigm and CNMP architecture choices."},{"cited_title":"ACNMP: Skill Transfer and Task Extrapolation through Learning from Demonstration and Reinforcement Learning via Representation Sharing","cited_arxiv_id":"2003.11334","evidence_quote":"Documents the extrapolation limitation of CNMP that motivates the comparison."}],"review_version":1}