{"id":"1b20b481-0c5e-4952-a6cb-6ff1f32ffa26","arxiv_id":"2512.13009","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"K-VARK combines a kernelized movement-primitive model of residual joint torques with an adaptive Kalman filter to estimate external forces on a 6-DoF robot without force sensors.","lead":"This paper introduces a filter that lets a robot estimate external contact forces without a force sensor by learning the mean and variability of unmodeled joint torques. If it holds up, it could make force-controlled polishing, assembly, and human-robot contact safer and cheaper.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Residual-torque validity under load is the load-bearing condition: Eq. (28) attributes whatever the velocity-only KMP mean cannot explain to external torque, so payload/configuration sensitivity must be demonstrated before the headline claim can stand.","rationale":"The reader's weakest assumption is correct and I agree: the observer has no independent mechanism to distinguish actual external torque from residual-model error; the distinction is made entirely by the KMP mean. This is load-bearing because the claimed accuracy improvement depends on that distinction. The paper's own Remark 1 and ARD analysis concede that velocity-only is a simplification, and the one-sample lag in Eq. (28) compounds the possible mismatch. I do not raise circularity: the offline learning/online observer separation is legitimate. My proposed payload/config variation directly tests the invariance assumption; without it, the 20% claim is limited to the specific setup. The reader's conditional verdict should stand, so no verdict change is needed.","tokens_in":18485,"tokens_out":6484,"duration_ms":61702,"concrete_test":"Run the F/T-instrumented contact experiment (Section V-A) again with, say, three payload masses (0.5/1.0/1.5 kg) and two distinct end-effector configurations, keeping the excitation/contact trajectory fixed and retraining KMP only on free-motion data as described. Compute K-VARK's Cartesian RMSE against the F/T ground truth. If RMSE increases by more than ~20% relative to the reported 4.42 Nm, or scales with payload, the velocity-only residual model does not transfer and Eq. (28) is absorbing model error as external torque. As a secondary check, report paired per-trial RMSE for K-VARK vs GPADKF to see whether the claimed 'over 20% vs SOTA' survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Eq. (28), where ζ*_k = x_k − x_{k−1} − t_s u_{k−1} + t_s μ*_k is treated as −t_s τ_ext,k−1. This is valid only if the KMP mean μ*_k equals the true residual torque at that instant in the deployed condition. But μ* is trained offline from free-motion data on joint velocity alone (Section IV-A Remark 1, Section IV-B), and the ARD analysis (Fig. 13) shows non-velocity features are not irrelevant. Under a payload or during contact, the nominal dynamics error changes (inertia, gravity, load-dependent friction), so a velocity-only free-motion residual cannot absorb it; the unmodeled component enters ζ* and is estimated as external force. The first experiment's ground truth, τ_loaded − τ_free on the same trajectory, presupposes exactly the residual invariance that is in question; the F/T experiment provides only one payload/configuration. Additionally, Eq. (28) uses μ*_k while the momentum difference spans the interval k−1→k, an undiscussed one-sample lag that makes the stated equality inexact. Separately, the abstract's 'over 20%' is computed against the GMR-GP baseline; against GPADKF, the state-of-the-art baseline from [24], Experiment 1 shows ~7% improvement. These issues do not make the method circular or clearly wrong, but they make the quantitative and general validity claims conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes K-VARK, a sensorless external-torque observer for collaborative robots. A per-joint residual-torque model is learned offline with Kernelized Movement Primitives (KMP) from free-motion excitation data, providing a predictive mean and an input-dependent variance. The mean is used to correct a momentum-based residual, giving a virtual measurement of external torque; the variance is added to the measurement-noise covariance, and the process-noise covariance is adapted online via variational Bayes. Experiments on a 6-DoF robot compare the method against GMR-GP, GPADKF, and NN-based observers, reporting lower RMSE in joint-space and Cartesian wrench estimation.","tokens_in":18883,"tokens_out":5847,"duration_ms":60652,"significance":"If the reported accuracy gains hold under the stated operating conditions, integrating KMP-style heteroscedastic uncertainty into an adaptive Kalman filter is a plausible and practically useful contribution to sensorless force estimation. The work addresses a real limitation of GP-based residual models, which typically capture only epistemic uncertainty, and it provides real-robot comparisons against multiple baselines, including the GPADKF approach of [24] and a neural-network variant. The paper is also transparent about several limitations in its conclusion. However, as detailed below, the central quantitative claim is not yet supported by the experiments as written, and there is a formal inconsistency in the virtual-measurement update that must be resolved.","major_comments":[{"comment":"The virtual measurement is defined as ζ*_k = x_k − x_{k−1} − t_s u_{k−1} + t_s μ*_k = −t_s τ_ext,k−1, but Eq. (29) and the KF recursion (44) treat ζ*_k as a direct measurement of the current state τ_ext,k. This is an unflagged one-step lag. With the random-walk state model (32), the update (44) corrects the current external-torque estimate with a measurement of the previous external torque, which is not the stated equality. Either define ζ*_{k+1} = −t_s τ_ext,k, or set the measurement model to H ω_{k−1}. The same index issue applies to the KMP mean: Eq. (8) requires the residual torque at the previous sample, τ_r,k−1, not μ*_k.","section":"Section IV-C, Eqs. (28)-(29) and Algorithm 1, line 18"},{"comment":"The ground truth in the first experiment is τ_loaded − τ_free. This equals the true external torque only if the residual torque is identical with and without the load. The residual model is trained on free-motion data using only joint velocity as input (Remark 1, Section IV-A), and the ARD analysis in Section VI (Fig. 13) shows that non-velocity features are not wholly irrelevant. Under a payload or during contact, load-dependent changes in friction and other residual effects are not captured by μ*, so they are attributed to τ_ext by Eq. (28). The reported RMSE therefore does not isolate external-torque estimation error. This is the load-bearing assumption for the headline claim and should be validated with varying payloads/contact configurations, or the claim should be weakened.","section":"Section V-A, Eq. (46)"},{"comment":"The abstract's 'over 20% reduction in RMSE' is computed only against the GMR-GP baseline in Experiment 1 (1.49 → 1.18). Against GPADKF, the state-of-the-art baseline from [24], the improvement is about 7% (1.27 → 1.18). In Experiment 2, the Cartesian RMSE improvement over GMR-GP is about 4.7%. The quantitative claims should be reported per-domain and against the strongest baseline. In addition, the Cartesian average in Table IV is computed as a Euclidean norm over forces and moments without unit normalization, producing a mixed-unit scalar that is not a meaningful average; separate force and moment metrics should be reported.","section":"Abstract and Section V-C, Figs. 7-8, Table IV"}],"minor_comments":[{"comment":"Several typos: 'methods are often suffer', 'Generalized momentum observerss', and '64GP of RAM'.","section":"Section II and V-A"},{"comment":"The notation l^{-1} in the squared-exponential kernel is ambiguous; l is introduced as a vector in the hyperparameter list. Define whether l is a diagonal matrix or a vector and write the kernel accordingly.","section":"Section III-D, Eq. (21)"},{"comment":"The confidence intervals and method colors are hard to distinguish in black-and-white print. Consider using line styles or adding direct labels.","section":"Figures 5 and 7"},{"comment":"K-VARK is not the lowest in several columns (e.g., τ1, τ6, mz). The claim of highlighting the lowest value in each column is not realized, and pairwise improvement should be stated with per-column numbers rather than only averages.","section":"Table IV"},{"comment":"The notation E[P_{k|k-1}^{-1}] is unclear; the standard posterior precision update uses P_{k|k-1}^{-1}, not an expectation of it. Please clarify the operation.","section":"Section IV-D, Eq. (35)"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern raised in the stress-test note does not land: the KMP residual is learned from free-motion data without external force, so there is no direct rearrangement of the fitted model into the estimate. The two substantive issues are the index inconsistency in the virtual measurement and the validity of the residual-invariance assumption under load. Both are fixable, but the second may require additional experiments or a substantial weakening of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick read for you. The paper proposes K-VARK: train a KMP model of joint residual torque (mean + input-dependent heteroscedastic variance) from free-motion excitation data, then feed both into a disturbance Kalman filter, with the variance inflating the virtual measurement noise and a variational-Bayes recursion adapting the process noise. That specific combination is new, as far as I can tell; previous GPADKF used sparse GP with epistemic-only uncertainty. The authors actually build the thing and test it on a 6-DoF arm, including an F/T-instrumented polishing experiment. That earns real credit.\n\nWhere it gets soft: the abstract's 'over 20% RMSE reduction' is true only against the GMR-GP baseline in the first experiment (1.49 to 1.18 Nm). Against GPADKF, the better baseline, the improvement is about seven percent there and two to five percent in the Cartesian F/T experiment. No error bars, no trial-to-trial variation, no code or data. The filter equations are standard but Eq. (28) forms a virtual measurement from momentum differences over [k−1,k] while Eq. (29) treats it as a direct measurement at k; that's a one-sample lag that should at least be acknowledged.\n\nThe load-bearing weakness, though, is the transfer assumption. The residual model is trained on free motion with only joint velocity as input (Remarks 1 and 2), and it works because in free motion the residual is mostly friction-like. During contact or with a payload, inertia, gravity, and load-dependent friction change the nominal dynamics error. Eq. (28) subtracts the KMP mean and attributes everything left to external torque, so any unmodeled load-induced residual gets read as contact force. The first experiment's ground truth, τ_loaded − τ_free on the same trajectory, assumes exactly the residual invariance that's in question; the F/T experiment covers one payload and one configuration. That doesn't make the method circular — the model is trained on data with no external force — but it does mean the headline generalizability claim is not supported.\n\nThe paper itself admits single-robot evaluation, no ablation of the variance gate or VB adaptation, and near-zero-velocity issues. So it's an honest, plausible engineering contribution that would benefit from a proper revision: justify the lag, add variability metrics, and demonstrate at least one payload/configuration change. A serious referee should see it.\n\nRecommendation: send to peer review, with the expectation of major revision. It's a real idea, not a toy.","headline":"K-VARK is a genuinely new combination — KMP residual mean/variance into an adaptive Kalman filter — but the headline 20% gain shrinks to single digits against the stronger baseline, and the paper never tests whether a velocity-only free-motion residual model stays valid under load.","tokens_in":19415,"tokens_out":2664,"would_cite":true,"duration_ms":24937,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Kalman filter that trusts a learned residual-torque model only where data supports it cuts sensorless force-estimation error by more than 20 percent.","keywords":["sensorless force estimation","residual torque modeling","Kernelized Movement Primitives","adaptive Kalman filter","variational Bayes","heteroscedastic variance","collaborative robots","momentum observer"],"falsifier":"Run the robot along the training trajectories with a known end-effector mass attached, and compare K-VARK's estimates against a wrist F/T sensor. If the error grows systematically with payload size while the KMP variance stays approximately constant, the velocity-only residual model is missing load-dependent torque and the central claim is falsified.","tokens_in":18348,"feed_emoji":"🤖","tokens_out":4369,"duration_ms":42323,"temperature":0.7,"pith_summary":"The paper proposes K-VARK, a sensorless force observer that replaces the usual 'friction model' in a momentum-based Kalman filter with a kernelized probabilistic model of residual torques—everything the nominal dynamics fail to explain. The key move is to feed not just the predicted mean of that residual model into the filter, but also its input-dependent variance, which is added to the measurement noise covariance. In data-rich velocity regions the filter trusts the virtual measurement; in sparse or extrapolated regions the variance inflates and the filter falls back on its prior, while a variational-Bayes step adapts the process noise online. On a 6-DoF collaborative robot, the authors report over 20% lower RMSE than a Gaussian-process baseline for external torque estimation, with real-time per-sample cost, in both loaded-trajectory and F/T-instrumented polishing experiments. The authors explicitly list limitations: single robot, limited near-zero-velocity analysis, and no full ablation of the variance and VB components.","feed_headline":"Force sensing without a sensor: 20% lower error","feed_subtitle":"Learned residual-torque variance tells a Kalman filter when to trust its virtual measurement.","key_machinery":"The engine is the virtual measurement ζ*_k = x_k − x_{k−1} − t_s u_{k−1} + t_s μ*_k (Eq. 28), which converts the momentum residual into a noisy observation of external torque after subtracting the KMP mean. Its noise covariance is Σ_{ν,k} = t_s² Σ*_{k} + Σ_{emp,k}: the KMP predictive covariance—capturing both aleatoric data variability and epistemic distance-to-training effects—plus an innovation-driven empirical term. A variational-Bayes inverse-Wishart update adapts the process noise covariance online, so the filter can respond to contact transitions while the KMP variance decides how much to trust the measurement.","core_discovery":"K-VARK's central claim is that residual torques—the mismatch between commanded motor torque and the nominal rigid-body model—should be treated as a probability distribution, not a fixed friction curve, and that distribution's variance is as informative as its mean. The paper shows that when the KMP-predicted mean is subtracted from the momentum residual, the leftover is a clean virtual measurement of external torque only if the residual model is right; when the KMP variance is added to the measurement noise, the Kalman gain automatically down-weights that virtual measurement in regions where the residual model is uncertain. On a 6-DoF arm, this yields joint-space and Cartesian force/torque e","pith_inferences":["The same variance-gating principle could set safety thresholds in impedance control or collision detection: whenever KMP variance is large, a controller could automatically reduce interaction stiffness or raise a contact alarm.","A natural falsifying extension is to carry known payloads during training-free operation; if the velocity-only residual model cannot explain residuals under load—their remaining bias exceeding the KMP variance—the central assumption fails.","Because KMP's uncertainty hyperparameters (λ2, σ_f²) can be tuned without changing the predictive mean, transferring the model to a new robot may only require recalibrating variance scaling, not re-fitting the mean—an unverified but plausible consequence.","Near-zero velocity is the weak spot the authors themselves flag: stick-slip and temperature-dependent friction are exactly the regimes where a velocity-only heteroscedastic model is most likely to under-represent the true residual distribution."],"forward_implications":["Joint-space and Cartesian external wrench estimates improve over GP- and NN-based observers, most clearly against the Gaussian-process baseline, while per-sample compute stays below real-time limits.","In data-rich velocity regions the filter trusts the learned residual model; in extrapolated regions it automatically down-weights the virtual measurement, which should make the observer more robust to novel motions.","The KMP variance encodes both aleatoric and epistemic uncertainty, so unlike GP-based observers the filter carries a calibrated confidence signal into the estimation loop.","Because no force/torque sensor is needed, the approach bears directly on cost-sensitive collaborative applications like polishing, assembly, and surface finishing.","The variational-Bayes process-noise adaptation couples to the measurement variance, so the filter can track time-varying disturbances without hand-tuned gain scheduling."],"fun_headline_variants":["Variance-aware Kalman filter cuts sensorless force error 20%","Residual torque variance drives 20% better force estimates","Residual uncertainty guides Kalman filter to 20% lower force error","Kalman filter learns residual variance to boost force estimates 20%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The offline residual-torque model, trained only in free motion as a function of joint velocity, is assumed unchanged during loaded and contact motion; if residual torque also depends on payload, configuration, or contact state, the virtual measurement is biased and the force estimate inherits that bias.","fun_headline_variants_meta":{"raw":{"variants":["Variance-aware Kalman filter cuts sensorless force error 20%","Residual torque variance drives 20% better force estimates","Residual uncertainty guides Kalman filter to 20% lower force error","Kalman filter learns residual variance to boost force estimates 20%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001223,"raw_usage":{"total_tokens":4857,"prompt_tokens":728,"completion_tokens":4129,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":4063}},"tokens_in":472,"tokens_out":4129,"duration_ms":24528,"temperature":1.0,"reasoning_tokens":4063,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:30:36.304320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the robot along the training trajectories with a known end-effector mass attached, and compare K-VARK's estimates against a wrist F/T sensor. If the error grows systematically with payload size while the KMP variance stays approximately constant, the velocity-only residual model is missing load-dependent torque and the central claim is falsified.","supporting_citations":[],"review_version":1}