{"id":"574ad690-d2c7-4944-85a3-795f11a614f9","arxiv_id":"2506.05812","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"A simulation of force-only Velcro peeling with a particle filter, whose headline optimality and success claims conflict with its own tables.","lead":"The paper models robotic Velcro peeling as a quasi-static geometric problem and proposes a particle-filter-based controller that uses only force and position feedback. The authors claim near-optimal performance, but the optimal solution is not actually derived and the headline success and cost numbers contradict the results tables.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised closed-form optimal solution is never derived; the Full-Obs baseline is only a heuristic, and the abstract's 100%/80% claims contradict the paper's own Table I/II.","rationale":"In good faith, the paper aims to solve force-only Velcro peeling with a tractable analytic model, particle-filter state estimation, and a heuristic controller; the system model and estimator are reasonable engineering contributions. However, the central advertised contribution is the closed-form optimal solution. For that claim to hold, the fully-observable optimization problem defined by Eqs. (3), (8), and (9) must be solved and proven, and the resulting controller must be the Full-Obs baseline used in the experiments. The manuscript lacks any such derivation or proof; Section IV.C describes Full-Obs only as a heuristic. The abstract's numerical claims can be checked directly against the tables and fail: Corner success is 98.5%, not 100%, and the Corner energy ratio is about 105%, not under 80%. The reader's weakest-assumption identification (tautness) is a real modeling limitation, but it is not the most load-bearing issue because the paper explicitly builds on it and maintains it in the controller. The decisive problem is that the central claim is unsupported by derivation and contradicted by the paper's own reported metrics. I therefore keep the reader's REJECT verdict.","tokens_in":9932,"tokens_out":7286,"duration_ms":64645,"concrete_test":"Independently derive the claimed closed-form optimal policy from Eqs. (3), (8), and (9) for the fully-observable case and check whether the Full-Obs controller in Sec. IV.C is exactly that minimizer. If no closed-form minimizer can be exhibited, or if Full-Obs is not it, the optimality claim is unsupported. Also recompute the abstract's performance claims from Table I and II: the Corner row must satisfy 103.7/50.6 < 1.8 and 98.5% = 100% for the headline numbers to hold.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim requires both a closed-form policy that minimizes total energy cost in the fully-observable case and the reported 100% success with less than 80% energy increase in the partially observable case. Neither is established. The manuscript defines the state and action models in Sec. III.A and the energy cost in Eq. 9, but no theorem, equation, or proof presents the optimal policy or its closed-form value. The fully-observable baseline in Sec. IV.C is explicitly called a simple heuristics whose cost may act as the lower bound, not a derived optimum, and Algorithm 1 is heuristic without an optimality guarantee. The numbers also contradict the tables: Table II reports 98.5% success for the Corner case, not 100%, and Table I gives 103.7/50.6 ~ 2.05, a 105% cost increase over the Full-Obs baseline, not less than 80%. The taut-peeled-part assumption in Sec. III.A is a genuine scope condition, but it is stated and maintained by the controller; the missing optimality derivation and the internal numerical inconsistency are the load-bearing defects.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a model-based approach for robotic Velcro peeling using force and position feedback. It models the peeling state under a taut-peeled-part assumption, designs a particle filter with state-space decomposition, and presents a heuristic controller that alternates peeling and non-peeling actions. The central claimed contributions are a closed-form optimal solution for the fully observable case and a partial-observability controller achieving 100% success with less than 80% energy increase over that optimum.","tokens_in":10133,"tokens_out":5015,"duration_ms":45152,"significance":"If the central claims were established, the paper would provide a useful theoretical baseline for a challenging deformable-object manipulation task with limited sensing. The state estimation scheme with decomposed particle filter updates is a reasonable engineering choice, and the empirical comparison against RL baselines suggests practical advantages. However, the closed-form optimality claim is never substantiated, and the headline numbers in the abstract are directly contradicted by the paper's own tables. The analytical model is a strength, but the paper does not support its strongest claims.","major_comments":[{"comment":"The claimed 'closed-form optimal solution which minimizes the total energy cost' is never derived or stated in the manuscript. Section IV.C explicitly describes the Full-Obs baseline as 'a simple heuristics' whose cost 'may act as the lower bound', and Algorithm 1 is heuristic. No theorem, equation, or policy expression is given for an optimal solution, so the abstract's comparison 'less than 80% increase compared to the optimal solution' is not meaningful.","section":"Abstract and Section I"},{"comment":"The abstract and Section I claim a 100% success rate, but Table II reports 98.5% for the Corner case. This is a direct numerical contradiction that must be resolved.","section":"Table II vs Abstract"},{"comment":"The abstract claims 'less than 80% increase in energy cost compared to the optimal solution'. Using Table I, the Corner case has 103.7/50.6 = 2.05, i.e., a 105% increase. Even if the intended claim is an average over shapes, the wording is unconditional and is violated by the published numbers.","section":"Table I vs Abstract"},{"comment":"The cost function U(X) is constructed to penalize deviation from phi - theta = pi/2, and Algorithm 1's first step enforces exactly that configuration. The manuscript does not show that this potential function corresponds to the physical energy of peeling or that the proposed policy minimizes the total cost in Eq. (9). Without an independent derivation, the 'optimality' claim is partly circular.","section":"Section III.C, Eq. (8)-(9)"},{"comment":"The model relies on the peeled part being always taut and on a zero-mean Gaussian random walk for theta (Eq. 7), which limits the surface curvature. The abstract states that 'the surface geometry is arbitrary and unknown', which is stronger than these model assumptions. The paper should either restrict the claim to geometries satisfying the tautness and mild-curvature conditions or provide evidence that violations do not affect the reported performance.","section":"Section III.A and Eq. (7)"}],"minor_comments":[{"comment":"The reported costs and success rates are point estimates over 200 configurations; standard deviations or confidence intervals would strengthen the quantitative claims, especially given the exact percentages in the abstract.","section":"Tables I and II"},{"comment":"The axis labels and units in panels (c), (d), (g), and (h) are unclear; the time-step axis is not labeled in all subplots, making it difficult to interpret the evolution of theta, phi, and r.","section":"Figure 6"},{"comment":"Reference [8] is a duplicate of reference [6]; the paper should use distinct references for distinct works.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript overclaims its central contributions: the closed-form optimal solution is never presented, and the abstract's success-rate and cost-increase numbers are contradicted by the paper's own Tables I and II. These are load-bearing issues that cannot be fixed by minor edits; the claims must either be substantially reworked or the paper's contribution redefined."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful part: the paper builds a clean quasi-static model of Velcro peeling under a taut-peeled-part assumption, designs a particle filter that deliberately updates only subsets of the state depending on action type, and shows in simulation that this estimator plus a heuristic controller beats two RL baselines on flat, arc, and corner shapes. That is a real, if incremental, contribution to force-only deformable-object manipulation. The state-space decomposition and the auxiliary PCA/arc fitting for the hidden hinge orientation are the most thought-through pieces. The citation pattern is fair: related work covers contact manipulation and active perception, and the self-citation to their own RL peeling paper is relevant.\n\nNow the soft spots, in rough order of size. The central advertised result, a closed-form optimal solution for the fully observable case, is never actually presented. No theorem, equation, or derivation gives the optimal policy or its value. When you get to the evaluation, the Full-Obs baseline is described, in the authors' own words, as 'a simple heuristics' whose cost 'may act as the lower bound.' Algorithm 1 is heuristic. So the abstract's phrase 'optimal solution in closed-form which minimizes the total energy cost' is not supported by anything in the paper. This is the load-bearing claim.\n\nSecond, the headline numbers conflict with the tables. The abstract says '100% success rate with less than 80% increase in energy cost compared to the optimal solution.' Table II gives Corner success as 98.5%, not 100%. Table I gives Corner cost 103.7 versus Full-Obs 50.6, which is about a 105% increase, not less than 80%. Flat and Arc are under 80% when measured against the Full-Obs heuristic, but the blanket statement is false.\n\nThird, there are no error bars or confidence intervals on any of the cost or success numbers across the 200 random configurations, and several parameters needed to reproduce the simulation (the c1 and c2 weights, measurement noise sigmas, action step d, health-index thresholds, forbidden-zone angles) are not listed. The paper is simulation-only; the accompanying video is mentioned but no hardware results appear in the text.\n\nThe taut-peeled-part assumption is a genuine scope condition, but it is stated clearly and the controller actively maintains it. I would not call that a flaw; it is a limitation that belongs in the abstract.\n\nWho is this for: people working on deformable-object manipulation and active perception with force/tactile feedback. They will find the state-estimation design useful even if they do not take the optimality claim seriously. The paper deserves a serious referee, because the method is concrete and the experiments are non-trivial. But the referee should require that the claims be corrected: drop 'optimal' or actually derive it, fix the abstract numbers, and report the missing parameters and variances. As written, the central claim does not hold up.","headline":"A worthwhile force-only peeling pipeline with a decomposed particle filter, but the advertised closed-form optimality is never derived and the abstract's success/cost claims contradict the paper's own tables.","tokens_in":10663,"tokens_out":3731,"would_cite":false,"duration_ms":35290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Force feedback alone lets a robot peel Velcro straps near-optimally.","keywords":["velcro peeling","deformable object manipulation","force feedback","particle filter","state estimation","quasi-static model","energy optimization","robot manipulation"],"falsifier":"Run the controller on a compliant or loosely-tensioned strap so the peeled segment visibly sags, and record whether the particle-filter hinge estimates diverge from the true peel line while the success rate drops below the reported 100%: if they do, the tautness premise is the failure point.","tokens_in":9665,"feed_emoji":"🤖","tokens_out":4471,"duration_ms":42988,"temperature":0.7,"pith_summary":"This paper shows that a robot can peel a Velcro strap from an unknown, arbitrarily-shaped surface using only force and end-effector position feedback. The authors derive a closed-form optimal peeling strategy for the fully-observable case, minimizing a total energy cost defined by keeping the peeled part perpendicular to the attached part. For the realistic partially-observable case, they build a particle-filter state estimator that recovers the hidden surface orientation from noisy force and position readings, and a heuristic controller that alternates exploratory and exploitative actions to keep the estimate healthy. In simulation with complex geometries and sensor noise, the method achieves a 100% success rate while staying within 80% of the optimal energy cost, far ahead of reinforcement-learning baselines.","feed_headline":"Force-only robot peels Velcro near-optimally","feed_subtitle":"Closed-form optimal solution sets the limit; particle-filter controller stays within 80% energy with full success.","key_machinery":"The load-bearing device is the taut-peeled-part geometry: the peeled segment is treated as a straight line of length $r$ at angle $\\phi$, so the five-dimensional state $X = [h_x, h_y, \\theta, \\phi, r]$ and the action $A = [\\alpha, d, s, \\delta\\phi]$ satisfy the algebraic transition constraints of Eq. (3). This geometry yields the closed-form fully-observable solution, the observation model in Eqs. (5–7), and the energy cost of Eq. (9). The particle filter's state-space decomposition and the controller's health index, which triggers non-peeling exploratory actions when the weight distribution signals sample impoverishment, carry the partial-observability result.","core_discovery":"Once the peeled segment of Velcro is assumed taut, the peeling configuration at any time is captured by five parameters: hinge position, surface orientation angle $\\theta$, peel angle $\\phi$, and peeled length $r$. The paper's central claim is that for this model the fully-observable peeling problem admits a closed-form optimal solution minimizing the energy cost $U(X) = 1 + \\|\\phi - \\theta - \\pi/2\\|^2$, which serves as a theoretical performance limit. In the partially-observable case, a particle filter with state-space decomposition—updating $\\phi$ under peeling actions, $\\phi$ and hinge position under non-peeling actions, and $\\theta$ via a line/arc fitting auxiliary estimator—combined with a heuristic controller yields 100% success with less than 80% energy increase over the optimal baseline across flat, arc, and corner geometries.","pith_inferences":["The tautness-maintenance requirement could be relaxed by adding a slackness-aware state or by actively controlling tension, an extension the paper does not pursue.","The unobservability of $\\theta$ in the transition model suggests the particle filter needs the auxiliary line/arc fitting step; a similar auxiliary estimator may be necessary in other tasks where a hidden parameter is Markov but unconstrained by actions.","If the 80% energy gap holds on hardware with real force-sensor noise, tactile-only peeling could be integrated into shoe-removal and garment-assist robots operating in low-light or occluded conditions."],"forward_implications":["A robot with only a wrist force sensor and proprioception can reliably peel straps from curved and cornered surfaces, removing the need for visual feedback during the peeling phase.","The closed-form fully-observable solution provides a reusable performance lower bound for any future Velcro-peeling controller.","The state-estimation design—updating only observable subsets per action type—is a template for other deformable-object manipulation tasks with mixed observable and hidden state parameters.","The method's success suggests that quasi-static modeling plus Bayesian filtering can outperform model-free reinforcement learning in this contact-rich manipulation regime."],"supporting_citations":[{"why":"Supplies the particle filter update and resampling scheme that the state estimator is based on.","marker":"[26]"},{"why":"Prior RL Velcro-peeling work that this method improves upon; provides the motivation and the RL baseline comparison.","marker":"[1]"},{"why":"PPO algorithm used to train the PF+RL baseline controller.","marker":"[28]"},{"why":"Transformer architecture used by the PF+RL and RL baseline agents.","marker":"[29]"}],"fun_headline_variants":["Velcro peeling robot uses force only, hits optimal limit","Closed-form optimal strategy for robotic Velcro peel","Robot peels Velcro blind, matches optimal energy","Force-feedback robot peels Velcro 100% success","Particle filter robot peels Velcro near-optimum"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The peeled part of the Velcro must remain taut for the entire task: if the strap goes slack, the straight-line geometric model, the transition equations, and the particle filter's measurement models no longer describe reality.","fun_headline_variants_meta":{"raw":{"variants":["Velcro peeling robot uses force only, hits optimal limit","Closed-form optimal strategy for robotic Velcro peel","Robot peels Velcro blind, matches optimal energy","Force-feedback robot peels Velcro 100% success","Particle filter robot peels Velcro near-optimum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000524,"raw_usage":{"total_tokens":2523,"prompt_tokens":929,"completion_tokens":1594,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":1514}},"tokens_in":545,"tokens_out":1594,"duration_ms":11027,"temperature":1.0,"reasoning_tokens":1514,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:12:43.826340+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the controller on a compliant or loosely-tensioned strap so the peeled segment visibly sags, and record whether the particle-filter hinge estimates diverge from the true peel line while the success rate drops below the reported 100%: if they do, the tautness premise is the failure point.","supporting_citations":[{"cited_title":"Multi-step recurrent q-learning for robotic velcro peeling,","cited_arxiv_id":null,"evidence_quote":"Prior RL Velcro-peeling work that this method improves upon; provides the motivation and the RL baseline comparison."}],"review_version":1}