{"id":"57aa456b-5d80-4b67-933c-0718ba7df829","arxiv_id":"2507.09857","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AdvGrasp generates adversarial object deformations that increase gravitational torque and reduce wrench-space stability margins, often degrading robot grasp performance in simulation and on two real objects.","lead":"This paper introduces AdvGrasp, a method that subtly deforms an object's 3D shape so that a robot's grasp becomes harder to lift and more unstable. It matters because it offers a physically grounded way to stress-test robotic grasping systems beyond attacking neural networks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of systematic degradation is undercut by a proxy mismatch: Tables 1-2 show the optimized LC/GS does not consistently transfer to MinGF/MaxLM/MaxED, with multiple backfire cases in the paper's own data.","rationale":"The reader's weakest assumption identifies the same load-bearing concern I would raise: the paper optimizes Ferrari's LC and GS but evaluates with MinGF, MaxLM, and MaxED, and the tables contain non-negligible backfire cases. This is not a stylistic issue; it blocks the central claim that the attack 'systematically' degrades grasp performance. The optimization could be reducing a static wrench-space quantity while the dynamic simulation metric actually improves because the deformation changes contact geometry, friction coupling, or mass distribution in ways not captured by the proxy. The paper's real-world validation is additionally confounded by inserting lead balls into hollowed 3D-printed objects, which changes mass distribution differently from the simulated uniform-density deformation; however, the proxy mismatch is the more fundamental gap because it affects the simulated evidence directly. I do not think the core idea should be rejected: the physical perspective is plausible, and many table entries do move in the intended direction. But the current evidence supports only a conditional acceptance with a required empirical demonstration of the LC/GS-to-MinGF/MaxLM/MaxED transfer. Since the reader already returned CONDITIONAL, my stress-test does not change that verdict.","tokens_in":12477,"tokens_out":8174,"duration_ms":90467,"concrete_test":"For all 20 objects under AdvGrasp and the ALC/AGS ablations, report the optimized LC and GS values alongside MinGF, MaxLM, and MaxED, then compute rank correlations between ΔLC and ΔMinGF and between ΔGS and ΔMaxED. If the correlations are not significant and in the expected direction, or if the backfire cases show decreased LC/GS while MinGF also decreases, the assumed transfer fails and the 'systematic' claim should be withdrawn or replaced with a per-object characterization. A complementary check: verify that each gradient step actually reduces LC/GS; if the optimizer reduces its own objective yet the simulation metric moves the wrong way in several cases, the proxy mismatch is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The method minimizes static wrench-space metrics LC (Eq. 4) and GS (Eq. 6), but the headline claim is validated with different dynamic simulation metrics: MinGF, MaxLM, and MaxED. The transfer from one set to the other is assumed, not demonstrated. Tables 1 and 2 already contradict the word 'systematically': for two-finger MinGF, AdvGrasp makes the grasp worse for MUSTARD BOTTLE (22.6→20.7), TOY AIRPLANE (29.2→14.6), CAMEL (15.5→9.7), and SHAMPOO (21.9→14.0); for three-finger MaxED, CRACKER BOX (24→36), KNIFE (7→13), BAOKE MARKER (15→40), and CAMEL (23→27) improve rather than degrade. If adversarial objects are supposed to compromise grasp performance, these are failures of the central claim, not mere noise. The paper provides no correlation or per-object analysis linking decreased LC/GS to increased MinGF or decreased MaxLM/MaxED, and no error bars to rule out noise as the explanation for the successes. Thus the abstract's 'systematically degrades these two key grasping metrics' is unsupported by the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AdvGrasp, a framework that generates adversarial 3D object shapes by deforming an object to reduce two analytic grasp metrics: Ferrari's lift capability LC (Eq. 4) and grasp stability score GS (Eq. 6). The authors introduce AdvGrasp-20, a benchmark of 20 objects with two- and three-finger grasps, and evaluate attack success using simulation-based metrics MinGF, MaxLM, and MaxED in PyBullet, plus two real-world 3D-printed objects. The stated central claim is that AdvGrasp 'systematically degrades' grasp performance.","tokens_in":12749,"tokens_out":6479,"duration_ms":64315,"significance":"If the central claim held, this would be a valuable contribution: it addresses an underexplored class of physical adversarial attacks and provides a benchmark and real-world validation. The use of physically grounded attack objectives (LC/GS) rather than network scores is conceptually sound, and the paper presents per-object tables rather than only qualitative examples. However, the paper's own data do not currently support the word 'systematically': a non-negligible fraction of entries in Tables 1 and 2 move opposite to the claimed direction. The manuscript therefore needs substantial additional evidence or a sharpened claim before the contribution can be accepted.","major_comments":[{"comment":"The abstract and Section 7 state that AdvGrasp 'systematically degrades' grasp performance, but Table 1 shows multiple counterexamples: for two-finger MinGF, MUSTARD BOTTLE changes 22.6→20.7, RACQUETBALL 21.9→20.8, TOY AIRPLANE 29.2→14.6, CAMEL 15.5→9.7, and SHAMPOO 21.9→14.0, i.e., the minimal grasp force decreases (grasp becomes easier). Table 2 shows MaxED increases after AdvGrasp for several entries (e.g., two-finger CHIPS CAN 33→35; three-finger CRACKER BOX 24→36, KNIFE 7→13, BAOKE MARKER 15→40, CAMEL 23→27). These cases are not negligible noise; the paper provides no per-object analysis, no rate of 'successful' degradation, and no statistical test. To support the central claim, the authors should report the failure rate, add error bars/confidence intervals, and either refine the method or restrict the claim to the majority of cases with quantified confidence.","section":"§6.2, Tables 1–2"},{"comment":"The optimization objective is a sum of LC and GS (plus Laplacian regularization), but the evaluation never reports these optimized quantities. The abstract's phrase 'degrades these two key grasping metrics' is therefore not directly verified: LC and GS are not in Tables 1–2. Instead, the paper reports transfer metrics (MinGF, MaxLM, MaxED), and the mapping from reduced LC/GS to these simulation outcomes is assumed. I ask the authors to include a table or scatter plot of LC and GS before/after attack, and a correlation or per-object comparison showing that reductions in LC/GS translate to lower MinGF, lower MaxLM, and lower MaxED.","section":"§4.1–4.3, Eq. (8); §6.2"},{"comment":"The evaluation protocol is deterministic in description, but no repeated runs or seeds are reported; the increments (0.2 N force steps, 1 N disturbance steps, 0.1 kg mass steps) also discretize the metrics coarsely. For example, many MaxLM changes are 0.1–0.3, which is at the step granularity. Without variance estimates or a significance test, the observed improvements in some rows could be attributed to quantization or simulation noise. I request error bars over multiple optimization runs (or at least multiple simulation seeds) and a statistical comparison (e.g., paired test over objects) for the aggregate claim.","section":"§6.1, Evaluation Metrics"},{"comment":"The real-world validation reports only two objects (DABAO SOD and TOMATO SOUP CAN) with a single qualitative outcome (lifted versus dropped) and no repeated trials, force calibration details, or measured slip/pose values. This is too thin to support the abstract's 'robustness and practical applicability.' Either add quantitative repeated trials or soften the claim.","section":"Abstract; §6.3"}],"minor_comments":[{"comment":"The definition of LC is ambiguous: please define the set G(w_gravity) explicitly and specify whether the norm in the numerator and denominator is taken in the wrench space or the force space.","section":"Eq. (4)"},{"comment":"The notation Disp2s(0, S) is not defined; specify the distance function and the space in which the convex hull is computed.","section":"Eq. (6)"},{"comment":"The post-processing step retains only grasps that successfully lift objects, but the paper does not report how many of the initially generated grasps were filtered out per object; this could bias the benchmark toward easy grasps and should be stated.","section":"§5, Grasp Generation"},{"comment":"The table captions do not state units for MinGF, MaxLM, and MaxED; the units appear only in the text of Section 6.1, and repeating them in the captions would greatly improve readability.","section":"Tables 1–2"},{"comment":"The figure lacks axis labels, a definition of the thresholds for 'good' and 'bad' under each metric, and error bars; since the DNN-based and physical metrics use different failure criteria, the comparison is hard to interpret without this information.","section":"Figure 7"},{"comment":"The reference to Alharthi and Brandão [2024] lists the page range as '1907–1902', which is inconsistent; please correct it.","section":"References"},{"comment":"The experimental comparisons are limited to the three proposed variants (ALC, AGS, AdvGrasp); adding a comparison with existing adversarial grasp attacks (e.g., Alharthi and Brandão, and Wang et al.) would make the benchmark more informative.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically coherent and introduces a useful benchmark, but the central claim of systematic degradation is not supported by the presented data. The editor may wish to require a revision that either substantially strengthens the evidence (reporting LC/GS values, per-object statistics, error bars, correlation analysis) or narrows the claims to match the observed mixed results. I see no evidence of misconduct; the issue is evidentiary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first paper I've seen that attacks a specific grasp configuration by deforming the object's shape to target wrench-space metrics (Ferrari's lift capability and stability score). That's a real addition to the adversarial-object literature, distinct from Wang et al. 2019's universally hard objects and from network-level attacks. Second, the central claim that AdvGrasp 'systematically degrades' grasp performance is not backed by the paper's own data. The optimization minimizes LC and GS, but the evaluation uses different simulation metrics (MinGF, MaxLM, MaxED), and the transfer between the two is assumed, not demonstrated. Tables 1 and 2 show multiple backfire cases: for two-finger MinGF, MUSTARD BOTTLE goes from 22.6 to 20.7, TOY AIRPLANE from 29.2 to 14.6, CAMEL from 15.5 to 9.7, and SHAMPOO from 21.9 to 14.0—all improved after the attack. For three-finger MaxED, CRACKER BOX (24 to 36), KNIFE (7 to 13), BAOKE MARKER (15 to 40), and CAMEL (23 to 27) improve. That's not noise; it directly contradicts 'systematically.' The paper offers no correlation or per-object analysis linking lower LC/GS to worse transfer metrics, and there are no error bars.\n\nWhat the paper does well: it builds a clean benchmark (AdvGrasp-20) with 20 objects and mixed grasp planners, and it validates in simulation and on two real 3D-printed objects. The real-world experiments are suggestive, though the protocol is confounded—adding lead balls to reach 1 kg changes mass distribution, not just geometry, so the failure could be due to altered inertial properties rather than the shape deformation.\n\nMy overall read: the core idea is plausible and worth pursuing as a robustness benchmark, but the evidence currently supports a proof-of-concept, not a 'systematic' attack. A revision should either weaken the claim, demonstrate the proxy transfer, or report per-object correlations with error bars. I'd send it to peer review with the expectation of major revision. It deserves referee time because the angle is new and the benchmark will be useful to the community. I wouldn't cite it in its current form, but I'd bring it to a reading group to discuss the proxy-metric problem.","headline":"A new physically grounded attack idea with a solid benchmark, but the 'systematically degrades' claim is undercut by the paper's own tables showing the optimized metrics don't transfer to the evaluation metrics.","tokens_in":13312,"tokens_out":2338,"would_cite":false,"duration_ms":24189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deforming an object's shape can make robot grasps fail.","keywords":["adversarial attacks","robotic grasping","grasp stability","lift capability","wrench space","shape deformation","grasp quality metrics","AdvGrasp-20"],"falsifier":"A decisive check is to compute, for every object in AdvGrasp-20, the change in the two optimized analytic scores alongside the change in minimal grasp force, maximal lifting mass, and maximal external disturbance; the central claim would be falsified if a substantial share of objects show physical metrics staying flat or improving while the analytic scores drop. A second check is to 3D-print adversarial versions of additional objects and run the same lift-and-disturbance protocol used for the two reported physical objects.","tokens_in":12268,"feed_emoji":"🤖","tokens_out":10613,"duration_ms":105821,"temperature":0.7,"pith_summary":"This paper sets out to show that a robot's grasp can be broken by changing the object itself rather than the perception network. AdvGrasp deforms the object's mesh so that gravity produces more torque on the held object and so that the set of wrenches the fingers can exert moves closer to the edge of feasibility, degrading both lift capability and grasp stability. These two objectives are optimized together with a smoothing term, producing adversarial objects whose overall shape is preserved. The attack is evaluated in simulation on a 20-object benchmark and in physical tests with a real gripper, where deformed objects slip and drop while the original objects lift cleanly. This matters because prior attacks mainly fool the network that scores grasps, leaving the physical grasp itself untouched.","feed_headline":"Deforming an object's shape can make robot grasps fail","feed_subtitle":"AdvGrasp warps shape to boost gravitational torque and shrink stability margin; real gripper tests confirm dropped objects.","key_machinery":"The machinery is the wrench-space formulation of grasp quality. For each contact, a wrench combines force and torque about the object's centroid, and a grasp is stable when the origin lies inside the convex hull of the wrenches the contacts can generate; lift capability is the ratio of the gravitational wrench to the normal force needed to balance it. AdvGrasp places control points on a bounding box around the object and deforms the mesh through iterative mean-value-coordinate warping, using simulated annealing to minimize the combined objective while the Laplacian term suppresses implausible geometry. The deformation acts on exactly the two quantities that change with shape: the center of mass that sets the gravitational torque and the surface normals that set the direction of each contact wrench.","core_discovery":"On the paper's own terms, the discovery is that two analytic grasp-quality quantities can serve as adversarial objectives: lift capability, the minimal normal gripper force needed to balance gravity, and grasp stability, the distance from the origin to the convex hull of the wrenches the contacts can exert. AdvGrasp perturbs vertices near the contacts and elsewhere on the object to minimize a weighted sum of these two quantities plus a Laplacian regularization term. Reducing lift capability is achieved by shifting geometry so the gravitational wrench becomes harder to counter, and reducing stability is achieved by changing surface normals near the contacts so the wrench hull passes closer to the origin. In simulation the resulting adversarial objects require more force to lift, lower the maximum liftable mass, and survive smaller external disturbances; in physical tests a two-finger gripper drops the deformed objects while lifting the originals. The paper also reports that a learned grasp evaluator often still judges the adversarial grasps as good, while the physical metric marks them as bad.","pith_inferences":["A testable extension is combining this physical attack with an image- or point-cloud-based attack on the perception stage, since the two failures may compound end-to-end.","The reported tables show the analytic-to-physical transfer is not perfect for every object, so adversary success may be object-dependent and per-object correlations would clarify where AdvGrasp can be relied on.","The same wrench-space objectives could be inverted for defense, for instance by choosing contacts or adding surface features that increase the stability margin without changing the intended grasp.","Extending the real-world validation to all twenty objects under varied friction and gripper stiffness would map where the simulation-based transfer holds."],"forward_implications":["Adversarial objects can be generated without access to the robot's perception network, so the attack applies to analytic grasp planners as well as learned ones.","A grasp that succeeds on the original object can be made to fail by a localized deformation that preserves the object's overall shape.","The AdvGrasp-20 benchmark gives later work a fixed set of objects and grasp configurations for comparing physical attacks on two- and three-finger grippers.","Because a learned grasp evaluator often approves adversarial grasps that the physical metric rejects, grasp-quality networks may need physics-based training signals to remain trustworthy under attack."],"supporting_citations":[{"why":"Supplies the lift-capability ratio and stability-margin formulas that define the two attack objectives.","marker":"[Ferrari et al., 1992]"},{"why":"Supplies the source objects and the learned grasp detector used to create two-finger grasp configurations for the benchmark.","marker":"[Fang et al., 2020]"},{"why":"Supplies the dataset protocol used to select and validate objects and grasps from the same object set.","marker":"[Fang et al., 2023]"},{"why":"Supplies the analytic grasp planner whose grasps fill the two-finger half of the benchmark.","marker":"[Mahler et al., 2017]"},{"why":"Defines the soft-finger contact model with torsional friction used in simulation.","marker":"[Mahler et al., 2018]"},{"why":"Provides the mean-value-coordinate deformation scheme used to warp the object mesh.","marker":"[Ju et al., 2023]"},{"why":"Provides the simulator in which grasp success, slippage, and the evaluation metrics are measured.","marker":"[Coumans and Bai, 2016]"},{"why":"Provides the learned grasp evaluator whose good/bad judgments are compared with the physical metric.","marker":"[Liang et al., 2019]"}],"fun_headline_variants":["Shape sabotage breaks robot grasps","Adversarial object shapes compromise robotic grasping","Physical adversarial attacks cause robot grasp failures","Warp shapes to foil robot grippers","Deform objects to undermine grasp stability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that reducing the analytic lift-capability and stability scores will make simulated and physical grasps fail more often, even though the evaluation uses different physical metrics and the reported tables show this transfer is not clean for every object.","fun_headline_variants_meta":{"raw":{"variants":["Shape sabotage breaks robot grasps","Adversarial object shapes compromise robotic grasping","Physical adversarial attacks cause robot grasp failures","Warp shapes to foil robot grippers","Deform objects to undermine grasp stability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000555,"raw_usage":{"total_tokens":2611,"prompt_tokens":880,"completion_tokens":1731,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":1669}},"tokens_in":496,"tokens_out":1731,"duration_ms":13236,"temperature":1.0,"reasoning_tokens":1669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:45:09.061756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to compute, for every object in AdvGrasp-20, the change in the two optimized analytic scores alongside the change in minimal grasp force, maximal lifting mass, and maximal external disturbance; the central claim would be falsified if a substantial share of objects show physical metrics staying flat or improving while the analytic scores drop. A second check is to 3D-print adversarial versions of additional objects and run the same lift-and-disturbance protocol used for the two reported physical objects.","supporting_citations":[{"cited_title":"Planning optimal grasps","cited_arxiv_id":null,"evidence_quote":"Supplies the lift-capability ratio and stability-margin formulas that define the two attack objectives."},{"cited_title":"Graspnet-1billion: A large-scale benchmark for general object grasping","cited_arxiv_id":null,"evidence_quote":"Supplies the source objects and the learned grasp detector used to create two-finger grasp configurations for the benchmark."},{"cited_title":"Robust grasping across diverse sensor qualities: The graspnet-1billion dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the dataset protocol used to select and validate objects and grasps from the same object set."},{"cited_title":"Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics","cited_arxiv_id":null,"evidence_quote":"Supplies the analytic grasp planner whose grasps fill the two-finger half of the benchmark."},{"cited_title":"Dex-net 3.0: Computing robust vacuum suction grasp targets in point clouds using a new analytic model and deep learning","cited_arxiv_id":null,"evidence_quote":"Defines the soft-finger contact model with torsional friction used in simulation."},{"cited_title":"Pybullet, a python module for physics simulation for games, robotics and machine learning","cited_arxiv_id":null,"evidence_quote":"Provides the simulator in which grasp success, slippage, and the evaluation metrics are measured."},{"cited_title":"Pointnetgpd: Detecting grasp configurations from point sets","cited_arxiv_id":null,"evidence_quote":"Provides the learned grasp evaluator whose good/bad judgments are compared with the physical metric."}],"review_version":1}