{"id":"3e42fdbc-0704-42c3-b988-3f13b646fa2e","arxiv_id":"2608.11769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A humanoid pick-and-place policy exhibits initial-pose-dependent hand preference, which can be suppressed or induced by swapping two specific arm joints, and broader initial-pose training data reduces this dependence.","lead":"This paper shows that robot policies with vision and language can develop a starting-pose-dependent habit of choosing the wrong hand, and that changing just two arm joints in the starting pose can turn that habit on or off. It offers a new score to measure this hand preference and shows that training from more varied starting poses makes the robot more robust.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal handle on R_arm,3/R_arm,5 is not established: the joints were selected post hoc as the largest pose differences, and the control is unmatched in perturbation size, so the intervention may show only that large joint changes flip the prior.","rationale":"The reader's weakest assumption identifies post-hoc joint selection and statistical reliability as the main threat to the causal-handle conclusion; I agree with that concern and sharpen it: the post-hoc selection is not merely a statistical issue, because the chosen joints are also the only joints with large position differences, so the control condition does not match the treatment in perturbation magnitude. That confound is sufficient to invalidate the strong causal-reading of Tables III-IV. However, the paper's descriptive contributions do not depend on this causal claim: the pose-dependence maps in Fig. 2, the diversification result (5.8% to 63.3% for A1), and the targeted augmentation result (30% to 75%) are larger-scale, mechanistically coherent, and plausible. The reader's CONDITIONAL verdict is therefore appropriate, and my read does not move it: the causal handle should be re-verified before the central claim is accepted, but the paper is not unsupported. My agreement is 'partial' because I frame the decisive weakness as a selection/perturbation-magnitude confound rather than purely as lack of confidence intervals.","tokens_in":13562,"tokens_out":4291,"duration_ms":46473,"concrete_test":"Re-run the Section V-E bidirectional swap protocol treating every right-arm joint identically in perturbation magnitude: for each of the seven right-arm joints, move that joint from its teleop default value to its right sim value (i.e., apply a perturbation comparable to the 1.27-1.39 rad changes used for R_arm,3/R_arm,5) while holding all other joints fixed, and collect at least 40 left-target rollouts per condition with Wilson confidence intervals. If only R_arm,3 and R_arm,5 suppress the right-hand prior (baseline ~87% to near 0%) while all other equal-magnitude single-joint perturbations leave it high, the causal handle is confirmed; if several joints or a matched control pair also flip the behavior, the Table III result is a perturbation-magnitude artifact rather than a joint-specific causal handle.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal claim (Section V-E, Tables III-IV) is that the local configuration of R_arm,3 and R_arm,5 is a 'causal handle' on the pose-conditioned hand prior. The evidence has two coupled problems. First, the two joints were selected post hoc as the largest position differences between teleop default and right sim (1.269 and 1.389 vs. less than 0.39 for all other right-arm joints). The control swap uses R_arm,0/R_arm,6, whose swapped values differ by less than 0.39. The treatment therefore differs from the control in both joint identity and perturbation magnitude; a non-specific result (any sufficiently large joint perturbation flips hand choice) would look identical to the reported pattern. Second, the sample sizes are too small to support the claim: the decisive forward swap was tested on only 4 rollouts (0/4), and the 95% Wilson upper bound for 0/4 is about 60%, so the result is consistent with a substantial residual right-hand rate. The reverse and cross-policy interventions use 7-14 episodes and are inconsistent (B1 reverse swap: 0/12 to 0/12; C2 reverse swap: 2/7 to 4/10). The paper's own concluding paragraph acknowledges 'small and unequal sample counts.' Until joint selection is pre-specified or an equal-magnitude control is used and sample sizes are increased, the causal-handle claim is not supported; the descriptive hand-prior results and the coverage/augmentation findings are independent of this weakness and remain plausible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies how the initial configuration of a Unitree G1 humanoid affects early hand selection in a vision-language-action policy for a PickApple task. The authors define a HandPriorScore based on normalized relative joint actions in an initial decision window, decompose it into a residual bias term and a target-responsiveness term, and evaluate nine policy variants across 17 initial poses. They report strong initial-pose–policy interactions, show that wrist-camera observations modulate hand preference and task performance, and demonstrate that expanding initial-pose coverage from two to eight training poses, plus targeted augmentation around a low-performing pose, substantially improves success rates. The central causal claim is that swapping the two most divergent right-arm joints (R_arm,3 and R_arm,5) between the teleop_default and right_sim poses suppresses or induces an asymmetric right-hand preference, identifying a localized 'causal handle' on the pose-conditioned hand prior. The paper also examines how auxiliary real data, wrist observations, and simulation-ratio reweighting affect robustness.","tokens_in":1676,"tokens_out":1525,"duration_ms":38342,"significance":"The descriptive part of the paper is a useful and fairly complete empirical characterization: it quantifies pose-conditioned hand bias in a way that goes beyond aggregate success, shows that pose coverage and targeted augmentation are effective mitigation levers, and demonstrates that wrist-camera observations have interpretable but partial effects on hand selection. The pose-diversification result (A1 from 5.8% to 63.3% on six evaluation poses) and the targeted augmentation result (30% to 75% at right_real) are concrete and actionable findings. The causal-handle claim, however, is the paper's strongest advertised contribution and it is not yet supported by the evidence as presented, because the joint selection is post hoc, the control is unmatched in perturbation magnitude, and the decisive intervention uses very small sample sizes. If the causal claim can be backed by a pre-specified or equal-magnitude control and larger samples, the paper would make a strong contribution; the rest of the findings are independently plausible and interesting.","major_comments":[{"comment":"The causal-handle conclusion is not established because the two swapped joints were selected post hoc as the largest position differences between teleop_default and right_sim (1.269 and 1.389 vs. less than 0.39 for all other right-arm joints), while the control swap involved joints with differences less than 0.39. The treatment therefore differs from the control in both joint identity and perturbation magnitude, so the observed pattern is also consistent with a non-specific effect in which any sufficiently large joint perturbation flips the prior. An equal-magnitude control (e.g., swapping R_arm,0 and R_arm,6 at a comparable displacement) or a pre-specified selection criterion is needed before the paper can claim that the local configuration of R_arm,3 and R_arm,5 is a causal handle.","section":"Section V-E, Tables III–IV"},{"comment":"The decisive forward swap rests on four rollouts (0/4); the 95% Wilson upper bound for 0/4 is about 60%, so the result is compatible with a substantial residual right-hand rate. The reverse swap (7/7) and the cross-policy rows in Table IV (7–14 episodes) also lack confidence intervals, and the paper's concluding paragraph acknowledges 'small and unequal sample counts.' Without larger samples or interval estimates, the bidirectional suppression/induction claim is not statistically supported.","section":"Section V-E, Table III"},{"comment":"The reverse induction is inconsistent across policies (B1 0/11 to 0/12; C2 2/7 to 4/10), so the paper's own data show that the effect is policy-dependent in the reverse direction. The forward direction (4/14 to 0/12; 11/13 to 1/9) is more suggestive but still based on small counts without interval estimates. I recommend softening the language from 'causal handle' to 'localized sensitivity' unless the forward effect is replicated with larger samples.","section":"Section V-E, Table IV"},{"comment":"The four-joint intervention for pose08–teleop_default changes four joints at once, including two left-arm joints, so the conclusion that 'the prior-removal effect transferred across both policies' is not tied to R_arm,3/R_arm,5 specifically in the cross-policy setting. The paper should clarify whether the cross-policy result is meant to support the same two-joint causal claim or a broader sensitivity to large arm-configuration changes.","section":"Section V-E, 'Joint-level intervention' and Table III"}],"minor_comments":[{"comment":"The values of the hand-activity weight lambda, the stability constant epsilon, and the decision-window length T=120 are not specified in the text; please report them and, ideally, a sensitivity check showing that H_prior and H_target do not change qualitatively under reasonable variations.","section":"Section III-B"},{"comment":"The normalization scale sigma is estimated from the full-observation evaluation rollouts and then applied to those same rollouts. This in-sample normalization is acknowledged, but a leave-one-policy-out or leave-one-pose-out normalization would strengthen the claim that the cross-policy comparisons are not an artifact of the shared scale.","section":"Eqs. (3)–(4) and the surrounding text"},{"comment":"Success is manually assessed, which can introduce bias in a study whose central claims are about specific pose-dependent failures; please state whether the assessment was blind to policy and pose, or consider an automated success criterion.","section":"Section IV-C and Fig. 4"},{"comment":"Entries such as 'N(45)' and 'alpha(45.5)' are not defined in the table caption; please spell out that 'N' denotes the natural simulation ratio determined by dataset size and that 'alpha' denotes the reweighted sampling ratio, with exact formulas or examples.","section":"Table II"},{"comment":"The reference appears to contain a typo ('Unifolm g1 dex3 dataset') and should be corrected to the Unitree G1-Dex3 collection name. The same reference also lacks a URL that is clearly usable in the printed text.","section":"Reference [17]"},{"comment":"The figure is dense, with many small policy labels and repeated legends; separating the success heatmap from the two hand-prior panels or adding per-polygon annotations would improve readability.","section":"Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The paper's descriptive and mitigation findings are solid and likely publishable after revision. The main risk is the causal-handle claim: the joint selection is post hoc and the control is unmatched in magnitude, so the claim as written is not yet convincing. I would ask for an equal-magnitude control or pre-specified joint selection, larger samples with confidence intervals for Tables III–IV, and a softening of the causal language unless the evidence improves. The rest of the paper does not need major restructuring."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candid take: the descriptive core is solid and useful. The paper shows convincingly that a GR00T-based policy's success and early hand choice vary strongly across 17 initial poses, that a particular teleop-default pose induces a right-hand bias even on left targets, and that broadening initial-pose coverage in training (from two to eight poses) plus targeted augmentation around a weak pose lifts success from 30% to 75%. The HandPriorScore decomposition into residual bias and target responsiveness is a reasonable quantification, and the wrist-masking analysis, while interpretable only as sensitivity, is carefully hedged. These contributions are real and will be of interest to anyone working on VLA or humanoid manipulation.\n\nThe soft spot is the causal-handle claim in Section V-E. The two joints R_arm,3 and R_arm,5 were selected post hoc because they had the largest position differences between teleop_default and right_sim. The control swap uses smaller-difference joints, so the treatment and control differ in both joint identity and perturbation magnitude. With the decisive forward swap at 0/4 and the reverse at 7/7, no error bars and no pre-specification, the result is consistent with chance or with a generic 'large joint perturbation flips the prior' effect. The paper honestly notes the small samples in the conclusion, but the abstract still presents the causal handle as a main finding. That overstates the evidence.\n\nThe cross-policy four-joint swap is also underpowered and inconsistent on the reverse direction. So I'd treat the localization as suggestive, not established. The descriptive and mitigation results do not depend on this causal claim and remain plausible.\n\nWho is this for? Researchers evaluating VLA policies under varied initializations, and humanoid manipulation folks thinking about data coverage. It deserves a serious referee, but the causal-handle conclusion should be softened or supported with more rollouts and an equal-magnitude control. I'd send it to review and ask for heavy revision on Section V-E.","headline":"Useful descriptive study of initial-pose dependence in VLA humanoid manipulation, but the causal-handle claim is not established by the evidence.","tokens_in":14402,"tokens_out":1889,"would_cite":false,"duration_ms":18010,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A vision-language-action policy for humanoid dual-arm manipulation develops an initial-pose-dependent hand prior, and two specific right-arm joints can suppress or induce that preference.","keywords":["vision-language-action policy","hand prior","initial-pose robustness","dual-arm manipulation","humanoid robot","proprioception","data augmentation","causal intervention"],"falsifier":"Repeat the $R_{\\mathrm{arm},3}$ / $R_{\\mathrm{arm},5}$ swap at the teleop-default pose with at least 50 rollouts per condition, pre-registering which joints to swap and reporting confidence intervals; if the right-hand attempt rate is not significantly lower under the targeted swap than under the control swap, the causal-handle claim is refuted.","tokens_in":13336,"feed_emoji":"🤖","tokens_out":6449,"duration_ms":58711,"temperature":0.7,"pith_summary":"This paper argues that a vision-language-action (VLA) policy trained for humanoid dual-arm manipulation develops a 'policy-induced hand prior': the robot's initial configuration biases which hand the policy commits to during the first seconds of a task, even before the target object's location is fully used. The authors show that aggregate success rates hide this effect, since the same policy can succeed at one starting pose and fail at a nearby one, while different policies behave differently at the same pose. They provide evidence that a localized part of the initial arm configuration, involving two joints of the right arm, can suppress or induce the asymmetric hand preference, and that targeted training-data coverage around a failing pose repairs the failure. If the claim holds, evaluating and training humanoid manipulation policies should account for initial-pose coverage rather than relying on average task success alone.","feed_headline":"Two arm joints can flip a robot's hand preference","feed_subtitle":"Starting posture biases which hand a humanoid policy uses; adding nearby training poses fixes the worst failures.","key_machinery":"The load-bearing instrument is the HandPriorScore, defined as $H^{(T)} = (A_R^{(T)} - A_L^{(T)}) / (A_R^{(T)} + A_L^{(T)} + \\epsilon)$, where $A_S^{(T)}$ is the average normalized commanded-motion magnitude for side $S$ during an initial decision window. It is decomposed into a residual hand bias $H_{\\mathrm{prior}} = (\\mu_\\ell + \\mu_r)/2$, which survives averaging over target locations, and a target responsiveness $H_{\\mathrm{target}} = (\\mu_r - \\mu_\\ell)/2$, which measures how strongly the preference follows the object. The paper pairs this diagnostic with two interventions: masked wrist-camera inputs to test modality sensitivity, and bidirectional joint-value swaps between behaviorally distinct initial poses to localize which physical coordinates drive the prior. These tools let the authors separate pose-driven bias from target-driven response and test causality by changing only the initial physical state.","core_discovery":"The central discovery is that the early hand choice of the policy is conditioned on the initial arm configuration in a way that can override the target object's location. At a pose called teleop default, the policy showed a strong right-hand bias even when the apple was on the left, and this bias was accompanied by weak target responsiveness in the HandPriorScore decomposition. Two specific right-arm joints, $R_{\\mathrm{arm},3}$ and $R_{\\mathrm{arm},5}$, differed most between this pose and a neutral one, and bidirectional swapping of just these two joint values changed the right-hand attempt rate from 7/8 to 0/4 in one direction and from 0/12 to 7/7 in the other, while control swaps of less divergent joints had no effect. A four-joint version of the swap transferred the bias-removal effect to two other policies. The paper reads this as evidence that the localized configuration of these joints acts as a causal handle on the pose-conditioned hand prior, while wrist-camera observations modulate but do not fully explain the behavior.","pith_inferences":["If the causal-handle claim generalizes, one testable implication is that initial-pose robustness could be engineered by decorrelating the initial-pose distribution from target location rather than by adding more data indiscriminately.","The HandPriorScore decomposition could serve as a model-agnostic diagnostic: policies trained with pose-target decorrelation should show higher $H_{\\mathrm{target}}$ and smaller pose-to-pose variance in $H_{\\mathrm{prior}}$, a direct checkable prediction beyond the paper's own experiments.","The fact that two joints can flip hand choice suggests the policy may rely on a low-dimensional proprioceptive subspace; if so, a small classifier could predict failure-prone hand choices from the initial pose alone, enabling pre-rollout warnings."],"forward_implications":["Aggregate success rates will continue to hide pose-specific hand-selection failures unless evaluations include multiple initial configurations and report hand-prior metrics alongside success.","Broadening initial-pose coverage, especially pairing every pose with both left and right target placements, can lift success from about 6% to 63% for a high-camera-only policy and reduce pose-target coupling.","When a specific pose fails, adding training data from nearby initial configurations can be an effective repair; targeted augmentation raised success at one failing pose from 30% to 75%.","Wrist-camera observations generally improve task performance but can strengthen a pose-conditioned hand bias, so adding them is not a universal fix for wrong-hand behavior.","Mixing real or auxiliary data into training helps only when it preserves enough exposure to the target simulation task and keeps observation formats consistent; otherwise robustness declines."],"supporting_citations":[{"why":"Supplies the base humanoid policy model whose checkpoint is fine-tuned for all experimental policies.","marker":"[6]"},{"why":"Motivates the robustness question by showing that initial-state and viewpoint perturbations degrade VLA performance.","marker":"[10]"},{"why":"Provides the vision-proprioception imbalance hypothesis that policies favor concise proprioceptive signals, which the paper extends to hand selection.","marker":"[11]"},{"why":"Supports the claim that policies can overfit to proprioceptive trajectories and that relative actions reduce this dependence.","marker":"[12]"},{"why":"Attributes inconsistent policy behavior to state-dominant fusion, the prior work most directly extended to the hand-prior mechanism.","marker":"[13]"},{"why":"Provides the pretrained model release used as the starting point for policy fine-tuning.","marker":"[14]"},{"why":"Defines the robot embodiment used in simulation and interventions.","marker":"[15]"},{"why":"Defines the seven-joint hand structure whose action magnitudes feed the HandPriorScore.","marker":"[16]"},{"why":"Provides the real-robot teleoperation task dataset groups used in the training-composition comparisons.","marker":"[17]"}],"fun_headline_variants":["Two arm joints flip a humanoid's hand preference","Swapping two joint angles reverses robot hand bias","Hand prior traced to two specific arm joint angles","Two-joint tweak flips which hand a robot uses","Pose-conditioned hand prior pinned to two arm joints"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The causal-handle conclusion assumes that the joint-swap results, which come from as few as 4 to 12 rollouts per condition and carry no confidence intervals, are not chance outcomes, and that the two joints tested were not selected after the fact in a way that guarantees a difference.","fun_headline_variants_meta":{"raw":{"variants":["Two arm joints flip a humanoid's hand preference","Swapping two joint angles reverses robot hand bias","Hand prior traced to two specific arm joint angles","Two-joint tweak flips which hand a robot uses","Pose-conditioned hand prior pinned to two arm joints"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1586,"prompt_tokens":1006,"completion_tokens":580,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":622,"tokens_out":580,"duration_ms":6218,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:26:41.445409+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the $R_{\\mathrm{arm},3}$ / $R_{\\mathrm{arm},5}$ swap at the teleop-default pose with at least 50 rollouts per condition, pre-registering which joints to swap and reporting confidence intervals; if the right-hand attempt rate is not significantly lower under the targeted swap than under the control swap, the causal-handle claim is refuted.","supporting_citations":[{"cited_title":"Libero-plus: A progressive robustness benchmark for visual-language-action models,","cited_arxiv_id":null,"evidence_quote":"Motivates the robustness question by showing that initial-state and viewpoint perturbations degrade VLA performance."},{"cited_title":"When would vision-proprioception policies fail in robotic manipulation?","cited_arxiv_id":null,"evidence_quote":"Provides the vision-proprioception imbalance hypothesis that policies favor concise proprioceptive signals, which the paper extends to hand selection."},{"cited_title":"Isaac GR00T: N1.5 release,","cited_arxiv_id":null,"evidence_quote":"Provides the pretrained model release used as the starting point for policy fine-tuning."},{"cited_title":"Unitree g1 humanoid robot,","cited_arxiv_id":null,"evidence_quote":"Defines the robot embodiment used in simulation and interventions."},{"cited_title":"Unitree dex3-1 power control dexterous hand,","cited_arxiv_id":null,"evidence_quote":"Defines the seven-joint hand structure whose action magnitudes feed the HandPriorScore."},{"cited_title":"Unifolm g1 dex3 dataset,","cited_arxiv_id":null,"evidence_quote":"Provides the real-robot teleoperation task dataset groups used in the training-composition comparisons."}],"review_version":1}