{"id":"91f83ab8-1078-4342-a35a-2dddc1d6e230","arxiv_id":"2512.07248","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A physics-based score (MDS) predicts how hard a motion is for a humanoid to imitate by measuring how much joint torques must change under small pose perturbations.","lead":"This paper introduces a score that estimates how hard a motion is for a simulated humanoid to imitate, before any policy is trained, by measuring how violently required joint torques change under small pose perturbations. It could make imitation-learning benchmarks fairer and help build curricula and clean up motion-capture data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MDS is not policy-free as defined: Eq. 12 computes floating-base torques with no contact/ground-reaction model, so the metric depends on an unspecified modeling choice.","rationale":"The paper's central claim is that MDS is a policy-free, physics-grounded scalar that explains and predicts imitation error. The strongest obstacle to that claim is not the empirical correlation itself, which the paper supports with a large sample and two strong baselines, but the definition of the quantity being correlated. Eq. 12 treats all N generalized coordinates as torque-actuated, yet a floating-base humanoid has unactuated base coordinates whose generalized forces must originate from contacts. The paper does not state how contacts are handled, what f_ext is, or how RBDL was configured for floating-base inverse dynamics. Because mocap motions are not dynamically consistent, the computed per-frame torques—and therefore the torque-space volume, variance, temporal variability, and final MDS—depend on modeling choices that are not disclosed. The correlation with imitation error could survive many of these choices, in which case the concern would be a reproducibility gap rather than a conceptual failure; but it could also vanish under a physically correct contact model, which would mean MDS is not measuring what the paper claims. Either way, the current manuscript does not provide enough information to determine which is true. I therefore agree with the reader that the weakest assumption is around Eq. 12, and I recommend keeping the verdict CONDITIONAL: the authors should specify the contact model and demonstrate robustness of MDS rankings and error correlations across plausible inverse-dynamics conventions. I am not arguing the central idea is false; the paper provides a plausible mechanism and real empirical associations. But the load-bearing definitional gap must be closed before the metric can be accepted as intrinsic or policy-free.","tokens_in":17604,"tokens_out":5207,"duration_ms":54161,"concrete_test":"Recompute MDS for 500 clips spanning the MDS range under three inverse-dynamics variants: (i) free-floating with f_ext=0 and root generalized forces included; (ii) floating-base with fixed-foot contact constraints and contact forces solved via rigid-contact inverse dynamics; (iii) an alternative but plausible humanoid mass distribution. Compare per-clip MDS rankings (Spearman/Kendall) and correlations with UHC/PHC+ MPJPE-G. If the variants disagree materially (e.g., rank correlation below 0.8) or only one variant reproduces the observed error correlations, MDS is an artifact of the unspecified contact/mass model. If all variants preserve rankings and correlations, the concern does not land.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Definition 4.1 and Eq. 12 define motion difficulty through τ=F(S)=M(q)q̈+h(q,q̇)−f_ext. But for a floating-base humanoid, the root DoFs (r_root, base orientation) are not actuated; the standard equation is M(q)q̈+h = S^T τ + J_c^T f_c, with contact forces f_c unknown. The paper specifies neither the contact model nor f_ext. Mocap motions are generally not dynamically consistent, so the 'unique torque' asserted in Sec. 4.1 is not unique: identical joint torque commands can produce the same observed motion with different ground-reaction forces, and computed base generalized forces are artifacts of the inverse-dynamics convention (for example, treating all N=3+3J DoFs as actuated). The cited mass-distribution protocol from [40] does not determine contact forces. Hence MDS measures torque variation under an unstated modeling choice, not an intrinsic property of the motion. This also undermines the reward-landscape argument: RL policies command joint torques (UHC additionally uses residual root forces; PHC+ does not), while Eq. 12 includes root forces. Without knowing which torque space is being perturbed, the claim that high MDS implies flat reward landscapes is unsupported. The undisclosed 'empirically derived' weights in Eq. 9 additionally make the reported correlations hard to interpret, but the contact-model gap is prior: if the base quantity is not well-defined, no choice of weights can fix it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Motion Difficulty Score (MDS), a policy-independent metric that measures a motion's intrinsic imitation difficulty through the torque variation induced by small pose perturbations. MDS is computed as a weighted combination of three terms: Spectral Diversity, Variance Diversity, and Segment Diversity, and is used to build MD-AMASS, a difficulty-labeled repartitioning of AMASS. The authors validate MDS by showing correlations with imitation error of UHC and PHC+, and introduce two derived metrics: Maximum Imitable Difficulty (MID) and Difficulty-Stratified Joint Error (DSJE). They also report applications to curriculum learning and flawed-motion detection.","tokens_in":17950,"tokens_out":2852,"duration_ms":32937,"significance":"If validated, MDS would be a practically valuable tool: it offers a policy-free scalar for comparing motions across training conditions, enables difficulty-stratified evaluation, and could support curriculum learning and mocap curation. The paper provides a clear dynamical intuition, an explicit computational pipeline, and a large-scale empirical study (MD-AMASS, >30,000 clips). The correlation analysis, ablations, and cross-robot generalization study are useful steps. However, the central claim rests on several under-specified and potentially circular empirical choices; these need to be resolved before the metric can be trusted as a predictive, architecture-independent difficulty measure.","major_comments":[{"comment":"The inverse-dynamics definition of torque is underdetermined for a floating-base humanoid. Eq. (12) defines τ = M(q)q̈ + h(q,q̇) − f_ext as a unique per-frame torque, but for a floating base the root degrees of freedom are not actuated and contact/ground-reaction forces are unknown. The paper neither specifies a contact model nor states how f_ext is computed. Since UHC applies residual root forces while PHC+ does not, the torque space being measured may not match the policy's action space. This makes MDS a function of an unstated modeling choice rather than an intrinsic motion property. The authors should specify the contact model, clarify which DoFs are considered actuated, and show MDS is robust to these choices.","section":"§4.1, Eq. (12), Appendix A"},{"comment":"The aggregation weights w_i in Eq. (9) are described only as 'empirically derived', with no fitting procedure, no reported values, and no validation split. The claimed correlations in Table 1 may therefore be in-sample fits rather than predictions. The authors should state the weight-selection method, report the weights, and evaluate MDS on a held-out set of motions or with cross-validation. The perturbation radius ε in Eq. (14) is also never stated; its value directly affects all three diversity terms and should be reported and varied in sensitivity analysis.","section":"§4.2.4, Eq. (9); §5.1"},{"comment":"The MID definition in Eq. (10) selects the threshold c that maximizes the error gap on the very same policy data that MID is then used to characterize. This is circular: the 'onset of performance collapse' is a fitted maximum, not an independent boundary. The resulting MID values (UHC: 308.22, PHC+: 320.50) are unsurprisingly ordered and provide no statistical evidence of a real capability difference. The authors should evaluate MID on held-out motion clips or use a pre-defined threshold selection rule (e.g., cross-validated or based only on MDS, not on error).","section":"§6.1, Eq. (10); Fig. 4"},{"comment":"The validation is correlational and uses samples drawn from the policies' training sets, so it does not establish that MDS predicts generalization to unseen motions. The exclusion of 'extreme outliers' (MPJPE-G>250 and MDS>350, fewer than 100 clips) is post hoc and its effect on the reported correlations is not quantified. The authors should report correlations with and without exclusions, and ideally evaluate on a held-out subset of AMASS that was not used in policy training.","section":"§5.1, Fig. 4; Table 1"}],"minor_comments":[{"comment":"Minor typos: 'Maximun' in Fig. 1, 'Temprol' in Fig. 2, 'mimicing' in §3, 'polices' in §4.2.4.","section":"Throughout"},{"comment":"The reference to AMASS appears as 'AMASS dataset []'—the citation placeholder needs to be filled.","section":"§4.3"},{"comment":"Notation is inconsistent: Eq. (1) uses τ ∈ R^N, while Definition 4.1 and Appendix A use Y = R^{J×t}; the state space X = R^{N×3×t} is also hard to reconcile with per-frame states s_i ∈ R^{3N}. Please clarify dimensions.","section":"§4.2 / Appendix A"},{"comment":"K=4 is chosen empirically with no sensitivity analysis. Since Segment Diversity depends on K, the authors should report how MDS and the correlations change for other K.","section":"§4.2.3"},{"comment":"The curriculum learning results in Table 4 are presented in the appendix but not discussed in the main text; the authors should either move this result to the main text or clarify its status.","section":"Appendix C.1"}],"recommendation":"major_revision","confidential_remarks":"The core idea is promising and the experimental effort is substantial, but the current form has two load-bearing gaps: an unspecified contact model in the definition of MDS and an undisclosed/possibly circular weight and threshold selection procedure. Both are fixable in a revision. I would not recommend rejection, but the empirical claims need to be re-run with explicit modeling choices and out-of-sample validation before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the MDS idea is genuinely new and domain-appropriate — measuring motion difficulty by how much inverse-dynamics torques vary under small pose perturbations is the right kind of object for this problem. The paper does something useful with it: MD-AMASS, the difficulty-stratified repartitioning, and the preliminary curriculum learning result in the appendix are concrete payoffs. Correlations with UHC/PHC+ error are real, especially Spearman 0.82 on PHC+, and the appendix scales the scatter to 10,000 clips. Credit where due.\n\nSoft spots, in order of severity. First, the metric is not policy-free as defined. Eq. 12 writes τ = M(q)q̈ + h − f_ext for all N DoFs, including the floating base. For a floating-base humanoid the root DoFs aren't directly actuated; the equation should involve a selection matrix and contact forces J_c^T f_c. The paper never says how f_ext is obtained or which contact model is used. Mocap motions aren't dynamically consistent, so the computed 'unique torque' is an artifact of an unstated modeling choice. That doesn't kill the idea — you could pick a contact model, say fixed feet or a soft-contact model, and re-derive MDS — but as written the central scalar is not well-defined. This problem precedes the other issues.\n\nSecond, the aggregation weights in Eq. 9 are 'empirically derived' with no held-out validation. The paper reports correlations on the same data used to set weights; that's in-sample. Third, the perturbation radius ε in Eq. 14 is never given, and Fig. 4 excludes outliers post hoc with no sensitivity analysis. Fourth, the MID threshold is defined as the argmax over the error data it then explains, so it's a descriptive split, not evidence of a regime change.\n\nSome minor issues: the appendix claims extensive validation with over 10,000 samples, but the main text shows 3,000; both are fine, but the discrepancy should be reconciled. The flawed-motion detection result is suggestive but qualitative.\n\nOverall: the central argument is plausible and the empirical association is probably real, but the paper currently overclaims 'policy-free' and 'principled' for a quantity that depends on unspecified modeling choices and in-sample fitted weights. These are fixable with disclosure and held-out analysis. I'd send it to review — the idea deserves referee time — but I'd expect major revision.\n\nRecommendation: accept for peer review, with the contact-model specification and held-out validation as non-negotiable conditions.","headline":"A genuinely new difficulty metric with a real contact-model gap and undisclosed fitting details; worth refereeing but needs hard revision.","tokens_in":18444,"tokens_out":2527,"would_cite":true,"duration_ms":25394,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the intrinsic difficulty of imitating a motion can be measured, independently of any policy, by the torque variation induced by small pose perturbations — larger torque swings flatten the reward landscape and make a mo","keywords":["motion imitation","reinforcement learning","rigid-body dynamics","motion difficulty","torque variation","humanoid control","reward landscape","curriculum learning"],"falsifier":"Compute MDS for the same motion set under two different contact models (e.g., full ground-reaction-force optimization versus simple residual-force compensation) and check whether the difficulty rankings, not just the scale, change; if the rankings scramble, the 'intrinsic' claim fails. Alternatively, hold out 20% of motion clips, fit the MDS aggregate weights on the rest, and test whether the correlation with imitation error survives on the held-out clips.","tokens_in":17489,"feed_emoji":"🤖","tokens_out":4115,"duration_ms":39943,"temperature":0.7,"pith_summary":"This paper argues that how hard it is to imitate a motion is not just a property of the learning algorithm — it is a measurable property of the motion itself. The proposed Motion Difficulty Score (MDS) measures the torque variation triggered by small pose perturbations: motions that demand wildly different torques for nearly identical poses give reinforcement learning a flat reward landscape, so they are intrinsically hard to imitate. The paper validates MDS on large-scale motion data with two state-of-the-art imitation policies, showing that MDS strongly correlates with imitation error. If this holds, evaluation of humanoid control can separate policy failures from motion-inherent challenges, and motion datasets can be curated and stratified by a physics-grounded difficulty label.","feed_headline":"Torque swings under tiny pose changes reveal a motion's true difficulty","feed_subtitle":"A physics-based score separates policy failures from motions that are intrinsically hard to learn.","key_machinery":"The central object is the Motion Difficulty Score (MDS), a policy-free scalar computed from the rigid-body inverse-dynamics map τ = M(q)q̈ + h(q,q̇) − f_ext. The paper perturbs each frame's state inside a small ball, maps the perturbed neighbourhood through the dynamics to a torque set, and characterizes that set by its volume (via singular values of the Jacobian), its joint-wise variance, and its temporal variability across four segments. The score is the weighted sum of these three diversities; the weights are empirically derived.","core_discovery":"The central claim is that imitation difficulty can be defined independently of any policy as the magnitude of torque variation induced within a bounded pose-error neighborhood, per Definition 4.1. High torque-to-pose variation collapses the reward landscape into a sharp spike surrounded by a plateau, so gradient-based reinforcement learning receives almost no directional signal. MDS operationalizes this via three complementary terms — spectral diversity (log-volume of the feasible torque space, derived through the coarea formula), variance diversity (per-joint variation of the Jacobian), and segment diversity (temporal uniformity of spectral diversity) — aggregated into a single scalar. Expe","pith_inferences":["The same principle could drive reward shaping or exploration heuristics: agents could be biased toward low-MDS regions early and toward high-MDS regions once the plateau is navigable — an extension the paper's curriculum experiment only begins to explore.","Because MDS depends on the chosen mass distribution and contact handling, the same motion could receive different scores on different morphologies; the paper's retargeting experiment suggests the ordering is stable, but the score's invariance is not proven.","The reported correlations are computed on training data; a stronger test would freeze the aggregate weights and measure correlation on a held-out set of motion clips, which the paper does not report.","If MDS truly captures landscape flatness, it should also predict sample efficiency rather than only final error — a testable extension that would connect the metric to learning-curve dynamics."],"forward_implications":["MDS gives a quantitative answer to 'why did the policy fail?' — high error on low-MDS motions points at the policy, while high error on high-MDS motions reflects an intrinsic ceiling.","MID (Maximum Imitable Difficulty) locates the difficulty threshold beyond which a policy's error explodes, turning a scatter plot into a single robustness boundary.","DSJE (Difficulty-Stratified Joint Error) exposes difficulty-regime reversals invisible in aggregate error, such as a policy that beats another overall but loses on easy motions.","MDS-based curriculum learning improves final policy performance across all difficulty groups in the paper's experiments.","MDS flags corrupted or physically implausible motion sequences by assigning them unusually high difficulty, enabling automated mocap quality control."],"fun_headline_variants":["Torque variation score reveals if failure is policy or motion","New metric separates policy flaws from hard-to-learn motions","Torque variation under small perturbations flags intrinsic difficulty","Why some motions fail to imitate: it's the torque, not the policy","Torque variation score pinpoints when imitation failure isn't policy's fault"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The score is computed from torques obtained via Eq. 12, but for a floating-base humanoid the inverse-dynamics problem requires a ground-contact model that the paper never specifies; if mocap motions are not dynamically consistent, the torque magnitudes — and therefore the score — could depend on arbitrary modeling choices. Also, the aggregation weights in Eq. 9 are empirically derived with no validation split, so the reported correlations could be partly in-sample.","fun_headline_variants_meta":{"raw":{"variants":["Torque variation score reveals if failure is policy or motion","New metric separates policy flaws from hard-to-learn motions","Torque variation under small perturbations flags intrinsic difficulty","Why some motions fail to imitate: it's the torque, not the policy","Torque variation score pinpoints when imitation failure isn't policy's fault"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00054,"raw_usage":{"total_tokens":2434,"prompt_tokens":762,"completion_tokens":1672,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":1587}},"tokens_in":506,"tokens_out":1672,"duration_ms":9728,"temperature":1.0,"reasoning_tokens":1587,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T17:58:53.176405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute MDS for the same motion set under two different contact models (e.g., full ground-reaction-force optimization versus simple residual-force compensation) and check whether the difficulty rankings, not just the scale, change; if the rankings scramble, the 'intrinsic' claim fails. Alternatively, hold out 20% of motion clips, fit the MDS aggregate weights on the rest, and test whether the correlation with imitation error survives on the held-out clips.","supporting_citations":[],"review_version":1}