{"id":"2bc9c3eb-5cc2-4574-a06b-7ff1e18bc91e","arxiv_id":"2608.13555","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"HumanTracker introduces a 153-hour categorized humanoid tracking benchmark and a preference-trained metric, HumanScore, that agrees with human judgments better than kinematic error metrics.","lead":"HumanTracker is a new 153-hour motion-capture benchmark and a learned score, HumanScore, for judging how well humanoid robots track human motions. It is designed so robot evaluations match what people actually see as unstable, sliding, or unnatural motion.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The alignment claim rests on single-judgment preference labels from six annotators with no inter-annotator reliability; if those labels are idiosyncratic, the reported Align Rate does not generalize to human preferences at large.","rationale":"The reader's weakest assumption is the same one I find most load-bearing: the entire HumanScore training and evaluation pipeline is anchored to preference labels that have no measured reliability. The paper is otherwise methodologically careful: the motion-disjoint split by source motion, the balanced tracker pairings, the masked window aggregation, and the bootstrap uncertainty estimates in Table 9 all support the internal validity of the experiments. The benchmark's scale and categorization are genuine contributions, and the paper is candid about the annotator pool and other limitations in Sec. G. Secondary concerns exist, such as the possible completion-selection confound in the Table 3 Ground results and the absence of comparisons against existing learned reward models, but these would matter most after label reliability is established. A focused inter-annotator reliability study is the single check that would settle whether the alignment numbers generalize beyond the six original annotators. Since this is the same concern the reader identified and it is addressable without invalidating the core contribution, the verdict remains CONDITIONAL, matching the reader's verdict.","tokens_in":15906,"tokens_out":10423,"duration_ms":112370,"concrete_test":"Re-annotate a random subset of roughly 200 test preference pairs with at least three independent annotators, including annotators outside the original six, and compute pairwise inter-annotator agreement (e.g., Cohen's kappa). Additionally, train HumanScore on labels from each original annotator separately and measure cross-annotator Align Rate on test pairs labeled by the others. If inter-annotator agreement is high (kappa above 0.6) and cross-annotator Align Rate remains near the reported 0.90, the concern does not land; if agreement is low or Align Rate drops toward the kinematic-baseline level, the central claim requires stronger label-quality evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the 6,000 original pairwise preference labels (mirrored to 12,000 records) collected from six doctoral researchers in Sec. 3.3 and Appendix C are a reliable, stable ground truth for human perception of tracking quality. Each pair receives one primary judgment, and no inter-annotator agreement or label-quality analysis is reported. HumanScore is trained and evaluated on these labels, so the Table 4 Align Rate (90.83% versus 84.05% for keypoint position MAE, with Table 9 bootstrap intervals) measures agreement with these six individuals, not with a broad population of viewers. The paper itself acknowledges in Sec. G that one primary judgment per pair 'does not quantify uncertainty through repeated independent labels.' If the six annotators are idiosyncratic, fatigued, or systematically biased toward a particular tracker's style, the learned metric can overfit those biases, and the central claim that HumanScore better predicts human preferences is substantially weakened. This is an addressable limitation, but until label reliability is demonstrated, the headline alignment numbers remain unvalidated as a general statement about human preference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HumanTracker, a large-scale benchmark for humanoid motion tracking containing approximately 153 hours of optical motion capture from 24 professional performers, organized into four motion families (Daily, Highly Dynamic, Interaction, Ground) with text labels and a standardized MuJoCo evaluation protocol. The paper also proposes HumanScore, a temporal Transformer reward model trained on 6,000 original pairwise preference judgments (mirrored to 12,000 records) collected from six doctoral researchers, and evaluated on a motion-disjoint test split. The central claim is that HumanScore better predicts human preferences than individual kinematic diagnostics such as MPJPE and keypoint-position MAE (Table 4: 90.83% vs 84.05% Align Rate), and that it reveals contact and stability failures that kinematic metrics miss. The paper reports bootstrap confidence intervals clustered by source motion (Table 9), sensitivity analyses of input features and temporal context (Figure 5), and a discussion of limitations including the single-judgment nature of the preference labels.","tokens_in":16279,"tokens_out":5984,"duration_ms":64831,"significance":"If the central claim holds, HumanTracker would be a valuable community resource: it provides a substantially larger and more diverse evaluation suite than the commonly used AMASS test set, a standardized protocol that controls rollout accounting and metric implementation, and a learned perceptual metric that can complement MPJPE-style diagnostics for contact-rich humanoid tracking. The paper has several concrete strengths: the motion-disjoint test split prevents adjacent clips from the same sequence crossing partitions; the bootstrap procedure in Appendix F clusters by source motion; the sensitivity analysis in Figure 5 directly tests the contribution of contact features; and the limitations section (Sec. G) is unusually candid about the scope of the claims, including the single-judgment annotation design and the restriction to one embodiment. These features make the benchmark itself reproducible and the evaluation protocol well specified. However, the headline preference-alignment result currently rests on label reliability and statistical evidence that are not fully established in the manuscript.","major_comments":[{"comment":"The preference ground truth is load-bearing for the central claim, yet it rests on one primary judgment per pair from six doctoral annotators, with no inter-annotator reliability check reported. Both the training labels and the held-out test labels come from the same six-annotator panel, so the Align Rate in Table 4 measures agreement with this specific panel rather than with human preferences broadly. The manuscript itself acknowledges in Sec. G that repeated independent labels were not collected. To support the claim that HumanScore is human-aligned, please report agreement on a double-annotated subset (e.g., Cohen's or Fleiss' kappa), annotator-level agreement rates, and ideally a leave-one-annotator-out training/evaluation analysis. Without such evidence, it is not possible to rule out that the reported 90.83% Align Rate reflects idiosyncratic or systematic biases of the six annotators rather than a general perceptual ground truth.","section":"Sec. 3.3, Appendix C, Sec. G"},{"comment":"The headline numerical advantage of HumanScore over the best kinematic diagnostic is not supported by the reported uncertainty. In Table 9, the 95% bootstrap interval for HumanScore is [87.36, 93.83] and the interval for KPT Position MAE is [79.67, 88.04]; these intervals overlap (between 87.36 and 88.04). Although overlapping marginal intervals do not automatically prove non-significance, the paper does not report a paired test of the difference in Align Rate. To establish that HumanScore is better than the best individual diagnostic, please provide a paired bootstrap or permutation test over source-motion clusters that directly tests the difference, or explicitly state that the current evidence does not show a statistically significant improvement. This is necessary because the 6.8 percentage-point difference is the quantitative basis for the paper's central claim.","section":"Table 9, Sec. 4.3"},{"comment":"The claim that HumanScore 'reveals contact and stability failures that kinematic metrics often miss' is not directly demonstrated in the experiments. Figure 5 shows that removing measured contact features degrades alignment most on Ground, and Table 3 shows that HumanScore ranks SONIC above Humanoid-GPT on Ground despite higher MPJPE. However, there is no qualitative or systematic analysis showing a specific contact or stability failure that HumanScore identifies and MPJPE misses. Please add a case study (e.g., video frames with contact and support annotations) or a quantitative error analysis of disagreements between HumanScore and MPJPE, to substantiate this component of the central claim.","section":"Abstract, Sec. 4, Fig. 5"}],"minor_comments":[{"comment":"Several entries in Table 3 are missing spaces between numbers, e.g., '97.60.128' and '0.23126.5'; please fix these formatting issues.","section":"Table 3"},{"comment":"The term 'zero-shot generalization' in the Introduction is stronger than what is demonstrated; the current evaluation is motion-level zero-shot only, as Sec. G correctly states. Please qualify the claim in the Introduction to avoid overstating the scope.","section":"Sec. 1, Sec. G"},{"comment":"The 'Foot Contact Accuracy' diagnostic in Table 4 is not defined in Sec. 3.2. Please specify how contact agreement is computed (e.g., frame-level binary accuracy, tolerance, and which contact states are compared).","section":"Sec. 4.1, Table 4"},{"comment":"The caption for Figure 5 refers to a 'baseline' that is not explicitly defined in the main text; please state that the baseline is the full model described in Sec. 3.4 and Table 6.","section":"Fig. 5"},{"comment":"Appendix C clarifies that the 12,000 preference records are mirrored variants of 6,000 original annotated pairs. The abstract's '12K motion pairs' could be misread as 12,000 independent judgments; please make the original/mirrored distinction explicit at first mention.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The annotator pool is described only as 'six doctoral researchers specializing in humanoid robotics'; it would be useful for the editor to know whether these annotators are independent of the author team and whether the annotation procedure received any ethical or consent review. This does not affect my technical assessment but is relevant for a benchmark paper that claims human-aligned ground truth."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read on HumanTracker. The headline: this is a genuinely useful benchmark plus a competent RLHF-style metric, and the central claim that HumanScore beats kinematic diagnostics is plausible, but the preference data is the weak load-bearing point and the paper doesn't compare against other learned reward models.\n\nWhat's actually new: the 153-hour four-family optical mocap benchmark with text labels is a real step up from the AMASS-140 test set. They make a sensible effort to standardize evaluation across trackers—same robot model, qpos reference, termination rule, metric implementation—while preserving each tracker's policy interface. That controlled protocol is valuable on its own. HumanScore is a reasonable application of Bradley-Terry reward learning to tracking evaluation, with a motion-disjoint test split and bootstrap CIs. The ablation showing contact features matter most on Ground and longer context helps is believable.\n\nThe soft spots, in order. First, the preference ground truth is six doctoral researchers, one judgment per pair, no inter-annotator agreement. The paper admits this in Sec. G but doesn't fix it. For a metric whose whole purpose is aligning with human perception, that's a real generalization gap. It's addressable—collect multiple labels on a subset, report agreement—but until then, the 90.83% vs 84.05% figure measures agreement with these six people, not with humanity. I don't think it's fatal; the motion-disjoint split and balanced pairing rules out some shortcuts, and expert annotators may be fine for domain-specific quality. But the claim in the abstract should be softened or backed with reliability data.\n\nSecond, no comparison against existing learned reward models. They cite RoboReward, Robometer, MotionCritic but only compare against kinematic diagnostics. If the contribution is a preference-aligned metric, you need to show it adds value over a generic pretrained reward model or at least explain why those don't transfer. That's a missing experiment, not a flaw in what they did.\n\nThird, the benchmark table numbers have no error bars. The preference CIs are there, but Succ and MPJPE are point estimates. Given that the benchmark is meant to be the standard, spread over rollouts matters.\n\nFourth, minor: the HumanScore values in Table 3 are very low across the board, which suggests the mapping is not well calibrated; it's fine for ranking but readers will ask what 54.7 versus 49.5 means.\n\nOverall, the paper is honest about limitations, the benchmark is well-documented, and the central argument mostly holds. This deserves a serious referee; I'd send it out with a request for annotation reliability analysis, learned-baseline comparisons, and error bars on benchmark results. I'd cite it if I worked in this area.","headline":"Solid benchmark and a plausible preference metric, but the six-annotator single-judgment ground truth and missing learned baselines keep the alignment claim from being fully settled.","tokens_in":16723,"tokens_out":2365,"would_cite":true,"duration_ms":23638,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces HumanTracker, a 153-hour motion-tracking benchmark, and HumanScore, a preference-trained metric that agrees with human judges 90.8% of the time versus 84.1% for the best kinematic diagnostic.","keywords":["humanoid motion tracking","preference-aligned metric","HumanScore","motion tracking benchmark","contact and stability evaluation","reward model","optical motion capture","zero-shot tracking evaluation"],"falsifier":"Re-run the preference alignment study with a fresh and larger annotation panel, giving each pair multiple independent labels and first measuring inter-annotator agreement; if HumanScore's agreement with those labels does not beat the best kinematic diagnostic, or if the original six annotators fail to agree with the new panel, the claim that HumanScore predicts human preferences collapses.","tokens_in":15720,"feed_emoji":"🤖","tokens_out":8258,"duration_ms":77096,"temperature":0.7,"pith_summary":"HumanTracker sets out to fix how humanoid motion tracking is judged. It argues that standard kinematic errors average pose differences and miss physical artifacts that people immediately notice, such as foot sliding, unstable support, and mistimed contacts. To close that gap, the paper builds a 153-hour, roughly 25K-clip benchmark of optical motion capture organized into four motion families, and proposes HumanScore, a learned score trained on 12K motion pairs with 24K motions. On a motion-disjoint test split, HumanScore agrees with human preferences 90.8% of the time, while the best individual kinematic diagnostic reaches 84.1%. If the claim holds, tracking evaluation can be reported per motion family with a perceptual metric alongside completion rate and joint error, exposing contact and stability failures that kinematic metrics miss.","feed_headline":"HumanScore beats pose-error metrics at matching human judges","feed_subtitle":"90.8% agreement with expert preferences vs 84.1% for the best kinematic diagnostic, on a 153-hour benchmark.","key_machinery":"The load-bearing object is HumanScore, a reward model built from a temporal Transformer. Each frame becomes a 539-dimensional token made of the current reference state plus simulated robot state, actions, measured contact dynamics, root motion, and keypoint kinematics; a padding mask confines attention and mean pooling to real frames, and the pooled token is mapped through an MLP to a scalar reward. Training uses a paired-comparison objective on strict preferences, a symmetric loss for similar pairs, and no cannot-compare pairs; at inference, window rewards pass through a sigmoid and are averaged with frame-length weights to produce a 0-to-100 score. The benchmark side of the paper is a standardized evaluation apparatus: one 29-DoF humanoid, a common simulator entry point, a common success criterion based on pelvis, ankle, and wrist vertical error plus pelvis rotation, and family-level reporting.","core_discovery":"HumanScore is a learned trajectory-level score trained to reproduce human comparisons of synchronized tracking rollouts. On the paper's test set, it reaches an Align Rate of 0.9083 with a 95% bootstrap interval from 0.8736 to 0.9383, while the best individual kinematic diagnostic, keypoint position MAE, reaches 0.8405. HumanScore's advantage is largest in contact-rich motions, and it scores Ground and Highly Dynamic rollouts differently than joint error would predict. The paper pairs this metric with a 153-hour benchmark of optical motion capture, retargeted to a 29-DoF humanoid and organized into four motion families, so that tracking quality can be reported per family instead of as one aggregate number. The authors' conclusion is that a preference-aligned metric and a categorized benchmark together reveal contact and stability failures that kinematic metrics miss.","pith_inferences":["A natural next test is applying HumanScore to real-hardware rollout trajectories; because its current input includes privileged simulator contact and force features, a real-robot study requires an estimated feature set, a boundary the paper itself flags.","Scaling the preference pipeline from six expert annotators to a larger, more diverse panel would test whether the 90.8% alignment is stable; a plausible outcome is that crowd preferences are noisier but still favor the same contact and stability cues.","Using HumanScore as a reinforcement-learning reward for tracking policies could improve perceived quality directly, but the paper warns that unregularized optimization of a learned score can be gamed; a controlled experiment with independent human evaluation would settle that.","The four-family taxonomy invites a failure-regime decomposition: attributing HumanScore to specific events such as impacts, support switches, and recoveries would turn a scalar score into a diagnostic report for controller design."],"forward_implications":["Tracking results can be reported by motion family, so a tracker that dominates daily upright motions may still collapse on ground-level transitions; aggregate scores alone hide this distinction.","A common reference representation, rollout accounting, and metric implementation make results across different tracking policies directly comparable rather than confounded by test-set differences.","HumanScore can be read alongside completion rate and joint error to rank trackers by perceived quality, catching foot sliding and mistimed contacts that pose error misses.","Longer evaluation windows, up to five seconds, improve alignment with human judgment, indicating that trajectory-level evaluation is needed rather than per-frame error.","The text-labeled 25K-clip benchmark supports fine-grained failure diagnosis in contact-rich and highly dynamic regimes."],"supporting_citations":[{"why":"Defines the widely used 140-sequence test set that the paper argues is too small and undiverse, motivating the new benchmark.","marker":"[26]"},{"why":"Retargets recorded human motion to the benchmark humanoid, creating the robot-space reference all trackers share.","marker":"[1]"},{"why":"One of the four trackers whose aligned rollouts form the preference pool and whose performance is compared on the benchmark.","marker":"[4]"},{"why":"Supplies the tracking success criterion used by the evaluator and is one of the four preference-pool trackers.","marker":"[24]"},{"why":"One of the four preference-pool trackers and the strongest overall method in the family-level benchmark results.","marker":"[29]"},{"why":"One of the four preference-pool trackers whose rollouts are paired for human comparison.","marker":"[44]"},{"why":"Provides the paired-comparison model used in the preference objective for HumanScore training.","marker":"[3]"},{"why":"Establishes the learning-from-human-preferences paradigm that HumanScore's data collection and training follow.","marker":"[5]"},{"why":"Supplies the Transformer encoder architecture that processes the per-frame trajectory tokens.","marker":"[35]"}],"fun_headline_variants":["HumanScore outranks pose errors at matching human judges","New metric aligns with human judges 90.8% vs 84.1%","Preference-trained score reveals contact failures pose errors miss","153-hour benchmark plus human-aligned metric for tracking eval","HumanScore sees what kinematic errors miss in contact-rich motions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on treating the pairwise judgments of six expert annotators, with a single label per pair and no measured agreement between annotators, as a reliable and stable ground truth for human perception of tracking quality.","fun_headline_variants_meta":{"raw":{"variants":["HumanScore outranks pose errors at matching human judges","New metric aligns with human judges 90.8% vs 84.1%","Preference-trained score reveals contact failures pose errors miss","153-hour benchmark plus human-aligned metric for tracking eval","HumanScore sees what kinematic errors miss in contact-rich motions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1390,"prompt_tokens":878,"completion_tokens":512,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":426}},"tokens_in":494,"tokens_out":512,"duration_ms":5512,"temperature":1.0,"reasoning_tokens":426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:04.049571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the preference alignment study with a fresh and larger annotation panel, giving each pair multiple independent labels and first measuring inter-annotator agreement; if HumanScore's agreement with those labels does not beat the best kinematic diagnostic, or if the original six annotators fail to agree with the new panel, the claim that HumanScore predicts human preferences collapses.","supporting_citations":[{"cited_title":"Gomez, Lukasz Kaiser, and Illia Polosukhin","cited_arxiv_id":null,"evidence_quote":"Supplies the Transformer encoder architecture that processes the per-frame trajectory tokens."},{"cited_title":"Rank analysis of incomplete block designs: I","cited_arxiv_id":null,"evidence_quote":"Provides the paired-comparison model used in the preference objective for HumanScore training."},{"cited_title":"Christiano, Jan Leike, Tom B","cited_arxiv_id":null,"evidence_quote":"Establishes the learning-from-human-preferences paradigm that HumanScore's data collection and training follow."}],"review_version":1}