{"id":"797032fe-2ebd-4f8a-8f51-f21d97f5da01","arxiv_id":"2607.11481","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A single-stage RL co-tracking controller trained on consecutive human-derived hand–object subgoals achieves ~75% real-robot success on long-horizon dexterous teleoperation where baselines fail.","lead":"TeleDexter is a learned hand–object co-tracking controller that turns operator fingertip and object targets into real-time multi-contact finger control for in-hand reorientation and tool use. It reports ~75% success on seven hard teleoperation tasks where kinematic and generative baselines fail, and the demos train autonomous policies.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"DexGen baseline reimplementation substitutes TeleDexter's own co-tracking controller for the missing AnyGrasp-to-AnyGrasp stage, so the claim that 'all baselines consistently fail' is not fully independent.","rationale":"The reader's weakest assumption correctly flags MoCap dependence, object-specific policies, free-space training gaps (impact dynamics, no tactile, compliance), and incomplete DexGen fidelity. The single most load-bearing concern for the central comparative claim is the last of these: the explicit substitution of TeleDexter's own controller into DexGen's training pipeline (Supp. D). That substitution is concrete, documented, and directly undercuts the independence of the strongest baseline that was supposed to possess a learned contact prior. Absolute performance (75% SR, stage-wise survival, autonomy-from-teleop) is still well-supported by the kinematic baselines and SimToolReal, so the verdict remains CONDITIONAL rather than REJECT; the concern simply makes the comparative language less secure than the absolute numbers. A clean re-run of DexGen without the co-tracking substitution would settle whether the gap is real or partly an artifact of the reimplementation choice.","tokens_in":24075,"tokens_out":569,"duration_ms":6268,"concrete_test":"Re-train the DexGen diffusion prior exclusively on rollouts from an independent contact-rich policy (e.g., a pure dense-tracking RL controller or the original paper's synthetic grasp transitions if recoverable) that does not use consecutive subgoal co-tracking; re-evaluate the resulting DexGen controller on the same 7 tasks / 15 trials. If its SR remains near zero, the comparative claim holds; if it rises substantially, the original 'all baselines fail' statement overstates the gap.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on TeleDexter's ~75% SR where DexRT, GeoRT, DexGen, and SimToolReal 'consistently fail' (near-zero SR). Supp. D states that DexGen has no official code and that the authors 'substitute it with our co-tracking controller to generate the simulation rollouts' for the AnyGrasp-to-AnyGrasp RL stage, then train the diffusion prior on those rollouts. Because the generative prior is therefore trained on trajectories produced by the same family of co-tracking policy being evaluated, DexGen is no longer an independent generative baseline; any failure may reflect distribution mismatch or incomplete reproduction rather than a pure comparison against the published DexGen method. This weakens the comparative half of the headline claim even if absolute TeleDexter numbers remain strong. The other baselines (pure kinematic retargeting and SimToolReal) are cleaner, but the paper repeatedly groups 'all baselines' together.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces TELEDEXTER, a hand–object co-tracking controller for dexterous teleoperation. The operator specifies synchronized fingertip and object pose targets; a low-level RL policy, trained in simulation on consecutive co-tracking subgoals derived from geometry-aware retargeted human HOI motions, realizes multi-contact dynamics. Training uses a hybrid sparse subgoal / dense tracking reward, curriculum annealing, domain randomization, and random action masking, and is claimed to transfer zero-shot. Real-world evaluation on seven reorientation and long-horizon tool-use tasks across SharpaWave and LeapHand reports ~75% average success (75.2% SR / 87.1% TP on SharpaWave) where kinematic retargeting, a reimplemented generative prior, and an object-centric tool policy largely fail. Teleoperated demos are further used to train Diffusion Policies for autonomous execution of contact-intensive stages.","tokens_in":24483,"tokens_out":1537,"duration_ms":23447,"significance":"If the absolute real-world results hold, this is a substantial systems contribution: continuous in-hand reorientation, finger gaiting, and multi-stage tool use under teleoperation remain largely out of reach for pure kinematic retargeting, and the paper shows a single-stage, reference-driven RL controller can close much of that gap on two hand morphologies without per-task reward engineering. Strengths include stage-wise SR/TP reporting, honest failure-mode analysis (Supp. A.4), ablations of sparse vs. dense tracking (sim) and action masking (real), any-to-any reposition stress tests, and a demonstrated path from teleop demos to autonomous BC. Random action masking as an action-space regularizer is a concrete, transferable sim-to-real idea. The work is a credible step toward scalable collection of contact-rich dexterous data, even though controllers remain object-specific and MoCap-dependent.","major_comments":[{"comment":"Supp. D (DexGen): The paper states that DexGen has no official code and that the authors “substitute it with our co-tracking controller to generate the simulation rollouts” for the AnyGrasp-to-AnyGrasp stage before training the diffusion prior. The generative baseline is therefore trained on trajectories from the same co-tracking family being evaluated, so it is not an independent reproduction of published DexGen. Tab. 1 and the abstract’s claim that “all baselines consistently fail” group this compromised baseline with cleaner ones (DexRT, GeoRT, SimToolReal). Please either (i) re-implement the missing stage without TELEDEXTER rollouts, (ii) drop DexGen from the main comparison, or (iii) clearly caveat Tab. 1 / abstract / §4.2 so that the comparative claim rests only on independent baselines. Absolute TELEDEXTER numbers can still stand.","section":"Supp. D; Tab. 1; Abstract; §4.2"},{"comment":"§3.1–3.2 and §6: Controllers are object-specific and trained only on free-space hand–object HOI (no tool–environment impact). Supp. A.4 correctly identifies interaction perturbation under hammering as a dominant failure mode. HammerUse still reports 66.7% SR, so the method is partially effective, but the abstract’s framing of “long-horizon tool use” and “human-level” contact transitions should be tightened to match the training distribution and the disclosed impact gap (e.g., quantify how often nail-driving succeeds vs. fails due to impulsive reaction). This is needed so readers do not over-read free-space co-tracking as sufficient for impact-rich tool application.","section":"§3.1–3.2; §6; Supp. A.4; Abstract"},{"comment":"§4.2 Protocol: Each task uses 15 trials and a skilled operator in the loop with MoCap. SR/TP therefore conflate operator skill, interface latency, and controller robustness. Stage-wise plots (Fig. 6) help, but the paper should report operator protocol more tightly (same operator across methods? practice trials? stopping rules) and, where possible, inter-operator or inter-session variance, so that the large gap vs. DexRT/GeoRT is attributable to the learned contact prior rather than unequal human adaptation. Without this, the comparative half of the headline claim is harder to interpret even for the clean kinematic baselines.","section":"§4.2; Fig. 6; Supp. B.3"}],"minor_comments":[{"comment":"Title and abstract use “human-level” while §6 and Supp. A.4 document object-specificity, MoCap dependence, and three systematic failure modes. Soften or define the phrase (e.g., “toward human-like in-hand contact transitions under teleoperation”).","section":"Title; Abstract; §5–6"},{"comment":"Eq. (2)–(3) and Supp. C.2 list many free reward/curriculum parameters (α_dense, β’s, w_step rules, N_stay, σ schedule). A short sensitivity note or default-transfer statement would help reproducibility claims for new objects/hands.","section":"Eq. (2)–(3); Supp. C.2–C.3"},{"comment":"Tab. 2 reports only three reorientation tasks on LeapHand; tool-use results for LeapHand are absent. Either add them or state explicitly that tool-use evaluation is SharpaWave-only.","section":"Tab. 2; §4.2"},{"comment":"Fig. 2 “86% in real world” is unclear relative to Tab. 1’s 75.2% SR / 87.1% TP; align figure callouts with table metrics.","section":"Fig. 2"},{"comment":"SimToolReal is correctly labeled non-teleoperation, but Tab. 1 averages it over three tasks only while TELEDEXTER is averaged over seven; footnote this more prominently when stating “all baselines.”","section":"Tab. 1"},{"comment":"Typographical: “arXiv:2607.11481v1” date line and occasional spacing (e.g., “hand–object” consistency) should be cleaned in camera-ready.","section":"Front matter"}],"recommendation":"major_revision","confidential_remarks":"The absolute real-world results and stage-wise analysis are the paper’s main asset; the DexGen reimplementation is the one issue that, if left unaddressed, would make me push back on the comparative headline in a high-visibility venue. I would accept after a revision that either fixes or clearly demotes DexGen and tightens the tool-impact / operator-protocol language. Fit for a top robotics journal is good if those points are handled; novelty is systems-level rather than a new learning theorem."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The headline result is real: a single-stage RL co-tracking controller that takes synchronized fingertip + object subgoals, trains with hybrid sparse/dense rewards plus random action masking, and zero-shots to two hands on seven multi-stage reorientation and tool-use tasks at ~75% SR / 87% TP while pure kinematic retargeting and SimToolReal mostly collapse. That is a useful systems step for collecting contact-rich demos that existing teleop cannot produce, and they show the demos train Diffusion Policies with non-trivial stage-wise success.\n\nWhat is actually new is the package, not any single ingredient. Consecutive subgoal co-tracking (reach then advance, not frame-wise imitation) plus the hybrid reward is the design that lets one stage cover translation, reorientation, and gaiting without per-skill reward engineering. Geometry-aware retargeting of MoCap HOI into feasible subgoals, curriculum, and especially random action masking are the practical enablers. Real evaluation is careful: 15 trials, stage-wise SR/TP, two embodiments, sim sparse-vs-dense ablation, real masking ablation, and honest failure modes (impact, jam, stall).\n\nSoft spots, in proportion. The DexGen comparison is not independent: they reimplemented without official code and substituted their own co-tracking controller for the missing AnyGrasp-to-AnyGrasp stage, so “all baselines fail” overstates the generative half of the claim. Kinematic baselines and SimToolReal are cleaner and still fail where it matters. Policies are object-specific; deployment needs MoCap; free-space training omits tool–environment impacts and tactile. Those are stated limitations, not hidden. Free parameters are many but typical for this class of work; no circularity in the success metrics.\n\nThis is for people building dexterous teleop or data pipelines for multi-finger autonomy. Math is standard RL; data and citations look solid. I would send it to peer review. Engage with the absolute system and the autonomy-from-teleop path; discount the DexGen line until someone runs a cleaner reimplementation.","headline":"Strong real-robot co-tracking teleop result (~75% SR on hard in-hand/tool tasks); the DexGen baseline is compromised by using their own controller for rollouts, but absolute numbers and other baselines still carry the paper.","tokens_in":25149,"tokens_out":533,"would_cite":true,"duration_ms":6426,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A hand–object co-tracking controller trained on consecutive human-motion subgoals delivers real-robot in-hand and tool teleoperation at about 75% average success where prior systems fail.","keywords":["dexterous teleoperation","in-hand manipulation","hand-object co-tracking","sim-to-real","reinforcement learning","finger gaiting","tool use","robot manipulation"],"falsifier":"Retrain and redeploy the co-tracking policy with consecutive subgoals, hybrid reward, and random action masking on the same seven real tasks and two hands; if success on stages that need in-hand reorientation or sustained tool contact remains near the near-zero rates of kinematic and generative baselines, the central claim is false.","tokens_in":24939,"feed_emoji":"✋","tokens_out":1068,"duration_ms":22886,"temperature":0.7,"pith_summary":"The paper claims that human-level in-hand robot teleoperation is reachable if the operator gives synchronized fingertip and object pose targets and a learned low-level controller realizes the contact physics. TeleDexter trains that controller in one reinforcement-learning stage on consecutive co-tracking subgoals taken from geometry-aware retargeted human hand–object motions, using a hybrid reward that mixes sparse subgoal success with light dense tracking so the policy can invent feasible contact strategies instead of copying frames. Random action masking plus domain randomization let the same policy transfer zero-shot to two real dexterous hands. On seven reorientation and long-horizon tool-use tasks it reaches roughly 75% average success while kinematic and generative baselines nearly always fail, and the teleoperated traces train autonomous policies. A sympathetic reader would care because the method both unlocks contact-rich teleoperation and supplies a practical way to collect the demonstration data those skills require.","feed_headline":"Robot hands hit 75% on hard in-hand tool teleoperation","feed_subtitle":"Co-tracking from human subgoals transfers zero-shot; kinematic and generative baselines fail.","key_machinery":"Consecutive subgoal co-tracking: ordered fingertip-and-object pose targets derived from human HOI motions that the policy must reach before advancing, trained with a hybrid sparse subgoal-reaching plus dense tracking reward, and regularized by random action masking for zero-shot sim-to-real transfer.","core_discovery":"TeleDexter shows that casting dexterous teleoperation as hand–object co-tracking—operator-specified fingertip positions and object poses executed by a single-stage RL controller trained on consecutive subgoals from human reference motions—yields real-world in-hand reorientation, finger gaiting, and multi-stage tool use at about 75% average success across seven tasks and two hand embodiments, where pure kinematic retargeting and prior learned action priors consistently fail.","pith_inferences":["Object-specific controllers and motion-capture pose streams remain the main deployment bottlenecks; a vision-conditioned multi-object co-tracker is the natural next system.","Documented failure modes—impact perturbation, contact jam, tracking stall—imply that adding tactile sensing and tool–environment impacts in training could close remaining long-horizon gaps.","If ordered fingertip–object subgoals are the right intermediate representation, the same formulation may extend to bimanual or multi-object in-hand tasks without new reward design.","High-quality teleop data from this interface could become a standard substrate for imitation learning of skills pure vision-based retargeting cannot demonstrate."],"forward_implications":["Operators can teleoperate contact-rich skills—in-hand reorientation, finger gaiting, hammering, screwdriving, bulb install—that kinematic retargeting cannot stabilize.","The same human references, after geometry-aware retargeting, train controllers for both four-finger and five-finger hands without recollecting motions.","Teleoperation traces collected with TeleDexter can train autonomous diffusion policies on dexterous subtasks from tens of demonstrations.","Diverse in-hand modalities can be learned in a single RL stage without per-task reward engineering when goals are consecutive co-tracking subgoals rather than frame-wise imitation.","Random action masking is presented as a necessary action-space regularizer for zero-shot transfer of contact-rich hand policies."],"fun_headline_variants":["TeleDexter hits 75% on seven hard in-hand teleop tasks","Co-tracking RL delivers 75% where kinematic baselines fail","Zero-shot hand-object co-tracking enables real tool teleop","Single-stage RL from human subgoals reaches 75% success","In-hand reorientation and tool use at 75% average success"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim rests on free-space hand–object contact skills learned from retargeted human subgoals in simulation, without tool–environment impact forces or touch sensing, being enough for long real-world tool use once action masking and domain randomization are applied.","fun_headline_variants_meta":{"raw":{"variants":["TeleDexter hits 75% on seven hard in-hand teleop tasks","Co-tracking RL delivers 75% where kinematic baselines fail","Zero-shot hand-object co-tracking enables real tool teleop","Single-stage RL from human subgoals reaches 75% success","In-hand reorientation and tool use at 75% average success"]},"model":"grok-4.5","effort":"low","cost_usd":0.00697,"raw_usage":{"total_tokens":1765,"prompt_tokens":809,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":69700000,"prompt_tokens_details":{"text_tokens":809,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":878,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":809,"tokens_out":78,"duration_ms":6858,"temperature":1.0,"reasoning_tokens":878,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T05:18:10.396296+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain and redeploy the co-tracking policy with consecutive subgoals, hybrid reward, and random action masking on the same seven real tasks and two hands; if success on stages that need in-hand reorientation or sustained tool contact remains near the near-zero rates of kinematic and generative baselines, the central claim is false.","supporting_citations":[],"review_version":1}