{"id":"1d1b2168-e08a-45e5-8387-fe6e1b7013a1","arxiv_id":"2607.06989","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A robot arm generates ITTF-compliant table tennis serves with measured spin and speed matching or exceeding professional players, validated in official tournaments.","lead":"A robot arm now serves table tennis balls at professional level — spins above 550 rad/s and speeds of 6.7 m/s — using motion planning, Bayesian optimization, and physics simulation. The system was validated in umpire-officiated tournaments against elite and professional players, scoring direct serve points in up to 21% of serves.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'surpassing elite players' comparison is based on a pruned, human-curated competition library; Table III shows 40–60% of generated plans fail zero-shot, so the full generative pipeline's spin/speed and legality are not established.","rationale":"The reader's weakest assumption was the simulator-to-real gap and toss representativeness. That gap is real and honestly conceded in §V, but the system is externally filtered by umpires and real match play, and the paper's feasibility claim survives it. I see the more load-bearing issue as the comparison set: the headline 'matching and even surpassing elite players' is supported only by statistics from a pruned, hand-curated competition library, while the professional comparison data are unfiltered match serves. The paper never reports the full generative yield, so the reader cannot distinguish a method that produces professional-level serves from a method that occasionally produces one after many failures and human curation. This is not an internal inconsistency, and the real-match evidence gives genuine support for the central capability claim; the issue is specifically that the quantified superiority claim is under-supported. The conditional verdict remains appropriate, but for a reason different from the reader's stated weakest assumption.","tokens_in":11784,"tokens_out":5345,"duration_ms":57861,"concrete_test":"Re-run the four task experiments with N=25 per task and report metrics for all 25 plans, including the 10–13 that fail zero-shot transfer, treating failed zero-shot transfers as invalid serves. Compare the mean and upper quantiles of spin/velocity and the ITTF-legal rate of this full batch against the professional serve distribution underlying Fig. 4(b) and the Table II serving/ace statistics. If the full batch no longer matches or exceeds elite spin/speed, or if ace/win rates drop to the receiving baseline, the 'surpassing' claim is a selection artifact; if it still matches, the concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that the framework generates serves matching or surpassing elite players (Abstract, §IV-A). The supporting evidence compares not the full output of the generation pipeline, but a competition library that was iteratively pruned by real-world performance: §III-E states serves whose rallies tend to end in opponent points are removed, and §IV-A says serves are 'strategically prune[d] and replace[d]' over tournaments. Table III further shows that in the task-generation experiment only 12–15 of 25 planned serves per task transferred zero-shot, and the real-world statistics are computed only over those survivors ('The statistics reported in the real world are thus based on the subset of serves that actually transferred zero-shot without issues'). Thus the reported real spin/speed values (e.g., real topspin mean 263 rad/s vs. simulated 218 rad/s) and the Table II ace rates characterize a hand-selected subset, not the automatically generated set. The professional-player comparison in Fig. 4(b) uses 1243 serves executed by the professionals in matches—presumably all their serves—while the robot's markers come from the pruned library. Without reporting the yield, legality rate, and spin/speed distribution of the full 25-plan batch, the 'matching and even surpassing' claim may be an artifact of selection rather than of the generation method. The real-match evidence (licensed umpires, professional opponents, counted aces) supports feasibility, but it does not, as reported, support the comparative headline as quantified.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a framework for generating ITTF-compliant table tennis serves with an 8-DoF robot arm. The pipeline combines a jerk-minimizing Model Predictive Control motion planner with the HEBO Bayesian optimizer to search over racket strike parameters in simulation, followed by offline filtering and iterative real-world selection into a competition library. The authors claim that this framework produces serves with spin up to 550 rad/s and speeds up to 6.7 m/s, matching or surpassing elite human players, and report match statistics against professional players (including Olympian Miu Hirano) with ace rates up to 21%. The central claim is that this is the first framework capable of generating professional-level, ITTF-compliant serve motion plans for a robot arm.","tokens_in":12088,"tokens_out":4767,"duration_ms":50063,"significance":"If the claim is substantiated, this is a significant advance in robot table tennis and dynamic manipulation. The real-world validation is unusually strong for a robotics paper: licensed umpires, professional opponents, and counted direct serve points. The combination of an optimization-based motion planner with Bayesian-optimized strike parameters is a technically plausible and potentially generalizable approach. The paper also provides quantitative data on spin, velocity, and aiming error, which is valuable to the community. However, the headline claim of 'matching and surpassing elite players' is weakened by selection effects in the reported data, as detailed in the major comments.","major_comments":[{"comment":"The real-world spin/velocity statistics and the comparison to elite players are computed only over the subset of serves that transferred zero-shot (N=12-15 of 25). The paper states: 'The statistics reported in the real world are thus based on the subset of serves that actually transferred zero-shot without issues.' This selection bias makes the 'matching and even surpassing elite players' claim an artifact of curation rather than a property of the generative framework. Please report the yield, legality rate, and full distributions for all 25 generated plans, or explicitly restrict the claim to the curated library.","section":"Section IV-A, Table III"},{"comment":"The improvement narrative in Table II confounds algorithmic pipeline changes with strategic library pruning. The March and April 2026 rows explicitly include 'strategically prune and replace serves whose historical performance shows a tendency to lead to opponent's points.' This means the increase in ace rate could be due solely to the removal of poor serves, not to the HEBO, toss-robustness, or dynamic-limit changes. To support the claim that the pipeline improves, provide an ablation that separates the effect of library curation from generation-side improvements.","section":"Section III-E, Table II"},{"comment":"The statement 'strong correlation between the serve property encouraged through the reward function and the corresponding physical metric' is circular. The reward functions in Eqs. (3)-(4) directly include the physical metric being measured (e.g., r_top maximizes topspin). Finding that topspin-rewarded runs produce higher topspin is a sanity check, not a validation of the framework. Please rephrase or provide a non-circular validation, such as comparing against a baseline without the corresponding reward term.","section":"Section IV-B, Eqs. (3)-(4)"}],"minor_comments":[{"comment":"The abstract claims spin 'up to 550 rad/s' but I could not find a table or figure in the provided text that reports this maximum. Please cite the specific figure/table that supports the 550 rad/s value.","section":"Abstract, Section IV-A"},{"comment":"The text mentions 'n_l = 323rd degree order polynomials' and later 'n_l = 32'. This appears to be a formatting error; likely it should read 'n_l = 32 third-degree polynomials.' Please clarify.","section":"Section III-C, Eq. (2)"},{"comment":"The claim of 'slightly higher winning probabilities while serving' compared to the elite human statistics from [27] should be qualified: the 95% Wilson confidence intervals in April 2026 (51.5-58.7%) overlap with the elite values (52.78% and 53.28%). A formal statistical comparison would strengthen the claim.","section":"Section IV-A, Table II"},{"comment":"The paper relies heavily on the authors' own prior Nature paper [1] as the 'baseline method.' Please explicitly state the novel contributions of this manuscript relative to [1] to avoid any ambiguity about the incremental advance.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is strong but currently supported by selective evidence. The full generative pipeline produces 40-60% zero-shot failures in Table III; the professional-level metrics come from a hand-curated subset. This is fixable if the authors report full-batch statistics or explicitly scope the claim to the competition library. I would also request an ablation for Table II to disentangle curation from algorithmic improvement. The paper is otherwise of high quality and would be a strong contribution after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a real achievement. The authors have a robot arm serving with measured spin up to 550 rad/s and speeds up to 6.7 m/s, under ITTF rules, in umpire-officiated matches against professional players, with aces counted. That's a first in the cited literature—[15] does kinesthetic teaching without elite spin, [16] uses a non-ITTF overhead drop, [17] is simulation-only. The MPC+HEBO pipeline is sensible, clearly described, and the real-world validation is stronger than most robotics papers: licensed umpires, professional opponents, ball tracking, direct points. The central capability claim—this system can produce legal, high-spin, professional-level serves in reality—holds up.\n\nWhat doesn't fully hold up is the quantitative comparison to elite players. The stress-test note is right. The real spin/speed values in Table III come only from the 12–15 of 25 plans per task that transferred zero-shot; the rest failed and are excluded. And the competition library used for the professional-player comparison in Fig. 4(b) was iteratively pruned—serves that lost points were removed. So the robot's markers represent a hand-curated best subset, while the professional data (1243 serves) are presumably all their serves. That makes 'matching and even surpassing' a claim about the curated system, not about the raw generative pipeline. The paper should report yield, legality rate, and full batch statistics, or drop the comparative headline.\n\nOther soft spots: the 'strong correlation' between reward and physical metric in Section IV-B is partly circular—the reward literally maximizes that metric. It's a sanity check, not independent validation. Table II's win-rate improvements sit in overlapping confidence intervals and are confounded by library pruning. No code or data are released, which makes the statistics hard to check. On the positive side, the paper is honest about the sim-to-real gap in Section V, and the methodology is detailed enough to reproduce.\n\nWho this is for: robotics researchers working on precision planning, black-box optimization in motion control, and robot sports. It deserves a serious referee—this is not a desk reject. But the authors should be pushed to release artifacts, report all transferred and failed plans, and soften the 'surpassing' claim to something like 'comparable in the curated set.' The core result stands either way.","headline":"Real ITTF-legal professional-level serves from a robot arm, genuinely new, but the 'surpassing elite' headline overstates a pruned and curated subset.","tokens_in":12698,"tokens_out":2069,"would_cite":true,"duration_ms":23442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T40"],"pacs":[],"model":"deepseek-v4-flash","headline":"A robot arm can serve table tennis at professional level — spins up to 550 rad/s and speeds up to 6.7 m/s — while obeying official ITTF rules, according to this motion-planning framework.","keywords":["table tennis","robot serve","motion planning","model predictive control","Bayesian optimization","ITTF rules","spin generation","sim-to-real"],"falsifier":"Run the competition library's serves on the real robot and measure the ball's linear and angular velocity at the second bounce with an independent multi-camera tracking system: if mean spin falls below the reported task ranges (e.g., below 200 rad/s for the topspin task) or if more than 20% of serves fail umpire legality review in a fresh session, the central claim would be contradicted.","tokens_in":11609,"feed_emoji":"🏓","tokens_out":5247,"duration_ms":48778,"temperature":0.7,"pith_summary":"This paper tries to establish that a robot arm can serve in table tennis at a professional level: generating high spin (up to 550 rad/s) and speed (up to 6.7 m/s) from a spinless free-falling ball, while remaining fully compliant with ITTF service rules. The authors argue their pipeline — a parametrized minimum-jerk motion planner plus a Bayesian optimizer that searches over racket contact state in simulation — is the first to achieve this. In officiated matches against elite and professional players, the robot's serves won points directly (aces) in 8 to 23 percent of its serves, and serving win probabilities reached 51.5–58.7%, slightly above published elite human stats. The claim matters because the serve is the only fully controllable action in table tennis, and previous robot serve work focused on placement without high spin or a legal toss. The paper rests on zero-shot open-loop execution: plans are trained in simulation and deployed without real-world adaptation, with a conceded sim-to-real gap that reduces transfer efficiency.","feed_headline":"Robot arm serves table tennis with elite-level spin and speed","feed_subtitle":"New pipeline produces ITTF-legal serves scoring aces in up to 23% of tries.","key_machinery":"The key machinery is a two-stage pipeline: (1) a parametrized model-predictive motion planner that solves a constrained minimum-jerk optimization (Eq. 2) to move the racket to a specified pose, orientation, and velocity at a specified time, then return smoothly to rest; and (2) HEBO, a heteroscedastic and evolutionary Bayesian optimizer, which searches over up to ten normalized parameters — hit timing, racket velocity vector, racket yaw/roll orientation, return duration, and optional position offset — using a simulator rollout with a legality-check chain to score each candidate. The racket's desired intercept point is tied to the expected tossed-ball trajectory, so the optimizer effectively","core_discovery":"The central discovery is that by combining a hybrid end-effector/joint-space minimum-jerk motion planner with a sample-efficient Bayesian optimizer (HEBO), the robot can find racket contact states (time, velocity, orientation, return duration) that turn a tossed ball into a legal, hard-to-return serve. The optimizer rolls each candidate plan out in a physics simulator, applies a sequential chain of ITTF legality checks, and rewards spin, velocity, placement, and low trajectory height. The resulting serve library, pruned by real tournament outcomes, produced measured spin and velocity that the authors report match or exceed those of professional players, and that won direct points against Oly","pith_inferences":["One testable extension would be to add a per-toss closed-loop adjustment of hit timing based on online ball-tracking; if the reported ace rates are driven partly by toss prediction errors, closing the loop should raise the zero-shot transfer rate well above the current 48–60%.","The reported ace rates vary from 1.6% to 21% across tournaments, and the authors attribute gains to pruning and spin/topspin prioritization; this suggests the library composition, not the motion planner alone, carries much of the matchup effectiveness — a hypothesis that could be checked by ablating the pruning rule.","The abstract's headline spin of 550 rad/s is an optimized extreme, while task-specific tables show lower, more variable typical values (e.g., topspin task real spin 263±70 rad/s); readers should distinguish the claimed ceiling from the library's typical performance."],"forward_implications":["If the claim holds, robot table tennis can now include a professional-caliber serve, closing the last fully 'human' phase of the game.","The legality-check chain in simulation plus open-loop deployment suggests that simulation-based training can produce rule-compliant behaviors without real-world trial and error, at least for compact, high-speed sports actions.","Because the same motion planner handles the toss and the strike, the framework extends to other tasks requiring precise, high-speed interception with secondary objectives (e.g., placement or spin).","The reported serving win probabilities (51.5–58.7%) slightly exceed published elite male/female averages, implying the serve is at least as effective as a human pro's serve.","The zero-shot transfer rate (12–15 of 25 plans per task) defines a concrete ceiling: a quarter to half of simulated serves still need to be filtered out by real-world trials."],"fun_headline_variants":["Robot arm's serves hit pro-level spin and ace 23% of the time","ITTF-legal robot serves top elite players' spin and speed","Robot teaches table tennis serve: 550 rad/s spin, 6.7 m/s speed","Ace-making robot arm matches pro table tennis serves","Bayesian-optimized robot serves score aces against pro-level play"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole serve library is planned against a physics simulator and executed open-loop, so everything rests on the simulator's ball-flight, bounce, and impact models being accurate enough that ITTF legality and high spin/speed carry over to the real table, and on the N=15 recorded toss trajectories remaining representative of future tosses.","fun_headline_variants_meta":{"raw":{"variants":["Robot arm's serves hit pro-level spin and ace 23% of the time","ITTF-legal robot serves top elite players' spin and speed","Robot teaches table tennis serve: 550 rad/s spin, 6.7 m/s speed","Ace-making robot arm matches pro table tennis serves","Bayesian-optimized robot serves score aces against pro-level play"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000463,"raw_usage":{"total_tokens":2126,"prompt_tokens":692,"completion_tokens":1434,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":1336}},"tokens_in":436,"tokens_out":1434,"duration_ms":10337,"temperature":1.0,"reasoning_tokens":1336,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:09:12.190244+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the competition library's serves on the real robot and measure the ball's linear and angular velocity at the second bounce with an independent multi-camera tracking system: if mean spin falls below the reported task ranges (e.g., below 200 rad/s for the topspin task) or if more than 20% of serves fail umpire legality review in a fresh session, the central claim would be contradicted.","supporting_citations":[],"review_version":2}