{"id":"234b7ffd-7622-4d78-9c1e-4735107a6601","arxiv_id":"2412.03252","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Variable-speed teaching-playback as real-world data augmentation improves variable-speed, contact-rich imitation learning from two demonstrations.","lead":"This paper proposes using teaching-playback at different playback speeds to gather real robot reactions and use them as extra training data for imitation learning with force control. It reports higher success rates and better speed matching than software-only speed changes, starting from just two fixed-speed human demonstrations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline dataset size and train/validation split are unreported, so the Tables 1–4 gains may reflect more training data rather than real-world speed-dependent reactions.","rationale":"The reader's weakest_assumption identifies exactly this issue: the baseline may not be matched to the proposed method in sample count and training conditions, so the observed gains could be a data-quantity effect rather than an effect of real-world reactions. I agree with that diagnosis and consider it the most load-bearing concern because the paper's title, abstract, and conclusion all attribute the improvement to 'real-world data augmentation'—the comparison to a simple off-line duplication baseline is the only evidence that isolates this mechanism. The proposed method's sample counts are reported in detail, but the baseline's are not, despite both being needed for a controlled comparison. I also note the success-filtering confound: keeping only successful playbacks may improve training data quality independently of speed-related environmental reactions, and the paper does not explain whether the baseline received an equivalent filter. This does not undermine the plausibility of the method or the direction of the pick-and-place evidence; it means the current manuscript does not yet establish the claimed causal effect. The verdict should remain CONDITIONAL rather than REJECT because the concern is addressable with additional reporting or a matched-baseline experiment, and the existing evidence is directionally supportive. The limitations stated in §5.5 bound the scope but do not resolve the baseline-construction issue.","tokens_in":11294,"tokens_out":4080,"duration_ms":42215,"concrete_test":"Obtain or reconstruct the baseline dataset construction: for each task, create the simple-duplication baseline with exactly the same number of training sequences per label as the playback condition (42 pick-and-place sequences: 7 per object per speed; 18 wiping sequences: 3 per height per speed), the same train/validation split, and the same label balance, while still using the original command/response values rather than playback responses. Retrain the same LSTM on this matched baseline and rerun the Tables 1–4 evaluations. If the matched baseline closes the success-rate gap, the real-world-response effect is not established; if the gap persists, the central claim is supported. Even without new experiments, the authors should report the per-condition sample counts and splits for the baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal claim—that gathering real-world reactions at variable speeds improves success—rests entirely on the comparison in Tables 1–4 between the proposed playback dataset and a baseline that 'simply duplicates' the human demonstrations and changes their speed (§4). The text states that both methods share the same command values, but it never reports how many baseline trajectories were used, how they were split into training/validation, or how duplicates were balanced across labels. The proposed method's counts are explicit (§5.3.1: 42 training playbacks for pick-and-place; §5.3.2: 18 for wiping), while the baseline counts are absent. If the baseline was built from only the two original demonstrations per speed (6 trajectories for pick-and-place), the 53% vs 30% overall gap and 88% vs 31% interpolation gap could be driven by dataset size and label coverage rather than by the real-world environmental reactions that the paper claims are the active ingredient. A second, related confound is that playback data were filtered to 'successful' playbacks only, while the text does not state whether any analogous selection was applied to the baseline; success-based filtering can itself improve training data quality independently of speed-dependent reactions. This is not a claim that the authors acted improperly—the omission may be an oversight—but it is the point where the central inference is least secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using teaching-playback at variable speeds as real-world data augmentation for bilateral-control-based imitation learning. Starting from two fixed-speed human demonstrations per task, the authors generate additional real-robot trajectories by replaying recorded motion at 0.5x, 1x, and (for pick-and-place) 2x speed, collecting the resulting real force/torque reactions, and retaining only 'successful' playbacks. They compare this dataset against a baseline in which the same human demonstrations are simply duplicated and their speed is changed offline, with identical command values but without the real-world responses. Experiments on pick-and-place and wiping report success rates and label-following accuracy for speed commands inside and outside the training range. The main claims are that the proposed real-world augmentation improves overall task success (53% vs 30% for pick-and-place; 79% vs 74% for wiping) and improves accuracy along the duration/frequency command, especially for interpolation.","tokens_in":11558,"tokens_out":3699,"duration_ms":36224,"significance":"The idea is practically motivated and squarely within the journal's scope: it addresses a real bottleneck in force-controlled imitation learning, namely the scarcity of hard-to-simulate contact data. A notable strength is that the method is evaluated on a physical robot with two distinct contact-rich tasks and unseen objects, rather than only in simulation. The authors are also explicit about limitations (no position diversity, no closed-loop feedback) and about the relationship to prior fast-forward collection work. If the comparison against the within-paper baseline is properly controlled, the pick-and-place results are substantial and would be a useful data-augmentation recipe for the bilateral-control imitation-learning community. The paper does not provide code or data, but the experimental protocol is largely reproducible from the text once the missing baseline details below are supplied.","major_comments":[{"comment":"The central comparison against the 'simple duplication and changes in speed' baseline is under-specified. The proposed method's dataset size and train/validation split are explicit (pick-and-place: 42 training and 18 validation playbacks; wiping: 18 training and 12 validation), but the text never states how many baseline trajectories were generated, how many times the human demonstrations were duplicated at each speed, or how the baseline was split into training and validation. Since the paper's main claim is that real-world reactions at variable speeds, rather than simply more samples or different label coverage, produce the gains in Tables 1–4, the baseline must be matched in sample count, duplication structure, and train/validation proportions. Without these numbers, the 53% vs 30% and 88% vs 31% gaps could in part reflect a data-quantity or label-coverage effect. Please report the full baseline construction and, ideally, run a matched-sample-count baseline.","section":"§5.2, §5.3.1, §5.3.2, Tables 1–4"},{"comment":"Every success-rate cell in Tables 1–4 is based on only five trials, and no confidence intervals or significance tests are reported. The pick-and-place improvement is large and fairly consistent across objects, but the wiping result is mixed: the proposed method improves at 15 cm (94% vs 68%) while worsening at 12 cm (63% vs 86%), for an overall 79% vs 74% over 70 trials. Given the small per-cell sample size, the strength of the success-rate claims should be supported with at least binomial confidence intervals or a simple test (e.g., Fisher's exact test on the pooled overall counts), and the mixed height-dependent effect in wiping should be acknowledged in the conclusions rather than only in the results section.","section":"§5.4.1, §5.4.2, Tables 1–4"},{"comment":"The text says that playbacks were repeated until a 'certain number of successful playbacks' was reached and that only successful data were used for training, but the success criterion for a playback is never defined for either task. For pick-and-place, the trial-level success criterion (object inside the circle within 40 s) is given, but it is not stated that the same criterion was applied to playbacks; for wiping, no playback-level success definition appears at all. Because success-based filtering can improve training-data quality independently of speed-dependent reactions, the paper should specify the playback success criteria and state whether any analogous selection was applied to the baseline dataset.","section":"§4 and §5.3.2"}],"minor_comments":[{"comment":"The phrase 'a maximum 55% increase in success rate' is ambiguous: it is not clear whether the increase is in percentage points or relative percentage, and the value does not obviously match any single cell in Tables 1–4 (several cells show larger point differences). Please report the exact source of the 55% figure.","section":"Abstract and §5.4.1"},{"comment":"The statement that scaling time 'clearly collides with the law of cause and effect, rendering it infeasible' is vague; a more precise technical explanation of why temporal scaling is not well posed in four-channel bilateral control would be more informative.","section":"§3.1"},{"comment":"The figure caption and the text use both 'simple duplication and speed adjustments' and 'simple fast-forward' to name the baseline; please use one consistent term throughout.","section":"Figure 4"},{"comment":"The text says the baseline's failures were mainly due to starting periodic movement before pressing, but Table 4 shows the baseline succeeding more at 12 cm than the proposed method (86% vs 63%); a sentence explaining this opposite pattern would help the reader interpret the height-dependent failure modes.","section":"§5.4.2, Table 4"},{"comment":"When discussing Sakaino et al.'s fast-forward data collection [11], the paper says its 'effect and feasibility for variable-speed tasks remain unclear'; since the current work directly builds on that method, it would be helpful to state more concretely what new evidence the current experiments add relative to [11].","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: the baseline dataset construction is the load-bearing missing piece. The paper's positive contribution is real, and the omissions appear to be reporting gaps rather than fundamental errors, so major revision rather than rejection seems appropriate. The five-trial-per-cell issue is common in hardware papers, but given that the entire evidence is empirical, the authors should at least report binomial confidence intervals or raw counts. The wiping task's mixed height result should be more prominently reflected in the discussion and conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is genuinely useful: instead of simulating speed-dependent contact forces, use motion-copying playback at variable speeds to collect real force reactions for imitation learning. That is a legitimate and practical contribution, and the pick-and-place results are directionally strong—53% overall vs 30% baseline, and 88% vs 31% in the interpolation range. The authors also deserve credit for testing on two tasks with different speed commands and for reporting actual completion times and frequencies, not just success counts. The paper is honest about the method's limitations: no environmental feedback during playback, and only speed diversity, not positional diversity.\n\nBut the central comparison has a load-bearing gap. The text says the baseline was built by simply duplicating human demonstrations and changing speed, and that both methods share the same command values. It never reports how many baseline trajectories were generated, how they were split into training and validation, or whether a success-filtering step was applied. The proposed method's counts are explicit—42 playbacks for pick-and-place, 18 for wiping—while the baseline counts are absent. If the baseline came from only the two original demonstrations per speed, the gains in Tables 1–4 could be a dataset-size effect rather than the effect of real-world speed-dependent reactions. The stress-test note lands on exactly this point, and the paper itself does not refute it. This is an omission, not necessarily misconduct, but it is the weakest link in the inference.\n\nThe other soft spots are smaller but real. Every cell is five trials, with no confidence intervals or significance tests; the wiping result is mixed, with the 12 cm condition actually worse under the proposed method (63% vs 86%). The paper says only successful playbacks were used for training, but does not say whether the baseline underwent any analogous selection. No code or data are released, so the hardware experiments cannot be independently checked.\n\nOn balance, the idea is coherent and the evidence is suggestive rather than conclusive. The paper deserves a serious referee, but the referee should ask for a full description of the baseline dataset, a statistical treatment of the success rates, and ideally a release of data or code.\n\nRecommendation: send to peer review with a request for major revision. I would bring it to a reading group as a discussion piece, and I would cite it if I were working on data augmentation for force-controlled imitation learning, but only after the baseline question is answered.","headline":"A sensible real-world augmentation idea with strong pick-and-place evidence, but the baseline comparison needs to be reported before the central claim holds.","tokens_in":12101,"tokens_out":1130,"would_cite":false,"duration_ms":12845,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that physically replaying taught motions at different speeds—rather than just rescaling the recorded data—provides real-world augmentation that lets force-controlled imitation learning generalize across speeds from two…","keywords":["imitation learning","motion-copying system","data augmentation","bilateral control","force control","teaching-playback","variable speed","robot manipulation"],"falsifier":"Train the simple-duplication baseline on the same number of samples, with identical train/validation splits and speed labels, then test on the same held-out objects and speeds; if the success-rate gap disappears, physical reactions are not the cause.","tokens_in":11104,"feed_emoji":"🤖","tokens_out":7205,"duration_ms":63474,"temperature":0.7,"pith_summary":"This paper tries to show that a robot can learn contact-rich manipulation at many speeds from only two human demonstrations recorded at one speed, if the missing speeds are generated by physically replaying the taught motion and recording what actually happens. The central comparison is between this real-world playback augmentation and simply duplicating or time-scaling the original demonstration data. Across pick-and-place and wiping experiments with a bilateral control-based imitation-learning model, the real-world replay data gave higher task success and better tracking of the commanded duration or frequency, especially inside the trained speed range. The approach matters because force and contact reactions are hard to simulate, so the only reliable way to get variable-speed force data is to collect it from the real environment.","feed_headline":"Real-world replays at new speeds lift imitation success by 55%","feed_subtitle":"Physical playback at varied speeds feeds real force reactions to the learner, so two demos cover many speeds.","key_machinery":"The central mechanism is the motion-copying system, a teaching–playback method that replays recorded position and force commands at altered speeds while the real robot interacts with the environment. During replay the robot is not controlled by a neural network, so the recorded force and position responses are genuine physical reactions to the sped-up or slowed-down motion; these responses are then labeled by the commanded time or frequency and added to the training set of a bilateral control-based imitation-learning LSTM. The load-bearing step is the physical replay: it converts a software speed change into real-world reaction data that the network can learn from.","core_discovery":"The paper's claim is that variable-speed teaching–playback works as data augmentation for imitation learning with position–force control, and that the real-world reaction data it collects are worth more than the same command data modified in software. Using the motion-copying system, two fixed-speed demonstrations were replayed at 0.5x, 1x, and 2x speed for pick-and-place and at 0.5x, 1x, and 1.5x speed for wiping, with the resulting follower responses—including contact forces—recorded and labeled by the commanded duration or frequency. Compared with a baseline that simply duplicated and rescaled the original demonstrations, training on these playbacks raised overall pick-and-place success from 30% to 53% (and interpolation-range success from 31% to 88%), and wiping success from 74% to 79%, while also keeping actual completion times and wiping frequencies closer to the label. The paper interprets this as evidence that speed changes in a nonlinear physical environment produce reactions that cannot be reproduced by downsampling or simulation, and that collecting those real reactions is what improves variable-speed imitation.","pith_inferences":["A natural test of the mechanism would be to train the simple-duplication baseline on the same number of samples, with identical train/validation splits and speed labels, to confirm that the success gap is caused by real-world reactions rather than by a difference in data quantity.","Because the playback method keeps only successful replays as training data, the augmented dataset is also a filtered, higher-quality subset of trajectories; replicating that filter on the baseline would isolate the contribution of physical reactions from the contribution of trajectory selection.","The method's stated limitation is that it only varies speed, not position, so a plausible extension is to combine the same real-world reaction-collection idea with spatial variation or with simulation-based augmentation for variable positions.","The wiping results show the largest gain at the higher surface height, while the lower-height condition still fails on force-contact detection, suggesting the method's benefit is strongest when replay data make the contact phases of the task learnable."],"forward_implications":["Contact-rich manipulation at variable speeds becomes learnable from as few as two fixed-speed demonstrations, without simulation.","Interpolation between trained speeds benefits most: pick-and-place success in the interpolated range rose from 31% to 88%.","Task success and adherence to the commanded duration or frequency both improve when training data include real environmental reactions at the target speeds.","The augmentation applies to distinct contact-rich tasks—grasping, carrying, placing, and continuous wiping—suggesting it generalizes across manipulation procedures.","Adding more diverse playbacks or combining with self-supervised learning is a stated path to finer speed control, particularly for extrapolation beyond the trained speeds."],"supporting_citations":[{"why":"Defines the motion-copying system, the teaching–playback method with variable-speed position–force replay that the proposed augmentation is built on.","marker":"[10]"},{"why":"Introduces bilateral control-based imitation learning with position and force information, the learning method the augmentation is tested with.","marker":"[3]"},{"why":"Supplies the prior formulation of imitation learning for variable-speed contact motion under bilateral control that this work extends.","marker":"[8]"},{"why":"Provides the self-supervised learning approach for speed variation and the controller parameters reused in the experiments.","marker":"[9]"},{"why":"Establishes fast-forward data collection of real environmental reactions through teaching–playback, the effect of which this paper evaluates at multiple speeds.","marker":"[11]"},{"why":"Supplies the downsampling and rearrangement procedure used to convert 500 Hz robot data into 50 Hz training samples for the LSTM model.","marker":"[25]"},{"why":"Provides the F2FL autoregressive network structure with command compensation that the imitation-learning model is based on.","marker":"[36]"}],"fun_headline_variants":["Variable-speed playback gives imitation a 55% boost","Real-world replay at new speeds lifts robot imitation","Speed-shifted demos improve imitation learning accuracy","Teach robots with variable-speed real reactions","Physical playback at variable speeds aids imitation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the baseline, simple duplication and speed changes applied to the original demonstrations, was matched to the real-world playback method in sample count and training conditions, since the paper does not report how many baseline samples were used.","fun_headline_variants_meta":{"raw":{"variants":["Variable-speed playback gives imitation a 55% boost","Real-world replay at new speeds lifts robot imitation","Speed-shifted demos improve imitation learning accuracy","Teach robots with variable-speed real reactions","Physical playback at variable speeds aids imitation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000325,"raw_usage":{"total_tokens":1835,"prompt_tokens":970,"completion_tokens":865,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":797}},"tokens_in":586,"tokens_out":865,"duration_ms":9395,"temperature":1.0,"reasoning_tokens":797,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:37:00.123873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the simple-duplication baseline on the same number of samples, with identical train/validation splits and speed labels, then test on the same held-out objects and speeds; if the success-rate gap disappears, physical reactions are not the cause.","supporting_citations":[{"cited_title":"Motion copying system based on real-world haptics in variable speed","cited_arxiv_id":null,"evidence_quote":"Defines the motion-copying system, the teaching–playback method with variable-speed position–force replay that the proposed augmentation is built on."},{"cited_title":"Imitation Learning for Object Manipulation Based on Position/Force Information Using Bilateral Control","cited_arxiv_id":null,"evidence_quote":"Introduces bilateral control-based imitation learning with position and force information, the learning method the augmentation is tested with."},{"cited_title":"Imitation learning for variable speed contact motion for operation up to control bandwidth","cited_arxiv_id":null,"evidence_quote":"Supplies the prior formulation of imitation learning for variable-speed contact motion under bilateral control that this work extends."},{"cited_title":"Imitation Learning for Nonprehensile Manipulation Through Self-Supervised Learning Considering Motion Speed","cited_arxiv_id":null,"evidence_quote":"Provides the self-supervised learning approach for speed variation and the controller parameters reused in the experiments."},{"cited_title":"Practical implementations of bilateral control-based imitation learning at irex2023","cited_arxiv_id":null,"evidence_quote":"Establishes fast-forward data collection of real environmental reactions through teaching–playback, the effect of which this paper evaluates at multiple speeds."},{"cited_title":"From virtual demonstration to real- world manipulation using lstm and mdn","cited_arxiv_id":null,"evidence_quote":"Supplies the downsampling and rearrangement procedure used to convert 500 Hz robot data into 50 Hz training samples for the LSTM model."},{"cited_title":"A New Autoregressive Neural Network Model with Command Compensation for Imitation Learning Based on Bilateral Control","cited_arxiv_id":null,"evidence_quote":"Provides the F2FL autoregressive network structure with command compensation that the imitation-learning model is based on."}],"review_version":1}