{"id":"80c172c5-9235-496a-b860-160112cab935","arxiv_id":"2412.10631","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Live augmented-reality feedback from a virtual robot raises the hardware replay success of barehanded human demonstrations from 1.3% to 71.1%, enabling robot-free data collection for imitation learning.","lead":"A new Apple Vision Pro app shows users a virtual robot that mirrors their hand movements in real time, letting them record robot-ready demonstrations without touching a physical robot. In a 15-person study, trajectories collected with this live feedback replayed successfully on a real robot 71% of the time, versus 1% without feedback, pointing toward a cheaper way to build large robot-training datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The No Feedback vs Feedback comparison in Section IV-A changes both AR feedback and task instructions, so the causal role of the visualization is confounded by instruction.","rationale":"The reader's weakest_assumption identified the fixed order of conditions as the main threat to causal attribution, which is valid but secondary. The more fundamental design flaw is that the No Feedback and Feedback conditions differ not only in feedback but also in task instructions: No Feedback asks for natural human motion, while Feedback asks the user to control the virtual robot. Thus the observed jump from 1.3% to 71.1% cannot be cleanly attributed to the AR visualization. The instruction confound invalidates the headline comparison even if order were counterbalanced. The Feedback vs Post Feedback comparison partially controls for instruction, but because Post Feedback always follows Feedback, it cannot isolate the live-feedback effect from practice. The system is clearly described, and the substantial Feedback vs Post Feedback gap suggests that live visualization does help, so the paper need not be rejected. However, the central causal claim is not established by the reported experiments, and the manuscript should either add the proposed control condition or substantially weaken its wording. This supports the reader's CONDITIONAL verdict, now for a different and more fundamental reason.","tokens_in":10842,"tokens_out":7690,"duration_ms":73503,"concrete_test":"See above.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-A defines the No Feedback condition as instructing users 'to demonstrate the task with natural human motion,' while the Feedback condition instructs users 'to demonstrate the task such that the virtual robot is controlled to execute the task.' This changes two variables simultaneously: the presence of AR feedback and the instruction to internalize robot constraints. The near-zero No Feedback replay success (1.3% average) may therefore reflect that participants were told to act naturally, not that they lacked feedback. The Post Feedback condition holds the instruction constant, but it always follows Feedback, so order and practice effects confound its comparison with Feedback. The Feedback vs Post Feedback gap (71.1% vs 33.3%) suggests live feedback provides real benefit, but the absolute 'critical' claim is supported primarily by the confounded No Feedback vs Feedback contrast. Qualitative results (slowing down, adjusting poses, accounting for gripper) could be responses to the instruction to control the robot rather than to the visualization. This is a load-bearing internal validity issue for the paper's central causal assertion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ARMADA, an Apple Vision Pro application that overlays a real-time simulated robot (a digital twin) on the user's view of their own hands, so that a human can provide barehanded manipulation demonstrations that are meant to be compatible with a physical robot. The authors report a user study with 15 participants, 3 tasks (Pick Tissue, Declutter, Bimanual Wipe), 5 starting states per task, and 3 feedback conditions (No Feedback, Feedback, Post Feedback), yielding 675 demonstrations that are directly replayed on physical Interbotix ViperX 300 arms. The central claim is that live AR feedback of the virtual robot is critical for collecting robot-free demonstrations that are directly replayable on real hardware, supported by Table I, which shows average replay success of 1.3% for No Feedback, 71.1% for Feedback, and 33.3% for Post Feedback. The paper also reports survey data suggesting the visualization is intuitive and useful, and describes the system architecture, IK-based control from hand tracking, and AR constraint visualizations (singularity, workspace, speed).","tokens_in":11065,"tokens_out":3270,"duration_ms":31657,"significance":"If the central claim is correct, ARMADA would be a notable step toward scalable imitation-learning data collection without physical robot access: it uses a consumer headset, requires no robot hardware during data collection, and the reported 1.3% to 71.1% improvement in direct replay success is a striking and practically meaningful effect. The evaluation has real strengths: an external, physical-hardware benchmark with success criteria defined independently of the system, a within-subject design with 15 participants, and 675 total demonstrations. The inclusion of a Post Feedback condition and qualitative reports is also valuable. However, the causal claim that feedback itself is critical is currently undermined by two design confounds: the No Feedback condition changes both feedback and task instruction, and the fixed condition order without counterbalancing confounds feedback with practice and fatigue. The absence of statistical inference further weakens the quantitative claims. The potential significance is high, but the current evidence does not yet cleanly separate the effect of feedback from these confounds.","major_comments":[{"comment":"The comparison that carries the paper's central claim contrasts No Feedback with Feedback while simultaneously changing the task instruction: participants in No Feedback are told 'to demonstrate the task with natural human motion,' whereas participants in Feedback are told 'to demonstrate the task such that the virtual robot is controlled to execute the task.' The near-zero No Feedback replay success (1.3% average in Table I) could therefore reflect the instruction to behave naturally rather than the absence of visual feedback. Please add a control condition that holds the instruction constant, for example by instructing participants to move as if controlling a virtual robot while showing them no visualization, or by counterbalancing the instruction wording across feedback conditions.","section":"IV-A"},{"comment":"All participants experience the three conditions in the same fixed order (No Feedback, then Feedback, then Post Feedback), and the task order, although randomized per participant, is held fixed across conditions. As a result, Feedback is always performed after 15 prior demonstrations and Post Feedback after 30, so practice and fatigue effects are perfectly confounded with feedback type. The interpretation that Post Feedback improves over No Feedback due to learning from feedback is confounded by the additional demonstration experience, and the Feedback vs Post Feedback gap may be inflated by fatigue. A counterbalanced design, or a control group that performs repeated No Feedback blocks with the same number of demonstrations, is needed to support the causal role of feedback.","section":"IV-C"},{"comment":"The text repeatedly states that Feedback 'significantly' outperforms No Feedback and Post Feedback, but no statistical tests are reported. With 15 participants and only 5 trials per condition-task cell, the standard deviations are large (e.g., 21.7% for Declutter under Feedback), so the observed differences may not be statistically reliable without a paired analysis. Please report within-participant paired tests (e.g., Wilcoxon signed-rank) or a mixed-effects model with participant random effects, and provide effect sizes. This is needed to support the quantitative 'significantly higher' claims in Section V-A.","section":"V-A, Table I"}],"minor_comments":[{"comment":"The sentence 'their utility is bottlenecked the ability to translate' is missing the word 'by'; it should read 'bottlenecked by the ability to translate.'","section":"II-A"},{"comment":"Figure references are inconsistently capitalized: the text uses 'fig. 2' and 'fig. 3' in some places and 'Figure 4' and 'Figure 5' in others; please standardize.","section":"III-A, III-C"},{"comment":"The paper states that code 'will be available on the project website when finalized' but provides no repository or release timeline in the current version. For a systems paper whose contribution includes the data collection pipeline, making code and the 675 collected trajectories available at the time of publication would substantially aid reproducibility.","section":"I"},{"comment":"The survey results report Likert means with standard deviations but do not include the full distribution or the exact number of responses per item; a supplementary table with the per-item response counts would be clearer.","section":"V-A"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a robotics or human-robot interaction venue. The core idea is interesting and the physical replay evaluation is a good choice, but the current experimental design does not separate the feedback visualization from instructions or practice, which are load-bearing for the 'feedback is critical' claim. I would encourage the authors to run a counterbalanced follow-up or a No-Feedback-with-robot-control-instruction control condition, and to add statistical inference. If the authors prefer not to run a new study, they should substantially soften the causal language and present the results as suggestive rather than definitive. The paper would also benefit from releasing the code and data to back the scalability claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-executed systems paper with a genuine empirical contribution—live AR feedback raises direct hardware replay success from about 1% to 71%—and the main threat to the headline claim is real but narrower than it first appears. The reader's conditional verdict is fair, and the stress-test note adds a valid point: the No Feedback condition in Section IV-A also changes the instruction, not just the visualization. That said, the Feedback vs Post Feedback contrast holds instructions fixed and still shows a large drop (71% to 33%), so live feedback is clearly doing causal work. What's actually new is the specific quantitative finding: naive natural human motion without any robot feedback is essentially unreplayable (1.3%), while the same motion shaped by an AR digital twin becomes replayable on physical hardware. The paper also deserves credit for disclosing the fixed order (No Feedback → Feedback → Post Feedback) in Section IV-C rather than hiding it. The evaluation is concrete—15 participants, 3 tasks, 675 demonstrations, physical replay with independent success criteria—and there is no circularity from fitted parameters. The soft spots: the lack of counterbalancing means the No Feedback baseline likely underestimates what practice alone could achieve, and the instruction change makes the 1.3% figure not a pure \"no feedback\" condition. A revision could add a counterbalanced control group, analyze per-trial learning curves, or at minimum discuss the magnitude of the likely practice effect. Code is not released, and some hand-tuned parameters (gripper threshold, wrist offset) limit reproducibility until the code ships. Survey results are self-report but not load-bearing. Overall: the central claim—that live AR feedback is critical for collecting robot-free replayable data—is supported in direction, though the \"critical\" part is partly confounded in magnitude. This paper deserves a serious referee. I would send it to review and ask for the order analysis or an explicit acknowledgment of the confound's limits. Who benefits: anyone working on scalable data collection for imitation learning, AR/VR teleoperation, or human-robot interfaces. I would cite it for the empirical result and the system design.","headline":"Solid systems paper with a real empirical result, but the headline causal claim is partly confounded by a fixed condition order and an instruction change; worth serious review.","tokens_in":11565,"tokens_out":1398,"would_cite":true,"duration_ms":14078,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ARMADA claims that real-time AR feedback of a simulated robot is critical for collecting robot-free demonstrations that replay directly on physical hardware, lifting average replay success from 1.3% to 71.1%.","keywords":["augmented reality","robot-free data collection","imitation learning","digital twin","teleoperation","Apple Vision Pro","replay success","human demonstrations"],"falsifier":"Counterbalance the condition order so half the participants receive Feedback first, then No Feedback; if the feedback effect is causal, the No Feedback condition should still yield near-zero replay success regardless of prior practice, and the Feedback condition should still exceed 70%.","tokens_in":10670,"feed_emoji":"🥽","tokens_out":6730,"duration_ms":54337,"temperature":0.7,"pith_summary":"ARMADA is a system for collecting robot-manipulation demonstrations without a physical robot: a demonstrator wears an Apple Vision Pro, performs tasks barehanded, and watches a simulated robot arm mirror their hand in real time in augmented reality. The paper's central claim is that this live feedback is what makes the collected trajectories usable, since direct replay on physical hardware succeeds 71.1% of the time on average with feedback versus 1.3% without it. It also reports that demonstrations performed after a feedback session, with the overlay hidden, still replay 33.3% of the time, suggesting some learned transfer. If true, this removes the teleoperation-hardware bottleneck and opens a route to large-scale human data collection for imitation learning.","feed_headline":"Robot-free demos jump from 1.3% to 71.1% with AR feedback","feed_subtitle":"A Vision Pro app overlays a simulated robot on bare hands, making human demos replayable on real hardware.","key_machinery":"The load-bearing mechanism is the closed 30 Hz feedback loop: the headset's camera-based hand tracking estimates wrist and knuckle poses, an inverse-kinematics solver maps that pose to joint commands for a simulated six-degree-of-freedom arm, and the resulting robot state is rendered as an AR overlay in the headset. Gripper open/close is mapped from thumb-index distance, so the human hand becomes a two-finger pincher in simulation. The system also renders constraint cues — a yellow color shift near singularities, a red wall at workspace limits, and a 'Slow Down' alert — that train the demonstrator to stay inside robot-compatible motion.","core_discovery":"On the paper's own terms, the discovery is that a real-time augmented-reality digital twin of a robot, overlaid onto a human's bare hands, turns otherwise unusable human motion into hardware-compatible robot trajectories. With ARMADA, 15 participants gave 675 demonstrations across three tasks; average replay success rose from 1.3% without feedback to 71.1% with feedback, with per-task gains of 76, 48, and 85 percentage points. After feedback was removed, success remained at 33.3%, which the authors read as evidence that people internalize some of the robot's constraints from the feedback session. The failures that remain with feedback are mostly small position and orientation errors from AR depth perception.","pith_inferences":["If the causal role of live feedback holds up under counterbalancing, the feedback itself could be made adaptive — for instance, showing only corrective cues — to push replay success beyond the 71% reported here.","A natural next experiment is to train a policy on ARMADA-collected trajectories and compare it against a policy trained on teleoperation data of the same tasks, which would test whether robot-free data closes the quality gap in downstream imitation learning.","The Post Feedback retention suggests a two-stage pipeline: teach demonstrators with AR once, then collect barehanded without headsets, which could scale collection beyond the number of available Vision Pro headsets."],"forward_implications":["Anyone with a Vision Pro can contribute demonstrations without owning a robot, removing the hardware bottleneck that limits imitation-learning data collection.","Because trajectories replay on real arms directly, ARMADA sidesteps the embodiment-gap problem faced by human-video approaches.","The Post Feedback result implies a short AR training session can improve barehanded data collection even when the overlay is later switched off, potentially lowering the cost of large-scale collection.","The plug-and-play interface means the same system can serve different robot arms by swapping 3D models and reconfiguring the communication layer, not just the specific ViperX arms tested."],"supporting_citations":[{"why":"The closest prior system for collecting demonstrations without a robot; it lacks real-time robot feedback, so it serves as the baseline this work improves upon.","marker":"[33]"},{"why":"Concurrent augmented-reality data collection system aimed at simulation, used to contrast with the replay-on-hardware result.","marker":"[49]"},{"why":"Concurrent AR feedback system that relies on extra hardware; used to highlight the barehanded, headset-only advantage.","marker":"[51]"},{"why":"Supplies the inverse-kinematics solver that maps tracked hand pose to robot joint commands in the feedback loop.","marker":"[52]"},{"why":"Existing teleoperation dataset that motivates the scalability claim for robot-free collection.","marker":"[14]"},{"why":"Recent technique for transforming egocentric observations into robot observations, cited as the path to policy learning from ARMADA data.","marker":"[53]"}],"fun_headline_variants":["AR feedback turns 1.3% robot-free demos into 71.1%","Barehanded demos become robot-replayable via AR twin","Vision Pro AR: robot-free demos jump to 71.1% success","AR boosts demo replay from 1.3% to 71.1% on real robots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes the jump from 1.3% to 71.1% replay success is caused by the AR feedback itself, but every participant experienced No Feedback first, then Feedback, then Post Feedback, so practice and growing familiarity with the tasks could also explain part of the gain.","fun_headline_variants_meta":{"raw":{"variants":["AR feedback turns 1.3% robot-free demos into 71.1%","Barehanded demos become robot-replayable via AR twin","Vision Pro AR: robot-free demos jump to 71.1% success","AR boosts demo replay from 1.3% to 71.1% on real robots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":1896,"prompt_tokens":829,"completion_tokens":1067,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":977}},"tokens_in":445,"tokens_out":1067,"duration_ms":8564,"temperature":1.0,"reasoning_tokens":977,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:45:05.660521+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Counterbalance the condition order so half the participants receive Feedback first, then No Feedback; if the feedback effect is causal, the No Feedback condition should still yield near-zero replay success regardless of prior practice, and the Feedback condition should still exceed 70%.","supporting_citations":[{"cited_title":"Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system,","cited_arxiv_id":null,"evidence_quote":"The closest prior system for collecting demonstrations without a robot; it lacks real-time robot feedback, so it serves as the baseline this work improves upon."},{"cited_title":"Su, X.-Q","cited_arxiv_id":null,"evidence_quote":"Concurrent AR feedback system that relies on extra hardware; used to highlight the barehanded, headset-only advantage."},{"cited_title":"Slow Down","cited_arxiv_id":null,"evidence_quote":"Supplies the inverse-kinematics solver that maps tracked hand pose to robot joint commands in the feedback loop."},{"cited_title":"Bert: Pre-training of deep bidirectional transformers for language understanding,","cited_arxiv_id":null,"evidence_quote":"Existing teleoperation dataset that motivates the scalability claim for robot-free collection."},{"cited_title":"Virtual kinesthetic teaching for bimanual telemanipulation,","cited_arxiv_id":null,"evidence_quote":"Recent technique for transforming egocentric observations into robot observations, cited as the path to policy learning from ARMADA data."}],"review_version":1}