{"id":"70e06f02-6fdb-497c-8082-9a34feb2f370","arxiv_id":"2411.09929","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A diffusion-policy visuomotor controller trained on 300 handheld demonstrations achieved 28.95% success harvesting peppers in an unprotected outdoor field, with a 31.7s cycle time.","lead":"This paper reports a robot that harvests peppers in outdoor fields using imitation learning from 300 human demonstrations. It achieved a 28.95% harvest success rate, comparable to greenhouse systems, and shares an open dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The success rate claimed as 'comparable to greenhouse systems' is measured only through the policy-controlled approach/cut/grasp/retract segment; the open-loop placement phase and manual setup are excluded, so Table II compares different task scopes.","rationale":"The paper's strongest point is that 300 handheld demonstrations can transfer to a UR5e/Husky and produce measurable harvest attempts in an outdoor field; the released dataset and the 221-trial evaluation are real contributions. I did not find an internal inconsistency in the diffusion-policy pipeline. The load-bearing weakness is comparative, not internal: the headline 'comparable' uses Table II, but Table II's rows are not known to measure the same thing. Even if every difficulty label is accepted, our success count stops at retraction and our cycle time stops before placement and excludes manual setup. If prior systems count only detachment, the comparison might survive; if they count deposit into a container, our end-to-end rate is lower. This is a testable question, so the honest verdict is CONDITIONAL: accept the feasibility/manipulation claim, but require the authors to either measure end-to-end success and cycle time on the same boundary as the baselines or explicitly scope the claim to the manipulation stage. This matches the reader's conditional verdict, hence UNCHANGED.","tokens_in":10697,"tokens_out":8330,"duration_ms":94361,"concrete_test":"Determine the success criteria and cycle-time boundaries in the methods sections of [18], [19], and [20]. If any baseline requires the harvested pepper to be placed in a container or includes detection/approach in its cycle time, recompute our comparable numbers on that same boundary: multiply the 64 policy successes by the grasp detector's recall (0.78 from Fig. 10) and by the open-loop placement success rate (unreported; measure it or re-run the placement phase), and add manual setup time to the 31.71s cycle time. If the corrected success rate leaves the range of baseline unmodified rates (4-25%) or the corrected cycle time exceeds the baselines, the claim must be downgraded to 'manipulation-stage success comparable,' not harvesting success.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. VI-B.1 defines success for the policy as 'approach, cut, and grasp a pepper ... without considering the open-loop placement phase.' Sec. V-C states each trial starts with a manual setup in which the Husky is positioned and the end-effector is adjusted by joystick until the target pepper is in view. So 28.95% and 31.71s describe only the policy-controlled manipulation segment on pre-selected, roughly centered targets, not autonomous harvesting. The prior rows in Table II are harvesting-success rates for greenhouse systems, and nothing in the paper shows their success criteria and cycle times exclude the detection, positioning, or placement stages. The abstract's 'comparable to existing systems tested under more controlled conditions like greenhouses' therefore rests on different task boundaries. The Sec. VI-C difficulty mapping (easy=modified, medium/hard=unmodified) is an additional unvalidated assumption, but the metric-scope mismatch alone is enough to make the headline comparison unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an imitation-learning-based robotic system for autonomous pepper harvesting in an unprotected outdoor field. The authors collect 300 human demonstrations with a custom handheld shear-gripper, train a diffusion policy visuomotor controller, and deploy it on a UR5e arm mounted on a Husky UGV. They evaluate the policy on 221 trials, reporting a 28.95% success rate (64/221) and an average cycle time of 31.71 s, and they compare these numbers with prior greenhouse systems in Table II. The paper also describes a robust fiducial-cube tracking method for data collection, a grasp detector for triggering open-loop placement, and a public dataset release.","tokens_in":10918,"tokens_out":3404,"duration_ms":32488,"significance":"If the reported results are taken at face value, the paper makes a useful contribution by demonstrating that imitation learning with a low-cost handheld data-collection device can produce a visuomotor policy that generalizes across outdoor field conditions, lighting variability, and crop diversity. The evaluation is transparent in reporting raw trial counts and held-out test rows, and the release of the 300-demonstration dataset is a concrete benefit to the community. The central feasibility claim—that a diffusion policy can perform the approach-cut-grasp-retract manipulation phase in an unstructured field—is supported by the data. However, the headline comparability to greenhouse systems is not yet established, because the task scope and difficulty definitions differ from the prior works in ways the paper does not validate.","major_comments":[{"comment":"The success metric reported as 28.95% is defined only over the policy-controlled approach, cut, grasp, and retract phases, explicitly excluding the open-loop placement phase and starting after a manual setup in which a human positions the Husky and joysticks the end-effector until the target pepper is in view. The prior systems in Table II report harvesting success rates for complete harvest cycles that include detection, positioning, and placement. The abstract's statement that the system is 'comparable to existing systems tested under more controlled conditions like greenhouses' therefore rests on different task boundaries. The authors should either present a like-for-like comparison by including the placement and setup phases in the success definition, or explicitly qualify the comparison as applying only to the manipulation phase.","section":"Sec. VI-B.1 and Sec. V-C"},{"comment":"The mapping of the authors' 'easy' difficulty category to the 'modified peppers' of prior works and of 'medium/hard' to 'unmodified peppers' is an unvalidated assumption. The prior works [18]–[20] modified peppers through pruning and repositioning, while the authors' difficulty categories are based on peduncle shape and occlusion. The paper provides no evidence that these two axes of difficulty correspond across studies. Because the entire comparability claim in Table II depends on this equivalence, the authors should either justify the mapping with quantitative similarity measures (e.g., comparable occlusion statistics or peduncle orientations) or remove the cross-study comparison and report only the internal difficulty breakdown.","section":"Sec. VI-C and Table II"}],"minor_comments":[{"comment":"The hard-difficulty row reports only 2 successes out of 27 trials (7.31%); a binomial confidence interval would help the reader judge the stability of this estimate and would strengthen the internal comparison across difficulty levels.","section":"Sec. VI-B.1"},{"comment":"The entry for [19] gives '25%' without a trial count; for consistency, provide the raw count or state explicitly that it is not available.","section":"Table II"},{"comment":"Equation (1) omits the variance schedule of the DDPM sampling process; clarifying the notation and the schedule would improve reproducibility.","section":"Sec. III-A"},{"comment":"The SSIM threshold is a free parameter; state its value and, if possible, its sensitivity, since it directly affects the pose-tracking quality used for demonstration data.","section":"Sec. IV-B"},{"comment":"There are minor grammatical errors, e.g., 'Advancements in robotics and deep learning is driving' should be 'are driving', and in Sec. III-A 'Diffusion policy offer' should be 'Diffusion policy offers'.","section":"Sec. II"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the dataset release is a tangible asset. The central feasibility result is plausible, but the abstract's comparability claim needs to be reconciled with the task-scope mismatch; this is a substantive correctness issue for the headline contribution, not merely a wording concern. The authors may resolve it by either qualifying the comparison or expanding the evaluation to include placement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first genuine outdoor field evaluation of a diffusion-policy visuomotor controller for pepper harvesting, and it ships a 300-demonstration dataset. The headline claim of being 'comparable to greenhouse systems' does not survive a close reading of what is actually being measured, but the core feasibility claim does.\n\nWhat is genuinely new and useful: the handheld shear-gripper with the fiducial-cube pose refinement pipeline is a practical workaround for SLAM failure in dynamic agricultural scenes, and the quantitative tracking comparison (MSE dropping from 0.276/0.358 to 0.003/0.018) is credible. The held-out-row evaluation with raw trial counts (64/221, with per-difficulty breakdown) is transparent. The released dataset is a real contribution for a domain where data is scarce.\n\nThe soft spots are real but localized. The success metric for the policy is defined as 'approach, cut, and grasp a pepper' without considering the open-loop placement phase, and each trial begins with manual positioning and joystick adjustment to bring the pepper into view. So 28.95% and 31.71s describe only the policy-controlled manipulation segment on pre-selected, roughly centered targets. Table II then compares this to greenhouse systems that likely include detection and placement in their success counts. Nothing in the paper shows those prior systems exclude those stages, so the comparison is apples-to-oranges. The difficulty mapping (easy=modified, medium/hard=unmodified) is asserted without validation; it may be reasonable, but it is not demonstrated. No confidence intervals on the success rates either. These issues do not sink the feasibility claim—the raw data speak for themselves—but they do mean the abstract's comparability sentence is overreach.\n\nWho should read this: anyone working on imitation learning for agriculture or on evaluation standards for field robotics. It is a useful data point and a cautionary example of how task scope can shift when comparing field and greenhouse results. A serious referee should engage with it, not desk-reject it, but with the expectation that the comparison be redone or reframed and the metric scope stated honestly.\n\nMy recommendation: send it to review, with a request for a revised Table II and a more careful statement of what the metric covers.","headline":"First real outdoor field study of diffusion-policy pepper harvesting with a released dataset; the feasibility claim holds, but the headline comparability to greenhouse systems rests on mismatched task boundaries.","tokens_in":11357,"tokens_out":1332,"would_cite":true,"duration_ms":15792,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trained on 300 human demonstrations with a custom handheld shear-gripper, a diffusion-policy robot autonomously approaches, cuts, and grasps peppers in an unprotected field, reaching 28.95% success at 31.71s per cycle.","keywords":["pepper harvesting","imitation learning","diffusion policy","visuomotor policy","agricultural robotics","outdoor field deployment","fiducial cube tracking","grasp detection"],"falsifier":"Re-run the same system while counting a trial as successful only when the pepper ends up in the crate, and have independent annotators label peppers using the greenhouse studies' criteria for modified versus unmodified; if the easy-only success rate no longer tracks the modified-pepper rates, or the medium/hard rate no longer tracks the unmodified rates, the 'comparable to greenhouse systems' claim fails.","tokens_in":1385,"feed_emoji":"🌶️","tokens_out":1352,"duration_ms":61389,"temperature":0.7,"pith_summary":"The paper claims that imitation learning can carry robotic pepper harvesting out of the greenhouse and into an open, unstructured field. The authors collect 300 human demonstrations with a custom handheld shear-gripper, train a diffusion-policy visuomotor controller, and deploy it on a mobile manipulator in a working pepper field. The reported total success rate is 28.95% (64/221 trials) with a cycle time of 31.71 seconds, which they argue is comparable to previous greenhouse-based pepper-harvesting systems. The comparison rests on tagging easy peppers as equivalent to greenhouse studies' modified peppers and medium/hard peppers as equivalent to unmodified ones, and on defining success as approach/cut/grasp rather than pepper-in-crate.","feed_headline":"Robot learns to harvest peppers outdoors at 29% success","feed_subtitle":"Trained on 300 human demos, the system equals greenhouse harvesters despite wind, glare, and heavy occlusion.","key_machinery":"The core mechanism is the diffusion policy, a conditional denoising diffusion model that maps observation sequences (RGB images, gripper pose, and gripper actuation) to action sequences of gripper pose and actuation. To make outdoor demonstration collection feasible, the authors replace SLAM-based localization with a fiducial cube tracked by an external camera, using a six-step pose-refinement pipeline with SSIM filtering and corner refinement that reduces pose estimation error by about two orders of magnitude. A custom handheld shear-gripper, designed to mirror the robot's end-effector, collects the demonstrations, and a small CNN grasp detector decides when to hand control to an open-loop placement motion.","core_discovery":"The central claim is that a diffusion-policy visuomotor model trained on 300 demonstrations collected with a custom handheld shear-gripper can autonomously approach, cut, and grasp peppers in an unprotected outdoor field, achieving a 28.95% success rate with a 31.71-second cycle time. The authors argue this is comparable to existing greenhouse systems, which operate under more controlled conditions. Success degrades with task difficulty: 42.57% for easy peppers, 20.21% for medium peppers, and 7.31% for hard peppers. The system also includes a separate CNN grasp detector that triggers an open-loop placement phase; the reported success metric covers approach, cut, and grasp but not the placement of the pepper into the crate.","pith_inferences":["A consequence the paper leaves implicit is that the released 300-demonstration dataset may become the more durable contribution, giving the field a shared outdoor benchmark for imitation-learning harvesters.","The paper's failure analysis suggests a concrete next test: condition the policy on an explicit peduncle detection signal; most of the 157 unsuccessful trials should shift into the success column if the bottleneck is fruit-targeting rather than actuation.","The reported emergence of re-grasping behavior not present in demonstrations implies diffusion policies can synthesize recovery strategies from demonstration data alone, a testable hypothesis for other deformable-object manipulation tasks such as berry picking."],"forward_implications":["A single diffusion policy trained on 300 demonstrations generalizes across plots, lighting conditions, and plant variability in an outdoor field.","Policy success drops as difficulty rises: 42.57% easy, 20.21% medium, 7.31% hard, identifying peduncle shape and occlusion as the main bottlenecks.","The overall 28.95% success and 31.71-second cycle time are comparable to greenhouse pepper harvesters, despite the added difficulty of an unprotected field.","Replacing SLAM-based localization with refined fiducial-cube tracking makes demonstration collection feasible outdoors, reducing pose error by about two orders of magnitude.","The grasp detector's 0.83 accuracy, 0.71 precision, and 0.78 recall allow the policy to hand off to an open-loop placement controller."],"supporting_citations":[{"why":"Supplies the handheld Universal Manipulation Interface demonstration-collection framework that the authors adapt with a shear-gripper and robust fiducial-cube tracking.","marker":"[3]"},{"why":"Supplies the diffusion-policy visuomotor architecture that maps observations to action sequences and is the core learned controller.","marker":"[4]"},{"why":"Provides the greenhouse harvesting baseline with modified and unmodified success rates that Table II uses for comparison.","marker":"[18]"},{"why":"Provides the glasshouse harvesting baseline whose reported modified/unmodified success rates anchor the comparison in Table II.","marker":"[19]"},{"why":"Provides the double-row greenhouse baseline whose success rates and cycle time are compared with the outdoor field results.","marker":"[20]"},{"why":"Supplies ORB-SLAM3, the visual-inertial SLAM baseline whose trajectories are used to evaluate the accuracy of the refined fiducial-cube pose tracking.","marker":"[29]"}],"fun_headline_variants":["Robot trained on 300 demos picks peppers outdoors at 29%","Outdoor pepper harvesting: imitation learning scores 29% success","Autonomous pepper picker learns from humans, hits 29% in fields","Field robot uses imitation learning to harvest peppers, 29% success","Imitation learning gets robot to pick peppers outdoors, 29% success"],"cache_read_input_tokens":13696,"weakest_assumption_plain":"The comparability claim depends on the assumption that the paper's 'easy' peppers are equivalent to greenhouse studies' 'modified' peppers and that 'medium/hard' peppers are equivalent to 'unmodified' ones, an equivalence asserted without validation; success also excludes the open-loop placement phase.","fun_headline_variants_meta":{"raw":{"variants":["Robot trained on 300 demos picks peppers outdoors at 29%","Outdoor pepper harvesting: imitation learning scores 29% success","Autonomous pepper picker learns from humans, hits 29% in fields","Field robot uses imitation learning to harvest peppers, 29% success","Imitation learning gets robot to pick peppers outdoors, 29% success"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000859,"raw_usage":{"total_tokens":3665,"prompt_tokens":816,"completion_tokens":2849,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":2755}},"tokens_in":432,"tokens_out":2849,"duration_ms":18976,"temperature":1.0,"reasoning_tokens":2755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:08:36.361836+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same system while counting a trial as successful only when the pepper ends up in the crate, and have independent annotators label peppers using the greenhouse studies' criteria for modified versus unmodified; if the easy-only success rate no longer tracks the modified-pepper rates, or the medium/hard rate no longer tracks the unmodified rates, the 'comparable to greenhouse systems' claim fails.","supporting_citations":[{"cited_title":"Performance improve- ments of a sweet pepper harvesting robot in protected cropping environments,","cited_arxiv_id":null,"evidence_quote":"Provides the glasshouse harvesting baseline whose reported modified/unmodified success rates anchor the comparison in Table II."},{"cited_title":"Develop- ment of a sweet pepper harvesting robot. j field robot 37: 1027–1039,","cited_arxiv_id":null,"evidence_quote":"Provides the double-row greenhouse baseline whose success rates and cycle time are compared with the outdoor field results."}],"review_version":1}