{"id":"09d4a273-4448-4b60-b002-1f6639f27c7b","arxiv_id":"2505.09477","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A closed-loop LLM planning framework was field-tested on ground and aerial robots for kilometer-scale missions, including an onboard UAV planner built from a distilled small language model.","lead":"This paper reports field deployments of SPINE, a language-model planning framework, on ground and aerial robots in urban and rural environments. It also presents a preliminary on-device language model for UAV planning and shares lessons learned for deploying foundation models in the field.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Success counts include two UGV missions finished only after manual takeover; 12/14 does not demonstrate fully autonomous completion.","rationale":"The reader's weakest assumption—permissive success counts after manual takeover—is exactly the condition on which the central demonstration rests. The paper is otherwise transparent about its limitations, explicitly notes the 'brief manual takeover' and the obstacle-detection failure modes, and self-reports the small sample sizes and short-horizon distillation restriction. I considered whether the deeper concern is that the 'first' claims are unverifiable because they rely on self-cited [4],[18] and no released artifacts; that is real but secondary for a workshop field report and does not change the conditional verdict. The manual-takeover issue is more directly load-bearing because it changes the numerical evidence for the headline claim: 12/14 becomes 10/14 under a stricter autonomy criterion. The UAV 3/4 result has the same permissive structure (user-correctable third mission), which strengthens the concern but does not alter the conclusion. The appropriate action is to keep the CONDITIONAL verdict and require the authors to disambiguate assisted versus autonomous successes; no harsher verdict is warranted because the paper discloses the interventions and does not overstate them elsewhere.","tokens_in":11192,"tokens_out":6963,"duration_ms":71089,"concrete_test":"Recompute the UGV results from Table I separating fully autonomous completions from assisted completions: mark S7 and S8 as 'assisted' (and flag any other interventions, including communication or odometry recovery, from the run logs in [4] and [18]). Report the autonomous success rate; if it falls below 12/14, qualify the 'completed twelve of fourteen missions' sentence and soften the 'first large-scale demonstration' claim accordingly. For the UAV claim, the decisive check is a one-sentence clarification in the paper: state which model (GPT-4o or distilled Llama-3.2-3B) was running on the UAV during the four missions and whether the 3/4 count is strictly first-try autonomous execution or includes outputs that a human would need to correct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-A reports that SPINE completed twelve of fourteen missions, but also states that obstacle detection failed during S7 and S8 and that 'SPINE was able to complete the mission after brief manual takeover.' Table I lists these outcomes as 1/1 with 'Obst. det.' as the failure mode. If the success criterion is autonomous completion without human intervention, these two runs should be counted as assisted or failed, reducing the UGV success rate from 12/14 (86%) to 10/14 (71%). This matters because the paper's headline contribution is a 'first demonstration of large-scale LLM-enabled robot planning in unstructured environments,' and that demonstration is defined by these counts. The claim is not that a human-robot team accomplished the missions; it is that SPINE's LLM-based planning did. The same permissive logic appears in the UAV section: the third mission confused northern and southern parking lots and is described as 'easily corrected' by a follow-up command, so the 3/4 'first try' figure is not a strict autonomy metric either. Field robotics commonly reports interventions, and the disclosure is honest; the problem is that the success counting equates human-assisted completion with autonomous completion, which is the load-bearing assumption behind the 'first demonstration' and '12/14' claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports field deployments of SPINE, an LLM-based planning framework, on ground robots and a UAV in large-scale unstructured environments. The UGV portion summarizes missions from prior work by the same authors, reporting 12/14 missions completed in urban, semi-urban, and rural sites. The UAV portion presents preliminary distillation of Llama-3.2-3B using GPT-4o-generated planning data, with 3/4 short missions fulfilled on the first try and a planning comparison on 11 specifications among GPT-4o, distilled Llama-3.2-3B, and off-the-shelf Llama-3.2-3B. The paper claims the first demonstration of kilometer-scale LLM-enabled robot planning in unstructured environments and the first language-driven UAV planner using on-device language models, and it concludes with lessons learned and open challenges.","tokens_in":11481,"tokens_out":5948,"duration_ms":60682,"significance":"If the reported results are taken at face value, the kilometer-scale UGV deployments are a useful step toward FM-enabled autonomy outside prior-map, closed-world settings, and the distillation result is a promising proof of concept for SWaP-limited onboard planning. The paper is honest about failure modes (communication loss, odometry drift, obstacle detection) and explicitly discloses the manual takeovers in S7 and S8. The headline claims are nonetheless moderated by very small sample sizes, the conflation of autonomous and human-assisted successes, and the fact that the main UGV evidence is inherited from self-cited prior work rather than re-derived or fully described here. The negative result for off-the-shelf Llama-3.2-3B (9.2%) is a useful data point for the community, and the lessons-learned section is practical and appropriately cautious.","major_comments":[{"comment":"Table I lists S7 and S8 as successful outcomes with \"Obst. det.\" as the failure mode, and Section III-A states that \"SPINE was able to complete the mission after brief manual takeover.\" Because the abstract and the first contribution bullet define the headline result by the 12/14 success count, counting these two runs as unqualified successes conflates human-assisted completion with autonomous completion. The paper should report both the strict autonomy rate (10/14 if S7 and S8 are excluded) and the assisted-completion rate, and should state the operational definition of success used for each specification.","section":"III-A and Table I"},{"comment":"The comparison among GPT-4o (100%), distilled Llama-3.2-3B (72.7%), and off-the-shelf Llama-3.2-3B (9.2%) is based on eleven specifications with no trial-by-trial listing, no confidence intervals, and no repeated sampling. With n=11, the 8/11 versus 11/11 difference is not statistically robust, so the unqualified claim that distillation \"leads to a significant performance gain\" is stronger than the evidence supports. The authors should provide per-specification results, repeated trials or a statistical test, and should soften the wording to something like \"a promising improvement in this preliminary set.\"","section":"III-B and Table II"},{"comment":"The UAV evaluation covers only four short missions, and the third mission (construction in the northern parking lot) required a follow-up user command after the planner confused southern and northern parking lots. Presenting this as 3/4 \"first try\" and claiming \"the first language-driven UAV planner using on-device language models\" goes beyond what four missions can establish. The authors should define \"first try\" explicitly, report the corrected-outcome rate separately, or restrict the novelty claim to the specific distillation-plus-onboard-execution configuration that was demonstrated.","section":"III-B, UAV missions"},{"comment":"The paper's central \"first demonstration\" claim rests on UGV deployments that are summarized from [4] and [18] rather than fully described in this manuscript. Because the reader cannot verify trial conditions, intervention criteria, or the exact definition of success from the text alone, the authors should either include the trial protocol and per-trial data in an appendix or explicitly reframe this paper as a summary of prior UGV work whose new contribution is limited to the UAV distillation results.","section":"III-A and I"}],"minor_comments":[{"comment":"The text and Table II refer to Llama-3.2-3B, but the Fig. 7 captions refer to Llama 3.1 8B; please reconcile the model names and sizes.","section":"III-B and Fig. 7"},{"comment":"The abstract contains a typo: \"FM-enabled robots primary operate\" should read \"primarily operate.\"","section":"Abstract"},{"comment":"Specification S1 reports an outcome of 1/3 with an average distance of 1200 m; please clarify whether repeated attempts of the same specification are counted as separate trials and whether the distance is averaged over all attempts or only successful runs.","section":"III-A and Table I"},{"comment":"Fig. 5 is reproduced from [4], but this paper does not provide the experimental protocol (number of runs, prompt configuration, definition of \"unknown portion of map\") for the no-validation versus with-validation comparison; please summarize the protocol or clearly mark the figure as prior-work data.","section":"III-A and Fig. 5"},{"comment":"The phrase \"based on air router, first introduced in [22]\" should be written as \"based on Air Router, first introduced in [22]\" for consistency with the system name.","section":"II-B"}],"recommendation":"major_revision","confidential_remarks":"The success-counting issue is the main gate for publication: the paper's headline numbers mix autonomous runs with runs completed after manual takeover. The paper is otherwise an honest field report with useful negative results. I would support acceptance after the authors add a strict autonomy metric, trial-level details for the distillation evaluation, and appropriately softened novelty claims. The heavy reliance on self-cited prior work is acceptable, but it makes the requested separation between new results and summaries essential."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this is a workshop paper from the GRASP group. It is two things: a recap of SPINE's UGV deployments (already in their ICRA'25 paper and a Transactions on Field Robots submission) and a genuinely new preliminary result—distilling Llama-3.2-3B with LoRA and running it onboard a Falcon 4 UAV for language-specified missions. That part is real and does not appear in the cited literature: 3 of 4 short missions fulfilled on the first try, and a 72.7% planning success rate on 11 specs versus 9.2% for off-the-shelf Llama. It is thin, but it is a first, and it is reproducible in principle.\n\nWhat the paper does well: it names failure modes (communication loss, odometry drift, obstacle detection), discusses the communication infrastructure honestly, and makes a reasonable case for modular autonomy with a validation layer. Figure 5, showing the value of plan validation, is a nice point even if it is carried over from their prior work.\n\nThe soft spot is the headline counting. Section III-A says SPINE completed 12/14 UGV missions, but S7 and S8 required 'brief manual takeover' after obstacle detection failure, and Table I lists them as 1/1. If the success criterion is autonomous completion, that is 10/14 (71%), not 12/14. The disclosure is transparent—they do not hide the takeover—but the framing 'first demonstration of large-scale LLM-enabled robot planning' leans on the 12/14 number. The UAV section is a little cleaner: the third mission confused north and south parking lots and needed a follow-up command, and the paper counts it as not first-try, so the 3/4 figure is accurate.\n\nThe evaluation is small everywhere: 14 UGV missions across three sites, 4 UAV flights, 11 specs for the distillation comparison, no error bars, no trial-by-trial detail. That is normal for a field workshop report, but it means the claims are existence proofs, not measured performance. No code or data are released, which limits independent checking.\n\nThe citation pattern is heavily self-referential, but appropriately so: the framework and UGV results come from [4] and [18], and the paper says exactly that. This is not circular reasoning.\n\nWho is this for? People working on LLM-driven field robotics, especially anyone considering distillation for onboard planning. It deserves a serious referee as a systems report, not as an architectural contribution. I would send it to review, but with a request to fix the success counting or soften the 'first demonstration' language.\n\nRecommendation: engage with it, but read the success criteria first.","headline":"A genuinely first on-device LM UAV planner, but the headline 12/14 UGV success count includes two manual takeovers and is softer than it reads.","tokens_in":11995,"tokens_out":3788,"would_cite":true,"duration_ms":33923,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SPINE, a closed-loop LLM planner, completed 12 of 14 kilometer-scale field missions on ground robots, and a distilled 3-billion-parameter model planned UAV flights entirely onboard.","keywords":["foundation models","large language models","field robotics","semantic planning","model distillation","UAV planning","UGV navigation","closed-loop validation"],"falsifier":"Rerun the fourteen UGV missions with the same specifications and map conditions but forbid any manual takeover, and separately rerun the four UAV missions with no prior semantic map and no server connectivity; if the first-try success rates fall materially from 12/14 and 3/4, the demonstration claims would not survive.","tokens_in":11042,"feed_emoji":"🤖","tokens_out":10646,"duration_ms":93881,"temperature":0.7,"pith_summary":"This paper reports field deployments of SPINE, a closed-loop autonomy framework in which a large language model turns an incomplete natural-language mission into a sequence of robot behaviors. Across urban, semi-urban, and rural sites, the ground version completed twelve of fourteen missions, with failures attributed to communication loss and odometry drift; two missions involving obstacle-detection failures were completed only after a brief manual takeover. The paper also reports preliminary distillation results: a LoRA-fine-tuned Llama-3.2-3B model running on a UAV's onboard computer planned three of four short missions correctly on the first try, which the authors present as the first language-driven UAV planner with fully onboard computation. The underlying argument is that embedding foundation models in a validated closed loop, rather than generating plans from a known map, is what makes language-specified missions work in large, unstructured, partially known terrain. If the demonstration holds, foundation-model-enabled robots can move from indoor, map-known settings toward useful field autonomy.","feed_headline":"12 of 14 field missions completed by an LLM-guided robot","feed_subtitle":"A distilled 3B model also flew a drone from language alone, with no server connection.","key_machinery":"The central object is SPINE, a closed-loop LLM-enabled planner built from plan generation and plan validation. Plan generation uses an LLM to turn the user's natural-language mission into a task sequence expressed in the robot's behavior API, using chain-of-thought reasoning and receiving incremental map updates as text. Plan validation checks each proposed behavior against syntax, reachability, and explorability constraints and returns natural-language error feedback. Around this planner sits the field-deployment machinery: a semantic graph, continuously built and shared between ground and aerial platforms; the ground stack's LiDAR odometry and ground segmentation; the aerial stack's waypoint selection; and LoRA distillation, which fine-tunes a small Llama model on expert planner demonstrations so it can run onboard a compute-limited UAV.","core_discovery":"The central claim is that language-driven planning by foundation models can be moved from closed-world settings into large-scale unstructured field environments, provided the language model is embedded in a closed loop with a semantic mapper and a plan validator. The paper reports that SPINE, using GPT-4o as the planner, completed 12 of 14 UGV missions requiring 2–8 reasoning steps and traversals from 100 m to 1 km, with two failures from communication loss and odometry drift and two obstacle-detection failures resolved by brief manual takeover. It further claims the first language-driven UAV planner running entirely on-device: a LoRA-distilled Llama-3.2-3B model onboard a Falcon 4 UAV fulfilled three of four short missions on the first try, and in an offline planning benchmark the distilled model outscored an off-the-shelf 3B model (72.7% versus 9.2%) while still trailing GPT-4o (100%). The paper also documents that online plan validation is load-bearing: without it, mission success drops sharply as the unknown portion of the map grows.","pith_inferences":["If the two missions completed after a manual takeover are counted as non-autonomous, the reported 12/14 completion rate becomes 10/14; publishing both counts would make future field comparisons cleaner.","The distilled planner's north-versus-south parking-lot confusion suggests a spatial-grounding deficit; a testable extension is to inject coordinates or topological pointers into the model's prompt.","The distillation setup hints at a compute-aware division of labor, where a large model decomposes missions and small onboard models execute well-specified subtasks, an architecture the paper sketches but does not test.","Because visual foundation models mislabel aerial images (cars as construction vehicles), fine-tuning on aerial viewpoints is a natural next experiment before visual-language planning can run onboard."],"forward_implications":["Language-specified missions can be run at kilometer scales in unknown terrain when the planner is coupled to online validation and a continually updated semantic map.","On-device distilled models can replace server-based LLMs for short-horizon planning, removing the need for continuous connectivity in network-denied field settings.","The value of validation feedback is measurable: even minimal explanations of infeasibility substantially improve LLM planning success, so reliability depends on both the model and the validator.","Distilled small models retain useful planning ability (72.7% success) but not full parity with the expert (100%), making multi-iteration planning and complex reasoning the next performance bottleneck.","A shared hierarchical semantic graph can serve as the common representation that lets a UAV build mission-relevant maps and a UGV execute language-specified missions from them."],"supporting_citations":[{"why":"Defines SPINE, the closed-loop LLM planner whose generation and validation modules are the paper's central object.","marker":"[4]"},{"why":"Source of the field UGV results summarized in Table I and of the air-ground teaming framework with hierarchical semantic graphs.","marker":"[18]"},{"why":"LoRA is the low-rank adaptation method used to distill the small onboard language model.","marker":"[23]"},{"why":"Provides Llama 3.2, the base model that is distilled for onboard UAV planning.","marker":"[24]"},{"why":"Faster-LIO supplies the LiDAR-inertial odometry used in the ground autonomy stack.","marker":"[19]"},{"why":"GroundGrid supplies free-space estimation and trajectory planning for ground robot behaviors.","marker":"[20]"},{"why":"Describes the Falcon 4 UAV used as the aerial platform for onboard planning.","marker":"[21]"},{"why":"Introduces the air router waypoint-based aerial autonomy stack used on the UAV.","marker":"[22]"},{"why":"Grounded SAM2 is used for open-vocabulary detection and segmentation to build mission-relevant semantic maps.","marker":"[26]"}],"fun_headline_variants":["LLM robot completes 12 of 14 field missions","On-device language model pilots drone from text","Field-proven LLM planner: 12 of 14 missions","Distilled 3B model flies UAV with no server","SPINE: LLM robot navigates unstructured terrain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported 12-of-14 success rate counts missions finished after a brief manual takeover during obstacle-detection failures as successes, so the strength of the field demonstration depends on accepting that intervention as still a success.","fun_headline_variants_meta":{"raw":{"variants":["LLM robot completes 12 of 14 field missions","On-device language model pilots drone from text","Field-proven LLM planner: 12 of 14 missions","Distilled 3B model flies UAV with no server","SPINE: LLM robot navigates unstructured terrain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00026,"raw_usage":{"total_tokens":1604,"prompt_tokens":978,"completion_tokens":626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":546}},"tokens_in":594,"tokens_out":626,"duration_ms":6238,"temperature":1.0,"reasoning_tokens":546,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:29:49.228692+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the fourteen UGV missions with the same specifications and map conditions but forbid any manual takeover, and separately rerun the four UAV missions with no prior semantic map and no server connectivity; if the first-try success rates fall materially from 12/14 and 3/4, the demonstration claims would not survive.","supporting_citations":[{"cited_title":"Air-ground col- laboration for language-specified missions in unknown environments,","cited_arxiv_id":null,"evidence_quote":"Source of the field UGV results summarized in Table I and of the air-ground teaming framework with hierarchical semantic graphs."},{"cited_title":"EvMAPPER: High Altitude Orthomapping with Event Cameras","cited_arxiv_id":"2409.18120","evidence_quote":"Describes the Falcon 4 UAV used as the aerial platform for onboard planning."},{"cited_title":"Enabling Large-scale Heterogeneous Collaboration with Opportunistic Communications,","cited_arxiv_id":null,"evidence_quote":"Introduces the air router waypoint-based aerial autonomy stack used on the UAV."}],"review_version":1}