{"id":"785dd114-a72c-4e2e-9e3a-b666d70bbff4","arxiv_id":"2411.10038","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A remote life-support robot interface uses @template variables in language instructions, letting the robot collect on-site information and query the user to expand uncertain task details.","lead":"A robot system lets people give tasks like “bring me a drink” with blanks marked by @ symbols, and the robot fills in the blanks on site using its camera and by asking the user to choose. The paper shows two real-robot demonstrations, which work but reveal that the vision model sometimes invents items that are not actually there.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The system's core value depends on VLM-supplied options being trustworthy, but the only buy-task demonstration includes hallucinated options, so the central claim of reliable intention execution is not yet established.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing concern: the VLM's reading of on-site information is not reliable, and the paper's own Section IV-A documents a failure in one of the two demonstrations. My stress-test does not change the reader's conditional verdict. The interface concept with @-delimited template variables is plausible and the real-robot demonstrations are useful existence proofs, but the central claim is stronger than the evidence. A single successful trial per task, with one trial containing hallucinated options, is insufficient to establish that users can reliably execute tasks 'according to their intentions' when the user cannot verify the robot's information. The proposed concrete test would directly measure whether hallucinated options threaten task success. I find no basis for moving the verdict to reject, because the system architecture is coherent and the failure is attributable to an external component that could be improved or guarded; conditional acceptance with a demand for reliability evaluation is the appropriate outcome.","tokens_in":9176,"tokens_out":2775,"duration_ms":30537,"concrete_test":"Re-run the Section IV-A buy scenario at least 20 times with known ground-truth menus (e.g., printed menus or a real shop display), and compare every VLM-listed item and price against the ground truth. Count the number of trials in which (a) a hallucinated item appears among the options and (b) a user selects a hallucinated item. The concern is settled if, under varied lighting and camera angles, hallucinated options never appear in the selectable list or if a built-in verification step (e.g., OCR cross-check or a second VLM pass) removes them before presentation. If any user selection lands on a nonexistent item, the central claim of reliable intention execution fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section V) is that the template-variable interface lets users execute tasks according to their intentions with minimal intervention, even when on-site information is needed. For a user who lacks prior knowledge of the menu or fridge contents, the only channel for that information is the VLM's extraction from a single camera image. Section IV-A explicitly reports that the VLM listed 'Parmesan' and 'Spicy Italian' even though they were not on the actual menu, and that only one of the four displayed prices was correct. Since the user cannot verify the menu independently, selecting either hallucinated option would have caused the buy task to fail at the point of interaction with the shop staff. The successful completion of the demo depended on the user happening to choose 'Chili Chicken,' one of the two real items. This is not a peripheral implementation detail: it is the mechanism by which the user's intention is supposed to reach the robot. The paper acknowledges the hallucination in Section V but only suggests future remedies such as a stronger model or multi-VLM averaging; no filtering or verification step is part of the current system. The second experiment is a single trial with accurate VLM output, but it does not establish reliability either. Therefore the evidence does not yet support the broad claim that users can reliably execute tasks according to their intentions for tasks requiring on-site information.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a robot interface system in which users specify life-support tasks in natural language using @-delimited template variables, such as \"Go to the Subway and buy @food@,\" to explicitly mark information that must be collected on-site. The system uses an LLM to decompose instructions into action sequences, Dialogflow to convert them into executable EusLisp scripts, a VLM to populate the template variables from camera images, and chat/AR interfaces for the user to select among the presented options. Two real-robot trials on a PR2 are reported: buying food at a Subway restaurant and fetching a drink from a refrigerator. The buy trial contained VLM hallucinations (two nonexistent menu items and incorrect prices for most items), while the fridge trial produced accurate item lists. The paper claims the system allows users to execute tasks according to their intentions with minimal intervention, even for tasks requiring on-site information.","tokens_in":9352,"tokens_out":4315,"duration_ms":47816,"significance":"The template-variable notation is a simple, intuitive way to make uncertainty explicit in human-robot instructions, and the integration of LLMs, VLMs, and AR into a working service-robot pipeline is a useful feasibility demonstration. The real-robot experiments show that the end-to-end system can operate in realistic settings, including navigation with an elevator and physical interaction with a fridge. However, the evidence for the central claim is currently weak: the two trials are single demonstrations with no repeated runs, no quantitative metrics, no baselines, and no measurement of user intervention. Moreover, the first trial shows that the VLM-based expansion mechanism can present incorrect options, directly threatening the reliability of the core interaction. The idea is promising, but the system needs a robustness mechanism and a more rigorous evaluation to support the claimed benefit.","major_comments":[{"comment":"The VLM in the buy task produced two menu items (\"Parmesan\" and \"Spicy Italian\") that were not on the actual menu, and only one of the four displayed prices was correct. Because the user has no independent way to verify the menu, selecting either hallucinated option would have caused the task to fail at the staff interaction. The successful completion depended on the user happening to choose \"Chili Chicken,\" one of the real items. This directly undermines the Section V claim that the system allows users to execute tasks according to their intentions for tasks requiring on-site information. A verification or filtering step, or a success-rate measurement over repeated trials, is necessary to support the claim.","section":"Section IV-A, Section V"},{"comment":"Both experiments are single-shot demonstrations with no repeated trials, no quantitative success metrics, no baseline comparison, and no measurement of intervention time or reliability. The Abstract's statement that \"effectiveness was demonstrated\" is therefore stronger than the evidence supports. The paper should either temper the claim to a feasibility demonstration or add an evaluation protocol with multiple runs and a clear success criterion.","section":"Section IV"},{"comment":"The architecture has no mechanism to detect or reject VLM hallucinations before presenting options to the user. The prompt template for the buy function asks the VLM to list items with prices and descriptions, but nothing validates the output against the image content or against known constraints. The paper acknowledges this in Section V but only suggests future remedies such as stronger models or multi-VLM averaging. Because option reliability is load-bearing for the central claim, the current system is incomplete as a solution to the stated problem.","section":"Section III-B"},{"comment":"The claim of \"minimal intervention\" is not operationalized. The system still requires the user to select an item and, in the fridge task, to place a virtual drink in the AR interface. There is no measure of the number of queries, the time spent by the user, or the cognitive load, so the improvement over existing query/response systems is not quantified. Without such a measure, the asserted advantage of the proposed interface over prior chat-based systems remains unsubstantiated.","section":"Section V"}],"minor_comments":[{"comment":"The instruction \"Nouns enclosed in @@ should be output as they are enclosed in @@\" is confusing because the template notation is elsewhere defined with a single @ on each side; please use the same notation throughout and correct the apparent typo in this sentence.","section":"Section III-A"},{"comment":"Figure 2 is very dense and the text is small; splitting the two stages into separate diagrams or enlarging the fonts would improve readability.","section":"Figure 2"},{"comment":"The brand name \"Wonda Wonderful Coffee Morning Shot\" appears to be spelled inconsistently; please verify the correct product name and use it consistently in the text and figures.","section":"Section IV-B"},{"comment":"The paper reuses the authors' previous Dialogflow-based script generator and chat system but does not clearly state which components are newly implemented. A short sentence distinguishing reused versus novel components would help readers understand the contribution.","section":"Section III-A and Reference [13]"},{"comment":"The limitation that the robot must autonomously navigate to appropriate positions and orient its camera is somewhat contradicted by the experiments, where the robot navigated to pre-defined map symbols and pointed its camera at the menu or fridge; please clarify what additional autonomous capability would be needed beyond the demonstrated setup.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The paper reads more like an extended workshop/system-demonstration paper than a full journal article. The main concern is the disconnect between the strong claim in Section V and the hallucinated output in Section IV-A. For a journal, additional experiments with repeated trials and a clear success criterion are needed. The self-citations to the authors' previous work are appropriate and do not constitute circularity. I would not reject the idea, but the current validation is insufficient to support the stated claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe template-variable interface is the thing worth remembering here. The @variable@ syntax is a simple, genuine addition to the language-to-robot toolbox: the user marks the slots that are uncertain now, and the robot fills them on-site with VLM reading plus a user confirmation. That pattern is not in SayCan, Inner Monologue, or Robots That Ask For Help, and it is a clean idea for remote life-support tasks. The system integration is coherent: LLM generates the script, Dialogflow turns it into EusLisp, VLM prompts are generated per function, and the chat/AR UI handles the options. The authors also deserve credit for reporting the VLM hallucination in the buy task instead of hiding it.\n\nThe soft spot is exactly where the stress-test points. In the buy experiment, the VLM presented four menu options, two of which were not on the actual menu, and only one price was right. The user happened to select one of the two real items, so the task succeeded. But the user had no independent way to verify the menu; selecting either hallucinated option would have led the robot to ask for a non-existent sandwich. That means the mechanism by which the user's intention reaches the on-site action gave wrong options half the time in the only demonstration that required it. The fridge trial had accurate VLM output, but it is a single successful trial. No repeated trials, no baseline, no error analysis. The paper acknowledges the hallucination in Section V and suggests stronger models or multi-VLM averaging, but no filtering or verification is in the current system. So the broad claim in Section V, that users can execute tasks according to their intentions with minimal intervention even when on-site information is needed, is stronger than the evidence supports.\n\nThis is not a takedown. The interface pattern is a solid practical contribution, and the paper is honest about its limitations. The citation pattern is fine: the self-citations are prior system components, not circular support for the new idea. But the empirical validation is a proof-of-concept, not a demonstration of reliable performance. A serious referee should ask for a handful of repeated trials per task, a baseline (for instance, asking the user to type the item name without a VLM-provided list), and some discussion of how hallucinated options get filtered or flagged.\n\nRead this if you work on language-conditioned robot interfaces or remote operation with foundation models. The community would benefit from having this design pattern on the record, and the flaws are correctable in revision. I would send it out for review, with the expectation that the evaluation gets tightened.","headline":"A genuinely useful interface pattern, but the only buy-task demo includes VLM hallucinations that undercut the paper's central reliability claim.","tokens_in":9939,"tokens_out":3410,"would_cite":true,"duration_ms":30791,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a task instruction with @-marked placeholders can be completed by a robot that gathers the missing options on site and asks the user to choose.","keywords":["template variables","language-based robot instruction","vision-language models","large language models","remote life support robots","augmented reality interface","task planning","human-robot interaction"],"falsifier":"Run the buying task at a menu with known ground truth, have a user select each option the vision-language model offers, and check whether the robot can complete the purchase for every choice; the paper already provides evidence for failure, since 'Parmesan' and 'Spicy Italian' were presented but absent from the actual menu, so selecting either would make the robot attempt an unavailable purchase.","tokens_in":8937,"feed_emoji":"🤖","tokens_out":9474,"duration_ms":93386,"temperature":0.7,"pith_summary":"Users of this system can instruct a remote life-support robot with phrases such as 'buy @food@' or 'bring @drink@', where the @-marks flag information that only exists on site. The robot keeps those placeholders in its generated task script, collects the missing information from its camera using a vision-language model, and presents the user with concrete choices in a chat interface or AR glasses. The user picks one option, and the robot carries out the rest of the task, such as buying the selected sandwich or fetching the selected drink from the fridge and delivering it. The paper argues this gives users minimal-intervention control over tasks that cannot be fully specified in advance.","feed_headline":"Robot turns @food@ into real menu choices at the task site","feed_subtitle":"Users leave unknown details as @variables; the robot fills them in on site and asks the user to pick.","key_machinery":"The load-bearing object is the template variable: text surrounded by @, such as @food@ or @user@, that explicitly marks information unknown at instruction time. The mechanism is a two-stage expansion: function-specific vision-language prompt templates turn each placeholder into a question about the robot's current camera image, and a feedback loop (chat option buttons and an AR object menu) brings the resulting candidates back to the user for selection, with AR placement adding the object's pose for delivery. This carries the argument because it converts an under-specified instruction into a script with explicit decision points, so uncertainty is handled at runtime rather than guessed by the planner.","core_discovery":"The paper's claim is that writing explicit placeholder variables, like @food@ or @drink@, inside a natural-language instruction converts an otherwise under-specified task into one the robot can finish. The placeholder is preserved through the generated robot script, and at execution time it is expanded by two mechanisms: the robot points its camera at the site, feeds the image with a function-specific prompt to a vision-language model, and converts the model's answer into JSON; the user then selects one of the presented options through chat buttons or AR glasses. In AR, placing a virtual copy of the selected object also transmits its pose, resolving the 'where to deliver' part of the task. In two real-robot experiments the system bought 'Chili Chicken' from a sandwich shop and fetched 'Georgia' canned coffee from a fridge and delivered it to the user. The authors conclude that users can execute tasks according to their intentions with minimal intervention even when the task depends on on-site information.","pith_inferences":["The @ notation is a small, human-readable way to mark unknowns; it could plausibly be extended to numeric constraints (@budget@), timing (@when@), or preference (@spicy@) using the same two-stage expansion, though the paper only demonstrates object-identity placeholders.","Because the sandwich-shop experiment offered two menu items that were not actually on the menu, a practical deployment would likely need a verification step, such as multiple camera views or a consistency check, before showing options to the user.","The stored task database suggests the system could accumulate a library of templated tasks, so a user might teach new tasks by analogy to old ones with different @variables@; the paper does not test this."],"forward_implications":["A user can order errands such as buying food or fetching a drink without knowing what is available, because the robot collects that information and asks only when a choice matters.","Task instructions can be registered once: a similar later instruction loads the stored script from the database, so the user does not have to re-approve the decomposition.","Geometric uncertainty such as the user's location can be resolved through AR: placing a virtual object transmits both the item choice and its pose to the robot.","The two experiments demonstrate end-to-end execution of purchase and delivery tasks with only a tap or AR placement as user input."],"supporting_citations":[{"why":"established the baseline of grounding natural-language commands in robot affordances, which this system extends with template variables.","marker":"[2]"},{"why":"provided the embodied reasoning-through-planning approach that supports generating action sequences from language with feedback.","marker":"[4]"},{"why":"showed robots querying humans when the planner is uncertain, the direct precedent for this system's runtime user feedback.","marker":"[7]"},{"why":"the natural-language engine that converts each action sentence into an executable EusLisp function without hallucination.","marker":"[12]"},{"why":"supplied the chat interface and the trained action-name/argument model used to generate task scripts.","marker":"[13]"},{"why":"provides the map and navigation symbols that resolve location arguments such as fridge-front into navigable coordinates.","marker":"[14]"},{"why":"documents the GPT-4 model family used for action decomposition and JSON formatting.","marker":"[16]"},{"why":"the PR2 mobile manipulator on which the two real-world experiments were carried out.","marker":"[18]"}],"fun_headline_variants":["Robot fills @food@ and @drink@ with real options on site","User picks from camera options to finish robot task","@placeholders@ let robots buy food and fetch drinks","Robot vision and AR help complete vague instructions","Robot asks you to pick @drink@ from the fridge view"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system works only if the vision-language model reads the on-site information from a single camera image accurately enough that the options it presents are real and available, and in the sandwich-shop experiment two of the four offered menu items were not actually on the menu.","fun_headline_variants_meta":{"raw":{"variants":["Robot fills @food@ and @drink@ with real options on site","User picks from camera options to finish robot task","@placeholders@ let robots buy food and fetch drinks","Robot vision and AR help complete vague instructions","Robot asks you to pick @drink@ from the fridge view"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3156,"prompt_tokens":849,"completion_tokens":2307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":2226}},"tokens_in":465,"tokens_out":2307,"duration_ms":18810,"temperature":1.0,"reasoning_tokens":2226,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:01:30.934530+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the buying task at a menu with known ground truth, have a user select each option the vision-language model offers, and check whether the robot can complete the purchase for every choice; the paper already provides evidence for failure, since 'Parmesan' and 'Spicy Italian' were presented but absent from the actual menu, so selecting either would make the robot attempt an unavailable purchase.","supporting_citations":[{"cited_title":"Ahn, et al","cited_arxiv_id":null,"evidence_quote":"established the baseline of grounding natural-language commands in robot affordances, which this system extends with template variables."},{"cited_title":"Huang, et al","cited_arxiv_id":null,"evidence_quote":"provided the embodied reasoning-through-planning approach that supports generating action sequences from language with feedback."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"showed robots querying humans when the planner is uncertain, the direct precedent for this system's runtime user feedback."},{"cited_title":"https://cloud.google.com/dialogflow","cited_arxiv_id":null,"evidence_quote":"the natural-language engine that converts each action sentence into an executable EusLisp function without hallucination."},{"cited_title":"Obinata, et al","cited_arxiv_id":null,"evidence_quote":"supplied the chat interface and the trained action-name/argument model used to generate task scripts."},{"cited_title":"Kunze, et al","cited_arxiv_id":null,"evidence_quote":"provides the map and navigation symbols that resolve location arguments such as fridge-front into navigable coordinates."},{"cited_title":"Bohren, et al","cited_arxiv_id":null,"evidence_quote":"the PR2 mobile manipulator on which the two real-world experiments were carried out."}],"review_version":1}