{"id":"d6a6da57-8200-4b11-b2d4-8b2be4e959d0","arxiv_id":"2507.12741","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 333 visitors who tried autonomous helper robots at Avatar Land found 74.7% willing to use them, with reliability the top concern.","lead":"Researchers ran a 19-day public demonstration of autonomous helper robots in Osaka and surveyed 2,285 visitors, 333 of whom tried the fully autonomous daily-life support zone. Most respondents were open to using such robots at home or work, and the biggest hesitation was reliable task completion, not cost or human-like appearance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'fully autonomous' label is undermined by disclosed preloaded inputs and precomputed plans; survey results may measure a scripted demonstration, not autonomous operation.","rationale":"The reader's weakest_assumption is exactly the load-bearing concern I identify: the disclosed preloaded offline inputs (Section III.A) and precomputed task allocations (Section III.C) contradict the 'fully autonomous' framing. This concern is not merely a matter of degree of autonomy; it determines what construct the survey actually measured. If the system was running on a fixed script, then the 74.7% willingness-to-use figure and the related scenario/aversion analyses describe reactions to a staged demonstration, not to an autonomous system that interprets and plans for individual users. That would invalidate the generalization of the findings to fully autonomous CAs, even though the perception data themselves are real and the paper is transparent about its implementation. The reader's CONDITIONAL verdict is appropriate: the paper should either soften the autonomy claims, re-frame the study as evaluating a scripted demonstration, or provide evidence of real-time user-driven operation. I agree with the reader's assessment and recommend no change to the verdict, because the concern is already captured and the authors' own disclosure supports a conditional reading rather than outright rejection.","tokens_in":9139,"tokens_out":2527,"duration_ms":30779,"concrete_test":"Inspect event logs or video recordings from the 333 interactions to count distinct user commands, pointing directions, and target objects; if all interactions used the same preloaded instruction and object, the demonstration was scripted. Alternatively, run a live visitor giving a novel instruction and pointing to a different object without preloading; failure would confirm the lack of real-time autonomy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that public perception of fully autonomous CAs is broadly positive (74.7% willing-to-use) depends on the demonstration actually being fully autonomous. Section III.A discloses that, to ensure robustness, 'we preloaded a user's pose skeleton and transcribed instruction data prepared offline, and manually recorded object locations as the model inputs.' Section III.C similarly states 'we precomputed task allocations using a set of predefined user instructions.' This indicates that the exophora resolution and LLM-based multi-robot planning did not process each visitor's live speech, pointing, and environment in real time; instead, the system executed a pre-scripted task. Consequently, respondents evaluated a demonstration where the core autonomous perception and planning components were bypassed. The paper's framing (e.g., 'fully autonomous robotic CAs that were not teleoperated in any form') describes execution autonomy, but not the interactive autonomy visitors experienced. Therefore, the survey results support positive perception of a scripted or simulated autonomous CA, not of a system that resolves novel user instructions and plans accordingly. This is the load-bearing weakness because the abstract, introduction, and conclusions all generalize to fully autonomous CAs for physical daily-life support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on a public demonstration and survey conducted at the Avatar Land event in Osaka, Japan, in which a daily-life support zone featured three robotic cybernetic avatars (CAs) described as fully autonomous. Among 2,285 survey volunteers, 333 respondents reported participating in this zone. The central finding is that 74.7% of these 333 respondents expressed a positive attitude toward using the demonstrated CAs (39.3% 'Very likely' plus 35.4% 'Use if conditions are right'), with the most common intended use scenarios being daily life (47.8%) and work (32.5%). The paper further reports that among the 8 respondents who declined, the most frequently cited reason was 'Could not use well' (37.5%), leading to the conclusion that task-execution reliability is the primary public concern. The authors position the work as a large-scale evaluation of public perception of fully autonomous CAs for physical daily-life support.","tokens_in":9286,"tokens_out":3382,"duration_ms":40954,"significance":"If the demonstration truly reflected fully autonomous operation, this study would provide a rare and valuable large-scale data point on public acceptance and concerns for domestic service robots. The scale (2,285 visitors, with 333 for the target demonstration) is a genuine strength, and the authors are commendably transparent about many implementation details, including the use of offline components. The internally consistent reporting of percentages and sample sizes is also a positive. However, the paper's central claim—that the survey measured perceptions of fully autonomous CAs—is directly weakened by the authors' own disclosure that key autonomous components were preloaded or precomputed. Consequently, the significance of the result is contingent on a substantial revision of the claim's scope or on additional evidence that visitors experienced live autonomous processing. Even with that caveat, the survey data on a scripted or partially autonomous demonstration could still be useful if framed accurately.","major_comments":[{"comment":"The characterization of the demonstrated system as 'fully autonomous' is contradicted by the disclosed implementation details. Section III.A states that 'we preloaded a user's pose skeleton and transcribed instruction data prepared offline, and manually recorded object locations as the model inputs,' and Section III.C states that 'we precomputed task allocations using a set of predefined user instructions.' These statements indicate that the exophora resolution model and the LLM-based multi-robot planner did not process live, arbitrary user inputs during the demonstration. The abstract, introduction, and conclusions generalize the survey results to 'fully autonomous CAs' that resolve novel instructions and plan accordingly. As written, the survey likely measured reactions to a pre-scripted or heavily constrained interaction, not to the autonomous interactive behavior claimed. The authors must either (a) clarify what each visitor actually experienced (e.g., whether they observed a scripted sequence or engaged in live interaction), and revise the claims to match that experience, or (b) provide evidence that preloaded data were only used as a fallback and that a substantial portion of interactions used live input. Without this clarification, the central claim that public perception of fully autonomous CAs is broadly positive is not supported.","section":"III.A and III.C"},{"comment":"The conclusion that 'hesitation primarily centered on whether the robots could consistently complete tasks successfully' and that 'cost and human-like interaction were not dominant concerns' is based on only 8 respondents who selected 'Do not want to use' (n− = 8). Within this tiny sample, 'Could not use well' was chosen by 3 respondents (37.5%), 'Prefer human face-to-face interaction' by 2 (25%), and 'Seemed expensive' by 1 (12.5%). With such a small n, the margin of error is extremely large, and the relative ordering of reasons is not statistically robust. The paper should explicitly quantify this limitation (e.g., 95% confidence intervals) and temper the language in the Conclusion. The current statements overstate the precision with which the reasons for non-adoption are known.","section":"IV, Q4 analysis and Conclusion"},{"comment":"The survey's external validity is limited by its self-selected, open-event sampling. The paper acknowledges that the total number of visitors is unknown and that non-respondents likely outnumbered respondents by an order of magnitude, but it still draws general conclusions about 'public perception' and 'public interest' from the 333 self-selected respondents in the daily-life support zone. The paper should explicitly list this as a limitation and avoid framing the results as representative of the broader Japanese or global population. A discussion of potential self-selection bias (e.g., tech-enthusiastic visitors being more likely to participate and respond) would strengthen the interpretation.","section":"II and IV"}],"minor_comments":[{"comment":"The paper states that 74.7% of respondents had a positive attitude, but 249/333 = 74.77%, which rounds to 74.8%. Please ensure the reported percentage is consistent with the stated counts.","section":"IV, first paragraph"},{"comment":"The observation that 'Daily life' was selected by only 47.8% of respondents despite the demonstration focusing on daily-life assistance is interesting, but the paper does not provide the corresponding values for the other demonstrations. A brief comparison or a note on whether this gap is statistically meaningful would help interpret the claim.","section":"IV, Q3 analysis"},{"comment":"The figure labels in Fig. 3(Q1) include the notation '(n = 333)' on the bar for the daily-life support zone, which is the sample size for that zone but is visually similar to a participation rate. Please clarify in the caption or axis label that this is a count, not a percentage.","section":"Table I and Fig. 3"},{"comment":"The sentence describing the preloading appears in the middle of a technical description without a transition. Consider moving this disclosure to a dedicated 'Demonstration limitations' subsection or integrating it explicitly into the discussion so readers immediately understand its implications.","section":"III.A"},{"comment":"The paper reports 'No answer' as 25.0% in Q4 (2 respondents), which is not commented on. Since n=8 is already very small, the 2 'No answer' responses reduce the effective denominator further; please address this in the analysis.","section":"IV, Q4"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is contingent on the demonstration genuinely reflecting fully autonomous operation. The authors' own disclosures in Sections III.A and III.C make this questionable. This is not a case of disagreement with consensus but rather an internal inconsistency between the claimed autonomy and the reported implementation. I recommend major revision to reframe the claims or provide additional evidence about the actual visitor experience. The paper may be better suited to a venue focused on human-robot interaction field studies rather than a robotics systems journal, but that is the editor's call. The authors self-cite extensively, but that alone is not a concern given the topic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper for what it is: a rare large-scale field survey (n=333 at the focal condition) of public perception of cybernetic avatars claimed to be fully autonomous, with a write-up that is unusually transparent about the system's real limitations. The survey data are real and the descriptive statistics are internally consistent: 74.7% willing-to-use, top scenarios daily life (47.8%) and work (32.5%). The authors also explicitly acknowledge that positive responses may reflect first impressions rather than long-term judgments.\n\nThe soft spot is structural. The abstract and conclusions frame the study as evaluating \"fully autonomous robotic CAs,\" but Sections III.A and III.C disclose that exophora resolution ran on preloaded pose skeletons and offline-transcribed instructions, and that task allocations were precomputed from predefined instructions. Visitors watched a scripted execution of an autonomous pipeline, not a system that resolved their live speech and pointing in real time. The authors deserve credit for disclosing this, but the disclosure sits in the system section while the title, abstract, and conclusion generalize without that qualification. The stress-test note gets this right.\n\nA second, smaller soft spot: the conclusion that \"hesitation primarily centered on whether the robots could consistently complete tasks successfully\" comes from Q4, where n=8 and only 3 people (37.5%) answered \"Could not use well.\" That is a very thin thread for such a central claim. The paper notes the small sample, but the conclusion states it as if it were robust.\n\nOther concerns are minor and mostly acknowledged: the event was open and self-selecting, demographic analysis is deferred, and raw data are not released. The citation pattern looks reasonable; the self-cited prior components are used as implementation tools, not as evidence for the survey claims.\n\nWho this is for: researchers in human-robot interaction, avatar-symbiotic society planning, and public acceptance of service robots. It deserves a serious referee, but the referee should push the authors to either re-frame the study as evaluating a scripted or semi-autonomous demonstration or substantially weaken the autonomy claims in the abstract and conclusion. The perception data remain useful either way.\n\nMy recommendation: send it to peer review with a request for major revision on the framing, and require confidence intervals or raw data for the perceptual findings. It is a worthwhile dataset that needs to be presented more carefully.","headline":"A genuinely useful large-scale perception dataset, but the paper's 'fully autonomous' claim runs ahead of a system that was partly scripted, and the headline reliability concern rests on only 8 respondents.","tokens_in":9981,"tokens_out":2261,"would_cite":true,"duration_ms":26960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that 74.7% of 333 surveyed visitors to a public demonstration of fully autonomous cybernetic avatars expressed willingness to use them for daily-life physical support, with task reliability as the main concern.","keywords":["cybernetic avatars","fully autonomous robots","public perception survey","physical daily-life support","human-robot interaction","object retrieval","large-scale demonstration","Avatar Land"],"falsifier":"A log audit of the Avatar Land event showing that every visitor instruction was replaced by pre-recorded transcriptions and every task allocation matched a precomputed template, with no novel input accepted by the system, would falsify the paper's 'fully autonomous' framing; conversely, evidence that previously unseen visitor instructions produced live exophora resolution and task planning in real time would support it.","tokens_in":8913,"feed_emoji":"🤖","tokens_out":7350,"duration_ms":69226,"temperature":0.7,"pith_summary":"This paper tries to establish how the general public perceives fully autonomous cybernetic avatars—robotic stand-ins that carry out physical tasks like retrieving objects without a human teleoperator—when they are deployed for daily-life support. At a 19-day public event in Osaka, the authors set up a replicated home environment where three robots worked together to fetch objects based on a user's pointing gesture and spoken instruction, and surveyed 2,285 visitors overall, 333 of whom interacted with the fully autonomous system. The survey found that 74.7% of those 333 respondents said they would use such avatars in daily life (39.3% 'very likely' plus 35.4% 'use if conditions are right'), with the most desired settings being daily-life tasks (47.8%) and work (32.5%). Among the few who declined, the main reason (37.5%) was doubt about whether the robots could reliably complete tasks, while cost and human-like interaction were not dominant concerns, so the paper argues that improving task success and stability is the key to adoption.","feed_headline":"74.7% of visitors would use fully autonomous helper avatars","feed_subtitle":"A 19-day public demo finds daily-life support wanted, with reliability the top worry over cost or human-likeness.","key_machinery":"The argument is carried by a public field demonstration paired with a short multiple-choice survey. The demonstrated system was a household-like environment with three fully autonomous cybernetic avatars (robotic agents that act on a user's behalf without teleoperation): one Fetch mobile manipulator and two Kachaka shelf-carrying robots. The interaction pipeline started with the user pointing at an object and saying something like 'bring that'; the robot's camera tracked the user's eyes and wrists, speech was transcribed by Whisper, an exophora resolution model combined pointing direction, demonstrative words, and object-category context to infer the target object, and a GPT-4o-based planner allocated retrieval and disposal tasks among the three robots. For the survey, the load-bearing instrument was the four-question questionnaire (Q1 participation, Q2 usage likelihood, Q3 usage scenarios, Q4 usage aversions) administered to 2,285 visitors, with a subset of 333 responses tied to the fully autonomous demonstration.","core_discovery":"On the paper's own terms, the central discovery is that a large, non-expert public audience reacts positively to fully autonomous cybernetic avatars performing physical support, and that the main perceived obstacle is functional reliability, not cost or uncanniness. Of the 333 survey respondents who reported engaging with the daily-life support demonstration, 249 (74.7%) answered 'Very likely to use' or 'Use if conditions are right' when asked whether they would use the demonstrated CAs in their daily life. Asked where they would use them, 47.8% picked 'Daily life' and 32.5% picked 'Work.' Among the eight respondents who said they would not want to use them, 37.5% cited 'Could not use well,' 25.0% preferred human face-to-face interaction, and only one respondent said they seemed expensive. The authors read this as evidence that public acceptance is contingent on proven task reliability, with cost and human-like interaction playing secondary roles, and note that the under-50% figure for 'Daily life' scenarios reveals a perceived applicability gap for home use.","pith_inferences":["Editorial inference: because the paper discloses that pose skeletons and instruction transcriptions were preloaded offline and task allocations were precomputed, the 74.7% willingness figure should be read as a reaction to a well-rehearsed demonstration rather than to real-time autonomous reasoning; a live system with frequent failures could yield lower willingness.","Editorial inference: with only eight 'Do not want to use' responses, the 37.5% reliability concern and the claim that cost is not dominant rest on very small counts; a targeted follow-up with a larger reluctant sample is needed before treating those proportions as stable.","Editorial inference: a natural extension is a controlled comparison where one group interacts with the genuinely live autonomy pipeline and another with the pre-scripted version, holding the visible behavior inside the replicated home constant, to isolate how much of the positive response comes from the idea of autonomy versus the actual working system.","Editorial inference: because the demonstration was ranked 8th out of 11 zones in participation (14.7% of survey respondents), self-selection may skew the sample toward visitors already curious about robots; the reported willingness may be an upper bound for the broader population."],"forward_implications":["If the 74.7% willingness-to-use figure holds, there is measurable public demand for fully autonomous robotic avatars that fetch and deliver objects in the home and workplace.","The dominant aversion being 'could not use well' implies that field demonstrations and product development should prioritize consistent task success and recovery from failures over human-likeness or cost reduction.","The finding that 'Daily life' was chosen by under half of interested respondents, despite the demo being a daily-life scenario, points to a gap between the technology's intended setting and how applicable it feels at home.","The event format—public, open access, with a voluntary survey—can yield perception data from non-experts at scale even when the total visitor count is unknown.","If reliability concerns are the main barrier, then reporting objective success rates from such demonstrations could raise adoption expectations in future surveys."],"supporting_citations":[{"why":"Exophora resolution model that maps pointing gestures and demonstrative words to a target object from contextual information, forming the core of the instruction-understanding pipeline.","marker":"[14]"},{"why":"Whisper speech recognition transcribes user instructions into text before exophora resolution.","marker":"[12]"},{"why":"MediaPipe extracts the user's pose skeleton (eyes and wrists) from RGB-D images to define the pointing direction probability p1.","marker":"[13]"},{"why":"GPT-4o, cited as the language model used for multi-robot task planning from user commands.","marker":"[15]"},{"why":"LLM-based approach that generates action sequences for each robot given capabilities and example task planning, adapted for precomputed task allocation.","marker":"[20]"},{"why":"Detic detects the target object before grasping to correct for self-localization error and changed object positions.","marker":"[17]"}],"fun_headline_variants":["74.7% open to fully autonomous helper avatars","Reliability, not cost, is top barrier to autonomous helper avatars","Public embraces autonomous helper avatars, with reliability caveat","2,285 surveyed: 74.7% would use autonomous helper avatars","Fully autonomous helper avatars: public wants them, worries about reliability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the visitors were actually experiencing fully autonomous operation, but Section III.A discloses that pose skeletons and transcribed instructions were preloaded offline and Section III.C that task allocations were precomputed, so the demonstration may have been effectively scripted rather than autonomous in real time.","fun_headline_variants_meta":{"raw":{"variants":["74.7% open to fully autonomous helper avatars","Reliability, not cost, is top barrier to autonomous helper avatars","Public embraces autonomous helper avatars, with reliability caveat","2,285 surveyed: 74.7% would use autonomous helper avatars","Fully autonomous helper avatars: public wants them, worries about reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1387,"prompt_tokens":990,"completion_tokens":397,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":606,"tokens_out":397,"duration_ms":3609,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:39:12.887743+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A log audit of the Avatar Land event showing that every visitor instruction was replaced by pre-recorded transcriptions and every task allocation matched a precomputed template, with no novel input accepted by the system, would falsify the paper's 'fully autonomous' framing; conversely, evidence that previously unseen visitor instructions produced live exophora resolution and task planning in real time would support it.","supporting_citations":[{"cited_title":"Exophora Resolution of Linguistic Instructions with a Demonstrative based on Real-World Multimodal Information,","cited_arxiv_id":null,"evidence_quote":"Exophora resolution model that maps pointing gestures and demonstrative words to a target object from contextual information, forming the core of the instruction-understanding pipeline."},{"cited_title":"Robust Speech Recognition via Large-Scale Weak Supervision,","cited_arxiv_id":null,"evidence_quote":"Whisper speech recognition transcribes user instructions into text before exophora resolution."},{"cited_title":"MediaPipe: A Framework for Perceiving and Processing Reality,","cited_arxiv_id":null,"evidence_quote":"MediaPipe extracts the user's pose skeleton (eyes and wrists) from RGB-D images to define the pointing direction probability p1."},{"cited_title":"Language Models are Few-Shot Learners,","cited_arxiv_id":null,"evidence_quote":"GPT-4o, cited as the language model used for multi-robot task planning from user commands."},{"cited_title":"Reducing cost of on-site learning by multi-robot knowledge integration and task allocation via large language models,","cited_arxiv_id":null,"evidence_quote":"LLM-based approach that generates action sequences for each robot given capabilities and example task planning, adapted for precomputed task allocation."},{"cited_title":"Detecting Twenty-Thousand Classes Using Image- Level Supervision,","cited_arxiv_id":null,"evidence_quote":"Detic detects the target object before grasping to correct for self-localization error and changed object positions."}],"review_version":1}